An inter-organizational relationship recognition method and device based on web text mining

By fine-tuning and incremental training on large language models, combined with Word2Vec technology, the problem of poor unstructured text processing capabilities in the existing technology is solved, and accurate identification of inter-organizational relationships is achieved.

CN119088978BActive Publication Date: 2025-06-24UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411109002.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-06-24
Estimated Expiration
2044-08-13

AI Technical Summary

Technical Problem

When processing unstructured text, it is difficult to accurately identify complex relationships between organizations, resulting in low recognition accuracy.

Method used

Using a method based on network text mining, the initial training set is constructed through Word2Vec technology, and the large language model is fine-tuned using the low-rank adaptation fine-tuning method, and the model is further trained incrementally through the incremental training set to obtain the second organizational relationship recognition model.

Benefits of technology

It effectively improves the accuracy of identification of inter-organizational relationships, can identify multiple possible relationships between different organizations from a large number of irregular texts, and achieves accurate and efficient identification of unstructured texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119088978B_ABST
    Figure CN119088978B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of natural language processing, and particularly to a method and device for identifying inter-organization relationships based on web text mining. The method includes: obtaining web texts; preprocessing the web texts to obtain processed texts; constructing triples based on the processed texts to obtain an initial training set; fine-tuning a preset large language model using a low-rank adaptation fine-tuning method according to the initial training set to obtain a first inter-organization relationship recognition model; constructing an incremental training set based on the processed texts using the first inter-organization relationship recognition model; performing incremental training on the first inter-organization relationship recognition model using the incremental training set based on a preset number of training rounds to obtain a second inter-organization relationship recognition model; obtaining the web texts to be recognized; and inputting the web texts to be recognized into the second inter-organization relationship recognition model to obtain the recognized inter-organization relationships. The present invention is an accurate and efficient method for identifying inter-organization relationships for unstructured texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method and device for identifying inter-organization relationships based on web text mining. Background Art

[0002] Relationship identification is a key task in the field of natural language processing, aiming to automatically identify and extract entities and their relationships from unstructured text data. Each relationship is represented in the form of a triple <entity a, relationship, entity b>. The relationships between organizations usually include various forms, such as cooperation, competition, subordination, shareholding, mergers and acquisitions, etc. Organization relationship identification shows significant importance and broad application prospects in different application fields such as business intelligence, risk management, and policy making.

[0003] In the business field, organization relationship identification helps enterprises understand the market structure and competition pattern. By analyzing texts such as public reports, business agreements, and partnership relationships, enterprises can identify potential business partners and competitors, and thus formulate corresponding market strategies to effectively avoid business risks. Government agencies and regulatory authorities can use organization relationship identification technology to monitor and analyze the interactions between different organizations to formulate effective policies and regulations. The application of organization relationship identification is not limited to the above fields, and it can also be extended to many interdisciplinary fields such as supply chain management, crisis response, and cultural research.

[0004] With the continuous improvement of computing power and the increasing abundance of data resources, large language models have developed rapidly. These models are pre-trained on large-scale corpora to capture rich language representations by learning the patterns, structures, and grammars of languages. However, existing organization relationship identification technologies still face some challenges. Due to the complexity of natural language, conventional deep learning models are difficult to fully understand the nuances and complex relationships in the text. Different industries and different types of texts also have different requirements for organization relationship identification. There is insufficient ability to process unstructured data.

[0005] In the prior art, there is a lack of an accurate and efficient method for identifying inter-organization relationships for unstructured text. Summary of the Invention

[0006] In order to solve the technical problems of poor unstructured text processing ability and low accuracy in the prior art, embodiments of the present invention provide a method and device for identifying inter-organization relationships based on web text mining. The technical solutions are as follows:

[0007] On the one hand, a method for identifying inter-organization relationships based on web text mining is provided. This method is implemented by an inter-organization relationship identification device, and the method includes:

[0008] Perform text crawling according to a preset list of organization names to obtain original network text; preprocess the original network text to obtain processed text;

[0009] Based on the Word2Vec technology, construct an initial training set according to a preset list of organization names, the processed text, preset training prompt words, and preset benchmark organizational relationships;

[0010] According to the initial training set, use the low-rank adaptation fine-tuning method to fine-tune a preset large language model to obtain a first organizational relationship recognition model;

[0011] Based on the first organizational relationship recognition model, construct an incremental training set according to preset training prompt words and the processed text;

[0012] Based on a preset number of training rounds, use the incremental training set to perform incremental training on the first organizational relationship recognition model to obtain a second organizational relationship recognition model;

[0013] Obtain a list of organization names to be recognized; perform text crawling and preprocessing according to the list of organization names to be recognized to obtain network text to be recognized; input the network text to be recognized into the second organizational relationship recognition model to obtain organizational recognition relationships.

[0014] On the other hand, provided is a device for recognizing relationships between organizations based on network text mining. This device is applied to the method for recognizing relationships between organizations based on network text mining. The device includes:

[0015] A text acquisition module, configured to perform text crawling according to a preset list of organization names to obtain original network text; preprocess the original network text to obtain processed text;

[0016] An initial training set construction module, configured to construct an initial training set based on the Word2Vec technology according to a preset list of organization names, the processed text, preset training prompt words, and preset benchmark organizational relationships;

[0017] A preliminary training module, configured to fine-tune a preset large language model according to the initial training set using the low-rank adaptation fine-tuning method to obtain a first organizational relationship recognition model;

[0018] An incremental training set construction module, configured to construct an incremental training set based on the first organizational relationship recognition model according to preset training prompt words and the processed text;

[0019] An incremental training module, configured to perform incremental training on the first organizational relationship recognition model using the incremental training set based on a preset number of training rounds to obtain a second organizational relationship recognition model;

[0020] An organizational relationship recognition module, configured to obtain a list of organization names to be recognized; perform text crawling and preprocessing based on the list of organization names to be recognized to obtain network text to be recognized; input the network text to be recognized into the second organizational relationship recognition model to obtain organizational recognition relationships.

[0021] On the other hand, an inter-organizational relationship recognition device is provided. The inter-organizational relationship recognition device includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned inter-organizational relationship recognition method based on network text mining is implemented.

[0022] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned inter-organizational relationship recognition method based on network text mining.

[0023] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0024] The present invention proposes an inter-organizational relationship recognition method based on network text mining, which captures rich language representations by learning the patterns, structures, and grammars of languages through large language models. A pre-trained large language model is fine-tuned using a relatively small dataset related to a specific task, enabling it to better complete specific application tasks. The present invention can identify various possible relationships between different organizations from a large number of unstructured texts, and effectively improve the recognition accuracy of inter-organizational relationships with the help of large language models. The present invention is an accurate and efficient inter-organizational relationship recognition method for unstructured texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a flowchart of an inter-organizational relationship recognition method based on network text mining provided by an embodiment of the present invention;

[0027] Figure 2 It is a block diagram of an inter-organizational relationship recognition device based on network text mining provided by an embodiment of the present invention;

[0028] Figure 3 It is a schematic structural diagram of an inter-organizational relationship recognition device provided by an embodiment of the present invention. Detailed implementation manners

[0029] The technical solutions in the present invention will be described below with reference to the accompanying drawings.

[0030] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.

[0031] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0032] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0033] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0034] The embodiments of the present invention provide a method for identifying inter-organization relationships based on web text mining. This method can be implemented by an inter-organization relationship identification device, and the inter-organization relationship identification device can be a terminal or a server. As Figure 1 shown in the flowchart of the method for identifying inter-organization relationships based on web text mining, the processing flow of this method can include the following steps:

[0035] S1. Perform text crawling according to a preset list of organization names to obtain the original web text; preprocess the original web text to obtain the processed text.

[0036] Optionally, preprocessing the original web text to obtain the processed text includes:

[0037] Perform missing value processing on the original web text to obtain the first text;

[0038] Perform data normalization processing on the first text to obtain the second text;

[0039] For the second text, organize and save it in JSON format to obtain the processed text.

[0040] In a feasible implementation manner, according to the list of organization names in a specific area obtained in advance, the present invention searches for the name of each organization in the list one by one through a search engine, so as to obtain the network text data corresponding to the organization name. And organize it into CSV format, that is, a row-column structure. Each row corresponds to an organization, and each row includes four columns, which respectively represent the organization name, organization profile, organization directory, and organization details.

[0041] Perform data cleaning on the obtained network text data. Among them, it includes:

[0042] Since the information of some organizations is incomplete or not even included on the Internet, resulting in the lack of some columns in the organization description text retrieved by the search engine, the rows with missing organization profiles and organization details are deleted.

[0043] Delete the unnecessary spaces and special characters in the organization name, organization profile, organization directory, and organization details to ensure the standardization and neatness of the text.

[0044] For the obtained detailed organization text, organize it in JSON format to ensure a clear and definite mapping relationship between the organization name and its text description. Write the constructed JSON format data into a file to achieve data persistent storage.

[0045] S2. Based on the Word2Vec technology, construct an initial training set according to the preset list of organization names, the processed text, the preset training prompt words, and the preset benchmark organization relationship.

[0046] Optionally, based on the Word2Vec technology, constructing an initial training set according to the preset list of organization names, the processed text, the preset training prompt words, and the preset benchmark organization relationship includes:

[0047] Obtain the organization name according to the preset list of organization names;

[0048] Obtain the detailed text description corresponding to the organization name according to the processed text and the organization name;

[0049] Based on the Word2Vec technology, construct triples according to the organization name, the preset list of organization names, the detailed text description, and the preset benchmark organization relationship to obtain organization relationship triples;

[0050] Construct training triples according to the organization relationship triples, the preset training prompt words, and the detailed text description;

[0051] Determine the training triples as the initial training set.

[0052] In a feasible implementation, Word2Vec used in the present invention is a word embedding technology that maps words to a high-dimensional vector space through a neural network. It can capture the semantic and context relationships between words. By converting each word into a vector of a fixed dimension, Word2Vec makes words with similar semantics closer in the vector space, enabling more effective text analysis and processing.

[0053] Among them, the Skip-gram model is a model for training Word2Vec word embeddings. Its goal is to predict the context words given a center word. Specifically, the Skip-gram model learns the word embedding vector of the center word, making this vector maximize the occurrence probability of the context words around the center word, thereby generating high-quality word embedding representations.

[0054] On this basis, an initial training set of the large language model is constructed based on the organization relationship triples. Each piece of training data can be represented as a triple <training prompt word, detailed text description, organization relationship triple>.

[0055] The "training prompt word" of the training triple can be set as "You are now an organization relationship extraction model. Do you think there is an alias, subordination, cooperation, competition, merger and acquisition, or shareholding relationship between two organizations in the following text? If so, please output it in the form of a triple, and connect the inside of the triple with ','. If the above relationships are not included, please output unknown."

[0056] Optionally, based on the Word2Vec technology, organization relationship triples are constructed according to the organization name, the preset list of organization names, the detailed text description, and the preset benchmark organization relationships, including:

[0057] Use a natural language tokenization tool to tokenize the detailed text description to obtain a noun set and a verb set;

[0058] Based on the preset list of organization names, in the noun set, synonym screening is performed through the Word2Vec technology to obtain the recognized organization names;

[0059] Based on the preset benchmark organization relationships, in the verb set, synonym screening is performed through the Word2Vec technology to obtain the first recognized relationship;

[0060] Construct triples according to the organization name, the recognized organization name, and the first recognized relationship to obtain organization relationship triples.

[0061] In a feasible implementation, the present invention is based on the types of common relationships between organizations determined by domain experts, including but not limited to cooperation, competition, mergers and acquisitions, shareholding, aliases, and affiliations, etc. These relationships are used as the benchmark organizational relationships in the organizational relationship recognition task.

[0062] Based on a list of organization names within a certain area obtained in advance, use web crawler technology to crawl the name of each organization in the search list on the Internet and obtain the detailed text description corresponding to the organization name. Use natural language tokenization tools to tokenize the original web text, and extract the noun set and verb set corresponding to each piece of text.

[0063] Since in the original web text, the same organization may have multiple different aliases, and the same relationship may also have multiple different expressions, therefore, we need to perform synonym analysis on the noun set and verb set in each piece of web text. Use the Word2Vec word embedding technology to input the ordered tokenization results of all crawled web texts into the model for training.

[0064] Through the Word2Vec model, convert the vocabulary in the noun set and verb set corresponding to each piece of web text into corresponding word embedding vectors. For the embedding vector of each noun in the noun set, calculate its cosine similarity with the embedding vector of the organization name in the pre-obtained organization name list. If the similarity is greater than a specific threshold, the two nouns are recognized as the same organization; for each verb in the verb set, calculate its cosine similarity with the embedding vector of the common relationship between organizations determined in advance. If the similarity is greater than a specific threshold, the two verbs are recognized as the same relationship.

[0065] Usually, a piece of web text corresponds to two organizations and one relationship. If more than two organizations or more than one relationship are initially identified, manual screening is required to ensure the correctness of the annotation.

[0066] S3. According to the initial training set, use the low-rank adaptation fine-tuning method to fine-tune the pre-set large language model to obtain the first organizational relationship recognition model.

[0067] Optionally, according to the initial training set, use the low-rank adaptation fine-tuning method to fine-tune the pre-set large language model to obtain the first organizational relationship recognition model, including:

[0068] According to the pre-set large language model, obtain the original parameter matrix;

[0069] During the forward propagation process, freeze the original parameter matrix, train the pre-set large language model, and obtain the updated gradient matrix A and the updated gradient matrix B;

[0070] Calculate based on the original parameter matrix, the updated gradient matrix A, and the updated gradient matrix B to obtain the updated parameter matrix; the updated parameter matrix is the sum of the original parameter matrix and the gradient matrix; the gradient matrix includes the gradient matrix A and the gradient matrix B;

[0071] Assign the updated parameter matrix to a pre-set large language model to obtain the first organizational relationship recognition model.

[0072] In a feasible implementation, the present invention uses the open-source model ChatGLM3-6B Base in the new generation of dialogue pre-training model ChatGLM3 series jointly released by Zhipu AI and the KEG Laboratory of Tsinghua University. On the basis of retaining many excellent features of the previous two generations of models such as smooth dialogue and low deployment threshold, this model also has a more powerful basic model, more complete function support, and a more comprehensive open-source sequence.

[0073] Select appropriate model hyperparameters, such as model size, number of iterations, learning rate, batch size, decoding type, number of training rounds, etc., and use the training data set to fine-tune the large language model so that the model can better adapt to the organizational relationship recognition task.

[0074] The present invention uses the Low-Rank Adaptation (LoRA) fine-tuning method to fine-tune the large language model. The optimization of the training process of LoRA is mainly reflected in its strategy for updating model parameters. In traditional fine-tuning methods, all parameters of the model are updated according to the data of the downstream task, which not only requires a large amount of computing resources but also may lead to overfitting due to too many parameters. To solve this problem, LoRA proposes to add a low-rank decomposition matrix to the parameter matrix of the pre-trained model to approximate the parameter update of each layer, thereby reducing the parameters required to adapt to the downstream task. Given a parameter matrix , its update process can be generally expressed as the following formula (1):

[0075] (1)

[0076] where, is the original parameter matrix, is the updated gradient matrix.

[0077] The weight matrix of the pre-trained model, is frozen during the fine-tuning process and does not directly participate in the gradient update. Instead, LoRA introduces a low-rank decomposition matrix for each weight matrix, that is, the update of the weight is approximated by the product of two small matrices and , where, , , and J is much smaller than I. During the training process, only and the parameters in need to be trained and updated, while the original weight matrix remains unchanged. During the forward propagation process, the formula for the original calculation of the intermediate state is as follows in Equation (2):

[0078] (2)

[0079] After the training is completed, further combine the original parameter matrix and the matrix obtained from training and the matrix to obtain the updated parameter matrix . Therefore, the LoRA method not only significantly reduces the number of parameters that need to be updated during model training, but also reduces the model's demand for computing resources.

[0080] S4. Based on the first organizational relationship recognition model, construct an incremental training set according to the preset training prompt words and the processed text.

[0081] Optionally, constructing an incremental training set based on the first organizational relationship recognition model according to the preset training prompt words and the processed text includes:

[0082] Obtain unlabeled network text according to the processed text;

[0083] Perform organizational relationship recognition on the unlabeled network text through the first organizational relationship recognition model to obtain a second recognition relationship;

[0084] Correct the second recognition relationship to obtain a third recognition relationship;

[0085] Construct incremental triples according to the third recognition relationship, the preset training prompt words, and the processed text;

[0086] Determine the incremental triples as the incremental training set.

[0087] In a feasible implementation manner, since the initial training set is constructed by combining the Word2Vec algorithm and manual verification, the scale of the initial training set is relatively small, resulting in poor performance of the large language model in organizational relationship recognition after the first round of fine-tuning.

[0088] To overcome the above defects, the present invention performs incremental training on the large language model. Fine-tune the large language model using the initial training set to enable it to achieve good recognition performance on the initial training set. Use the fine-tuned large language model to automatically recognize a part of the unlabeled text data, manually verify the recognition results, manually correct the misclassified data, and use the corrected samples as the incremental training set.

[0089] S5. Based on the preset number of training rounds, use the incremental training set to perform incremental training on the first organizational relationship recognition model to obtain the second organizational relationship recognition model.

[0090] In a feasible implementation, use the incremental training set to perform incremental training on the first organizational relationship recognition model. The incremental training process is repeated. When the number of training rounds reaches the preset threshold, stop the training, and use the organizational relationship recognition model of this round as the second organizational relationship recognition model.

[0091] S6. Obtain the list of organization names to be recognized; perform text crawling and preprocessing according to the list of organization names to be recognized to obtain the network text to be recognized; input the network text to be recognized into the second organizational relationship recognition model to obtain the organizational recognition relationship.

[0092] In a feasible implementation, the present invention fine-tunes a pre-trained large language model using a relatively small dataset related to a specific task, enabling it to better complete specific application tasks. Applying the large language model to organizational relationship recognition can effectively improve the accuracy of organizational relationship recognition. The accuracy and efficiency of organizational relationship recognition will directly affect the quality and effect of decision-making in related fields, which is of great significance for promoting social progress and knowledge innovation.

[0093] The present invention proposes a method for identifying inter-organizational relationships based on network text mining, which captures rich language representations by learning the patterns, structures, and grammars of language through a large language model. A pre-trained large language model is fine-tuned using a relatively small dataset related to a specific task, enabling it to better complete specific application tasks. The present invention can identify various possible relationships between different organizations from a large amount of unstructured text, and effectively improve the recognition accuracy of inter-organizational relationships with the help of a large language model. The present invention is an accurate and efficient method for identifying inter-organizational relationships for unstructured text.

[0094] Figure 2 It is a block diagram of an apparatus for identifying inter-organizational relationships based on network text mining shown according to an exemplary embodiment. This apparatus is used for the method of identifying inter-organizational relationships based on network text mining. Refer to Figure 2 , this apparatus includes a text acquisition module 210, an initial training set construction module 220, a preliminary training module 230, an incremental training set construction module 240, an incremental training module 250, and an organizational relationship recognition module 260. Among them:

[0095] The text acquisition module 210 is used to perform text crawling according to the preset list of organization names to obtain the original network text; perform preprocessing on the original network text to obtain the processed text;

[0096] The initial training set construction module 220 is used to construct an initial training set based on the Word2Vec technology, according to a preset list of organization names, processed text, preset training prompt words, and preset benchmark organization relationships;

[0097] The preliminary training module 230 is used to fine-tune a preset large language model according to the initial training set by using the low-rank adaptation fine-tuning method to obtain a first organization relationship recognition model;

[0098] The incremental training set construction module 240 is used to construct an incremental training set based on the first organization relationship recognition model, according to preset training prompt words and processed text;

[0099] The incremental training module 250 is used to perform incremental training on the first organization relationship recognition model based on a preset number of training rounds, using the incremental training set, to obtain a second organization relationship recognition model;

[0100] The organization relationship recognition module 260 is used to obtain a list of organization names to be recognized; perform text crawling and preprocessing according to the list of organization names to be recognized to obtain network text to be recognized; input the network text to be recognized into the second organization relationship recognition model to obtain organization recognition relationships.

[0101] Optionally, the text acquisition module 210 is further used for:

[0102] Perform missing value processing on the original network text to obtain a first text;

[0103] Perform data normalization processing on the first text to obtain a second text;

[0104] Perform structured organization and storage on the second text in JSON format to obtain processed text.

[0105] Optionally, the initial training set construction module 220 is further used for:

[0106] Obtain organization names according to a preset list of organization names;

[0107] Obtain a detailed text description corresponding to the organization name according to the processed text and the organization name;

[0108] Perform triple construction based on the Word2Vec technology, according to the organization name, preset list of organization names, detailed text description, and preset benchmark organization relationships, to obtain organization relationship triples;

[0109] Construct training triples according to the organization relationship triples, preset training prompt words, and detailed text description;

[0110] Determine the training triples as the initial training set.

[0111] Optionally, the initial training set construction module 220 is further configured to:

[0112] Use a natural language tokenization tool to tokenize the detailed text description to obtain a noun set and a verb set;

[0113] Based on a preset list of organization names, perform synonym screening on the noun set through the Word2Vec technique to obtain recognized organization names;

[0114] Based on a preset benchmark organization relationship, perform synonym screening on the verb set through the Word2Vec technique to obtain a first recognized relationship;

[0115] Construct a triple according to the organization name, recognized organization name, and first recognized relationship to obtain an organization relationship triple.

[0116] Optionally, the preliminary training module 230 is further configured to:

[0117] Obtain an original parameter matrix according to a preset large language model;

[0118] During the forward propagation process, freeze the original parameter matrix and train the preset large language model to obtain an updated gradient matrix A and an updated gradient matrix B;

[0119] Calculate according to the original parameter matrix, updated gradient matrix A, and updated gradient matrix B to obtain an updated parameter matrix; the updated parameter matrix is the sum of the original parameter matrix and the gradient matrix; the gradient matrix includes gradient matrix A and gradient matrix B;

[0120] Assign the updated parameter matrix to the preset large language model to obtain a first organization relationship recognition model.

[0121] Optionally, the incremental training set construction module 240 is further configured to:

[0122] Obtain unlabeled network text according to the processed text;

[0123] Perform organization relationship recognition on the unlabeled network text through the first organization relationship recognition model to obtain a second recognized relationship;

[0124] Correct the second recognized relationship to obtain a third recognized relationship;

[0125] Construct an incremental triple according to the third recognized relationship, preset training prompt words, and processed text;

[0126] Determine the incremental triple as the incremental training set.

[0127] The present invention proposes a method for identifying inter-organization relationships based on web text mining, which captures rich language representations by learning the patterns, structures, and grammars of languages through large language models. The pre-trained large language model is fine-tuned using a relatively small dataset related to a specific task to enable it to better complete specific application tasks. The present invention can identify various possible relationships between different organizations from a large amount of irregular texts, and effectively improve the accuracy of identifying inter-organization relationships with the help of large language models. The present invention is an accurate and efficient method for identifying inter-organization relationships for unstructured texts.

[0128] Figure 3 FIG. is a schematic structural diagram of an inter-organization relationship recognition device provided by an embodiment of the present invention, as Figure 3 shown, the inter-organization relationship recognition device may include the above-mentioned Figure 2 inter-organization relationship recognition device based on web text mining shown. Optionally, the inter-organization relationship recognition device 310 may include a first processor 2001.

[0129] Optionally, the inter-organization relationship recognition device 310 may further include a memory 2002 and a transceiver 2003.

[0130] Among them, the first processor 2001, the memory 2002, and the transceiver 2003, such as, may be connected through a communication bus.

[0131] Next, in conjunction with Figure 3 each component of the inter-organization relationship recognition device 310 will be specifically introduced:

[0132] Among them, the first processor 2001 is the control center of the inter-organization relationship recognition device 310, which may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, for example: one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs).

[0133] Optionally, the first processor 2001 may execute various functions of the inter-organization relationship recognition device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0134] In a specific implementation, as an example, the first processor 2001 may include one or more CPUs, such as Figure 3 the CPU0 and CPU1 shown in

[0135] In a specific implementation, as an example, the inter-organization relationship recognition device 310 may also include multiple processors, such as Figure 3 the first processor 2001 and the second processor 2004 shown in. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0136] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 to execute. The specific implementation manner may refer to the above method embodiments and will not be elaborated here.

[0137] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic storage medium such as a disk storage, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the inter-organization relationship recognition device 310 ( Figure 3 not shown in). The embodiments of the present invention do not make specific limitations on this.

[0138] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0139] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not separately shown in). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0140] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit (not shown in Figure 3 ) of the inter-organization relationship recognition device 310. The embodiments of the present invention do not make specific limitations on this.

[0141] It should be noted that Figure 3 the structure of the inter-organization relationship recognition device 310 shown in [] does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0142] In addition, the technical effects of the inter-organization relationship recognition device 310 may refer to the technical effects of the inter-organization relationship recognition method based on network text mining described in the above method embodiments, which will not be elaborated here.

[0143] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0144] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0145] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0146] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0147] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0148] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0149] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0150] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0151] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0152] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0154] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0155] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for identifying inter-organizational relationships based on network text mining, characterized in that: The method comprises: Crawling text according to a preset list of organization names to obtain original network text; preprocessing the original network text to obtain processed text; Based on Word2Vec technology, an initial training set is constructed according to a preset list of organization names, the processed text, preset training prompt words, and preset benchmark organization relationships; According to the initial training set, a low-rank adaptation fine-tuning method is used to fine-tune the preset large language model to obtain a first organizational relationship recognition model; Based on the first organizational relationship recognition model, constructing an incremental training set according to preset training prompt words and the processed text; Based on a preset training round, using the incremental training set, incrementally training the first organizational relationship recognition model to obtain a second organizational relationship recognition model; Obtain a list of organization names to be identified; perform text crawling and preprocessing based on the list of organization names to be identified to obtain network text to be identified; input the network text to be identified into the second organization relationship identification model to obtain organization identification relationships.

2. The method for identifying inter-organizational relationships based on network text mining according to claim 1, characterized in that: The preprocessing of the original network text to obtain the processed text includes: Performing missing value processing on the original network text to obtain a first text; Performing data normalization processing on the first text to obtain a second text; The second text is structured and organized and saved using the JSON format to obtain a processed text.

3. The method for identifying inter-organizational relationships based on network text mining according to claim 1, characterized in that: The Word2Vec technology is used to construct an initial training set according to a preset list of organization names, the processed text, preset training prompt words, and preset benchmark organization relationships, including: Get the organization name according to the preset organization name list; According to the processed text and the organization name, obtaining a detailed text description corresponding to the organization name; Based on the Word2Vec technology, triples are constructed according to the organization name, the preset organization name list, the detailed text description and the preset benchmark organization relationship to obtain an organization relationship triple; Constructing a training triplet according to the organizational relationship triplet, the preset training prompt words and the detailed text description; The training triples are determined as an initial training set.

4. The method for identifying inter-organizational relationships based on network text mining according to claim 3 is characterized in that: The Word2Vec technology is used to construct triples according to the organization name, the preset organization name list, the detailed text description and the preset benchmark organization relationship to obtain the organization relationship triples, including: Using a natural language word segmentation tool, performing word segmentation processing on the detailed text description to obtain a noun set and a verb set; Based on a preset list of organization names, in the noun set, synonym screening is performed using the Word2Vec technology to obtain recognized organization names; Based on the preset benchmark organizational relationship, in the verb set, synonym screening is performed by using the Word2Vec technology to obtain a first recognition relationship; A triple is constructed according to the organization name, the identification organization name and the first identification relationship to obtain an organization relationship triple.

5. The method for identifying inter-organizational relationships based on network text mining according to claim 1, characterized in that: The method of fine-tuning the preset large language model using a low-rank adaptation fine-tuning method according to the initial training set to obtain a first organizational relationship recognition model includes: According to the preset large language model, the original parameter matrix is ​​obtained; In the forward propagation process, the original parameter matrix is ​​frozen, and the preset large language model is trained to obtain an updated gradient matrix A and an updated gradient matrix B; Calculate according to the original parameter matrix, the updated gradient matrix A and the updated gradient matrix B to obtain an updated parameter matrix; the updated parameter matrix is ​​the sum of the original parameter matrix and the gradient matrix; the gradient matrix includes the gradient matrix A and the gradient matrix B; The updated parameter matrix is ​​assigned to a preset large language model to obtain a first organizational relationship recognition model.

6. The method for identifying inter-organizational relationships based on network text mining according to claim 1, characterized in that: The step of constructing an incremental training set based on the first organizational relationship recognition model and according to preset training prompt words and the processed text includes: According to the processed text, obtaining unlabeled network text; Using the first organizational relationship recognition model, the unlabeled network text is subjected to organizational relationship recognition to obtain a second recognition relationship; Correcting the second identification relationship to obtain a third identification relationship; constructing an incremental triple according to the third recognition relationship, the preset training prompt word and the processed text; The incremental triples are determined as an incremental training set.

7. An inter-organizational relationship identification device based on network text mining, the inter-organizational relationship identification device based on network text mining is used to implement the inter-organizational relationship identification method based on network text mining as claimed in any one of claims 1 to 6, characterized in that: The device comprises: A text acquisition module is used to crawl text according to a preset list of organization names to obtain original network text; pre-process the original network text to obtain processed text; An initial training set construction module, used to construct an initial training set based on Word2Vec technology according to a preset list of organization names, the processed text, preset training prompt words and preset benchmark organization relationships; A preliminary training module, used to fine-tune a preset large language model using a low-rank adaptation fine-tuning method according to the initial training set to obtain a first organizational relationship recognition model; An incremental training set construction module, used to construct an incremental training set based on the first organizational relationship recognition model, according to preset training prompt words and the processed text; An incremental training module, configured to perform incremental training on the first organizational relationship recognition model based on a preset training round and using the incremental training set to obtain a second organizational relationship recognition model; The organization relationship identification module is used to obtain a list of organization names to be identified; perform text crawling and preprocessing based on the list of organization names to be identified to obtain network text to be identified; input the network text to be identified into the second organization relationship identification model to obtain organization identification relationships.

8. The device for identifying inter-organizational relationships based on network text mining according to claim 7, characterized in that: The preliminary training module is further used to: According to the preset large language model, the original parameter matrix is ​​obtained; In the forward propagation process, the original parameter matrix is ​​frozen, and the preset large language model is trained to obtain an updated gradient matrix A and an updated gradient matrix B; Calculate according to the original parameter matrix, the updated gradient matrix A and the updated gradient matrix B to obtain an updated parameter matrix; The updated parameter matrix is ​​the sum of the original parameter matrix and the gradient matrix; the gradient matrix includes the gradient matrix A and the gradient matrix B; The updated parameter matrix is ​​assigned to a preset large language model to obtain a first organizational relationship recognition model.

9. An inter-organizational relationship identification device, characterized in that: The inter-organization relationship identification device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Entity subordination relation extraction and identification method, system and device based on NLP and trigger and storage medium

    CN114625885A

  • Multi-value chain data management aided decision-making model construction method based on knowledge graph

    CN114911945A