Method and device for processing entity relationship in text information
By processing Chinese text through a bidirectional encoder and a long short-term memory network and dynamically adjusting the loss weight, the problems of repetition and inaccurate recognition in Chinese entity relationship extraction are solved, achieving more efficient entity relationship extraction.
Patent Information
- Application Number
- CN202510848462.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
Smart Images

Figure CN120764658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a method and device for processing entity relationships in text information. Background Art
[0002] Entity relationship extraction (ERE) is a key research direction in natural language processing and a popular topic in text information processing. It aims to extract valuable information from large amounts of unstructured text. Entity relationship extraction (ERE) involves two tasks: named entity recognition (NER) and relationship extraction (RE). Currently, supervised ERE based on deep learning is divided into two approaches: end-to-end learning and joint learning. The first approach treats NER and RER as two independent tasks, which are completed incrementally without sharing learning parameters. The second approach, joint learning, shares learning parameters for both tasks. The former avoids conflicting optimization objectives, while the latter leads to poor performance in balancing the two tasks, resulting in performance degradation in one task. Furthermore, the black-box nature of the joint model makes debugging and improvement difficult when the results are erroneous. Although end-to-end learning methods have made significant progress in Chinese entity relationship extraction, overlapping relationship triplets remain common in real-world scenarios with massive amounts of text. The accuracy of subject and object entity recognition is also low, which can easily lead to errors in entity and relationship alignment, significantly increasing the complexity of the extraction task. Summary of the Invention
[0003] The present invention provides a method and device for processing entity relationships in text information, which solves the problems of repeated appearance of a single entity in multiple triples, repeated appearance of entity pairs under different relationships, and low accuracy of Chinese entity recognition during the extraction of Chinese entity relationships.
[0004] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0005] An embodiment of the present invention provides a method for processing entity relationships in text information, comprising:
[0006] Obtain target text data to be processed;
[0007] Performing bidirectional encoder representation encoding on the target text data to obtain a target text vector;
[0008] Performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data;
[0009] Injecting main entity information into the target text vector to obtain an intermediate text vector;
[0010] Performing object entity feature boundary prediction on the intermediate text vector to obtain an object entity prediction boundary set in the target text data;
[0011] The main entity prediction boundary set, the object entity prediction boundary set, and the preset relationship between the main entity features and the object entity features are output in the form of triples.
[0012] Optionally, performing bidirectional encoder representation encoding on the target text data to obtain a target text vector includes:
[0013] Perform word segmentation vector processing on the target text data to obtain multiple word segmentation vectors;
[0014] Encode multiple word segmentation vectors according to the dimensions of the preset text vector processing model to obtain the target text vector.
[0015] Optionally, performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data includes:
[0016] Performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector;
[0017] The main entity start position and the main entity end position are predicted according to the first enhancement vector to obtain a main entity prediction boundary set.
[0018] Optionally, performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector includes:
[0019] Performing forward encoding on the words of the target text vector according to the forward long short-term memory network to obtain a first forward encoding vector;
[0020] Reversely encode the words of the target text vector according to the backward long short-term memory network to obtain a first reverse encoding vector;
[0021] The first forward coding vector and the first reverse coding vector are concatenated to obtain a first enhanced vector.
[0022] Optionally, predicting a main entity start position and a main entity end position according to the first enhancement vector to obtain a main entity prediction boundary set includes:
[0023] Predicting a starting position of the main entity according to the first enhancement vector to obtain a first prediction probability;
[0024] Predicting the end position of the main entity according to the first enhancement vector to obtain a second prediction probability;
[0025] Obtaining a first labeled value according to the first predicted probability and a first preset threshold;
[0026] Obtaining a second marked value according to the second predicted probability and a second preset threshold;
[0027] A main entity prediction boundary set is obtained according to the first label value and the second label value.
[0028] Optionally, injecting main entity information into the target text vector to obtain an intermediate text vector includes:
[0029] Performing weighted average pooling processing on the target text vector and the main entity prediction boundary set to obtain weighted information;
[0030] The weighted information is embedded into the target text vector to obtain an intermediate text vector.
[0031] Optionally, performing object entity feature boundary prediction on the intermediate text vector to obtain an object entity prediction boundary set in the target text data includes:
[0032] Performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector;
[0033] The guest entity start position and the guest entity end position are predicted according to the second enhanced vector to obtain a guest entity prediction boundary set.
[0034] Optionally, performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector includes:
[0035] Performing transformer encoding on the intermediate text vector to generate context-dependent representation data;
[0036] Performing forward encoding on the words of the context-related representation data according to a forward long short-term memory network to obtain a second forward encoding vector;
[0037] Reverse encoding the words of the context-related representation data according to a backward long short-term memory network to obtain a second reverse encoding vector;
[0038] The second forward coding vector and the second reverse coding vector are concatenated to obtain a second enhanced vector.
[0039] Optionally, predicting a guest entity start position and a guest entity end position according to the second enhanced vector to obtain a guest entity prediction boundary set includes:
[0040] Predicting a starting position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a third prediction probability;
[0041] Predicting an end position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a fourth prediction probability;
[0042] Obtaining a third annotation value according to the third predicted probability and a third preset threshold;
[0043] Obtaining a fourth marked value according to the fourth predicted probability and a fourth preset threshold;
[0044] Obtaining a predicted boundary set of an object entity according to the third annotation value and the fourth annotation value;
[0045] Among them, during the customer entity prediction process, the customer entity weight is adjusted in real time through dynamic weighting.
[0046] An embodiment of the present invention further provides a device for processing entity relationships in text information, comprising:
[0047] An acquisition module, used to acquire target text data to be processed;
[0048] A processing module is used to perform bidirectional encoder representation encoding on the target text data to obtain a target text vector; perform main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data; integrate main entity information into the target text vector to obtain an intermediate text vector; perform guest entity feature boundary prediction on the intermediate text vector to obtain a guest entity prediction boundary set in the target text data; and output the main entity prediction boundary set, the guest entity prediction boundary set, and the preset relationship between the main entity feature and the guest entity feature in the form of a triple.
[0049] An embodiment of the present invention further provides a device for processing entity relationships in text information, comprising:
[0050] An acquisition module, used to acquire target text data to be processed;
[0051] A processing module is used to perform bidirectional encoder representation encoding on the target text data to obtain a target text vector; perform main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data; integrate main entity information into the target text vector to obtain an intermediate text vector; perform guest entity feature boundary prediction on the intermediate text vector to obtain a guest entity prediction boundary set in the target text data; and output the main entity prediction boundary set, the guest entity prediction boundary set, and the preset relationship between the main entity feature and the guest entity feature in the form of a triple.
[0052] The technical solution of the present invention includes at least the following effects:
[0053] The above scheme of the present invention obtains the target text data to be processed; performs a bidirectional encoder representation encoding on the target text data to obtain a target text vector; performs main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data; performs main entity information injection on the target text vector to obtain an intermediate text vector; performs guest entity feature boundary prediction on the intermediate text vector to obtain a guest entity prediction boundary set in the target text data; outputs the main entity prediction boundary set, the guest entity prediction boundary set, and the preset relationship between the main entity feature and the guest entity feature in the form of triples. The technical scheme of the present invention can significantly reduce the probability of triple overlap in massive text scenarios, so that entities and relationships clearly correspond, reduce the complexity of Chinese extraction tasks, and improve the accuracy of main entity recognition and guest entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flow chart of a method for processing entity relationships in text information provided by an embodiment of the present invention;
[0055] Figure 2 It is a roadmap for a method for processing entity relationships in text information provided by an embodiment of the present invention;
[0056] Figure 3 It is a model structure diagram of the method for processing entity relationships in text information provided by an embodiment of the present invention;
[0057] Figure 4 It is a structural diagram of an apparatus for processing entity relationships in text information provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0058] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0059] like Figure 1 As shown, an embodiment of the present invention provides a method and apparatus for processing entity relationships in text information, including:
[0060] Step 11, obtaining target text data to be processed; the target text data can be obtained by the system through the data interface for interacting with other data sources to collect text data to be processed from different data sources, and the text data includes various types of Chinese text data such as news articles, social media posts, academic papers and the like, and the obtained target text data will be used as the basis for subsequent processing;
[0061] Step 12, encoding the target text data by using a bidirectional encoder representation to obtain a target text vector;
[0062] Step 13, performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data;
[0063] Step 14, performing main entity information injection on the target text vector to obtain an intermediate text vector;
[0064] Step 15, performing guest entity feature boundary prediction on the intermediate text vector to obtain a guest entity prediction boundary set in the target text data;
[0065] Step 16, outputting the main entity prediction boundary set, the guest entity prediction boundary set and a preset relationship between the main entity feature and the guest entity feature in the form of a triple.
[0066] In this embodiment, the main entity and the guest entity are recognized by a named entity recognition method to be entities with specific meanings in the text, mainly including names, place names, organization names and proper nouns. The preset relationship between the main entity feature and the guest entity feature is a semantic relationship between entities detected and classified from the text or other data sources;
[0067] This method accurately extracts the relevant information of the main entity and the guest entity from the given target text data, and presents the information together with the preset relationship between the main entity feature and the guest entity feature in the form of a structured triple, which is convenient for subsequent data processing and application.
[0068] The target text data will be encoded by the trained model, so as to fully consider the context information of each word in the text and model the long-distance dependence with the information at other positions in the text. The specific process can include:
[0069] The target text data is input into the model, and the model will encode each word in the text to generate a corresponding word vector;
[0070] The generated word vector not only contains the semantic information of the word itself, but also fuses the semantic features of the word in the context of the text;
[0071] Finally, all word vectors are combined to form a target text vector that can fully represent the semantic information of the target text.
[0072] After obtaining the target text vector, the system will use the text feature enhancement method to enhance the main entity features of the target text vector, thereby generating a first enhanced vector;
[0073] The model predicts the starting and ending positions of the main entity of the first enhanced vector, that is, predicts the boundaries of the main entity: for example, in a sentence, it determines whether a noun or noun phrase is the main entity and marks its specific position range in the text, thereby obtaining the predicted boundaries of all main entities in the target text data, thereby forming a set of main entity prediction boundaries.
[0074] Principal entity information injection involves incorporating information related to the identified principal entity into the target text vector. This process involves concatenating or fusing the weighted principal entity information with the target text vector. Injecting principal entity information allows subsequent processing to focus more on guest entity information related to the principal entity, improving guest entity recognition accuracy. After principal entity information injection, an intermediate text vector containing the principal entity information is generated.
[0075] Using a prediction model similar to the one used to predict the feature boundaries of the main entity, the intermediate text vector is processed. Based on the injected main entity information, the prediction model searches for guest entities related to the main entity in the text. Specifically, it predicts the starting and ending positions of the guest entity and determines its boundaries. Finally, it obtains the predicted boundaries of all guest entities in the target text data, forming a set of predicted guest entity boundaries. During the guest entity prediction process, loss weights are adjusted in real time through dynamic weighting, allowing the model to adaptively focus on different tasks. This strengthens the accuracy of the prediction model, especially the guest entity prediction model, and thus improves the performance of extracting Chinese entity relationships.
[0076] After obtaining the main entity prediction boundary set and the guest entity prediction boundary set, the system will extract the specific feature information of the main entity and the guest entity, and organize and output this information in the form of triples based on the preset relationship between the main entity features and the guest entity features. The specific output format is a typical triple form, namely (main entity, relationship, guest entity). This form will facilitate subsequent knowledge storage, reasoning and application.
[0077] This method makes full use of the model's learning and growth capabilities, language understanding and context perception capabilities, and can accurately capture the semantic information in the text. It effectively improves the accuracy of entity boundary recognition by adjusting the loss weight in real time, and reduces the inaccurate entity information extraction caused by boundary recognition errors. At the same time, it greatly reduces the probability of overlapping output triples by presetting the relationship between entities and predicting the guest entity through the association of the main entity.
[0078] In an optional embodiment of the present invention, step 11 may include:
[0079] Step 111, determining the target text data to be processed currently;
[0080] Step 112, reading the target text data;
[0081] Wherein, the target text data is Chinese text.
[0082] In this embodiment, step 111 represents the formal start of the processing process for specific Chinese text data in the data processing flow. The system will identify the data source that stores the data, including a database, file system, network service or any other system that can store and provide text data; after identifying the data source, the data is screened according to certain conditions or rules to find the Chinese target text data that needs to be processed, ensuring that the selected target text data meets the requirements of subsequent processing; after determining the data that needs to be processed, the data is accurately located in the data source through indexing, querying or other data retrieval technologies to ensure that the target text data can be quickly and accurately obtained. After locating the target text data, in order to prevent the data from being operated by other processes during the processing process, the data will be locked first. The specific locking mechanism varies depending on the data source, but the basic principle is to ensure the stability and consistency of the data during data processing.
[0083] After determining the target text data to be processed, the data is read. The reading process is to load the text data stored in the data source into the memory or other processing environment for subsequent analysis, processing or conversion. According to the type and characteristics of the data source, select the appropriate reading interface or tool, and consider the conversion of data format to convert different storage formats into plain text format; and design a complete error handling mechanism to ensure that when problems are encountered during the reading process, these errors can be captured and handled in time to avoid affecting the subsequent data processing process; for large-scale data sets, the reading process will become a performance bottleneck. Therefore, performance optimization issues also need to be considered during the reading process, and try to use caching technology and parallel reading strategies to improve reading efficiency.
[0084] In an optional embodiment of the present invention, step 12 may include:
[0085] Step 121, performing word segmentation vector processing on the target text data to obtain multiple word segmentation vectors;
[0086] Step 122: Encode the multiple word segmentation vectors according to the dimensions of the preset text vector processing model to obtain a target text vector.
[0087] In this embodiment, the target text data is first segmented. The purpose of segmentation is to split the continuous text sequence into meaningful units, which will be converted into vector representations later. Since the target text data is Chinese data, and Chinese text is usually segmented in units of words or characters, this technical solution chooses to use characters for segmentation to facilitate subsequent operations, and each character is used as a segmentation. After the segmentation is completed, each segmentation will be mapped to a vector space to form a segmentation vector. The segmentation vector contains the semantic information of the segmentation, so that segments with similar semantics are closer in the vector space. It can be obtained by looking up a predefined word vector table or using an embedding layer to learn in the model training process, so that the target text data is converted into a series of segmentation vectors, each segmentation corresponds to a vector.
[0088] The pre-set text vector processing model has specific input and output dimensions. The input dimensions match the dimensions of the word segmentation vectors, while the output dimensions are the dimensions of the text representation learned by the model. During the encoding process, the model receives multiple word segmentation vectors as input and processes these vectors through a series of neural network layers. The neural network layers capture the contextual relationships between word segments and then dynamically adjust the representation of each word segmentation vector using a self-attention mechanism, making the vector representation contain richer contextual information. After the model processing is completed, a vector is extracted from a specific position in the model as the vector representation of the entire text. This vector represents the model's comprehensive understanding of the target text data and includes the text's semantics, grammar, and contextual information.
[0089] In an optional embodiment of the present invention, in step 122: encoding the multiple word segmentation vectors according to the dimensions of the preset text vector processing model to obtain the target text vector includes:
[0090] Step 1221, establishing a text vector processing model;
[0091] Step 1222: training the text vector processing model;
[0092] Step 1223: Obtain a target text vector by inputting multiple word segmentation vectors into the text vector processing model.
[0093] In this embodiment, a suitable bidirectional encoder representation model is selected to capture bidirectional contextual information in text. Based on the selected model, a corresponding neural network structure is constructed, including an input layer, multiple encoder layers, and an output layer. The input layer is used to receive word segmentation vectors, the multiple encoder layers are used to extract text features, and the output layer is used to generate target text vectors or other task-specific outputs.
[0094] The specific training process of the model is as follows: According to L MLM =-∑ i∈masked logP(w i |w \i ) Randomly mask some words in the text and let the model predict the word, and infer the masked word through context information, thereby learning bidirectional context information;
[0095] Among them, i is the position index of the masked token, masked is the set of all masked positions, P(w i |w \i ) indicates that the context w \i Predict the mask word w i probability.
[0096] The coherence between different sentences is then trained. Given two sentences, the model can determine whether they are adjacent or related. This helps the model correctly identify the coherence of contextual information, thereby improving its understanding of paragraphs. Furthermore, to directly model long-distance dependencies and effectively capture the semantic or grammatical dependencies between distant words or sentence components within a sentence or paragraph, a self-attention mechanism is used to calculate the association weights of each word with all other words in the sentence during encoding.
[0097] After the model training is completed, according to Convert the text into a target text vector, where S is the text; n is the number of tokens formed after a single line of text is segmented; d is the hidden dimension of the model, which defaults to 768, indicating that a line of Chinese text S is segmented into n tokens, and each token is extended into a vector of 1xd dimensions, that is, a vector of 1x768 dimensions; H (B) is the target text vector; BERT() is the model processing process.
[0098] like Figure 3 As shown, in an optional embodiment of the present invention, step 13 may include:
[0099] Step 131, performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector;
[0100] Step 132: predict the main entity start position and the main entity end position according to the first enhancement vector to obtain a main entity prediction boundary set.
[0101] In this embodiment, in step 131, the target text vector is sent to a bidirectional long short-term memory network for processing. This structure performs forward and backward propagation on each word segmentation vector to capture the semantic relationship between it and the context. By splicing the forward and backward outputs, an enhanced representation of each word segmentation is obtained after considering the context information, that is, the first enhanced vector. The first enhanced vector not only contains the semantic information of the word segmentation itself, but also incorporates its context in the text, making the subsequent main entity boundary prediction more accurate.
[0102] In step 132, the first enhancement vector obtained in step 131 is used to predict the boundary of the main entity. The first enhancement vector is received as input, and the probability of the start position or end position of the main entity is output through linear training and function processing. The final main entity prediction boundary set is obtained by comparing the probability with the preset threshold.
[0103] In an optional embodiment of the present invention, in step 131: performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector includes:
[0104] Step 1311, forward encoding the words of the target text vector according to the forward long short-term memory network to obtain a first forward encoding vector;
[0105] Step 1312, reverse encoding is performed on the words of the target text vector according to the backward long short-term memory network to obtain a first reverse encoding vector;
[0106] Step 1313: Concatenate the first forward coding vector and the first reverse coding vector to obtain a first enhanced vector.
[0107] In this embodiment, in order to more accurately capture the semantic and contextual information in the text, it is necessary to perform feature enhancement processing on the original target text vector, and use a bidirectional long short-term memory network to perform feature vector enhancement on the target text vector to generate a first enhanced vector.
[0108] In step 1311, the target text vector is sequentially fed into the forward long short-term memory network. For each word, the forward long short-term memory network considers the information of all the words preceding it and generates an encoding vector that contains the semantic and contextual information of the word and all the words preceding it. As the processing proceeds, the forward long short-term memory network gradually updates its internal state to reflect the information of the currently processed part of the text. After processing by the forward long short-term memory network, the corresponding first forward encoding vector is obtained, which captures the forward dependency in the text, that is, by encoding the contextual information from the first word to the i-th word, recorded as
[0109] In step 1312, the target text vector is sequentially fed into the backward long short-term memory network and processed in reverse order. The backward long short-term memory network considers the information of all the words behind it and generates an encoding vector that contains the semantic and contextual information of the word and all the words behind it. As the processing proceeds, the backward long short-term memory network gradually updates its internal state to reflect the information of the currently processed part of the text. After being processed by the backward long short-term memory network, the corresponding first reverse encoding vector is obtained, which captures the backward dependency in the text, that is, by encoding the contextual information from the i-th word to the n-th word, recorded as
[0110] In step 1313, the first forward encoding vector and the first reverse encoding vector are concatenated according to The concatenated feature vector is obtained, where i is the position index of the word segmentation; n is the number of tokens formed after the word segmentation of a single line of text; d b The unidirectional hidden dimension of the model, the default is 256, 2d b is the output dimension, that is, the vector is 2x256 dimensions; h i The function of the feature vector is to identify the main entity of the enhanced text feature vector, so the feature vector is recorded as H(sub) for subsequent understanding. This process can be summarized as Among them, H (B) is the target text vector; H(sub) is the first enhanced vector after enhancement of the main entity features; BILSTM() is the model processing. After the concatenation operation, the corresponding first enhanced vector is obtained. This vector contains both forward and backward dependency information, thus having stronger expressive power and effectively capturing the semantic and contextual information in the text, providing strong support for subsequent natural language processing tasks.
[0111] In an optional embodiment of the present invention, in step 132: predicting the main entity start position and the main entity end position according to the first enhancement vector to obtain a main entity prediction boundary set includes:
[0112] Step 1321: predict the starting position of the main entity according to the first enhancement vector to obtain a first prediction probability;
[0113] Step 1322: predict the end position of the main entity based on the first enhancement vector to obtain a second prediction probability;
[0114] Step 1323: Obtain a first annotation value according to the first predicted probability and a first preset threshold;
[0115] Step 1324: Obtain a second annotation value according to the second predicted probability and a second preset threshold;
[0116] Step 1325: Obtain a main entity prediction boundary set based on the first annotation value and the second annotation value.
[0117] In this embodiment, in step 1321, when predicting the starting position of the main entity, the Get the first predicted probability;
[0118] In step 1322, when predicting the end position of the main entity, the Obtain a second predicted probability;
[0119] Among them, W sh and W st is the weight parameter, b sh and b st is the bias, H(sub) is the first enhancement vector, σ() represents the function processing process, represents the first predicted probability, represents the second predicted probability. The range of the first predicted probability and the second predicted probability is [0,1].
[0120] In step 1323, the first predicted probability is compared with the first preset threshold to obtain a first labeling value. The first preset threshold is preferably 0.5. When the first predicted probability is equal to or exceeds the first preset threshold, the probability of this position is marked as 1 through the binary labeling method; when the first predicted probability is less than the first preset threshold, the probability of this position is marked as 0 through the binary labeling method.
[0121] In step 1324, the second predicted probability is compared with the second preset threshold to obtain a second labeling value. The second preset threshold is preferably 0.5. When the second predicted probability is equal to or exceeds the second preset threshold, the probability of this position is marked as 1 using the binary labeling method; when the second predicted probability is less than the second preset threshold, the probability of this position is marked as 0 using the binary labeling method.
[0122] It should be noted that the first preset threshold and the second preset threshold may be the same or different, and both the first preset threshold and the second preset threshold are preferably 0.5.
[0123] In step 1325 , after steps 1323 and 1324 , the location of the main entity can be determined according to the first label value and the second label value, that is, the main entity predicted boundary set is obtained.
[0124] In an optional embodiment of the present invention, step 14 may include:
[0125] Step 141: performing weighted average pooling processing on the target text vector and the main entity prediction boundary set to obtain weighted information;
[0126] Step 142: embed the weighted information into the target text vector to obtain an intermediate text vector.
[0127] In this embodiment, according to The target text vector and the main entity prediction boundary set are weighted average pooled to obtain a pooled vector; where n is the number of tokens formed after word segmentation of a single line of text; d is the hidden dimension of the model, which defaults to 768, and A sub is the predicted boundary set of the main entity, H (B) is the target text vector, h avg is the pooling vector, Avgpool() represents weighted average pooling. sub =σ(W r h avg +b r ) processes the pooled vector to obtain weighted main entity information; where W r is the weight, b r is the bias, σ() is the function processing process, h sub is weighted information. According to E=H (B) +h sub The weighted information is embedded into the main entity information to obtain the intermediate text vector; where H (B) is the target text vector, h sub is the weighted information, and E is the intermediate text vector.
[0128] In an optional embodiment of the present invention, step 15 may include:
[0129] Step 151, performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector;
[0130] Step 152: predict the guest entity start position and the guest entity end position according to the second enhanced vector to obtain a guest entity prediction boundary set.
[0131] In this embodiment, in step 151, the target text vector is sent to a bidirectional long short-term memory network for processing. This structure performs forward and backward propagation on each word segmentation vector to capture the semantic relationship between it and the context. By splicing the forward and backward outputs, an enhanced representation of each word segmentation is obtained after considering the context information, that is, the second enhanced vector. The second enhanced vector not only contains the semantic information of the word segmentation itself, but also incorporates its context in the text, making the subsequent object entity boundary prediction more accurate.
[0132] In step 152, the second enhancement vector obtained in step 151 is used to predict the boundary of the guest entity. The second enhancement vector is received as input, and through linear training and function processing, the probability of the starting position or the ending position of the main entity is output. By comparing the probability with the preset threshold, the final guest entity prediction boundary set is obtained.
[0133] In an optional embodiment of the present invention, in step 151: performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector includes:
[0134] Step 1511 , performing converter encoding on the intermediate text vector to generate context-related representation data;
[0135] Step 1512: forward-encode the words of the context-related representation data using a forward long short-term memory network to obtain a second forward-encoded vector;
[0136] Step 1513: reversely encode the words of the context-related representation data according to the backward long short-term memory network to obtain a second reverse encoding vector;
[0137] Step 1514: Concatenate the second forward coding vector and the second reverse coding vector to obtain a second enhanced vector.
[0138] In this embodiment, in step 1511, in order to further enhance the integration of the main entity information, the self-attention mechanism in the transformer encoding is used to assign different weights to the features of each part of the intermediate text vector, deeply extracting key information so that the model can fully utilize the features of the text. According to Q=EWQ ,K=EW K ,V=EW V , respectively, linearly transform the input intermediate text vector into query matrix, key matrix and value matrix, where Q represents query matrix, K represents key matrix, V represents value matrix, E represents intermediate text vector, W Q is the query weight matrix, W K is the key weight matrix, W V is the value weight matrix. Calculate the dot product of the transpose of the query matrix and the key matrix to obtain an attention score matrix. Each element in the attention score matrix is uniformly reduced to prevent the value from being too large. Apply the function to the scaled attention score matrix to obtain an attention weight matrix. Multiply the attention weight matrix with the value matrix to obtain the attention output matrix, which is a header data, where K T is the transpose of the key matrix; Attention() is the execution process of the attention mechanism; Sogtmax() is the function processing process; is the scaling factor, and head is the head data. Multiple head data are concatenated along the feature dimension according to x=MultiHead(Q,K,V)=Concat(head1,head2,……,head k )W 0 , apply linear transformation to the concatenated output to obtain multi-head data, where x is the multi-head data, Concat() is the function processing process, and W 0 is a linear transformation, and MultiHead() is the execution process of the multi-head self-attention mechanism. Tra =FFN(x)=ReLU(xW1+b1)W2+b2 Further transform each position independently, where W1, W2, b1, b2 are learnable parameters, ReLU() is the function execution process, FFN() is the feedforward neural network processing process, H Tra Represent data in a context-sensitive manner.
[0139] In step 1512, the context-related representation data is sequentially fed into the forward long short-term memory network. For each word, the forward long short-term memory network considers the information of all the words preceding it and generates an encoding vector that contains the semantic and contextual information of the word and all the words preceding it. As the processing proceeds, the forward long short-term memory network gradually updates its internal state to reflect the information of the currently processed part of the text. After processing by the forward long short-term memory network, the corresponding second forward encoding vector is obtained, which captures the forward dependency in the text, that is, by encoding the contextual information from the first word to the i-th word, recorded as
[0140] In step 1513, the context-related representation data is sequentially fed into the backward long short-term memory network and processed in reverse order. The backward long short-term memory network considers the information of all the words behind it and generates an encoding vector that contains the semantic and contextual information of the word and all the words behind it. As the processing proceeds, the backward long short-term memory network gradually updates its internal state to reflect the information of the currently processed part of the text. After processing by the backward long short-term memory network, the corresponding second reverse encoding vector is obtained, which captures the backward dependency in the text, that is, by encoding the contextual information from the i-th word to the n-th word, recorded as
[0141] In step 1514, the first forward encoding vector and the first reverse encoding vector are concatenated. The concatenated feature vector is obtained, where i is the position index of the word segmentation; n is the number of tokens formed after the word segmentation of a single line of text; d b The unidirectional hidden dimension of the model, the default is 256, 2d b is the output dimension, that is, the vector is 2x256 dimensions; h i The function of the feature vector is to identify the object entity of the enhanced text feature vector, so the feature vector is recorded as H(obj) for subsequent understanding. This process can be summarized as Among them, H Tra represents the context-dependent representation data; H(obj) is the second enhanced vector after object entity feature enhancement; and BILSTM() represents the model processing. After the concatenation operation, the corresponding second enhanced vector is obtained. This vector contains both forward and backward dependency information, thus having stronger expressive power and effectively capturing the semantic and contextual information in the text, providing strong support for subsequent natural language processing tasks.
[0142] In an optional embodiment of the present invention, in step 152: predicting the guest entity start position and the guest entity end position based on the second enhanced vector to obtain a guest entity predicted boundary set includes:
[0143] Step 1521: predict the starting position of the object entity based on the preset relationship between the second enhanced vector and the main entity feature and the object entity feature to obtain a third predicted probability;
[0144] Step 1522: predict the end position of the guest entity based on the preset relationship between the second enhancement vector and the main entity feature and the guest entity feature to obtain a fourth prediction probability;
[0145] Step 1523: Obtain a third annotation value according to the third predicted probability and a third preset threshold;
[0146] Step 1524: Obtain a fourth marked value according to the fourth predicted probability and a fourth preset threshold;
[0147] Step 1525: Obtain a predicted boundary set of an object entity according to the third annotation value and the fourth annotation value;
[0148] Among them, during the customer entity prediction process, the customer entity weight is adjusted in real time through dynamic weighting.
[0149] In this embodiment, in step 1521, when predicting the starting position of the guest entity, the Get the third predicted probability;
[0150] In step 1522, when predicting the end location of the guest entity, the Get the fourth predicted probability; where, and is the weight parameter, and are the corresponding biases, H(obj) is the second enhancement vector, σ() represents the function processing process, represents the third predicted probability, represents the fourth predicted probability, r i Represents a preset relationship between the subject and object entities, and the ranges of the third prediction probability and the fourth prediction probability are both [0, 1].
[0151] In step 1523, the third predicted probability is compared with a third preset threshold to obtain a third labeling value. The third preset threshold is preferably 0.5. When the third predicted probability is equal to or exceeds the third preset threshold, the probability of this position is labeled as 1 using a binary labeling method; when the third predicted probability is less than the third preset threshold, the probability of this position is labeled as 0 using a binary labeling method.
[0152] In step 1524, the fourth predicted probability is compared with the fourth preset threshold to obtain a fourth marking value, where the fourth preset threshold is preferably 0.5. When the fourth predicted probability is equal to or exceeds the fourth preset threshold, the probability of this position is marked as 1 using a binary marking method; when the fourth predicted probability is less than the fourth preset threshold, the probability of this position is marked as 0 using a binary marking method.
[0153] It should be noted that the third preset threshold and the fourth preset threshold may be the same or different, and the third preset threshold and the fourth preset threshold are preferably both 0.5;
[0154] The first preset threshold, the second preset threshold, the third preset threshold and the fourth preset threshold may be the same or different, and are preferably all 0.5.
[0155] In step 1525 , after steps 1523 and 1524 , the location of the guest entity can be determined according to the third annotation value and the fourth annotation value, that is, the guest entity prediction boundary set is obtained.
[0156] During the training process, due to the difference in difficulty between the main entity prediction and the guest entity prediction, static weights will cause the model to converge slowly on more difficult tasks and converge prematurely on simple tasks in multi-task learning, thus affecting the overall performance. Since the main entity prediction is easier than the guest entity prediction, premature convergence will result in a high loss in the guest entity prediction, which affects the guest entity prediction loss. Therefore, a dynamic weighting method is used to adjust the loss weight in real time, so that the model can adaptively focus on different tasks, especially the guest entity prediction module.
[0157] Specific loss weighting process:
[0158] Measure the difference between the probability distribution predicted by the model and the true label according to L = -[y·log(p)+(1-y)·log(1-p)];
[0159] Among them, y is the true label, p is the predicted probability, and L is the specific loss calculation process.
[0160] according to Get the overall loss after dynamic weighting;
[0161] Among them, L s is the loss of the main entity; L o is the guest entity loss; ω is a fixed hyperparameter used to adjust the weight of the auxiliary task loss; L total is the overall loss after dynamic weighting; γ is an adjustable hyperparameter that controls the overall proportion of dynamic weights; ε is a small constant to prevent division by zero, usually e -9 , to prevent division by 0; It is a dynamic weight, which is adjusted in real time according to the loss values of the main task and auxiliary task. When the loss of the main entity prediction task is large, the weight of the guest entity prediction task loss will decrease; conversely, when the loss of the main entity prediction task is small, the weight of the guest entity prediction task loss will increase, thereby balancing different tasks.
[0162] In an optional embodiment of the present invention, step 16 may include:
[0163] Step 161: Filter and convert the information of the main entity prediction boundary set and the guest entity prediction boundary set to obtain first prediction data and second prediction data respectively;
[0164] Step 162: Output the first prediction data, the preset relationship, and the second prediction data as a triple.
[0165] In this embodiment, the text at the location marked with the value 1 in the predicted boundary set of the primary entity and the predicted boundary set of the secondary entity is extracted to obtain primary entity information and secondary entity information, namely the first predicted data and the second predicted data, respectively. The first predicted data, the preset relationship, and the second predicted data are output as a triple of (primary entity, relationship, secondary entity).
[0166] A specific embodiment of the method for processing entity relationships in text information provided by an embodiment of the present invention is as follows:
[0167] Step 1: Obtain the target text data to be processed, for example, "Smartphones have high-performance processors and large-capacity batteries, suitable for long-term use. Tablets have lightweight designs and high-definition screens, making them easy to carry."
[0168] Step 2: According to L MLM =-∑ i∈masked logP(w i |w \i ) Randomly mask some words in the text and let the model predict the word, and infer the masked word through context information, so as to learn bidirectional context information; where i is the position index of the masked token, masked is the set of all masked positions, P(w i |w \i ) indicates that the context w \i Predict the mask word w i . The coherence between different sentences is then trained so that, given two sentences, it can determine whether they are adjacent or related. This helps the model correctly identify whether contextual information is coherent, thereby improving the model's understanding of paragraphs. Furthermore, to directly model long-distance dependencies and enable the model to effectively capture the semantic or grammatical dependencies between distant words or sentence components in a sentence or paragraph, a self-attention mechanism is used. This allows each word in a sentence to calculate its association weight with all other words in the sentence during encoding.
[0169] After the model training is completed, according to Convert the text into a target text vector, where S is the text; n is the number of tokens formed after a single line of text is segmented; d is the hidden dimension of the model, which defaults to 768, indicating that a line of Chinese text S is segmented into n tokens, and each token is extended into a vector of 1xd dimensions, that is, a vector of 1x768 dimensions; H (B) is the target text vector; BERT() is the model processing process.
[0170] Step 3: After the forward long short-term memory network is processed, the corresponding first forward encoding vector is obtained to capture the forward dependency in the text, that is, by encoding the context information from the first word to the i-th word, which is recorded as After the processing of the backward long short-term memory network, the corresponding first reverse encoding vector is obtained to capture the backward dependency in the text, that is, by encoding the context information from the i-th word to the n-th word, recorded as The first forward coding vector and the first reverse coding vector are concatenated according to The concatenated feature vector is obtained, where i is the position index of the word segmentation; n is the number of tokens formed after the word segmentation of a single line of text; d b The unidirectional hidden dimension of the model, the default is 256, 2d b is the output dimension, that is, the vector is 2x256 dimensions; h i The function of the feature vector is to identify the main entity of the enhanced text feature vector, so the feature vector is recorded as H(sub) for subsequent understanding. This process can be summarized as Among them, H (B) is the target text vector; H(sub) is the first enhanced vector after the main entity feature is enhanced; BILSTM() is the model processing process.
[0171] Predict the starting position of the main entity according to Get the first prediction probability; predict the end position of the main entity, according to Get the second predicted probability; where W sh and W st is the weight parameter, b sh and b st is the bias, H(sub) is the first enhancement vector, σ() represents the function processing process, represents the first predicted probability, Represents the second predicted probability. The first predicted probability is compared with a first preset threshold to obtain a first labeled value, which is 0 or 1. The second predicted probability is compared with a second preset threshold to obtain a second labeled value, which is 0 or 1. The first and second labeled values are used to determine the location of the main entity, thereby obtaining the predicted boundary set of the main entity.
[0172] Step 4, according to The target text vector and the main entity prediction boundary set are weighted average pooled to obtain a pooled vector; where n is the number of tokens formed after word segmentation of a single line of text; d is the hidden dimension of the model, which defaults to 768, and A sub is the predicted boundary set of the main entity, H (B) is the target text vector, h avg is the pooling vector, Avgpool() represents weighted average pooling. sub =σ(W r h avg +b r ) processes the pooled vector to obtain weighted main entity information; where W r is the weight, b r is the bias, σ() is the function processing process, h sub is weighted information. According to E=H (B) +h sub The weighted information is embedded into the main entity information to obtain the intermediate text vector; where H (B) is the target text vector, h sub is the weighted information, and E is the intermediate text vector.
[0173] Step 5: Predict the starting position of the customer entity according to Get the third prediction probability; predict the end position of the guest entity, according to Get the fourth predicted probability; where, and is the weight parameter, and are the corresponding biases, H(obj) is the second enhancement vector, σ() represents the function processing process, represents the third predicted probability, represents the fourth predicted probability, r i represents the preset relationship between the host and guest entities. The third predicted probability and the fourth predicted probability are both in the range of [0, 1]. The third predicted probability is compared with a third preset threshold to obtain a third labeled value, which is 0 or 1. The fourth predicted probability is compared with a fourth preset threshold to obtain a fourth labeled value, which is 0 or 1. The location of the guest entity is determined based on the third and fourth labeled values, thereby obtaining the guest entity prediction boundary set.
[0174] The dynamic weight assignment method is used to adjust the loss weight in real time, and the specific loss weighting process is: the difference between the probability distribution predicted by the model and the real label is measured according to L = - [y log (p) + (1-y) log (1-p)], wherein y is the real label, p is the predicted probability, and L is the specific loss calculation process. The overall loss after dynamic weighting is obtained, wherein L s is the main entity loss; L o is the guest entity loss; ω is a fixed hyperparameter for adjusting the weight of the auxiliary task loss; L total is the overall loss after dynamic weighting; γ is an adjustable hyperparameter, which controls the overall proportion of the dynamic weight; ε is a small constant to prevent division by zero, usually e -9 , to prevent division by zero; is the dynamic weight, which is adjusted in real time according to the loss values of the main task and the auxiliary task; when the main entity prediction task loss is large, the weight of the guest entity prediction task loss will decrease; on the contrary, when the main entity prediction task loss is small, the weight of the guest entity prediction task loss will increase, so as to balance different tasks.
[0175] Step 6, the text of the position of the labeled value 1 in the main entity prediction boundary set and the guest entity prediction boundary set is extracted, and the main entity information and the guest entity information are obtained, that is, the first prediction data and the second prediction data. The first prediction data, the preset relationship and the second prediction data are output in the form of a triple (main entity, relationship, guest entity), and the final output data is (smartphone, has, high-performance processor), (smartphone, has, large-capacity battery), (tablet computer, has, lightweight design), (tablet computer, has, high-definition screen).
[0176] The entity relationship processing method provided by the application can help optimize the prediction performance of the process of extracting triplets from massive Chinese data, and the scheme has strong expandability and can adapt to the characteristics of different texts, improve the accuracy of text entity relationship extraction, and greatly reduce the overlap probability of output triplets.
[0177] As shown in Figure 4 , the embodiment of the application further provides an entity relationship processing apparatus 60 in text information, comprising:
[0178] An acquisition module 61 is configured to acquire target text data to be processed;
[0179] The processing module 62 is used to perform bidirectional encoder representation encoding on the target text data to obtain a target text vector; perform main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data; integrate main entity information into the target text vector to obtain an intermediate text vector; perform guest entity feature boundary prediction on the intermediate text vector to obtain a guest entity prediction boundary set in the target text data; and output the main entity prediction boundary set, the guest entity prediction boundary set, and the preset relationship between the main entity feature and the guest entity feature in the form of a triple.
[0180] Optionally, performing bidirectional encoder representation encoding on the target text data to obtain a target text vector includes:
[0181] Perform word segmentation vector processing on the target text data to obtain multiple word segmentation vectors;
[0182] Encode multiple word segmentation vectors according to the dimensions of the preset text vector processing model to obtain the target text vector.
[0183] Optionally, performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data includes:
[0184] Performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector;
[0185] The main entity start position and the main entity end position are predicted according to the first enhancement vector to obtain a main entity prediction boundary set.
[0186] Optionally, performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector includes:
[0187] Performing forward encoding on the words of the target text vector according to the forward long short-term memory network to obtain a first forward encoding vector;
[0188] Reversely encode the words of the target text vector according to the backward long short-term memory network to obtain a first reverse encoding vector;
[0189] The first forward coding vector and the first reverse coding vector are concatenated to obtain a first enhanced vector.
[0190] Optionally, predicting a main entity start position and a main entity end position according to the first enhancement vector to obtain a main entity prediction boundary set includes:
[0191] Predicting a starting position of the main entity according to the first enhancement vector to obtain a first prediction probability;
[0192] predicting, according to the first enhanced vector, a main entity end position to obtain a second prediction probability;
[0193] obtaining a first label value according to the first prediction probability and a first preset threshold;
[0194] obtaining a second label value according to the second prediction probability and a second preset threshold;
[0195] obtaining a main entity prediction boundary set according to the first label value and the second label value.
[0196] Optionally, the target text vector is subjected to main entity information injection to obtain an intermediate text vector, including:
[0197] performing weighted average pooling processing on the target text vector and the main entity prediction boundary set to obtain weighted information;
[0198] embedding the weighted information into the target text vector to obtain the intermediate text vector.
[0199] Optionally, the intermediate text vector is subjected to guest entity feature boundary prediction to obtain a guest entity prediction boundary set in the target text data, including:
[0200] performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector;
[0201] predicting a guest entity start position and a guest entity end position according to the second enhanced vector to obtain the guest entity prediction boundary set.
[0202] Optionally, the intermediate text vector is subjected to feature vector enhancement processing through a bidirectional long short-term memory network to obtain a second enhanced vector, including:
[0203] performing converter encoding on the intermediate text vector to generate context-related representation data;
[0204] forwardly encoding words of the context-related representation data according to a forward long short-term memory network to obtain a second forward encoding vector;
[0205] reversely encoding words of the context-related representation data according to a reverse long short-term memory network to obtain a second reverse encoding vector;
[0206] splicing the second forward encoding vector and the second reverse encoding vector to obtain the second enhanced vector.
[0207] Optionally, predicting a guest entity start position and a guest entity end position according to the second enhanced vector to obtain a guest entity prediction boundary set includes:
[0208] Predicting a starting position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a third prediction probability;
[0209] Predicting an end position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a fourth prediction probability;
[0210] Obtaining a third annotation value according to the third predicted probability and a third preset threshold;
[0211] Obtaining a fourth marked value according to the fourth predicted probability and a fourth preset threshold;
[0212] Obtaining a predicted boundary set of an object entity according to the third annotation value and the fourth annotation value;
[0213] Among them, during the customer entity prediction process, the customer entity weight is adjusted in real time through dynamic weighting.
[0214] It should be noted that this device is a device corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0215] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0216] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0217] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0218] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0219] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0220] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, ROM, RAM, a magnetic disk, or an optical disk.
[0221] In addition, it should be noted that, in the apparatus and method of the present invention, it is obvious that each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. Moreover, the steps of performing the above-mentioned series of processing can naturally be performed in chronological order according to the order of description, but it is not necessary to perform them in chronological order, and some steps can be performed in parallel or independently of each other. For those of ordinary skill in the art, it will be understood that all or any steps or components of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in hardware, firmware, software or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.
[0222] Therefore, the purpose of the present invention can also be achieved by running a program or a group of programs on any computing device. The computing device can be a well-known general-purpose device. Therefore, the purpose of the present invention can also be achieved simply by providing a program product containing program code for implementing the method or device. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be pointed out that in the device and method of the present invention, it is obvious that each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. In addition, the steps of performing the above-mentioned series of processing can naturally be performed in chronological order according to the order of description, but do not necessarily need to be performed in chronological order. Certain steps can be performed in parallel or independently of each other.
[0223] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for processing entity relationships in text information, characterized in that: include: Obtain target text data to be processed; Performing bidirectional encoder representation encoding on the target text data to obtain a target text vector; Performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data; Injecting main entity information into the target text vector to obtain an intermediate text vector; Performing object entity feature boundary prediction on the intermediate text vector to obtain an object entity prediction boundary set in the target text data; The main entity prediction boundary set, the object entity prediction boundary set, and the preset relationship between the main entity features and the object entity features are output in the form of triples.
2. The method for processing entity relationships in text information according to claim 1, characterized in that: The target text data is encoded by a bidirectional encoder representation to obtain a target text vector, including: Perform word segmentation vector processing on the target text data to obtain multiple word segmentation vectors; Encode multiple word segmentation vectors according to the dimensions of the preset text vector processing model to obtain the target text vector.
3. The method for processing entity relationships in text information according to claim 1, characterized in that: Performing main entity feature boundary prediction on the target text vector to obtain a main entity prediction boundary set in the target text data includes: Performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector; The main entity start position and the main entity end position are predicted according to the first enhancement vector to obtain a main entity prediction boundary set.
4. The method for processing entity relationships in text information according to claim 1, characterized in that: Performing feature vector enhancement processing on the target text vector through a bidirectional long short-term memory network to obtain a first enhanced vector, including: Performing forward encoding on the words of the target text vector according to the forward long short-term memory network to obtain a first forward encoding vector; Reversely encode the words of the target text vector according to the backward long short-term memory network to obtain a first reverse encoding vector; The first forward coding vector and the first reverse coding vector are concatenated to obtain a first enhanced vector.
5. The method for processing entity relationships in text information according to claim 4, characterized in that: Predicting a main entity start position and a main entity end position according to the first enhancement vector to obtain a main entity prediction boundary set includes: Predicting a starting position of the main entity according to the first enhancement vector to obtain a first prediction probability; Predicting the end position of the main entity according to the first enhancement vector to obtain a second prediction probability; Obtaining a first labeled value according to the first predicted probability and a first preset threshold; Obtaining a second marked value according to the second predicted probability and a second preset threshold; A main entity prediction boundary set is obtained according to the first label value and the second label value.
6. The method for processing entity relationships in text information according to claim 1, characterized in that: Injecting the main entity information into the target text vector to obtain an intermediate text vector includes: Performing weighted average pooling processing on the target text vector and the main entity prediction boundary set to obtain weighted information; The weighted information is embedded into the target text vector to obtain an intermediate text vector.
7. The method for processing entity relationships in text information according to claim 1, characterized in that: Performing object entity feature boundary prediction on the intermediate text vector to obtain an object entity prediction boundary set in the target text data includes: Performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector; The guest entity start position and the guest entity end position are predicted according to the second enhanced vector to obtain a guest entity prediction boundary set.
8. The method for processing entity relationships in text information according to claim 1, characterized in that: Performing feature vector enhancement processing on the intermediate text vector through a bidirectional long short-term memory network to obtain a second enhanced vector, including: Performing transformer encoding on the intermediate text vector to generate context-dependent representation data; Performing forward encoding on the words of the context-related representation data according to a forward long short-term memory network to obtain a second forward encoding vector; Reverse encoding the words of the context-related representation data according to a backward long short-term memory network to obtain a second reverse encoding vector; The second forward coding vector and the second reverse coding vector are concatenated to obtain a second enhanced vector.
9. The method for processing entity relationships in text information according to claim 8, characterized in that: Predicting the guest entity start position and the guest entity end position based on the second enhanced vector to obtain a guest entity prediction boundary set includes: Predicting a starting position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a third prediction probability; Predicting an end position of the object entity based on a preset relationship between the second enhancement vector and the main entity feature and the object entity feature to obtain a fourth prediction probability; Obtaining a third annotation value according to the third predicted probability and a third preset threshold; Obtaining a fourth marked value according to the fourth predicted probability and a fourth preset threshold; Obtaining a predicted boundary set of an object entity according to the third annotation value and the fourth annotation value; Among them, during the customer entity prediction process, the customer entity weight is adjusted in real time through dynamic weighting.
10. A device for processing entity relationships in text information, characterized in that: include: An acquisition module, used to acquire target text data to be processed; A processing module, configured to perform bidirectional encoder representation encoding on the target text data to obtain a target text vector; The target text vector is subjected to main entity feature boundary prediction to obtain a main entity prediction boundary set in the target text data; the target text vector is integrated with main entity information to obtain an intermediate text vector; the intermediate text vector is subjected to guest entity feature boundary prediction to obtain a guest entity prediction boundary set in the target text data; the main entity prediction boundary set, the guest entity prediction boundary set, and the preset relationship between the main entity feature and the guest entity feature are output in the form of a triple.