Content entity and relation extraction method based on self-attention model

By establishing a self-attention model at the last layer of the encoder, calculating the eigenvalue correlation degree and assigning attention weights, the extraction of long-distance dependence and diversified entity relationships in complex texts is solved, and entity relationship extraction with high accuracy and generalization capabilities is achieved.

CN120386869APending Publication Date: 2025-07-29珠海必优科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219942.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture long-distance dependencies and diversified entity relationships in complex texts, resulting in incomplete or errors in entity relationship extraction, which affects the accuracy and generalization capabilities of the model, especially in scenarios with multiple entities and multiple relationship types.

Method used

Using a method based on the self-attention model, a self-attention model is established at the last layer of the encoder, the correlation degree of eigenvalues of different positions is calculated, and the correlation degree is quantized by JS divergence. After the information is aggregated, the linear layer is input to the feedforward calculation to judge the termination word closest to the target starting word, and adapt to multi-entity or multi-relational type tasks.

Benefits of technology

It improves the accuracy and generalization ability of entity relationship extraction, can effectively capture long-distance dependencies in text, and is suitable for complex multi-entity multi-relationship scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386869A_ABST
    Figure CN120386869A_ABST
Patent Text Reader

Abstract

The invention provides a content entity and relation extraction method based on a self-attention model, and the method comprises the steps: S101, building a self-attention model at the last layer of an encoder, and inputting parameters outputted by the encoder into the self-attention model for calculation; s102, in the self-attention model, carrying out correlation degree calculation on the characteristic value of each position and the characteristic values of other positions; s103, calculating and adding the attention input at other positions obtained in the step S102 and the information amount carried at the position, and taking the sum as the output of the target position; and S104, a linear layer is arranged on the rear side of the self-attention model, and parameters output by the self-attention model are input into the linear model for calculation. The method is applied to the fields of computer technology, artificial intelligence and image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for extracting content entities and relationships based on the self-attention model. Background Art

[0002] Entity relationship extraction plays an important role in natural language processing. However, traditional methods often struggle to accurately capture long-distance dependencies in complex texts. Especially in scenarios with multiple entities and multiple relationship types, how to effectively identify and extract the semantic associations between entities has become a major challenge. Existing technologies are prone to information loss when processing long texts, resulting in incomplete or incorrect relationship extraction. At the same time, different types of entities and relationships often have different semantic features, and how to model these differences specifically is also a difficult problem. In addition, the expression methods of entity relationships are diverse and may span multiple sentences or even paragraphs, increasing the complexity of extraction. While ensuring the generalization ability of the model, how to improve the extraction accuracy of texts in specific domains is also an urgent problem to be solved. These challenges not only affect the accuracy of entity relationship extraction but also limit its application effects in downstream tasks such as complex text analysis and knowledge graph construction. Therefore, developing an extraction method that can effectively handle long-distance dependencies, adapt to diverse entity relationships, and maintain high accuracy in complex scenarios has become the focus of current research. Summary of the Invention

[0003] The present invention aims to provide a method for extracting content entities and relationships based on the self-attention model.

[0004] The method of the present invention includes the following steps:

[0005] S101. Establish a self-attention model at the last layer of the encoder, and input the parameters output by the encoder into the self-attention model for calculation;

[0006] S102. In the self-attention model, the feature value at each position will be calculated for the correlation degree with the feature values at other positions to obtain the correlation value, and this correlation value determines how much attention the target position needs to invest in each other position;

[0007] S103. Calculate and sum the attention invested by other positions obtained in step S102 with the amount of information carried by this position as the output of the target position;

[0008] S104. Set a linear layer at the rear of the self-attention model, and input the parameters output by the self-attention into the linear model for calculation;

[0009] S105. Use a linear layer to perform a feed-forward calculation on the relationships between the eigenvalue of each position included in the self-attention model to make a judgment for the next step, and finally obtain the termination word with the closest relationship to the target start word, thereby establishing the relationship from the start word to the termination word and realizing the extraction of entity words and entity relationships.

[0010] Furthermore, for tasks with multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type, perform single-group entity and relationship type discrimination on each set of linear layer parameters, and perform a linear layer feed-forward calculation. Finally, obtain the termination word with the closest relationship to the target start word, and thereby establish the relationship from the start word to the termination word.

[0011] Furthermore, in step S102, the correlation degree between the eigenvalue of each position and the eigenvalues of other positions is calculated. The correlation degree of the eigenvalues is calculated by the JS divergence, and its calculation formula is:

[0012]

[0013] where P and Q represent two different distributions, and use d i,j to represent the correlation degree between the eigenvalue a i at the i-th position and the eigenvalue a j at the j-th position, then there is:

[0014]

[0015] Furthermore, the calculation formula for calculating and summing the attention input from other positions obtained in step S102 and the amount of information carried by this position in step S103 is:

[0016]

[0017] where n represents the sequence length, and 1 ≤ i, j ≤ n.

[0018] Furthermore, for tasks with multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type, perform single-group entity and relationship type discrimination on each set of linear layer parameters, and perform a linear layer feed-forward calculation. Finally, obtain the termination word with the closest relationship to the target start word. This process is calculated by the following formula:

[0019] t k = L k (D),

[0020] where k ∈ (1, m), m is the number of entity types, configure a set of linear layer parameters for each relationship type, L kThe linear layer parameters representing the k-th type of relationship type, where D is the sum of attention information, and through t k to determine whether there is a specific class relationship for the entity word at the corresponding position.

[0021] In addition, the method further includes the following steps:

[0022] For tasks of multiple entities or multiple relationship types, the number of target encoding values is increased exponentially.

[0023] The method further includes the following steps:

[0024] For tasks of multiple entities or multiple relationship types, an independent model is equipped for each type of entity and type, and multiple entities and multiple relationships are identified through the independent model.

[0025] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:

[0026] The present invention discloses an entity relationship extraction method based on the self-attention mechanism. This method establishes a self-attention model in the last layer of the encoder, calculates the correlation degree between feature values at different positions, and quantifies this correlation using the JS divergence. Then, according to the correlation degree, attention weights are assigned to aggregate information to obtain the output at the target position. Next, the self-attention output is passed into a linear layer for feed-forward calculation to determine the termination word closest to the target start word, thereby establishing the relationship between entities. For tasks of multiple entities or multiple relationship types, the present invention configures independent linear layer parameters for each type and performs discrimination and calculation separately. This method can effectively capture long-distance dependencies in the text, improve the accuracy and generalization ability of entity relationship extraction, and is especially suitable for complex multi-entity multi-relationship scenarios. Description of the Drawings

[0027] Figure 1 is a flowchart of the method of the present invention;

[0028] Figure 2 is a flow block diagram of the present invention. Detailed Embodiments

[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0030] As Figure 1 and Figure 2 shown, a method for content entity and relationship extraction based on a self-attention model in this embodiment may specifically include:

[0031] Step S101: Obtain the parameters output by the encoder. The parameters are input into the self-attention model established in the last layer of the encoder for calculation.

[0032] Specifically, a self-attention model is established in the last layer of the encoder, and the model parameters are set. Obtain the output data of the encoder, and extract the feature vectors. Input the feature vectors into the self-attention model to calculate the attention weights. According to the attention weights, perform weighted summation on the feature vectors to obtain the attention features. Normalize the attention features to obtain the normalized features. Input the normalized features into the linear layer for linear transformation. Perform activation function processing on the result of the linear transformation to obtain the activation features. Perform dimensionality reduction processing on the activation features to obtain the dimensionality-reduced features. Use the dimensionality-reduced features as the final output to complete the model calculation.

[0033] For example, a self-attention model is established in the last layer of the encoder, and the model parameters are set, including the number of heads of the multi-head attention mechanism being 8 and the hidden layer dimension being 512. Obtain the output data of the encoder, extract the feature vectors, and use matrix multiplication to map the input sequence to a feature vector with a dimension of 512. Input the feature vectors into the self-attention model to calculate the attention weights, using the scaled dot-product attention formula, where the dimensions of the query vector, key vector, and value vector are all 64. According to the attention weights, perform weighted summation on the feature vectors to obtain the attention features, and use the softmax function to normalize the weights to ensure that the sum of the weights is 1. Normalize the attention features using layer normalization to make the mean of the features 0 and the variance 1, obtaining the normalized features. Input the normalized features into the linear layer for linear transformation, setting the dimension of the weight matrix of the linear layer to 512×512 and the dimension of the bias term to 512. Perform activation function processing on the result of the linear transformation using the ReLU activation function to obtain the activation features, ensuring the non-linear expression ability of the features. Perform dimensionality reduction processing on the activation features, using a fully connected layer to reduce the dimension from 512 to 256, obtaining the dimensionality-reduced features. Use the dimensionality-reduced features as the final output to complete the model calculation, and output a feature vector with a dimension of 256 for subsequent tasks.

[0034] Step S102: For the eigenvalue at each position in the self-attention model, obtain its correlation degree with the eigenvalues at other positions. The correlation degree of the eigenvalues determines how much attention the target position needs to invest in each other position. The correlation degree of the eigenvalues is calculated by the JS divergence, and its calculation formula is:

[0035]

[0036] where P and Q represent two different distributions, and Use d i,j to represent the eigenvalue a at position i i and the eigenvalue a at position j j The correlation degree between them is as follows:

[0037]

[0038] Specifically: Obtain the eigenvalue at each position in the input sequence to form a feature matrix. Perform a linear transformation on the feature matrix to obtain a query matrix, a key matrix, and a value matrix respectively. According to the query matrix and the key matrix, calculate the correlation degree score between each position and other positions. Normalize the correlation degree score to obtain an attention weight matrix. Perform weighted summation according to the attention weight matrix and the value matrix to obtain the context representation of each position. Determine whether the context representation is the final output. If so, end; otherwise, proceed to the next step. Concatenate the context representation with the original feature matrix to form a new feature matrix. Perform a non-linear transformation on the new feature matrix to obtain the updated eigenvalue. Repeat the above eight steps until the preset number of iterations or convergence conditions are met.

[0039] For example, the length of the input sequence is 10, the feature dimension of each position is 512, and the dimension of the feature matrix is 10×512. Perform a linear transformation on the feature matrix to obtain a query matrix, a key matrix, and a value matrix respectively. The dimension of the weight matrix for the linear transformation is 512×64, and the dimensions of the query matrix, the key matrix, and the value matrix are all 10×64. According to the query matrix and the key matrix, calculate the correlation degree score between each position and other positions, using the dot product calculation method. For example, the correlation degree score between position 1 and position 2 is the dot product result of the first row of the query matrix and the second row of the key matrix. Normalize the correlation degree score to obtain an attention weight matrix, using the softmax function to normalize the correlation degree score so that the sum of the weights in each row is 1. Perform weighted summation according to the attention weight matrix and the value matrix to obtain the context representation of each position. For example, the context representation of position 1 is the weighted sum of the first row of the attention weight matrix and the value matrix. Determine whether the context representation is the final output. If the dimension of the context representation is consistent with the preset output dimension, output the result; otherwise, proceed to the next step. Concatenate the context representation with the original feature matrix to form a new feature matrix. For example, the dimension of the context representation is 10×64, and the dimension after concatenation with the original feature matrix is 10×576. Perform a non-linear transformation on the new feature matrix to obtain the updated eigenvalue, using the ReLU activation function to process the transformed eigenvalue. Repeat the above processes of linear transformation, correlation degree calculation, normalization, weighted summation, concatenation, and non-linear transformation until the preset number of iterations or convergence conditions are met, such as the number of iterations reaching 5 times or the change in the context representation being less than the threshold 0.001.

[0040] Step S103: Calculate and sum up the attention invested in other positions obtained from the above calculations and the information carried by this position to obtain the output of the target position. Here, the calculation formula for calculating and summing up the attention invested in other positions and the information carried by this position is as follows:

[0041]

[0042] where n represents the sequence length, and 1 ≤ i, j ≤ n.

[0043] Specifically: Obtain the attention values invested in other positions in Step S102 and extract the attention distribution data. Obtain the information carried by this position and extract the information quantity feature data. Normalize the attention distribution data to obtain the normalized attention values. Normalize the information quantity feature data to obtain the normalized information quantity values. Use a weighted algorithm to assign weights to the normalized attention values and the normalized information quantity values. Calculate the product of the attention value and the information quantity value according to the weight assignment result. Perform an accumulation operation on the product result to obtain the preliminary output value. If the preliminary output value exceeds the preset threshold, smooth the preliminary output value. Determine the final output value of the target position according to the smoothing result.

[0044] For example, given an input attention score of [3.0, 1.0, 0.2], an attention distribution of [0.84, 0.11, 0.05] is obtained. The information content carried at this position is acquired, and information quantity feature data is extracted. The information entropy is used to calculate the information content. For example, given an input probability distribution of [0.5, 0.3, 0.2], an information entropy value of 1.029 is obtained. The attention distribution data is normalized to obtain a normalized attention value. Using the min-max normalization method, the data is mapped to the interval [0, 1]. For example, given an input of [0.84, 0.11, 0.05], [1.0, 0.07, 0.0] is obtained. The information quantity feature data is normalized to obtain a normalized information quantity value. Using the z-score method, the data is converted into a distribution with a mean of 0 and a standard deviation of 1. For example, given an input of [1.029, 0.8, 1.2], [0.0, -1.14, 0.85] is obtained. A weighted algorithm is used to assign weights to the normalized attention value and the normalized information quantity value. For example, linear weighting is used, and the weight coefficients are set to 0.6 and 0.4. According to the weight assignment result, the product of the attention value and the information quantity value is calculated. For example, given a normalized attention value of [1.0, 0.07, 0.0] and a normalized information quantity value of [0.0, -1.14, 0.85], the product result is [0.0, -0.08, 0.0]. An accumulation operation is performed on the product result to obtain a preliminary output value. For example, given an input of [0.0, -0.08, 0.0], the accumulation result is -0.08. If the preliminary output value exceeds a preset threshold, the preliminary output value is smoothed. For example, the threshold is set to 0.1, and the preliminary output value of -0.08 does not exceed the threshold, so no smoothing is required. According to the smoothing result, the final output value of the target position is determined. For example, given a preliminary output value of -0.08, it is directly used as the final output value.

[0045] Step S104: Obtain the parameters output by self-attention, and input the parameters into a linear layer arranged at the rear side of the self-attention model for calculation.

[0046] Specifically: Obtain the parameters output by the self-attention model, where the parameters include feature representations. Align the feature representations with the weight matrix of the linear layer to determine the input dimension match. If the input dimension is consistent with the dimension of the linear layer weight matrix, perform matrix multiplication. If the input dimension is inconsistent with the dimension of the linear layer weight matrix, adjust the dimension of the feature representation to match the weight matrix. Through matrix multiplication, multiply the feature representation by the weight matrix to obtain an intermediate calculation result. Add the intermediate calculation result to the bias vector of the linear layer to obtain the feature representation after linear transformation. Input the feature representation after linear transformation into an activation function for non-linear transformation. Determine the final feature representation based on the output result of the activation function. Use the final feature representation as the input data for subsequent models or tasks.

[0047] For example, assume the dimension of the feature vector is 512. Align the feature representations with the weight matrix of the linear layer to determine the input dimension match. If the dimension of the weight matrix is 512×256, then the input dimension is consistent with the dimension of the weight matrix. If the input dimension is consistent with the dimension of the linear layer weight matrix, perform matrix multiplication, multiply the feature representation by the weight matrix to obtain an intermediate calculation result, for example, the dimension of the calculation result is 256. If the input dimension is inconsistent with the dimension of the linear layer weight matrix, adjust the dimension of the feature representation to match the weight matrix, for example, adjust the dimension of the feature vector from 512 to 256 through a fully connected layer. Through matrix multiplication, multiply the feature representation by the weight matrix to obtain an intermediate calculation result, for example, the dimension of the calculation result is 256. Add the intermediate calculation result to the bias vector of the linear layer to obtain the feature representation after linear transformation, for example, the dimension of the bias vector is 256. Input the feature representation after linear transformation into an activation function for non-linear transformation, for example, process it using the ReLU activation function. Determine the final feature representation based on the output result of the activation function, for example, the output dimension is still 256. Use the final feature representation as the input data for subsequent models or tasks, for example, input it into a fully connected layer for classification tasks.

[0048] In step S105, use the linear layer to perform a feed-forward calculation on the relationships between the feature values at each position included in the self-attention model, and determine the end word with the closest relationship to the target start word based on the calculation result, thereby establishing the relationship from the start word to the end word, and realizing the extraction of entity words and entity relationships.

[0049] Specifically: establish a self-attention model at the last layer of the encoder, and input the feature values output by the encoder into the self-attention model for calculation. Use the JS divergence to calculate the correlation degree between the feature values at each position in the self-attention model and the feature values at other positions. The specific calculation formula can be the one in the above step S102, and finally obtain the correlation value. Add the correlation degree calculation result to the attention information input at other positions to generate the information volume carried by each position. Set a linear layer behind the self-attention model to perform a feed-forward calculation on the relationship between the feature values at each position. If the task involves multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type. Use the linear layer to perform a feed-forward calculation on the feature values to determine whether the feature value at each position has a specific class relationship with the target start word. According to the feed-forward calculation result of the linear layer, judge the end word that is closest to the target start word. Establish the relationship between the start word and the end word to determine the association between the entity word and the entity relationship. For each set of linear layer parameters, perform single-group entity and relationship type discrimination to generate the extraction results of entities and relationships.

[0050] For example, input a sequence of length 128, where the feature dimension at each position is 768. The Jensen-Shannon (JS) divergence is used to calculate the correlation between the feature values at each position in the self-attention model and the feature values at other positions. For example, when calculating the correlation between position i and position j, the JS divergence formula is used, where P and Q are the feature distributions at position i and position j respectively, and the correlation value is obtained through the formula. The calculation results of the correlation are added to the attention information input from other positions. For example, the attention weight at position i is multiplied by the feature value at position j and then accumulated to generate the amount of information carried by each position. A linear layer is set behind the self-attention model to perform a feed-forward calculation on the relationships between the feature values at each position. For example, a linear transformation matrix is used to map the 768-dimensional features to 256 dimensions. If the task involves multiple entity or multiple relationship types, a set of linear layer parameters is configured for each type of entity or relationship. For example, 5 sets of linear layer parameters are configured for 5 entity types respectively, and the dimension of each set of parameters is 256×256. The linear layer is used to perform a feed-forward calculation on the feature values to determine whether the feature value at each position has a specific type of relationship with the target start word. For example, by calculating the feature similarity between the target start word and the candidate end word, it is determined whether there is a relationship. According to the feed-forward calculation results of the linear layer, the end word with the closest relationship to the target start word is determined. For example, the position with the highest similarity is selected as the end word. The relationship between the start word and the end word is established to determine the association between the entity word and the entity relationship. For example, the start word "Apple" is associated with the end word "Company" as a "brand-company" relationship. For each set of linear layer parameters, single-set entity and relationship type discrimination is performed to generate the extraction results of entities and relationships. For example, the feature values are classified through the linear layer to obtain the final discrimination results of entity types and relationship types.

[0051] Step S106, if there is a task involving multiple entities or multiple relationship types, then a set of linear layer parameters is configured for each type of entity or relationship. For each set of linear layer parameters, single-set entity and relationship type discrimination is performed, and a feed-forward calculation of the linear layer is carried out. Finally, the end word with the closest relationship to the target start word is obtained, and then the relationship from the start word to the end word is established. For a task involving multiple entities or multiple relationship types, a set of linear layer parameters is configured for each type of entity or relationship. For each set of linear layer parameters, single-set entity and relationship type discrimination is performed, and a feed-forward calculation of the linear layer is carried out. Finally, the end word with the closest relationship to the target start word is obtained, and this process is calculated through the following formula:

[0052] t k =L k (D),

[0053] where k ∈ (1, m), m is the number of entity types, a set of linear layer parameters is configured for each type of relationship, and L kThe linear layer parameters representing the k-th type of relationship, D is the sum of attention information, and t is used k to determine whether the entity word at the corresponding position has a specific class relationship.

[0054] For tasks with multiple entities or multiple relationship types, the number of target encoding values is increased exponentially. For tasks with multiple entities or multiple relationship types, an independent model is equipped for each type of entity and type, and multiple entities and multiple relationships are identified through the independent model.

[0055] Specifically: A self-attention model is established in the last layer of the encoder, and the parameters output by the encoder are input into the self-attention model for calculation. The self-attention model calculates the feature values at each position in the sequence to obtain the feature information at each position, which serves as the basic data for calculating the correlation degree. According to the feature values at each position in the sequence, the correlation degree between the feature values at other positions is calculated. The JS divergence is used to calculate the correlation degree between the feature value at position i and the feature value at position j, where P and Q represent the feature distributions at position i and position j respectively. According to the JS divergence formula, the correlation degree value between the feature value at position i and the feature value at position j is calculated. The calculated correlation degree value is used as the attention weight to represent the information correlation strength between position i and position j. The attention weights invested by other positions are calculated with the amount of information carried by this position and an addition operation is performed. A linear layer is used to perform a feed-forward calculation on the relationship between the feature values at each position in the self-attention model. The terminating word closest to the target starting word is judged, and the relationship from the starting word to the terminating word is established. For tasks with multiple entities or multiple relationship types, the linear layer parameters of each type of relationship are configured, and single-group entity and relationship type discrimination is performed to complete the extraction of entity words and entity relationships.

[0056] For example, for a sequence of length 10, the feature vectors at each position are extracted respectively, with a dimension of 512. According to the feature values at each position in the sequence, the correlation degree between the feature values at other positions is calculated. When calculating the correlation degree, the JS divergence formula is used to calculate the correlation degree value. For example, the JS divergence of the feature values P and Q at position i and position j is calculated, and the formula is

[0057] JSD(P||Q) = 0.5 * (KL(P||M) + KL(Q||M)),

[0058] where P and Q represent two different distributions, and M = 0.5 * (P + Q).

[0059] Use d i,j to represent the correlation degree between the feature value a at position i i and the feature value a at position j j Then there is:

[0060]

[0061] Take the calculated correlation degree value as the attention weight, which is used to represent the information correlation strength between position i and position j.

[0062] Sum up the calculated correlation degree value and the amount of information carried by each position to obtain the attention information sum D. For example, the attention information sum for position i is

[0063]

[0064] where n represents the sequence length, and 1 ≤ i, j ≤ n.

[0065] If there are tasks with multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type. For example, if the number of entity types is 3 and the number of relationship types is 2, then configure 3 sets of entity linear layer parameters and 2 sets of relationship linear layer parameters respectively. Perform single-group entity and relationship type discrimination on each set of linear layer parameters. For example, for the linear layer parameters W_k of the k-th relationship type, calculate their product with the eigenvalue W_k * D to determine the entity or relationship type corresponding to each set of parameters. Use the linear layer to perform feed-forward calculation on the relationship between the eigenvalues of each position included in the self-attention model. For example, perform feed-forward calculation on the eigenvalue of position i, F_i = W * D_i + b, to obtain the calculation result of the relationship between the eigenvalues. According to the feed-forward calculation result, determine the end word that is closest to the target start word. For example, if the start word is at position 3, calculate its feed-forward calculation results F_3j with all positions, and select the position j with the largest F_3j as the end word to establish the association relationship between the start word and the end word. Complete the establishment of the relationship between the start word and the end word to achieve accurate extraction of entity words and entity relationships.

[0066] Let the eigenvalues of each position in the sequence of length n be x1, x2, … As the basic data for correlation degree calculation. Use JS divergence to calculate the correlation degree between the eigenvalue of position i and the eigenvalue of position j, where P and Q represent the feature distributions of position i and position j respectively. According to the JS divergence formula, calculate the correlation degree value between the eigenvalue of position i and the eigenvalue of position j. For example, when P = [0.3, 0.7] and Q = [0.4, 0.6], the JS divergence value is 0.056. Take the calculated correlation degree value as the attention weight, which is used to represent the information correlation strength between position i and position j. For example, the attention weight between position i and position j is 0.8. Calculate the attention weight invested by other positions and the amount of information carried by this position, and perform the summation operation. For example, for position i, its weighted information amount is (Attention weight × Information amount). Use a linear layer to perform a feed-forward calculation on the relationship between the feature values at each position in the self-attention model. For example, through a linear transformation Wx + b, where W is the weight matrix and b is the bias term. Determine the end word that is closest to the target start word. For example, by calculating the cosine similarity between the start word and the end word, and select the end word with the highest similarity. For tasks with multiple entities or multiple relationship types, configure the linear layer parameters for each type of relationship. For example, configure the parameters for relationship type k as and Perform single-group entity and relationship type discrimination to complete the extraction of entity words and entity relationships.

[0067] Another example is that the input sequence length is n = 10 text data. Obtain the feature value at each position in the sequence. The feature value includes the information amount carried by this position. For example, the feature value at position i is [0.1, 0.3, 0.5]. Calculate the correlation degree between the feature value at each position in the sequence and the feature values at other positions. Use the JS divergence as the correlation degree calculation method. For example, calculate the JS divergence between position i and position j, where P is the feature distribution at position i and Q is the feature distribution at position j, and get JS(P||Q) = 0.2. According to the JS divergence calculation formula, the correlation degree between the feature value at position i and the feature value at position j is 0.2. Use the correlation degree as the attention value input from other positions and perform a weighted calculation with the information amount carried by the current position. For example, the attention value at position i is 0.8 and the information amount is 0.5, and the weighted result is 0.8 * 0.5 = 0.4. Use the sequence length n = 10 to perform a summation calculation on the weighted attention value and the information amount. For example, add up the weighted results at all positions to get a total of 3.6. If there are multiple entities or multiple relationship types in the sequence, configure a set of linear layer parameters for each type of relationship. For example, configure 3 sets of linear layer parameters, and each set of parameters is [0.1, 0.2, 0.3]. Use a linear layer to perform a feed-forward calculation on the summed attention information and feature values to determine whether there is a specific type of relationship for the entity word at the corresponding position. For example, if the linear output is 0.7, it is determined that there is a first type of relationship at this position. Determine the end word that is closest to the target start word, establish the relationship from the start word to the end word, and complete the extraction of entity words and entity relationships. For example, the start word is "company" and the end word is "founder", and establish the "company-founder" relationship.

[0068] For another example, a text with a sequence length of 512 is input, and the feature dimension at each position is 768. The relationship between feature values is calculated through the self-attention mechanism. The feature values at each position in the self-attention model are obtained, and the Jensen-Shannon divergence is used to calculate the correlation degree between feature values. The specific formula is as follows, where P and Q respectively represent feature values of two different distributions, and the correlation degree between the feature value at position i and the feature value at position j is calculated. According to the correlation degree between feature values, the attention information between each position and other positions is calculated. For example, the attention weight between position 1 and position 2 is 0.85, and the attention weight between position 1 and position 3 is 0.72. The attention information is added to the amount of information carried by each position to obtain the attention information sum D. For example, the attention information sum of position 1 is D1 = Σ(attention weight × amount of information). The number of entity types m is determined. For example, m = 10, and a set of linear layer parameters is configured for each relationship type, representing the linear layer parameters of the kth relationship type. For example, the linear layer parameters of the first relationship type are W1, with a dimension of 768×768. The relationship between the feature values at each position included in the self-attention model is calculated through a feed-forward calculation using the linear layer. For example, the feature value of position 1 is input into the linear layer to obtain the output vector Y1 = W1 × X1. According to the feed-forward calculation result of the linear layer, it is determined whether there is a specific class relationship for the entity word at the corresponding position. For example, it is determined whether there is a first-class relationship between position 1 and position 2. If the calculation result is greater than the threshold of 0.8, there is a specific class relationship. If there is a specific class relationship, the end word with the closest relationship to the target start word is calculated. For example, starting from position 1, the correlation degree with position 2 is 0.85, and the correlation degree with position 3 is 0.72, and position 2 is determined as the closest end word. The relationship between the start word and the end word is established. For example, the relationship between position 1 and position 2 is marked as "person - place of birth", realizing the extraction of entity words and entity relationships.

[0069] Although the embodiments of the present invention have been shown and described above, it can be understood that the above examples are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0070] Finally, it should be emphasized that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for extracting content entities and relationships based on the self-attention model, characterized in that The method includes the following steps: S101. Establish a self-attention model in the last layer of the encoder, and input the parameters output by the encoder into the self-attention model for calculation; S102. In the self-attention model, the feature value at each position is calculated for the degree of association with the feature values at other positions to obtain an association value, which determines how much attention the target position needs to pay to each other position; S103. Calculate and sum the attention paid by other positions obtained in step S102 with the amount of information carried by this position as the output of the target position; S104. Set a linear layer at the back of the self-attention model, and input the parameters output by the self-attention into the linear model for calculation; S105. Use the linear layer to perform a feed-forward calculation on the relationship between the feature values of each position included in the self-attention model to make a judgment for the next step, and finally obtain the termination word closest to the target starting word, thereby establishing the relationship from the starting word to the termination word, and realizing the extraction of entity words and entity relationships.

2. The content entity and relationship extraction method based on the self-attention model according to claim 1, characterized in that, For tasks of multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type, perform single-group entity and relationship type discrimination on each set of linear layer parameters, and perform linear layer feed-forward calculation. Finally, obtain the termination word closest to the target starting word, and then establish the relationship from the starting word to the termination word.

3. A method for extracting content entities and relationships based on the self-attention model according to claim 1, characterized in that In step S102, the feature value at each position is calculated for the degree of association with the feature values at other positions. The degree of association of the feature values is calculated by the JS divergence, and its calculation formula is: where P and Q represent two different distributions, and using d i,j to represent the correlation degree between the eigenvalue a at the i-th position i and the eigenvalue a at the j-th position j we have:

4. A method for content entity and relationship extraction based on the self-attention model according to claim 1, characterized in that, The calculation formula for calculating and summing the attention paid by other positions obtained in step S102 with the amount of information carried by this position in step S103 is: Where n represents the sequence length, and 1 ≤ i, j ≤ n.

5. The content entity and relationship extraction method based on the self-attention model according to claim 2, wherein For tasks of multiple entities or multiple relationship types, configure a set of linear layer parameters for each entity or relationship type, perform single-group entity and relationship type discrimination on each set of linear layer parameters, and perform linear layer feed-forward calculation. Finally, obtain the termination word closest to the target starting word. This process is calculated by the following formula: t k = L k (D), Among them, k ∈ (1, m), where m is the number of entity types, and a set of linear layer parameters is configured for each relationship type, L k represents the linear layer parameters of the k-th relationship type, D is the sum of attention information, and through t k to determine whether there is a specific class relationship for the entity words at the corresponding positions.

6. A method for content entity and relationship extraction based on the self-attention model according to any one of claims 1-5, characterized in that The method further includes the following steps: For tasks of multiple entities or multiple relationship types, multiply the number of target coding values.

7. A method for extracting content entities and relationships based on the self-attention model according to any one of claims 1-5, characterized in that The method further includes the following steps: For tasks of multiple entities or multiple relationship types, equip each type of entity and type with an independent model, and perform multi-entity and multi-relationship recognition through the independent model.