Relationship extraction method for knowledge graph construction of low-annotation industrial corpus

By dynamically optimizing prompt templates and phased course learning, combined with industrial text features, the problem of insufficient generalization ability of relationship extraction in low-annotated industrial corpus is solved, the robustness and efficiency of the model in complex relationship recognition are improved, and the construction of high-quality knowledge graphs is supported.

CN120688598AActive Publication Date: 2025-09-23ZHEJIANG COLLEGE OF ZHEJIANG UNIV OF TECHOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510779151.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing relationship extraction methods lack generalization capabilities in low-labeled industrial corpora, making it difficult to effectively identify complex relationships. Furthermore, they are limited by high labeling costs and expertise requirements, resulting in insufficient performance and robustness of the models in industrial scenarios.

Method used

By adopting a dynamically optimized prompt template and a phased course learning strategy, combined with the contextual information of industrial text and the feature alignment mechanism, the model is optimized through cross-entropy loss to gradually improve the model's ability to discriminate complex relationships.

Benefits of technology

It improves the stability and generalization ability of relationship extraction in low-annotation environments, enhances the quality and efficiency of building industrial knowledge graphs, and adapts to multi-task relationship extraction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688598A_ABST
    Figure CN120688598A_ABST
Patent Text Reader

Abstract

A relation extraction method for construction of a knowledge graph of low-annotation industrial corpus comprises the steps that firstly, a dynamic prompt template is constructed in combination with context information of industrial text data, and a semantic collaboration mechanism is introduced to enhance semantic association between template marks; secondly, in combination with a feature alignment mechanism and staged course learning, the model can gradually adapt to multi-task setting with gradually increased difficulty, and the discrimination ability of the model to semantic differences is enhanced; and finally, calculating the matching probability of the relation representation and the label representation, and predicting the most probable relation type through a cross entropy loss optimization model. According to the method, the semantics of the context of the industrial text data is fully utilized, and a staged course learning strategy is combined, so that the model can gradually adapt to the complexity of a relationship extraction task, the requirement on the scale of annotated data is low, and the relationship extraction accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing technology, and in particular relates to a relationship extraction method for constructing a knowledge graph for low-annotated industrial corpus. Background Art

[0002] With the accelerating digital transformation of industrial manufacturing, equipment operation and maintenance, and production management, industrial knowledge graphs, as a key technology for semantic management of industrial data and intelligent decision-making, are being increasingly applied in key scenarios such as equipment fault diagnosis, process optimization, and knowledge reasoning. Relationship extraction, a core component of building knowledge graphs, identifies semantic relationships between entities in unstructured industrial text and plays a crucial role in improving the integrity and usability of knowledge graphs.

[0003] However, text data in the industrial field often features highly specialized terminology, complex linguistic styles, and high contextual dependencies, making relation extraction a significant challenge. In particular, high-quality, manually annotated data is extremely limited in practical applications. This high annotation cost and the need for specialized knowledge have led to the widespread availability of "poorly annotated corpus." This data scarcity severely limits the performance and generalization capabilities of traditional supervised learning methods in industrial scenarios.

[0004] Most existing relationship extraction methods rely on large-scale, high-quality annotated data, making them difficult to directly transfer to low-resource industrial corpora. Furthermore, the presence of a large number of long-tail, complex, and implicit relationships in industrial texts further complicates model learning. Rule-based methods have poor generalization capabilities when faced with complex sentence structures, while deep learning-based methods are prone to overfitting and lack robustness in small sample scenarios. Therefore, there is an urgent need for a relationship extraction method that can maintain good generalization and semantic understanding capabilities even with low-annotated corpora to support the construction and implementation of high-quality industrial knowledge graphs. Summary of the Invention

[0005] To overcome the challenges of scarce annotated data and complex relationship semantics in building industrial knowledge graphs, this paper proposes a relationship extraction method for low-annotated industrial corpus. This method uses dynamically optimized prompt templates to obtain industrial entity representations. In course learning tasks, a progressively more challenging training strategy is employed to gradually improve the model's ability to discern complex relationships. This enables efficient and robust relationship extraction in low-resource scenarios, supporting the construction of industrial knowledge graphs.

[0006] The technical solution adopted by the present invention to solve the technical problem is:

[0007] A relationship extraction method for constructing knowledge graphs from low-annotated industrial corpus. First, a dynamic prompt template is constructed based on the contextual information of industrial text data, and a semantic collaboration mechanism is introduced to enhance the semantic association between template tags. Second, a feature alignment mechanism and phased curriculum learning are combined to enable the model to gradually adapt to multi-task settings with increasing difficulty and enhance the model's ability to discriminate semantic differences. Finally, the matching probability between the relationship representation and the label representation is calculated, and the model is optimized through cross-entropy loss to predict the most likely relationship type.

[0008] Furthermore, the method comprises the following steps:

[0009] Step 1: Build entity annotation set E based on industry-specific terminology, build entity recognition template using existing industry terminology dictionary, and extract corresponding head and tail entities from text by exact matching. Use trainable continuous identifiers to replace fixed words in traditional prompts to build dynamic prompts T1 = [t1][t2]...[t m ][sub][MASK][obj], where [sub] and [obj] represent identified industrial entities, [t1][t2]...[t m ]∈V, V represents the pre-trained model vocabulary, m is the number of trainable tags, and [MASK] represents the complex semantic relationships that may exist between industrial entities in the context of industrial text data;

[0010] Step 2: Use the Word2Vec model to train word vectors for industrial entities, obtain words that are semantically similar to the target entity, and embed the first k related words into the vector e h and the head entity embedding e sub , tail entity embedding e obj The trainable label t in the mean initialization template T1 j Embedded representation of

[0011]

[0012] Step 3: The original relationship label set Y = {y1, y2, ..., y λ}, remove the symbols and interpret them into natural language expressions, and get the relationship description set Y'={y1',y2',...,y λ '}, and the manually constructed prompt T2 = means[MASK] form a new relation label sequence [CLS]y w '[SEP]means[MASK], w∈λ, embeds the mask position as the answer word to construct a continuous answer space Q={q1,q2,...,q λ},

[0013]

[0014] Step 4: Combine dynamic prompt T1 with industrial field text data X i =(x1,x2,x3,...,x n ), i∈N is merged, N is the number of samples in each batch. The merged sequence is passed through the encoder to obtain the corresponding word embedding representation Input into the pre-trained model, the output embedding F(χ i , y) and the output vector G(χ i ,y),

[0015] F(χ i ,y)=PLM(χ i );

[0016] G(χ i ,y)=W(F(χ i ,y))+b;

[0017] Where W∈R |V|×d is the output projection matrix of the model, b∈R |V| , |v| represents the vocabulary size;

[0018] Step 5: Input sequence X i Randomly mask a mark in the masked sequence X i ', replace the answer word corresponding to the real label in the prompt template T1 mask position to construct a new input sequence χ i ', use the pre-trained language model to predict the masked tag word word, through the cross entropy loss L f Maximize the conditional probability p(word|χ′ i ,y),

[0019]

[0020] Where M is the set of all masked locations and BCE() represents the binary cross entropy loss function.

[0021] The processing of steps 1 to 5 above is a specific implementation process of building a dynamic prompt template by combining the context information of industrial text data and introducing a semantic collaboration mechanism to enhance the semantic association between template tags.

[0022] Furthermore, the relationship extraction method further includes the following steps:

[0023] Step 6: Divide the training samples into two formats: masked format and unmasked format. In the masked format, the entity pair [sub], [obj] is used as the input of the model and forms a new input sequence χ with the prompt T1.entity =[CLS][sub][obj][SEP].T1[SEP], the output vector before classification of the mask position is used as the entity deviation G(χ entity ,y), blank input and prompt T1 constitute the input sequence χ prompt = [CLS] [SEP]. T1 [SEP], and obtain the output vector before classification of the mask position as the prompt deviation G (χ prompt ,y), subtract the bias from the pre-classification output vector of the original input,

[0024] x out =G(χ i ,y)-G(χ entity ,y)-G(χ prompt ,y);

[0025] The fine-tuning process is achieved through the cross entropy loss L mlm optimization,

[0026]

[0027] p(y|x)=w*ReLU(x out )+c;

[0028] Where w∈R |V|×d is the output projection matrix of the model, c∈R |V| is the bias variable;

[0029] Step 7: In the answer space Q = {r1, r2, ..., r n}Select the current input X i True relation label y i The corresponding answer word r i As a positive sample, randomly select i Any answer word other than r is used as a negative sample i ', by calculating the cosine similarity, the model prediction result F(χ i ,y) is closer to the real answer word and farther away from the irrelevant answer word,

[0030] L rel =max{cos(F(χ i ,y),r i )-cos(F(χ i ,y),r i ')+γ,0};

[0031] Step 8: In the unmasked format, replace the [MASK] token in the original prompt template with the relational answer word in the continuous space, and concatenate the output vectors corresponding to the positions of the two starting tokens to form the relational representation erel , and input into the classifier to output the probability distribution on the label set Y. The fine-tuning process is achieved through the cross entropy loss L cls Optimize,

[0032] e rel =[e [CLS] ;e [SEP] ];

[0033] p i =softmax(W*e rel +b);

[0034] L cls =-logp i .

[0035] The processing of steps 6 to 8 above is a specific implementation process that combines feature alignment with a phased curriculum learning strategy to enable the model to gradually adapt to the task difficulty and improve its ability to discriminate semantic differences.

[0036] Furthermore, the relationship extraction method further includes the following steps: Step 9, calculating the overall loss L,

[0037] L=α*(L mlm +L rel )+(1-α)*L cls +β*L f ;

[0038] Among them, when the input is in mask format, α is 1, otherwise it is 0, and β is a hyperparameter;

[0039] Step 10: At the beginning of training, a lower proportion of masked format samples is used and a higher proportion of unmasked format data is retained. As the training process progresses, the proportion of masked samples is gradually increased, and finally transitions to the fully masked format;

[0040] Step 11: Repeat steps 4 to 10. When L is less than the specified minimum loss value, the calculation ends and the relationship corresponding to the position index with the maximum probability in the prediction result is used as the final result of the industrial entity relationship in the industrial field text.

[0041] The process of steps 9 to 11 above is a specific implementation process for dynamically adjusting model parameters through a joint loss and using the converged model to infer relationships between entities. The technical concept of this invention is to construct dynamic prompts to learn knowledge features of annotation sets in industrial fields such as equipment, production processes, and quality control; to enhance semantic associations between templates using semantic coherence constraints; to enhance the model's adaptability to multi-task settings through phased curriculum learning; and to enhance the model's perception of relational semantics through a feature alignment mechanism.

[0042] The beneficial effects of the present invention are: it can integrate multi-type annotation knowledge in the industrial field and construct a dynamic prompt template with enhanced semantic association, which can effectively adapt to multi-task relationship extraction scenarios; it overcomes the problems of traditional methods such as insufficient understanding of domain semantics and insensitivity to changes in task difficulty, and improves the stability and generalization ability of relationship recognition in low-annotation environments, thereby significantly enhancing the quality and efficiency of industrial knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a relationship extraction method for constructing a knowledge graph from low-annotated industrial corpus. DETAILED DESCRIPTION

[0044] The present invention will be further described below with reference to the accompanying drawings.

[0045] Reference Figure 1 , a relation extraction method for constructing a knowledge graph for low-annotated industrial corpus, comprising the following steps:

[0046] Step 1: Build entity annotation set E based on industry-specific terminology, build entity recognition template using existing industry terminology dictionary, extract corresponding head entity and tail entity from text by exact matching, and use trainable continuous identifier to replace fixed words in traditional prompts to build dynamic prompt T1=[t1][t2]...[t m ][sub][MASK][obj], where [sub] and [obj] represent identified industrial entities, [t1][t2]...[t m ]∈V, V represents the pre-trained model vocabulary, m is the number of trainable tags, and [MASK] represents the complex semantic relationships that may exist between industrial entities in the context of industrial text data;

[0047] In this embodiment, the text data in the field of automobile maintenance, "Fault phenomenon of automobile fault report No. 363: Fault phenomenon: abnormal noise of steering gear when driving on bumpy road", is constructed by using GPT or automobile manual to construct entity annotation set E containing automobile domain nouns "engine" and "accelerator". The entities "steering gear" and "abnormal noise" in the text data are identified by rule matching, and dynamic prompts [t1][t2]...[t m ]Steering gear [MASK] abnormal noise;

[0048] Step 2: Use the Word2Vec model to train word vectors for industrial entities, obtain words that are semantically similar to the target entity, and embed the first k related words into the vector e h and the head entity embedding e sub , tail entity embedding e obj The trainable label t in the mean initialization template T1j Embedded representation of

[0049]

[0050] In this example, a large number of documents related to industrial equipment, such as maintenance manuals, operating procedures, and technical reports, are collected and trained to obtain a Word2Vec model that can learn the word vector representation of each industrial term. Through Word2Vec, k words similar to "steering gear" and "abnormal noise" are obtained, and the average embedding of these similar words is used to initialize the tags [t i ];

[0051] Step 3: The original relationship label set Y = {y1, y2, ..., y λ}, remove the symbols and interpret them into natural language expressions, and get the relationship description set Y'={y1',y2',...,y λ '}, and the manually constructed prompt T2 = means[MASK] form a new relation label sequence [CLS]y w '[SEP]means[MASK], w∈λ, embeds the mask position as the answer word to construct a continuous answer space Q={q1,q2,...,q λ},

[0052]

[0053] In this example, the original relation label "component failure" can be briefly expanded into a natural language expression of "faulty component". This is then combined with the manually constructed prompt template T2 and input into the pre-trained model. The resulting mask position embedding is used to initialize the answer word corresponding to the relation label.

[0054] Step 4: Combine dynamic prompt T1 with industrial field text data X i =(x1,x2,x3,...,x n ), i∈N is merged, N is the number of samples in each batch. The merged sequence is passed through the encoder to obtain the corresponding word embedding representation Input into the pre-trained model, the output embedding F(χ i , y) and the output vector G(χ i ,y),

[0055] F(χ i ,y)=PLM(χ i );

[0056] G(χ i ,y)=W*F(χi ,y)+b;

[0057] Where W∈R |V|×d is the output projection matrix of the model, b∈R |V| is the bias variable, |v| is the vocabulary size;

[0058] In this embodiment, the prompt T1 is spliced ​​with the original industrial text data "Fault Report No. 363, Fault Phenomenon: When the vehicle is driving on a bumpy road, the steering wheel makes an abnormal noise" to form a complete input. During the model training process, such input is packaged into a batch, and each batch contains several such prompt-enhanced texts. These merged text sequences are sent to the encoder for processing, and the encoder converts each word in the entire sentence into a corresponding embedding vector. These embeddings are then sent to a pre-trained language model (such as BERT or RoBERTa), and the model outputs a vector representation at the position of [MASK] and calculates the probability of each word being the answer;

[0059] Step 5: Input sequence X i Randomly mask a mark in the masked sequence X i ', replace the answer word corresponding to the real label in the prompt template T1 mask position to construct a new input sequence χ i ', use the pre-trained language model to predict the masked tag word word, through the cross entropy loss L f Maximize the conditional probability p(word|χ′ i ,y),

[0060]

[0061] Where M is the set of all masked positions, and BCE() represents the binary cross entropy loss function;

[0062] In this embodiment, a word is randomly selected from the original industrial text data "No. 363 car fault report fault phenomenon: when the vehicle is driving on a bumpy road, the steering gear makes an abnormal noise". For example, "car" can be masked to obtain a new text input "No. 363 [MASK] fault report fault phenomenon: when the vehicle is driving on a bumpy road, the steering gear makes an abnormal noise". At the [MASK] position in the prompt template T1, the answer word r corresponding to the relationship label between "steering gear" and "abnormal noise" is used. i Replace to get new prompt template [t1][t2]...[t n Steering gear iThese inputs are then fed into a pre-trained language model, which is then tasked with predicting the randomly masked word "car." This entire training process maximizes the probability of the correct word appearing, thereby continuously optimizing the model's ability to understand context.

[0063] Step 6: Divide the training samples into two formats: masked format and unmasked format. In the masked format, the entity pair [sub], [obj] is used as the input of the model and forms a new input sequence χ with the prompt T1. entity =[CLS][sub][obj][SEP].T1[SEP], the output vector before classification of the mask position is used as the entity deviation G(χ entity ,y), blank input and prompt T1 constitute the input sequence χ prompt = [CLS] [SEP]. T1 [SEP], and obtain the output vector before classification of the mask position as the prompt deviation G (χ prompt ,y), subtract the bias from the pre-classification output vector of the original input,

[0064] x out =G(χ i ,y)-G(χ entity ,y)-G(χ prompt ,y);

[0065] The fine-tuning process is achieved through the cross entropy loss L mlm optimization,

[0066]

[0067] p(y|x)=w*ReLU(x out )+c;

[0068] Where w∈R |V|×d is the output projection matrix of the model, c∈R |V| is the bias variable;

[0069] In this embodiment, the original text input and template T1 are directly combined as the input of the model in the mask format, and the predicted score logits of the label can be obtained at the [MASK] position. In addition, another blank format input containing only prompts is constructed, and only "[t1][t2]...[t n ] Steering gear [MASK] abnormal noise" is input into the model. At this time, the model's prediction result for [MASK] mainly reflects the preference of the template itself or the guiding effect of the prompt, which is called prompt bias. The entity pair "steering gear" and "abnormal noise" are used as input and combined with model T1 to form a new input "[CLS] steering gear, abnormal noise [SEP][t1][t2]...[t n] Steering wheel [MASK] abnormal noise [SEP]" is input into the model. The model's prediction for [MASK] primarily reflects the guiding effect of the entity, known as entity bias. To eliminate the influence of bias on the results, the bias score is subtracted from the original prediction score to retain the true semantic differences related to the entity. Finally, during training, this adjusted output is optimized using cross-entropy loss, allowing the model to focus more on the true semantic relationships between entities rather than the influence of the template itself.

[0070] Step 7: In the answer space Q = {r1, r2, ..., r n}Select the current input X i True relation label y i The corresponding answer word r i As a positive sample, randomly select i Any answer word other than r is used as a negative sample i ', by calculating the cosine similarity, the model prediction result F(χ i ,y) is closer to the real answer word and farther away from the irrelevant answer word,

[0071] L rel =max{cos(F(χ i ,y),r i )-cos(F(χ i ,y),r i ')+γ,0};

[0072] In this embodiment, the relationship between the entity pairs "steering gear" and "abnormal noise" in the original industrial text data "Fault Report No. 363: When the vehicle is driving on a bumpy road, the steering gear makes abnormal noises" is "component failure". During training, the word corresponding to the relationship label "component" in the answer space is used as a positive sample, and at the same time, another relationship word, such as "failure" or "dependency", is randomly selected from the answer space as a negative sample. After the prompt template processing and context encoding, the model will output a vector in the answer space. Calculate the cosine similarity between this vector and the word vector of the positive sample ("component failure"), and also calculate the similarity with the negative sample. The training goal is to make the vector output by the model closer to "component failure" and away from irrelevant relationships such as "failure";

[0073] Step 8: In the unmasked format, replace the [MASK] token in the original prompt template with the relational answer word in the continuous space, and concatenate the output vectors corresponding to the positions of the two starting tokens to form the relational representation e rel , and input into the classifier to output the probability distribution on the label set Y. The fine-tuning process is achieved through the cross entropy loss L cls Optimize,

[0074] e rel =[e [CLS] ;e [SEP] ];

[0075] p i =soft max(W*e rel +b);

[0076] L cls =-log p i ;

[0077] Where W∈R |V|×d is the output projection matrix of the model, b∈R |V| is the bias variable;

[0078] In this embodiment, the relationship between the entity pairs "steering gear" and "abnormal noise" in the original industrial text data "Fault report of automobile No. 363, Fault phenomenon: when the vehicle is driving on a bumpy road, the steering gear makes abnormal noises" is "component failure". In the unmasked format, [MASK] is no longer used as a placeholder in the template T1, but the relational answer word "component failure" is directly filled in the [MASK] position in the prompt template. Next, the output vectors of the two starting positions of the input are taken out and spliced ​​as the semantic representation of the entire relationship. This relational representation is then fed into a classifier, which outputs a probability distribution over all possible relational labels (such as "component failure", "control", "dependency", etc.). Finally, the cross-entropy loss is used to optimize the model;

[0079] Step 9: Calculate the overall loss L.

[0080] L=α*(L mlm +L rel )+(1-α)*L cls +β*L f ;

[0081] Among them, when the input is in mask format, α is 1, otherwise it is 0, and β is a hyperparameter;

[0082] Step 10: In the early stages of training, use a lower proportion of masked samples and retain a higher proportion of unmasked data. As the training process progresses, gradually increase the proportion of masked samples, and eventually transition to a fully masked format.

[0083] In this embodiment, for the task of extracting industrial text relationships, the initial stage of model training mainly uses samples in an unmasked format, with a ratio of approximately 80% unmasked and 20% masked. This stage helps the model quickly establish basic semantic connections between entities and relationships by providing clear relationship prompt words. As the training progresses, the model's understanding of the task gradually improves, and the number of masked format samples is gradually increased by approximately 20% per stage, reducing dependence on explicit relationship words. Finally, in the later stages of training, the training data is completely transitioned to a masked format, allowing the model to judge the relationship between entities based solely on contextual information in the absence of relationship prompt words, thereby better adapting to the relationship prediction needs in practical applications;

[0084] Step 11: Repeat steps 4 to 10. When L is less than the specified minimum loss value, the calculation ends and the relationship corresponding to the position index with the maximum probability in the prediction result is used as the final result of the industrial entity relationship in the industrial field text.

[0085] In this embodiment, the model iteratively executes the relationship prediction steps, updating parameters, calculating losses, and evaluating the difference between the current output and the true relationship. Training ends when the loss value gradually decreases until it falls below a set minimum threshold. At this point, the relationship label corresponding to the position with the highest probability in the predicted probability distribution of all possible relationship labels is selected as the final output for subsequent applications such as knowledge graph construction and fault analysis.

[0086] The embodiments of this specification are merely examples of implementations of the invention and are provided for illustrative purposes only. The scope of protection of the present invention should not be considered limited to the specific embodiments described in these embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by a person of ordinary skill in the art based on the invention.

Claims

1. A relation extraction method for constructing a knowledge graph of low-labeled industrial corpus, characterized in that: First, a semantic collaboration mechanism is introduced to optimize the dynamic prompt template constructed with contextual information. Second, feature alignment is combined with a phased curriculum learning strategy to enable the model to gradually adapt to the task difficulty and improve its ability to discriminate semantic differences. Finally, the model parameters are dynamically adjusted through a joint loss, and the converged model is used to infer the relationship between entities.

2. A relationship extraction method for constructing a knowledge graph of low-labeled industrial corpus according to claim 1, characterized in that: The relationship extraction method comprises the following steps: Step 1: Build entity annotation set E based on industry-specific terminology, build entity recognition template using existing industry terminology dictionary, extract corresponding head entity and tail entity from text by exact matching, and use trainable continuous identifier to replace fixed words in traditional prompts to build dynamic prompt T1=[t1][t2]...[t m ][sub][MASK][obj], where [sub] and [obj] represent identified industrial entities, [t1][t2]...[t m ]∈V, V represents the pre-trained model vocabulary, m is the number of trainable tags, and [MASK] represents the complex semantic relationships that may exist between industrial entities in the context of industrial text data; Step 2: Use the Word2Vec model to train word vectors for industrial entities, obtain words that are semantically similar to the target entity, and embed the first k related words into the vector e h and the head entity embedding e sub , tail entity embedding e obj The trainable label t in the mean initialization template T1 j Embedded representation of Step 3: The original relationship label set Y = {y1, y2, ..., y λ }, remove the symbols and interpret them into natural language expressions, and get the relationship description set Y'={y1',y2',...,y λ '}, and the manually constructed prompt T2 = means[MASK] form a new relation label sequence [CLS]y w '[SEP]means[MASK], w∈λ, embeds the mask position as the answer word to construct a continuous answer space Q={q1,q2,...,q λ }, Step 4: Combine dynamic prompt T1 with industrial field text data X i =(x1,x2,x3,...,x n ), i∈N is merged, N is the number of samples in each batch, and the merged sequence is passed through the encoder to obtain the corresponding word embedding representation Input into the pre-trained model, the output embedding F(χ i , y) and the output vector G(χ i ,y), F(x i ,y)=PLM(x i ); G(x i ,y)=W*F(x i ,y))+b; Where W∈R |V|×d is the output projection matrix of the model, b∈R |V| is the bias variable, |v| is the vocabulary size; Step 5: Input sequence X i Randomly mask a mark in the masked sequence X i ', replace the answer word corresponding to the real label in the prompt template T1 mask position to construct a new input sequence χ i ', use the pre-trained language model to predict the masked tag word word, through the cross entropy loss L f Maximize the conditional probability p(word|χ i ′,y), Where M is the set of all masked locations and BCE() represents the binary cross entropy loss function.

3. A relationship extraction method for constructing a knowledge graph of low-labeled industrial corpus according to claim 2, characterized in that: The relationship extraction method further comprises the following steps: Step 6: Divide the training samples into two formats: masked format and unmasked format. In the masked format, the entity pair [sub], [obj] is used as the input of the model and forms a new input sequence χ with the prompt T1. entity =[CLS][sub][obj][SEP].T1[SEP], the output vector before classification of the mask position is used as the entity deviation G(χ entity ,y), blank input and prompt T1 constitute the input sequence χ prompt = [CLS] [SEP]. T1 [SEP], and obtain the output vector before classification of the mask position as the prompt deviation G (χ prompt ,y), subtract the bias from the pre-classification output vector of the original input, x out =G(x i ,y)-G(x entity ,y)-G(x prompt ,y); The fine-tuning process is achieved through the cross entropy loss L mlm optimization, p(y|x)=w*ReLU(x out )+c: Where w∈R |V|×d is the output projection matrix of the model, c∈R |V| is the bias variable; Step 7: In the answer space Q = {r1, r2, ..., r n }Select the current input X i True relation label y i The corresponding answer word r i As a positive sample, randomly select i Any answer word other than r is used as a negative sample i ', by calculating the cosine similarity, the model prediction result F(χ i ,y) is closer to the real answer word and farther away from the irrelevant answer word, L rel =max{cos(F(χ i ,y),r i )-cos(F(x i ,y),r i ')+γ,0}; Step 8: In the unmasked format, replace the [MASK] token in the original prompt template with the relational answer word in the continuous space, and concatenate the output vectors corresponding to the positions of the two starting tokens to form the relational representation e rel , and input into the classifier to output the probability distribution on the label set Y, the fine-tuning process is carried out through the cross entropy loss L cls Optimize, And rel =[and [CLS] ;And [SEP] ]; p i =softmax(W*e rel +b); L cls =-logp i ; Where W∈R |V|×d is the output projection matrix of the model, b∈R |V| is the bias variable.

4. A relationship extraction method for constructing a knowledge graph of low-labeled industrial corpus according to claim 3, characterized in that: The relationship extraction method further comprises the following steps: Step 9: Calculate the overall loss L. L=α*(L mlm +L rel )+(1-a)*L cls +β*L f ; Among them, when the input is in mask format, α is 1, otherwise it is 0; β is a hyperparameter; Step 10: At the beginning of training, a lower proportion of masked format samples is used and a higher proportion of unmasked format data is retained. As the training process progresses, the proportion of masked samples is gradually increased, and finally the full mask format is transitioned. Step 11: Repeat steps 4 to 10. When L is less than the specified minimum loss value, the calculation ends and the relationship corresponding to the position index with the maximum probability in the prediction result is used as the final result of the industrial entity relationship in the industrial field text.

Citation Information

Patent Citations

  • Knowledge perception prompt learning-based few-sample relation extraction method

    CN117010392A

  • Complex semantic relationship extraction method for industrial knowledge graph construction

    CN119558325A

  • Multi-modal knowledge graph completion method based on dynamic prompt learning and multi-granularity aggregation

    CN119783799A

  • Rapid labeling method based on artificial intelligence large model

    CN119848549A