A Method for Extracting Complex Semantic Relations for Industrial Knowledge Graph Construction
Through dynamic prompts and contrast learning methods, the problem of complex semantic relationship extraction in the construction of industrial knowledge graphs is solved, more efficient and accurate relationship extraction is achieved, and high-quality data support is provided.
Patent Information
- Application Number
- CN202510115274.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
It is difficult for the prior art to effectively extract complex semantic relationships in the construction of industrial knowledge graphs, especially when data labeling is scarce and relationship semantic expression is complex.
Using dynamic prompts and implicit structural constraints, the in-class aggregation and inter-class separation of relational features in relational storage queues is enhanced through comparative learning, and the training process is dynamically optimized to infer the relationship between industrial entities in the industrial knowledge graph.
It improves the accuracy and efficiency of extracting complex semantic relationships in the construction of industrial knowledge graphs, overcomes the limitations of traditional methods in low-resource environments, and provides high-quality data support.
Smart Images

Figure CN119558325B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method for extracting complex semantic relationships for industrial knowledge graph construction. Background Art
[0002] With the deep promotion of informatization and intelligentization in the industrial field, knowledge graph technology has been widely applied in fields such as industrial production, quality management, and equipment maintenance. Industrial knowledge graphs effectively present entities and their relationships in the industrial field in a structured data expression form. Combining real-time data analysis and reasoning helps decision-makers quickly identify problems, predict risks, and optimize processes, and is conducive to providing important data support for decision-making on complex issues such as equipment fault diagnosis and supply chain management. However, the complexity, diversity of industrial data, and the richness of relationship semantic expressions pose great challenges to the construction of knowledge graphs, especially in the extraction of complex semantic relationships.
[0003] Traditional rule-based relationship extraction methods are difficult to adapt to the diverse and long-tail distributed relationships in industrial text data; data-driven methods conflict with the scarcity of industrial corpus and high annotation costs in the need for large-scale labeled data; deep learning-based methods are sensitive to industrial domain-specific terms and context noise and are difficult to fully utilize the semantic information of complex relationship labels. Existing methods are insufficient in extracting complex semantic relationships from small-sample industrial text data, and there is an urgent need for a more efficient and robust solution to support the construction of industrial knowledge graphs. Summary of the Invention
[0004] In order to overcome the deficiencies in the existing industrial knowledge graph construction, such as scarce data annotation, complex semantic expressions of relationships between industrial entities, as well as problems such as diverse relationship types and semantic ambiguities in industrial domain knowledge, the present invention proposes a method for extracting complex semantic relationships for industrial knowledge graph construction, which obtains relationship representations between industrial entities by using dynamic prompts and implicit structure constraints, and infers the relationship categories between entities in industrial text data through clustering.
[0005] The technical solution adopted by the present invention to solve the technical problems is as follows:
[0006] A method for extracting complex semantic relationships for industrial knowledge graph construction, comprising the following steps:
[0007] Step 1, construct dynamic prompts in combination with the context of industrial text data to obtain relationship representations between industrial entities;
[0008] Step 2, adopt contrastive learning to enhance the intra-class aggregation and inter-class separation of relationship features in the relationship storage queue;
[0009] Step 3: Dynamically optimize the training process, and use the converged model to infer the relationships between industrial entities in the industrial knowledge graph.
[0010] Further, in the above step 1, the process of constructing dynamic prompts in combination with the context of industrial text data to obtain the relationship representation between industrial entities includes the following steps:
[0011] Step 11: Introduce professional vocabulary in the industrial field as the entity annotation set E, analyze the entity types of equipment names, failure types, and process flow in E, construct the entity type set H, and enhance the set E by using the strong correlation and timeliness characteristics of industrial field knowledge; according to the industrial entity set E, use rules to identify the head and tail industrial entities in the given industrial field text data, and construct the dynamic prompt Γ = [T 1 [E 1 [MASK][T 2 [E 2 , where E 1 and E 2 represent the identified industrial entities, T 1 and T 2 are the pseudo-types of the corresponding industrial entities, initialized by the set H, and [MASK] is the possible complex semantic relationship of the industrial entity in the context of industrial text data.
[0012] Step 12: Combine the dynamic prompt Γ with the industrial field text data X i = (x 1 , x 2 , x 3 ,..., x n ), i ∈ N, where N is the number of samples in each batch, and then map it into a sequence of word vectors word by word through the embedding layer. The unified sequence length is L. If the mapped sequence length is less than L, pad with 0 at the end; if the sequence length is greater than L, directly truncate the redundant characters at the end. Obtain the encoded representation χ i = (e 1 , e 2 , e 3 ,..., e n ) of the input sequence through the encoder, and add the absolute position encoding POS to χ i ,
[0013]
[0014] where d model is equal to the dimension of the embedding layer;
[0015] Step 13: Select a deep learning model for the industrial field with the hidden layer M, and select the [MASK] embeddings from the second layer to the M - 1 layer to construct the industrial entity relationship matrix χ know ,
[0016]
[0017] Step 14, calculate the industrial entity relationship matrix χ know Take the tanh of
[0018] u i = tanh(χ know );
[0019] Step 15, calculate the similarity between u i and the context representation w of the industrial domain text data u ;
[0020]
[0021] α i is the importance measure of the industrial entity relationship representation predicted by each feature layer of the model compared to the industrial domain text data. Use α i as the global weighted sum on χ know to generate the final relationship representation χ i between industrial entities;
[0022]
[0023] Furthermore, in the above Step 2, the process of adopting contrastive learning to enhance the intra-class aggregation and inter-class separation of the relationship features in the relationship storage queue includes the following steps:
[0024] Step 21, construct an instance-level contrastive loss for a small amount of manually labeled and unlabeled industrial domain text data sets
[0025]
[0026] Among them, identify industrial entity pairs through rules and the industrial entity annotation set E and the relationship representation χ i predicted by the model to form a positive sample pair Select two different arbitrary-span segments from the current industrial text data as pseudo-entities E 1 i′ and E 2 i′ . Among them, for a small amount of manually labeled industrial domain text data sets, select the true labeled relationship labels; for unlabeled industrial text data sets, select the central representation R i of their relationship clusters to form a negative sample pair
[0027] Step 22, calculate the probability distribution of the possible relationships between industrial entities based on χ i ;
[0028] p i = W T (ReLU(χ i )) + b;
[0029] Where W is the classification weight and b is the bias parameter;
[0030] Step 23, calculate the relationship category of the industrial entity in the industrial data context as
[0031] y i = argmax(p i );
[0032] Step 24, construct a classification cross - entropy loss for the labeled industrial domain text dataset
[0033]
[0034] Step 25, construct a relationship queue set R = {R 1 , R 2 ,..., R o} for the unlabeled industrial text data, where o represents the number of unlabeled categories, the size of the queue R i is b·N, where b is the batch size. For the positive sample relationship representation with the pseudo - queue label as the contrast set is The contrast set is
[0035] Step 26, calculate the cluster - center semantic similarity between the industrial text data and each relationship queue, and minimize the cross - entropy L between the relationship - queue semantic similarity i and the classification probability p CO ,
[0036]
[0037] where τ is the temperature coefficient. After each round of training, update the queue label where the industrial entity relationship in the industrial text data in the iteration period using maximum likelihood estimation
[0038]
[0039] After each backpropagation ends, add to
[0040] Furthermore, in step 3, the process of dynamically optimizing the training process and using the converged model to infer the relationships between industrial entities in the industrial knowledge graph includes the following steps:
[0041] Step 31: In adjacent training rounds, record the number of changes in the predicted categories of industrial entity relationships in industrial text data as the sample allocation weight w i , and co-optimize the instance-level contrast loss
[0042]
[0043] where is the relationship of the industrial entity predicted in the k-th round in the given industrial domain text data is χ i Adjust the parameters in the k-th round;
[0044] Step 32: Calculate the overall loss L
[0045] L = L CO + L EN + λ · L CE ;
[0046] where λ is a hyperparameter;
[0047] Step 33: If L is less than the specified minimum loss value or the maximum number of training rounds is reached, terminate the training, and use as the final result of the relationships between industrial entities in the industrial domain text, otherwise repeat steps 13 to 32.
[0048] The technical concept of the present invention is: using prompt learning to fuse the knowledge features of industrial domain annotation sets such as equipment, production processes, and quality control; using contrast loss to reduce the interference of noise in industrial text data; constructing a relationship repository for unannotated industrial text instances, co-optimizing the annotation set for dynamic clustering, improving the quality of industrial domain relationship extraction, and providing reliable data support for the construction of a high-quality industrial knowledge graph.
[0049] The beneficial effects of the present invention are: it can process the multi-level and multi-dimensional association information between entities in industrial text data, overcome the limitations of traditional methods in processing complex domain data in a low-resource environment, and improve the accuracy and efficiency of industrial knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flow schematic diagram of a complex semantic relationship extraction method for industrial knowledge graph construction. DETAILED DESCRIPTION OF THE INVENTION
[0051] The present invention will be further described below with reference to the accompanying drawings.
[0052] Refer toFigure 1 , A complex semantic relation extraction method for industrial knowledge graph construction, comprising the following steps:
[0053] Step 1. Construct dynamic prompts in combination with the context of industrial text data to obtain the relationship representation between industrial entities. The processing process includes the following steps:
[0054] Step 11. Introduce professional vocabulary in the industrial field as the entity annotation set E, analyze the entity types of equipment names, fault types, and process flows in E, construct the entity type set H, and enhance the set E by using the strong relevance and timeliness characteristics of industrial field knowledge; according to the industrial entity set E, use rules to identify the head and tail industrial entities in the given industrial field text data, and construct the dynamic prompt Γ = [T 1 [E 1 [MASK][T 2 [E 2 , where E 1 , E 2 represent the identified industrial entities, T 1 , T 2 are the pseudo-types of the corresponding industrial entities, initialized by the set H, and [MASK] is the complex semantic relationship that may exist between industrial entities in the context of industrial text data;
[0055] In this embodiment, the labeled data text X i in the automotive field is "Comprehensive situation of automotive fault report No. 62: Fault phenomenon: After acceleration, when the throttle is released, the engine stalls". Use GPT or automotive manuals to construct the entity annotation set E containing automotive field nouns such as "engine" and "throttle", and the entity type set H containing "fault" and "parts", and initialize the virtual type marker [T i with the average embedding of the entity types in H. Use rule matching to identify the entities "engine" and "stall" in X i with the help of the entity annotation set, and construct the dynamic prompt Γ = [T 1 engine[MASK][T 1 stall;
[0056] Step 12. Combine the dynamic prompt Γ with the industrial field text data X i = (x 1 , x 2 , x 3 ,..., x n ), i ∈ N, where N is the number of samples in each batch, and then map each word to a sequence of word vectors through the embedding layer. The unified sequence length is L. If the length of the mapped sequence is less than L, pad with 0 at the end; if the sequence length is greater than L, directly truncate the redundant characters at the end, and obtain the encoded representation χ i = (e1 , e 2 , e 3 ,..., e n ), add the absolute position encoding POS in χ i , where d
[0057]
[0058] equals the dimension of the embedding layer; model
[0059] In this embodiment, after Γ and X i are merged, the input sequence is [CLS]X i [SEP].Γ[SEP]. Using the RoBERTa deep learning model, if the length of the example in step 11 after preprocessing and passing through the tokenizer is less than 256, then fill the tail of the sequence with "0" markers until the maximum length, and finally sum the word vector encoding and position encoding generated by the model;
[0060] Step 13: Select a deep learning model with the hidden layer being M for the industrial field, and select the [MASK] embeddings from the second layer to the (M - 1)-th layer to construct the industrial entity relationship matrix χ know ,
[0061]
[0062] In this embodiment, the size of the hidden layer of the RoBERTa deep learning model used in the example of step 12 is 25, and the [MASK] positions in the second to the 25th hidden layers are selected to construct the relationship matrix;
[0063] Step 14: Calculate the tanh of the industrial entity relationship matrix χ know ,
[0064] u i = tanh(χ know );
[0065] Step 15: Calculate the similarity between u i and the context representation w u of the industrial field text data,
[0066]
[0067] α i is the importance measure of the industrial entity relationship representation predicted by each feature layer of the model compared to the industrial field text data. Use α i as the global weighted sum on χ know to generate the final relationship representation χ i between industrial entities,
[0068]
[0069] In this embodiment, w u is the embedding at the [CLS] position in the last hidden layer of RoBERTa. The similarities with the embeddings at the [MASK] positions in the selected hidden layers from the 2nd to the 25th are calculated respectively. For example, q = (0.2, 0.3,..., 0.3). Using q as the weight of the [MASK] embeddings in the selected hidden layers, 0.2·χ 1_mask + 0.3·χ 2_mask +,...+ 0.3·χ 23_mask is used as the representation of the relationship between "engine" and "flameout".
[0070] Step 2: Use contrastive learning to enhance the intra-class aggregation and inter-class separation of the relationship features in the relationship storage queue. The processing includes the following steps:
[0071] Step 21: Construct an instance-level contrastive loss for a small amount of manually annotated and unannotated industrial domain text datasets,
[0072]
[0073] where industrial entity pairs are identified through rules and the industrial entity annotation set E and the relationship representation χ i predicted by the model form positive sample pairs Select two different arbitrary-span segments from the current industrial text data as pseudo-entities E 1 i′ , E 2 i′ . Among them, for a small amount of manually annotated industrial domain text datasets, select the true annotated relationship labels; for unannotated industrial text datasets, select the central representation R i of their relationship clusters to form negative sample pairs
[0074] In this embodiment, the annotation set is "Fault phenomenon of vehicle fault report No. 795, abnormal noise at the right front wheel when driving on an uneven road", (right front wheel, χ i , abnormal noise) is a positive sample, and the negative sample pair is (vehicle, component fault, flat road); for the unannotated set "Comprehensive situation of vehicle fault report No. 62: Fault phenomenon: After acceleration, release the throttle, and the engine flames out", the positive sample pair is (engine, χ i , flameout), and select any span in the instance to generate pseudo-entities, such as (report, cluster central representation, throttle) as the negative sample pair;
[0075] Step 22: Calculate the probability distribution of the possible relationships between industrial entities based on χ i ,
[0076] pi = W T (ReLU(χ i )) + b;
[0077] Where W is the classification weight and b is the bias parameter;
[0078] Step 23, calculate that the relationship category of the industrial entity in the industrial data context is
[0079] y i = argmax(p i );
[0080] In this embodiment, the relationship label annotation set y = {component failure, assembly, no relationship}, and the probability distribution calculated in step 22 is p = {0.6, 0.2, 0.2}, then the final relationship category of the industrial entity in the industrial data context is "component failure";
[0081] Step 24, construct a classification cross-entropy loss for the labeled industrial domain text dataset
[0082]
[0083] Step 25, construct a relationship queue set R = {R 1 , R 2 ,..., R o} for the unlabeled industrial text data, where o represents the number of unlabeled categories, and the size of the queue R i is b·N, where b is the batch size. For the positive sample relationship representation with the pseudo queue label the comparison set is The comparison set is
[0084] In this embodiment, an empty relationship queue set R = {R 1 , R 2 ,..., R 20} with a size of 7 is constructed for the unlabeled dataset. The maximum length of the queue is 240, where the batch size is 8 and the size of each batch is 30. Among them, for the relationship representation belonging to queue 2, its comparison set is all relationship queue sets except queue 2;
[0085] Step 26, calculate the cluster center semantic similarity between the industrial text data and each relationship queue, and minimize the cross-entropy L between the semantic similarity based on the relationship queue i and the classification probability p CO calculated based on the classification weight parameter,
[0086]
[0087] Among them, τ is the temperature coefficient. After each round of training, the industrial entity relationships in the industrial text data in the iteration period are updated using maximum likelihood estimation. The queue label where it is located
[0088]
[0089] After each backpropagation ends, is added to
[0090] In this embodiment, for the unlabeled data "Comprehensive situation of vehicle 62 fault report: Fault phenomenon: After acceleration, when releasing the throttle, the engine stalls", the relationship embedding between the entity pair "engine" and "stalls" is obtained through steps 11 - 15. The cluster center embedding of the relationship queue is calculated using the clustering algorithm. Next, the similarity between the relationship embedding and the center embeddings of different queues is calculated respectively. If the similarity with queue 2 is the largest, the relationship label between the entity pairs is the result decoded from the cluster center embedding of queue 2;
[0091] Step 3: Dynamically optimize the training process, and use the converged model to infer the relationships between industrial entities in the industrial knowledge graph. The processing process includes the following steps:
[0092] Step 31: In adjacent training rounds, record the number of times the predicted category of the industrial entity relationship in the industrial text data changes as the sample allocation weight w i , and co - optimize the instance - level contrastive loss,
[0093]
[0094] where is the relationship of the industrial entity predicted in the k - th round in the given industrial domain text data, is χ i Adjust the parameter in the k - th round;
[0095] In this embodiment, taking the unlabeled data exemplified in step 26 as an example, if the queue where the relationship label between the entity pair "engine" and "stalls" is located changes from R 2 to R 3 in the training from the first round to the second round, then the cross - entropy coefficient in the second round is Otherwise, it is
[0096] Step 32: Calculate the overall loss L,
[0097] L = L CO + L EN + λ·L CE ;
[0098] Among them, λ is a hyperparameter;
[0099] Step 33: If L is less than the specified minimum loss value or the maximum number of training rounds is reached, the calculation ends, and is used as the final result of the industrial entity relationship in the industrial field text. Otherwise, repeat steps 13 to 32.
[0100] In this embodiment, taking the unlabeled data exemplified in step 26 as an example, the trained RoBERTa clusters the relationship prediction result between the entity pair "engine" and "flameout" into queue R 2 , then the relationship between the entity pairs is the result decoded from the cluster center embedding of queue 2.
[0101] The content described in the embodiments of this specification is only a list of implementation forms of the inventive concept and is only for illustrative purposes. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.
Claims
1. A complex semantic relationship extraction method for industrial knowledge graph construction, characterized in that: The method comprises the following steps: Step 1: Combine the industrial text data context to construct dynamic prompts to obtain the relationship representation between industrial entities; Step 2: Use contrastive learning to enhance the intra-class aggregation and inter-class separation of relation features in the relation storage queue; Step 3: Dynamically optimize the training process and use the converged model to infer the relationship between industrial entities in the industrial knowledge graph; In step 1, the process of constructing dynamic prompts in combination with industrial text data context to obtain the representation of the relationship between industrial entities includes the following steps: Step 11, introduce professional vocabulary in the industrial field as the entity annotation set E, analyze the equipment name, fault type, and process flow entity type in E, construct the entity type set H, and enhance the set E by using the strong correlation and temporal characteristics of industrial field knowledge; according to the industrial entity set E, use the rules to identify the head and tail industrial entities in the given industrial field text data, and construct a dynamic prompt Γ = [T1][E1][MASK][T2][E2], where E1 and E2 represent the identified industrial entities, T1 and T2 are the pseudo types of the corresponding industrial entities, initialized by the set H, and [MASK] is the complex semantic relationship that the industrial entity may have in the context of industrial text data; Step 12: Combine the dynamic prompt Γ with the industrial field text data X i =(x1,x2,x3,...,x n ), i∈N, N is the number of samples in each batch, and then the words are mapped to word vector sequences through the embedding layer, with a unified sequence length of L. If the length of the mapped sequence is less than L, the end is padded with 0; if the sequence length is greater than L, the redundant characters at the end are directly truncated, and the encoded representation of the input sequence χ is obtained through the encoder i =(e1,e2,e3,...,e n ), in χ i Add absolute position code POS, Among them, d model Equal to the dimension of the embedding layer; Step 13: Select a deep learning model for the industrial field with M hidden layers, and select [MASK] embedding from the 2nd to the M-1th layer to construct the industrial entity relationship matrix χ know Step 14: Calculate the industrial entity relationship matrix χ know tanh, u i =tanh(χ know ); Step 15: Calculate u i and contextual representation of industrial text data w u The similarity between α i α is used to measure the importance of the industrial entity relationship representation predicted by each feature layer of the model compared to the industrial field text data. i As χ know The global weighted summation on the final industrial entity relationship representation χ i , In step 2, the process of using contrastive learning to enhance intra-class aggregation and inter-class separation of relational features in the relational storage queue includes the following steps: Step 21: Construct instance-level contrastive loss for a small number of manually annotated and unannotated industrial text datasets. Among them, the industry entity pairs are identified through rules and industrial entity annotation set E. Relationship with model predictions χ i Constitute a positive sample pair Select two different arbitrary span segments from the current industrial text data as pseudo-entities E1 i′ 、E2 i′ , where for a small number of manually annotated industrial text datasets, the real annotated relationship labels are selected; for unlabeled industrial text datasets, the central representation R of the relationship cluster is selected i Constitute a negative sample pair Step 22: Based on χ i Compute the probability distribution of possible relationships between industrial entities, p i =In T (ReLU(χ i ))+b; Among them, W is the classification weight and b is the bias parameter; Step 23: Calculate the relationship category of the industrial entity in the industrial data context: y i =argmax(p i ); Step 24: Construct a classification cross entropy loss for annotating industrial text datasets. Step 25: Construct a relational queue set R = {R1, R2, ..., R o }, o represents the number of unlabeled categories, queue R i The size of is b·N, where b is the batch size, and the pseudo queue label is The positive sample relationship representation The comparison set is Step 26: Calculate the semantic similarity between the cluster center of the industrial text data and each relational queue, and minimize the semantic similarity based on the relational queue. and the classification probability p calculated based on the classification weight parameter i The cross entropy L between CO , Among them, τ is the temperature coefficient. After each round of training, the maximum likelihood estimation is used to update the industrial entity relationship in the industrial text data in the iteration cycle. The queue tag After each back propagation, Add to 2. A complex semantic relationship extraction method for industrial knowledge graph construction according to claim 1, characterized in that: In step 3, the process of dynamically optimizing the training process and using the converged model to infer the relationship between industrial entities in the industrial knowledge graph includes the following steps: Step 31: In adjacent training rounds, record the number of changes in the predicted category of industrial entity relationships in the industrial text data and assign weights w to the samples. i , collaboratively optimize instance-level contrast loss, in is the relationship between the industrial entities predicted in the kth round in the given industrial field text data, is i Adjust parameters in round k; Step 32, calculate the overall loss L, L=L CO +L EN +λ·L CE ; Among them, λ is a hyperparameter; Step 33: When L is less than the specified minimum loss value or the maximum number of training rounds is reached, the training is terminated. As the final result of the industrial entity relationship in the industrial domain text, otherwise repeat steps 13 to 32.
Citation Information
Patent Citations
Knowledge reasoning method based on industrial mechinery fault diagnosis knowledge graph
CN113961718A
Dialogue text relation extraction method based on knowledge assistance
CN117573878A