A large model-assisted knowledge graph construction method for self-lifting multimodal industrial equipment

Through the large-model self-improvement multimodal industrial equipment knowledge graph construction method, the problems of low efficiency and poor effectiveness of industrial equipment knowledge graph construction are solved, and high-quality knowledge graph construction is achieved, supporting fault diagnosis and process optimization of industrial equipment.

CN119577159BActive Publication Date: 2025-05-16ZHEJIANG UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510134451.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

In the prior art, the construction efficiency of industrial equipment knowledge graphs is low and the effect is poor, mainly due to the difficulty in uniform processing and integration of multi-source heterogeneous data, and the lack of deep understanding of domain knowledge and contextual reasoning capabilities, resulting in fuzzy and inaccurate semantics of the generated graphs.

Method used

The knowledge graph construction method of self-improvement multimodal industrial equipment assisted by large model is adopted. By collecting and preprocessing multimodal data, using few sample prompts for knowledge extraction, combining large model self-improvement technology to optimize triplet quality, perform entity alignment, and knowledge graph completion is carried out through retrieval enhancement generation method.

Benefits of technology

It realizes efficient and automated industrial equipment knowledge graph construction, and the generated map is of high quality, which can effectively support downstream tasks such as fault diagnosis and process optimization of industrial equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577159B_ABST
    Figure CN119577159B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a self-improving multimodal industrial equipment knowledge graph assisted by a large model, the method comprising: collecting and preprocessing multimodal data; obtaining structured knowledge data in the form of triples based on a few sample prompts; and then gradually self-optimizing the triples from three evaluation dimensions in combination with thinking chain skills to ensure the standardization, integrity and reliability of knowledge expression; using a triple-ladder entity alignment method and utilizing the self-improvement capability of a large model to solve data redundancy and ambiguity problems in the knowledge graph; based on a knowledge graph completion template and a retrieval enhancement mechanism, applying a large model to complete relationship prediction and entity completion tasks to ensure the integrity and applicability of the knowledge graph. The present invention does not require a large amount of manually annotated data, and through the self-improvement idea of ​​a large model and multiple iterative optimizations, it improves the automation level and knowledge quality of knowledge graph construction, providing strong support for industrial equipment management, process optimization and intelligent diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial production operation and maintenance technology, and in particular to a large-model-assisted self-lifting multimodal industrial equipment knowledge graph construction method. Background Art

[0002] During the operation and maintenance of industrial equipment, data sources such as equipment operation and maintenance sheets, introductory materials, and fault reports contain a large amount of expert experience and equipment domain knowledge. These data cover equipment operation specifications, fault diagnosis methods, maintenance strategies, etc. However, these data have not yet been effectively integrated, and there are serious data islands between them, which makes it impossible to fully utilize expert knowledge and affects the intelligent management and optimization of industrial equipment.

[0003] As an effective knowledge representation and management tool, knowledge graphs are widely used in knowledge domain visualization and knowledge domain mapping in the field of library and information science. Knowledge graphs graphically display the development process and structural relationship of knowledge, use visualization technology to describe knowledge resources and their carriers, and realize knowledge mining, analysis, construction, drawing and display as well as the mutual connection between knowledge. In the field of industrial equipment, knowledge graphs can integrate knowledge from multiple sources, including fault forms, industry manuals, operation and maintenance data, etc.; at the same time, they can be applied to fault diagnosis and process optimization of industrial equipment, and guide factory operation and maintenance personnel to provide knowledge assistance, risk warning and alarm intervention. Therefore, how to construct a high-quality knowledge graph in the field of industrial equipment has become an urgent problem to be solved.

[0004] In actual engineering applications, industrial knowledge graphs have problems such as low construction efficiency and poor application effects, which are mainly affected by the following aspects:

[0005] First, the construction of the knowledge graph requires the collection of relevant knowledge from multiple data sources. Due to the diversity of industrial equipment knowledge, the collected data is usually heterogeneous data from multiple sources, which places high demands on the unified processing and integration of data.

[0006] Secondly, the source of the collected data may not be reliable, and there are problems such as redundancy and duplication. Knowledge fusion is needed, that is, merging entities representing the same concept and merging data from different sources into a unified data set, while establishing the relationship between entities. The process of knowledge fusion requires semantic understanding and alignment of the data to ensure the consistency and accuracy of the knowledge graph finally constructed.

[0007] Third, general large models lack a deep understanding of domain knowledge and contextual reasoning capabilities when constructing knowledge graphs, resulting in the generated graphs being semantically ambiguous, inaccurate, or unable to meet the needs of specific scenarios.

[0008] Therefore, in order to address the problem of unified processing and integration of multi-source heterogeneous data in existing technologies, an efficient data processing and fusion method is urgently needed to solve the task of building an industrial equipment knowledge graph. Summary of the invention

[0009] The purpose of the present invention is to provide a large model-assisted self-enhancing multimodal industrial equipment knowledge graph construction method to address the deficiencies of the prior art. The present invention applies the Large Language Model (LLM) technology to the automatic and efficient construction of multimodal industrial equipment knowledge graphs, ultimately providing support for downstream tasks in the field of industrial equipment, including but not limited to fault diagnosis, process optimization, etc.

[0010] The object of the present invention is to achieve the following technical solution: a large model-assisted self-lifting multimodal industrial equipment knowledge graph construction method, comprising the following steps:

[0011] (1) Collect multimodal data of industrial equipment knowledge graph and preprocess it to obtain preprocessed multimodal data;

[0012] (2) Based on the few-sample prompts, knowledge is extracted from the preprocessed multimodal data to obtain industrial domain knowledge data in the form of triples; where the triple form is head entity-relationship-tail entity;

[0013] (3) Based on the self-improvement of the big model, the quality of the industrial domain knowledge data in the form of triples is optimized. Specifically, from multiple evaluation dimensions, the skills of the big model thinking chain are combined to gradually adjust the industrial domain knowledge data in the form of triples through a step-by-step optimization method to standardize, complete and accurately adjust them, so as to generate optimized high-quality triples;

[0014] (4) Performing knowledge graph entity alignment based on large model self-improvement, specifically: for the optimized high-quality triples, a three-step entity alignment method is used to align the entities to obtain the entity-aligned triples;

[0015] (5) Based on the triples after entity alignment, the multimodal data is formatted and imported into the graph database to construct a multimodal industrial equipment knowledge graph;

[0016] (6) For the multimodal industrial equipment knowledge graph, a retrieval-enhanced generation method is used to complete the knowledge graph based on the self-enhancement of the large model. Specifically, a prompt word template based on the knowledge graph completion template is designed to solve the three tasks of relationship prediction, head entity prediction and tail entity prediction; combined with the retrieval-enhanced generation method, the context is dynamically recalled and injected into the prompt word template, which is then input into the large model to obtain the completed multimodal industrial equipment knowledge graph.

[0017] Furthermore, the multimodal data of the industrial equipment knowledge graph is collected and preprocessed, specifically including:

[0018] First, multimodal data on industrial equipment is collected with the assistance of web crawlers and industrial equipment operation and maintenance personnel, wherein the multimodal data includes text modal data, image modal data, and time series table modal data;

[0019] The collected multimodal data are then preprocessed, specifically: the text and the corresponding title, pictures and the picture titles, the table numbers of the time series table files and the corresponding titles are organized into one-to-one structured files, and the text is segmented using a recursive segmentation strategy that combines characters and semantics to obtain the preprocessed multimodal data.

[0020] Furthermore, the text is segmented by using a recursive segmentation strategy combining characters and semantics, specifically including:

[0021] Recursively split by characters and semantics. First, set the maximum length of each text block to And set the character priority list for segmentation, the priority from high to low is: first level is chapter title, second level is paragraph, third level is sentence, fourth level is word; if the segmentation with higher priority does not meet the required text block size, the text block output after segmentation will be re-input into the segmenter and further segmented according to the segmentation requirements of the lower level;

[0022] After the character segmentation is completed, the semantic chunking method is applied to analyze the paragraphs or sentences of the initially segmented text blocks to obtain the semantic chunking paragraphs or sentences; the semantic chunking sentences or paragraphs are then input into the pre-trained large model to generate the corresponding embedding vector to represent the semantic information of each sentence or paragraph; the cosine similarity between adjacent sentences or paragraphs is calculated as the semantic similarity score , to evaluate the semantic coherence of adjacent sentences or paragraphs; based on the calculated semantic similarity score and the preset semantic similarity threshold , merge sentences or paragraphs with similar semantics to form semantically reorganized text segments;

[0023] For a semantically reorganized text segment, determine whether its length is greater than , if the length of the semantically reorganized text segment is greater than , the text segments that have not undergone semantic reorganization are retained; otherwise, the semantically reorganized text segments are merged again, and then the length is returned to determine whether it is greater than steps.

[0024] Furthermore, the knowledge extraction of the preprocessed multimodal data based on the few-sample prompt to obtain industrial domain knowledge data in the form of triples specifically includes:

[0025] Based on the preprocessed multimodal data, when processing its unstructured text, we first select sample text from the text and annotate it with corresponding structured triples; then, on this basis, we design prompt words based on few-sample learning to extract text knowledge, and specify the triple output format for the large model; then, we merge the examples and prompt words and input them into the large model. Through the sample input and corresponding output provided to the large model, we guide the large model to extract structured triples from the unstructured text and obtain industrial domain knowledge data in the form of triples.

[0026] Furthermore, the evaluation dimensions include three evaluation dimensions: grammatical standardization, semantic completeness and information accuracy; wherein the grammatical standardization is used to evaluate whether the entities and relations in the triples conform to the grammatical rules of natural language; the semantic completeness is used to evaluate whether the triples fully express a knowledge unit, that is, whether the three elements of the triples can convey a complete semantics; the information accuracy is used to evaluate whether the information contained in the triples is true and correct;

[0027] According to the three evaluation dimensions of grammatical standardization, semantic completeness and information accuracy, combined with the skills of the big model thinking chain, the optimization process is divided into multiple steps. Prompt words are designed for the big model, and then the prompt words and the industrial domain knowledge data in the form of triples to be optimized are input into the big model together. Combined with the self-improvement of the big model, the industrial domain knowledge data in the form of triples are gradually standardized, complete and accurate to ensure that the triple structure is complete, the semantic expression is clear, and the information is authentic and reliable, so as to obtain the final optimized high-quality triples output by the big model.

[0028] Furthermore, the optimized high-quality triples are aligned using a triple-step entity alignment method, specifically including:

[0029] The first entity alignment method is character-level entity alignment, the second entity alignment method is vector-level entity alignment, and the third entity alignment method is model-level entity alignment;

[0030] For the optimized high-quality triples, we first merge identical entities through character-level entity alignment, and use attributes to prevent incorrect merging of ambiguous entities; then, based on the character-level entity alignment, we use the embedding model to encode the entities into entity embedding vectors, calculate the similarity between the entity embedding vectors, and merge the entity pairs corresponding to the similarity greater than the preset entity similarity threshold to achieve vector-level entity alignment; finally, based on the vector-level entity alignment, we use the embedding model to obtain the entity embedding vector, use the unsupervised clustering algorithm to cluster the entity embedding vectors, and use the large model to judge similar entities based on the context of a few sample prompts and merge them.

[0031] Furthermore, the first entity alignment method is character-level entity alignment, which specifically includes:

[0032] For the optimized high-quality triples, first, define the character same relationship function , if and only if the entity and entities When the characters are exactly the same, ,otherwise ;

[0033] At the same time, define the polysemy judgment function , the attribute As a conditional input, if and only if the entity and entities Attributes When it is judged as "not polysemy", ,otherwise ;

[0034] Get the character-level entity alignment operation function , expressed as:

[0035]

[0036] in, represents the i-th entity in character form, Represents the jth entity in character form; Represents the merged entity, that is, when the entity and entities Characteristically identical and entity and entities Attributes When it is judged as "non-polysemous", the input entity and entities Merge to get the merged entity ; Otherwise, the input entity and entities No merging is performed. Retain entity and entities No merging is performed.

[0037] Furthermore, the second entity alignment method is a vector-level entity alignment, which specifically includes:

[0038] For the triples that have completed character-level entity alignment, we first use the embedding model to encode different entities into corresponding entity embedding vectors, expressed as:

[0039]

[0040] in, represents the entity embedding vector, Represents a character entity, represents the embedding model composed of the encoding layer; then the cosine similarity between two entities is calculated based on the entity embedding vector, which is expressed as:

[0041]

[0042] in, represents the cosine similarity between entity i and entity j, represents the entity embedding vector of entity i, represents the entity embedding vector of entity j, represents the two-norm;

[0043] Finally, the entity similarity threshold is set and entity pairs whose cosine similarity is greater than the entity similarity threshold are merged.

[0044] Furthermore, the third entity alignment method is model-level entity alignment, which specifically includes:

[0045] For the triplets that have completed the entity alignment at the vector level, we first use the embedding model to obtain the entity embedding vectors of different entities, expressed as:

[0046]

[0047] in, represents the entity embedding vector, Represents a character entity, Indicates that the pre-trained BERT model is selected as the embedding model;

[0048] Then, an unsupervised clustering algorithm is used to cluster the entity embedding vectors. Specifically, for the obtained entity embedding vectors , initialize and generate Initial cluster centers ,in represents the i-th entity embedding vector, n represents the total number of entity embedding vectors, represents the kth cluster center; assign each entity embedding vector to the nearest cluster center, and use the following formula to embed each entity vector Find the cluster center closest to it to embed each entity into a vector Assigned to the cluster corresponding to the cluster center:

[0049]

[0050] in, Represents entity embedding vector The cluster index to which it belongs; according to the allocation result, each cluster center is recalculated to update the cluster center. The updated cluster center is equal to the mean of all entity embedding vectors in the cluster, expressed as:

[0051]

[0052] in, represents the set of entity embedding vectors of the kth cluster, Represents the total number of entity embedding vectors in the kth cluster; repeat the two steps of assigning each entity embedding vector to the nearest cluster center and updating the cluster center until the preset maximum number of iterations is reached or the change value of the cluster center is less than the set change threshold, expressed as:

[0053]

[0054]

[0055] in, Represents clusters The distance between the new and old cluster centers, Represents clusters The new cluster center of the tth iteration, Represents clusters The old cluster center of the t-1th iteration, Indicates the set change threshold;

[0056] Finally, the clustered entity clusters are input into the large model, and the large model is guided to make similarity judgments by constructing a few-sample prompt based on context learning. Specifically, multiple samples are selected from the entity cluster to construct a prompt template, and these samples are input into the large model together with the current entity to be processed to guide the large model to judge which entities can be merged based on context information. By repeating this process, similar entities in the cluster are gradually screened and processed to obtain similar entities that can be merged in the entity cluster.

[0057] Furthermore, the combined retrieval enhancement generation method dynamically recalls the context and injects it into the prompt word template, which is then input into the large model to obtain a completed multimodal industrial equipment knowledge graph, specifically including:

[0058] Input contains and or and The query vector ;

[0059] In the knowledge sample library, search for vectors containing and or and Context documents; Generate query vectors using embedding models The embedding vector , and each context document The embedding vector , and calculate the query vector and each context document The similarity score between calculations is calculated as follows:

[0060]

[0061] in, Represents the query vector and context documents Calculate the similarity score between them; then filter out the top M context documents with the largest similarity scores, expressed as:

[0062]

[0063] in, Represents the filtered context document;

[0064] The context documents that will be filtered out Insert it into the prompt word template to get the final prompt word of the generated task;

[0065] The final prompt word of the generation task is input into the large model, and the large model outputs the completed multimodal industrial equipment knowledge graph in the form of head entity-relationship-tail entity. The generation process of the large model is represented by the following conditional probability model:

[0066]

[0067] in, represents the final output sequence of the large model, represents all possible output sequences, Indicates the final prompt word for the generated task, For modeling the generation process of the large model, the input is The output of the large model is probability.

[0068] The beneficial effects of the present invention are as follows: the present invention applies the large language model iteratively multiple times to gradually optimize the results of the previous stage, and efficiently and high-quality converts semi-structured and unstructured data into structured knowledge graphs; the present invention not only provides a solid foundation for downstream tasks in the field of industrial equipment such as fault diagnosis and process optimization, but also has important significance for the automatic construction of large-scale, high-quality, multi-modal knowledge graphs by large language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A flowchart of a method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model of the present invention;

[0070] Figure 2 is a flow chart of the architecture of the triple-step physical alignment method of the present invention;

[0071] Figure 3 A schematic diagram of the storage of multimodal nodes in the industrial equipment knowledge graph of the present invention;

[0072] Figure 4 This is an architectural flow chart of the present invention for completing the knowledge graph based on a large model using a retrieval enhancement generation method;

[0073] Figure 5 A schematic diagram of a partial knowledge graph extracted from a thermal power plant fault analysis report of the present invention;

[0074] Figure 6 It is a schematic diagram of the knowledge map extracted from the metallurgical process introduction teaching material of the present invention;

[0075] Figure 7 This is a schematic diagram of the knowledge map extracted from the chemical materials introduction textbook of the present invention. DETAILED DESCRIPTION

[0076] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0077] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0078] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0079] The present invention is described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the features of the following embodiments and implementations can be combined with each other.

[0080] The large-model-assisted self-enhancing multimodal industrial equipment knowledge graph construction method of the present invention integrates multimodal data, utilizes the information extraction, self-verification and self-optimization capabilities of the large language model, and gradually optimizes the results of the previous stage through multiple iterative applications of the large model to construct a high-quality knowledge graph.

[0081] Taking a data set from a real thermal power plant as an example, the large model-assisted self-lifting multimodal industrial equipment knowledge graph construction method of the present invention is described in detail. The process is as follows: Figure 1 As shown, the method specifically comprises the following steps:

[0082] (1) Collect multimodal data of industrial equipment knowledge graph and preprocess it to obtain preprocessed multimodal data.

[0083] Furthermore, the multimodal data of the industrial equipment knowledge graph is collected and preprocessed, specifically including: first, multimodal data with industrial equipment as the theme is collected through web crawlers and the assistance of industrial equipment operation and maintenance personnel, wherein the multimodal data includes text modal data, image modal data, time series table modal data and other multimodal data, the source of text modal data includes but is not limited to operation and maintenance texts such as operation and maintenance orders, equipment introduction materials, and fault reports in industrial fields such as thermal power plants, the source of image modal data includes but is not limited to industrial introduction materials in the industrial field, papers in the field of industrial equipment and other data with pictures, the source of time series table modal data includes but is not limited to factory operation and maintenance history databases, monthly fault reports, log data, etc. in the field of industrial equipment, and the sources of the above multimodal data include but are not limited to online forums, industrial textbooks, sensor data and papers, etc. Then, the collected multimodal data is preprocessed, taking the equipment introduction document as an example, the text and corresponding title in the equipment introduction document, the equipment picture and the picture title, the table number of the time series table file of the equipment parameters and the article title corresponding to the parameter table are all organized into one-to-one corresponding structured files, such as Figure 2 As shown, a recursive segmentation strategy combining characters and semantics is used to segment the text, ensuring that the text input into the large model has high-quality structure and semantic consistency, obtaining preprocessed multimodal data, and realizing the normalization of multimodal data.

[0084] Furthermore, a recursive segmentation strategy combining characters and semantics is used to segment the text into blocks, specifically including: recursive segmentation by characters and semantics, first setting the maximum length of each text block to And set the character priority list for segmentation, the priorities from high to low are: the first level is chapter title, the second level is paragraph, the third level is sentence, and the fourth level is word; if the higher priority segmentation does not meet the required text block size, the output text block after segmentation will be re-input into the segmenter and further segmented according to the lower level segmentation requirements. After the character segmentation is completed, the semantic chunking method is applied to analyze the paragraphs or sentences of the preliminary segmented text blocks to obtain the paragraphs or sentences after semantic chunking; then the sentences or paragraphs after semantic chunking are input into a pre-trained large model such as the BERT model to generate the corresponding embedding vector to represent the semantic information of each sentence or paragraph; the cosine similarity between adjacent sentences or paragraphs is calculated as the semantic similarity score , to evaluate the semantic coherence of adjacent sentences or paragraphs; based on the calculated semantic similarity score and the preset semantic similarity threshold , merge sentences or paragraphs with similar semantics to form semantically reorganized text segments. For the semantically reorganized text segments, determine whether their length is greater than , if the length of the semantically reorganized text segment is greater than , the text segments that have not undergone semantic reorganization are retained; otherwise, the semantically reorganized text segments are merged again, and then the length is returned to determine whether it is greater than steps.

[0085] It should be understood that semantic chunking is a natural language processing technique that processes and analyzes text by dividing it into smaller, more semantic chunks. Greater than or equal to the preset semantic similarity threshold , the adjacent sentences or paragraphs are considered to be semantically similar, and the semantically similar sentences or paragraphs are merged into a semantically reorganized text segment; otherwise, the adjacent sentences or paragraphs are considered to be not semantically similar and are not merged.

[0086] (2) Based on the few-sample prompts, knowledge is extracted from the preprocessed multimodal data to obtain industrial domain knowledge data in the form of triples. The triples are in the form of head entity-relationship-tail entity.

[0087] Furthermore, in the knowledge graph, data representing industrial domain knowledge is stored in the form of triples. Mathematically, triples are It can be expressed as:

[0088]

[0089] in, and Represent the head entity and the tail entity respectively, Represents the relationship between the head entity and the tail entity, and Represents a collection of entities and relations, respectively. For example, the triple (coal-burning emissions, environmental impact, severe pollution) represents the fact that "coal-burning emissions are a factor that has a serious impact on the environment."

[0090] Furthermore, knowledge extraction is performed on the preprocessed multimodal data based on few-sample prompts to obtain industrial domain knowledge data in the form of triples, specifically including: based on the preprocessed multimodal data, when processing its unstructured text, first select a small amount of sample text from the text and annotate it with corresponding structured triples; then on this basis, design appropriate prompt words based on few-sample learning to extract text knowledge, and specify the triplet output format for the large model; then merge the examples and prompt words and input them into the large model, and guide the large model to extract structured triples from unstructured text in the absence of large-scale annotated data through the example input and corresponding output provided to the large model, so as to obtain industrial domain knowledge data in the form of triples, which can effectively enhance the performance of the large model.

[0091] Furthermore, specifying the triple output format for the large model can facilitate subsequent automated processing. An example is: "The output format of the triple is: head entity-relationship-tail entity. Different triples are separated by '\n'. Please strictly follow the specified format for output, output in Chinese, and do not add numbers or other content." This formatted output reduces the uncertainty of large model generation and reduces the burden of the large model's attention mechanism.

[0092] It should be noted that the large model-assisted self-enhancing multimodal industrial equipment knowledge graph construction method described in the present invention is based on the initial structured triples obtained through steps (1) and (2), and utilizes the information extraction ability and self-verification and optimization ability of the large model to further self-enhance triple quality optimization, entity alignment and knowledge graph completion, see steps (3), (4) and (6) respectively.

[0093] (3) Based on the self-improvement of the big model, the quality of the industrial domain knowledge data in the form of triples is optimized. Specifically, from multiple evaluation dimensions, combined with the skills of the big model thinking chain, the industrial domain knowledge data in the form of triples is gradually standardized, completed and accurately adjusted through a step-by-step optimization method to ensure that the triple structure is complete, the semantic expression is clear, and the information is authentic and reliable, so as to generate optimized high-quality triples.

[0094] Furthermore, in order to effectively improve the quality of industrial domain knowledge data in the form of triples, three evaluation dimensions are designed to evaluate the quality of optimized triples, including grammatical standardization, semantic completeness, and information accuracy. Among them, grammatical standardization is used to evaluate whether the entities and relationships in the triples conform to the grammatical rules of natural language; for example, "coal-fired power generation process-application field-thermal power generation process" is compared with "coal-fired power generation process-application field-thermal power", thermal power can refer to the energy form of thermal power, but in the context of this triple knowledge, thermal power generation process is the field. Semantic completeness is used to evaluate whether the triple fully expresses a knowledge unit, that is, whether the three elements of the triple can convey a complete semantics; for example, "boiler-main equipment-steam generator" is more semantically complete than "boiler-related equipment-steam generator" because the scope of "main equipment" is smaller and conveys a clearer meaning. Information accuracy is used to evaluate whether the information contained in the triplet is true and correct; for example, "coal-fired power generation-fuel source-coal" has higher accuracy than "coal-fired power generation-fuel source-natural gas" because the former is consistent with the actual fuel source of coal-fired power generation, while the latter is wrong information and cannot reflect the actual fuel composition of coal-fired power generation.

[0095] Furthermore, for the three evaluation dimensions of grammatical standardization, semantic completeness and information accuracy, combined with the skills of the big model thinking chain, the optimization process is divided into multiple steps to design prompt words for the big model. The designed prompt words are as follows: "You are an expert in the field of thermal power. Please optimize the triples from the perspective of grammatical standardization. You need to identify the language expressions of entities and relationships in the triples to ensure that the entity names and relationship descriptions conform to the standard language grammatical rules. Please follow my steps to optimize the provided triples step by step:

[0096] In the first step, the entity name is checked for completeness and formality;

[0097] The second step is to check whether the relation words clearly express the semantic relationship between entities;

[0098] The third step is to adjust the non-standard or ambiguous relationship expression according to the context. "After that, the prompt words and the industrial domain knowledge data in the form of triples to be optimized are input into the big model together. Combined with the self-improvement of the big model, the industrial domain knowledge data in the form of triples are gradually standardized, complete and accurate to ensure that the triple structure is complete, the semantic expression is clear, and the information is authentic and reliable, so as to obtain the final optimized high-quality triples output by the big model.

[0099] (4) Perform knowledge graph entity alignment based on large model self-improvement. Specifically, for the optimized high-quality triples, a three-step entity alignment method is used to align the entities to obtain the entity-aligned triplets.

[0100] Specifically, Figure 2 As shown, the three-step entity alignment method specifically includes: the first entity alignment method is character-level entity alignment, the second entity alignment method is vector-level entity alignment, and the third entity alignment method is model-level entity alignment. For the optimized high-quality triples, firstly, through character-level entity alignment, identical entities are merged, and attributes are used to prevent the incorrect merging of polysemous entities; then, based on the character-level entity alignment, the embedding model is used to encode the entities into entity embedding vectors, and the similarity between the entity embedding vectors is calculated, and the entity pairs corresponding to the similarity greater than the preset entity similarity threshold are merged to achieve vector-level entity alignment; finally, based on the vector-level entity alignment, the embedding model is used to obtain the entity embedding vector, the entity embedding vector is clustered using an unsupervised clustering algorithm, and similar entities are merged based on the context judgment of the large model based on the few sample prompts. The efficiency of entity alignment can be effectively improved through multi-level alignment strategies.

[0101] It should be understood that the description of a node may have multiple dimensions, such as node type and node attributes. Node attributes include name, model, etc. Node attributes can prevent incorrect merging of ambiguous entities.

[0102] Furthermore, the first entity alignment method is character-level entity alignment, which specifically includes: for the optimized high-quality triples, first, define the character same relationship function , if and only if the entity and entities When the characters are exactly the same, ,otherwise ,Right now:

[0103]

[0104] in, represents the i-th entity in character form, represents the jth entity in character form, Representing Entities Characters, Representing Entities At the same time, define a polysemy judgment function , the attribute As a conditional input, if and only if the entity and entities Attributes When it is judged as "not polysemy", ,otherwise In summary, the character-level entity alignment operation function It is expressed as:

[0105]

[0106] in, Represents the merged entity, that is, when the entity and entities Characteristically identical and entity and entities Attributes When it is judged as "non-polysemous", the input entity and entities Merge to get the merged entity ; Otherwise, the input entity and entities No merging is performed. Retain entity and entities No merging is performed. The first entity alignment method can merge entities that are exactly the same in terms of characters into one entity. Adding attributes to nodes with multiple meanings can effectively prevent entities with the same characters but different meanings from being merged.

[0107] Furthermore, the second entity alignment method is vector-level entity alignment, which specifically includes: for the triples that have completed character-level entity alignment, firstly use the embedding model to encode different entities into corresponding entity embedding vectors, expressed as:

[0108]

[0109] in, represents the entity embedding vector, Represents a character entity, Represents the embedding model composed of the encoding layer. Then the cosine similarity between two entities is calculated based on the entity embedding vector, which is expressed as:

[0110]

[0111] in, represents the cosine similarity between entity i and entity j, represents the entity embedding vector of entity i, represents the entity embedding vector of entity j, Represents the bi-norm. Finally, set the entity similarity threshold and merge entity pairs whose cosine similarity is greater than the entity similarity threshold. In summary, the vector-level entity alignment function It can be expressed as:

[0112]

[0113] in, Represents the entity similarity threshold, that is, when the entity and entities The cosine similarity between Greater than the preset entity similarity threshold When the input entity and entities Merge to get the merged entity ; Otherwise, the input entity and entities No merging is performed.

[0114] Furthermore, the third entity alignment method is model-level entity alignment, which specifically includes: for the triples that have completed the vector-level entity alignment, firstly use the embedding model such as the BERT model to obtain the entity embedding vectors of different entities, expressed as:

[0115]

[0116] in, represents the entity embedding vector, Represents a character entity, Indicates that the pre-trained BERT model is used as the embedding model. Of course, the embedding model can also use models such as RoBERTa, T5, GPT series, as well as methods based on word bags and distributed word vectors to obtain entity embedding vectors. Then, an unsupervised clustering algorithm is used to cluster the entity embedding vectors. Specifically, for the obtained entity embedding vectors , initialize and generate Initial cluster centers ,in represents the i-th entity embedding vector, n represents the total number of entity embedding vectors, represents the kth cluster center; assign each entity embedding vector to the nearest cluster center, and use the following formula to encode each embedding vector Find the cluster center closest to it to embed each entity into a vector Assigned to the cluster corresponding to the cluster center:

[0117]

[0118] in, Represents entity embedding vector The cluster index to which it belongs; according to the allocation result, each cluster center is recalculated to update the cluster center. The updated cluster center is equal to the mean of all entity embedding vectors in the cluster, expressed as:

[0119]

[0120] in, represents the set of entity embedding vectors of the kth cluster, Represents the total number of entity embedding vectors in the kth cluster; repeat the two steps of assigning each entity embedding vector to the nearest cluster center and updating the cluster center until the preset maximum number of iterations is reached or the change value of the cluster center is less than the set change threshold, expressed as:

[0121]

[0122]

[0123] in, Represents clusters The distance between the new and old cluster centers, Represents clusters The new cluster center of the tth iteration, Represents clusters The old cluster center of the t-1th iteration, Represents the set change threshold. Finally, the clustered entity clusters are input into the large model, and the large model is guided to make similarity judgments by constructing a few-sample prompt based on context learning. Specifically, multiple samples are selected from the entity cluster to construct a prompt template, and these samples are input into the large model together with the current entity to be processed to guide the large model to judge which entities can be merged according to the context information. By repeating this process, similar entities in the clusters are gradually screened and processed, and similar entities that can be merged in the entity clusters are obtained, which facilitates the further merging and optimization of entities.

[0124] It should be understood that in this embodiment, an existing commonly used unsupervised clustering algorithm can be used to cluster the entity embedding vectors, such as a K-means clustering algorithm.

[0125] (5) Based on the triples after entity alignment, the multimodal data is formatted and imported into the graph database to build a large-scale multimodal industrial equipment knowledge graph to support downstream applications such as process optimization and fault diagnosis.

[0126] Furthermore, if Figure 3 As shown in the figure, the multimodal data is formatted and imported into the graph database to build a large-scale multimodal industrial equipment knowledge graph, which includes: for the image data, it is structured and organized into file, with the title as the node name (i.e. Figure 3 The entity name in the ), the URL source address as an attribute (i.e. Figure 3 Entity attributes in ); for tabular data, it is structured and node names are generated by row (i.e. Figure 3 Entity names in ), column contents as attributes (i.e. Figure 3 Entity attributes in ); For time series data, it is structured and features are extracted by frequency, time and statistical features. The fault representation is used as the node name, and the extracted features are used as attributes (i.e. Figure 3 Then, a large-scale multimodal industrial equipment knowledge graph is constructed based on the node names and corresponding attributes of the multimodal data to support downstream applications such as process optimization and fault diagnosis.

[0127] (6) For the multimodal industrial equipment knowledge graph, a retrieval enhancement generation method is used to complete the knowledge graph based on the self-enhancement of the large model. Specifically, a prompt word template based on the knowledge graph completion template is designed to solve the three tasks of relationship prediction, head entity prediction and tail entity prediction; combined with the retrieval enhancement generation method, the context is dynamically recalled and injected into the prompt word template, which is then input into the large model to obtain the completed multimodal industrial equipment knowledge graph, such as Figure 4 shown.

[0128] It should be understood that step (5) is to import the titled pictures and table nodes into the graph database. In the corresponding multimodal industrial equipment knowledge graph, they are still isolated nodes. After the knowledge graph is completed in step (6), these multimodal nodes and the text nodes that account for the majority of the number will form a complete knowledge graph after establishing relationships.

[0129] In this embodiment, the knowledge graph completion task is divided into two tasks: relationship prediction task and entity prediction task. Among them, the relationship prediction task is specifically: provide the head entity and the tail entity to the big model, the big model determines the relationship between the two entities, and outputs the form of "head entity-relationship-tail entity". The entity prediction task is divided into the head entity prediction task and the tail entity prediction task. The head entity prediction task is specifically: provide the tail entity and relationship to the big model, the big model generates the head entity, and outputs the form of "head entity-relationship-tail entity"; the tail entity prediction task is specifically: provide the head entity and relationship to the big model, the big model generates the tail entity, and outputs the form of "head entity-relationship-tail entity".

[0130] Furthermore, by designing a prompt word template based on the knowledge graph completion template, we can solve the three tasks of relationship prediction, head entity prediction and tail entity prediction. The three sets of prompt word templates corresponding to these three tasks are:

[0131] ① The example of the prompt word template corresponding to the relationship prediction task is: "The content in the square brackets is and The context of [the context of the recall], please answer the questions according to my requirements,' and What is the relationship between 'Please generate the most appropriate relationship based on the context, and the output is' ' form. ".

[0132] ② The example of the prompt word template corresponding to the head entity prediction task is: "The content in the square brackets is and The context of [recall context], please answer the question according to my request, 'Which entity is and Header entity? 'Please generate the most appropriate head entity based on the context, and output it as' ' form. ".

[0133] ③ The prompt word template example corresponding to the tail entity prediction task is: "The content in the square brackets is and The context of [recall context], please answer the question according to my request, 'Which entity is and tail entity? 'Please generate the most appropriate tail entity based on the context, and output it as' ' form. ".

[0134] In this embodiment, the contextual content required for recall in the three knowledge graph completion task prompt word templates is achieved through a retrieval enhancement generation method, such as Figure 4 As shown in the figure, combined with the retrieval enhancement generation method, the context is dynamically recalled and injected into the prompt word template, which is then input into the large model to obtain the completed multimodal industrial equipment knowledge graph, which specifically includes the following steps:

[0135] (6.1) Input contains and or and The query vector .

[0136] (6.2) Search the knowledge sample library by vector and or and context document; using embedding models such as Or BERT model to generate query vector The embedding vector , and each context document The embedding vector , and calculate the query vector and each context document The similarity score between calculations is calculated as follows:

[0137]

[0138] in, Represents the query vector and context documents Calculate the similarity score between them; then select the top M context documents with the largest similarity scores as the selected M most relevant context documents, expressed as:

[0139]

[0140] in, Represents the filtered context documents. It is easy to understand that these filtered context documents are related to the query vector The most relevant document is the context of dynamic recall.

[0141] (6.3) The filtered context documents Insert it into the prompt word template to get the final prompt word for the generated task.

[0142] (6.4) The final prompt word of the generated task is input into the large model, and the large model outputs the generated head entity-relationship-tail entity form of the completed multimodal industrial equipment knowledge graph; the generation process of the large model is represented by the following conditional probability model:

[0143]

[0144] in, represents the final output sequence of the large model, represents all possible output sequences, Indicates the final prompt word for the generated task, For modeling the generation process of the large model, the input is The output of the large model is probability.

[0145] The multimodal node insertion, triple extraction, entity alignment, and knowledge graph completion tasks described in the present invention can be used in any industrial equipment knowledge graph construction process, and the knowledge graph construction method described in the present invention can be used for any industrial process object that has a knowledge graph construction demand in fields including but not limited to thermal power, energy, metallurgy, injection molding, chemical industry, etc. The present invention lists three examples including the construction of a knowledge graph with data from a real thermal power plant, the construction of a knowledge graph with data from metallurgical powder sintering, and the construction of a knowledge graph with data from plastic injection molding to verify the effect of the knowledge graph construction method described in the present invention.

[0146] Example 1: Knowledge graph of industrial equipment in thermal power plants

[0147] In this embodiment, the task of constructing a knowledge graph of a thermal power plant using operation and maintenance texts such as an operation and maintenance form, equipment introduction, and fault report of a high-pressure heater of a thermal power plant industrial equipment is taken as an example to illustrate the effectiveness of the method described in the present invention.

[0148] As a key equipment in the thermal system, the high-pressure heater of a thermal power plant undertakes the important functions of improving thermal efficiency, reducing fuel consumption and optimizing heat energy distribution. However, in the long-term operation process, the high-pressure heater often faces risk events such as steam leakage, water hammer, and pipeline corrosion, which may lead to reduced heating efficiency, equipment damage, heat supply interruption, and even threaten the safe operation of the entire unit. The early warning log and fault report of the high-pressure heater record a wealth of event data, which includes the evolution process of the event and its causal relationship, providing an important basis for analyzing the propagation law of the event, predicting the fault trend, and formulating preventive and emergency measures.

[0149] In this example, based on the above text data, 17 fault analysis reports of high-pressure heaters in thermal power plants were used as examples to extract knowledge using the Zhipu ChatGLM model, and a knowledge graph containing 1,773 triples, 2,120 entities, and 2,537 relationships was obtained. This knowledge graph effectively presents the logical relationship between the operation and faults of thermal power plant equipment, such as Figure 5 It provides important data and method support for event diagnosis, fault prediction and optimization decision of high-voltage heater.

[0150] First, according to the above step (1), the collected operation and maintenance texts such as the operation and maintenance order, equipment introduction, fault report, etc. of the high-voltage heater are preprocessed, and the recursive segmentation strategy combining characters and semantics is performed, and the multimodal data is structured, finally obtaining a text block with high-quality structure and semantic consistency.

[0151] Furthermore, for the above semi-structured text blocks, a large language model knowledge extraction is performed based on a few sample prompts to obtain industrial domain knowledge in the form of triples. Examples of prompt words for knowledge extraction are as follows:

[0152] "Please fully extract the entities and relations in the text below to form triples for constructing a knowledge graph. Different triples are separated by '\n'. Please strictly follow the specified format for output, the output is in Chinese, and do not add other content such as numbers." Sample input: (previous separator) High-pressure heater is a key auxiliary equipment to improve the thermal economy and overall operating efficiency of thermal power generating units. Its main function is to preheat the feed water with medium and high-pressure steam extracted from the steam turbine, thereby raising its temperature before the feed water enters the boiler, reducing subsequent fuel consumption and enhancing the thermodynamic efficiency of the power generation process. High-pressure heaters are usually made of high-pressure and high-temperature resistant alloy materials. Their application history can be traced back to the mid-20th century, and they are constantly optimized with the continuous progress of new material science, advanced manufacturing processes and automatic control technologies. (After separator) Sample output: (Before separator) High-pressure heater-is-a key auxiliary equipment to improve the thermal economy and overall operating efficiency of thermal power generating units' / n'High-pressure heater-uses-medium and high-pressure steam extracted from the steam turbine to preheat the feed water' / n'High-pressure heater-is used to-raise the temperature of the feed water before it enters the boiler' / n'High-pressure heater-reduces-subsequent fuel consumption' / n'High-pressure heater-enhances-the thermodynamic efficiency of the power generation process' / n'High-pressure heater-is made of-high-pressure resistant high-temperature alloy materials' / n'High-pressure heater-can be traced back to-the mid-20th century' / n'High-pressure heater-is continuously optimized with-the continuous progress of new material science, advanced manufacturing processes and automatic control technology (after separator). The text content to be extracted is as follows: [text block of triples to be extracted]. ". The large model extracts triples based on the prompt word, and the results of some extracted triples are shown in Table 1.

[0153] Table 1: Some triples extracted from the high-voltage heater warning report

[0154]

[0155] Furthermore, the triples extracted above are optimized for triple quality based on the self-improvement of the large language model according to the above step (3). The triples initially extracted by the large model are further structured, semantically clarified, and information verified.

[0156] Furthermore, based on the triple-step enhanced entity alignment method of the large model self-lifting in the above step (4), the same entity concepts are merged and unified into one identification. An example of the third entity alignment prompt is as follows: "I will give you a list of entities. The entities in the list are separated by commas and marked with single quotes. Please judge the similarity of the entities based on the similarity, and find no more than three entities that are similar in pairs. Output them separated by commas. Do not output redundant information. If there are no similar entities, output null. Please note that the output only contains commas and similar entities. Do not add other content. The following are input examples: ['High-pressure heater maintenance process', 'High-pressure heater operating procedures', 'Improve feed water temperature efficiency', 'High-pressure heater maintenance work instructions', 'New alloy heat transfer material', 'The role of high-pressure heater in steam-water circulation system', 'Automated monitoring and sensing system', 'High-pressure heater performance test method', 'Adjustable high-pressure heating element', 'High-pressure heater operating procedures manual'], output examples: High-pressure heater operating procedures, High-pressure heater operating procedures manual. The list content that you need to process is as follows: [clustered entity clusters]".

[0157] Furthermore, according to step (5)-step (6), the structured multimodal nodes obtained by step (1) and the text nodes after entity alignment in step (4) are inserted into the graph database; after recalling the required context through the retrieval enhancement generation step and inputting the head and tail entities or entities and relations to be completed, the completed knowledge graph is obtained by self-enhancement of the large language model.

[0158] Example 2: Metallurgical process equipment knowledge graph

[0159] Metallurgical technology is a technical science that extracts and processes metals from ores or other metal-bearing resources through physical, chemical and thermodynamic processes. It covers smelting, refining and forming, and its core is to optimize metal properties to meet industrial needs. In this example, three textbooks on metallurgical plant processes were used as examples to extract 824 triples using the Zhipu ChatGLM model, and a metallurgical process knowledge graph was constructed to show the structured knowledge association from concepts to process flows. Some entities and relationships were visualized as follows: Figure 6 As shown, some of the triples extracted in this embodiment are shown in Table 2. This case effectively explores the technical method of converting metallurgical process text into knowledge graphs, and provides new ideas for industrial process analysis and optimization.

[0160] Table 2: Some triples extracted from the metallurgical plant process introduction textbook

[0161]

[0162] Example 3: Knowledge graph of injection molding process equipment

[0163] The injection molding process is a technology that heats and melts polymer materials, injects them into molds, and cools and solidifies them into products. It covers material selection, mold design, and process optimization. Its core is to use mechanics, thermodynamics, and fluid mechanics to achieve precise molding and performance optimization of materials. In this embodiment, 1539 triples were extracted based on the Zhipu ChatGLM model using 4 injection molding process material textbooks as an example, and an injection molding process knowledge map was constructed. The system shows the relationship between the process flow and key elements, and some entities and relationships are visualized as follows: Figure 7 As shown, some of the triples extracted in this embodiment are shown in Table 3, which provides structured knowledge support for the research and optimization of injection molding process.

[0164] Table 3: Some triples extracted from the injection molding process introduction textbook

[0165]

[0166] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model-assisted self-lifting multimodal industrial equipment knowledge graph construction method, characterized in that: The following steps are involved: (1) Collect multimodal data of industrial equipment knowledge graph and preprocess it to obtain preprocessed multimodal data; (2) extracting knowledge from the preprocessed multimodal data based on few-sample prompts to obtain industrial domain knowledge data in the form of triples; where the triples are in the form of head entity-relationship-tail entity; The method of extracting knowledge from the preprocessed multimodal data based on the few-sample prompt to obtain industrial domain knowledge data in the form of triples specifically includes: Based on the preprocessed multimodal data, when processing its unstructured text, first select sample text from the text and annotate the corresponding structured triples for it; then, based on this, design prompt words based on few-sample learning to extract text knowledge, and specify the triple output format for the large model; then merge the sample and prompt words and input them into the large model, and guide the large model to extract structured triples from the unstructured text through the sample input and corresponding output provided to the large model, and obtain industrial domain knowledge data in the form of triples; (3) Based on the self-improvement of the big model, the quality of the industrial domain knowledge data in the form of triples is optimized. Specifically, from multiple evaluation dimensions, the skills of the big model thinking chain are combined to gradually standardize, complete and accurately adjust the industrial domain knowledge data in the form of triples through a step-by-step optimization method to generate optimized high-quality triples; (4) Performing knowledge graph entity alignment based on large model self-boosting, specifically: for the optimized high-quality triples, a three-step entity alignment method is used to perform entity alignment to obtain entity-aligned triples; (5) Based on the triples after entity alignment, the multimodal data is formatted and imported into the graph database to construct a multimodal industrial equipment knowledge graph; (6) For the multimodal industrial equipment knowledge graph, a retrieval-enhanced generation method is used to complete the knowledge graph based on the self-enhancement of the large model. Specifically, a prompt word template based on the knowledge graph completion template is designed to solve the three tasks of relationship prediction, head entity prediction and tail entity prediction; combined with the retrieval-enhanced generation method, the context is dynamically recalled and injected into the prompt word template, which is then input into the large model to obtain the completed multimodal industrial equipment knowledge graph.

2. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 1 is characterized in that: The multimodal data of the industrial equipment knowledge graph is collected and preprocessed, specifically including: First, multimodal data on industrial equipment is collected with the assistance of web crawlers and industrial equipment operation and maintenance personnel, wherein the multimodal data includes text modal data, image modal data, and time series table modal data; The collected multimodal data are then preprocessed, specifically: the text and the corresponding title, pictures and the picture titles, the table numbers of the time series table files and the corresponding titles are organized into one-to-one structured files, and the text is segmented using a recursive segmentation strategy that combines characters and semantics to obtain the preprocessed multimodal data.

3. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 2 is characterized in that: The recursive segmentation strategy combining characters and semantics is used to segment the text, specifically including: Recursively split by characters and semantics. First, set the maximum length of each text block to L. max And set the character priority list for segmentation, the priority from high to low is: first level is chapter title, second level is paragraph, third level is sentence, fourth level is word; if the segmentation with higher priority does not meet the required text block size, the text block output after segmentation will be re-input into the segmenter and further segmented according to the segmentation requirements of the lower level; After the character segmentation is completed, the semantic chunking method is applied to analyze the paragraphs or sentences of the initially segmented text blocks to obtain the paragraphs or sentences after semantic chunking; the sentences or paragraphs after semantic chunking are then input into the pre-trained large model to generate the corresponding embedding vector to represent the semantic information of each sentence or paragraph; the cosine similarity between adjacent sentences or paragraphs is calculated as the semantic similarity score S sim , to evaluate the semantic coherence of adjacent sentences or paragraphs; according to the calculated semantic similarity score S sim and the preset semantic similarity threshold T sim , merge sentences or paragraphs with similar semantics to form semantically reorganized text segments; For the semantically reorganized text segment, determine whether its length is greater than L max , if the length of the semantically reorganized text segment is greater than L max , then retain the text segment that has not undergone semantic reorganization; otherwise, continue to merge the semantically reorganized text segment, and then return to determine whether its length is greater than L max steps.

4. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 1 is characterized in that: The evaluation dimensions include grammatical standardization, semantic completeness and information accuracy. The grammatical standardization is used to evaluate whether the entities and relations in the triples conform to the grammatical rules of natural language. The semantic completeness is used to evaluate whether the triples fully express a knowledge unit, that is, whether the three elements of the triples can convey a complete semantics. The information accuracy is used to evaluate whether the information contained in the triples is true and correct. According to the three evaluation dimensions of grammatical standardization, semantic completeness and information accuracy, combined with the skills of the big model thinking chain, the optimization process is divided into multiple steps. Prompt words are designed for the big model, and then the prompt words and the industrial domain knowledge data in the form of triples to be optimized are input into the big model together. Combined with the self-improvement of the big model, the industrial domain knowledge data in the form of triples are gradually standardized, complete and accurate to ensure that the triple structure is complete, the semantic expression is clear, and the information is authentic and reliable, so as to obtain the final optimized high-quality triples output by the big model.

5. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 1 is characterized in that: For the optimized high-quality triples, a triple-step entity alignment method is used to perform entity alignment, specifically including: The first entity alignment method is character-level entity alignment, the second entity alignment method is vector-level entity alignment, and the third entity alignment method is model-level entity alignment; For the optimized high-quality triples, we first merge identical entities through character-level entity alignment, and use attributes to prevent incorrect merging of ambiguous entities; then, based on the character-level entity alignment, we use the embedding model to encode the entities into entity embedding vectors, calculate the similarity between the entity embedding vectors, and merge the entity pairs corresponding to the similarity greater than the preset entity similarity threshold to achieve vector-level entity alignment; finally, based on the vector-level entity alignment, we use the embedding model to obtain the entity embedding vector, use the unsupervised clustering algorithm to cluster the entity embedding vectors, and use the large model to judge similar entities based on the context of a few sample prompts and merge them.

6. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 5 is characterized in that: The first entity alignment method is character-level entity alignment, which specifically includes: For the optimized high-quality triples, first, define the character same relationship function C (Entity i ,Entity j ), if and only if the entity Entity i and Entity j When the characters are identical, C(Entity i ,Entity j )=1, otherwise C(Entity i ,Entity j )=0; At the same time, we define a polysemy judgment function M(Entity i ,Entity j ,P), taking attribute P as the condition input, if and only if entity Entity i and Entity j When attribute P determines that a word is not polysemous, M(Entity i ,Entity j ,P)=1, otherwise M(Entity i ,Entity j ,P)=0; Get the character-level entity alignment operation function A1 (Entity i ,Entity j ), expressed as: Among them, Entity i Represents the i-th entity in character form, Entity j Represents the jth entity in character form; Entity merged Represents the merged entity, that is, when the entity Entity i and Entity j Character-identical and entity i and Entity j When attribute P determines that a word is not polysemous, the input entity Entity i and Entity j Merge to get the merged entity Entity merged ; Otherwise, the input entity Entity i and Entity j No merging is performed, Entity i ,Entity j remain means retaining the entity i and Entity j No merging is performed.

7. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 5 is characterized in that: The second entity alignment method is vector-level entity alignment, which specifically includes: For the triples that have completed character-level entity alignment, we first use the embedding model to encode different entities into corresponding entity embedding vectors, expressed as: v=Embedding(Entity) Among them, v represents the entity embedding vector, Entity represents the entity in character form, and Embedding() represents the embedding model composed of the encoding layer; then the cosine similarity between two entities is calculated based on the entity embedding vector, which is expressed as: Among them, S ij represents the cosine similarity between entity i and entity j, v i The entity embedding vector v represents entity i. j represents the entity embedding vector of entity j, ‖·‖ represents the bi-norm; Finally, the entity similarity threshold is set and entity pairs whose cosine similarity is greater than the entity similarity threshold are merged.

8. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 5 is characterized in that: The third entity alignment method is model-level entity alignment, which specifically includes: For the triplets that have completed the entity alignment at the vector level, we first use the embedding model to obtain the entity embedding vectors of different entities, expressed as: v=BERT(Entity) Where v represents the entity embedding vector, Entity represents the entity in character form, and BERT() represents the use of the pre-trained BERT model as the embedding model; Then, an unsupervised clustering algorithm is used to cluster the entity embedding vectors. Specifically, for the obtained entity embedding vector v = {v1, v2, ..., v i ,…,v n }, initialize to generate K initial cluster centers μ1,μ2,…,μ k ,…,μ K , where v i represents the i-th entity embedding vector, n represents the total number of entity embedding vectors, μ k represents the kth cluster center; assign each entity embedding vector to the nearest cluster center, and use the following formula to embed each entity vector v i Find the cluster center with the closest distance to embed each entity into vector v i Assigned to the cluster corresponding to the cluster center: c i =argmin k ||v i -μ k || 2 Among them, c i Represents the entity embedding vector v i The cluster index to which it belongs; according to the allocation result, each cluster center is recalculated to update the cluster center. The updated cluster center is equal to the mean of all entity embedding vectors in the cluster, expressed as: Among them, C k represents the set of entity embedding vectors of the kth cluster, |C k ∣ represents the total number of entity embedding vectors in the kth cluster; the two steps of assigning each entity embedding vector to the nearest cluster center and updating the cluster center are repeated until the preset maximum number of iterations is reached or the change value of the cluster center is less than the set change threshold, which is expressed as: Among them, Δ j Represents cluster c j The distance between the new and old cluster centers, Represents cluster c j The new cluster center of the tth iteration, Represents cluster c j The old cluster center of the t-1th iteration, ∈ represents the set change threshold; Finally, the clustered entity clusters are input into the large model, and the large model is guided to make similarity judgments by constructing a few-sample prompt based on context learning. Specifically, multiple samples are selected from the entity cluster to construct a prompt template, and these samples are input into the large model together with the current entity to be processed to guide the large model to judge which entities can be merged based on context information. By repeating this process, similar entities in the cluster are gradually screened and processed to obtain similar entities that can be merged in the entity cluster.

9. The method for constructing a knowledge graph of self-lifting multimodal industrial equipment assisted by a large model according to claim 1 is characterized in that: The combined retrieval enhancement generation method dynamically recalls the context and injects it into the prompt word template, which is then input into the large model to obtain the completed multimodal industrial equipment knowledge graph, specifically including: Input contains <entity1>and <entity2>or <entity1>and <relation> The query vector x of< / relation> In the knowledge sample library, search for vectors containing <entity1>and <entity2>or <entity1>and <relation>The context document of query vector x; use the embedding model to generate the embedding vector v of query vector x x , and each context document d r The embedding vector And calculate the query vector x and each context document d r The similarity score between the calculations is calculated as follows:< / relation> Among them, Sim(x,d r ) represents the query vector x and the context document d r Calculate the similarity score between them; then filter out the top M context documents with the largest similarity scores, expressed as: C * =top M {Sim(x,d r )∣ <entity1> ∈d r , <entity2> ∈d r}< / entity2> < / entity1> Among them, C * Represents the filtered context documents; The filtered context document C * Insert it into the prompt word template to get the final prompt word of the generated task; The final prompt word of the generation task is input into the large model, and the large model outputs the completed multimodal industrial equipment knowledge graph in the form of head entity-relationship-tail entity. The generation process of the large model is represented by the following conditional probability model: Among them, y represents the final output sequence of the large model, y′ represents all possible output sequences, T(x,C * ) represents the final prompt word of the generation task, p(y′|T(x,C * )) is the modeling of the generation process of the large model, indicating that the input is T(x,C * ) is the probability that the output of the large model is y′.

Citation Information

Patent Citations

  • Intelligent media asset multi-modal knowledge graph management method based on large model

    CN119294506A