Industrial innovation knowledge graph dynamic construction method based on large language model
Through a method based on a large language model, multi-source data is integrated and entity relationship extraction and attribute extraction are performed, the problems of low efficiency and poor generalization ability of traditional methods when building and updating industrial innovation knowledge graphs are solved, and efficient and comprehensive knowledge graph construction and update are achieved.
Patent Information
- Application Number
- CN202510637149.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional methods have low efficiency, poor generalization ability and limited coverage of long-tail knowledge when building and updating industrial innovation knowledge graphs.
A dynamic construction method based on a large language model is adopted, and multi-source data is collected and integrated, and a triple extraction model is used to perform entity relationship extraction and attribute extraction, and data quality is optimized through entity alignment and reference digestion technology, and finally the knowledge graph is constructed and updated.
It has achieved efficient and comprehensive construction and update of industrial innovation knowledge graphs, overcome the field limitations of traditional methods and insufficient long-tail knowledge coverage, and provides stronger generalization capabilities and timeliness.
Smart Images

Figure CN120179832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for dynamically constructing an industrial innovation knowledge graph based on a large language model. Background Art
[0002] As a structured semantic network, a knowledge graph represents entities and their relationships through nodes and edges respectively, and describes the attribute values of entity features through attributes, which can efficiently express knowledge and support complex reasoning tasks. Traditional knowledge graph construction and update rely on manual feature engineering, rule bases or specific machine learning models, involving steps such as entity extraction, relationship extraction, knowledge fusion and reasoning. For example, the method and device for constructing an industrial chain knowledge graph disclosed in CN117540031A. However, in the current context of promoting technological innovation and industrial development, in the face of various data resources such as patent information, academic literature and policy systems scattered on different platforms, traditional methods face many challenges. First, industrial innovation requires processing complex text data across different fields and languages, such as scientific and technological literature, patent documents and industry reports. Since traditional methods rely on domain experts to formulate rules or design templates, they are not only costly but also difficult to cope with the rapidly changing knowledge environment. Second, with the increasing complexity of the industrial environment, for interdisciplinary research and technological development involving multiple fields, traditional methods are often limited to specific fields and difficult to achieve cross-field migration and generalization, restricting their application scope. For example, a method for constructing and analyzing a ship industry knowledge graph disclosed in CN113987210A, a method for constructing a knowledge graph of a biomedical industry disclosed in CN118278507A, etc. Finally, traditional methods have limited coverage of long-tail knowledge (i.e., knowledge that is not common but crucial for a specific field), resulting in the constructed knowledge graph lacking comprehensiveness and being unable to fully meet the needs of industrial innovation analysis.
[0003] In contrast, in recent years, large models represented by GPT, DeepSeek, etc. have demonstrated excellent capabilities in natural language processing tasks, providing new solutions for the construction and update of knowledge graphs. Leveraging the pre-training and fine-tuning framework, large models possess deep language understanding capabilities through learning from massive amounts of text corpora. They can not only accurately extract entities and relationships from unstructured text but also excel in context reasoning and complex semantic understanding, and can better cover long-tail knowledge. This ability is particularly suitable for integrating and analyzing industrial innovation data distributed across different platforms, such as patent information, academic literature, and policy systems. Large models can directly utilize prompts or context information to complete tasks without relying on complex feature engineering and rule design. These characteristics have greatly improved the efficiency, flexibility, and knowledge comprehensiveness of knowledge graph construction and update methods based on large models compared to traditional methods, providing strong support for industrial knowledge graph application scenarios such as intelligent search, question-answering systems, and decision-making support, thus more effectively promoting technological innovation and industrial development. Summary of the Invention
[0004] Aiming at the problems of low efficiency, poor generalization ability, and limited coverage of long-tail knowledge in the existing construction and update of industrial innovation knowledge graphs, the present invention provides a method for dynamically constructing an industrial innovation knowledge graph based on a large language model.
[0005] To achieve the above object, the present invention provides the following technical solutions: The present invention provides a method for dynamically constructing an industrial innovation knowledge graph based on a large language model, comprising the following steps: S1. Collect and integrate multi-source industrial innovation data in structured or unstructured form, and perform structured preprocessing to form a standardized text data stream; the multi-source industrial innovation data includes multi-dimensional information such as patent information, academic literature, policy systems, news media, and industrial reports, and its data types include multi-modal data such as text, charts, audio, and structured tables; the structured preprocessing includes standardizing and converting multi-modal data into a unified text format that can be input into the large language model. S2. Perform zero-shot entity relation extraction and attribute extraction on the text data stream based on a triple extraction model trained through supervised fine-tuning and reinforcement learning, generate triple information including entities, relationships, and attributes, and implement reasoning and completion of temporal attributes, and then improve the stability of the extraction results through a multi-round verification mechanism. S3. Use entity alignment and coreference resolution techniques to perform semantic consistency processing on the extraction results, store the processed data in a graph database to construct an industrial innovation knowledge graph, and extract events based on core entities.
[0006] Furthermore, the process of step S2 includes: training a triple extraction model to adapt it to the triple extraction task, where the training process includes data annotation and supervised fine-tuning, and reinforcement learning for the triple extraction task; designing a structured prompt template to guide the large model to perform zero-shot extraction tasks, concatenating the template with the preprocessed text and inputting it into the large model to generate triple information, performing four independent samplings for each piece of text and retaining the result with the highest confidence, performing secondary reasoning to complete the entity attributes with fuzzy time descriptions, and storing the final results in a triple file, an entity attribute file, and a relationship file in JSON format for standardization.
[0007] Furthermore, the entity alignment process in step S3 includes: using a pre-trained model fine-tuned with domain data to generate entity vector representations, calculating semantic correlation degrees through cosine similarity and a multi-layer perceptron, and merging entities with similarity exceeding the threshold.
[0008] Furthermore, the anaphora resolution process in step S3 includes: determining the referent object of a pronoun entity based on the context analysis of the large model, and selecting the optimal candidate entity for replacement through correlation scoring.
[0009] Furthermore, constructing an industrial innovation knowledge graph in step S3 includes: storing triples based on the Neo4j graph database to build a visual knowledge network, defining an event as a structured dictionary containing core entity attributes and their associated relationships, and implementing event extraction and storage through the Cypher query language to form an industrial innovation knowledge graph that supports complex query analysis.
[0010] Furthermore, the above method for dynamically constructing an industrial innovation knowledge graph based on a large language model further includes the step of updating the knowledge graph: S4. Use the dynamic perception layer to continuously obtain new information from external data sources, combine the change detection and conflict marking mechanism driven by the large model, and perform incremental maintenance on the knowledge graph in an atomic transaction update manner to achieve incremental update and version tracing of events and the knowledge graph; external data sources include but are not limited to industry news APIs, patent databases, academic journals, policy release platforms, news websites, and industry reports.
[0011] Furthermore, continuously obtaining new information from external data sources by using the dynamic perception layer in step S4 includes: real-time monitoring of multi-source data streams from industry news, patent databases, and policy platforms through API interfaces and web crawlers, and using the large model to screen text fragments containing high-value update signals, where the update signals include first release statements, strategic transformation announcements, and key technology breakthrough indicators.
[0012] Furthermore, the change detection and conflict marking mechanism in step S4 includes: extracting the core entities of new events using a large model and matching them with the existing nodes in the knowledge graph, comparing the differences between the old and new version data and marking the attribute conflict points, evaluating the change confidence through multi-source cross-verification, directly triggering the creation operation for non-conflicting events, and generating trigger modification or deletion operations for conflicting events.
[0013] Furthermore, the atomic transaction update method in step S4 includes: dividing the update operations into three types of transaction groups, namely creation, modification, and deletion, according to the event type, simulating the execution of the full process change in a sandbox environment and generating a JSON difference file, implementing batch update using a two-phase commit protocol, marking the timeliness of nodes through timestamps and creating a version snapshot, and supporting transaction rollback and historical version tracing.
[0014] Furthermore, the incremental update and version tracing of events and the knowledge graph in step S4 include: parsing the core entities and associated relationships in the update event, performing create, delete, update, and query operations on the nodes of the graph database through the Cypher query language, synchronously updating the entity attribute file and relationship file, and recording the version change log containing the operation type, timestamp, and scope of influence.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The method for dynamically constructing an industrial innovation knowledge graph based on a large language model provided by the present invention overcomes the domain limitations of traditional methods in entity relationship extraction by integrating multi-source heterogeneous industrial data and utilizing the semantic understanding and reasoning capabilities of the large model. Through zero-shot prompt templates and multi-round verification mechanisms, high-precision triple extraction is achieved, and a comprehensive industrial knowledge graph is constructed. This knowledge graph integrates all elements of the industry through a structured semantic network, covering various technical fields and industry dynamics, forming a dynamic knowledge base to support strategic judgment, and providing multi-dimensional decision-making basis for industry trend prediction and industrial policy formulation.
[0016] The method for dynamically constructing an industrial innovation knowledge graph based on a large language model provided by the present invention realizes the incremental update of the knowledge graph by real-time monitoring data through a dynamic perception layer, combined with a conflict detection mechanism and atomic transaction processing. Data consistency is guaranteed through a two-phase commit protocol, and full-cycle tracing is achieved using version snapshots and operation logs, enabling efficient response to regular update scenarios. Compared with traditional methods, the present invention realizes efficient, accurate, and traceable updates of the knowledge graph, shows stronger generalization ability in scenarios such as policy timeliness tracking and technology evolution analysis, and ensures the currency and accuracy of the content of the knowledge graph. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0018] Figure 1 This is the overall flowchart of the method for dynamically constructing an industrial innovation knowledge graph based on a large language model provided by the embodiments of the present invention.
[0019] Figure 2 This is the flowchart of the method for creating an industrial innovation knowledge graph based on a large model provided by the embodiments of the present invention.
[0020] Figure 3 This is the flowchart of the method for updating an industrial innovation knowledge graph based on a large model provided by the embodiments of the present invention. Detailed implementation manners
[0021] To enable those skilled in the art to better understand the technical solutions of the present invention, the following will further introduce the present invention in detail in conjunction with the accompanying drawings and embodiments.
[0022] As Figure 1 shown, the present invention provides a method for dynamically constructing an industrial innovation knowledge graph based on a large language model, including the following steps: Step 1: Collect and integrate multi-source industrial innovation data in structured or unstructured form, and perform structured preprocessing to form a standardized text data stream; the multi-source industrial innovation data includes multi-dimensional information such as patent information, academic literature, policy systems, news media, and industrial reports, and its data types include multi-modal data such as text, charts, audio, and structured tables; the structured preprocessing includes standardizing and converting multi-modal data into a unified text format that can be input into a large language model. Step 2: Based on a triple extraction model trained through supervised fine-tuning and reinforcement learning, perform zero-shot entity relationship extraction and attribute extraction on the preprocessed data (text data stream), generate triple information including entities, relationships, and attributes, and implement inference and completion of temporal attributes, and then improve the stability of the extraction results through a multi-round verification mechanism. Step 3: Optimize the data quality through entity alignment and coreference resolution, perform semantic consistency processing on the extraction results, import the structured data into a graph database to construct an industrial innovation knowledge graph, and extract events based on core entities. Step 4: Obtain the newly added information from external data sources, and combine the large model-driven change detection and conflict marking mechanism to incrementally maintain the knowledge graph in an atomic transaction update manner, realizing the incremental update and version traceability of events and the knowledge graph; the external data sources include but are not limited to industry news APIs, patent databases, academic journals, policy release platforms, news websites, and industry reports.
[0023] Specifically, as Figure 2 shown, the implementation process of each step is described in detail as follows.
[0024] Step 1: Collect and integrate multi-source industrial innovation data such as patent information, academic literature, policy systems, news media, and industry reports, and perform structured preprocessing. In this embodiment, open-source data is collected by using crawlers and calling relevant data platform APIs, and some closed-source data is obtained through commercial cooperation to achieve the aggregation of industrial technology-related data and form an industrial text data stream.
[0025] The data preprocessing stage involves converting data in different formats (such as pdf, cxv, doc, html, jpg, etc.) into plain text format, and performing a series of data cleaning and normalization operations to ensure data quality. The specific steps are as follows.
[0026] Step 11: Textualization of multi-modal data In addition to the conventional pure text format (such as doc, html, etc.), industrial innovation data also includes other modal data, and it is necessary to uniformly convert other format data (such as csv, jpg, etc.) into a text format that can be semantically parsed. Other modal data includes: image data in patent charts and scientific and technological literature illustrations; embedded tabular data in policy reports and industrial yearbooks; speech data in expert interviews and industrial forum audio and video materials; captions and scene information in news reports and enterprise promotional videos. To achieve unified extraction and processing of different modal data at the semantic level, the present invention converts multi-source heterogeneous data such as images, tables, audio, and video into a unified text format that can be input into a large language model through a variety of multi-modal standardization conversion methods. The specific process includes: for image data such as patent charts and scientific and technological illustrations, an OCR system is used to identify the title, legend, labels, and annotation areas, and a visual semantic enhancement model BLIP is introduced to generate a structured scene description, and natural language fragments are uniformly constructed in combination with the metadata of the image content (including figure numbers, cited chapters, etc.). For the embedded tables in policy reports and industrial yearbooks, based on the table structure perception model TableNet, the table header, row and column structures, and cell contents are parsed, and semantic reconstruction is completed according to the expression method of "subject-attribute-value-unit" in combination with a large model to form text fragments with context-dependent relationships, while retaining the semantic hierarchy of the original table. For speech materials such as expert interviews and industrial forums, an automatic speech recognition (ASR) system is used to extract the original speech text, and in combination with video captions and speaker annotations, the dialogue paragraphs are reorganized according to dialogue turns or topic paragraphs to be reconstructed into clearly readable text segments. After the semantic reconstruction of all modal information is completed, it will be encapsulated into text segments in accordance with a unified structured prompt format, and at the same time, the modal type, content location, and source time index are marked to provide highly consistent semantic input for subsequent entity relationship extraction tasks and realize the generalization and understanding ability of the large language model on heterogeneous data.
[0027] Step 12: Processing of pure text format data For data in pure text format, its data processing process includes: removing redundant content such as special characters, extra spaces, and tags in the text; detecting and correcting incorrect data in the text; using the functions provided by data processing tools and programming languages to identify and delete duplicate data records; and unifying the representation forms of information such as date formats and currency units. The processed data will be used for knowledge graph construction and knowledge graph update.
[0028] Step 2: Based on the triple extraction model trained through supervised fine-tuning and reinforcement learning, perform zero-shot triple extraction on the preprocessed data, complete the joint extraction of entities, relationships, and attributes, and implement the inference and complementation of temporal attributes, and then improve the stability of the extraction results through a multi-round verification mechanism; the specific process is as follows.
[0029] Step 21: Training of the triple extraction model In this embodiment, a fine-tuning strategy and reinforcement learning are adopted to make the large model more effectively adapt to the triple extraction task. The specific implementation is as follows.
[0030] Step 211: Data annotation and supervised fine-tuning Select 1000 text paragraphs from the preprocessed data. Apply a rule-based named entity recognition and relation extraction method to each text segment to extract entities and their relationships, thereby constructing triples. Manually proofread the extracted triples and unify their formats. Package each text segment and its corresponding triples into a JSON object to form the training dataset required for fine-tuning. Use this data to perform supervised fine-tuning (SFT) on the DeepSeek-R1-Distill-Qwen-32B model to improve the model's ability to generate and label entities and relationships.
[0031] Step 212: Reinforcement learning for the triple extraction task After a small amount of data fine-tuning, the large model has the preliminary ability to extract triples. To break through the bottleneck of limited labeled data on the model performance, this embodiment introduces reinforcement learning to perform secondary optimization on the fine-tuned model. For the triple extraction task of this embodiment, the reward function is defined to include: content accuracy reward (accuracy of triple content), structure compliance reward (correctness of triple structure), and generation diversity reward (penalize repeated generation and encourage the generation of new triples). A reward signal is formed by weighting them in the ratio of 0.6:0.3:0.1. Use the fine-tuned model as the initial policy model and adopt the PPO algorithm to iteratively optimize the policy model.
[0032] Step 22: Zero-shot extraction of entity relationships and attributes This embodiment designs a structured prompt template to guide the identification of entities, relationships, and attributes (it should be noted that the task to be performed by the large model and the expected output format should be clearly stated in the prompt template. In all the following steps involving the large model, the focus is on constructing the prompt template, and after splicing the prompt with the data to be processed, the expected output is obtained by calling the large model. Only according to the different tasks to be performed and expected outputs, the prompt words should be adjusted accordingly, so the following relevant content will not be elaborated). The specific operation is to splice this prompt template with the industrial text data stream as the input of the large model, and the large model outputs the extracted triple information, including entities, relationships, and attributes; the large model outputs triple information. To facilitate the processing of entities and relationships, the entities, attributes, and relationships are stored in the entity attribute file and the relationship file respectively (it should be noted that when modifying the entity attribute file in the following steps, the corresponding entity attribute information in the triples should also be modified synchronously); mark the core entity in each piece of data in the entity file (generally the subject part of this piece of data) to prepare for the subsequent event extraction step.
[0033] Given the possible instability of the large model output, this embodiment adopts a self-consistency verification mechanism, using the large model to independently sample and extract each piece of text in the data stream four times and evaluate the confidence level, and only retains the result with the highest confidence level to ensure the accuracy and reliability of the data. Finally, all triple output results will be standardized in JSON format to reduce the impact of free text noise.
[0034] Step 23, Reasoning and Completion of Temporal Attributes In industrial data, time attributes such as the effective time and expiration time of policies are crucial for the construction and update of the knowledge graph. Based on the large model's understanding ability of temporal information, this embodiment specifically performs secondary processing on those entities with fuzzy time expressions in the entity attribute file. For example, a description like "the second quarter of next year" will be accurately mapped to specific years and months (such as "April - June 2025", assuming the data source release time is 2024), so as to make the construction and update process of the knowledge graph more accurate. This meticulous processing of time dimension information not only improves the accuracy of the data but also provides a solid foundation for subsequent time series analysis.
[0035] Step 3, Optimize the data quality through entity alignment and coreference resolution, import the structured data into the graph database to construct the industrial innovation knowledge graph, and extract events based on the core entities; the specific process is as follows.
[0036] Step 31, Entity Alignment and Coreference Resolution Given the diversity of data sources for knowledge graphs, it is possible that the names of the same entity may be inconsistent in different data sources, resulting in the repeated extraction of entities with the same meaning. This situation usually stems from different naming or representation methods for the same entity in different data sources. In addition, the use of pronouns may also make the extracted entities less specific, affecting the accuracy of the data. To ensure the consistency and accuracy of entities within the knowledge graph, this implementation adopts entity alignment and coreference resolution techniques.
[0037] Entity alignment in this implementation generates vector representations for each entity through a domain data fine-tuned Bert model, and then jointly measures the semantic association degree between entities through cosine similarity and a multi-layer perceptron (MLP). The specific operation is as follows: Calculate the cosine value of the embeddings of two entities to obtain similarity score 1; then concatenate the embeddings of the two entities, expand the dimension of the concatenated embeddings to twice the original through a linear transformation, restore it to the initial dimension after passing through the Relu activation function and a linear transformation, and finally map the embeddings processed by the linear layer to the range from -1 to 1 through the Tanh function as similarity score 2; calculate the average of similarity score 1 and similarity score 2 as the semantic association degree between the two entities. When the semantic association degree exceeds the manually set threshold, these two entities are regarded as the same entity and merged.
[0038] For coreference resolution in this implementation, other entities in the text where the entity with the part-of-speech of demonstrative pronoun is located are identified as potential entity candidates. Use a large model for context analysis to determine the most appropriate referent object and evaluate the association degree score of each candidate. Finally, select the entity with the highest score as the best referent object to replace the pronoun entity.
[0039] Step 32, Knowledge graph construction This implementation uses the Neo4j graph database to store triples, thereby constructing a visual industrial innovation knowledge graph. As an efficient graph database, Neo4j supports the Cypher query language and can provide strong data support for subsequent applications based on the knowledge graph. This knowledge graph can not only visually display the relationship network between entities, but also facilitate the execution of complex query and analysis tasks, greatly improving the efficiency and accuracy of information retrieval and update.
[0040] Step 33, Event extraction In a knowledge graph, an event is regarded as a collection of a special entity and a series of relationships. It contains multiple elements, such as participants (including individuals and organizations), time, location, etc. By incorporating events into the knowledge graph, not only can triple information be integrated, which helps to clearly sort out the development context of events, but also in-depth data analysis can be promoted, enhancing the organization and accessibility of information, and providing strong support for understanding and decision-making. In this embodiment, an event is a structured dictionary representation composed of all the attribute information in the core entity of each piece of data, as well as the relationships of the entity and the associated objects. After event extraction, it will be stored in JSON dictionary format and database for subsequent update and maintenance.
[0041] Step 4: Use the dynamic perception layer to capture industry dynamics, and adopt the atomic transaction mechanism to achieve incremental update and version tracing of events and the knowledge graph; the specific process is as Figure 3 shown, and the detailed process is described as follows.
[0042] Step 41: Dynamic perception and data collection To ensure the timeliness and accuracy of the industrial innovation knowledge graph, this embodiment introduces a dynamic perception layer to monitor and collect the latest information from multiple data sources in real time, including but not limited to industry news APIs, patent databases, academic journals, policy release platforms, news websites, and industry reports. By setting keyword monitoring, performing regular web crawler scraping, and calling API interfaces, data streams containing high-value update signals such as "first release" and "strategic transformation" are specifically screened out. These signals often indicate key market changes or corporate strategy adjustments. The new data streams are preprocessed by the method in Step 1 to be converted into plain text format. With the powerful natural language understanding ability of the large model, information that has a significant impact on a specific industry is accurately identified from the new data streams, ensuring that the collected data is both comprehensive and targeted.
[0043] Step 42: Change detection and conflict marking Before integrating newly collected data into the existing knowledge graph, a detailed change detection is required. This step utilizes large models to identify the differences between the new data and existing events, especially those critical changes that may affect the structure and content of the knowledge graph. Specifically, in this embodiment, the large model is first used to extract the events and core entities of the new data, and through core entity matching and large model assistance, it is determined whether the new events conflict with the existing events. Once it is found that the relevant attributes, relationships, or objects of a certain core entity have changed (such as a company name change or a technology upgrade), these changes are marked, and their impact on the current knowledge graph is evaluated. This process also includes marking the conflict points for subsequent review and resolution to ensure data consistency. If the core entity of the new event cannot be found in the existing events and knowledge graph or is found but identified as non-conflicting, the event will be marked as non-conflicting and will be directly used for knowledge graph update.
[0044] Step 43, Incremental Processing and Atomic Transaction Update After clarifying the scope of influence of the data changes, the incremental update process will be initiated. This process performs update operations on an event-by-event basis and synchronously marks the updated events during the execution. To ensure data consistency and integrity, this embodiment adopts an atomic transaction processing mechanism to ensure that each update operation is either fully executed successfully or completely rolled back to the initial state. It should be noted that the atomic transaction processing mechanism is executed in the database, and then the event dictionary is modified based on the database synchronization. Specifically, when implementing, all content to be updated needs to be first simulated and tested in a sandbox environment to verify the feasibility of the changes through a rehearsal. Only after the simulation environment is confirmed to be error-free will the updated content be officially submitted to the event dictionary in the production environment. This hierarchical mechanism of "pre-check - simulation - verification - implementation" not only avoids data chaos caused by partial update failures but also significantly improves the update efficiency and stability. The specific process is as follows.
[0045] Step 431, Update Pre-check and Transaction Grouping In this embodiment, the large model is first called to classify the events to be updated into three operation types: creation, modification, and deletion. Among them, the new events marked as conflicting will perform modification or deletion operations, and the new events marked as non-conflicting will perform creation operations. After creating a complete copy of the event file in the sandbox environment, inject the changed data and simulate the execution of the full process operation, and automatically detect potential risks such as time window overlap and version conflicts. The change operations that pass the pre-check will generate a transaction group containing a JSON difference file. The transaction group divides the submission units according to the association relationships between entities (for example, combining "new event A" and "revised related event B" into an atomic group) to ensure the coherence of the business logic.
[0046] Step 432, Atomic Submission and Version Snapshot Management In the formal submission stage of this embodiment, a two-phase commit protocol (2PC) is adopted to achieve batch updates. In the preparation stage, relevant file resources are locked and the current status is verified. In the commit stage, the changes that pass the pre-check are written into the production environment. If any link fails, a global rollback is triggered. After each successful submission, a version tag with a timestamp is automatically generated, and the old version is archived and stored to form a complete version chain. This implementation method supports tracing historical changes through timestamps or version tags and provides a one-key rollback function. During the submission process, the data writer adopts a batch operation and exponential backoff retry mechanism to ensure the execution reliability in high-concurrency scenarios.
[0047] Step 433, Continuous maintenance and exception recovery In this embodiment, the background service continuously monitors the status of event files, automatically marks expired events as invalid and triggers the update process. To improve processing efficiency, a hash sharding strategy is adopted to distribute high-concurrency requests, and the transaction operation logs are persistently stored according to the operation time sequence. If an exception occurs and interrupts the update process, the progress can be restored by parsing the transaction logs to ensure the integrity and consistency of the data during breakpoint resumption. The transaction logs also support the impact scope analysis in exception scenarios. Combined with the archiving mechanism of version snapshots, problems can be quickly located and data repair can be completed when necessary.
[0048] Step 44, Knowledge graph update After completing the necessary inspections and preparations, the new or modified data will be officially incorporated into the existing industrial innovation knowledge graph. This process involves performing operations such as creation, deletion, and modification on the corresponding nodes and edges in the graph database based on the update events, and all the above operations are completed through the Cypher query language. The specific process is as follows.
[0049] Step 441, Event parsing This implementation method uses a large model to identify entities, core entities, attributes, and relationships in the update event, and identifies whether there is a corresponding node for the core entity in the existing knowledge graph. Specifically, in this embodiment, the identified core entity is compared with the entities in the entity attribute file, and the large model is used to judge whether there are semantically consistent entities, so as to judge whether there are corresponding nodes in the knowledge graph.
[0050] Step 442, Perform create, delete, update, and query operations on the knowledge graph based on the update event For the query operation of this implementation method, corresponding entities, relationships, and attributes in the knowledge graph are found based on the update event.
[0051] For the create operation of this implementation method, if the entity relationship mentioned in the event does not exist in the knowledge graph, the corresponding points and edges need to be created in the knowledge graph.
[0052] For the deletion operation of this embodiment, if a certain relationship, object, or attribute in the event is no longer valid, the corresponding nodes and edges in the knowledge graph need to be deleted.
[0053] For the modification operation of this embodiment, if there are conflict points in the event, the corresponding node and edge information in the knowledge graph needs to be updated.
[0054] It should be noted that when modifying the knowledge graph in this embodiment, the entity attribute file and the relationship file are also modified to facilitate subsequent iterative updates.
[0055] In addition, for the convenience of tracking and management, each update of the knowledge graph records a detailed version history, thus facilitating future auditing and traceability work.
[0056] The method for dynamically constructing an industrial innovation knowledge graph based on a large language model provided by the present invention integrates multi-source heterogeneous data, performs cleaning, format conversion, and normalization processing, uses zero-shot extraction technology driven by a large model to automatically extract entities, relationships, and attributes, constructs a high-precision triple library by combining entity embedding alignment and context anaphora resolution, and generates a knowledge graph based on a graph database. The dynamic perception layer captures incremental data in real time, realizes the dynamic update of the graph through an atomic update mechanism, and combines timestamps and multi-source verification to ensure data timeliness. This method improves the construction efficiency and coverage rate of the knowledge graph, effectively alleviates the defects of high manual dependence, lack of long-tail knowledge, and update lag in traditional methods, and provides accurate support for technological innovation decision-making.
[0057] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for dynamically constructing an industrial innovation knowledge graph based on a large language model, characterized in that: The following steps are involved: S1. Collect and integrate structured or unstructured multi-source industrial innovation data, and perform structured preprocessing to form a standardized text data stream; the multi-source industrial innovation data includes multi-dimensional information such as patent information, academic literature, policies and systems, news media and industry reports, and its data type includes multi-modal data such as text, charts, audio and structured tables; the structured preprocessing includes standardizing the multi-modal data into a unified text format that can be input into a large language model; S2. Based on the triple extraction model trained by supervised fine-tuning and reinforcement learning, zero-shot entity relationship extraction and attribute extraction are performed on the text data stream to generate triple information containing entities, relationships and attributes, and implement reasoning completion of temporal attributes, and then the stability of the extraction results is improved through a multi-round verification mechanism; S3. Use entity alignment and coreference resolution techniques to process the semantic consistency of the extraction results, store the processed data in the graph database to build an industrial innovation knowledge graph, and extract events based on core entities.
2. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 1 is characterized in that: The process of step S2 includes: training the triple extraction model to adapt it to the triple extraction task, the training process includes data labeling and supervised fine-tuning, and reinforcement learning for the triple extraction task; designing a structured prompt template to guide the large model to perform the zero-sample extraction task, splicing the template with the preprocessed text and inputting it into the large model to generate triple information, performing four independent sampling extractions on each text and retaining the result with the highest confidence, performing secondary reasoning and completion on entity attributes with fuzzy time descriptions, and storing the final results in standardized JSON format as triple files, entity attribute files, and relationship files.
3. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 1 is characterized in that: The entity alignment process in step S3 includes: using a pre-trained model fine-tuned by domain data to generate entity vector representation, calculating semantic relevance through cosine similarity and multi-layer perceptron, and merging entities whose similarity exceeds a threshold.
4. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 1 is characterized in that: The reference resolution process in step S3 includes: determining the referent of the pronoun entity based on the large model context analysis, and selecting the best candidate entity for replacement through the relevance score.
5. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 1 is characterized in that: Constructing the industrial innovation knowledge graph in step S3 includes: constructing a visual knowledge network based on the Neo4j graph database to store triples, defining events as structured dictionaries containing core entity attributes and their association relationships, and implementing event extraction and storage through the Cypher query language to form an industrial innovation knowledge graph that supports complex query analysis.
6. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 1 is characterized in that: It also includes the steps of updating the knowledge graph: S4. Use the dynamic perception layer to obtain new information from external data sources in real time, combine it with the change detection and conflict marking mechanism driven by the big model, and perform incremental maintenance on the knowledge graph in an atomic transaction update manner to achieve incremental updates and version traceability of events and knowledge graphs; external data sources include but are not limited to industry information APIs, patent databases, academic journals, policy release platforms, news websites, and industry reports.
7. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 6 is characterized in that: In step S4, the new information of the external data source is obtained in real time by using the dynamic perception layer, including: monitoring the multi-source data streams of industry information, patent database and policy platform in real time through API interface and web crawler, and using large models to screen text fragments containing high-value update signals, wherein the update signals include the first release statement, strategic transformation announcement and key technology breakthrough mark.
8. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 6 is characterized in that: The change detection and conflict marking mechanism in step S4 includes: using the big model to extract the core entities of the new event and match them with the existing nodes of the knowledge graph, comparing the data differences between the new and old versions and marking the attribute conflict points, evaluating the change confidence through multi-source cross-validation, directly triggering the creation operation for non-conflicting events, and triggering the modification or deletion operation for conflicting events.
9. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 6 is characterized in that: The atomic transaction update method in step S4 includes: dividing the update operation into three types of transaction groups: creation, modification, and deletion according to the event type, simulating the execution of the whole process change in the sandbox environment and generating a JSON difference file, using a two-phase commit protocol to implement batch updates, marking the node timeliness with a timestamp and creating a version snapshot, supporting transaction rollback and historical version tracing.
10. The method for dynamically constructing an industrial innovation knowledge graph based on a large language model according to claim 6, characterized in that: The incremental update and version tracing of events and knowledge graphs in step S4 include: parsing the core entities and relationships in the update events, performing add, delete, modify and query operations on the graph database nodes through the Cypher query language, synchronously updating the entity attribute files and relationship files, and recording the version change log containing the operation type, timestamp and impact scope.
Citation Information
Patent Citations
Ship industry knowledge graph construction and analysis method
CN113987210A
Method and device for constructing industrial chain knowledge graph
CN117540031A
Construction method of knowledge graph of biomedical industry
CN118278507A
Fault knowledge graph construction method and device
CN115858796A
Generative AI large language model-based knowledge base construction method, system and equipment
CN118885465A
Cited By
Knowledge graph construction method and system based on large language model technology
CN120523966A
A knowledge graph construction method and system based on large language model technology
CN120523966B
DOM-based enterprise knowledge network construction method and system
CN120596685A
Fragmented semantic understanding order building system and method based on knowledge graph
CN120725026A
Automatic construction method and system for dynamic mode knowledge graph
CN120745784A