Industrial knowledge graph construction method, device, equipment, medium and program product

By constructing an industry knowledge graph and utilizing a large language model and quintuple representation, the problem of existing technologies being unable to analyze the vulnerability of the industrial chain is solved, enabling the temporal evolution and risk assessment of industrial chain relationships and supporting dynamic decision-making.

CN122154888APending Publication Date: 2026-06-05CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-02-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing knowledge graph construction methods are insufficient to support vulnerability analysis of entity relationships in the industrial chain, cannot reflect complex industrial chain dynamics, and cannot provide decision-making basis for supply chain resilience assessment and risk identification.

Method used

An industry knowledge graph is constructed by acquiring multiple sets of industry information, generating knowledge units and building a graph, including entities, entity relationships, time-series parameters and risk indicators. A large language model is used for entity recognition, relationship extraction and parameter extraction, generating quintuples and attaching risk indicators to achieve multi-dimensional representation of edge attributes.

Benefits of technology

It enables time-evolution analysis and risk assessment of supply chain relationships, supports supply chain resilience assessment and risk identification, and provides a basis for dynamic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122154888A_ABST
    Figure CN122154888A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computers and provides an industry knowledge graph construction method, device, equipment, medium and program product. The method comprises the following steps: acquiring multiple groups of industry information; each group of industry information comprises entity information, relationship information and relationship attribute parameters; the relationship attribute parameters comprise a time sequence parameter and a risk parameter corresponding to the relationship; based on the multiple groups of industry information, multiple knowledge units are generated, the knowledge units comprise a first entity and a second entity, an entity relationship, a time sequence parameter and a risk index corresponding to the entity relationship; the entity relationship is used for representing the relationship between the first entity and the second entity, and the risk index is a statistical result of the risk parameter; and based on the multiple knowledge units, an industry knowledge graph is constructed. The industry knowledge graph construction method, device, equipment, medium and program product provided by the application can reflect complex industry dynamics and provide decision basis for industry-related supply chain resilience evaluation and risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, equipment, medium, and program product for constructing an industrial knowledge graph. Background Technology

[0002] Currently, knowledge graphs are typically constructed using triples to represent facts and describe the relationships between entities. However, for industries related to critical technologies, this method of representing abstract factual knowledge often struggles to support vulnerability analysis of relationships between entities, fails to reflect complex supply chain dynamics, and cannot provide decision-making support for industry-related supply chain resilience assessments and risk identification. Summary of the Invention

[0003] This application provides an industrial knowledge graph construction method, apparatus, equipment, medium, and program product to solve the technical problem that the way knowledge graphs represent abstract factual knowledge is difficult to support the vulnerability analysis of relationships between entities.

[0004] In a first aspect, embodiments of this application provide a method for constructing an industry knowledge graph, including: Acquire multiple sets of industry information; each set of industry information includes entity information, relationship information, and relationship attribute parameters; the relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship. Based on multiple sets of industry information, multiple knowledge units are generated. Each knowledge unit includes a first entity and a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. Entity relationships are used to represent the relationship between the first entity and the second entity, and risk indicators are the statistical results of risk parameters. An industry knowledge graph is constructed based on multiple knowledge units. Nodes in the industry knowledge graph correspond to entities in the knowledge units, edges in the industry knowledge graph correspond to entity relationships in the knowledge units, and the attributes of the edges correspond to time-series parameters and risk indicators.

[0005] In one embodiment, the timing parameters include at least one of start time, end time, and time interval.

[0006] In one embodiment, risk indicators include at least one of concentration, dependence, redundancy, substitutability, performance differences, cost differences, maturity, investment size, cooperation intensity, duration, geopolitical risk, and export control risk.

[0007] In one embodiment, multiple sets of industry information are acquired, including: Collect structured and unstructured information from multi-source, heterogeneous industry data; Structured information, unstructured information, and predefined prompts are input into a large language model for extraction and processing to obtain multiple sets of industry information; the predefined prompts include entity naming prompts, relation extraction prompts, and parameter extraction prompts. Among them, entity naming prompts are used to identify entities corresponding to a predetermined entity type through contextual examples; relation extraction prompts are used to extract the predetermined relation types that conform between entities through contextual examples; and parameter extraction prompts are used to provide numerical parameters related to the relation through contextual examples.

[0008] In one embodiment, multiple knowledge units are generated based on multiple sets of industry information, including: Risk indicators are obtained by statistically analyzing the risk parameters in each group of industry information according to predetermined statistical rules. Based on the entity information, relation information, and time series parameters of each set of industry information, a quintuple is generated, and the risk index is attached to the relation attribute corresponding to the entity relation in the quintuple to obtain the knowledge unit corresponding to each set of industry information. The quintuple includes the first entity and the second entity, the entity relation, the start time in the time series parameters, and the end time in the time series parameters.

[0009] In one embodiment, the industry knowledge graph construction method further includes: The problem of receiving natural language input from users; The natural language problem and the predefined user intent recognition example are input into the large language model for processing to obtain the problem type of the natural language problem; The natural language problem and the predetermined information extraction example are input into the large language model for processing to obtain the information extraction result. The information extraction result includes at least one of the following in the natural language problem: entity, relation type, time range, indicator field and threshold condition. The extracted information is used to populate the query template corresponding to the question type, generating a query statement. Execute a query statement in the graph database of the storage industry knowledge graph to obtain the query results.

[0010] Secondly, embodiments of this application provide an industrial knowledge graph construction apparatus, comprising: The acquisition module is used to acquire multiple sets of industry information; each set of industry information includes entity information, relationship information, and relationship attribute parameters; the relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship. The processing module is used to generate multiple knowledge units based on multiple sets of industry information. The knowledge units include a first entity and a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. Entity relationships are used to represent the relationship between the first entity and the second entity, and risk indicators are the statistical results of risk parameters. The processing module is also used to construct an industry knowledge graph based on multiple knowledge units; the nodes in the industry knowledge graph correspond to the entities in the knowledge units, the edges in the industry knowledge graph correspond to the entity relationships in the knowledge units, and the attributes of the edges correspond to time series parameters and risk indicators.

[0011] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the industry knowledge graph construction method described in the first or second aspect.

[0012] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the industry knowledge graph construction method described in the first or second aspect.

[0013] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the industry knowledge graph construction method described in the first or second aspect.

[0014] The industry knowledge graph construction method, apparatus, equipment, medium, and program products provided in this application can construct an industry knowledge graph that can express the time dimension, retain the time evolution information of relationships such as supply, dependence, and cooperation, and introduce industry-specific risk indicators such as concentration, dependence, substitutability, geopolitical risk, and export control risk for each relationship to achieve multi-dimensional representation of edge attributes, thereby reflecting complex industry dynamics and providing decision-making basis for industry-related supply chain resilience assessment and risk identification. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is one of the flowcharts illustrating the industry knowledge graph construction method provided in this application embodiment.

[0017] Figure 2 This is the second flowchart illustrating the industry knowledge graph construction method provided in this application embodiment.

[0018] Figure 3 This is the third flowchart illustrating the industry knowledge graph construction method provided in this application embodiment.

[0019] Figure 4This is a schematic diagram of the structure of the industry knowledge graph provided in the embodiments of this application.

[0020] Figure 5 This is the fourth flowchart illustrating the industry knowledge graph construction method provided in this application embodiment.

[0021] Figure 6 This is a schematic diagram of the structure of the industrial knowledge graph construction device provided in the embodiments of this application.

[0022] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0024] Figure 1 This is one of the flowcharts illustrating the industry knowledge graph construction method provided in this application embodiment. (Refer to...) Figure 1 This application provides a method for constructing an industry knowledge graph, which may include: Step 110: Obtain multiple sets of industry information.

[0025] Each set of industry information includes entity information, relationship information, and relationship attribute parameters. Furthermore, a set of industry information may include two or more entities, as well as the relationships between each pair of entities, without limitation.

[0026] Entity information can include entity type and entity attributes. Entity type refers to the type of element an entity belongs to in the industry chain, such as enterprise, product, technology, patent, supplier, customer, factory, policy, or event. Entity attributes refer to the static characteristics of an entity, such as basic information like name and nationality.

[0027] Relationship information can be used to describe the types of relationships between entities. Relationship types can be one or more of the following: supply relationship, dependency relationship, investment relationship, cooperation relationship, substitution relationship, regional transfer, and patent reference, used to characterize the multidimensional connections between entities.

[0028] Relationship attribute parameters are used to statistically determine the attributes that define the relationships between entities. These parameters may include time-series parameters and risk parameters corresponding to the relationship. Time-series parameters describe the evolution of the relationship between entities over time and may include the start and / or end times of the relationship.

[0029] Risk parameters are used to conduct vulnerability analysis on relationships between entities and may include supplier output, upstream supplier supply, downstream enterprise total procurement, number of alternative suppliers, product performance score, product cost score, product technology maturity score, performance of alternative technologies, performance of mainstream technologies, cost of alternative solutions, cost of existing solutions, product development stage, joint R&D investment cost, number of jointly filed patents, number of joint papers, number of joint projects, duration, regional risk index, export control score, and political stability, etc.

[0030] Specifically, in one embodiment, multiple sets of industry information can be extracted from multi-source heterogeneous industry data. For example, industry data of various structures can be extracted from data sources such as policy documents, corporate annual reports, news reports, bidding announcements, patent documents, and scientific research papers.

[0031] Step 120: Generate multiple knowledge units based on multiple sets of industry information.

[0032] The knowledge unit includes a first entity and a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. Entity relationships represent the relationship between the first entity and the second entity, and risk indicators are the statistical results of the risk parameters.

[0033] In one embodiment, the timing parameters include at least one of start time, end time, and time interval.

[0034] For example, time-series parameters may include start and end times. In this case, the first and second entities, entity relationships, start and end times can be combined into a quintuple, and a risk indicator can be attached as a relational attribute of the entity relationships within that quintuple. That is, a knowledge unit includes a quintuple and a risk indicator attached to that quintuple.

[0035] For example, time-series parameters can include time intervals. In this case, the first entity and the second entity, the entity relationship, and the time interval can be combined into a quadruple, and the risk indicator can be attached as a relational attribute of the entity relationship in this quadruple. That is, a knowledge unit includes a quadruple and a risk indicator attached to that quadruple.

[0036] In one embodiment, multi-dimensional quantitative indicators can be designed and calculated for different types of relationships in the industrial chain to comprehensively characterize the strength and risk features of upstream and downstream relationships, facilitating vulnerability analysis. For example, risk indicators may include at least one of concentration, dependence, redundancy, substitutability, performance differences, cost differences, maturity, investment scale, cooperation intensity, duration, geopolitical risk, and export control risk.

[0037] Among these metrics, concentration can be used to measure whether market supply is overly concentrated. Dependence can reflect the strength of downstream enterprises' dependence on specific upstream nodes. Redundancy can characterize the number and distribution of alternative suppliers. Substitutability can measure the feasibility of alternative solutions in terms of performance and cost. Performance difference can reflect the extent to which the performance of an alternative technology improves or declines compared to the mainstream technology. Cost difference can reflect the percentage change in cost after adopting an alternative solution. Maturity can represent the level of technology maturity. Investment scale refers to the scale of funds invested in cooperation or R&D. Cooperation strength can reflect the comprehensive strength of the cooperative relationship in multiple aspects, and is a comprehensive indicator based on the number of jointly filed patents, joint papers, and joint projects. Duration refers to the time span of the relationship. Geopolitical risk can reflect the level of risk that the supply chain is exposed to geopolitical events. Export control risk can reflect the level of risk that the supply chain is affected by export controls.

[0038] Specifically, in one embodiment, each set of industry information can be organized into a five-tuple format, such as (first entity, entity relationship, second entity, start time, end time), and attribute fields such as {index name: index value} can be added to the entity relationship to obtain multiple knowledge units.

[0039] Step 130: Construct an industry knowledge graph based on multiple knowledge units.

[0040] In this industry knowledge graph, nodes correspond to entities within knowledge units. Edges correspond to entity relationships within knowledge units. Edge attributes correspond to time-series parameters and risk indicators.

[0041] Industry knowledge graphs possess a specific knowledge graph model that is industry-oriented, verifiable, and scalable. Specifically, industry knowledge graphs employ formal modeling methods based on knowledge units (such as quintuples and relation attributes) to structurally represent industry elements and their relationships. An industry knowledge graph can include four parts: node type, node attributes, relation type, and relation attributes.

[0042] Node types are populated by entity types in the entity information, such as enterprise, product, technology, patent, supplier, customer, factory, policy, and event, and are used to characterize the core elements of the industry chain. Node attributes are populated by entity attributes in the entity information, such as name, nationality, and other basic information, and are used to describe the static characteristics of the node. Furthermore, node types and node attributes are used to construct nodes in the industry knowledge graph.

[0043] Relationship types include supply relationships, dependency relationships, investment relationships, cooperation relationships, substitution relationships, regional transfer relationships, and patent citation relationships, used to characterize the multidimensional connections between nodes. Relationship attributes include time-series parameters such as start time and end time, as well as industry chain-specific risk indicators such as concentration, dependency, substitutability, and geopolitical risk, used to support dynamic analysis and risk assessment. Furthermore, relationship types and attributes are used to construct edges in the industry knowledge graph.

[0044] Specifically, in one embodiment, entities can be created as nodes, and the node type can be specified according to the aforementioned knowledge graph pattern. Furthermore, relationships can be created as edges, with edge attributes corresponding to relationship attributes, which may include start time, end time, and indicator value. Further, it may also include relevant industry information, the source of the corresponding industry information, and confidence scores. If a Resource Description Framework (RDF) is used for storage, it can be transformed into a five-tuple and attribute appending form, ultimately forming an industry knowledge graph that supports time-series queries, risk perception, and aggregation analysis.

[0045] Based on this, an industry knowledge graph that can express the time dimension can be constructed to retain the time evolution information of relationships such as supply, dependence, and cooperation. Industry-specific risk indicators such as concentration, dependence, substitutability, geopolitical risk, and export control risk can be introduced for each relationship to achieve multi-dimensional representation of edge attributes.

[0046] Figure 2 This is the second flowchart illustrating the industry knowledge graph construction method provided in this application embodiment. (Refer to...) Figure 2 This application provides a method for constructing an industry knowledge graph, which can acquire multiple sets of industry information, including: Step 210: Collect structured and unstructured information from multi-source heterogeneous industry data.

[0047] Specifically, in one embodiment, after collecting structured and unstructured information from multi-source heterogeneous industry data, the text data can be processed by deduplication, sentence segmentation, word segmentation, stop word removal, spelling standardization, and unit conversion. And / or, time expressions can also be standardized. And / or, entities such as enterprises, technologies, and products can also undergo alias disambiguation and unified naming processing. In this way, clean text corpora and structured metadata that can be directly used for natural language processing can be formed.

[0048] Step 220: Input structured information, unstructured information and predetermined prompt words into the large language model for extraction processing to obtain multiple sets of industry information.

[0049] The predefined prompts include entity naming prompts, relation extraction prompts, and parameter extraction prompts.

[0050] Entity naming prompts are used to identify entities corresponding to a predefined entity type by providing contextual examples. For example, entity naming prompts can contain several labeled examples to guide large language models in outputting standardized entity recognition results.

[0051] As an example, entity naming prompts can be designed as follows: Task: Identify entities appearing in the text (such as companies, products, technologies, patents, regions, times, etc.).

[0052] Output format: JavaScript Object Notation (JSON), where each entity includes name, category, location of occurrence, and time context.

[0053] Example: The text is "Company 1 manufactures chip b for Company 2 in year a, with a capacity share of 70%", and the output is {"Name": "Company 1", "Type": "Enterprise"}, {"Name": "Company 2", "Type": "Enterprise"}, {"Name": "chip b", "Type": "Product"}, {"Name": "Year a", "Type": "Time"}.

[0054] Now please process the following text: Enter text.

[0055] In this way, few-shot prompting technology can be used to identify entities such as enterprises, products, technologies, patents, events, and regions through contextual example prompts on a large language model.

[0056] Relation extraction prompts are used to suggest, through contextual examples, the predefined relationship types that exist between entities.

[0057] As an example, entity naming prompts can be designed as follows: Task: Identify the types of relationships between entities in a text (such as supply, dependence, cooperation, substitution, investment, patent citation, etc.).

[0058] Output format: JSON. Each relationship contains Entity 1, Relationship Type, Entity 2, Time Information, and Numerical Attributes (if any).

[0059] Example: The text is "Company 1 manufactures chip b for Company 2 in year a, and its production capacity share is 70%", the output is {"Subject": "Company 1", "Relationship": "Supply", "Object": "Company 2", "Start Time": "January 1, year a", "End Time": "December 31, year a", "Attribute": {"Supply Share": 0.7}.

[0060] Now please process the following text: Enter text.

[0061] Thus, few-shot hinting techniques can be used to automatically extract relationships from large language models through contextual example prompts, identifying supply relationships, dependencies, collaborations, substitutions, investment relationships, patent citations, regional transfer relationships, etc., between entities. Furthermore, post-processing such as validity checks can be performed on the output of the large language model to improve knowledge accuracy.

[0062] Parameter extraction prompts are used to suggest numerical parameters related to a relation through contextual examples. For example, parameter extraction prompts can be designed for each relation type, requiring a large language model to output results in structured JSON format.

[0063] As an example, the parameter extraction prompts for supply relationships could be designed as follows: Task: Extract parameters related to supply relationships from the following text and output them in JSON format.

[0064] Required output fields: supply relationship start time, supply relationship end time, supply share, total purchase volume of downstream enterprises, and original text fragment.

[0065] Example: The text is "Company 1 manufactures chip b for Company 2 in year a, and its capacity share is 70%", and the output is {start time: January 1, year a, end time: December 31, year a, supply share: 0.7, total purchase volume of downstream enterprises: empty, source: Company 1 manufactures chip b for Company 2 in year a, and its capacity share is 70%}.

[0066] Now please process the following text: Enter text.

[0067] In this way, the few-sample prompting technique can be used to extract key parameters for risk indicator calculation through a large language model, such as supply share, total purchase volume of downstream enterprises, performance indicators of alternative solutions, cost change ratio, amount of cooperative investment, and contract start and end time.

[0068] Furthermore, after completing entity recognition, relation extraction, and attribute parameter extraction, the extracted results can be parsed, standardized, and confidence evaluated to ensure data consistency and reliability.

[0069] Based on this, by combining large language models with few-shot prompting technology, entities can be automatically extracted, relationships identified, and indicator parameters extracted from multi-source heterogeneous industrial data. This enables semantic understanding and information extraction processes driven by large language models, accurately extracting structured and unstructured data such as policy documents, corporate annual reports, news reports, bidding announcements, patent documents, and scientific research papers, thereby obtaining multiple sets of industrial information that can be directly used for the construction of industrial knowledge graphs.

[0070] Figure 3This is the third flowchart illustrating the industry knowledge graph construction method provided in this application embodiment. (Refer to...) Figure 3 This application provides a method for constructing an industry knowledge graph, which can generate multiple knowledge units based on multiple sets of industry information, including: Step 310: Statistically analyze the risk parameters in each group of industry information according to the predetermined statistical rules to obtain risk indicators.

[0071] Predefined statistical rules may include a unified calculation method for each attribute. That is, the calculation method, required parameters, and data sources are predefined for each risk indicator to ensure that the results are repeatable and verifiable.

[0072] For example, concentration is calculated by dividing the output of the top N suppliers by the total industry output. Dependency is calculated by dividing the supply of a single upstream supplier by the total purchases of downstream companies. Redundancy is calculated by the number of alternative suppliers. Substitutability is calculated by weighted summing of performance score, cost score, and technology maturity score. Performance difference is calculated by subtracting the performance of the mainstream technology from the performance of the alternative technology, and then dividing by the performance of the mainstream technology. Cost difference is calculated by subtracting the cost of the existing solution from the cost of the alternative solution, and then dividing by the cost of the existing solution. Maturity ranges from 1 to 9, representing the stages from basic research to mature mass production. Investment scale is the total amount of joint R&D investment. Cooperation intensity is calculated by weighted summing of the number of jointly filed patents, the number of joint papers, and the number of joint projects. Duration is calculated by subtracting the start time from the end time of cooperation. Geopolitical risk can be quantified by functional relationships between the regional risk index, export control score, and political stability. Export control risks can be classified based on export control scores, such as low risk corresponding to a score range of 0 to 29, medium risk corresponding to a score range of 30 to 59, high risk corresponding to a score range of 60 to 79, and very high risk corresponding to a score range of 80 to 100.

[0073] Specifically, in one embodiment, risk parameters in each group of industry information can be statistically analyzed according to a risk indicator calculation method defined by predetermined statistical rules to obtain risk indicator values ​​such as concentration, dependence, redundancy, substitutability, performance differences, and cost differences. In this way, quantitative calculation methods can be clearly defined for different attributes of the industry chain relationship, enabling automated indicator calculation and dynamic updating.

[0074] Step 320: Based on the entity information, relation information and time series parameters of each group of industry information, generate a quintuple, and attach the risk index as the relation attribute corresponding to the entity relation in the quintuple to obtain the knowledge unit corresponding to each group of industry information.

[0075] The quintuple includes the first entity and the second entity, the entity relation, the start time in the timing parameters, and the end time in the timing parameters.

[0076] As an example, Figure 4 This is a schematic diagram of the structure of the industry knowledge graph provided in this application embodiment. For the knowledge that "Company 1 manufactures chip b for Company 2 in year a, with a 70% capacity share," the generated quintuple is (Company 1, supplier, Company 2, January 1, year a, December 31, year a), with the additional attribute fields {Concentration: 0.7; Dependency: 0.85; Redundancy: 1; Substitutability: Low; Product: chip b; Source: Company 1 manufactures chip b for Company 2 in year a, with a 70% capacity share}. Furthermore, referring to... Figure 4 It can be built in the industry knowledge graph, such as Figure 4 The nodes corresponding to Company 1 and Company 2 are shown, and the supply relationship is attached to the edge between the nodes, i.e., start time = January 1, year a, end time = December 31, year a, concentration = 0.7, dependence = 0.85, redundancy = 1, substitutability = low, product = chip b, source = Company 1 manufactures chip b for Company 2 in year a, and its capacity share is 70%.

[0077] Based on this, compared with the static triplet storage method used in related technologies, this application can calculate key risk indicators of the industrial chain. By modeling with five-tuples and adding attribute fields, it can combine the time dimension with industrial risk indicators. It not only records the time evolution trajectory of relationships such as supply, dependence, and cooperation, but also automatically calculates indicators such as concentration, dependence, substitutability, cost difference, and maturity, which can directly support time series analysis, trend prediction, and risk warning.

[0078] Furthermore, by integrating risk indicator calculation with knowledge graphs, integrated processing of indicator calculation and graph processing is achieved. This enables the real-time writing of risk indicator calculation results back to the attributes of edges, allowing for subsequent direct condition filtering, aggregation analysis, and threshold judgment. For example, it can directly query supply relationships with a dependency greater than 0.8 or calculate the concentration trend of a specified enterprise.

[0079] Figure 5 This is the fourth flowchart illustrating the industry knowledge graph construction method provided in this application embodiment. (Refer to...) Figure 5 This application provides a method for constructing an industry knowledge graph, which enables user question answering based on the knowledge industry graph, including: Step 510: Receive natural language input from the user.

[0080] For example, a natural language question input by the user could be about the dependence of battery company A on lithium mining company B within time 1. Another example is listing the battery suppliers for new energy vehicle company C.

[0081] Step 520: Input the natural language problem and the pre-defined user intent recognition example into the large language model for processing to obtain the problem type of the natural language problem.

[0082] Examples of predefined user intent recognition may include the input user question and the output user intent. The question type of a natural language question refers to the user intent in the natural language question, such as retrieval, query, statistics, sorting, or prediction.

[0083] For example, in the example of pre-defined user intent recognition, the input user question is to list the main suppliers of company 1 in year m, and the output user intent is {intent: retrieval, confidence level: 0.98}. Another example is that in the example of pre-defined user intent recognition, the input user question is how the market concentration of lithography machines will change from year m to year n, and the output user intent is {intent: statistics, confidence level: 0.95}.

[0084] Specifically, in one embodiment, after receiving a natural language question input by the user, the natural language question input by the user and a predetermined user intent recognition example can be concatenated into an intent recognition prompt word, and the concatenated intent recognition prompt word can be input into a large language model for processing.

[0085] Furthermore, intent recognition prompts can require large language models to output intent recognition results in strictly structured JSON. For example, intent recognition prompts could include a predefined example of user intent recognition and a natural language question input by the user, along with a request to return only structured JSON output.

[0086] Step 530: Input the natural language problem and the predetermined information extraction example into the large language model for processing to obtain the information extraction result.

[0087] The information extraction results include at least one of the following in natural language problems: entity, relation type, time range, indicator field, and threshold condition.

[0088] For example, multiple pre-defined information extraction examples can be set. Each pre-defined information extraction example corresponds to a user intent, and each pre-defined information extraction example may include one or more examples. The pre-defined information extraction example may display the input text and the expected output JSON field, and prompt for filling in the blank if the field does not exist.

[0089] Specifically, in one embodiment, the natural language question input by the user and a predetermined information extraction example corresponding to the question type can be concatenated into an information extraction prompt word, and the concatenated information extraction prompt word can be input into a large language model for processing.

[0090] As an example, information extraction prompts can be shown below: Task: Extract the following fields from the input question: entity array, relation type, time range, and threshold (if present).

[0091] Example: Input is the dependency of Company 1 on Company 2 in year y, output is {Entities: Company 1, Company 2, Relationship: Dependency, Time range: January 1, y to December 31, y, Metric: Dependency, Threshold: Empty}.

[0092] Now addressing: the natural language issue of user input.

[0093] Step 540: Fill the information extraction results into the query template corresponding to the question type and generate the query statement.

[0094] For example, corresponding query structures can be designed for different relationship types such as supply, dependence, substitution, cooperation, investment, regional transfer, and patent citation. Time constraints can be added to the query template, using start and end times to filter relationships, supporting both point-in-time and time-range queries. Furthermore, indicator fields, comparison operators, and threshold placeholders can be reserved in the query template to support dynamic filtering of attributes such as concentration, dependence, substitutability, and geopolitical risk, and to support ascending or descending sorting by indicator values. Additionally, entity name placeholders and relationship type placeholders can be reserved in the query template for dynamically populating specific companies, technologies, products, and their corresponding relationship types.

[0095] Furthermore, common problems can be abstracted into types such as query, statistics, comparison, sorting, aggregation, and prediction, and different reusable query templates can be designed for each type to ensure consistency and traceability in subsequent data entry, calculation, analysis, and write-back.

[0096] Specifically, in one embodiment, the text parameters such as entities and relation types in the natural language questions included in the information extraction results can be linked with standardized nodes in the industry knowledge graph. Through techniques such as alias disambiguation and fuzzy matching, it can be ensured that the final entity names correspond one-to-one with the node names in the industry knowledge graph. Furthermore, the time range, indicators, and other fields included in the information extraction results can be standardized to output a structured and standardized parameter set, which serves as input for subsequent template filling.

[0097] Furthermore, based on the identified natural language question type, a suitable query template can be selected, and the structured parameters obtained from the information extraction results can be filled into the placeholders of the selected query template to generate a specific graph database query statement.

[0098] Step 550: Execute the query statement in the graph database storing the industry knowledge graph to obtain the query results.

[0099] Specifically, in one embodiment, the generated query statement can be executed in a graph database storing the industry knowledge graph. The graph database returns node, relationship, and attribute data that meet the conditions, and the results are aggregated, sorted, or threshold-filtered to obtain the query results. Furthermore, the query results can be returned as input to the result integration module for further semantic interpretation, indicator calculation, or predictive analysis, thereby answering the user's natural language questions.

[0100] Thus, upon receiving a natural language question from a user, a large language model can be invoked first, combined with few-shot hints technology, to perform semantic parsing of the input question, identify the user's intent, and extract key parameters. Then, the user question can be categorized into a preset task type, such as retrieval, comparison, statistics, or prediction, determining the appropriate query logic. Furthermore, elements corresponding to the industry knowledge graph can be identified from the user input, such as entities, relationship types, time ranges, indicator fields, threshold conditions, and ranking requirements, providing complete parameters for filling in subsequent query templates. Finally, the parsing results can be mapped to the corresponding query template to generate a specific graph database query statement, which is then executed to return result data matching the user question, providing complete input for subsequent result interpretation, statistical analysis, or predictive calculations.

[0101] Based on this, compared with traditional methods that have poor generalization ability and rely on a large amount of manual annotation and rule matching, this application can automatically complete entity recognition, relation extraction and indicator parameter extraction through intelligent information extraction based on large language models and few-shot prompting technology. This significantly reduces the need for annotation data, improves the recognition accuracy and robustness on industry chain professional corpora, and outputs structured JSON results, which are easy to write directly into graph databases.

[0102] Furthermore, by using predefined query templates and intent recognition logic, it can automatically convert users' natural language questions into graph database query statements, enabling advanced analysis tasks such as dynamic question answering in the industrial chain, trend prediction, and alternative path recommendation. This breaks through the limitation of traditional knowledge graphs that only support static queries, and enables dynamic question answering and predictive reasoning.

[0103] The industry knowledge graph model proposed in the above embodiments of this application can be well adapted to the industrial chain scenario, supports the construction of industry knowledge graphs that can express the time dimension, expands the traditional triple relationship to a quintuple, and retains the time evolution information of supply, dependence, cooperation and other relationships. It introduces industry-specific risk indicators such as concentration, dependence, substitutability, geopolitical risk, and export control risk for each relationship, realizing multi-dimensional representation of edge attributes. Secondly, by using a large language model and few-shot prompting technology, entities such as enterprises, products, and technologies can be automatically extracted from multi-source texts such as policies, news, corporate annual reports, bidding announcements, and patents, and relationships such as supply, dependence, substitution, investment, and patent citations can be identified, and the time range and intensity can be extracted. Finally, based on graph database aggregation operations, industry-specific attribute systems and other indicators can be automatically calculated, and relationship attributes can be written back in real time for risk analysis and decision support.

[0104] The following describes the industrial knowledge graph construction apparatus provided in the embodiments of this application. The industrial knowledge graph construction apparatus described below and the industrial knowledge graph construction method described above can be referred to in correspondence.

[0105] Figure 6 This is a schematic diagram of the structure of the industrial knowledge graph construction device provided in the embodiments of this application. Figure 6 The illustrated industrial knowledge graph construction device includes: The acquisition module 610 is used to acquire multiple sets of industry information. Each set of industry information includes entity information, relationship information, and relationship attribute parameters. The relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship.

[0106] The processing module 620 is used to generate multiple knowledge units based on multiple sets of industry information. Each knowledge unit includes a first entity, a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. Entity relationships represent the relationship between the first entity and the second entity, and risk indicators are statistical results of the risk parameters.

[0107] The processing module 630 is also used to construct an industry knowledge graph based on multiple knowledge units. Nodes in the industry knowledge graph correspond to entities in the knowledge units, edges in the industry knowledge graph correspond to entity relationships in the knowledge units, and the attributes of the edges correspond to time-series parameters and risk indicators.

[0108] In one embodiment, the timing parameters include at least one of start time, end time, and time interval.

[0109] In one embodiment, risk indicators include at least one of concentration, dependence, redundancy, substitutability, performance differences, cost differences, maturity, investment size, cooperation intensity, duration, geopolitical risk, and export control risk.

[0110] In one embodiment, the acquisition module 610 is specifically used for: Collect structured and unstructured information from multi-source, heterogeneous industry data; Structured information, unstructured information, and predefined prompts are input into a large language model for extraction and processing to obtain multiple sets of industry information; the predefined prompts include entity naming prompts, relation extraction prompts, and parameter extraction prompts. Among them, entity naming prompts are used to identify entities corresponding to a predetermined entity type through contextual examples; relation extraction prompts are used to extract the predetermined relation types that conform between entities through contextual examples; and parameter extraction prompts are used to provide numerical parameters related to the relation through contextual examples.

[0111] In one embodiment, the processing module 620 is specifically used for: Risk indicators are obtained by statistically analyzing the risk parameters in each group of industry information according to predetermined statistical rules. Based on the entity information, relation information, and time series parameters of each set of industry information, a quintuple is generated, and the risk index is attached to the relation attribute corresponding to the entity relation in the quintuple to obtain the knowledge unit corresponding to each set of industry information. The quintuple includes the first entity and the second entity, the entity relation, the start time in the time series parameters, and the end time in the time series parameters.

[0112] In one embodiment, the industry knowledge graph construction method further includes: The acquisition module 610 is also used to receive natural language questions input by the user; The processing module 620 is also used to input natural language questions and predefined user intent recognition examples into a large language model for processing to obtain the question type of the natural language question; The processing module 620 is also used to input the natural language problem and the predetermined information extraction example into the large language model for processing, and to obtain the information extraction result. The information extraction result includes at least one of the entities, relation types, time ranges, indicator fields and threshold conditions in the natural language problem. The processing module 620 is also used to fill the information extraction results into the query template corresponding to the question type and generate a query statement; The processing module 620 is also used to execute query statements in the graph database storing the industry knowledge graph and obtain query results.

[0113] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call a computer program in the memory 730 to execute the steps of the industry knowledge graph construction method, such as including: Acquire multiple sets of industry information. Each set includes entity information, relationship information, and relationship attribute parameters. Relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship. Based on these multiple sets of industry information, generate multiple knowledge units. Each knowledge unit includes a first entity, a second entity, entity relationships, corresponding time-series parameters, and risk indicators. Entity relationships represent the relationship between the first and second entities, and risk indicators are statistical results of the risk parameters. Based on these multiple knowledge units, construct an industry knowledge graph. Nodes in the industry knowledge graph correspond to entities within knowledge units, edges correspond to entity relationships within knowledge units, and edge attributes correspond to time-series parameters and risk indicators.

[0114] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0115] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the industry knowledge graph construction method provided in the above embodiments, such as including: Acquire multiple sets of industry information. Each set includes entity information, relationship information, and relationship attribute parameters. Relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship. Based on these multiple sets of industry information, generate multiple knowledge units. Each knowledge unit includes a first entity, a second entity, entity relationships, corresponding time-series parameters, and risk indicators. Entity relationships represent the relationship between the first and second entities, and risk indicators are statistical results of the risk parameters. Based on these multiple knowledge units, construct an industry knowledge graph. Nodes in the industry knowledge graph correspond to entities within knowledge units, edges correspond to entity relationships within knowledge units, and edge attributes correspond to time-series parameters and risk indicators.

[0116] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to perform the steps of the methods provided in the above embodiments, such as including: Acquire multiple sets of industry information. Each set includes entity information, relationship information, and relationship attribute parameters. Relationship attribute parameters include time-series parameters and risk parameters corresponding to the relationship. Based on these multiple sets of industry information, generate multiple knowledge units. Each knowledge unit includes a first entity, a second entity, entity relationships, corresponding time-series parameters, and risk indicators. Entity relationships represent the relationship between the first and second entities, and risk indicators are statistical results of the risk parameters. Based on these multiple knowledge units, construct an industry knowledge graph. Nodes in the industry knowledge graph correspond to entities within knowledge units, edges correspond to entity relationships within knowledge units, and edge attributes correspond to time-series parameters and risk indicators.

[0117] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for constructing an industry knowledge graph, characterized in that, include: Obtain multiple sets of industry information; Each set of industry information includes entity information, relationship information, and relationship attribute parameters; The relation attribute parameters include time-series parameters and risk parameters corresponding to the relation; Based on the multiple sets of industry information, multiple knowledge units are generated. Each knowledge unit includes a first entity and a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. The entity relationships are used to represent the relationship between the first entity and the second entity, and the risk indicators are the statistical results of the risk parameters. Based on the aforementioned multiple knowledge units, an industry knowledge graph is constructed; The nodes in the industry knowledge graph correspond to the entities in the knowledge unit, the edges in the industry knowledge graph correspond to the entity relationships in the knowledge unit, and the attributes of the edges correspond to the time series parameters and the risk indicators.

2. The method for constructing an industry knowledge graph according to claim 1, characterized in that, The timing parameters include at least one of start time, end time, and time interval.

3. The method for constructing an industry knowledge graph according to claim 1, characterized in that, The risk indicators include at least one of the following: concentration, dependence, redundancy, substitutability, performance differences, cost differences, maturity, investment scale, cooperation intensity, duration, geopolitical risk, and export control risk.

4. The method for constructing an industry knowledge graph according to claim 1, characterized in that, The acquisition of multiple sets of industry information includes: Collect structured and unstructured information from multi-source, heterogeneous industry data; The structured information, the unstructured information, and the predetermined prompt words are input into a large language model for extraction processing to obtain the multiple sets of industry information; the predetermined prompt words include entity naming prompt words, relation extraction prompt words, and parameter extraction prompt words. The entity naming prompt is used to identify entities corresponding to a predetermined entity type through contextual example prompts; The relation extraction prompts are used to suggest the predefined relation types that the entities being extracted conform to, through contextual examples. The parameter extraction prompts are used to provide contextual examples to suggest numerical parameters related to the relationship.

5. The method for constructing an industry knowledge graph according to claim 1, characterized in that, Based on the multiple sets of industry information, multiple knowledge units are generated, including: Risk indicators are obtained by statistically analyzing the risk parameters in each set of industry information according to predetermined statistical rules. Based on the entity information, the relationship information, and the time series parameters of each group of industry information, a quintuple is generated, and the risk index is attached as the relationship attribute corresponding to the entity relationship in the quintuple to obtain the knowledge unit corresponding to each group of industry information. The quintuple includes the first entity and the second entity, the entity relationship, the start time in the time series parameters, and the end time in the time series parameters.

6. The method for constructing an industry knowledge graph according to claim 1, characterized in that, The method further includes: The problem of receiving natural language input from users; The natural language problem and the predefined user intent recognition example are input into a large language model for processing to obtain the problem type of the natural language problem. The natural language problem and the predetermined information extraction example are input into the large language model for processing to obtain the information extraction result. The information extraction result includes at least one of the entities, relation types, time ranges, indicator fields and threshold conditions in the natural language problem. The extracted information is used to fill the query template corresponding to the question type, and a query statement is generated. The query statement is executed in the graph database storing the industry knowledge graph to obtain the query results.

7. An industrial knowledge graph construction device, characterized in that, include: The acquisition module is used to acquire multiple sets of industry information; Each set of industry information includes entity information, relationship information, and relationship attribute parameters; The relation attribute parameters include time-series parameters and risk parameters corresponding to the relation; The processing module is used to generate multiple knowledge units based on the multiple sets of industry information. The knowledge units include a first entity and a second entity, entity relationships, time-series parameters corresponding to the entity relationships, and risk indicators. The entity relationships are used to represent the relationship between the first entity and the second entity, and the risk indicators are the statistical results of the risk parameters. The processing module is also used to construct an industry knowledge graph based on the multiple knowledge units; The nodes in the industry knowledge graph correspond to the entities in the knowledge unit, the edges in the industry knowledge graph correspond to the entity relationships in the knowledge unit, and the attributes of the edges correspond to the time series parameters and the risk indicators.

8. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the industrial knowledge graph construction method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the industry knowledge graph construction method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the industrial knowledge graph construction method according to any one of claims 1 to 6.