An industry chain design method and device and a computer storage medium

By optimizing node graphs using multi-source heterogeneous datasets and machine learning, the problems of insufficient data processing and poor graph quality in traditional supply chain design are solved, realizing the automation, intelligence and precision of supply chain design, and improving the completeness and accuracy of supply chain graphs.

CN122114767APending Publication Date: 2026-05-29BEIJING RUYITANG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING RUYITANG TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-05-29

Smart Images

  • Figure CN122114767A_ABST
    Figure CN122114767A_ABST
Patent Text Reader

Abstract

The application discloses an industry chain design method and device and a computer storage medium, relates to the technical field of industry chain management, and comprises the following steps: firstly, collecting multi-source industry data legally and preprocessing and storing the multi-source industry data; then, determining candidate nodes through knowledge framework construction, entity relationship extraction, score calculation and screening verification; setting a multi-task optimization target value strategy or an ideal value strategy function; training a learning model or screening potential chain nodes through multi-target optimization; then, constructing an initial knowledge graph, complementing missing correlations and optimizing the structure through a clustering algorithm; finally, constructing an evaluation system and forming a complete industry chain design process. Through the fusion of multi-source heterogeneous data processing, knowledge graph construction, reinforcement learning screening and whole life cycle analysis technology, the application solves the problems of insufficient data processing, inaccurate node screening, poor graph quality, incomplete chain and unstable generated quality in traditional industry chain design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of supply chain management, artificial intelligence, and knowledge graph technology, and specifically to a supply chain design method, apparatus, and computer storage medium. Background Technology

[0002] The digital and intelligent design of the industrial chain is a core driving force for industrial upgrading. Industry reports, corporate annual reports, and other texts contain rich industrial chain information, but manual extraction is inefficient and makes it difficult to quickly construct a complete industrial chain map. The integrated application of knowledge graphs, big data, machine learning, and large models provides a new solution for industrial chain analysis, management, map construction, and quality evaluation. Related research mainly focuses on four aspects: first, data mining and analysis, extracting entity relationships through web crawlers and NER models to generate and optimize triplet maps; second, knowledge screening and intelligent management, relying on models such as XGBoost to predict nodes and accurately screen enterprises through semantic matching; third, automatic map generation, combining graph databases, NER, clustering, and BERT to improve construction efficiency; and fourth, quantitative quality evaluation, using weighted association and random walk to achieve scientific assessment. In the future, the integration of multiple technologies will further promote the intelligent development of the industrial chain.

[0003] Existing technology, such as the invention patent application with announcement number CN118297543A, discloses a collaborative service platform system and method for the coal mining machinery industry chain. This system includes a resource layer for connecting and exchanging information between upstream and downstream enterprises in the coal mining machinery industry chain, as well as for data fusion across the industry chain, to achieve efficient collaboration across all aspects of the coal mining machinery industry chain, including supply chain, design, production, and operation and maintenance. A service module layer supports the application and optimization of various service modules, each providing different services, and at least possesses platform parameter optimization, intelligent matching, and risk warning functions. A service enhancement layer is used for collaborative development, autonomous learning, adaptive perception of the industry chain situation, and improvement of collaborative services among enterprises within the industry chain. This invention helps the coal mining machinery industry chain form a complementary and integrated adaptive evolution solution for a complex collaborative enterprise group, achieving enhanced collective intelligence and sustainable ecosystem construction.

[0004] As can be seen from the above solutions, traditional supply chain design methods have multiple limitations: First, they rely on human experience, making it difficult to adapt to the needs of industrial data analysis, mining, and design under large-scale data; second, they lack entity relationships in industrial knowledge, especially the deep relationships between technical entities are not fully explored, which cannot support enterprises' refined production decisions; third, data processing is limited by scale and accuracy, resulting in unreliable supply chain construction data and difficulty in accurately selecting qualified enterprises; fourth, the map construction relies on manual organization, which is time-consuming and labor-intensive, and the classification is vague, making it impossible to quickly generate multi-perspective supply chains with universal applicability; fifth, the evaluation methods rely on only a single indicator, which is difficult to reflect the rationality of the supply chain structure and the strength of its connections. Summary of the Invention

[0005] To address the aforementioned technical shortcomings, the present invention aims to provide a supply chain design method, apparatus, and computer storage medium.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In the first aspect, the present invention provides an industry chain design method, including: screening industry nodes from a constructed multi-source heterogeneous dataset, confirming potential chain nodes based on an ideal value strategy, completing and optimizing the node graph, and completing the industry chain design based on full life cycle analysis.

[0007] The data sources of the multi-source heterogeneous dataset include open-source or closed-source industry, enterprise, market, sector, product data and product functional semantic data; the multi-source heterogeneous data must contain the mapping relationship between industry, product and functional semantic data; the construction process of the dataset is to preprocess the original data, perform field mapping and format alignment according to a unified knowledge ontology, and complete data fusion and standardization.

[0008] The selection of industry nodes is based on a set industry knowledge framework to complete node selection and optimization. The industry knowledge framework revolves around the goals or needs of the industrial chain and is a combination of one or more knowledge systems from the perspective of industry, sector or field characteristics, including the category, product, technology and parameters. The selection is based on the industry knowledge framework selected by the user, combined with historical industry data statistics and weight analysis to determine the preset knowledge set. The optimization includes manual review and correction or machine learning model correction based on feedback data, and includes the construction of potential industrial chain levels and the generation of knowledge combinations of industry direction nodes.

[0009] The ideal value strategy is used to train the learning model and identify potential chain nodes. It includes at least one or a combination of learning reward function, energy index function, and value function. The selection strategy for the node knowledge combination is optimized through multiple iterations of training. The learning model includes various models or combinations of machine learning, deep learning, or reinforcement learning.

[0010] The ideal value strategy can also be a multi-objective value function, which performs multi-objective optimization and screening of potential industry chain node combinations. The multi-objective value function includes multi-dimensional task objective modeling of chain node value, relevance, coverage, relationship accuracy, completeness, foresight and timeliness, and defines quantitative weights for the corresponding indicators of each objective to construct a weighted value function. The multi-objective optimization includes at least optimizing the coverage and ideal value of node knowledge combinations, and the indicator proportions can be dynamically adjusted.

[0011] The completion and optimization of the node map is based on the upstream and downstream relationship analysis to confirm the chain nodes. It can introduce initial industrial chain node knowledge or node knowledge set and integrate it with existing node knowledge. The upstream and downstream relationship analysis is to label the relationship type and train the model to predict potential industrial chain node entities and entity relationships. The completion of chain node confirmation is to obtain the industrial chain node set at different levels from a certain perspective.

[0012] The aforementioned supply chain design based on full life cycle analysis involves dividing the life cycle into various stages around the goals or needs of the supply chain, ensuring the integrity of the supply chain's formation.

[0013] When completing the industrial chain design, it is necessary to construct a key element evaluation model to obtain quantitative results of the industrial chain design quality, and form multi-perspective industrial chain design results and ranking optimization results. The key element evaluation model comprehensively evaluates the rationality of the distribution of upstream and downstream nodes in the industrial chain based on the weight of the correlation relationship and the dependency relationship between nodes. The multi-perspective industrial chain design results include macro perspective, industry perspective, micro perspective, product perspective, domain perspective, problem perspective and concept perspective, forming a multi-perspective industrial chain design set and ranking results.

[0014] In a second aspect, the present invention provides an industry chain design device, comprising: a data acquisition module for collecting industry-related data from compliant multi-source channels and preprocessing it to form a multi-source heterogeneous dataset.

[0015] Potential chain node screening module: used to build an industry knowledge framework, extract entities and relationships, calculate enterprise industry matching scores, and obtain a set of potential chain nodes through threshold screening and cross-validation.

[0016] The completion and optimization module is used to construct ideal value strategy functions or set multi-task optimization target value strategies. Through learning models or multi-objective optimization decisions, it filters potential chain node sets and is also used to construct an initial industrial chain knowledge graph, complete missing associations, and optimize the graph structure through clustering.

[0017] The full lifecycle analysis module is used to divide the lifecycle into various stages based on the goals or needs of the industry chain, ensuring the integrity of the industry chain's formation.

[0018] The supply chain design quality assessment module is used to generate multi-perspective supply chain design results and ranking optimization results, providing decision support for the generated supply chain results.

[0019] In a third aspect, the present invention provides a computer storage medium for supply chain design, wherein a computer program is burned into the computer storage medium for supply chain design, and the computer program implements the supply chain design method described above when it runs in the memory of a server.

[0020] The beneficial effects of this invention are as follows: 1. This invention provides a supply chain design method, apparatus, and computer storage medium. First, multi-source industry data is legally collected, pre-processed, and stored. Then, candidate nodes are determined through knowledge framework construction, entity relationship extraction, score calculation, and screening verification. A multi-task optimization objective value strategy or ideal value strategy function is set. Potential chain nodes are screened through learning model training or multi-objective optimization. An initial knowledge graph is then constructed, missing associations are filled in, and the structure is optimized through clustering algorithms. Finally, an evaluation system is constructed, forming a complete supply chain design process. This solution, by integrating multi-source heterogeneous data processing, knowledge graph construction, reinforcement learning screening, and full lifecycle analysis technologies, solves the problems of insufficient data processing, inaccurate node screening, poor graph quality, incomplete chains, and unstable generation quality in traditional supply chain design, achieving automation, intelligence, and precision in multi-perspective supply chain design.

[0021] 2. By mining deep relationships between entities, using reinforcement learning to screen potential chain nodes, and employing a dynamic update mechanism, the completeness, accuracy, and timeliness of the industry chain map have been improved, reducing omissions and errors in recording. Based on full life cycle analysis and a multi-dimensional quality evaluation system, quantitative assessment of industry chain design has been achieved, providing precise support for enterprise decision-making and industrial policy formulation. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram illustrating the implementation steps of the supply chain design method of the present invention; Figure 2 This is a schematic diagram of the multi-source heterogeneous data processing flow of the present invention; Figure 3 This is a mapping diagram of the data in this invention; Figure 4 This invention establishes a good mapping relationship diagram between industries and sectors; Figure 5 This is a table showing the industry mapping results of the vehicle body components and materials and standard classifications selected for this invention. Figure 6 This is the original data preprocessing flowchart for this invention; Figure 7 This invention provides a relational mapping diagram for data fusion and standardization. Figure 8 This is a visualization of the fused dataset achieved in this invention. Figure 9 This is a schematic diagram of the process for selecting industry nodes in this invention; Figure 10 This is a visualization diagram of industrial knowledge clustering for this invention; Figure 11 This is a schematic diagram illustrating the process of identifying potential chain nodes in this invention. Figure 12 This is a schematic diagram of the ideal value strategy calculation method of the present invention; Figure 13 This is a schematic diagram of the optimal path selection for the reinforcement learning strategy of this invention; Figure 14 This is a diagram showing the parameter settings for the reinforcement learning strategy of this invention; Figure 15 This is a comparison chart of the generation effects of the industrial chain of the present invention; Figure 16 This is a diagram showing the training results of the reinforcement learning model of this invention; Figure 17 To confirm potential chain nodes and combination tables for this invention; Figure 18 This is a preliminary rendering of the industry chain diagram of the present invention; Figure 19 This is the prospective feature evaluation table for chain nodes in this invention; Figure 20 This is a table showing the effect of the industrial chain knowledge graph obtained after multi-objective optimization in this invention; Figure 21 This is a schematic diagram illustrating the relationship between nodes determined through upstream and downstream relationship analysis in this invention. Figure 22 This is a table showing the results of the knowledge processing and hierarchical analysis of the industry chain nodes in this invention. Figure 23 This is a schematic diagram of the upstream and downstream results of the industrial chain of this invention; Figure 24 This is a table showing the hierarchical division of the new energy vehicle industry chain nodes in this invention; Figure 25 This is a table evaluating the prediction effect and reliability of the hierarchical relationship of the industrial chain nodes in this invention. Figure 26 This is a table showing the results of the industry chain life cycle analysis of this invention; Figure 27 This is a schematic diagram illustrating the process of constructing and optimizing the industry chain map in this invention; Figure 28 This invention employs weighted processing of the relationships between entities in the industrial chain and designs an evaluation index table; Figure 29 This is a schematic diagram of the new energy vehicle industry chain from a microscopic perspective, as presented in this invention. Figure 30 This is a schematic diagram of the industrial chain distribution from a macroscopic perspective, as presented in this invention. Figure 31This is a flowchart illustrating the process of quality evaluation in the supply chain design of this invention. Figure 32 Output the industry chain result table for this invention; Figure 33 This is the table showing the basis for assigning weights to chain nodes in this invention; Figure 34 This is a summary table of node hierarchy statistics for this invention; Figure 35 This is a quality quantitative evaluation index table for the hierarchical division of the industrial chain nodes in this invention; Figure 36 This is a schematic diagram of the modular structure of the industrial chain design device of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] See Figure 1 As shown, an industry chain design method includes: screening industry nodes from a constructed multi-source heterogeneous dataset, identifying potential chain nodes based on an ideal value strategy, completing and optimizing the node graph, and completing the industry chain design based on full life cycle analysis.

[0026] See Figure 2 As shown, in a specific embodiment, the data sources of the multi-source heterogeneous dataset include open-source or closed-source industry, enterprise, market, sector, product data and product functional semantic data; the multi-source heterogeneous data must contain the mapping relationship between industry, product and functional semantic data; the construction process of the dataset is to preprocess the original data, perform field mapping and format alignment according to a unified knowledge ontology, and complete data fusion and standardization.

[0027] It should be noted that a unified knowledge ontology refers to a standardized data framework predefined to achieve data fusion, which clarifies the entity types, attribute fields, relationship definitions, and format specifications of the data, serving as a unified reference standard for aligning heterogeneous data.

[0028] Field mapping and format alignment: This refers to associating various fields in heterogeneous data according to the definition of a unified knowledge ontology (field mapping) and converting data of different formats into a unified format specified by the ontology (format alignment). This is the core step in achieving data fusion.

[0029] In one specific embodiment, the multi-source heterogeneous data originates from different acquisition devices, systems, platforms, or channels. Heterogeneity refers to data that differs fundamentally in type, structure, scale, and semantics. The open-source or closed-source data includes industry-related documents, company annual reports, and industry analysis reports, etc.

[0030] The multi-source heterogeneous data can be obtained from legitimate channels such as industry reports (PDF), government open data (CSV), corporate websites (HTML), and academic papers (text). The specific implementation process is as follows: 1. Data source selection: Determine legitimate data sources, including the National Enterprise Credit Information Publicity System, CNINFO, Tianyancha public interface, etc.; 2. Anti-crawling strategy design: Adopt IP proxy pool rotation, random switching of User-Agent, and CAPTCHA recognition technology (including OCR recognition or API call); 3. Data crawling: Use the Scrapy framework to write crawler scripts, based on preset field extraction rules, to selectively crawl corporate credit information (including but not limited to registered capital, credit rating) and annual report data (including but not limited to financial indicators, shareholder structure).

[0031] The functional semantic data includes, but is not limited to, data describing the functions, associated objects, attributes, and relationships of industries, sectors, and products.

[0032] The mapping relationship between the data is as follows: Figure 3 As shown, data from different sources can be mapped and connected through string matching, vector similarity calculation, semantic association analysis, etc. This mapping relationship can be persistently stored or temporarily stored; established industry and sector mapping relationships are as follows: Figure 4 As shown; taking wheels as an example, the industry mapping results of the selected vehicle body components and materials and standard classifications are as follows: Figure 5 As shown (the parts in parentheses are related to the vehicle body components or materials).

[0033] In a specific embodiment, the construction of the multi-source heterogeneous dataset can adopt multiple implementation paths, as shown in the following examples: Example 1 (Heterogeneous data preprocessing and fusion process): 1. Data cleaning: Perform deduplication, missing value imputation (using mean or median for numerical data, and manual completion for key text data) and noise data filtering on the collected raw data to eliminate data redundancy and errors.

[0034] 2. Data parsing: BeautifulSoup is used to parse HTML data, the built-in JSON library is used to parse JSON data, and PyPDF2 is used to parse PDF data to extract structured or semi-structured information.

[0035] 3. Data Fusion: Map heterogeneous data to a unified knowledge ontology (defining industry entities, attributes, and relationship standards) to achieve information alignment across data sources.

[0036] 4. Data standardization: Unify data format, date format is YYYY-MM-DD, monetary unit is uniformly converted to ten thousand yuan / one hundred million yuan, and encoding format is uniformly UTF-8.

[0037] 5. Intermediate storage: Store the preprocessed data in a data lake or relational database for subsequent knowledge extraction and graph construction.

[0038] Example 2 (Entity Linking and Core Entity Filtering Process): 1. Data Cleaning: Perform deduplication, missing value filling, and outlier filtering on the basic data to ensure data integrity.

[0039] 2. Data standardization: Unify entity naming conventions, data formats, and units of measurement to eliminate ambiguity in expression.

[0040] 3. Preliminary entity identification: Based on rule-based or industry dictionary-based NER methods, candidate entities such as enterprises, products, and technologies are extracted.

[0041] 4. Entity Linking: Map candidate entities to existing authoritative knowledge bases such as Qichacha to achieve entity normalization.

[0042] 5. Candidate entity screening: retain core entities in the industrial chain (such as core enterprises and key products) and eliminate irrelevant and redundant entities.

[0043] Example 3 (Model-Driven Entity Relationship Extraction Process): 1. Text Preprocessing: Convert text data such as annual report summaries and credit reports into the model input format, and perform word segmentation, part-of-speech tagging, and stop word removal.

[0044] 2. Model selection: BERT-based NER models (such as ERNIE-3.0) are used for entity extraction, and REBERT models are used to extract the relationships between entities.

[0045] Training data annotation: Manually annotate the enterprise entities (names and types) and their relationships (shareholders, partners, and suppliers) in the text to build an annotated dataset.

[0046] Model training and tuning: Train the model using labeled data, set the learning rate range, and the batch size to 16 or 32. Monitor the performance and iteratively tune the parameters using the validation set F1-score.

[0047] Entity and Relationship Extraction: Input the preprocessed text into the trained model to extract standardized enterprise entities (such as "Huawei Technologies Co., Ltd.") and relational triples (such as "Huawei-Supplier-CATL").

[0048] It should be noted that the BERT-based NER model refers to a named entity recognition model that is improved based on the BERT pre-trained model. By capturing the semantic information of the text context, it can accurately identify and extract the core entities in the text, such as company names, product names, and technical terms.

[0049] Relationship extraction model: refers to an algorithm model used to identify and extract the relationships between core entities in a text, such as shareholder relationships, cooperative relationships, and supply relationships. After training and optimization, the automatic extraction of relationships can be achieved.

[0050] Annotated data: refers to text data that has been manually or semi-automatically annotated with core entities and entity relationships. It is used to train and optimize BERT-based NER models and relation extraction models to improve the model's extraction accuracy.

[0051] In one specific embodiment, the raw data preprocessing can also be performed as follows: Figure 6 The process shown includes steps such as data acquisition, NLP processing, pattern mining, and knowledge storage, to extract and store components or functional entities and initial relationships.

[0052] After the multi-source heterogeneous data preprocessing is completed, the specific steps for unified knowledge ontology design and initial knowledge graph construction are as follows: 1. Define entity, relation and attribute types: Clarify entity types (enterprises, products and people), relation types (shareholders, partners and suppliers) and attributes (enterprise name, establishment time and registered capital) to provide a unified standard for aligning entities and relations across data sources.

[0053] 2. Triple Transformation: Converts entities and relations extracted by NER into triples (subject-relation-object) that conform to the RDF specification. An example is <Huawei, Supplier, CATL>.

[0054] 3. Database Import: Use Neo4j's Cypher statement to import triple data in batches. For example, write the statement CREATE(a:Company{name:"Huawei"})-[r:Supplier]->(b:Company{name:"CATL"}) to realize relational entry.

[0055] 4. Initial graph generation: Based on the imported data, construct a visual graph of enterprise entity nodes, relationship edges and attributes. Gephi or Neo4j Browser can be used to visualize the graph, intuitively presenting the entity relationship logic, which is convenient for subsequent graph verification and optimization.

[0056] The aforementioned integration and standardization: Existing knowledge can be categorized and organized according to dimensions such as industry, sector, product, and technology, and cross-data source relationship mappings can be established (such as semantic associations between industry classifications and product functions) to achieve data integration and standardization. Specific relationships are as follows: Figure 7 As shown; the visualization effect of the merged dataset is as follows. Figure 8 As shown.

[0057] See Figure 9 As shown, in a specific embodiment, the selection of industry nodes is completed based on a set industry knowledge framework to complete node selection and optimization correction; the industry knowledge framework revolves around the generation goals or needs of the industrial chain, and is a combination of one or more knowledge systems among the categories, products, technologies and parameters set from the perspective of industry, sector or field characteristics; the selection is based on the industry knowledge framework selected by the user, combined with historical industry data statistics and weight analysis to determine a preset knowledge set; the optimization correction includes manual review correction or machine learning model correction based on feedback data, and includes the construction of potential industrial chain levels and the generation of knowledge combination of industry direction nodes.

[0058] In a specific embodiment, the set industry knowledge framework refers to a predefined industry chain knowledge graph structured framework, which clearly defines the entity types (such as enterprises, products and technologies), relationship types (such as supply and cooperation), and attribute fields of various entities (such as enterprise establishment time and product specifications).

[0059] It should be noted that entity type refers to the classification of core elements in the industry chain knowledge graph, such as enterprise entities, product entities, technology entities, and industry link entities, which are key identifiers for distinguishing different core elements.

[0060] Relationship type: refers to the type of association between different entities in the graph, such as upstream and downstream supply relationships, enterprise cooperation relationships, and technology application relationships, which are used to represent the business logic associations between entities.

[0061] Attributes: These are specific information items that describe the inherent characteristics of an entity, such as the registered capital, main business, and establishment time of a business entity, and the functional characteristics and specifications of a product entity. They are the core information carriers of an entity.

[0062] Triples are standardized data formats used to represent relationships between entities. They consist of three parts: "subject-relationship-object", such as <Company A-Supplier-Company B>. They are the core data units for knowledge graph storage and representation.

[0063] Initial graph: refers to the basic knowledge graph generated by converting the extracted entities and relations into triples and importing them into a graph database. It includes core entities and relationships, but the deep associations and optimized structure have not yet been completed.

[0064] In a specific embodiment, based on the completed industrial chain knowledge graph, the acquisition of the industrial knowledge framework specifically includes the following steps: 1. Determine the industrial chain level and core variables: Divide the industrial chain level into upstream raw materials, midstream manufacturing, and downstream sales, and define the core variables (market share, growth rate, and capacity utilization rate) for each level to provide a structured standard for framework construction.

[0065] 2. Data Extraction: Using Neo4j's Cypher query statement, enterprise data at the corresponding level (such as production data of upstream lithium mining companies and capacity data of midstream battery manufacturing companies) are accurately extracted from the graph database.

[0066] 3. Parameter filling: The extracted standardized data is filled into the specified positions of the industry knowledge framework according to the correspondence between hierarchy and variables to form structured framework data.

[0067] 4. Framework Validation: First, the script automatically checks the integrity and logical consistency of variables (e.g., upstream raw material output ≥ midstream manufacturing capacity), and then manually reviews and verifies abnormal data to ensure the rationality and accuracy of the framework data.

[0068] In a specific embodiment, based on the industry knowledge framework constructed above, the construction of the industry keyword library and the calculation of the enterprise industry matching score specifically include the following steps: 1. Construction of the industry keyword library: Based on the core links of the industry knowledge framework, determine the core keywords of the target industry (such as "lithium battery, charging pile and autonomous driving" for new energy vehicles); assign differentiated weights according to the core degree of the corresponding links of the keywords (such as lithium battery 0.5, charging pile 0.3 and autonomous driving 0.4) to provide a quantitative standard for subsequent matching.

[0069] 2. Enterprise Data Field Extraction: Extract core text fields such as business description, product name, patent abstract, and main business scope of enterprises from the industry chain knowledge graph.

[0070] 3. Keyword matching: The TF-IDF algorithm is used to calculate the frequency and importance of each core keyword in the enterprise text field, and to quantify the relevance of the keywords to the enterprise business.

[0071] 4. Cumulative score calculation: The weights of the matched keywords are summed to obtain the enterprise industry matching score (e.g., enterprise A matches lithium battery (0.5) + autonomous driving (0.4) = 0.9). The higher the score, the higher the fit between the enterprise and the target industry.

[0072] It should be noted that a graph database refers to a database specifically designed for storing and managing graph data such as entities, relationships, and attributes. It supports efficient relational queries and graph traversal, and is adapted to the structured storage requirements of knowledge graphs.

[0073] Assigning weights: High weights are assigned to feature words corresponding to core links of the industry (such as core technologies and main products), and low weights are assigned to feature words with weak auxiliary and related characteristics. The weight gradient is refined with reference to the hierarchical relationship of industry standards and industry knowledge framework. The settings are made by professionals according to industry needs, and no specific numerical restrictions are imposed here.

[0074] TF-IDF stands for Term Frequency-Inverse Document Frequency. It quantifies the importance of keywords to text by calculating the frequency of keyword occurrences in target text (term frequency) and their scarcity in all texts (inverse document frequency). It is the core algorithm for calculating the matching degree between enterprises and industries.

[0075] In a specific embodiment, the industry knowledge framework can also be obtained through the following methods: using large models, web retrieval tools, deepResearch tools, agent tools, or other model services to collect industry knowledge through question-and-answer interaction, keyword retrieval, intelligent recommendation, etc.; or by resolving conflicts, removing redundancies, and fusing information from the return results of multiple models to form a logically consistent and comprehensive feasible industry knowledge framework.

[0076] The selection process, specifically entity or relation matching and cumulative score calculation, includes: setting differentiated weight values ​​for target keywords in each field based on the priority of core links in the industry (keywords for core links have higher weights than keywords for auxiliary links); calculating the cumulative score by weighted summation of the number of matched keywords and their corresponding weight values ​​in the statistics, thereby quantifying the fit between entities or relations and the industry knowledge framework.

[0077] The variable parameter filling includes setting at least one variable parameter based on the characteristics of the industry sector: industry classification parameters (such as industry category code and sub-sector identifier), product parameters (such as product specifications and functional characteristics indicators), and technical parameters (such as technology maturity level and performance compliance standards).

[0078] In one specific embodiment, the matching threshold is a preset score threshold, which is determined in the following way: based on the core requirements of the industry knowledge framework, combined with the statistical analysis results of historical industry data, and after multiple cross-validations and expert review and calibration, the critical score used to screen qualified nodes is finally determined.

[0079] The following specific implementation examples can be adopted in the process of selecting industry nodes and matching relationships: Example 1 (Application of supply chain knowledge throughout the entire process): 1. Needs analysis: Clarify the goals of supply chain analysis (such as risk warning and collaborative optimization) to provide guidance for subsequent data collection and knowledge processing.

[0080] 2. Multi-source data collection: Collects various types of data, including company annual reports, industry reports, supply chain data, and public APIs, to ensure that the data covers the core links of the upstream and downstream of the industry chain.

[0081] 3. Knowledge Extraction: Using rule templates or BERT-NER models, entity extraction, relation extraction, and attribute extraction are carried out to extract the core elements of the industry chain.

[0082] 4. Knowledge Integration: Through entity alignment, conflict resolution, and knowledge supplementation, data redundancy and contradictions are eliminated, forming a unified industry chain knowledge system.

[0083] 5. Knowledge Storage: Select storage media based on business needs, such as graph databases like Neo4j to store relationships, or relational databases like MySQL to store structured data.

[0084] 6. Knowledge Application: Apply the constructed knowledge system to scenarios such as decision support, risk warning, and quality evaluation to empower intelligent analysis of the industrial chain.

[0085] Example 2 (Semantic Vector Construction and Storage): 1. Domain Corpus Preparation: Collect industry chain-related texts (technical documents and corporate reports), build a domain-specific corpus, and provide data support for model training.

[0086] 2. Model selection: Pre-trained models such as Sentence-BERT and RoBERTa are selected to meet the semantic feature extraction requirements of industry chain text.

[0087] 3. Model fine-tuning: Optimize model parameters using labeled industry chain data (entity and relationship labeled data) to improve vector generation accuracy.

[0088] 4. Sentence Vector Generation: Input the preprocessed text into the model and output a fixed-dimensional semantic vector (e.g., 768-dimensional).

[0089] 5. Vector storage: Use FAISS or Milvus to build vector indexes to improve the efficiency of subsequent similarity queries.

[0090] Example 3 (Entity Relationship Matching and Validation): 1. Obtain Entity Text: Collect the descriptive text of the project or item (product manual, business introduction).

[0091] 2. Generate sentence vectors: Call the fine-tuned pre-trained model to generate semantic vectors for the text.

[0092] 3. Calculate similarity: Use cosine distance to calculate semantic relevance (suitable for high-dimensional vectors) or Euclidean distance (suitable for low-dimensional features).

[0093] 4. Threshold filtering: Retain high similarity relationships with similarity ≥ preset threshold (e.g., 0.7).

[0094] 5. Manual verification (optional): Manual sampling verification is conducted for high-value relationships such as core suppliers and key technology collaborations.

[0095] 6. Relationship storage: Store the verified entity relationships into the graph database.

[0096] Example 4 (Technology entity and item association matching): 1. Technology data collection: Collect data such as patent documents and technical standards to obtain relevant information about technology entities.

[0097] 2. Technical Entity Recognition: The NER model is used to extract technical terms such as "solid-state battery technology" and "autonomous driving algorithm".

[0098] 3. Vector matching: Calculate the semantic vector similarity between item entities and technology entities to identify potential associations.

[0099] 4. Relationship type determination: Determine the relationship type (e.g., "item-application-technology" or "item-adoption-technology") using rule templates or relationship extraction models.

[0100] 5. Verification and Storage: After manual or third-party data verification, the relationships are stored in the database.

[0101] Example 5 (Construction and Maintenance of Industry Chain Knowledge Graph): 1. Demand Schema Design: Define the entity types (enterprises, products), relationship types (supply, cooperation), and attributes (registered capital, product specifications) of the industry chain.

[0102] 2. Multi-source data collection: Collect structured data (annual reports), semi-structured data (official website information), and unstructured data (academic papers) to ensure data diversity.

[0103] 3. Data cleaning and fusion: The collected data is deduplicated, missing values ​​are filled in, and the format is standardized to eliminate data conflicts.

[0104] 4. Entity Link Extraction: Align entities with the same name (e.g., unify "Huawei Technologies Co., Ltd." as "Huawei") and extract the relationships between entities.

[0105] 5. Knowledge Storage: Neo4j is used to store knowledge graph data, supporting related queries and visualization.

[0106] 6. Update and maintain: Regularly crawl incremental data (such as new business partnerships) on a daily or weekly basis and dynamically update the map.

[0107] 7. Interface Development: Develop RESTful APIs or SPARQL interfaces for industry chain decision-making queries.

[0108] Example 6 (Enterprise Industry Matching Screening): 1. Threshold setting: Based on the industry knowledge framework requirements and combined with historical data statistical analysis, determine the matching threshold (e.g., the threshold for the artificial intelligence industry is 0.6).

[0109] 2. Score Ranking: The industry matching scores of enterprises are sorted in descending order to highlight enterprises with high matching scores.

[0110] 3. Preliminary screening: Retain companies with scores ≥ the threshold to form a preliminary candidate list.

[0111] 4. Manual review (optional): Sample verification of the preliminary screening results (sampling ratio 30%), and adjust the threshold based on the results.

[0112] 5. Industry list generation: Outputs a list of eligible companies to provide a basis for selecting nodes in the industrial chain.

[0113] It should be noted that entity alignment refers to the operation of matching and associating entities extracted from the model with standard entities in the existing knowledge base to unify entity representations.

[0114] Conflict resolution: refers to the operation of identifying multiple different entities corresponding to the same name extracted from the model, combining contextual semantics and existing knowledge base information to determine the true target of the entity.

[0115] The cumulative score is calculated based on the number of keywords matched in the data and their corresponding weight values: Using core feature words in the industry knowledge framework as a basis, the frequency of each word in enterprise texts (such as main business and patent abstracts) is counted, and the ratio of this frequency to the total number of words in the enterprise text is calculated to obtain the term frequency (TF); the number of enterprise texts containing this feature word is counted, and the ratio of the total number of enterprise texts to the number of texts containing this word is calculated and its logarithm is taken to obtain the inverse document frequency (IDF); the term frequency and the inverse document frequency are multiplied to obtain the initial matching coefficient, which is then normalized to generate the standardized matching frequency; the matching frequency of each feature word is multiplied by its corresponding differentiated weight, and the sum is the quantified score for the matching between the enterprise and the industry.

[0116] The matching threshold is a critical value used to determine whether a company's cumulative score in industry matching is qualified. It is set by professionals according to industry needs, and no specific numerical limit is set here.

[0117] In a specific embodiment, the optimization correction specifically includes the following: 1. Basic correction: manual review correction or machine learning model correction based on feedback data is used to ensure the accuracy of candidate node data.

[0118] 2. Construction of Hierarchical and Knowledge Combinations: Complete the construction of potential hierarchical levels of the industrial chain model and generate knowledge combinations of industrial direction nodes to improve the structural logic of the industrial chain.

[0119] 3. Graph generation: Machine learning is used to identify relevant results from the target text, and the identified results are filled in to generate an industry chain graph.

[0120] In one specific embodiment, the machine learning includes a supervised learning model (such as BiLSTM-CRF and BERT-NER) or a semi-supervised learning model. The supervised learning model is suitable for scenarios with sufficient labeled data, while the semi-supervised learning model is suitable for scenarios with scarce labeled data. Both are used to extract industry node names, inter-node relationships, and relationship strength information from the target text.

[0121] In a specific embodiment, the specific implementation examples of triple extraction, model training, cluster optimization and cross-validation in the optimization and correction process are as follows: Example 1 (Triple Extraction and Graph Update Process): 1. Triple Extraction: Extract subject-verb-object triples (such as <CATL, supply, Tesla>) from the target text to provide standardized basic data for graph update.

[0122] 2. Verification: Perform logical consistency verification on the extracted triples to eliminate contradictory or redundant relationships.

[0123] 3. Graph Update: Add the verified new triples to the industry chain knowledge graph to improve the node relationships.

[0124] 4. Production control: Set the map update frequency (daily incremental update, weekly full update) and establish a quality monitoring mechanism (e.g., triplet accuracy ≥ 95%).

[0125] 5. Interface Development: Develop SPARQL query interfaces or RESTful APIs to support the calling and application of map data.

[0126] Example 2 (Target Text Processing and Model Training Process): 1. Target Text Preprocessing: Perform cleaning, word segmentation and stop word removal operations on industry-related texts to reduce noise data interference.

[0127] 2. Construction of labeled datasets: Entities, relationships and events in text are manually labeled (e.g., CATL, supply, Tesla) to build datasets for model training and validation.

[0128] 3. Feature Engineering: Use TF-IDF algorithm or BERT embedding technology to extract text features and improve the model's recognition accuracy.

[0129] 4. Model Training: Named Entity Recognition (NER) uses BiLSTM-CRF or BERT-NER model, and relation extraction uses RE-BERT model; the learning rate is set and the batch size is 16 / 32 for training.

[0130] 5. Model evaluation and optimization: The model performance is evaluated using accuracy and F1-score as the core indicators, and the parameters are iteratively adjusted to optimize the model.

[0131] 6. Recognition Result Output: Outputs standardized entity, relation, and event triples for supplementing and updating the industry chain map.

[0132] Example 3 (Feature Engineering and Clustering Optimization Process): 1. Feature Extraction: Extract core attribute features of enterprises (such as revenue growth rate, registered capital, and industry keyword matching degree) from the graph database, and construct a clustering analysis feature set.

[0133] 2. Clustering algorithm selection: Select an appropriate algorithm based on the data distribution characteristics. When the data is distributed in spherical clusters, choose the K-means algorithm; when the data is distributed in uneven density, choose the DBSCAN algorithm.

[0134] 3. Model training: The clustering model is trained using the Scikit-learn library, and the optimal number of clusters for the K-means algorithm is determined by the elbow rule.

[0135] 4. Adding clustering results: Add the cluster label to which the enterprise belongs as a new attribute to the graph entity to achieve the classification management of nodes.

[0136] 5. Graph Update: Synchronously update enterprise attribute information in the graph database and improve the multi-dimensional features of nodes.

[0137] Example 4 (Cross-validation and parameter optimization process): 1. Consistency verification: Compare the preliminary industry list with authoritative third-party industry reports such as iResearch Consulting to verify the consistency of the node selection results.

[0138] 2. Data Correction: Manually verify and correct inconsistent data (such as changes in the company's business scope or updates to cooperative relationships) to ensure data authenticity.

[0139] 3. Parameter optimization: Adjust keyword weights or matching thresholds based on the verification results (e.g., increase the weight of "autonomous driving" from 0.4 to 0.5) to improve the accuracy of subsequent screening.

[0140] 4. Output Results: Generate a final industry data report that includes a list of companies, a visualized industry map, and statistics on key indicators.

[0141] 5. Map Synchronization: The corrected enterprise data is synchronized and updated to the map database to ensure the timeliness and accuracy of the map data.

[0142] It should be noted that cross-validation refers to the operation of comparing and verifying the preliminary candidate list with independent data sources such as authoritative third-party industry reports and government industry filing lists. This is used to screen for industries with a matching degree less than a preset matching screening threshold and improve the accuracy of the candidate list.

[0143] Adjusting keyword weights: This refers to the operation of optimizing and correcting the quantitative importance score of core keywords based on the results of manual review and cross-validation, in order to improve the accuracy of subsequent cumulative score calculations and reduce misjudgments.

[0144] Clustering algorithms: These are machine learning algorithms that automatically classify enterprise entities with common attributes into different categories based on the similarity of entity attributes. They include one or more combinations of algorithms such as K-means, hierarchical clustering, or DBSCAN.

[0145] Clustering algorithm selection: This refers to selecting suitable algorithms by setting specific numerical thresholds based on quantifying the data cluster shape fit, density matching, cluster number fit, and outlier tolerance. Specifically, the K-means algorithm must meet the following requirements: cluster shape fit ≥ 0.8, density matching ≥ 0.7, cluster number fit ≥ 0.9, and outlier tolerance ≥ 0.6; the DBSCAN algorithm must meet the following requirements: cluster shape fit ≥ 0.7, density matching ≥ 0.6, cluster number fit ≥ 0.8, and outlier tolerance ≥ 0.9; and the hierarchical clustering algorithm must meet the following requirements: cluster shape fit ≥ 0.7, density matching ≥ 0.6, cluster number fit ≥ 0.8, and outlier tolerance ≥ 0.8.

[0146] Clustering parameter thresholds are critical values ​​used to determine whether the preset core parameters of the clustering algorithm are qualified during training and application. They are set by professionals according to industry needs, and no specific numerical restrictions are imposed here.

[0147] In a specific embodiment, the following are examples of machine learning and model application generated by the potential hierarchical construction of the industrial chain model and the combination of knowledge of industrial direction nodes in the optimization and correction: Example 1 (model framework design and iteration process): 1. Model framework design: Select a hierarchical or network structure (such as the photovoltaic industrial chain hierarchy: silicon material → silicon wafer → cell → module → application) to provide a structured foundation for the construction of the industrial chain hierarchy.

[0148] 2. Node attribute definition: Clearly define node types (industrial links, enterprises and technologies) and core attributes (production capacity, output value and technology level), and standardize the unified standards of node data.

[0149] 3. Relationship rule construction: Define the rules for determining upstream and downstream, competitive and cooperative relationships (such as "upstream nodes provide core raw materials to downstream nodes"), and clarify the relationship logic.

[0150] 4. Knowledge combination data integration: The selected node knowledge combinations are mapped to the preset model framework according to the relationship rules to form a preliminary hierarchical structure.

[0151] 5. Model Validation and Iteration: Based on expert feedback and actual industry data, adjust the framework structure or relationship rules to ensure the rationality of the hierarchy.

[0152] Example 2 (Node Extraction and Knowledge Combination Generation Process): 1. Defining the Boundaries of the Industrial Chain: Clarify the research scope (such as the new energy vehicle industrial chain, including upstream raw materials to downstream after-sales service) and lock in the application boundaries of the technical solution.

[0153] 2. Multi-source data collection: Collect various types of data, including policy documents, corporate annual reports, industry reports, and patent documents, to ensure data coverage of the core links of the entire industry chain.

[0154] 3. Node entity extraction: Based on the BERT-NER model or rule template, extract key nodes such as industry links, core enterprises, products or technologies.

[0155] 4. Node Relationship Mining: Through cosine similarity semantic analysis and Apriori association rule mining, the supply and demand, upstream and downstream, and competitive relationships between nodes are identified.

[0156] 5. Generation of industry-oriented node knowledge combinations: Based on industry logic or semantic similarity, relevant nodes are combined to form targeted knowledge sets.

[0157] 6. Knowledge combination verification and optimization: After expert review (pass rate ≥90%) and comparison with actual industry data, the node combination and correlation relationship are adjusted.

[0158] Example 3 (Graph Construction and Quality Verification Process): 1. Identification Result Parsing: Transform the results of node extraction and relation mining into triple (subject-relation-object) or event tuple format to standardize the data structure.

[0159] 2. Graph node or relation mapping: Match the parsing results with existing model nodes to achieve normalization of node representation (e.g., unify "Huawei Technologies Co., Ltd." to "Huawei").

[0160] 3. Conflict detection and handling: Data conflicts such as duplicate nodes and contradictory relationships are resolved through rule-based judgment or voting mechanisms.

[0161] 4. Graph Filling and Update: Add nodes or relationships to the initial graph and update node attributes (such as enterprise capacity and technology level).

[0162] 5. Visualized storage of the graph: Neo4j is used to store the graph data, and Gephi is used to visualize it, making it easy to analyze the hierarchical structure intuitively.

[0163] 6. Spectrum quality verification: The reliability of the spectrum data is ensured by expert review and rule check (no logical contradictions).

[0164] In one specific embodiment, the industry chain model is generated through the combination of potential hierarchy construction and industry direction node knowledge. It can also utilize community discovery to achieve specific examples of industry chain node clustering or community discovery effects, such as... Figure 10 As shown, different industrial chain node communities of 0-3 can be formed, and a preliminary industrial chain map can be formed after optimization and correction.

[0165] See Figure 11 As shown, in a specific embodiment, the ideal value strategy is used to train a learning model and identify potential chain nodes. It includes at least one or a combination of learning reward function, energy index function, and value function, and optimizes the selection strategy of node knowledge combination through multiple iterations of training. The learning model includes various models or combinations of models such as machine learning, deep learning, or reinforcement learning.

[0166] Taking reinforcement learning as an example, a method for calculating the ideal value strategy is designed as follows: Figure 12 As shown, this method integrates configuration synergy benefits (such as the uniformity and coverage breadth of node combinations) with the intrinsic value of components through a benefit function, while introducing a configuration cost model to penalize unreasonable layouts (such as excessive redundant nodes); the learning process selects optimal paths in the value space and policy space, as shown in the figure. Figure 13 As shown, the optimal node combination is explored iteratively through the IPO Agent to maximize both ideality and spatial coverage; the reinforcement learning strategy parameters based on ideal value are set as follows. Figure 14 As shown, the value ranges of key parameters, including learning rate, reward coefficient, and iteration steps, are adapted to the needs of different industry scenarios.

[0167] Comparing the industry chain generation effects of different learning methods, such as Figure 15 As shown, the selection method of the ideal value strategy of this invention improves the ideality of the combination of industry chain nodes compared with random selection, and reinforcement learning can be preferred as the core learning method.

[0168] Comparing the training results of different reinforcement learning models, as follows: Figure 16 As shown, the convergence of the model is verified by the reward history and training loss curve. The ideality score of the node combination can be combined to select the optimal potential chain nodes. Taking the new energy vehicle industry chain as an example, the potential chain nodes output by the model include DC fast charging interface, four-wheel drive transfer case, permanent magnet synchronous motor, battery pack, DC converter, electronic brake force distribution and four-wheel steering system and wire connector. The above nodes have been verified by experts and compared with industry data to confirm them as potential chain nodes that meet the core needs of the industry chain.

[0169] The above method, through training and comparing the effects of reinforcement learning models, can not only evaluate the effect of reinforcement learning methods on the generation of industrial chains, but also select the best learning method to improve the accuracy and rationality of potential chain node selection.

[0170] The confirmation of potential chain nodes and combinations is as follows: Figure 17 As shown; the preliminary effect of the resulting industry chain map is as follows. Figure 18 As shown.

[0171] It should be noted that reinforcement learning models refer to machine learning models that optimize decision-making strategies through iterative cycles of "state-action-reward". The core idea is to allow the model to learn the optimal behavior through interaction with the environment, which in this case is the node selection strategy.

[0172] State: refers to the description of the current environment in the reinforcement learning model, specifically the current combination of industry node knowledge, including complete information such as node composition and relationships.

[0173] Actions: refer to the decision-making operations that a reinforcement learning model can perform. Here, it specifically refers to the operation of adding or deleting nodes in the knowledge combination of industry nodes, which is used to adjust the node combination structure.

[0174] Reward: refers to the quantization function used in reinforcement learning models to evaluate the effect of "actions". Here, the output value of the ideality policy function is directly used. The higher the ideality of the node combination after the action is executed, the higher the reward value.

[0175] In a specific embodiment, the ideal value strategy can also be a multi-objective value function, which performs multi-objective optimization and screening of potential industry chain node combinations. The multi-objective value function includes modeling multi-dimensional task objectives such as the value of chain nodes, the relevance, coverage, accuracy, completeness, foresight, and timeliness of node combinations, defining quantitative weights for the corresponding indicators of each objective, and constructing a weighted value function. The multi-objective optimization includes at least optimizing the coverage and ideal value of node knowledge combinations, and the proportion of indicators can be dynamically adjusted.

[0176] When the ideal value strategy employs a multi-objective value function, if forward-looking perspective is the core objective, it can be combined with, for example... Figure 19 The value assessment of the forward-looking characteristics of the chain nodes shown (such as technological innovation and adaptability to future market demand) is carried out. The specific optimization process is as follows: 1. Determine the multi-task objectives: Clarify the multi-dimensional optimization objectives, including coverage (complete coverage of core links), relationship accuracy (authentic node association), model simplicity (no redundant nodes) and forward-looking (adapting to future industry development trends), among which forward-looking is the core optimization objective.

[0177] 2. Quantification of Ideality Indicators: Each target indicator is quantitatively defined as follows: Coverage = Number of core links covered by potential chain nodes ÷ Total number of core links; Relationship Accuracy = Number of correct associations verified by a third party ÷ Total number of associations; Model Simplicity = 1 - Number of redundant nodes ÷ Total number of nodes (redundant nodes refer to auxiliary nodes that are not directly related to the core target); Foresight = Number of potential core nodes included ÷ Total number of potential chain nodes (potential core nodes refer to nodes that conform to the development trend of the industry in the next 3-5 years, such as emerging technology companies and nodes related to future market demand).

[0178] 3. Policy Function Construction: Constructing the weighted composite score function: = ×Coverage+ ×Accuracy+ ×Simplicity+ × Forward-looking; among which , , and These are weighting coefficients, summing to 1, when forward-looking perspective is the core objective. The value range is 0.3-0.5, and the remaining weights are dynamically allocated according to industry needs (e.g., when emphasizing technological innovation, the forward-looking weight can be increased, and when emphasizing stability, the relationship accuracy weight can be increased).

[0179] 4. RL Model Design: Define the state as the current potential chain node combination (including node composition and association relationships), the action as the node addition or deletion operation, and the reward as the ideality score calculated by the policy function; the model adopts the DDPG or PPO reinforcement learning algorithm, sets the number of iterations to 500-3000, and the convergence condition is that the ideality score fluctuation is ≤5% over 100 consecutive iterations to ensure that the model is trained sufficiently.

[0180] 5. RL Training and Selection: Input the quantified index data into the RL model, and select the node combination with the highest ideality score as the optimal potential chain node combination through multiple iterations of optimization; monitor the change of reward value in real time during training, and adjust the model parameters (such as learning rate and batch size) if the reward value continues to decline.

[0181] 6. Function Iteration Optimization: The weight coefficients are dynamically adjusted based on feedback data, which includes expert evaluation results, actual industry application data (such as the subsequent market performance of potential nodes and the implementation of technologies), and model prediction accuracy. If the forward-looking objectives are not met (such as the proportion of potential core nodes <20%), the weights are increased by 0.05-0.1 until the preset objectives are met, thus achieving dynamic adaptation of multi-objective optimization.

[0182] It should be noted that coverage refers to the degree to which the selected industry node knowledge combination covers the core links and key areas of the target industry chain. Coverage = number of core links covered divided by the total number of core links. First, the core links are defined as the key processes or areas that support the normal operation of the industry chain and determine the core value of the industry. Then, the number of coverages is counted: from the selected candidate nodes, each node is determined to correspond to a certain core link, and finally the total number of nodes that accurately match the core links is counted.

[0183] Relationship accuracy: refers to the degree to which the relationships between nodes (such as supply, cooperation and competition) in the knowledge combination of industry nodes match the actual industry scenario. Relationship accuracy = number of correct relationships divided by total number of relationships. The criteria for judging correct relationships are: verification through authoritative third-party data (government filing directory, enterprise announcement and authoritative industry report) or confirmation by manual verification to ensure that the relationship is real (e.g., company A is indeed a core supplier of company B) and that there is no mismatch in type (e.g., a competitive relationship is not mistakenly judged as a cooperative relationship).

[0184] Simplicity: Under the premise of satisfying coverage (complete coverage of core links) and relationship accuracy (real and reasonable association), the degree of simplification of the knowledge combination of industry nodes is judged by calculation to determine whether the node combination is redundant: redundant node ratio = number of redundant nodes divided by total number of nodes (redundant nodes refer to auxiliary nodes that are not directly related to core links).

[0185] Foresight: This refers to the ability of the selected industry node knowledge combination to predict and incorporate future industry development trends and potential nodes. It is calculated by dividing the proportion of potential core nodes by the number of potential nodes included and the total number of potential chain nodes (potential nodes refer to emerging technology companies that conform to future industry trends and nodes related to future market demand, etc.).

[0186] Weight: refers to the quantitative parameter used to adjust the proportion of each evaluation indicator in the ideality strategy function. It is set by professionals according to industry needs, and no specific numerical limit is imposed here.

[0187] Adjust weights based on feedback: After each reinforcement learning model training, calculate the deviation between the actual performance of each indicator and the preset target. For example, if the actual coverage value is 80%, the target value is 90%, and the deviation value is 10%, adjust the weight coefficients according to the deviation value. During the adjustment process, avoid excessive weight of a single indicator that could lead to function imbalance. For example, the weight of a certain indicator should not exceed 0.5 and should not be less than 0.05.

[0188] The multi-objective value function can also be implemented in the following specific ways:

[0189] in , The weighting of patents can be adjusted based on industry characteristics and research objectives. For example, if the emphasis is on the impact of technological innovation, the weighting of the number of patents can be increased accordingly.

[0190] The multi-objective optimization includes optimizing the coverage and redundancy of node knowledge combinations, and can dynamically adjust the indicator proportions to optimize node combinations. Based on multi-objective optimization, the resulting industry chain knowledge graph exhibits the following effect: Figure 20 As shown.

[0191] In a specific embodiment, the completion and optimization of the node graph is based on the upstream and downstream relationship analysis to confirm the chain nodes. Initialized industrial chain node knowledge or a set of node knowledge can be introduced and fused with existing node knowledge. The upstream and downstream relationship analysis is to label the relationship types and train the model to predict potential industrial chain node entities and entity relationships. The completion of chain node confirmation is to obtain a set of industrial chain nodes at different levels from a certain perspective.

[0192] The specific process for completing and optimizing the node graph is as follows: 1. Based on the pre-trained model, classify the nodes of the industrial chain in pairs, check cyclically whether a cycle is formed, and delete the relationship with the lowest confidence until the graph no longer forms a cycle.

[0193] 2. Find nodes with only out-degree and no in-degree and nodes with only in-degree and no out-degree as upstream and downstream nodes to be mapped.

[0194] 3. By using a graph theory-based centrality method to classify the remaining nodes into three categories, we can obtain the upstream, midstream, and downstream nodes, thus satisfying the requirements.

[0195] In a specific embodiment, the similarity between projects and item entities is determined based on the characteristics of the upstream and downstream of the industrial chain and semantic vectors, and the semantic association degree between entities is calculated using the cosine similarity algorithm; the relationship between item entities and technical data is established by matching the attribute features of item entities with the application scenarios of technical data; new ternary graph data is generated and applied to production control.

[0196] The analysis at different levels uses the cosine similarity algorithm to calculate the semantic association between entities based on the characteristics of the upstream and downstream of the industrial chain and the semantic vector to determine the similarity between project and item entities.

[0197] The relationship between the physical object and the technical data is established by matching the attribute characteristics of the physical object with the application scenario of the technical data.

[0198] After the new ternary spectral data is generated, the process also includes steps for verifying the accuracy of the spectral data and updating it in real time.

[0199] In one specific embodiment, the production control method includes: preprocessing the candidate entity data, and dynamically controlling the production process based on the new ternary graph data includes adjusting the raw material procurement quantity, optimizing the operating parameters of production equipment, or scheduling supply chain logistics.

[0200] The pre-trained vector model generates vectors that are used to improve the accuracy of determining the similarity between project and item entities.

[0201] The weighted processing of the relationships determines the weight value of the relationships based on the frequency of cooperation between entities, the degree of resource dependence, or the influence coefficient; alternatively, different weights can be assigned based on the transaction amount, cooperation duration, or risk transmission coefficient between entities. The importance score of each key element can be obtained through iterative calculation using the relationship weights and a random walk strategy, including setting initial node weights, transition probability matrices, and restart probability parameters. The relationship weights are determined based on the frequency of cooperation between entities, the degree of resource dependence, or the influence coefficient; alternatively, different weights can be assigned based on the transaction amount, cooperation duration, or risk transmission coefficient between entities.

[0202] In one specific embodiment, the construction of the key element evaluation model can be carried out by using a random walk strategy, including setting initial node weights, transition probability matrices, and restart probability parameters, and obtaining the importance score of each key element through iterative calculation.

[0203] The development quality score is obtained by evaluating key elements based on the assessment model, and then weighting and summing the importance scores of the key elements with the current development status indicators of each element. This results in a multi-perspective industrial chain design set and ranking results.

[0204] The upstream and downstream relationship analysis, and the schematic diagram of a specific embodiment for determining the relationship between nodes, are shown below. Figure 21 As shown; the completion of chain node confirmation obtains a set of industrial chain nodes at different levels from a certain perspective; the perspective can be divided according to user needs or actual needs, such as industry, sector, product, region, planning and concept, to guide the generation of industrial chain preference effects.

[0205] In a specific embodiment, the specific process of obtaining a set of industry chain nodes at different levels from a certain perspective and filtering key elements during the process of completing and optimizing the node graph is as follows: 1. Perspective node knowledge acquisition and weighted storage: Perspective judgment: Based on user needs input, perspective judgment is performed (such as macro perspective to adapt to industrial policy formulation, industry perspective to adapt to industry trend analysis, micro perspective to adapt to enterprise cooperation decision-making, product perspective to adapt to product supply chain design, domain perspective to adapt to specific technical field analysis, problem perspective to adapt to industry pain point solutions, and concept perspective to adapt to the implementation of emerging concepts), to adapt to different users' preferences for the output effect of the industry chain.

[0206] Entity recognition and extraction: The BERT-NER model, finely tuned with corpus data from the industry chain domain, is used to accurately extract core entities such as enterprises, products and technologies.

[0207] Relationship definition: Classify the relationships between entities into three categories: supply, cooperation and competition, and clarify the judgment criteria (e.g., supply relationship means that entity A provides core raw materials or products to entity B).

[0208] Weighting rules: Static weights are set based on fixed attributes such as supply ratio and cooperation depth, while dynamic weights are set based on real-time data such as transaction frequency and cooperation duration.

[0209] Dynamic weight updates: The weight values ​​are adjusted in real time based on incremental data (such as new transaction records, cooperation announcements, and policy adjustment information) to ensure the timeliness of the weights.

[0210] Weighted storage verification: Entities and relationships with weights are stored in the Neo4j graph database, and the accuracy and logical consistency of the data are verified by experts.

[0211] 2. Initial generation of the industry chain knowledge graph: Graph loading: Load the industry chain graph with weighted edges (including entity, relationship and weight information).

[0212] Parameter settings: Set the number of random walk steps (100-500 steps, adjust according to the size of the map), restart probability parameter (value range 0.1-0.3, balancing local correlation and global coverage) and starting node (select the core entity according to the perspective requirements, such as the target product as the starting node in the product perspective).

[0213] Path generation: Starting from the initial node, perform a random walk, traverse the associated nodes in the graph, and record the node visit paths.

[0214] Score Calculation: The importance score of each node is calculated by using the PageRank or Personalized PageRank algorithm, combined with the node weights.

[0215] Verification and parameter tuning: Compare the calculated score with the expert score, and adjust parameters such as the number of steps and the restart probability to ensure that the score matches the actual importance.

[0216] Output visualization: Sort by importance score in descending order, and intuitively display the importance of nodes by node size, generating a visual graph for user reference.

[0217] 3. Key element screening and decision application: Initial element selection: Select the Top N elements according to their importance scores (N is set according to the industry scale and perspective requirements, such as 5-20 elements).

[0218] Indicator system construction: A multi-level indicator system is constructed from the following dimensions: technology (technology maturity, innovation and compatibility), market (market share, growth rate and competitive intensity), supply chain (supply stability, dependence and difficulty of substitution) and financial (revenue scale, profit margin and return on investment).

[0219] Data standardization: The min-max normalization method is used to transform the index data of different orders of magnitude to the [0,1] interval to eliminate the influence of the units.

[0220] Overall score: The weights of each indicator are determined by the Analytic Hierarchy Process (AHP) or machine learning algorithms (such as random forest), and the overall score of each element is calculated using a weighted summation formula.

[0221] Grading: The elements are divided into four levels based on the overall score: ≥0.8 is excellent, 0.6-0.8 is good, 0.4-0.6 is average, and <0.4 is poor.

[0222] Feedback Application: Based on the grading results, guide the decision-making of the industrial chain, such as prioritizing the retention of excellent grade elements, optimizing the correlation of medium and lower grade elements, and supplementing missing key elements.

[0223] In a specific embodiment, the results of knowledge processing and hierarchical analysis of the industry chain nodes, such as Figure 22 As shown; a schematic diagram of the upstream and downstream results of the industrial chain is shown below. Figure 23 As shown.

[0224] The different levels of the industrial chain node set refer to the node division results at the upstream, midstream, and downstream levels of the industrial chain. The identified key node knowledge is then processed and subjected to hierarchical analysis.

[0225] The construction of the node graph, the construction of the industrial chain knowledge graph, includes collecting attribute data, business data and relationship data between entities in the industrial chain, and performing entity alignment and relationship fusion.

[0226] The processing includes constructing a knowledge graph of the industrial chain, identifying candidate industrial chain knowledge or entities, and the relationships between entities and technologies; generating new ternary graph data and applying it to production control; and preprocessing the entity data of the candidate industrial chain knowledge graph, including cleaning, deduplication, and standardization of the entity data.

[0227] Taking new energy vehicles as an example, the results of the industry chain node hierarchy are as follows: Figure 24 As shown; the prediction effect of the hierarchical relationship of the industrial chain nodes is as follows. Figure 25 As shown, a credibility assessment is given.

[0228] In a specific embodiment, the process of completing the industrial chain design based on full life cycle analysis involves dividing the life cycle into various stages around the goals or needs of the industrial chain to ensure the integrity of the industrial chain's generation.

[0229] The life cycle analysis divides the life cycle into various stages, such as industry cycle, product life, technology maturity and technology cycle. From the perspectives of production, sales, after-sales service and recycling, it can provide more detailed relationships between the various stages of the industrial chain.

[0230] Lifecycle analysis, such as utilizing product lifecycles and technology lifecycles and employing graph methods to extend downstream lifecycle nodes, can help the supply chain extend upstream and downstream, thereby ensuring the integrity of the supply chain. The specific effects of supply chain analysis include... Figure 26 As shown.

[0231] See Figure 27 As shown, in a specific embodiment, the specific process of dividing and dynamically evaluating the life cycle stages of the industrial chain in the full life cycle analysis stage is as follows: 1. Division of the life cycle stages of the industrial chain: Based on the core characteristics of the industrial chain, it is divided into four stages: the nascent stage, the growth stage, the maturity stage, and the decline stage. The core characteristics of each stage are clarified (nascent stage: immature technology, small market size; growth stage: rapid market expansion, accelerated technology iteration; maturity stage: stable market, saturated capacity; decline stage: declining demand, overcapacity), providing a classification basis for subsequent dynamic evaluation.

[0232] 2. Multi-source data acquisition and preprocessing: Collect data from all aspects of the industry chain (production data, logistics data, sales data, and R&D investment data), perform data cleaning, deduplication, and min-max normalization (mapping the data to the [0,1] interval) to eliminate noise and the influence of units, and provide standardized data for model training.

[0233] 3. Quality evaluation index system construction: Based on the four dimensions of efficiency, quality, risk and collaboration, a hierarchical evaluation index system is constructed: efficiency index (capacity utilization rate and logistics turnover speed), quality index (product qualification rate and technology maturity), risk index (supply chain dependence and market volatility coefficient) and collaboration index (frequency of upstream and downstream cooperation and data sharing rate).

[0234] 4. Intelligent evaluation of model training: Select the XGBoost algorithm or neural network model (such as CNN and LSTM), use preprocessed standardized data for model training and validation, set the learning rate and number of iterations, and optimize model parameters with evaluation accuracy as the core indicator.

[0235] 5. Dynamic monitoring and optimization: Real-time collection of new data from the industry chain and input into the trained model to update the evaluation results at each stage; setting early warning thresholds (e.g., triggering an early warning when the evaluation score is <0.5), and automatically triggering an early warning and generating preliminary adjustment suggestions when abnormal data occurs or the evaluation score is lower than the threshold.

[0236] 6. Results Visualization and Decision Support: Through the Dashboard visualization tool, the evaluation scores and core indicator performance of each life cycle stage are displayed intuitively; based on the visualization results, targeted optimization decision suggestions are output (such as increasing R&D investment during the growth stage and optimizing capacity layout during the decline stage), empowering the life cycle management of the industrial chain.

[0237] In one specific embodiment, the relationships between entities in the industrial chain are weighted, and different evaluation indicators are designed, such as... Figure 28 As shown; the results of the completed supply chain design implementation examples from different perspectives: the supply chain results formed from the micro perspective are as follows. Figure 29 As shown; from a macro perspective, the resulting industrial chain is as follows: Figure 30 As shown.

[0238] It should be noted that the intelligent evaluation model is trained based on machine learning algorithms such as XGBoost and neural networks, and is updated in real time with multi-dimensional data such as production and logistics. It dynamically adjusts the evaluation results of the supply chain design quality and triggers risk warnings.

[0239] See Figure 31 As shown, in a specific embodiment, when completing the industrial chain design, it is necessary to construct a key element evaluation model to obtain the quantitative results of the industrial chain design quality, and form multi-perspective industrial chain design results and ranking optimization results; the key element evaluation model comprehensively evaluates the rationality of the distribution of upstream and downstream nodes in the industrial chain based on the correlation weight and the dependency relationship between nodes; the multi-perspective industrial chain design results include macro perspective, industry perspective, micro perspective, product perspective, domain perspective, problem perspective and concept perspective, forming a multi-perspective industrial chain design set and ranking results.

[0240] It should be noted that the construction of the key element assessment model refers to a quantitative evaluation framework with a multi-level structure built based on the evaluation needs of the industrial chain. It starts from four core dimensions: efficiency (such as production efficiency and circulation efficiency), quality (such as product qualification rate and data accuracy), risk (such as supply chain disruption risk and policy compliance risk), and collaboration (such as inter-enterprise cooperation and data sharing efficiency). It is divided into a hierarchy of total indicators, primary indicators, and secondary indicators to form a logically clear and calculable evaluation system.

[0241] Random walk strategy: refers to a graph-based node importance assessment method that simulates random particles moving between entity nodes in an industry chain knowledge graph. It quantifies node importance by statistically analyzing particle dwell probabilities and is suitable for mining hidden key elements in an industry chain. This includes setting initial node weights, transition probability matrices, and restart probability parameters, and obtaining the importance score of each key element through iterative calculation.

[0242] Initial node weights: In a random walk strategy, the initial importance score is pre-set for each entity node. It can be set according to the size of the enterprise, its industry position, or its industry experience, and no specific numerical limit is imposed here.

[0243] Transition probability matrix: refers to the probability matrix that represents the transition of a random particle from one entity node to another associated node. The values ​​of the matrix elements are determined by the weighted association weights.

[0244] Restart probability parameter: refers to the probability threshold of a particle returning to the initial node from the current node during a random walk. It is used to balance local association mining and global node coverage, and to prevent particles from getting stuck in local node clusters.

[0245] In one specific embodiment, the process of completing the supply chain design involves weighting the relationships between supply chain entities; constructing a key element importance assessment model using a random walk strategy; and evaluating the key elements based on the assessment model to obtain a supply chain design quality score.

[0246] The multi-perspective industry chain design assesses perspectives based on input demands, including macro, industry, micro, product, domain, problem, and conceptual perspectives—all reflecting the specific application needs of the industry chain. This caters to the diverse needs of users regarding the output effects of the industry chain.

[0247] In a specific embodiment, the construction of the key element importance assessment model can be implemented by referring to the following steps: 1. Network representation: Model the chain system as a directed weighted graph. .

[0248] : A set of nodes, representing entities in the chain (such as suppliers, processes, and knowledge units).

[0249] : A set of directed edges, representing the associations or dependencies between nodes (such as material flow, information flow, and dependencies).

[0250] Weighting function Indicates from node arrive The strength of association or dependency weight (such as transaction amount, dependency degree, traffic).

[0251] 2. Core Objective: Objective 1 (Importance Assessment): Calculate the overall importance score of each node, taking into account not only its own attributes, but also its structural position and dependencies in the network.

[0252] Objective 2 (Hierarchical Division): Based on importance scores and network topology, automatically divide nodes into three levels: upstream, midstream, and downstream.

[0253] In a specific embodiment, the detailed steps of the key element importance assessment model are as follows: Step 1: Graph construction and preprocessing: Adjacency matrix construction: Based on the edge set and weight Construct a weighted directed adjacency matrix ,in If the edge It exists; otherwise, it is 0.

[0254] Symmetry transformation (optional): For some applications that emphasize association rather than direction, a symmetric adjacency matrix can be constructed. Used for partial centrality calculations.

[0255] Graph connectivity check: Ensure the graph is weakly or strongly connected. For disconnected graphs, consider processing them as components or introducing virtual nodes.

[0256] Step 2: Calculation of multi-dimensional node importance indicators: We will calculate the centrality indicators in four dimensions to capture different aspects of the importance of nodes.

[0257] 1. Weighted Degree Centrality: Measures the strength and number of direct connections between nodes.

[0258] Weighted out-degree centrality: It reflects the node's direct ability to "influence others".

[0259] Weighted in-degree centrality: It reflects the degree to which a node is "dependent on by others".

[0260] Application: High in-degree may be a key convergence point (downstream), and high out-degree may be a key distribution point (upstream).

[0261] 2. Weighted betweenness centrality: measures a node’s ability to act as a “bridge” to control the flow of resources.

[0262]

[0263] It is a node arrive The total number of shortest paths, It was through The quantity.

[0264] Key improvements Traditional betweenness calculations count the number of paths. We modify it to calculate the sum or minimum bottleneck weight of path weights, so that paths with larger weights contribute more to the betweenness. This can identify nodes with strong control over high-value / high-dependency flows.

[0265] 3. Weighted proximity centrality: measures how easy it is for a node to reach other nodes in the network (the reciprocal of the average distance).

[0266]

[0267] It is a node arrive The weighted shortest path distance is calculated. Here, distance is defined as a function of the weights of the edges on the path (such as summation or the reciprocal of the maximum value) to ensure that the greater the weight, the closer the "distance". Proximity to nodes with high centrality ensures the highest average efficiency in transmitting information or resources to the entire network.

[0268] 4. Eigenvector centrality: measures the degree to which a node is connected to important nodes, reflecting the long-term influence of the network.

[0269] Solve the equation Main eigenvectors .

[0270] The choice of this approach means we focus more on "who depends on it" (joining the chain). The importance of a node depends on the sum of the importance of all nodes that point to it. This identifies nodes at the core of the dependency relationship (usually key nodes in the middle and lower reaches).

[0271] 5. PageRank variant (PR considering weights): An improvement on feature vector centrality, adding a random jump factor for greater stability.

[0272]

[0273] This formula incorporates edge weights into the transition probability, and the votes (influence) cast by important nodes are allocated according to the proportion of their outgoing edge weights.

[0274] Step 3: Integration of Overall Importance Scores: Purpose: To combine the above multi-dimensional indicators into a single overall importance score. .

[0275] Methods: We used weighted geometric mean or TOPSIS multi-criteria decision-making methods.

[0276] Weighted geometric mean (recommended): .in It is a node In the Normalized values ​​for each indicator It is the weight of this indicator. .

[0277] Advantages: Geometric mean is less sensitive to extreme values ​​and requires nodes to be relatively balanced in importance across all dimensions.

[0278] Weight setting Betweenness centrality can be determined through expert scoring, the Analytic Hierarchy Process (AHP), or entropy weighting based on the correlation between indicators. For example, if "control flow" is considered more important than "direct connectivity," then betweenness centrality can be given a higher weight.

[0279] Step 4: Hierarchical division based on topology sorting and importance score: Hierarchical division needs to combine network flow and node importance.

[0280] 1. Calculate the basic coordinates of the topological hierarchy: Perform topological sorting on the directed acyclic graph (DAG). For graphs with cycles, first shrink the strongly connected components, treating each SCC as a super node, and then perform topological sorting on the DAG.

[0281] Assign a topology sequence number to each node (or SCC). (From 1 to L, where L is the longest chain length). This reflects the basic preceding and following positions of the nodes in the process.

[0282] 2. Construct a hierarchical two-dimensional decision diagram: by (Topological order) is the horizontal axis, with Plot a scatter plot of all nodes with the overall importance score as the vertical axis.

[0283] 3. Clustering to divide upstream, midstream and downstream: Method: Apply clustering algorithms (such as K-means, K=3) on a two-dimensional decision graph.

[0284] Explanation of cluster: Cluster A (high) ,high ): A key upstream node. Located at the starting point of the process and of high importance, it is the "leader" or "source of innovation" of the chain.

[0285] Cluster B (middle) ,high ): Key midstream node. Located in the middle of the process, but with the highest importance, it is often the core processing, assembly or hub node, and is the "waist and abdomen" and "bottleneck" of the chain.

[0286] Cluster C (low) Medium and high ): A key downstream node. Located at the end of the process, directly facing the end user, it is of high importance and is the "exit" and "value realization point" of the chain.

[0287] Nodes with low importance scores are non-critical nodes in the corresponding level.

[0288] Alternative method (rule-based method): If the network structure is clear, it can be directly divided according to the topological order quantile (e.g., the first 30% is upstream, the last 30% is downstream, and the middle 40% is midstream), and then sorted by importance score within each layer to identify key nodes.

[0289] Step 5: Evaluation and Output: Output Results: 1. Multidimensional centrality index value and overall importance score for each node. .

[0290] 2. Hierarchical labels for each node: upstream - critical, upstream - non-critical, midstream - critical, midstream - non-critical, downstream - critical, and downstream - non-critical.

[0291] 3. A visualization of the overall hierarchical distribution of the network (two-dimensional decision diagram). The output of the industry chain results is as follows: Figure 32 As shown.

[0292] Using the methods described above, we can explore the generation results of the industrial chain from different perspectives, integrate multiple importance dimensions, and score and rank the industrial chain results from multiple perspectives: By using directed edges and in-degree related metrics (eigenvector centrality and PageRank), we can deeply capture the "supply chain dependence" attribute. For example... Figure 33 The optional chain node weight allocation shown is based on (expert scoring - AHP method) to evaluate the tightness and separation of the hierarchical division.

[0293] In a specific embodiment, after completing the hierarchical division of the industry chain nodes, the following optional methods can be used to quantitatively evaluate the quality of the division and improve the credibility of the results: 1. Quantitative calculation: Quantify the quality of the hierarchical division using methods such as weighted geometric mean and multi-criteria decision method (TOPSIS), integrate multi-dimensional evaluation indicators, and improve the objectivity of the division results. An example of node hierarchy statistical summary is shown below. Figure 34 As shown.

[0294] 2. Consistency verification: Compare with the list of key nodes annotated by domain experts, and calculate the accuracy = number of correctly identified key nodes ÷ total number of key nodes annotated by experts, recall = number of correctly identified key nodes ÷ total number of key nodes identified by the model, and Kappa coefficient (to measure the degree of consistency between human and model results).

[0295] 3. Stability Testing: Randomly add or remove 10%-15% of edges or adjust the weights by ±5%-10% for the industry chain results. Observe the stability of the ranking and hierarchical division of important nodes. The quality quantification evaluation indicators for the industry chain node hierarchical division are as follows: Figure 35 As shown.

[0296] 4. Visual Support: A two-dimensional decision graph (with topological order as the horizontal axis and comprehensive importance score as the vertical axis) visually presents the relationship between the "structural position" and "importance" of nodes, facilitating the rapid identification of key-level nodes and supporting decision analysis.

[0297] It should be noted that accuracy refers to the ratio of the number of correctly identified key nodes by the model to the total number of key nodes labeled by domain experts. It is used to measure the accuracy of the model in identifying key nodes, and the value ranges from 0 to 1. The higher the value, the better the accuracy.

[0298] Recall: The ratio of the number of correctly identified key nodes by the model to the total number of key nodes identified by the model itself. It is used to measure the coverage of key nodes labeled by domain experts. The value ranges from 0 to 1, and the higher the value, the more comprehensive the coverage.

[0299] Kappa coefficient: A statistical index used to measure the consistency between the model output of the industry chain node hierarchy division results and the results of manual annotation by domain experts. It eliminates the influence of random consistency. Among them, 0.8~1 represents high consistency, 0.6~0.8 represents moderate consistency, 0~0.6 represents low consistency, and negative values ​​represent inconsistency.

[0300] See Figure 36 As shown, an industry chain design device includes: a data acquisition module: used to collect industry-related data from compliant multi-source channels and preprocess it to form a multi-source heterogeneous dataset.

[0301] Potential chain node screening module: used to build an industry knowledge framework, extract entities and relationships, calculate enterprise industry matching scores, and obtain a set of potential chain nodes through threshold screening and cross-validation.

[0302] The completion and optimization module is used to construct ideal value strategy functions or set multi-task optimization target value strategies. Through learning models or multi-objective optimization decisions, it filters potential chain node sets and is also used to construct an initial industrial chain knowledge graph, complete missing associations, and optimize the graph structure through clustering.

[0303] The full lifecycle analysis module is used to divide the lifecycle into various stages based on the goals or needs of the industry chain, ensuring the integrity of the industry chain's formation.

[0304] The supply chain design quality assessment module is used to generate multi-perspective supply chain design results and ranking optimization results, providing decision support for the generated supply chain results.

[0305] The database is used to store multi-source heterogeneous datasets, core keyword and weight data, entity and relation data, candidate potential chain node data, graph triple data and multi-dimensional evaluation data. It is also used to store matching thresholds and clustering parameter thresholds.

[0306] In a third aspect, the present invention provides a computer storage medium for supply chain design, wherein a computer program is burned into the computer storage medium for supply chain design, and the computer program implements the supply chain design method described above when it runs in the memory of a server.

[0307] The models and algorithms used in this invention are all existing technologies, such as: BERT-based NER model, relation extraction model and TF-IDF algorithm, clustering algorithm, reinforcement learning model, cosine similarity algorithm and intelligent evaluation model and random walk strategy. The specific calculation process can be found on the Internet and will not be described in detail here.

[0308] The examples described in this invention are not limited to the specific embodiments listed above. The examples are merely illustrative to facilitate understanding of the invention and do not constitute a limitation on the scope of protection of this invention. Any modifications, equivalent substitutions, etc., made within the spirit and principles of this invention should be included within the scope of protection.

[0309] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.

Claims

1. A supply chain design method, characterized in that, include: Industry nodes are screened from the constructed multi-source heterogeneous dataset, potential chain nodes are identified based on the ideal value strategy, the node graph is completed and optimized, and the industrial chain design is completed based on the full life cycle analysis.

2. The supply chain design method according to claim 1, characterized in that, The data sources of the multi-source heterogeneous dataset include open-source or closed-source industry, enterprise, market, sector, product data and product functional semantic data; the multi-source heterogeneous data must contain the mapping relationship between industry, product and functional semantic data; the construction process of the dataset is to preprocess the original data, perform field mapping and format alignment according to a unified knowledge ontology, and complete data fusion and standardization.

3. The supply chain design method according to claim 1, characterized in that, The selection of industry nodes is based on a set industry knowledge framework to complete node selection and optimization. The industry knowledge framework revolves around the goals or needs of the industrial chain and is a combination of one or more knowledge systems from the perspective of industry, sector or field characteristics, including the category, product, technology and parameters. The selection is based on the industry knowledge framework selected by the user, combined with historical industry data statistics and weight analysis to determine the preset knowledge set. The optimization includes manual review and correction or machine learning model correction based on feedback data, and includes the construction of potential industrial chain levels and the generation of knowledge combinations of industry direction nodes.

4. The supply chain design method according to claim 1, characterized in that, The ideal value strategy is used to train the learning model and identify potential chain nodes. It includes at least one or a combination of learning reward function, energy index function, and value function. The selection strategy for the node knowledge combination is optimized through multiple iterations of training. The learning model includes various models or combinations of machine learning, deep learning, or reinforcement learning.

5. A supply chain design method according to claim 1 or 4, characterized in that, The ideal value strategy can also be a multi-objective value function, which performs multi-objective optimization and screening of potential industry chain node combinations. The multi-objective value function includes multi-dimensional task objective modeling of chain node value, relevance, coverage, relationship accuracy, completeness, foresight and timeliness, and defines quantitative weights for the corresponding indicators of each objective to construct a weighted value function. The multi-objective optimization includes at least optimizing the coverage and ideal value of node knowledge combinations, and the indicator proportions can be dynamically adjusted.

6. The supply chain design method according to claim 1, characterized in that, The completion and optimization of the node map is based on the upstream and downstream relationship analysis to confirm the chain nodes. It can introduce initial industrial chain node knowledge or node knowledge set and integrate it with existing node knowledge. The upstream and downstream relationship analysis is to label the relationship type and train the model to predict potential industrial chain node entities and entity relationships. The completion of chain node confirmation is to obtain the industrial chain node set at different levels from a certain perspective.

7. The supply chain design method according to claim 1, characterized in that, The aforementioned supply chain design based on full life cycle analysis involves dividing the life cycle into various stages around the goals or needs of the supply chain, ensuring the integrity of the supply chain's formation.

8. A supply chain design method according to claim 1 or 7, characterized in that, When completing the industrial chain design, it is necessary to construct a key element evaluation model to obtain quantitative results of the industrial chain design quality, and form multi-perspective industrial chain design results and ranking optimization results. The key element evaluation model comprehensively evaluates the rationality of the distribution of upstream and downstream nodes in the industrial chain based on the weight of the correlation relationship and the dependency relationship between nodes. The multi-perspective industrial chain design results include macro perspective, industry perspective, micro perspective, product perspective, domain perspective, problem perspective and concept perspective, forming a multi-perspective industrial chain design set and ranking results.

9. A supply chain design apparatus that utilizes the supply chain design method according to any one of claims 1-8, characterized in that, include: Data acquisition module: used to collect industry-related data from compliant multi-source channels and preprocess it to form a multi-source heterogeneous dataset; Potential chain node screening module: used to build an industry knowledge framework, extract entities and relationships and calculate enterprise industry matching scores, and obtain a set of potential chain nodes through threshold screening and cross-validation; The completion and optimization module is used to construct ideal value strategy functions or set multi-task optimization target value strategies. Through learning models or multi-objective optimization decisions, it filters potential chain node sets and is also used to construct an initial industrial chain knowledge graph, complete missing associations, and optimize the graph structure through clustering. Full lifecycle analysis module: used to divide the lifecycle into various stages based on the goals or needs of the industrial chain, ensuring the integrity of the industrial chain's formation; The supply chain design quality assessment module is used to generate multi-perspective supply chain design results and ranking optimization results, providing decision support for the generated supply chain results.

10. A computer storage medium with an integrated supply chain design, characterized in that: The computer storage medium for the supply chain design is burned with a computer program, which, when run in the server's memory, implements the supply chain design method as described in any one of claims 1-8.