Method, device and equipment for enterprise industry chain association by multi-stage semantic matching

CN122796162APending Publication Date: 2026-09-22SHANGHAI BRANCH OF CHINA URBAN PLANNING & DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610819527.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]本申请实施例提供了一种多阶段语义匹配的企业产业链关联方法、装置及设备,以至少解决相关技术中产业链与企业匹配准确性不足、效率低的技术问题

Benefits of technology

本申请提供了一种多阶段语义匹配的企业产业链关联方法,利用大语言模型自动生成包含多级节点的产业链层级本体结构;并采用多级筛选匹配,结合基于全文检索的快速召回、基于向量相似度的多级精筛、以及融合企业多维度信息的深度语义匹配,构建了从粗到精的企业-产业链关联流水线,实现了对海量企业数据的高效、精准、可扩展的产业链环节自动化挂链,克服了传统人工标注效率低、规则匹配覆盖面窄、单一特征匹配精度不足的缺陷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122796162A_ABST
    Figure CN122796162A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, and equipment for multi-stage semantic matching to associate enterprises within a supply chain. The method includes: obtaining seed words for the target supply chain; automatically generating a hierarchical ontology structure of the supply chain containing multi-level nodes using a large language model; recalling a set of candidate enterprises from an enterprise database based on keywords of each node in the supply chain hierarchical ontology structure using full-text search, keyword indexing, or vector pre-recall methods; calculating the similarity between the supply chain node text and the candidate enterprise text; filtering the candidate enterprise set based on the similarity to obtain initial filtering results; and calculating a deep matching score using a semantic matching model based on the multi-dimensional information text of the candidate enterprises in the initial filtering results and the supply chain node text to obtain the matching result between the enterprise and the supply chain link. This application provides a method for linking enterprises within a supply chain based on multi-stage semantic matching, achieving efficient, accurate, and scalable automatic association of enterprise supply chains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and artificial intelligence technology, and more specifically, to a method, apparatus, and equipment for multi-stage semantic matching of enterprise supply chain association. Background Technology

[0002] Supply chain analysis, as a key basis for corporate strategic planning and management decisions, hinges on accurately identifying the relationships between enterprises and specific links in the supply chain. However, existing technologies have significant limitations: traditional manual annotation methods heavily rely on expert experience, resulting in low efficiency and difficulty in scaling when processing massive amounts of enterprise data; keyword-based rule matching methods are limited by the coverage of pre-set thesaurus, making it difficult to cope with the diversity and dynamic changes in enterprise business information; existing single-vector matching technologies only utilize shallow text such as enterprise names and business scopes, resulting in insufficient matching accuracy due to a single information dimension; more importantly, existing methods generally lack consideration for the hierarchical structure of the supply chain, failing to effectively handle the different roles that the same enterprise may play at different supply chain levels, making it difficult for matching results to meet the needs of refined supply chain analysis. Summary of the Invention

[0003] This application provides a method, apparatus, and equipment for multi-stage semantic matching of enterprise supply chain associations, in order to at least solve the technical problems of insufficient accuracy and low efficiency in matching supply chains and enterprises in related technologies.

[0004] According to one aspect of the embodiments of this application, a multi-stage semantic matching method for enterprise supply chain association is provided, including: Obtain seed words for the target industry chain and automatically generate an industry chain hierarchical ontology structure containing multi-level nodes using a large language model; Based on the keywords of each node in the aforementioned industry chain hierarchical ontology structure, a set of candidate enterprises is retrieved from the enterprise database using full-text search, keyword indexing, or vector pre-recall methods. Calculate the similarity between the text of the industry chain node and the text of the candidate enterprise, and filter the candidate enterprise set based on the similarity to obtain the initial filtering results; Based on the multidimensional information text of candidate companies in the initial screening results, as well as the text of industry chain nodes, a semantic matching model is used to calculate the deep matching score, thereby obtaining the matching results between companies and industry chain links.

[0005] In one implementation, seed words for the target industry chain are obtained, and a hierarchical ontology structure containing multi-level nodes is automatically generated using a large language model, including: The seed words of the target industry chain are obtained. Based on the seed words and prompts, the big language model breaks down the industry chain into a hierarchical framework of upstream, midstream and downstream. The hierarchical framework is recursively extended to construct a tree-like hierarchical ontology structure, where each node contains a node name, description information, keywords, child node information, and industry chain position identifier. The location identifier is used to characterize the upstream, midstream, and downstream location attributes of the node in the industry chain.

[0006] In one implementation, based on the keywords of each node in the industry chain hierarchy ontology structure, a set of candidate enterprises is recalled from the enterprise database using full-text search, keyword indexing, or vector pre-recall methods, including: Extract keywords from the descriptive information of each node in the hierarchical ontology structure of the industry chain; Using full-text search, keyword indexing, or vector pre-recall based on a pre-computed enterprise vector table, a preset number of candidate enterprises are recalled from the enterprise database.

[0007] In one implementation, the similarity between the text of industry chain nodes and the text of candidate enterprises is calculated respectively, and the candidate enterprise set is filtered based on the similarity to obtain an initial filtering result, including: The text of the industry chain node and the text of the candidate enterprise are respectively processed by feature concatenation and vectorization to obtain the industry chain node vector and the candidate enterprise vector. Calculate the cosine similarity between the industry chain node vector and the candidate enterprise vector to obtain the initial matching score of each candidate enterprise for each node; Based on the initial matching scores of each candidate enterprise for each node, the matching score distribution of each node is obtained. Based on the matching score distribution and the results of manual quality inspection, the differentiated first threshold and second threshold of each node are determined. When the initial matching score is greater than or equal to the first threshold, it is determined that the candidate enterprise has successfully matched with the node, and a matching result is obtained; When the initial matching score is greater than or equal to the second threshold and less than the first threshold, the candidate company is determined to be a candidate company that has passed the preliminary screening, and the initial screening result is formed.

[0008] In one implementation, based on the multidimensional information text of candidate enterprises in the initial screening results and the text of industry chain nodes, a semantic matching model is used to calculate a deep matching score to obtain the matching results between enterprises and industry chain links, including: Obtain multidimensional information of candidate companies from the initial screening results. The multidimensional information includes at least two of the following: company name, business scope, patent list, product list, software copyright list, winning bid record, and project record. Input the multidimensional information text of the enterprise and the text of the industrial chain nodes into the semantic matching model to obtain the semantic matching score; A dynamic judgment threshold is determined based on the location identifier of the target industry chain node. If the target industry chain node is a downstream application link, a first matching threshold is used; if the target industry chain node is an upstream or midstream application link, a second matching threshold is used. The first matching threshold is greater than the second matching threshold. When the semantic matching score is greater than or equal to the dynamic judgment threshold, the corresponding candidate enterprise is determined to be successfully matched with the target industrial chain node, and the matching result is output.

[0009] In one implementation, after obtaining the matching results between enterprises and links in the industrial chain, the process further includes: Based on the matching results, a set of enterprise identifiers matching each node of the target industry chain is obtained; Based on the enterprise relationship graph, query the list of associated enterprises for each enterprise in the enterprise identifier set; For each enterprise in the enterprise identifier set, determine whether its associated enterprise list intersects with the enterprise identifier set; If there is an overlap, the matching score of the corresponding enterprise will be adjusted according to the preset gain rules.

[0010] In one implementation, after obtaining the matching results between enterprises and links in the industrial chain, the process further includes: Output the matching result, which includes one or more of the following: enterprise ID, chain link, matching score, node path, threshold version, evidence summary, and manual quality inspection status. Based on the matching results, manual quality inspection is performed, and manual quality inspection annotation information is received. The quality inspection annotation information includes at least one of strong correlation, weak correlation, and mislabeling. Based on the quality inspection labeling information, the judgment thresholds used in the corresponding industrial chain links are iteratively adjusted and optimized.

[0011] In one implementation, it further includes: When matching the target industry chain for the first time, the text of each node in the industry chain ontology structure is vectorized to generate the corresponding node vector. The node vector is cached using the seed words of the target industry chain as keys; When performing subsequent matching on the same target industry chain, the node vector stored in the cache is read based on the seed word.

[0012] According to another aspect of the embodiments of this application, a multi-stage semantic matching enterprise supply chain association device is provided, comprising: The industry chain ontology construction module is used to obtain seed words of the target industry chain and automatically generate an industry chain hierarchical ontology structure containing multi-level nodes using a large language model. The first matching module is used to recall a set of candidate enterprises from the enterprise database based on the keywords of each node in the hierarchical ontology structure of the industry chain, using full-text search, keyword indexing or vector pre-recall methods. The second matching module is used to calculate the similarity between the text of the industry chain node and the text of the candidate enterprise respectively, and to filter the candidate enterprise set based on the similarity to obtain the initial filtering result; The third matching module is used to calculate the deep matching score based on the multi-dimensional information text of candidate enterprises in the initial screening results and the text of industrial chain nodes, using a semantic matching model to obtain the matching result between the enterprise and the industrial chain link.

[0013] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described multi-stage semantic matching enterprise supply chain association method through the computer program.

[0014] The technical solutions provided in this application embodiment may include the following beneficial effects: This application provides a multi-stage semantic matching method for enterprise supply chain association. It utilizes a large language model to automatically generate a hierarchical ontology structure of the supply chain containing multiple levels of nodes. It also employs multi-level filtering and matching, combining rapid recall based on full-text retrieval, multi-level fine filtering based on vector similarity, and deep semantic matching that integrates multi-dimensional enterprise information. This constructs a coarse-to-fine enterprise-supply chain association pipeline, enabling efficient, accurate, and scalable automated linking of supply chain links to massive amounts of enterprise data. This overcomes the shortcomings of traditional manual annotation, such as low efficiency, narrow coverage of rule matching, and insufficient accuracy of single feature matching. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a multi-stage semantic matching method for enterprise supply chain association according to an embodiment of this application; Figure 2 This is an architecture diagram of a multi-stage semantic matching enterprise supply chain association method according to an embodiment of this application; Figure 3 This is a schematic diagram of a data preparation process according to an embodiment of this application; Figure 4 This is a schematic diagram of a multi-stage semantic matching enterprise supply chain association device according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0018] The following is a detailed description of the enterprise supply chain association method for multi-stage semantic matching according to embodiments of this application, with reference to the accompanying drawings. For example... Figure 1 As shown, the method mainly includes the following steps: S101 obtains seed words for the target industry chain and uses a large language model to automatically generate an industry chain hierarchical ontology structure containing multi-level nodes.

[0019] In one implementation, seed words for the target industry chain are obtained, and a hierarchical ontology structure containing multi-level nodes is automatically generated using a large language model. This includes: obtaining seed words for the target industry chain; the large language model decomposes the industry chain into a hierarchical framework of upstream, midstream, and downstream based on the seed words and prompts; recursively expanding the hierarchical framework to construct a tree-shaped hierarchical ontology structure, where each node contains a node name, description information, keywords, child node information, and an industry chain position identifier; the position identifier is used to characterize the upstream, midstream, and downstream position attributes of the node in the industry chain.

[0020] Specifically, a three-step thinking chain method based on a large language model automatically constructs the hierarchical ontology structure of the industry chain. First, in the framework design phase, the model breaks down the industry into three major segments—upstream, midstream, and downstream—based on industry seed words, and identifies key first-level sub-segments for each segment, forming a preliminary hierarchical framework. Next, in the depth expansion phase, the model recursively expands the semantics of each first-level sub-segment, delving into the granularity of levels 3-4, detailing the specific core components, key materials, or key technologies required for each level. Finally, in the structure assembly phase, the expanded nodes at each level are assembled into a complete tree structure, where each node contains a name, description, keyword list, and child node information, and its industry chain position attribute is recorded in the node's metadata.

[0021] The automatic generation mechanism of the industrial chain hierarchy is entirely driven by a large language model, without the need for preset manual rules. Based on the semantic understanding of the technology transmission logic of the industrial chain, the large language model follows the industrial transmission path of "raw materials → components → complete machines → application services", autonomously determines the hierarchical position of each link in the industrial chain, and records the position information in the node meta-information, thereby realizing the automated and intelligent generation of the industrial chain structure.

[0022] In one exemplary embodiment, the user inputs seed terms for the industry chain, such as new energy vehicles and low-altitude economy. This embodiment does not impose specific limitations. Next, a large language model is invoked to automatically generate the hierarchical structure of the industry chain. The input prompt reads, "You are a senior industry planning expert. Your goal is to design a top-level framework for an industry chain, breaking the industry down into three main segments: upstream, midstream, and downstream, and their respective key first-level sub-segments." Then, in the deep expansion step, the model switches to the role of a supply chain engineer, recursively expanding each sub-segment to a granularity of 3-4 levels, detailing specific core components, key materials, and key technologies. Finally, in the structure assembly step, the framework and expanded information are integrated into a tree-shaped ontology structure containing complete node metadata (name, description, keywords, sub-nodes, and industry chain position identifiers), achieving automated construction from seed terms to a computable industry chain ontology.

[0023] S102 uses keywords from each node in the industry chain hierarchy ontology structure to recall a set of candidate companies from the company database using full-text search, keyword indexing, or vector pre-recall methods.

[0024] In one implementation, keywords are extracted from the descriptive information of each node in the industry chain hierarchy ontology structure; a preset number of candidate enterprises are recalled from the enterprise database using full-text search, keyword indexing, or vector pre-recall based on a pre-computed enterprise vector table.

[0025] Optionally, a GIN full-text search method can be used for initial recall. Keywords are extracted from the descriptive information of each node in the industry chain hierarchy ontology structure, and the extracted keywords are concatenated into a query string according to the full-text search syntax rules. In the database storing enterprise information, a GIN index is built for the enterprise business scope field. Based on the query string, a full-text search operation is performed using the GIN index to filter out a preset number of candidate enterprises.

[0026] Specifically, keywords are first extracted from the description information of each node in the industry chain hierarchy ontology structure. The keywords are then preprocessed, including removing special characters, and concatenated using logical OR operators to generate query strings that conform to the PostgreSQL full-text search syntax. Subsequently, based on the enterprise business scope field with the established GIN index, a full-text search query is executed to match enterprise records containing any keyword in the enterprise business scope. Finally, a preset number (e.g., 100) of candidate enterprises are recalled from the database to form the initial candidate set for the next stage of screening.

[0027] Optionally, keyword indexing or vector pre-recall based on a pre-calculated enterprise vector table can be used to recall a predetermined number of candidate enterprises from the enterprise database. For example, during the system initialization phase, high-dimensional semantic vectors of all enterprises are pre-calculated and stored. During the recall phase, the similarity between the target industry chain node vector and the enterprise vectors in the enterprise vector table is calculated, and the enterprises are sorted from high to low according to the similarity scores to directly recall a predetermined number of candidate enterprises, forming an initial candidate set for subsequent matching processes.

[0028] S103 calculates the similarity between the text of the industry chain node and the text of the candidate enterprise, and filters the candidate enterprise set based on the similarity to obtain the initial screening results.

[0029] First, feature concatenation and vectorization are performed on the text of the industry chain nodes and the text of the candidate enterprises respectively to obtain the vectors of the industry chain nodes and the vectors of the candidate enterprises. The cosine similarity between the vectors of the industry chain nodes and the vectors of the candidate enterprises is calculated to obtain the initial matching score of each candidate enterprise for each node.

[0030] Specifically, the candidate company's "name + business scope" text and the industry chain link's "name + description + keywords" text are first input into the semantic encoding model to generate high-dimensional vector representations. Then, the similarity mapping matrix between the industry chain node vector matrix and the candidate company vector matrix is ​​calculated based on the cosine similarity formula, thereby quantifying the semantic association strength between each company and each industry chain link, providing a quantitative basis for subsequent fine screening.

[0031] Furthermore, based on the initial matching scores of each candidate enterprise for each node, the matching score distribution of each node is obtained. Based on the matching score distribution and the results of manual quality inspection, the differentiated first threshold and second threshold of each node are determined.

[0032] This implementation method determines an independent differentiated threshold for each node in the industry chain. Specifically, after calculating the initial matching score of each candidate enterprise for a certain node, the system will statistically analyze the distribution of all matching scores corresponding to that node, and combine the manual quality inspection feedback of the matching results of that node to independently set a first threshold for high-confidence matching and a second threshold for entering in-depth verification for each node, thereby achieving a refined adaptation to the differentiated matching accuracy requirements of different links in the industry chain.

[0033] Furthermore, if the initial matching score is greater than or equal to the first threshold, the candidate enterprise is determined to have successfully matched with the node, and a matching result is obtained; if the initial matching score is greater than or equal to the second threshold but less than the first threshold, the candidate enterprise is determined to be a candidate enterprise that has passed the preliminary screening, and an initial screening result is formed.

[0034] In one implementation, a quantile-based dynamic threshold strategy is used to classify the preliminary screening results. First, the P95 percentile of the initial matching score set of all candidate companies is calculated as a high-confidence pass threshold. Candidate companies with scores higher than this threshold are directly judged as successful matches and skip the subsequent deep verification process. At the same time, the P80 percentile is calculated as an advancement threshold. Candidate companies with scores in the P80 to P95 range enter the subsequent deep semantic matching stage for further verification. This data-driven threshold setting method can adapt to the data distribution characteristics of different industry chains and significantly optimize the overall processing efficiency while ensuring matching quality.

[0035] S104 uses the multi-dimensional information text of candidate companies in the initial screening results, as well as the text of industry chain nodes, to calculate the deep matching score using a semantic matching model, and obtains the matching results between companies and industry chain links.

[0036] In one implementation, based on the multidimensional information text of candidate companies in the initial screening results and the text of industry chain nodes, a semantic matching model is used to calculate a deep matching score to obtain the matching result between the company and the industry chain link. This includes: obtaining multidimensional information of candidate companies in the initial screening results, including at least two of the following: company name, business scope, patent list, product list, software copyright list, bidding record, and project record; and inputting the multidimensional information text of the company and the text of industry chain nodes into the semantic matching model to obtain a semantic matching score.

[0037] Specifically, the process first integrates the names and business scopes of the shortlisted candidate companies, and then selectively merges fragments of multi-source business data such as their patents, products, software copyrights, bidding records, and project information to form a company description text rich in business context. This text is then combined with the "name + description" of the target industry chain link to form a semantic pair, which is then input into a deep semantic matching model to calculate a refined semantic matching score.

[0038] The semantic matching model used for deep semantic matching can be implemented in various forms, including but not limited to the Cross-Encoder dual-tower structure model, reordering model, cross-encoding model or a combination thereof. This application does not limit the specific architecture, training method and parameter configuration of the model. Any model that can perform semantic interaction between the text of the industry chain node and the multi-dimensional text of the enterprise and output the matching score is within the protection scope of this application.

[0039] Furthermore, a dynamic judgment threshold is determined based on the location identifier of the target industry chain node. If the target industry chain node is a downstream application link, a first matching threshold is used; if the target industry chain node is an upstream or midstream application link, a second matching threshold is used. The first matching threshold is greater than the second matching threshold. When the semantic matching score is greater than or equal to the dynamic judgment threshold, the corresponding candidate enterprise is determined to be successfully matched with the target industry chain node, and the matching result is output.

[0040] For example, after calculating the semantic matching score between the candidate enterprise and the target industrial chain node in the deep semantic matching stage, the system will automatically identify the industrial chain position attribute of the target node and dynamically apply differentiated thresholds according to the attributes of the industrial chain links. If the node belongs to the downstream link with high technology integration and clear application scenarios, a higher judgment threshold (such as 0.8) will be used; if the node belongs to the intermediate link of technology transmission or the upstream and midstream link of basic material supply, a relatively lower judgment threshold (such as 0.70) will be used. When the semantic matching score is greater than or equal to the dynamic threshold of the corresponding link, it is determined that the candidate enterprise has successfully matched the target industrial chain node.

[0041] In one implementation, after obtaining the matching results between enterprises and links in the industrial chain, the method further includes: obtaining a set of enterprise identifiers that match each node in the target industrial chain based on the matching results; querying the list of associated enterprises for each enterprise in the enterprise identifier set based on the enterprise relationship graph; determining whether the list of associated enterprises of each enterprise in the enterprise identifier set intersects with the enterprise identifier set; if there is an intersection, adjusting the matching score of the corresponding enterprise according to a preset gain rule.

[0042] Specifically, the semantic matching results are enhanced by using an enterprise relationship graph. First, the set of enterprise identifiers matched through multi-dimensional semantic stages is obtained, and the list of related enterprises of each enterprise is queried based on the enterprise relationship graph. Then, for each enterprise, it is determined whether there is an intersection between its list of related enterprises and the set of currently matched enterprises in the same industry chain. If there is an intersection, it indicates that the enterprise has business relationships with other enterprises in the same industry ecosystem in the industry network, thereby enhancing its matching credibility. This is achieved by adding a preset gain value (such as 0.1 points) to the matching score of the enterprise. Finally, an enhanced matching result with network topology correction is generated.

[0043] In one implementation, after obtaining the matching results between the enterprise and the link in the industrial chain, the method further includes: outputting the matching results, which include one or more of the following: enterprise ID, linked link, matching score, node path, threshold version, evidence summary, and manual quality inspection status.

[0044] When outputting the final matching result, this implementation method generates a structured data record that includes the enterprise ID, the identifier of the linked link, the semantic matching score, the node hierarchical path, the judgment threshold version, the key text evidence summary supporting the matching, and the current manual review status, providing a complete and traceable decision-making basis for subsequent quality inspection, analysis and application.

[0045] Furthermore, manual quality inspection is conducted based on the matching results, and manual quality inspection labeling information is received. The quality inspection labeling information includes at least one of strong correlation, weak correlation, and mislabeling. Based on the quality inspection labeling information, the judgment thresholds used in the corresponding industrial chain links are iteratively adjusted and optimized.

[0046] Based on the output matching results, a manual quality inspection process is carried out, receiving categorized feedback information marked by quality inspectors. The feedback information includes at least one category among "strong correlation" indicating matching accuracy, "weak correlation" indicating marginal matching, and "false positive" indicating no relevance. Subsequently, based on the collected quality inspection marking information, the system performs statistical analysis on the judgment thresholds used in the corresponding industrial chain links, and automatically or semi-automatically iteratively calibrates and optimizes the thresholds according to indicators such as false positive rate and recall rate, so as to continuously improve the matching accuracy.

[0047] In one implementation, the method further includes: when matching the target industry chain for the first time, vectorizing the text of each node in the industry chain ontology structure to generate corresponding node vectors; caching the node vectors using the seed words of the target industry chain as keys; and reading the stored node vectors from the cache based on the seed words when matching the same target industry chain subsequently.

[0048] This implementation optimizes matching efficiency through an industry chain ontology vector caching mechanism. When matching a specific industry chain for the first time, the system uses the seed word of the industry chain as the key to call the semantic coding model to calculate the vector representation of each node text in its ontology structure, and stores the node information and corresponding vectors in the memory cache. For subsequent matching requests for the same industry chain, the system directly reads the calculated node vectors from the cache through the seed word key for reuse, avoiding repeated vectorization calculations for the same industry chain links, and achieving significant performance optimization of "one calculation, multiple uses".

[0049] Alternatively, matching efficiency can be improved through a full enterprise vector pre-computation system, such as... Figure 3 As shown, the process includes building a cache table, pre-computing vectors, and outputting data. First, a rich, multi-dimensional text is constructed for each enterprise, containing its name, business scope, patents, products, software copyrights, bidding records, and project information. Then, the continuous batch processing optimization technology of the vLLM framework is used to process the enterprise text obtained by cursor pagination concurrently with a fixed batch size. High-dimensional semantic vector representations are generated for all enterprises through asynchronous concurrent calls to the semantic encoding model, and the results are persistently stored to provide pre-computation support for efficient vector similarity calculation in the subsequent pipeline matching stage.

[0050] A high-performance vector pre-computation system was designed for enterprise data of hundreds of millions of records. By integrating heterogeneous data from six dimensions—enterprise name, business scope, patents, products, software copyrights, bidding records, and project information—a high-information-density enterprise semantic text was constructed. The system employs continuous batch processing technology of the vLLM framework, making full use of GPU computing power by sending concurrent requests in batches of 32 records. Combined with a cursor paging mechanism, the system avoids the performance degradation of traditional OFFSET paging and enables parallel processing by using 8 worker processes, significantly reducing the time overhead of full enterprise vector computation.

[0051] This application proposes a four-stage pipeline architecture (L0-L3). The L0 stage uses full-text search based on GIN index to quickly recall candidate enterprises. The L1 stage uses a dual-tower model to calculate cosine similarity for vector coarse screening. The L2 stage combines multi-dimensional business information of enterprises for deep semantic fine matching. The L3 stage introduces enterprise relationship graphs for topology correction, forming a progressive matching process from massive recall, rapid filtering, accurate matching to network enhancement, realizing efficient and accurate industrial chain linking of hundreds of millions of enterprise data.

[0052] To facilitate understanding of the methods in the embodiments of this application, the following description is provided in conjunction with the appendix. Figure 2To further explain, firstly, the system receives seed words input by the user and automatically generates an industry chain hierarchical ontology structure based on a large language model. The data preparation layer completes the basic data construction through an enterprise information cache table and a full-scale enterprise vector pre-computation module. Secondly, the core processing layer constructs a multi-stage matching pipeline from L0 to L3, sequentially using full-text search for candidate recall, vector cosine similarity for coarse screening, and multi-dimensional information enhancement to achieve semantic precision matching. It also combines enterprise relationship graphs for topological correction and introduces a percentile dynamic threshold mechanism to optimize the screening accuracy. Finally, the system outputs batch results including enterprise ID, chain links, matching scores, and credibility tags.

[0053] According to another aspect of the embodiments of this application, a multi-stage semantic matching enterprise supply chain association apparatus for implementing the above-described multi-stage semantic matching enterprise supply chain association method is also provided. For example... Figure 4 As shown, the device includes: The industry chain ontology construction module 401 is used to obtain the seed words of the target industry chain and automatically generate an industry chain hierarchical ontology structure containing multi-level nodes using a large language model. The first matching module 402 is used to recall a set of candidate enterprises from the enterprise database based on the keywords of each node in the industry chain hierarchy ontology structure, using full-text search, keyword indexing or vector pre-recall methods. The second matching module 403 is used to calculate the similarity between the text of the industry chain node and the text of the candidate enterprise respectively, and to filter the candidate enterprise set based on the similarity to obtain the initial filtering results; The third matching module 404 is used to calculate the deep matching score based on the multi-dimensional information text of candidate enterprises in the initial screening results and the text of industrial chain nodes, and to obtain the matching results between enterprises and industrial chain links.

[0054] It should be noted that the multi-stage semantic matching enterprise supply chain association device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the multi-stage semantic matching enterprise supply chain association method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the multi-stage semantic matching enterprise supply chain association device and the multi-stage semantic matching enterprise supply chain association method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0055] According to another aspect of the embodiments of this application, an electronic device corresponding to the multi-stage semantic matching enterprise supply chain association method provided in the foregoing embodiments is also provided, so as to execute the above-mentioned multi-stage semantic matching enterprise supply chain association method.

[0056] Please refer to Figure 5 This illustrates a schematic diagram of an electronic device provided by some embodiments of this application. For example... Figure 5 As shown, the electronic device includes: a processor 500, a memory 501, a bus 502, and a communication interface 503. The processor 500, the communication interface 503, and the memory 501 are connected via the bus 502. The memory 501 stores a computer program that can run on the processor 500. When the processor 500 runs the computer program, it executes the multi-stage semantic matching enterprise supply chain association method provided in any of the foregoing embodiments of this application.

[0057] The memory 501 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 503 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.

[0058] Bus 502 can be an ISA bus, PCI bus, or EISA bus, etc. Buses can be divided into address buses, data buses, control buses, etc. Memory 501 is used to store programs. After receiving execution instructions, processor 500 executes the program. The multi-stage semantic matching enterprise supply chain association method disclosed in any of the aforementioned embodiments of this application can be applied to processor 500, or implemented by processor 500.

[0059] The processor 500 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 500 or by instructions in software form. The processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 501. The processor 500 reads the information in memory 501 and, in conjunction with its hardware, completes the steps of the above method.

[0060] The electronic device provided in this application embodiment and the enterprise supply chain association method for multi-stage semantic matching provided in this application embodiment are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.

[0061] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0062] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.

Claims

1. A multi-stage semantic matching method for enterprise supply chain association, characterized in that, include: Obtain seed words for the target industry chain and automatically generate an industry chain hierarchical ontology structure containing multi-level nodes using a large language model; Based on the keywords of each node in the aforementioned industry chain hierarchical ontology structure, a set of candidate enterprises is retrieved from the enterprise database using full-text search, keyword indexing, or vector pre-recall methods. Calculate the similarity between the text of the industry chain node and the text of the candidate enterprise, and filter the candidate enterprise set based on the similarity to obtain the initial filtering results; Based on the multidimensional information text of candidate companies in the initial screening results, as well as the text of industry chain nodes, a semantic matching model is used to calculate the deep matching score, thereby obtaining the matching results between companies and industry chain links.

2. The method according to claim 1, characterized in that, Obtain seed words for the target industry chain, and automatically generate a hierarchical ontology structure of the industry chain containing multi-level nodes using a large language model, including: The seed words of the target industry chain are obtained. Based on the seed words and prompts, the big language model breaks down the industry chain into a hierarchical framework of upstream, midstream and downstream. The hierarchical framework is recursively extended to construct a tree-like hierarchical ontology structure, where each node contains a node name, description information, keywords, child node information, and industry chain position identifier. The location identifier is used to characterize the upstream, midstream, and downstream location attributes of the node in the industry chain.

3. The method according to claim 1, characterized in that, Based on the keywords of each node in the aforementioned industry chain hierarchical ontology structure, a set of candidate companies is retrieved from the company database using full-text search, keyword indexing, or vector pre-recall methods, including: Extract keywords from the descriptive information of each node in the hierarchical ontology structure of the industry chain; Using full-text search, keyword indexing, or vector pre-recall based on a pre-computed enterprise vector table, a preset number of candidate enterprises are recalled from the enterprise database.

4. The method according to claim 1, characterized in that, Calculate the similarity between the text of each node in the industry chain and the text of each candidate company. Based on the similarity, filter the candidate company set to obtain an initial filtering result, including: The text of the industry chain node and the text of the candidate enterprise are respectively processed by feature concatenation and vectorization to obtain the industry chain node vector and the candidate enterprise vector. Calculate the cosine similarity between the industry chain node vector and the candidate enterprise vector to obtain the initial matching score of each candidate enterprise for each node; Based on the initial matching scores of each candidate enterprise for each node, the matching score distribution of each node is obtained. Based on the matching score distribution and the results of manual quality inspection, the differentiated first threshold and second threshold of each node are determined. When the initial matching score is greater than or equal to the first threshold, it is determined that the candidate enterprise has successfully matched with the node, and a matching result is obtained; When the initial matching score is greater than or equal to the second threshold and less than the first threshold, the candidate company is determined to be a candidate company that has passed the preliminary screening, and the initial screening result is formed.

5. The method according to claim 1, characterized in that, Based on the multidimensional information text of candidate companies in the initial screening results, and the text of industry chain nodes, a semantic matching model is used to calculate the deep matching score, thereby obtaining the matching results between companies and industry chain links, including: Obtain multidimensional information of candidate companies from the initial screening results. The multidimensional information includes at least two of the following: company name, business scope, patent list, product list, software copyright list, winning bid record, and project record. Input the multidimensional information text of the enterprise and the text of the industrial chain nodes into the semantic matching model to obtain the semantic matching score; A dynamic judgment threshold is determined based on the location identifier of the target industry chain node. If the target industry chain node is a downstream application link, a first matching threshold is used; if the target industry chain node is an upstream or midstream application link, a second matching threshold is used. The first matching threshold is greater than the second matching threshold. When the semantic matching score is greater than or equal to the dynamic judgment threshold, the corresponding candidate enterprise is determined to be successfully matched with the target industrial chain node, and the matching result is output.

6. The method according to claim 1, characterized in that, After obtaining the matching results between enterprises and links in the industrial chain, the following also applies: Based on the matching results, a set of enterprise identifiers matching each node of the target industry chain is obtained; Based on the enterprise relationship graph, query the list of associated enterprises for each enterprise in the enterprise identifier set; For each enterprise in the enterprise identifier set, determine whether its associated enterprise list intersects with the enterprise identifier set; If there is an overlap, the matching score of the corresponding enterprise will be adjusted according to the preset gain rules.

7. The method according to claim 1, characterized in that, After obtaining the matching results between enterprises and links in the industrial chain, the following also applies: Output the matching result, which includes one or more of the following: enterprise ID, chain link, matching score, node path, threshold version, evidence summary, and manual quality inspection status. Based on the matching results, manual quality inspection is performed, and manual quality inspection annotation information is received. The quality inspection annotation information includes at least one of strong correlation, weak correlation, and mislabeling. Based on the quality inspection labeling information, the judgment thresholds used in the corresponding industrial chain links are iteratively adjusted and optimized.

8. The method according to claim 1, characterized in that, Also includes: When matching the target industry chain for the first time, the text of each node in the industry chain ontology structure is vectorized to generate the corresponding node vector. The node vector is cached using the seed words of the target industry chain as keys; When performing subsequent matching on the same target industry chain, the node vector stored in the cache is read based on the seed word.

9. A multi-stage semantic matching device for enterprise supply chain association, characterized in that, include: The industry chain ontology construction module is used to obtain seed words of the target industry chain and automatically generate an industry chain hierarchical ontology structure containing multi-level nodes using a large language model. The first matching module is used to recall a set of candidate enterprises from the enterprise database based on the keywords of each node in the hierarchical ontology structure of the industry chain, using full-text search, keyword indexing or vector pre-recall methods. The second matching module is used to calculate the similarity between the text of the industry chain node and the text of the candidate enterprise respectively, and to filter the candidate enterprise set based on the similarity to obtain the initial filtering result; The third matching module is used to calculate the deep matching score based on the multi-dimensional information text of candidate enterprises in the initial screening results and the text of industrial chain nodes, using a semantic matching model to obtain the matching result between the enterprise and the industrial chain link.

10. An electronic device, characterized in that, It includes a processor and a memory storing program instructions, the processor being configured to, when executing the program instructions, perform the enterprise supply chain association method of multi-stage semantic matching as described in any one of claims 1 to 8.