Industrial chain enterprise identification and analysis method

By building an industrial chain knowledge base and adopting a layered identification model, identifying industrial chain enterprises and optimizing the knowledge base, the problems of poor identification results and inefficiency in the existing technology have been solved, and efficient and accurate identification of industrial chain enterprises have been achieved.

CN120011823APending Publication Date: 2025-05-16JIANGSU UNITED CREDIT REFERENCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411871505.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When identifying enterprises in the industrial chain, it is difficult to ensure the recognition effect, and the algorithm efficiency is low when there are many enterprises.

Method used

Build an industrial chain knowledge base, collect descriptions of the main stages and related technology and relevance information, and form knowledge factors; collect and mark enterprise lists through multiple channels, generate quantitative indicators and textual descriptions of the correlation between specific enterprises and knowledge factors, and form a collection of enterprise characteristics. The industrial chain matching model and stage identification model are used to identify candidate enterprises and their own stages, and the knowledge base is optimized through verification data sets to improve the recognition effect.

Benefits of technology

Through the layered identification model and knowledge base optimization mechanism, the efficiency and accuracy of enterprise identification in the industrial chain are improved, missed judgments and misjudgments are reduced, and the recognition effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011823A_ABST
    Figure CN120011823A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial chain enterprise identification and analysis method. The method specifically comprises the steps of 1, constructing an industrial chain knowledge base and a verification data set; 2, collecting enterprise information, and carrying out correlation calculation on the enterprise information and the industry chain knowledge base to generate an enterprise feature set; step 3, identifying a candidate enterprise set on the target industrial chain by adopting an industrial chain matching model; 4, judging an industrial chain stage to which the enterprise belongs by adopting an industrial chain stage identification model; 5, the industrial chain stage recognition language model is adopted to discriminate the industrial chain enterprises of which the stages are not divided again; 6, comparing the verification data set with an identification result, performing attribution analysis of misjudgment and missed judgment on the industry chain knowledge, and optimizing the industry chain knowledge base; and 7, updating the identification result until the deviation is smaller than a preset value. A feedback mechanism is introduced into industrial chain recognition, and the recognition effect is improved through iterative optimization of a knowledge base; a hierarchical model mechanism is adopted in industrial chain enterprise identification, coarse screening is performed firstly, and then fine division is performed, so that the model has relatively high identification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial chain data technology, and specifically relates to an industrial chain enterprise identification and analysis method. Background Art

[0002] There are many application scenarios for identifying enterprises in the industrial chain, such as optimizing the configuration of the industrial chain, investment decision-making and strategic planning, supply chain management, etc. These scenarios cover different industries and economic activities, aiming to optimize resource allocation, promote industrial upgrading and improve market efficiency. There are more than 50 million registered enterprises in my country. How to accurately identify enterprises related to the industrial chain based on massive information is a difficult problem. One of the key points is the need to collect accurate knowledge in the field of industrial chain to assist in accurately identifying related enterprises.

[0003] The existing technology constructs evaluation models based on specific dimensions (indicators of preset information) to form scores, and forms comprehensive scores for each stage (upper, middle and lower reaches) through different weights. The stage is determined by taking the highest score or comparing the threshold. The existing technology currently has the following problems: 1) The knowledge in the industry chain covers a wide range of content and detailed items, and the preset model built based on independent training is difficult to guarantee the recognition effect; 2) Each enterprise must be judged by multiple models, and the algorithm efficiency is relatively low when there are a large number of enterprises. Summary of the invention

[0004] To achieve the above purpose, the technical solution of the present invention is as follows: a method for identifying and analyzing enterprises in an industrial chain, specifically comprising the following steps: Step 1: Build an industrial chain knowledge base, including collecting the main stages of the industrial chain and related technical descriptions and relevance information, and forming a number of knowledge factors to form the industrial chain knowledge base; build a verification data set, collect and annotate enterprise lists through multiple channels; Step 2: Collect enterprise information, perform correlation calculation with the industry chain knowledge base, generate quantitative indicators and textual descriptions of the correlation between specific enterprises and knowledge factors, and form an enterprise feature set; Step 3: Use the industrial chain matching model to identify the set of candidate enterprises in the target industrial chain. The input of the industrial chain matching model is the index in the feature set. The candidate enterprises in the target industrial chain are selected based on the correlation index output by the industrial chain matching model. Step 4: Use the industrial chain stage identification model to determine the industrial chain stage to which the enterprise belongs. The industrial chain stage identification model contains several sub-models, which correspond to different stages of the industrial chain. The input of the industrial chain stage identification model is the indicator in the feature set. Based on the predicted probability of the model output, it is determined whether the enterprise belongs to a specific industrial chain stage. Step 5: Use the industry chain stage recognition language model to re-identify the industry chain enterprises that have not been divided into stages. The input of the industry chain stage language recognition model is the textual description of the enterprise, and it is judged whether the enterprise belongs to a specific industry chain based on natural language processing technology; Step 6: Use the verification data set to compare the recognition results, conduct attribution analysis on the misjudgment and missed judgment of the industrial chain knowledge, identify the existing knowledge factors that cause misjudgment in the knowledge base, extract new knowledge factors that reduce missed judgment, and optimize the industrial chain knowledge base; Step 7: Update the recognition result until the deviation is less than the preset value.

[0005] As an improvement of the present invention, the industrial chain knowledge base in step 1 refers to the domain knowledge required to determine whether an enterprise belongs to the target industrial chain and the stage of the industrial chain it is in; the verification data set refers to a list of specific enterprises used to verify the effect of the industrial chain model. The industrial chain knowledge base is composed of a number of knowledge factors. The knowledge factors have a variety of organizational forms. They can be keywords describing raw materials, technologies, products and services related to the industrial chain, or they can be industrial access rules (such as administrative licenses and qualification certificates, etc.). Optionally, the knowledge factors can correspond to weights and industrial chain stage identifiers. The weight is used to indicate the contribution of the knowledge factor. The larger the weight, the higher the correlation with the industrial chain; the industrial chain stage identifier is used to indicate one or more industrial chain stages corresponding to the knowledge factor, such as upstream, midstream or downstream. The industrial chain knowledge base is composed of a number of knowledge factors, and the verification data set includes a positive data set and a negative data set. The positive data set contains a number of enterprises that have been clearly identified as belonging to the target industrial chain, and the negative data includes a number of enterprises that have been clearly identified as not belonging to the target industrial chain or not belonging to a specific industrial chain stage.

[0006] As an improvement of the present invention, in step 2, the enterprise feature set processing forms numerical indicators and / or unstructured content by comparing enterprise information and industrial chain knowledge base. Enterprise information refers to the content that can describe the business scope, development field, main products and services of the enterprise; feature set refers to the content that reflects the correlation between the enterprise and the industrial chain formed after processing based on enterprise information.

[0007] As an improvement of the present invention, the input of the industrial chain matching model in step three is the indicator in the feature set, and the output value of the model is a score or probability. The higher the output value, the higher the correlation between the enterprise and the industry. By setting a threshold, candidate enterprises in the target industrial chain are screened out from all enterprises to form a set of candidate enterprises.

[0008] As an improvement of the present invention, the industrial chain stage identification model in step 4 can be composed of several sub-models, corresponding to different stages of the industrial chain, the input of the sub-model is the index in the feature set, and the output value of the sub-model is the score or probability. The higher the output value, the higher the correlation between the enterprise and the corresponding industrial chain stage. It is possible to determine whether a candidate enterprise belongs to a specific industrial chain stage by setting a threshold. A candidate enterprise can correspond to multiple industrial chain stages, and it is also possible that the candidate enterprise cannot be classified into any industrial chain stage.

[0009] As an improvement of the present invention, in step 5, the industry chain stage identification language model is used to re-identify the candidate enterprises that have not been divided into the industry chain stage, and the output is whether they belong to a specific industry chain. Candidate enterprises that still cannot be divided into the industry chain stage are considered to be non-industry chain enterprises, and the industry chain enterprise identification results are formed. The input of the industry chain stage identification language model can be unstructured information in the feature set, and the natural language processing and other technologies are used to comprehensively identify the industry chain stage of the enterprise.

[0010] As an improvement of the present invention, in step 6, for enterprises in the positive data set but not in the industrial chain enterprise identification results, the information of the missed enterprises in the feature set can be compared with the knowledge base, and the newly added knowledge factors can be refined to complete the knowledge base. The negative data set and the industrial chain enterprise identification results can be compared, and the knowledge factors that cause the misjudgment of the misjudged enterprises can be located and deleted or weighted down.

[0011] As an improvement of the present invention, step 7 repeats steps 3 to 6 to update the recognition result until the verification deviation is less than the convergence preset value. Here, the verification deviation can be the weighted sum of the number of enterprises that have been missed or misjudged.

[0012] Compared with the prior art, the beneficial effects of the present invention are: 1) A method for constructing an industrial chain knowledge base is proposed, which can support a hierarchical industrial chain identification model; 2) A feedback mechanism is introduced in the industrial chain identification to improve the identification effect through iterative optimization of the knowledge base; 3) In the identification of enterprises in the industrial chain, a hierarchical model mechanism is adopted, which is a rough screening followed by a fine division, so that the model has a higher recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flow chart of the prior art; Figure 2 This is a flow chart of the industrial chain enterprise identification and analysis method of the present invention; Figure 3 This is a schematic diagram of the device architecture involved in the industrial chain enterprise identification and analysis method of the present invention. DETAILED DESCRIPTION

[0014] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0015] Example: Figures 2 to 3 As shown, the specific steps of an industrial chain enterprise identification and analysis method implemented in this implementation are as follows: Step 1: Build an industry chain knowledge base and verify the data set.

[0016] Industry chain information can be collected through expert experience, the Internet, large language models and other methods or channels to build an industry chain knowledge base and verify data sets.

[0017] Taking the biopharmaceutical industry chain as an example, Table 1 below is an example of an organizational form of the knowledge base, where weights are used to represent the contribution of the knowledge factor. The larger the weight, the higher the correlation with the industry chain.

[0018] Table 1. Example of industry chain knowledge base

[0019] The verification data set may include a positive data set and a negative data set, as shown in Table 2 below: Table 2. Classification of verification datasets

[0020] Step 2: Collect enterprise information and generate feature sets.

[0021] Enterprise information includes business registration information, business information, intellectual property information, etc.

[0022] The feature set may include numerical indicators, such as the sum of the frequency and weight of knowledge factors in the business scope, the sum of the number and weight of patents related to knowledge factors, etc. The feature set may also include unstructured information, such as a text description formed by extracting key information from the business scope, intellectual property rights, etc.

[0023] Step 3: Use the industrial chain matching model to identify candidate companies in the target industrial chain. The input of the industrial chain matching model is the indicators in the feature set, and the output value of the model is the score or probability. The higher the output value, the higher the correlation between the enterprise and the industry. The candidate companies in the target industrial chain can be screened out from the full amount of companies by setting a threshold.

[0024] For example, the industrial chain matching model takes the indicators such as "the sum of the weights of the knowledge factors appearing in the business scope" and "the number of patents related to the knowledge factors" as input, and the output is the relevance to the target industrial chain. The model construction method can be but is not limited to expert model, logistic regression, decision tree and other algorithms. Enterprises whose model output is greater than the set threshold are regarded as candidate enterprises in the target industrial chain.

[0025] Step 4: Use the industrial chain stage identification model to determine the industrial chain stage of the candidate enterprise.

[0026] The industrial chain stage identification model can be composed of several sub-models, corresponding to different stages of the industrial chain (upstream, midstream, and downstream). Among them, the input indicators of the upstream sub-model are more correlated with the upstream knowledge factors in the knowledge base, that is, the knowledge factors of the industrial chain stage "upstream" in Table 1. The midstream and downstream sub-models are also set up similarly.

[0027] Whether a candidate enterprise belongs to a specific industrial chain stage can be determined by setting a threshold. A candidate enterprise can correspond to multiple industrial chain stages, and it is also possible that a candidate enterprise cannot be classified into any industrial chain stage.

[0028] Step 5: Use the industry chain stage identification language model to re-identify candidate companies that have not been divided into industry chain stages, and output the conclusion of whether they belong to a specific industry chain stage. Candidate companies that still cannot be divided into industry chain stages are considered to be non-industry chain companies, and the industry chain enterprise identification results are formed.

[0029] For example, if an enterprise is judged not to belong to the upstream, midstream, or downstream in the previous step, the content of the enterprise in the feature set is further input into the industrial chain stage identification language model. The industrial chain stage identification language model can be a large language model or other model that can process natural language. The input content includes the business scope of the enterprise, intellectual property description, industrial chain knowledge base, etc. The industrial chain stage identification language model comprehensively determines whether the enterprise belongs to the target industrial chain and industrial chain stage.

[0030] Step 6: Use the verification data set to verify the industrial chain enterprise identification results, and then optimize the industrial chain knowledge base based on the verification results.

[0031] Compare the forward data set and the results of the industrial chain enterprise identification, and extract new knowledge factors to complete the knowledge base. Assuming that "ABC New Drug Development Co., Ltd." in Table 2 is not in the industrial chain enterprise identification results, a natural language processing algorithm is used to extract some key information not included in the current knowledge base from the feature set of ABC New Drug Development Co., Ltd., such as "pharmaceutical intermediates", and add it to the knowledge base, giving it a default weight (such as 0.8). An example of the completed knowledge base is shown in Table 3 below.

[0032] Table 3. Example of industry chain knowledge base (new knowledge factor)

[0033] The negative data set and the results of the industrial chain enterprise identification can be compared, and the knowledge factors that cause the misjudgment can be located for the misjudged enterprises and deleted or reduced in weight. Assuming that "JKL Pesticide Co., Ltd." in Table 2 is in the industrial chain enterprise identification results, the company's business scope and winning bid information contain descriptions of "pesticide raw materials", and the knowledge factor that caused the misjudgment of the business can be located as "raw materials" in Table 1, and the weight of this knowledge factor can be lowered. The adjusted results based on Table 3 are shown in Table 4 below.

[0034] Table 4. Example of industry chain knowledge base (updated weights)

[0035] Step 7: Repeat steps 3-6 to update the recognition results until the verification deviation is less than the convergence preset value. Here, the verification deviation can be the weighted sum of the number of enterprises that have been missed or misjudged.

[0036] For example, the weight of missed judgment can be set to 1, and the weight of misjudgment can be set to 0.8. Assuming that there are 3 missed judgments and 4 misjudged enterprises in the verification data set, the deviation value is 3*1+4*0.8=6.2. Assuming that the convergence preset value is 6, it means that the next round of optimization is required until the deviation value is less than 6.

[0037] It should be noted that the above content only illustrates the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.

Claims

1. A method for identifying and analyzing enterprises in an industrial chain, characterized in that: The steps include: Step 1: Build an industrial chain knowledge base, including collecting the main stages of the industrial chain and related technical descriptions and relevance information, and forming a number of knowledge factors to form the industrial chain knowledge base; build a verification data set, collect and annotate enterprise lists through multiple channels; Step 2: Collect enterprise information, perform correlation calculation with the industry chain knowledge base, generate quantitative indicators and textual descriptions of the correlation between specific enterprises and knowledge factors, and form an enterprise feature set; Step 3: Use the industrial chain matching model to identify the set of candidate enterprises in the target industrial chain. The input of the industrial chain matching model is the index in the feature set. The candidate enterprises in the target industrial chain are selected based on the correlation index output by the industrial chain matching model. Step 4: Use the industrial chain stage identification model to determine the industrial chain stage to which the enterprise belongs. The industrial chain stage identification model contains several sub-models, which correspond to different stages of the industrial chain. The input of the industrial chain stage identification model is the indicator in the feature set. Based on the predicted probability of the model output, it is determined whether the enterprise belongs to a specific industrial chain stage. Step 5: Use the industry chain stage recognition language model to re-identify the industry chain enterprises that have not been divided into stages. The input of the industry chain stage language recognition model is the textual description of the enterprise, and it is judged whether the enterprise belongs to a specific industry chain based on natural language processing technology; Step 6: Use the verification data set to compare the recognition results, conduct attribution analysis on the misjudgment and missed judgment of the industrial chain knowledge, identify the existing knowledge factors that cause misjudgment in the knowledge base, extract new knowledge factors that reduce missed judgment, and optimize the industrial chain knowledge base; Step 7: Update the recognition result until the deviation is less than the preset value.

2. The method for identifying and analyzing industrial chain enterprises according to claim 1, characterized in that: The verification data set in step one includes a positive data set and a negative data set. The positive data set contains a number of enterprises that have been clearly identified as belonging to the target industrial chain, and the negative data set includes a number of enterprises that have been clearly identified as not belonging to the target industrial chain or not belonging to a specific industrial chain stage.

3. The method for identifying and analyzing industrial chain enterprises according to claim 1, characterized in that: In step 2, the enterprise feature set processing forms numerical indicators and / or unstructured content by comparing enterprise information and industrial chain knowledge base.

4. The method for identifying and analyzing industrial chain enterprises according to claim 1, characterized in that: The input of the industrial chain matching model in step three is the indicator in the feature set, and the output value of the model is a score or probability. The higher the output value, the higher the correlation between the enterprise and the industry. By setting a threshold, candidate enterprises in the target industrial chain are screened out from all enterprises to form a set of candidate enterprises.

5. The method for identifying and analyzing industrial chain enterprises according to claim 1, characterized in that: The industrial chain stage identification model in step 4 can be composed of several sub-models, corresponding to different stages of the industrial chain. The input of the sub-model is the indicator in the feature set, and the output value of the sub-model is the score or probability. The higher the output value, the higher the correlation between the enterprise and the corresponding industrial chain stage.

6. The method for identifying and analyzing industrial chain enterprises according to claim 1, characterized in that: In step five, the industry chain stage identification language model is used to re-identify candidate enterprises that have not been divided into industry chain stages, and the conclusion of whether they belong to a specific industry chain stage is output. Candidate enterprises that still cannot be divided into industry chain stages are considered to be non-industry chain enterprises, and the industry chain enterprise identification results are formed.