Method and related device for acquiring associated companies based on enterprise and public institution information
By combining the stable node processing model and the sensitive node processing model, high confidence and low confidence samples are processed respectively, the problem of insufficient adaptability of dynamic nodes in the existing technology is solved, and efficient correlation analysis of enterprise and institution information is achieved.
Patent Information
- Application Number
- CN202510780051.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-12
AI Technical Summary
When processing enterprise and institution information, it is difficult to balance the accuracy of high confidence samples and the adaptability of dynamic nodes. Especially when facing emerging enterprise information or unseen node types, the model is prone to overfitting, which reduces the ability to generalize low confidence samples.
The method of combining the stable node processing model with the sensitive node processing model is adopted to train for common nodes and new nodes respectively. The stable node processing model is used for high confidence samples, and the sensitive node processing model is used for low confidence samples. By introducing a multi-head attention mechanism and dynamic rule supplement mechanism, the joint loss function is shared for parameter adjustment.
It improves the generalization ability and adaptability of the system, and can quickly adapt to emerging node types while maintaining the accuracy of high confidence samples, enhancing the accuracy and expansion of correlation analysis.
Smart Images

Figure CN120336526A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of information processing technology, and particularly relates to a method for obtaining associated companies based on the information of enterprises and institutions. Background Art
[0002] In modern society, obtaining information on related enterprises and institutions is of great significance. By analyzing the association relationships between enterprises, more comprehensive and accurate services can be provided to users, which plays an indispensable role in aspects such as business cooperation, risk assessment, and industrial chain optimization.
[0003] In the prior art, usually a single model is adopted to process the input materials so as to obtain the associated companies of the target enterprises and institutions. This method establishes the matching relationship between enterprises and associated units through text analysis and semantic understanding of enterprise information. However, although a single model performs well in processing high-confidence samples (such as obvious associated enterprises), its adaptability to dynamic nodes (such as newly emerging enterprise information or unseen node types) has limitations. Especially when a single model needs to process common nodes and dynamic nodes simultaneously, due to the complexity of the task, the model is prone to overfitting on high-confidence samples, thus reducing the generalization ability for new nodes and low-confidence samples. Summary of the Invention
[0004] In order to balance the accuracy of high-confidence samples and the adaptability to dynamic nodes, this application provides a method and related device for obtaining associated companies based on the information of enterprises and institutions.
[0005] In a first aspect, this application provides a method for obtaining associated companies based on the information of enterprises and institutions, adopting the following technical solution:
[0006] A method for obtaining associated companies based on the information of enterprises and institutions includes the following steps:
[0007] S1. Obtain the original data, clean the information of enterprises and institutions and calculate the confidence level, filter out the low-confidence data, and obtain the high-confidence data;
[0008] S2. Conduct semantic analysis and confidence verification on the high-confidence data, and screen out the credible associated data;
[0009] S3. Input the trusted association data into the stable node processing language model to obtain the final association result, wherein the stable node processing model and the sensitive node processing model are trained with each other, the stable node processing model is used to train common nodes and high-confidence samples, and the sensitive node processing model is used to train new nodes or low-confidence samples. The stable node processing model and the sensitive node processing model share the same input data, and the sensitive node processing model introduces an attention mechanism to strengthen the feature learning of new nodes; the stable node processing model and the sensitive node processing model share a comprehensive loss function, and the results of the stable node processing model and the sensitive node processing model are interactively fed back, and the parameters are adjusted according to the joint loss function.
[0010] Optionally, the structural design step of the stable node processing model includes:
[0011] The embedding layer is used to convert the input text into a context-related high-dimensional vector representation using BERT or RoBERTa. The model input is the concatenated text to be processed, which includes enterprise information and node descriptions. The enterprise information includes the name of the enterprise, the introduction of the enterprise, and the list of main business keywords. The node description includes the node name and node details.
[0012] The feature extraction layer is composed of multiple layers of Transformers, which is used to capture the semantic association between enterprise information and node descriptions;
[0013] The output layer is used to output the classification head and the associated probability. The classification head uses the output [CLS] tag of BERT as the overall semantic representation, and the associated probability passes through a fully connected layer and a softmax function.
[0014] Optionally, the structural design step of the sensitive node processing model includes:
[0015] An embedding layer, which is used to convert the input text into a context-related high-dimensional vector representation using BERT or RoBERTa, and to embed the dynamic features of the new node into the model using a dynamic rule embedding module; wherein the model input is a sample whose probability of being output by a stable node processing model is lower than a certain threshold, including enterprise information and node descriptions, the enterprise information includes the enterprise name, enterprise profile, main business keyword list and historical confidence information, the node description includes the new node name, new node details and dynamic keywords, and the historical confidence information is the predicted distribution of low-confidence samples obtained from the output of the second language model;
[0016] The feature extraction layer, which is composed of multiple stacked Transformers, is used to strengthen the semantic association between the enterprise and institution information and the new node description by introducing the multi-head attention mechanism, and the association between the dynamic keywords and the enterprise information is processed by a dedicated attention head;
[0017] The output layer includes a classification task head and a dynamic rule generation head. The classification task head is used to determine whether an enterprise or institution belongs to a new node, and the dynamic rule generation head is used to generate dynamic features for the stable node processing model to learn.
[0018] Optionally, the step of strengthening the semantic association between the enterprise and institution information and the new node description by introducing the multi-head attention mechanism includes the following steps:
[0019] Split the input embedding into multiple subspaces, that is, multiple heads;
[0020] Each head independently calculates the attention weights to capture different semantic features;
[0021] Concatenate and linearly transform the outputs of all heads.
[0022] Optionally, the joint training steps of the sensitive node processing model and the stable node processing model include:
[0023] S301. Independently train the sensitive node processing model and the stable node processing model respectively;
[0024] S302. Input the current sample into the stable node processing model for prediction and calculate the classification loss;
[0025] S303. Transmit the low-confidence samples to the sensitive node processing model to generate dynamic rule features;
[0026] S304. Calculate the difference between the predicted probability and the true label based on the joint loss function;
[0027] S305. Calculate the gradient by backpropagation based on the joint loss function;
[0028] S306. The parameters of the sensitive node processing model and the stable node processing model are updated simultaneously by the optimizer, and the parameter sharing part is synchronously adjusted between the two.
[0029] Optionally, the joint loss function is , is a hyperparameter used to control the contribution of each part of the loss to the training;
[0030] is the classification loss used to optimize the classification task of whether an enterprise or institution belongs to a node, ; where, It is a real label indicating whether an enterprise or institution belongs to a certain target node; It is the probability output predicted by the model; N is the number of samples;
[0031] It is the consistency loss between the stable node processing model and the sensitive node output model, , where, It is the output distribution of the stable node processing model, It is the output distribution of the sensitive node processing model; It is the Kullback–Leibler divergence, which is used to measure the difference between two probability distributions;
[0032] It is the dynamic rule supplementary loss, , where, It is the node feature infiltration of the stable node processing model, It is the dynamic feature embedding of the sensitive node processing model.
[0033] Optionally, the S304 includes the following sub-steps:
[0034] S3041. Combine the output of the sensitive node processing model, compare it with the result of the stable node processing model, and calculate the consistency loss;
[0035] S3042. Use the dynamic rule supplementary feature to optimize the feature generation of the stable node processing model;
[0036] S3043. Dynamically adjust the weights of the consistency loss and the dynamic rule supplementary loss according to the model training progress.
[0037] Optionally, the S1 includes the following sub-steps:
[0038] S11. Obtain the original data and perform preprocessing;
[0039] S12. Use the first pre-trained model for semantic analysis, map enterprise information and the industrial chain to the same semantic space, and determine the initial association result according to the semantic similarity;
[0040] S13. Perform confidence verification on the initial association result and screen out the credible associated data.
[0041] In a second aspect, the present application provides a system for obtaining associated companies based on enterprise and institution information, adopting the following technical solution:
[0042] A system for obtaining associated companies based on enterprise and institution information includes a processor, and a program of the method for obtaining associated companies based on enterprise and institution information as described in any one of the above is run in the processor.
[0043] In a third aspect, the present application provides a storage medium, adopting the following technical solution:
[0044] A storage medium stores a program of the method for obtaining associated companies based on the information of enterprises and institutions described in any one of the above.
[0045] In summary, the present application includes at least one of the following beneficial technical effects:
[0046] First, by combining a stable node processing model and a sensitive node processing model, the present invention respectively performs targeted training and processing on high-confidence samples and new nodes or low-confidence samples, which can not only ensure a high classification accuracy on known nodes, but also quickly adapt to newly emerged node types, thereby effectively improving the generalization ability and adaptability of the overall system.
[0047] Second, the present invention introduces a multi-head attention mechanism and a dynamic rule supplementation mechanism, so that when dealing with the association between the information of enterprises and institutions and industrial chain nodes, the model can fully explore hidden semantic connections and continuously learn key features of new nodes. By comprehensively considering classification loss, consistency loss, and dynamic rule supplementation loss in a common loss function, the system can continuously absorb new feature information while maintaining overall stability, further enhancing the accuracy and expansibility of association analysis. Description of the Drawings
[0048] Figure 1 is a flowchart of the method for obtaining associated companies based on the information of enterprises and institutions in a certain embodiment of the present application. Detailed Embodiments
[0049] The following details the embodiments of the present application, and the examples of the embodiments are shown in the drawings.
[0050] In the description of this specification, the description referring to the terms "certain embodiments", "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0051] The embodiments of the present application disclose a method for obtaining associated companies based on the information of enterprises and institutions, referring to Figure 1 , and including the following steps S1 - S3.
[0052] S1. Obtain the original data, clean the information of enterprises and institutions, calculate the confidence level, filter out the data with low confidence level, and obtain the data with high confidence level.
[0053] Optionally, the S1 includes the following sub-steps S11 - S13.
[0054] S11. Obtain the original data and perform preprocessing.
[0055] S12. Use the first pre-trained model to perform semantic analysis, map the enterprise information and the industrial chain to the same semantic space, and determine the initial association result according to the semantic similarity.
[0056] S13. Verify the confidence level of the initial association result and screen out the reliable associated data.
[0057] In this step, it is first necessary to obtain the original enterprise information and industrial chain data, and perform preliminary preprocessing and confidence level calculation on these data. Enterprise information usually includes enterprise name, main products, enterprise profile, and other data that can reflect the core business and scale of the enterprise. Through cleaning and statistical analysis, obviously incomplete or logically incorrect information can be removed, and entries with too low confidence level can be filtered out. The confidence level mentioned here can be understood as a quantitative evaluation of the reliability of the data, and is measured by indicators such as historical records, source credibility, and internal consistency of the data. For example, for a newly established enterprise, its main business information is not yet perfect, and the source channel lacks authority, so the calculated confidence level may be relatively low. After completing this filtering step, there will often be a relatively reliable and high-quality data set, laying a foundation for subsequent semantic analysis. Through this process, it can be ensured that the model has less noise interference when encountering new data, and more accurate results can be achieved in subsequent association analysis.
[0058] After the preprocessing is completed, it is necessary to perform semantic analysis on enterprise information and the industrial chain with the help of a stable node processing model. The so-called "semantic analysis" refers to using large-scale pre-trained language models (such as BERT, GPT, etc.) to process text, mapping the text into a high-dimensional vector representation to capture the implicit semantic relationships therein. By first pre-training on a large-scale general corpus, the model learns the general distribution laws of vocabulary, phrases, and sentence semantics; subsequently, it is fine-tuned on a small-scale dataset related to the industry to make the model more adaptable to the characteristics of specific industry texts. The reason for converting both enterprise information and industrial chain descriptions into vectors and uniformly mapping them into the same semantic space is to enable the model to measure the similarity between the two with the same set of semantic scales. The so-called "semantic space" can be understood as a multi-dimensional vector space: text contents such as enterprise names, enterprise profiles, and product keywords are successively converted into vectors in the embedding layer of the model; while the description information of industrial chain nodes also becomes vectors in the same embedding manner, and the relative distances or directions between these vectors reflect the similarities and differences in semantics between the two.
[0059] In this semantic space, by calculating the semantic similarity between vectors (common methods include cosine similarity or Euclidean distance), the degree of association between a certain enterprise information and a certain industrial chain node can be judged. Specifically, if the name of a certain enterprise or the keywords of its main business have a very high similarity with the description of a certain node, the model will initially determine that there is a relatively close connection between the two. Behind this calculation process, the context semantic capture ability of the pre-trained model is utilized: the model does not simply match keywords, but based on the language knowledge learned from a large-scale corpus, it can understand the semantic similarities between different expressions of enterprise names and industrial chain nodes. For example, assume that the main business of an enterprise is "smart home", and the description of an industrial chain node uses the term "smart home devices". Although they are not exactly the same literally, since the model can capture that they are both related to "home automation", the distance in the semantic space will be relatively close, thus being judged to have a preliminary association.
[0060] After completing the above vector mapping and similarity calculation, the system will judge which enterprises have significant associations with which industrial chain nodes based on a set threshold, and regard these matching results as the initial association results. The purpose of this preliminary screening is to ensure that the data volume is controllable and to include the most likely relevant enterprise-industrial chain combinations in the subsequent processing scope first. After this step, users can use this part of the preliminary association data to carry out further feature extraction or detailed analysis, realize the visualization and precise mining of the connection between enterprises and the industrial chain, and achieve the effect of improving efficiency and accuracy in actual business cooperation and risk assessment. Combining the good adaptability of the stable node processing model to common and high-confidence samples, this step can capture conventional business connections to the greatest extent in the initial stage and prepare for subsequent dynamic node adaptation and more complex sensitive node processing.
[0061] S2. Conduct semantic analysis and confidence verification on high-confidence data to screen out reliable associated data.
[0062] After completing the preliminary semantic analysis and association matching, it is necessary to conduct confidence verification on the obtained association results. The so-called confidence verification is an evaluation process for the reliability of the results. Based on elements such as semantic similarity, multi-source data support, and historical records, a score between 0 and 1, that is, a confidence score, will be assigned to the generated association results. When this score exceeds the pre-set threshold, the result can be regarded as reliable associated data; on the contrary, it is marked as low-confidence or uncertain data, and may require manual review or dynamic learning by the sensitive node processing model for further processing in the future.
[0063] In the scenario of associating with enterprise and institution information and industrial chain nodes, the implementation methods of confidence verification usually include the quantitative calculation of semantic similarity and the evaluation of the authority and consistency of data sources. For example, if the semantic similarity is higher than a threshold widely used in practice (such as 0.75 or 0.9) based on the output results of the stable node processing model, it can be determined that there is a significant association between the enterprise and institution name and the industrial chain node description at the semantic level; at the same time, if this association result is cross-verified by multiple authoritative data sources or public data of relevant departments, its confidence will be further improved. This comprehensive calculation process helps the system give a relatively objective confidence score for each association record, and then combined with the user-defined industry requirements, select an appropriate threshold to distinguish reliable data from suspicious data.
[0064] To make the confidence score more accurate, various scoring factors are introduced in practice. The most direct factor is the semantic similarity score, which is usually output by the model and normalized. Secondly, the authority of the data source itself cannot be ignored: if the identified associations only come from anonymous online texts, the score will be correspondingly reduced; if they come from the official websites of enterprises, institutions or professional databases, the score will be increased. In addition, the model sometimes also counts historical matching records. If an enterprise or institution has often been correctly associated with a certain type of node in previous analyses, then the historical performance will also be a great help in the scoring. For different fields or different business scenarios, the setting of specific thresholds may also vary: in high-risk or high-value fields (such as scientific research institutions, intellectual property, etc.), usually higher thresholds are used to ensure the accuracy of the results; in general manufacturing or trade fields, the thresholds can be appropriately reduced to avoid filtering out a large number of potential association opportunities.
[0065] Through this confidence verification mechanism, the system can not only screen out reliable association data for subsequent analysis or display, but also effectively identify which association results have high uncertainty, preparing for the next dynamic processing. If some association scores do not reach the threshold but are not completely invalid, they will be regarded as potential low-confidence samples, and these samples often need to be further learned and corrected by the sensitive node processing model.
[0066] S3. Input the reliable association data into the stable node processing language model to obtain the final association result. Among them, the stable node processing model and the sensitive node processing model are trained with each other. The stable node processing model is used to train for common nodes and high-confidence samples. The sensitive node processing models are trained with each other for new nodes or low-confidence samples. The stable node processing model and the sensitive node processing model share the same input data. The sensitive node processing model introduces an attention mechanism to strengthen the learning of the characteristics of new nodes; the stable node processing model and the sensitive node processing model share a comprehensive loss function, and the results of the stable node processing model and the sensitive node processing model are interactively fed back, and the parameters are adjusted according to the joint loss function.
[0067] Specifically, the structural design steps of the stable node processing model include an embedding layer, a feature extraction layer, and an output layer.
[0068] The embedding layer is used to convert the input text into a context-related high-dimensional vector representation using BERT (Bidirectional Encoder Representations from Transformers) or RoBERTa (Robustly optimized BERT approach); where the model input is the concatenated text to be processed, and the text to be processed includes enterprise and institution information and node descriptions. The enterprise and institution information includes the name of the enterprise and institution, the introduction of the enterprise and institution, and the list of main business keywords, and the node description includes the node name and the node details.
[0069] The feature extraction layer is stacked by multiple Transformers and is used to capture the semantic associations between enterprise and institution information and node descriptions.
[0070] The output layer is used to output the classification head and the association probability. The classification head uses the output [CLS] token of BERT as the overall semantic representation, and the association probability is obtained through a fully connected layer and a softmax function.
[0071] The stable node processing model undertakes the core function of efficiently processing common nodes and high-confidence samples in this solution. Its overall structure is based on a pre-trained Transformer framework (such as BERT or RoBERTa), mainly including an embedding layer, a feature extraction layer, and an output layer, and realizes the joint understanding of the semantics of the two parts of the text by concatenating the enterprise and institution information with the industrial chain node description into an input sequence. In order to better handle various business scenarios, the model will be fine-tuned according to industry characteristics after pre-training, making it more accurate and stable when processing conventional nodes or high-confidence texts.
[0072] In the embedding layer part of the model, all input texts need to be first converted into context-related high-dimensional vector representations. The specific method is to concatenate the information of the enterprise and institution (such as name, introduction, list of main business keywords) with the industrial chain node description (such as node name, node details), and insert a special token at the beginning to generate the overall semantic vector. For example, if the name of an enterprise and institution is "Green Energy Institution", its main business includes "photovoltaic equipment" and "wind turbine generators", and the industrial chain node description contains information related to "new energy equipment manufacturing", the model will serialize all these texts and input them into the pre-trained embedding layer. Since BERT or RoBERTa can learn rich context semantics based on a large amount of corpus, the model can not only identify the specific industries pointed to by these keywords, but also capture their internal semantic associations with concepts such as "new energy" and "manufacturing industry", thus constructing a more accurate vector representation in the high-dimensional space.
[0073] After the embedding is completed, the model enters the feature extraction layer stacked by multiple layers of Transformers. The multi-head attention mechanism in this part calculates the dependencies between each position in the sequence and other positions, and extracts deeper semantic interactions between texts. For example, if there are similar or complementary keywords between the main business of an enterprise or institution and the description of a certain industrial chain node, the multi-head attention can accurately capture this correspondence and highlight it in the subsequent feature representation. In this way, the model can distinguish the subtle differences between concepts such as "photovoltaic" and "wind power", and can also understand that they belong to a higher-level semantic category such as "new energy equipment". Since the stable node processing model is mainly oriented towards common nodes and high-confidence data, its focus often lies in how to provide highly accurate classification results within the known business domain and minimize misjudgment or neglect of such data. The deep learning ability of multiple layers of Transformers exactly meets this requirement and can maintain stable association judgments in the texts of a large number of enterprises, institutions and industrial chain nodes.
[0074] When the feature extraction is completed, the model will concentrate the context representation of the entire sequence into the output layer, and use a fully connected layer and a softmax function to output the association probability. At this time, the hidden state vector of the [CLS] token will be passed into the classification head as the semantic aggregation vector representing the entire input sequence (including the information of the enterprise or institution and the description of the industrial chain node), and then calculate whether the current input combination has a high degree of association. For example, if the comprehensive semantic similarity between a "photovoltaic equipment manufacturer" and "new energy equipment manufacturing" exceeds the threshold, it will be determined by the model to have a high probability of belonging to the same industrial chain node. In addition, in many application scenarios, the cosine similarity between the text of the enterprise or institution and the text of the node can also be calculated to obtain a more refined association degree score, providing more multi-dimensional reference indicators for subsequent confidence verification and business decision-making.
[0075] Since this model has been pre-trained on a large-scale corpus and then fine-tuned with industry data, it can give relatively stable and accurate association results in most common cases. In this way, when the system processes high-confidence samples, the stable node processing model can often make accurate judgments immediately and directly feedback the results to the user or enter the downstream process. For dynamic node data with low confidence or containing new concepts, it can be further processed by the sensitive node processing model through reinforcement learning or refined processing, so as to achieve a balance between high accuracy and good adaptation to new nodes globally.
[0076] Specifically, the structural design steps of the sensitive node processing model include an embedding layer, a feature extraction layer and an output layer.
[0077] The embedding layer is used to convert the input text into a context-related high-dimensional vector representation using BERT or RoBERTa, and use the dynamic rule embedding module to embed the dynamic features of the new node into the model; wherein the model input is a sample whose probability of being output by the stable node processing model is lower than a certain threshold, including enterprise information and node description, the enterprise information includes the enterprise name, enterprise profile, main business keyword list and historical confidence information, the node description includes the new node name, new node details and dynamic keywords, and the historical confidence information is the predicted distribution of low-confidence samples obtained from the output of the second language model;
[0078] The feature extraction layer is composed of multiple layers of Transformer stacking, which is used to strengthen the semantic association between enterprise information and new node descriptions by introducing a multi-head attention mechanism. The association between dynamic keywords and enterprise information is processed by a dedicated attention head.
[0079] The output layer contains a classification task head and a dynamic rule generation head. The classification task head is used to determine whether an enterprise or institution belongs to a new node, and the dynamic rule generation head is used to generate dynamic characteristics that can be used for learning by the stable node processing model.
[0080] The sensitive node processing model is mainly aimed at those new industry chain nodes and low-confidence samples that have not been seen before, and supplements or strengthens the analysis ability of this part of uncertain data through dynamic learning and attention mechanism. In terms of overall structure, it is also based on the pre-trained Transformer model, but in the input design and feature extraction links, there are more processing methods specifically for the dynamic characteristics of new nodes, so as to quickly adapt the enterprise information and unknown or uncommon node descriptions.
[0081] In the input stage, the model receives low-confidence samples from the previous stage (i.e., output by the stable node processing model) as the main data source. These samples often have low confidence due to large semantic differences, new business areas, or lack of high-quality historical records. Among them, the information of enterprises and institutions will include the name, introduction, list of main business keywords, and optional historical confidence information; while the node part may be the name of the new node, the descriptive text of the node, and dynamic keywords (usually extracted from recent industry reports, news, or the latest research). The reason why they are called "sensitive nodes" is that these nodes may contain new technologies, new industry directions, or features that have not yet been incorporated into the existing rule system. If they are not processed separately, they are easily misjudged or ignored by traditional models.
[0082] In the embedding layer, the model utilizes pre-trained language models such as BERT, RoBERTa, or similar ones to convert text sequences into context-related high-dimensional vector representations. Different from conventional embedding methods, it additionally introduces a dynamic rule embedding module to learn new node features and incorporate them into the model. Briefly speaking, in addition to conventional text vectors, the model also generates specific "new node vectors" based on dynamic keywords, industry hotspots, and possible historical confidence features, and fuses them with the text vectors of enterprises and institutions. This approach helps the model capture emerging or less popular node concepts more quickly, making the subsequent semantic interaction process no longer limited to known fields.
[0083] In the feature extraction layer, the stacking of multiple Transformer layers provides the ability to learn the deep semantic associations between different text segments. Since the sensitive node processing model focuses on the potentially less obvious relationships between new nodes and enterprises and institutions, in the design of the attention mechanism, separate attention heads are used for dynamic keywords and enterprise profiles, main business keywords, ensuring that potential new technologies or business points can be highlighted. For example, if the dynamic keyword of a new node is "hydrogen fuel cell", and the enterprise profile mentions words such as "hydrogen production process" and "gas purification equipment", the model can capture the potential intersection signs between the two in the industrial chain through a dedicated attention mechanism.
[0084] Finally, in the output layer, the model generally includes two parts: one is the classification task head, which is used to determine whether the current enterprise or institution belongs to the given new node and outputs an association probability; the other is the dynamic rule generation head, which generates dynamic characteristics or rules related to the new node according to the new industry information captured by the model during training. These dynamic rules can be used to improve the description of similar new nodes and also feed back to the stable node processing model, enabling the overall system to gradually enrich the known node rule library in subsequent iterations, thereby improving the processing efficiency and accuracy of similar new nodes. Through cooperation with the stable node processing model, the sensitive node processing model can effectively share the data analysis work of those difficult or not yet finalized nodes.
[0085] Optionally, the step of strengthening the semantic association between enterprise and institution information and new node descriptions by introducing the multi-head attention mechanism includes the following steps a - c.
[0086] a. Split the input embedding into multiple subspaces, i.e., multiple heads.
[0087] b. Each head independently calculates attention weights to capture different semantic features; among them, the attention weight is: Q is the query vector, representing the feature representation in enterprise information; K is the key vector, representing the feature representation in the new node description; V is the value vector, representing the semantic information of the new node description; d is the dimension of the vector, and the scaling factor is used to balance the calculation. T represents the transpose of the matrix.
[0088] c. Concatenate and linearly transform the outputs of all heads, that is
[0089] . Among them, ; is the independent parameter matrix for each head, is the output transformation matrix.
[0090] When performing semantic interaction between enterprise and institution information and new node descriptions, complex and variable context-dependent relationships are often encountered. As new nodes emerge continuously, there may be implicit or cross-sentence connections between them and texts such as enterprise and institution profiles or main businesses. Relying solely on a single attention head often fails to capture all details completely. By having multiple attention heads work in parallel, the model can focus on multi-dimensional information such as keywords, syntactic structures, and upstream and downstream semantics from different perspectives, assign appropriate attention degrees to multiple aspects of the same input respectively, and make the overall semantic mapping more refined. In this process, the model can adaptively adjust the weights of each head, so as to flexibly highlight key points and weaken noise when facing different types of inputs.
[0091] In specific implementation, both the stable node processing model and the sensitive node processing model will split the input text into several vector sequences. Enterprise and institution information is regarded as the query vector Q, while the new node description and its dynamic keywords are used as the key vector K and the value vector V for the model to calculate attention. The core of multi-head attention is to perform dot product operation on Q and K and scale it, then obtain the attention distribution through the softmax function, and multiply it by V to generate the context vector. Since the parameters of each head are independent of each other, different heads will gradually learn to capture different semantic features during training. For example, the first head may focus on the direct matching of keywords, the second head may study syntactic dependencies or temporal information in the context, and the third head may pay more attention to long-distance cross-references. In this way, the same enterprise and institution information and new node description will generate several sub-features in different heads, and finally after concatenation and linear transformation, a more comprehensive context representation is output.
[0092] For example, in the profile of an enterprise or institution, it is stated that "focus on the research and development of energy storage systems", while the new node description mentions "new energy energy storage equipment". When the multi-head attention starts to calculate, the model first maps "research and development of energy storage systems" into multiple vector segments in Q, and texts such as "new energy energy storage equipment" enter K and V respectively. Each head independently calculates the matching degree between them, and judges the semantic relevance between the two texts through cosine similarity or other metrics, and then focuses the attention scores on the most relevant part. For example, when the attention mechanism realizes that "energy storage system" and "energy storage equipment" are very close in industry terms, it will give higher weights to highlight this association. After each head obtains its own attention distribution respectively, the model then combines their outputs into a unified vector representation and passes it to the subsequent classification layer or dynamic generation module to confirm the association probability between the enterprise or institution and the new node, and to explore new industrial characteristics.
[0093] On the one hand, multi-head attention can effectively capture the corresponding relationships between keywords, helping the system to emphasize those semantic points that can reflect the potential of business intersection or technical collaboration; on the other hand, it can also span multiple sentence ranges, track context information, enabling the model to recognize that concepts scattered in multiple places in the enterprise profile and new node description can be closely related.
[0094] Optionally, the joint training steps of the sensitive node processing model and the stable node processing model include:
[0095] S301. Independently train the sensitive node processing model and the stable node processing model respectively;
[0096] S302. Input the current sample into the stable node processing model for prediction and calculate the classification loss;
[0097] S303. Transmit the low-confidence samples to the sensitive node processing model to generate dynamic rule features;
[0098] S304. Based on the joint loss function, calculate the difference between the predicted probability and the true label;
[0099] S305. Based on the joint loss function, calculate the gradient through backpropagation;
[0100] S306. The parameters of the sensitive node processing model and the stable node processing model are updated simultaneously by the optimizer, and the parameter sharing part is synchronously adjusted between the two.
[0101] In this joint training process, the stable node processing model and the sensitive node processing model need to be trained separately first, so that they can obtain relatively mature initial weights on their respective tasks. The stable node processing model mainly learns common nodes and high-confidence samples, focusing on ensuring the classification accuracy and model stability of these data; while the sensitive node processing model focuses on dynamic adaptation of new nodes or low-confidence samples, aiming to improve the generalization ability of unknown scenarios through mechanisms such as active learning and rule generation.
[0102] When independent training is completed, the training process will enter the joint training stage (S302-S306). In this stage, the stable node processing model will first predict the current sample and calculate the classification loss; if there are low-confidence results, these samples will be passed to the sensitive node processing model for generating new dynamic rule features. The purpose of this step is to allow the sensitive node processing model to focus on those samples that the stable node processing model cannot accurately identify or have insufficient confidence, and to generate dynamic features that are instructive for new fields by mining more subtle semantic features or new node information in these data. Subsequently, the system will measure the deviation between the current prediction probability and the true label based on the joint loss function, and combine the consistency loss and the dynamic rule supplement loss for overall back propagation. In this joint loss function, the classification loss mainly guarantees the overall prediction performance, the consistency loss constrains the output distribution of the stable node processing model and the sensitive node processing model from completely deviating, and the dynamic rule supplement loss helps the stable node processing model gradually absorb the new node features captured by the sensitive node processing model, thereby achieving improvements for low-confidence samples in the next round of training.
[0103] As the loss backpropagates, the parameters of the two models are updated synchronously through the optimizer, and the shared parameters are adjusted in conjunction. This means that the experience gained by the stable node processing model when facing common node data will complement the adaptive ability of the sensitive node processing model to new nodes; on the other hand, the dynamic rule features extracted by the sensitive node processing model will also be reflected in the shared part, allowing the system to continuously integrate new knowledge and adapt to more changes. For example, assuming that a new "hydrogen energy storage" node appears in the manufacturing field, the stable node processing model may produce low-confidence predictions due to the lack of relevant prior information, but the sensitive node processing model can identify and extract keywords such as "hydrogen fuel cell" or "hydrogen production equipment", and then pass these dynamic characteristics back to the stable node processing model, so that it can gradually learn to make more accurate judgments on this type of new node.
[0104] Specifically, the joint loss function is , is a hyperparameter used to control the contribution of each part loss to training;
[0105] is the classification loss, which is used to optimize the classification task of whether enterprises and institutions belong to nodes. ; where is the true label, indicating whether an enterprise or institution belongs to a certain target node; is the probability output predicted by the model; N is the number of samples;
[0106] is the consistency loss between the stable node processing model and the sensitive node output model. , where is the output distribution of the stable node processing model. is the output distribution of the sensitive node processing model; is the Kullback–Leibler divergence, abbreviated as KL divergence, which is used to measure the difference between two probability distributions; minimizing aims to make the outputs of the two models consistent in semantic classification.
[0107] is the dynamic rule supplement loss. , where is the node feature infiltration of the stable node processing model. is the dynamic characteristic embedding of the sensitive node processing model.
[0108] To better integrate the collaborative training effects of the stable node processing model and the sensitive node processing model, the solution of the present invention introduces a joint loss function L composed of three parts of losses, and adopts a linear weighted form to synthesize the priorities of different objectives. The first item is the classification loss L1, which is used to measure the difference between the current model prediction and the true label, so as to optimize the classification accuracy of whether enterprises and institutions belong to the target node. This part usually adopts the form of cross-entropy, and improves the ability of the model to distinguish positive and negative samples by minimizing the negative log-likelihood on the training set. In practical applications, for example, when we want to judge whether a "new energy institution" belongs to the "photovoltaic industrial chain", we will compare the predicted probability given by the model with the true label. If the classification is incorrect or the uncertainty is high, a large loss value will be generated, and then the model will be prompted to adjust the parameters in the backpropagation, in order to reduce similar classification errors in subsequent training.
[0109] The second term is the consistency loss L2, which uses the KL divergence to measure the difference between the output distribution P1 of the stable node processing model and the output distribution P2 of the sensitive node processing model. Its core significance lies in encouraging the two models to maintain a certain consistency in the output distribution for the same input sample. Especially when encountering marginal cases between high confidence and low confidence, if the associated predictions of the stable node processing model and the sensitive node processing model differ greatly, L2 will increase. By adjusting the hyperparameter β to increase or decrease the constraint on consistency, the cooperation degree of the two models can be strengthened in the early stage of training, or they can be allowed to give full play to their respective advantages in the later stage of training, making the system more flexible. For example, if the stable node processing model has a high accuracy for common nodes such as the "hydrogen energy industry chain", and the sensitive node processing model also agrees with this result, the output probability distributions of the two will tend to be consistent; if the sensitive node processing model draws different conclusions based on new features or new semantic analysis, it will increase L2 and force the two to learn from each other in the backpropagation stage.
[0110] The third term is the dynamic rule supplement loss L3, which mainly aims at the differences between the stable node processing model and the sensitive node processing model at the feature representation level, and uses the vector distance to measure this difference. Specifically, is the vector representation of the stable node processing model in the node feature infiltration part, is the dynamic feature embedding generated by the sensitive node processing model. The squared Euclidean distance between the two reflects the deviation degree of the learned node semantics in the high-dimensional space. By minimizing this distance during training, the stable node processing model can gradually absorb the new node information or special features of low-confidence samples provided by the sensitive node processing model, helping it to better adapt to unseen scenarios in subsequent predictions. For example, when the sensitive node processing model discovers that "hydrogen production equipment" and "hydrogen fuel cell" are strongly related features, it will incorporate the corresponding dynamic keywords into ; if the of the stable node processing model is still insensitive to these features, the loss value will increase, driving the system to update the weights during iteration, and gradually converging the representations of the two.
[0111] Overall, the loss functions of these three parts have different roles: the classification loss L1 ensures the accuracy of the core task, the consistency loss L2 encourages the output distributions of the two models to learn from each other, and the dynamic rule supplement loss L3 helps to stabilize the node processing model as it continuously absorbs new information. By setting different hyperparameters α, β, and γ to regulate the contribution weights of the three in training, it can be flexibly switched according to the focus of the business scenario: if overall stability needs to be strengthened, the proportion of the classification loss can be increased; if new nodes appear frequently, the proportions of L2 and L3 need to be increased to make the best use of the dynamic rules generated by the sensitive node processing model. Ultimately, this multi-objective joint optimization method can significantly improve the adaptability of the system in practical applications, ensuring both performance on high-confidence data and the timely capture of low-confidence data or new node information.
[0112] Further, the step S304 includes the following sub-steps S3041 - S3043.
[0113] S3041. Combine the output of the sensitive node processing model and compare it with the result of the stable node processing model to calculate the consistency loss.
[0114] S3042. Use the dynamic rule supplement feature to optimize the feature generation of the stable node processing model.
[0115] S3043. Dynamically adjust the weights of the consistency loss and the dynamic rule supplement loss according to the model training progress.
[0116] To better complete the joint optimization between the stable node processing model and the sensitive node processing model, it is necessary to combine the output results of the two models and make mutual corrections at different levels. First, by comparing the prediction given by the sensitive node processing model for the same input sample with the classification result of the stable node processing model, it is possible to measure whether there are significant differences in classification or association determination between the two. If the difference is large, it means that there are obvious disagreements in their interpretations of the sample, and further convergence of their output distributions is needed through the consistency loss. For example, when the stable node processing model has high stability in processing traditional manufacturing nodes, while the sensitive node processing model outputs a different probability distribution because it captures recently updated industry information, it is necessary to calculate the KL divergence (i.e., the consistency loss) of their output results and moderately adjust the parameters during backpropagation to make the two models gradually reach relatively consistent judgments for similar samples.
[0117] After calculating the consistency loss, the dynamic rule features generated by the sensitive node processing model can be used to optimize the internal representation of the stable node processing model. Especially for low-confidence samples with new industrial trends or characteristics in unknown fields, the sensitive node processing model often extracts additional semantic clues from dynamic keywords or new business descriptions. These pieces of information will be supplemented into the stable node processing model in the form of embedding vectors or dynamic rules, thereby helping the latter improve the recognition accuracy of similar new nodes in future predictions. For example, if the sensitive node processing model learns certain key phrases or feature vectors from a batch of samples related to "new energy materials", these features can be incorporated into the node feature embedding part of the stable node processing model, enabling the model to gradually master the understanding ability of emerging directions such as "hydrogen energy storage" or "solid-state batteries".
[0118] During the training process, the weights of the consistency loss and the dynamic rule supplementation loss can also be dynamically adjusted according to the training progress. If the system mainly focuses on the cooperation between models in the early stage, the proportion of the consistency loss can be increased to ensure that the output distributions of the two models are as consistent as possible, thereby quickly constructing a globally stable classification or association framework. When the training enters the middle and late stages, as the understanding of the two models gradually converges, the proportion of the consistency loss can be gradually reduced, and instead, the dynamic rule supplementation loss can be strengthened, allowing the stable node processing model to have more room to receive and integrate new features, thus achieving more significant improvements in refined adaptation or new node expansion. Through such an adaptive adjustment mechanism, the system can ultimately strike a balance between stability and scalability and still maintain a high prediction accuracy and association discovery ability in the face of rapid changes or the continuous emergence of new nodes in different industrial fields.
[0119] The embodiment of the present application also discloses a system for obtaining associated companies based on enterprise and institution information, including a processor, in which a program of the method for obtaining associated companies based on enterprise and institution information described in any one of the above is run.
[0120] The embodiment of the present application also discloses a storage medium storing a program of the method for obtaining associated companies based on enterprise and institution information described in any one of the above.
[0121] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for obtaining associated companies based on the information of enterprises and institutions, characterized in that It includes the following steps: S1. Obtain the original data, clean the information of enterprises and institutions and calculate the confidence level, filter out the low-confidence data, and obtain the high-confidence data; S2. Conduct semantic analysis and confidence verification on the high-confidence data, and screen out the credible associated data; S3. Input the credible associated data into the stable node processing language model to obtain the final association result. Among them, the stable node processing model and the sensitive node processing model are trained with each other. The stable node processing model is used for training with respect to common nodes and high-confidence samples, and the sensitive node processing model is used for training with respect to new nodes or low-confidence samples. The stable node processing model and the sensitive node processing model share the same input data. The sensitive node processing model introduces an attention mechanism to strengthen the learning of the characteristics of new nodes; the stable node processing model and the sensitive node processing model share a comprehensive loss function and the results of the stable node processing model and the sensitive node processing model are interactively fed back, and the parameters are adjusted according to the joint loss function.
2. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 1, wherein The structural design steps of the stable node processing model include: The embedding layer is used to convert the input text into a context-related high-dimensional vector representation using BERT or RoBERTa; among them, the model input is the concatenated text to be processed, and the text to be processed includes the information of enterprises and institutions and the node description. The information of enterprises and institutions includes the name of the enterprise or institution, the introduction of the enterprise or institution, and the list of main business keywords, and the node description includes the node name and the node details; The feature extraction layer is stacked by multiple Transformers and is used to capture the semantic association between the information of enterprises and institutions and the node description; The output layer is used to output the classification head and the association probability. The classification head uses the [CLS] token output by BERT as the overall semantic representation, and the association probability passes through a fully connected layer and the softmax function.
3. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 2, wherein, The structural design steps of the sensitive node processing model include: The embedding layer is used to convert the input text into a context-related high-dimensional vector representation using BERT or RoBERTa, and the dynamic features of the new node are infiltrated into the model using the dynamic rule embedding module; among them, the model input is the sample with the probability output by the stable node processing model lower than a certain threshold, including the information of enterprises and institutions and the node description. The information of enterprises and institutions includes the name of the enterprise or institution, the introduction of the enterprise or institution, the list of main business keywords, and the historical confidence information. The node description includes the new node name, the new node details, and the dynamic keywords. The historical confidence information is the prediction distribution of the low-confidence samples obtained from the output of the second language model; The feature extraction layer is stacked by multiple Transformers and is used to strengthen the semantic association between the information of enterprises and institutions and the new node description by introducing the multi-head attention mechanism, and the association between the dynamic keywords and the enterprise information is processed by a special attention head; The output layer includes a classification task head and a dynamic rule generation head. The classification task head is used to judge whether the enterprise or institution belongs to the new node, and the dynamic rule generation head is used to generate the dynamic characteristics that can be learned by the stable node processing model.
4. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 3, wherein, The steps of strengthening the semantic association between the information of enterprises and institutions and the description of new nodes by introducing the multi-head attention mechanism include the following steps: The input embedding is divided into multiple subspaces, i.e., multiple heads; Each head independently calculates the attention weights to capture different semantic features; The outputs of all heads are concatenated and linearly transformed.
5. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 4, wherein The joint training steps of the sensitive node processing model and the stable node processing model include: S301. Independently train the sensitive node processing model and the stable node processing model respectively; S302. Input the current sample into the stable node processing model for prediction and calculate the classification loss; S303. Transmit the low-confidence samples to the sensitive node processing model to generate dynamic rule features; S304. Based on the joint loss function, calculate the difference between the predicted probability and the true label; S305. Based on the joint loss function, calculate the gradient through backpropagation; S306. The parameters of the sensitive node processing model and the stable node processing model are updated simultaneously by the optimizer, and the parameter sharing part is synchronously adjusted between the two.
6. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 5, characterized in that The combined loss function is , is a hyperparameter used to control the contribution of each part of the loss to the training; is the classification loss, which is used to optimize the classification task of whether enterprises and institutions belong to a node, ; where, is the true label, indicating whether an enterprise or institution belongs to a certain target node; is the probability output predicted by the model; N is the number of samples; To stabilize the consistency loss between the node processing model and the sensitive node output model, , where is the output distribution of the stable node processing model, is the output distribution of the sensitive node processing model; is the Kullback–Leibler divergence, which is used to measure the difference between two probability distributions; Supplement losses for dynamic rules , where is the node feature infiltration for stabilizing the node processing model is the dynamic characteristic embedding for the sensitive node processing model 7. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 6, wherein The said S304 includes the following sub-steps: S3041. Combine the output of the sensitive node processing model, compare it with the result of the stable node processing model, and calculate the consistency loss; S3042. Use the dynamic rule supplementary features to optimize the feature generation of the stable node processing model; S3043. Dynamically adjust the weights of the consistency loss and the dynamic rule supplementary loss according to the model training progress.
8. The method for obtaining associated companies based on the information of enterprises and institutions according to claim 7, wherein, The said S1 includes the following sub-steps: S11. Obtain the original data and perform preprocessing; S12. Use the first pre-trained model for semantic analysis, map the enterprise information and the industrial chain to the same semantic space, and determine the initial association result according to the semantic similarity; S13. Conduct confidence verification on the initial association result and filter out the credible association data.
9. A system for obtaining associated companies based on the information of enterprises and institutions, characterized in that, It includes a processor, and a program of the method for obtaining associated companies based on the information of enterprises and institutions as described in any one of claims 1-8 runs in the processor.
10. A storage medium, characterized in that, Store a program of the method for obtaining associated companies based on the information of enterprises and institutions as described in any one of claims 1-8.
Citation Information
Patent Citations
News interpretation method and system based on natural language processing
CN119025670A
Business opportunity recommendation method and device for B-type enterprises, equipment, medium and product
CN119295120A
Network intrusion detection method based on pre-training language model federal segmentation learning
CN119766574A
Key information extraction method and system based on multi-modal model
CN119892215A
Administrative institution internal control information processing method and system based on big data analysis
CN119989421A
Cited By
Open knowledge base-oriented concept dependency relationship recognition method, medium and equipment
CN120995052A
Conceptual dependency relationship identification method, medium, and device for open knowledge base
CN120995052B
Position searching method and device based on implicit intention recognition
CN121144373A