Enterprise portrait automatic generation system and method based on multi-source data fusion
Through multi-source data fusion, the multi-dimensional map is constructed, and the problems of multi-source heterogeneous data fusion and dynamic update in the enterprise portrait system are solved, real-time perception of dynamic changes of enterprises and the mining of implicit features, and the risk warning and business decision support capabilities of enterprise portraits are improved.
Patent Information
- Application Number
- CN202510840261.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing enterprise portrait system is difficult to integrate multi-source heterogeneous data, lacks real-time tracking of enterprise dynamic changes and incremental update capabilities, cannot reflect the evolutionary characteristics and potential risks of enterprise relationships, and lacks sufficient portrayal and evaluation of the implicit characteristics of enterprises.
Through multi-source data fusion, a multi-dimensional map is built, heterogeneous data acquisition, knowledge extraction and alignment, and incremental update mechanisms are adopted, combined with BiLSTM-CRF model and dependent syntax analysis, correlation weights are calculated, enterprise portraits are generated and visualized, correlation weights are updated in real time using the exponential attenuation model, implicit features are mined and dynamic labels are set.
It realizes the comprehensive integration and accurate extraction of multi-source data, enhances the perception of dynamic changes of enterprises, improves risk warning and intelligent analysis capabilities, provides a multi-dimensional perspective of enterprise operation characteristics, and enhances the practical value and business decision support of enterprise portraits.
Smart Images

Figure CN120338619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent enterprise portrait generation, and specifically to an automated enterprise portrait generation system and method based on multi-source data fusion. Background Art
[0002] Intelligent enterprise portrait generation refers to collecting, fusing, and modeling multi-source heterogeneous data of an enterprise to construct a multi-dimensional, structured, visual, and interactive digital enterprise portrait system, and realizing dynamic update and intelligent analysis; With the rapid development of big data, artificial intelligence, and knowledge graph technologies, enterprise portraits play an increasingly important role in scenarios such as financial risk control, industrial chain analysis, and market insight. Traditional enterprise portrait methods mainly rely on structured data, such as industrial and commercial registration information, financial statements, and transaction records, and the construction methods are mostly static and single-dimensional, making it difficult to reflect the multi-dimensional characteristics and dynamic evolution relationships of enterprises in actual operations. However, the business operations of enterprises in the real world involve a large number of heterogeneous data sources, such as regulatory records, news and public opinion, patent applications, supply chain collaboration platform data, bidding documents, etc. These data types are diverse and the formats are complex; there are large structural differences and inconsistent semantic expression methods among multi-source data, making it difficult to uniformly parse and establish effective association relationships; existing systems mostly use isolated indicators to describe enterprise characteristics, lacking a method based on unified entity relationship modeling, making it difficult to reflect the complex supply and demand, shareholding, technical dependence, etc. network structures and influence among enterprises; enterprise portraits are often constructed based on a snapshot at a certain point in time, lacking the ability to track real-time data and incrementally update the graph, unable to reflect the evolution characteristics of enterprise relationships and potential risk changes, and traditional portraits focus on displaying basic information, such as legal person, registered capital, business scope, etc., and do not fully depict and evaluate the hidden characteristics of enterprises in the industrial chain, such as roles, associated risks, break risks, etc. Therefore, there is an urgent need for an enterprise portrait generation method that can fuse multi-source heterogeneous data, construct a multi-dimensional association graph, support dynamic update, and calculate hidden characteristics. Summary of the Invention
[0003] The purpose of the present invention is to provide an automated enterprise portrait generation system and method based on multi-source data fusion to solve the problems raised in the prior art.
[0004] To achieve the above purpose, the present invention provides the following technical solution: An automated enterprise portrait generation method based on multi-source data fusion, the enterprise portrait automated generation method specifically includes the following steps: Step S100, heterogeneous data collection, performing knowledge extraction and alignment on the collected heterogeneous data; Step S200: Identify the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and the association weights of different association relationships; Step S300: Set an incremental update mechanism for the constructed multi-dimensional graph, and regularly update the nodes, association relationships, and the association weights of different association relationships; Step S400: According to the multi-dimensional graph updated regularly, extract the structural features of the multi-dimensional graph to generate an enterprise portrait, obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait through the dynamic labels; Step S500: Visually display the generated enterprise portrait.
[0005] In step S100, heterogeneous data collection is performed, specifically: Classify and collect heterogeneous data, including structured data, semi-structured data, and unstructured data; The structured data includes industrial and commercial registration, tax records, patent databases, and supply chain system logs; The semi-structured data includes bidding documents and company annual reports; The unstructured data includes news sentiment and industry research reports.
[0006] Perform knowledge extraction and alignment on the collected heterogeneous data, specifically: Directly extract data entities from structured data; Extract data entities from semi-structured data through manual format parsing and keyword matching methods; For unstructured data, use the method based on the BiLSTM-CRF model to extract enterprise, product, and legal person data entities; Perform knowledge extraction based on the semantic relationship of dependency syntax; Unify the time and location in heterogeneous data from different sources into a standardized format.
[0007] Optionally, the method for unifying time and location includes spatio-temporal association alignment, specifically: Enterprise - geography binding: Associate the enterprise to GeoHash through the registered address entity data; Event - time binding: Associate the data entities of supply chain events to a standardized timestamp (UTC time); In step S200, the data entities in the heterogeneous data are determined as nodes, the knowledge extraction results are determined as association relationships, and the association weights of different association relationships are calculated. Based on the nodes, association relationships, and the association weights of different association relationships, a multi-dimensional graph is constructed as follows: Step S201: Determine the data entities as nodes, where the nodes include enterprises, products, and legal persons; Step S202: Determine the knowledge extraction results as association relationships, where the association relationships include enterprise - production - product, enterprise - supply - enterprise, enterprise - holding - legal person, product - dependence - product, and geographical location - aggregation - enterprise; Preferably, the representation of the nodes is specifically as follows: Enterprise: Industrial and commercial registration number, industry classification code; Product: HS code, technical complexity level; Transaction: Transaction amount, frequency, timestamp; Investment: Shareholding ratio, investment amount; Legal person: Depth of shareholding chain, control coefficient; Geographical location: Latitude and longitude, economic location entropy; Step S203: Calculate the association weights of different association relationships respectively by fitting the functional relationship through historical data; For the calculation of the association weight of the transaction relationship, the fitted functional relationship is characterized as: ; where, w ij represents the association weight from node i to node j, and i and j represent different node identifiers; a ij represents the annual transaction amount from node i to node j; a total represents the total industry transaction amount; z represents the annual transaction frequency; Among them, the association weight of the investment relationship is represented by the shareholding ratio or the investment amount; For the legal person relationship, product supply relationship, etc., only their association relationships need to be represented for easy viewing, and there is no need to calculate the association weights; Among them, according to the actual situation and the update cycle, determine the time period of the transaction amount, and correspondingly adjust the total industry transaction amount and transaction frequency of the time period; Step S204: Based on the nodes, association relationships, and the association weights of different association relationships, connect the nodes through the association relationships and mark the association weights to form a multi-dimensional graph.
[0008] In step S300, an incremental update mechanism is set for the constructed multi-dimensional graph to regularly update the nodes, association relationships, and the association weights of different association relationships, specifically as follows: Regularly update the nodes, association relationships, and the association weights of different association relationships; In the continuous update cycle, the exponential decay model is used to update the association weight in real time. The specific steps are as follows: Step S301: Query the current association weight w t−1 and the last update time t last ; Step S302: Real-time acquisition of the time interval Δt=t of the update cycle current −t last ; where t current Indicates the current time; Step S303: adopt an exponential decay model, superimpose the newly added weights, and apply the decay formula to update the weights in real time, specifically: ; Among them, w t represents the updated association weight; w t−1 represents the current association weight; e represents the natural constant; λ represents the attenuation coefficient; Δt represents the time interval of the update cycle; Δw new Indicates the newly added weight; Optionally, the newly added weight Δw new By the method in step S203, the updated parameters are calculated and obtained; Optionally, the newly added weight Δw new It can also be represented by a normalized value of the transaction amount within the time interval of the update cycle; the normalization method adopts the Max-Min method.
[0009] In step S400, the structural features of the multi-dimensional graph are extracted according to the regularly updated multi-dimensional graph, specifically: The structural features of the multi-dimensional graph include node type, association relationship type and node feature vector; The node types include enterprise, product and legal person; The types of relationships include supply chain relationships (supply, procurement), investment relationships (holding, equity participation) and technology dependency relationships (patent citations); The node feature vector includes: Enterprise node: registered capital, number of patents and revenue growth rate data.
[0010] Product nodes: technical complexity (classified by HS code) and annual production data.
[0011] Legal person node: data on shareholding ratio and number of affiliated companies.
[0012] Generate an enterprise portrait through the structural characteristics of the multi-dimensional graph to display explicit information such as the enterprise's legal person, products, relationships with other enterprises, and technology dependence; Obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait in the way of dynamic labels; The implicit features of the multi-dimensional graph include the paths between nodes and the association weights between nodes; The dynamic labels include the single-sided break risk label; According to the single-sided break risk label, the stability of the enterprise in later cooperation can be displayed in real time; The calculation steps of the single-sided break risk are specifically as follows: Step S401: Obtain the shortest paths of all node pairs and calculate the centrality of the nodes, specifically: C i =Σ s≠i≠t [(σ st_i ) / (σ st )]; Among them, C i represents the node centrality of node i; s, i, and t represent different node identifiers; σ st_i represents the total number of shortest paths from node s to node t passing through node i; σ st represents the total number of shortest paths from node s to node t; Step S402: Find all paths from node i to node j, obtain the paths that do not pass through edge e ij , and through the association weights between nodes in different paths, obtain the weights of each path, and take the ratio of the sum of the weights of the paths that do not pass through edge e ij to the sum of the weights of all paths as the edge redundancy estimation; Specifically characterized as: ; Among them, ρ ij represents the redundancy estimation of edge e ij ; p represents the path identifier; represents the path starting from node i and ending at node j and not passing through edge e ij ; represents all paths starting from node i and ending at node j; represents the node identifier in path p; w uv represents the association weight between node u and node v; Step S403: Calculate the single-sided break risk according to the centrality of the node and the edge redundancy estimation.
[0013] The calculation of the single-sided break risk is specifically characterized as: R ij =w ij (t)×(C i +C j )×(1 - ρij ); Among them, R ij represents the unilateral break risk of edge e ij ; w ij (t) represents the real-time association weight between node i and node j; C j represents the node centrality of node j; Visualize and display the generated enterprise portrait, specifically: Display the equity structure through a tree model; Display the business scope of the enterprise and search keywords through a word cloud; Update dynamic labels in real time.
[0014] An automated enterprise portrait generation system based on multi-source data fusion, the enterprise portrait automated generation system includes a data collection module, a data processing module, an entity relationship modeling module, a graph incremental update module, a feature extraction and portrait generation module, and a visualization display module; The data collection module is used to collect heterogeneous enterprise-related data and perform classified collection, including structured data, semi-structured data, and unstructured data; The data processing module is used to perform knowledge extraction and semantic alignment on the collected heterogeneous data, and extract data entity and relationship information; The entity relationship modeling module is used to determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and association weights of different association relationships; The graph incremental update module is used to periodically update nodes, association relationships, and the association weights of different association relationships; within consecutive update cycles, an exponential decay model is used to update the association weights in real time; The feature extraction and portrait generation module is used to extract the structural features of the multi-dimensional graph according to the periodically updated multi-dimensional graph, generate an enterprise portrait, obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait in the form of dynamic labels; The visualization display module is used to display the enterprise portrait and the implicit data calculated according to the implicit features of the multi-dimensional graph in different ways.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. Through the classified collection and hierarchical processing of structured, semi-structured, and unstructured data, combined with natural language processing technologies such as the BiLSTM-CRF model and dependency syntax analysis, this method realizes the accurate extraction and semantic relationship alignment of core data entities such as enterprises, legal persons, and products, effectively improving the comprehensiveness and accuracy of data fusion; 2. This method models entity relationships as nodes and edges, introduces various semantic relationships such as supply, holding, and technological dependence, and combines quantitative indicators such as transaction weights and investment ratios to form a multi-dimensional graph with rich semantics, clear structure, and scalability, providing comprehensive structural support for enterprise profiling. 3. This method uses an exponential decay model to update the association weights in the graph in real time, effectively solving the problems of static and poor timeliness of traditional enterprise profiles, and significantly enhancing the system's perception and response capabilities to enterprise dynamic changes. 4. This method characterizes explicit information through the structural features of the graph, while mining implicit structural features such as path centrality and edge redundancy, and further calculates dynamic labels to improve risk warning and intelligent analysis capabilities. 5. The present invention visually expresses the legal person structure, business scope, and dynamic labels through methods such as tree diagrams and word clouds, facilitating users to intuitively understand the enterprise operation characteristics and its position in the industrial chain from a multi-dimensional perspective, and enhancing the practical value of enterprise profiling and the business decision-making support capabilities. 6. The present invention realizes the unified expression of enterprise, event, and time-space information through means such as GeoHash geocoding and UTC timestamp standardization, providing basic support for cross-regional and cross-cycle data alignment and event reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flow diagram of the method for automatically generating an enterprise profile based on multi-source data fusion according to the present invention; Figure 2 is a schematic structural diagram of the system for automatically generating an enterprise profile based on multi-source data fusion according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a method for automatically generating an enterprise profile based on multi-source data fusion. The method for automatically generating an enterprise profile specifically includes the following steps: Step S100, heterogeneous data collection, and knowledge extraction and alignment of the collected heterogeneous data; Step S200: Identify the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and the association weights of different association relationships; Step S300: Set an incremental update mechanism for the constructed multi-dimensional graph, and regularly update the nodes, association relationships, and the association weights of different association relationships; Step S400: According to the regularly updated multi-dimensional graph, extract the structural features of the multi-dimensional graph to generate an enterprise portrait, obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait in the form of dynamic labels; Step S500: Visually display the generated enterprise portrait.
[0019] In step S100, heterogeneous data collection is performed, specifically: Classify and collect heterogeneous data, including structured data, semi-structured data, and unstructured data; Structured data includes industrial and commercial registration, tax records, patent databases, and supply chain system logs; Semi-structured data includes bidding documents and company annual reports; Unstructured data includes news sentiment and industry research reports.
[0020] Perform knowledge extraction and alignment on the collected heterogeneous data, specifically: Directly extract data entities from structured data; Extract data entities from semi-structured data through manual format parsing and keyword matching methods; For unstructured data, use the method based on the BiLSTM-CRF model to extract enterprise, product, and legal person data entities; Perform knowledge extraction based on the semantic relationship of dependency syntax; S1: Generate a dependency tree through the spaCy tool, extract the dependency path, and determine the main path; S2: Manually formulate dependency template matching rules in advance; including holding relationship dependency template matching rules and supply relationship dependency template matching rules; S3: Extract the verbs ("invest", "supply"), entity types, and path lengths in the dependency path; obtain the context bag-of-words features ("purchase", "shareholding"); based on sequence classification of BERT, the input is the dependency path context: S4: Use the path between the data entities involved in the main path as the knowledge extraction result; Unify the time and location in heterogeneous data from different sources into a standardized format.
[0021] The method for unifying time and location includes spatio-temporal correlation alignment, specifically: Enterprise-Geography Binding: Associate the enterprise to GeoHash through the entity data of the registered address. Event-Time Binding: Associate the data entity of the supply chain event to the standardized timestamp (UTC time). In step S200, determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and association weights of different association relationships, specifically: Step S201: Determine the data entities as nodes, where the nodes include enterprises, products, and legal persons. Step S202: Determine the knowledge extraction results as association relationships, where the association relationships include enterprise-production-product, enterprise-supply-enterprise, enterprise-holding-legal person, product-dependency-product, and geographical location-aggregation-enterprise. The representation of the nodes is specifically: Enterprise: Industrial and Commercial Registration Number, Industry Classification Code; Product: HS Code, Technical Complexity Level; Transaction: Transaction Amount, Frequency, Timestamp; Investment: Shareholding Ratio, Investment Amount; Legal Person: Depth of Shareholding Chain, Control Right Coefficient; Geographical Location: Latitude and Longitude, Economic Location Entropy; Step S203: Calculate the association weights of different association relationships respectively through the historical data fitting function relationship. The calculation of the association weight of the transaction relationship, the fitting function relationship is characterized as: ; where, w ij represents the association weight from node i to node j, and i and j represent different node identifiers; a ij represents the annual transaction amount from node i to node j; a total represents the total industry transaction amount; z represents the annual transaction frequency; Among them, the investment relationship represents the association weight through the shareholding ratio or investment amount; For the legal person relationship, product supply relationship, etc., only represent their association relationships for easy viewing, and there is no need to calculate the association weights; Among them, according to the actual situation and the update cycle, determine the time period of the transaction amount, and correspondingly adjust the total industry transaction amount and transaction frequency of the time period; Optionally, the calculation of the location entropy index is characterized as: LQ i = (Eir / E r ) / (E i / E); where, LQ i represents the location quotient index of node i; E ir represents the pedestrian flow in area r where node i is located; E r represents the total pedestrian flow in area r; E i represents the total number of employees in the country corresponding to the industry of node i; E represents the total number of employees in the country; Step S204: Based on the nodes, association relationships, and association weights of different association relationships, connect the nodes through the association relationships and mark the association weights to form a multi-dimensional graph.
[0022] In step S300, an incremental update mechanism is set for the constructed multi-dimensional graph to regularly update the nodes, association relationships, and association weights of different association relationships. Specifically: Regularly update the nodes, association relationships, and association weights of different association relationships; Within consecutive update cycles, use the exponential decay model to update the association weights in real time. The specific steps are as follows: Step S301: Query the current association weight w t−1 and the last update time t last ; Step S302: Obtain the time interval Δt = t current −t last in real time; where, t current represents the current time; Step S303: Use the exponential decay model, superimpose the new weight, and apply the decay formula to update the weight in real time. Specifically: ; where, w t represents the updated association weight; w t−1 represents the current association weight; e represents the natural constant; λ represents the decay coefficient; Δt represents the time interval of the update cycle; Δw new represents the new weight; The new weight Δw new is obtained by updating the parameter calculation through the method in step S203; The new weight Δw new can also be represented by the normalized value of the transaction amount within the time interval of the update cycle; The normalization method uses the Max - Min method.
[0023] In step S400, according to the regularly updated multi-dimensional graph, extract the structural features of the multi-dimensional graph. Specifically: The structural features of the multi-dimensional graph include node types, association relationship types, and node feature vectors; Node types include enterprise, product, and legal person; The types of related relationships include supply chain relationships (supply, procurement), investment relationships (holding, equity participation) and technology dependency relationships (patent citations); The node feature vector includes: Enterprise node: registered capital, number of patents and revenue growth rate data.
[0024] Product nodes: technical complexity (classified by HS code) and annual production data.
[0025] Legal person node: data on shareholding ratio and number of affiliated companies.
[0026] Generate enterprise portraits through the structural characteristics of multi-dimensional graphs, showing explicit information such as the enterprise's legal person, products, relationships with other enterprises, and technology dependence; Acquire the implicit features of the multi-dimensional graph in real time, calculate and set dynamic tags based on the implicit features, and supplement the implicit data of the enterprise portrait through dynamic tags; The structural features of the multidimensional graph include node type, relationship type, and node feature vector; Dynamic labels include single-side fracture risk labels; The unilateral fracture risk label can be used to show the stability of the enterprise in the later cooperation in real time; The specific steps for calculating the risk of unilateral fracture are as follows: Step S401: Obtain the shortest paths of all node pairs and calculate the centrality of the nodes, specifically: C i =Σ s≠i≠t [(σ st_i ) / (σ st )]; Among them, C i represents the node centrality of node i; s, i and t represent different node identifiers; σ st_i represents the total number of shortest paths from node s to node t through node i; σ st represents the total number of shortest paths from node s to node t; Step S402: Find all paths from node i to node j and obtain all paths that do not pass through edge e. ij The paths, and the weight of each path is obtained through the association weights between the nodes in different paths, and the paths that do not pass through the edge e ij The ratio of the sum of the weights of the paths to the sum of the weights of all paths is used as the edge redundancy estimate; The specific characteristics are: ; Among them, ρ ij Represents edge e ijRedundancy estimation; p represents the path identifier; represents a path starting from node i and ending at node j without passing through edge e ij ; represents all paths starting from node i and ending at node j; represents the node identifier in path p; w uv represents the association weight between node u and node v; Step S403: Calculate the unilateral break risk according to the centrality of the node and the edge redundancy estimation.
[0027] The calculation of the unilateral break risk is specifically characterized as: R ij = w ij (t) × (C i + C j ) × (1 - ρ ij ); where R ij represents the unilateral break risk of edge e ij ; w ij (t) represents the real-time association weight between node i and node j; C j represents the node centrality of node j; Visualize and display the generated enterprise portrait, specifically: Display the equity structure through a tree model; Display the business scope of the enterprise and the search keywords through a word cloud; Update the dynamic tags in real time.
[0028] As Figure 2 shown, an automated enterprise portrait generation system based on multi-source data fusion, the automated enterprise portrait generation system includes a data collection module, a data processing module, an entity relationship modeling module, a graph incremental update module, a feature extraction and portrait generation module, and a visualization display module; The data collection module is used to collect heterogeneous enterprise-related data and perform classified collection, including structured data, semi-structured data, and unstructured data; The data processing module is used to perform knowledge extraction and semantic alignment on the collected heterogeneous data, and extract data entity and relationship information; The entity relationship modeling module is used to determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and association weights of different association relationships; The graph incremental update module is used to periodically update the nodes, association relationships, and association weights of different association relationships; within consecutive update cycles, the association weights are updated in real time using an exponential decay model; The feature extraction and portrait generation module is used to extract the structural features of the multi-dimensional atlas according to the regularly updated multi-dimensional atlas, generate an enterprise portrait, obtain the implicit features of the multi-dimensional atlas in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait in the form of dynamic labels; The visualization display module is used to display the enterprise portrait and the implicit data calculated according to the implicit features of the multi-dimensional atlas in different ways.
[0029] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. An automated enterprise portrait generation method based on multi-source data fusion, characterized in that: The method for automatically generating an enterprise portrait specifically includes the following steps: Step S100: Heterogeneous data collection, and knowledge extraction and alignment are performed on the collected heterogeneous data; Step S200: Determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and the association weights of different association relationships; Step S300: Set an incremental update mechanism for the constructed multi-dimensional graph, and regularly update the nodes, association relationships, and the association weights of different association relationships; Step S400: According to the multi-dimensional graph updated regularly, extract the structural features of the multi-dimensional graph to generate an enterprise portrait, obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait through the dynamic labels; the implicit features of the multi-dimensional graph include the paths between nodes and the association weights between nodes in the multi-dimensional graph; The dynamic labels include one-sided break risk labels; The one-sided break risk is calculated based on the centrality of nodes and the edge redundancy estimation in the multi-dimensional graph; Step S500: Visually display the generated enterprise portrait.
2. The automated generation method of an enterprise portrait based on multi-source data fusion according to claim 1, wherein: In step S100, for heterogeneous data collection, specifically: Classify and collect heterogeneous data, including structured data, semi-structured data, and unstructured data; The structured data includes industrial and commercial registration, tax records, patent databases, and supply chain system logs; The semi-structured data includes bidding documents and company annual reports; The unstructured data includes news public opinion and industry research reports.
3. The automated generation method of an enterprise portrait based on multi-source data fusion according to claim 2, characterized in that: For the knowledge extraction and alignment of the collected heterogeneous data, specifically: For unstructured data, use the method based on the BiLSTM-CRF model to extract enterprise, product, and legal person data entities; Perform knowledge extraction based on the semantic relationship of dependency syntax; Unify the time and location in heterogeneous data from different sources into a standardized format.
4. The automated generation method of an enterprise portrait based on multi-source data fusion according to claim 3, wherein: In step S200, determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and the association weights of different association relationships, specifically: Step S201: Determine the data entities as nodes, and the nodes include enterprises, products, and legal persons; Step S202: Determine the knowledge extraction results as association relationships, and the association relationships include enterprise - production - product, enterprise - supply - enterprise, enterprise - holding - legal person, product - dependency - product, and geographical location - aggregation - enterprise; Step S203: Calculate the association weights of different association relationships respectively through the historical data fitting function relationship; For the calculation of the association weight of the transaction relationship, the fitting function relationship is characterized as: ; where, w ij represents the association weight from node i to node j, and i and j represent different node identifiers; a ij represents the annual transaction volume of node i with respect to node j; a total represents the total industry transaction volume; z represents the annual transaction frequency; Step S204: Based on the nodes, association relationships, and the association weights of different association relationships, connect the nodes through the association relationships and mark the association weights to form a multi-dimensional graph.
5. The automated generation method of an enterprise portrait based on multi-source data fusion according to claim 4, characterized in that: In step S300, an incremental update mechanism is set for the constructed multi-dimensional graph, and nodes, association relationships, and association weights of different association relationships are updated regularly. Specifically: Regularly update nodes, association relationships, and association weights of different association relationships; In consecutive update cycles, an exponential decay model is used to update the association weights in real time. The specific steps are as follows: Step S301, query the current association weight w t−1 and the last update time t last ; Step S302: Obtain the time interval Δt = t of the update period in real time current − t last ; where t current represents the current time Step S303: Use the exponential decay model, superimpose the newly added weights, and apply the decay formula to update the weights in real time. Specifically: ; Among them, w t represents the updated correlation weight; w t−1 represents the current correlation weight; e represents the natural constant; λ represents the decay coefficient; Δt represents the time interval of the update period; Δw new represents the newly added weight.
6. The automated generation method of an enterprise portrait based on multi-source data fusion according to claim 5, characterized in that: In step S400, according to the regularly updated multi-dimensional graph, extract the structural features of the multi-dimensional graph. Specifically: The structural features of the multi-dimensional graph include node types, association relationship types, and node feature vectors; The node types include enterprises, products, and legal persons; The association relationship types include supply chain relationships, investment relationships, and technology dependence relationships; The node feature vectors include: Enterprise nodes: registered capital, number of patents, and revenue growth rate data; Product nodes: technical complexity and annual output data; Legal person nodes: shareholding ratio and number of associated enterprise data.
7. A method for automatically generating an enterprise portrait based on multi-source data fusion according to claim 6, characterized in that: Obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait through the dynamic labels. Specifically: The calculation steps of the unilateral break risk are specifically as follows: Step S401: Obtain the shortest paths of all node pairs and calculate the centrality of the nodes. Specifically: C i =Σ s≠i≠t [(σ st_i ) / (σ st )]; Among them, C i represents the node centrality of node i; s, i, and t represent different node identifiers; σ st_i represents the total number of shortest paths from node s to node t passing through node i; σ st represents the total number of shortest paths from node s to node t; Step S402: Find all paths from node i to node j, obtain the paths that do not pass through edge e ij and, based on the association weights between nodes in different paths, obtain the weight of each path. Take the ratio of the sum of the weights of the paths that do not pass through edge e ij to the sum of the weights of all paths as the edge redundancy estimation; Step S403: Estimate and calculate the unilateral break risk based on the centrality of the nodes and the edge redundancy.
8. An enterprise portrait automatic generation system based on multi-source data fusion, which is applied to an enterprise portrait automatic generation method based on multi-source data fusion as described in any one of claims 1-7, and is characterized in that: The enterprise portrait automatic generation system includes a data collection module, a data processing module, an entity relationship modeling module, a graph incremental update module, a feature extraction and portrait generation module, and a visualization display module; The data collection module is used to collect heterogeneous enterprise-related data and perform classified collection, including structured data, semi-structured data, and unstructured data; The data processing module is used to perform knowledge extraction and semantic alignment on the collected heterogeneous data, and extract data entity and relationship information; The entity relationship modeling module is used to determine the data entities in the heterogeneous data as nodes, determine the knowledge extraction results as association relationships, calculate the association weights of different association relationships, and construct a multi-dimensional graph based on the nodes, association relationships, and association weights of different association relationships; The graph incremental update module is used to regularly update nodes, association relationships, and association weights of different association relationships; In consecutive update cycles, an exponential decay model is used to update the association weights in real time; The feature extraction and portrait generation module is used to extract the structural features of the multi-dimensional graph according to the regularly updated multi-dimensional graph, generate an enterprise portrait, obtain the implicit features of the multi-dimensional graph in real time, calculate and set dynamic labels through the implicit features, and supplement the implicit data of the enterprise portrait through the dynamic labels; The visualization display module is used to display the enterprise portrait and the implicit data calculated based on the implicit features of the multi-dimensional graph in different ways.
Citation Information
Patent Citations
Holographic city big data model and knowledge graph enterprise portrait construction method
CN112131275A
Credit risk prediction method and device, equipment, medium and program product
CN112927082A
Molecular interaction hypergraph modeling method and system based on redundancy prevention mechanism
CN117438001A
Enterprise portrait calculation method, system and terminal based on big data and multi-dimensional features
CN117575148A
Enterprise portrait construction method based on industrial cloud
CN118427363A
Cited By
Electric power material supply subject portrait generation method and device
CN120707330A
Enterprise relation graph construction system
CN120975214A
Enterprise portrait visualization system and enterprise portrait construction method
CN121235727A