Enterprise Relationship Mining Method, Device, Terminal Device and Storage Medium

By building an enterprise relationship network and using genetic algorithms and community mining models to identify implicit associations between enterprises, the implicit association problems that cannot be expressed in the existing technology are solved, and more comprehensive enterprise management and decision-making support are achieved.

CN118709908BActive Publication Date: 2025-07-22SHENZHEN INSAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410851782.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-07-22
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

The existing technology cannot effectively express the implicit relationship between enterprises, affecting the management of enterprise business decisions and fund scheduling.

Method used

By collecting enterprise-related data, building an enterprise relationship network, using genetic algorithms to find the optimal path, combining community mining models to identify the neighboring communities of the enterprise, and mining and analyze the implicit correlation relationship between enterprises.

Benefits of technology

Accurately identify implicit relationships between enterprises, provide more comprehensive management and decision-making support, and help enterprises understand market structure, assess risks and formulate strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118709908B_ABST
    Figure CN118709908B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of computer technology, and provides an enterprise relationship mining method, apparatus, terminal device, and storage medium. The method includes: collecting enterprise-related data of multiple enterprises, and constructing an enterprise relationship network according to the enterprise-related data of the multiple enterprises; obtaining an enterprise connection relationship set of the multiple enterprises according to an enterprise relationship path finding model based on a genetic algorithm and the enterprise relationship network; obtaining an enterprise adjacent community set of the multiple enterprises according to a community mining model and the enterprise relationship network; and mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set, and outputting the relationships. The embodiments of this application can accurately identify the implicit association relationships existing between enterprises, and provide strong support for the management and decision-making of enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a method, device, terminal device and storage medium for enterprise relationship mining. Background Art

[0002] In recent years, the development of economic globalization and the intensification of market competition have promoted new trends in enterprise development, which are specifically manifested as agglomeration, industrialization and networking. The connections and cooperation between enterprises have become increasingly frequent and complex. Therefore, accurately identifying the associated relationships between enterprises is very important for enterprise management and decision-making.

[0003] However, the current expression of the associated relationships between enterprises is relatively single, and most of them can only show some superficial relationships. For example, the relationships between the controlling shareholders, actual controllers, directors, supervisors, senior management personnel of a company and the enterprises directly or indirectly controlled by them, as well as other relationships that may lead to the transfer of company interests, etc. And these relationships are only the explicit part of enterprise relationships; there are also deeper implicit associated relationships between enterprises that are sufficient to affect aspects such as enterprise operation decisions, capital scheduling, and production and operation, but these implicit associated relationships cannot be expressed by existing technologies. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a method, device, terminal device and storage medium for enterprise relationship mining to solve the problem that the deeper implicit associated relationships between enterprises cannot be shown in the existing technology.

[0005] The first aspect of the embodiments of this application provides a method for enterprise relationship mining, and the method for enterprise relationship mining includes:

[0006] Collect enterprise-related data of multiple enterprises, and construct an enterprise relationship network according to the enterprise-related data of the multiple enterprises;

[0007] Obtain the enterprise connection relationship set of the multiple enterprises according to the enterprise relationship path finding model based on the genetic algorithm and the enterprise relationship network;

[0008] Obtain the enterprise adjacent community set of the multiple enterprises according to the community mining model and the enterprise relationship network;

[0009] Mine and analyze the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set, and output the results.

[0010] The second aspect of the embodiments of this application provides an enterprise relationship mining device, and the enterprise relationship mining device includes:

[0011] A network construction module that collects enterprise-related data of multiple enterprises and constructs an enterprise relationship network based on the enterprise-related data of the multiple enterprises;

[0012] An enterprise relationship acquisition module for obtaining an enterprise connection relationship set of the multiple enterprises according to an enterprise relationship path finding model based on a genetic algorithm and the enterprise relationship network;

[0013] An enterprise community acquisition module for obtaining an enterprise adjacent community set of the multiple enterprises according to a community mining model and the enterprise relationship network;

[0014] An analysis module for mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises based on the enterprise connection relationship set and the enterprise adjacent community set and outputting the results.

[0015] In a third aspect of the embodiments of the present application, a terminal device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the enterprise relationship mining method described in the first aspect of the embodiments of the present application are implemented.

[0016] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the enterprise relationship mining method described in the first aspect of the embodiments of the present application are implemented.

[0017] The enterprise relationship mining method provided in the first aspect of the embodiments of the present application constructs an enterprise relationship network by collecting enterprise-related data of multiple enterprises and based on the enterprise-related data of the multiple enterprises; obtains an enterprise connection relationship set of the multiple enterprises according to an enterprise relationship path finding model based on a genetic algorithm and the enterprise relationship network; obtains an enterprise adjacent community set of the multiple enterprises according to a community mining model and the enterprise relationship network; and finally mines and analyzes the relationships between a preset number of target enterprises among the multiple enterprises based on the enterprise connection relationship set and the enterprise adjacent community set and outputs the results. In this way, the implicit association relationships existing between enterprises can be accurately identified, providing strong support for enterprise management and decision-making.

[0018] It can be understood that the beneficial effects of the above second aspect to the fourth aspect can refer to the relevant descriptions in the above first aspect and will not be repeated here. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0020] Figure 1 is the first flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0021] Figure 2 is the second flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0022] Figure 3 is the third flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0023] Figure 4 is the fourth flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0024] Figure 5 is the fifth flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0025] Figure 6 is the sixth flowchart of the enterprise relationship mining method provided by the embodiments of the present application;

[0026] Figure 7 is the structural schematic diagram of the enterprise relationship mining device provided by the embodiments of the present application;

[0027] Figure 8 is the structural schematic diagram of the terminal device provided by the embodiments of the present application. Detailed implementation manners

[0028] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0029] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0030] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0031] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that in one or more embodiments of the present application, specific features, structures or characteristics described in connection with that embodiment are included. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all of the embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways. "A plurality" means "two" or "more than two".

[0032] In recent years, the development of economic globalization and the intensification of market competition have promoted new trends in enterprise development, which are specifically manifested as agglomeration, industrialization and networking. The connections and cooperation between enterprises have become increasingly frequent and complex. Therefore, accurately identifying the association relationships between enterprises is very important for enterprise management and decision-making.

[0033] However, at present, the expression of the association relationships between enterprises is relatively single, and most of them can only show some superficial relationships. For example, the relationships between the controlling shareholders, actual controllers, directors, supervisors, senior management personnel of a company and the enterprises directly or indirectly controlled by them, and other relationships that may lead to the transfer of the company's interests, etc. And these relationships are only the explicit part of the enterprise relationships; there are also deeper implicit association relationships between enterprises that are sufficient to affect aspects such as the business decisions, fund scheduling, and production and operation of enterprises, but these implicit association relationships cannot be expressed by the existing technologies.

[0034] Based on this, the embodiments of the present application provide an enterprise relationship mining method. By collecting enterprise-related data of multiple enterprises, an enterprise relationship network is constructed according to the enterprise-related data of the multiple enterprises; according to the enterprise relationship path finding model based on the genetic algorithm and the enterprise relationship network, an enterprise connection relationship set of the multiple enterprises is obtained; according to the community mining model and the enterprise relationship network, an enterprise adjacent community set of the multiple enterprises is obtained; and finally, according to the enterprise connection relationship set and the enterprise adjacent community set, the relationships between a preset number of target enterprises among the multiple enterprises are mined and analyzed and output. In this way, the implicit association relationships existing between enterprises can be accurately identified, providing strong support for enterprise management and decision-making.

[0035] Embodiment 1

[0036] As Figure 1As shown in the figure, the enterprise relationship mining method provided by the embodiments of the present application includes the following steps S1 to S4:

[0037] Step S1: Collect enterprise-related data of multiple enterprises, construct an enterprise relationship network based on the enterprise-related data of the multiple enterprises, and enter step S2.

[0038] In applications, the enterprise relationship network can reveal multi-level relationships such as cooperation, competition, supply chain, and ownership among enterprises, which is of great significance for understanding the market structure, assessing risks, predicting trends, and formulating strategies. In this embodiment, an enterprise relationship network is constructed by collecting enterprise-related data of multiple enterprises, and the specific process is as follows.

[0039] In one embodiment, as Figure 2 shown in the figure, step S1 includes the following steps S11 to S15:

[0040] Step S11: Collect enterprise-related data of multiple enterprises, and enter step S12.

[0041] In applications, when collecting enterprise-related data, collect as much enterprise-related data of some enterprises as possible, and the number of enterprises is not limited here.

[0042] In applications, enterprise-related data includes but is not limited to enterprise registration data, bidding data, shareholder data, product data, patent data, and other data.

[0043] In applications, enterprise-related data can be collected from multiple channels such as government public databases (such as the National Enterprise Credit Information Publicity System), industry databases, bidding platforms, and patent databases. Specifically, it can be the basic data of the market supervision and administration bureau and the open-source data of the Internet, which is not limited here.

[0044] In applications, there are various methods for obtaining enterprise-related data. Specifically, a suitable method can be selected according to the type of data and actual needs, which is not limited here.

[0045] Step S12: Construct a feature information database of the multiple enterprises according to the unified social credit code of each enterprise and the corresponding enterprise-related data, and enter step S13.

[0046] In applications, the unified social credit code (USCC) is a national unified identification code established by China for legal persons, unincorporated organizations, and individual industrial and commercial households. Each enterprise corresponds to a unique unified social credit code.

[0047] In an application, after collecting enterprise-related data, to ensure the quality and consistency of the data, the enterprise-related data can be cleaned to remove duplicate, missing, or incorrect data; then, using the unified social credit code, data with the same unified social credit code from different sources can be fused to form a characteristic information database containing multi-dimensional information such as registration information (including enterprise type, registration address, registered capital, legal representative, business scope, business period, etc.), shareholder information (including shareholder name, shareholding ratio, contribution method, etc.), product information, bidding information, patent information (number of patents), etc.

[0048] Step S13: Based on the characteristic information database, obtain the direct association relationships and indirect association relationships of the multiple enterprises, and proceed to step S14.

[0049] In an application, based on the characteristic information databases of multiple enterprises obtained in step S12, the direct association relationships and indirect association relationships of the multiple enterprises can be obtained, specifically including: extracting the direct association relationships between enterprises from the registration information, shareholder information, and bidding information in the characteristic information database, including but not limited to direct association relationships such as the same legal person relationship, investment relationship, procurement relationship, etc.; extracting the indirect association relationships generated due to certain similar characteristics between enterprises from the registration information, product information, and patent information in the characteristic information database, including but not limited to indirect association relationships such as spatial relationship, industry relationship, field relationship, etc.

[0050] Step S14: Use the entropy method to obtain the first weight of the direct association relationship and the second weight of the indirect association relationship, and proceed to step S15.

[0051] In an application, to avoid the influence of subjective judgment, the entropy method can be used to calculate the information entropy of the direct association relationship or indirect association relationship between different enterprises, and then determine the first weight of the direct association relationship according to the information entropy of the direct association relationship, and determine the second weight of the indirect association relationship according to the information entropy of the indirect association relationship, so as to provide more accurate weight information for subsequent enterprise relationship mining.

[0052] Specifically, the information entropy calculation method for the direct association relationship is as follows:

[0053] H k =-p k,i ln(p k,i );

[0054] where, H k represents the information entropy of the direct association relationship k, and p k,i represents the probability distribution of the number of enterprises under the direct association relationship k.

[0055] Specifically, the information entropy calculation method for the indirect association relationship is as follows:

[0056]

[0057] Among them, H m represents the information entropy of the indirect association relationship m, and p m,i represents the probability distribution of the number of enterprises under the indirect relationship m.

[0058] According to the above information entropy calculation method for the direct association relationship and the information entropy calculation method for the indirect association relationship, the information entropy of various association relationships can be obtained. Based on this information entropy, the weights between enterprises in each association relationship can be determined.

[0059] Step S15: Standardize the first weight and the second weight to obtain the enterprise relationship network.

[0060] In application, by standardizing the first weight and the second weight to make them within a certain range (such as (0, 1)), the fairness and consistency of weight distribution can be effectively ensured. The specific processing method can be as follows:

[0061]

[0062] Among them, represents the information entropy in the association relationship k and the relationship weight between enterprises i and j.

[0063] In application, when using the standardized first weight and second weight to construct the enterprise relationship network, each enterprise is used as a node in the enterprise relationship network. When there is a certain relationship between enterprises, an edge is established between the two enterprise nodes, and the weight of the edge is the standardized weight corresponding to the association relationship type (direct association relationship or indirect association relationship). Finally, the enterprise nodes and the weight information corresponding to the edges are input into a graph database or network analysis software to construct the enterprise relationship network.

[0064] In application, the graph database includes but is not limited to Neo4j, JanusGraph, ArangoDB; the network analysis software includes but is not limited to Gephi, NetworkX, UCINET; specifically, the appropriate network construction software can be selected according to data scale, data type, function, budget, etc., and it is not limited here.

[0065] The embodiments of this application adopt multi-dimensional enterprise data, which includes not only data such as legal persons, shareholders, and bidding that can directly reflect the association relationship between enterprises, but also data such as products, patents, and spatial locations that can indirectly reflect the association relationship between enterprises. The data dimension is rich, and the obtained analysis results of enterprise association relationships are more comprehensive and accurate.

[0066] Step S2: Obtain the enterprise connection relationship set of the multiple enterprises according to the enterprise relationship path finding model based on the genetic algorithm and the enterprise relationship network, and proceed to Step S3.

[0067] In applications, the enterprise relationship path finding model based on the genetic algorithm is a method that applies the genetic algorithm to solve the problem of finding the optimal path in the enterprise relationship network. In the enterprise relationship network, the path finding problem may include finding the shortest path between enterprises, the optimal cooperation path, the resource flow path, etc. As a heuristic global optimization algorithm, the genetic algorithm is very suitable for solving the above path finding problems.

[0068] In applications, the specific process of obtaining the enterprise connection relationship set (i.e., the path set of enterprise connection relationships) of multiple enterprises according to the enterprise relationship path finding model based on the genetic algorithm and the already constructed enterprise relationship network is described as follows.

[0069] In one embodiment, as Figure 3 shown, Step S2 includes the following Steps S21 to S25:

[0070] Step S21: Extract n connected subgraphs of the enterprise relationship network, and proceed to Step S22.

[0071] In applications, a connected subgraph refers to a subset in the above enterprise relationship network, which consists of a part of enterprise nodes and the edges between them, and any two enterprise nodes can be directly or indirectly connected through the edges in the graph. In short, a connected subgraph is an independent and indivisible part of the enterprise relationship network, where there are path connections between the enterprise nodes, and there are no direct edge connections with the nodes in other parts.

[0072] In applications, a graph database or network analysis software (such as Gephi, NetworkX) can be used to load the already constructed enterprise relationship network to ensure the integrity and accuracy of the network data; then, a connected component algorithm can be used to identify n connected subgraphs in the enterprise relationship network, where n is a positive integer; specifically, in NetworkX, the networkx.connected_components(G) function can be used to identify the connected subgraphs, where G is the network graph object.

[0073] In applications, after identifying the connected subgraphs, the functions of the graph database or network analysis software can be used to extract the identified connected subgraphs; specifically, in NetworkX, the enterprise node list of the connected subgraph can be traversed, and then the G.subgraph(nodes) function can be used to extract the connected subgraph, where nodes represents the enterprise node list of the connected subgraph.

[0074] Step S22: Use the random walk algorithm based on the Markov process to obtain multiple reachable paths between any two enterprise nodes in each connected subgraph, and then proceed to Step S23.

[0075] In application, when using the random walk algorithm based on the Markov process to obtain multiple reachable paths between any two enterprise nodes in each connected subgraph, each enterprise node in each connected subgraph is regarded as a state of the Markov chain, and a transition probability matrix is defined. Then, starting from any enterprise node, perform a random walk according to the transition probability matrix (each time when selecting the next enterprise node from the current enterprise node, follow the transition probability in the transition probability matrix). During the random walk process, record the sequence of enterprise nodes passed through, which forms a reachable path from the starting enterprise node to the ending enterprise node. To obtain multiple reachable paths, the above random walk process can be repeated, and finally multiple reachable paths between any two enterprise nodes can be obtained.

[0076] In application, in some cases, it may be necessary to adjust the transition probability. For example, by introducing the damping factor in the PageRank algorithm or recalculating the transition probability according to the edge weights, so as to ensure the fairness and efficiency of the random walk.

[0077] In application, when generating multiple reachable paths from the starting enterprise node to the ending enterprise node through multiple independent random walk experiments, a preset threshold or stop condition can also be set, and stop the random walk experiment when the stop condition is met or the number of generated reachable paths reaches the preset threshold.

[0078] In application, since the random walk may generate duplicate paths or long loops, it is necessary to screen the generated reachable paths, remove the duplicate and unreasonable paths, and only retain a preset number of reachable paths with higher quality.

[0079] Step S23: Aggregate all the reachable paths generated in the n connected subgraphs as the initial population, and use the weighted distance of each reachable path as the fitness of each path individual in the initial population, and then proceed to Step S24.

[0080] In application, the weighted distance can be based on the sum of the weights of the edges on the path. When using the weighted distance of each reachable path as the fitness of each path individual in the initial population, the weighted distance can be processed accordingly according to the actual problem to convert it into a fitness value.

[0081] Step S24: Run the genetic algorithm, and simulate the evolution process of the initial population through the genetic operations in the genetic algorithm, and then proceed to Step S25.

[0082] In an application, after constructing the initial population, genetic operations in the genetic algorithm can be used to iteratively optimize the initial population. After each iteration, the fitness of each path individual in the population is recalculated, and the next-generation population is selected based on the fitness. Continue this population evolution process until a preset stopping condition (such as the optimal fitness of the offspring has not changed in n consecutive generations of evolution) is reached, and then stop the process.

[0083] In an application, genetic operations include selection operations, crossover operations, and mutation operations. Among them, the selection operation tends to retain path individuals with high fitness; the crossover operation is used to combine the characteristics of two path individuals to generate new offspring path individuals; and the mutation operation is used to add new path individuals to the population.

[0084] Step S25: When the genetic algorithm runs to termination, obtain the target population, and select the top k path individuals with the highest fitness in the target population as the enterprise connection relationship set for output.

[0085] In an application, when the genetic algorithm runs to termination, the current population obtained is the target population. According to the fitness values of each path individual in the target population, perform sorting and select the top k path individuals with the highest fitness as the enterprise connection relationship set for output, where k is a positive integer. Each of these k path individuals not only contains the sequence of enterprise nodes on the path, but may also contain additional information such as the weighted distance of the path, the key enterprise nodes on the path, the path type (such as supply chain path, cooperation path, etc.), which is convenient for subsequent enterprise relationship analysis.

[0086] The embodiment of this application adopts an enterprise relationship path finding model based on the genetic algorithm, which is more efficient when solving the top k connection paths between any two enterprises in a large-scale enterprise relationship network.

[0087] Step S3: According to the community mining model and the enterprise relationship network, obtain the enterprise neighboring community set of the multiple enterprises, and enter step S4.

[0088] In an application, community mining, also known as community detection, is an important method for identifying groups of enterprise nodes with close internal connections in enterprise relationship network analysis. In an enterprise relationship network, performing community mining helps to understand the interaction patterns between enterprises and their surrounding communities, thereby better understanding the cooperation patterns and competitive relationships between enterprises.

[0089] In an application, an appropriate community mining model can be selected according to the characteristics of the enterprise relationship network and the mining purpose, and based on the selected community mining model and the enterprise relationship network, the enterprise neighboring community set is mined for subsequent data analysis.

[0090] In one embodiment, the community mining model includes a first community mining model constructed based on a single clustering algorithm and a second community mining model constructed based on multiple clustering algorithms.

[0091] In application, constructing the first community mining model based on a single clustering algorithm can be constructing the first community mining model based on the K-means algorithm, or can be constructing the first community mining model based on a density-based clustering algorithm (such as the DBSCAN algorithm (Density-Based Spatial Clustering of Applications with Noise)). Specifically, the first community mining model can also be constructed according to other clustering algorithms, which is not limited here.

[0092] In application, constructing the second community mining model based on multiple clustering algorithms can be obtaining the second community mining model by fusing multiple clustering algorithms such as the Louvain Algorithm, Label Propagation Algorithm (LPA), Walktrap Algorithm, Infomap Algorithm, etc. Specifically, the second community mining model can also be constructed by fusing more clustering algorithms, which is not limited here.

[0093] In one embodiment, as Figure 4 shown, step S3 includes the following steps S31 to S34:

[0094] Step S31: Obtain the enterprise feature vector of each enterprise among the multiple enterprises according to the feature information database of the multiple enterprises, and proceed to step S32.

[0095] In application, based on the feature information database of the multiple enterprises constructed in step S12, a multi-dimensional enterprise feature vector including registration information such as enterprise type, establishment duration, registered capital, registered place, shareholder structure, business status, and registered address, business operation information such as main business, operating income, profit margin, assets and liabilities, number of employees, and enterprise recruitment, scientific and technological innovation information such as patents, copyrights, papers, and trademarks, and administrative supervision information such as business anomalies, administrative penalties, enterprise changes, and serious violations is constructed for each enterprise.

[0096] Step S32: Input the enterprise feature vector of each enterprise into the first community mining model to mine the enterprise community, and obtain the first enterprise neighboring community, and proceed to step S33.

[0097] In an application, when inputting the enterprise feature vectors into the first community mining model to mine enterprise communities and obtaining the first enterprise neighboring communities, taking the first community mining model constructed based on the K-means algorithm as an example, randomly select K from the constructed enterprise feature vectors as the initial centroids, calculate the distances from each enterprise feature vector to the K centroids, and assign each enterprise to the cluster represented by the centroid with the closest distance; then recalculate the centroids of each cluster, where the centroid is the average value of all enterprise feature vectors within the cluster, repeat the above iterative process until the centroids no longer move significantly or reach the maximum number of iterations, and obtain the first enterprise neighboring communities.

[0098] In an application, the method for calculating the distances from each enterprise feature vector to the K centroids is as follows:

[0099]

[0100] Step S33: Input each connected subgraph of the enterprise relationship network into the second community mining model to mine enterprise communities and obtain the second enterprise neighboring communities, and proceed to step S34.

[0101] In an application, each connected subgraph can be input into the second community mining model obtained by fusing multiple clustering algorithms such as the Louvain Algorithm, LabelPropagation Algorithm (LPA), Walktrap Algorithm, and Infomap Algorithm to mine enterprise communities and obtain the second enterprise neighboring communities.

[0102] Step S34: Output the first enterprise neighboring communities and the second enterprise neighboring communities as the enterprise neighboring community set of the multiple enterprises.

[0103] In an application, outputting the first enterprise neighboring communities and the second enterprise neighboring communities obtained above as the enterprise neighboring community set helps to better understand the complex relationships between enterprises and provides data support for the strategic planning, market analysis, etc. of enterprises.

[0104] In one embodiment, step S33 includes:

[0105] Input each connected subgraph into the second community mining model, apply multiple clustering algorithms in the second community mining model for community classification, and obtain multiple community division results for each connected subgraph;

[0106] Evaluate the multiple community division results of each connected subgraph according to the improved information criterion, and determine one of the multiple community division results of each connected subgraph as the second enterprise neighboring community according to the evaluation results.

[0107] In applications, the Improved Information Criteria (IIC) are developed based on the classical Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) to meet the needs of more diverse and complex data analysis. The improved information criteria include the Corrected AIC information criterion (AICc), Hannan-Quinn information criterion (HQIC), Network Information Criterion information criterion (NIC), and Extended Bayesian Information Criterion information criterion (EBIC).

[0108] In applications, when obtaining the neighboring communities of the second enterprise, it includes inputting each connected subgraph into the second community mining model, applying various clustering algorithms in the second community mining model for community classification to obtain various community division results for each connected subgraph; then evaluating the various community division results of each connected subgraph through the constructed improved information criterion method. The specific evaluation method can be:

[0109] AIC = 2 * N - 2ln(Q c )

[0110]

[0111] where N represents the number of communities divided after classification by a clustering algorithm, Q c represents modularity, m represents the number of edges, A i,j represents the adjacency matrix, which is 1 if there is a direct edge between enterprise node i and enterprise node j, otherwise 0; k i represents the degree of enterprise node i, and k j represents the degree of enterprise node j.

[0112] Through the above evaluation method, the evaluation result corresponding to each clustering algorithm of each connected subgraph can be obtained, and one of the community division results is selected as the final community division result of the connected subgraph and used and output as the neighboring community of the second enterprise.

[0113] The embodiment of this application adopts a fusion community mining model based on multiple clustering algorithms, which can meet different types of community classification problems. At the same time, an evaluation index that comprehensively considers the accuracy and simplicity of classification is proposed to ensure that the community classification result is more scientific and has robustness in the algorithm.

[0114] Step S4: According to the enterprise connection relationship set and the enterprise neighboring community set, mine and analyze the relationships between a preset number of target enterprises among the multiple enterprises and output them.

[0115] In an application, based on the enterprise connection relationship set obtained in step S2 and the enterprise adjacent community set obtained in step S3, the relationships of multiple target enterprises can be mined and analyzed and output to the client. Specifically, the number of target enterprises is not limited. Here, the first target enterprise and the second target enterprise are taken as examples to illustrate this step.

[0116] In an application, the client includes but is not limited to devices such as mobile phones, tablet computers, wearable devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs).

[0117] In one embodiment, as Figure 5 shown, step S4 includes the following steps S41 to S43:

[0118] Step S41: Based on the enterprise connection relationship set and the enterprise adjacent community set, obtain the path connection relationship and community relationship between the preset number of target enterprises, and enter step S42.

[0119] In an application, based on the obtained enterprise connection relationship set and enterprise adjacent community set, the path connection relationship and community relationship between the preset number of target enterprises can be obtained. According to actual needs, a privacy relationship graph between the preset number of target enterprises can also be output for users to view. Specifically, taking the first target enterprise and the second target enterprise as examples, a privacy relationship graph between the first target enterprise and the second target enterprise can be output.

[0120] Step S42: According to the path connection relationship and the community relationship, obtain the average reach distance of each target enterprise among the preset number of target enterprises, and enter step S43.

[0121] In an application, the average reach distance can be the average of the reach distances from one target enterprise node to all other enterprise nodes. The smaller the average reach distance, the closer the connection between the target enterprise node and other enterprise nodes in the network, and the stronger its external influence ability, and information, resources, or influence can spread from this enterprise node to other enterprise nodes in the network faster.

[0122] Step S43: Evaluate and output the external influence and importance of each target enterprise according to the average reach distance of each target enterprise.

[0123] In an application, the smaller the average reach distance of a target enterprise, the stronger its external influence. Correspondingly, the larger the average reach distance of a target enterprise, the weaker its external influence.

[0124] In an application, when evaluating the importance of a target enterprise, it can be carried out through the following formula:

[0125]

[0126] Where N represents the number of enterprise nodes covered within the average reach distance, and k i represents the degree of enterprise node i, and d i,j represents the distance between enterprise node i and enterprise node j.

[0127] In an application, by identifying the importance of an enterprise, it helps the enterprise find more valuable partners and make more informed strategic decisions.

[0128] In one embodiment, as Figure 6 shown, after step S41, it further includes steps S44 to S45:

[0129] Step S44: According to the path connection relationship and the community relationship, obtain the average reach distance of each target enterprise among the preset number of target enterprises, and enter step S45.

[0130] In an application, the average reach distance can be, for a target enterprise node, the average value of the shortest path distances from other enterprise nodes in the network to this target enterprise node. The smaller the average reach distance, the more easily this target enterprise node is affected by other enterprise nodes in the network, and information, risks, or shocks can spread to this enterprise node faster from other parts of the network.

[0131] Step S45: Evaluate the degree of influence and risk degree of each target enterprise according to the average reach distance of each target enterprise and output.

[0132] In an application, the smaller the average reach distance of a target enterprise, the greater its degree of influence. Correspondingly, the larger the average reach distance of a target enterprise, the smaller its degree of influence.

[0133] In an application, when evaluating the risk degree of a target enterprise, it can be carried out through the following formula:

[0134]

[0135] Where M represents the number of enterprise nodes covered within the average reach distance, and k j represents the degree of enterprise node j, and d i,jDenotes the distance between enterprise node i and enterprise node j.

[0136] In applications, by identifying which enterprises or communities are the main sources of risks, it helps enterprises take targeted risk prevention and control measures.

[0137] Through the enterprise relationship mining method provided by the embodiments of the present application, the implicit relationships existing between any two enterprises can be mined, which can help the managers and industrial decision-makers of enterprises to more comprehensively understand the enterprise relationship network, accurately identify the association relationships between enterprises, and provide decision-making support for constructing a more comprehensive enterprise relationship management system.

[0138] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0139] Embodiment Two

[0140] The embodiments of the present application further provide an enterprise relationship mining device for executing the method steps in the method embodiments of the above enterprise relationship mining method. This device can be a virtual appliance in a terminal device, run by the processor of the terminal device, or the terminal device itself.

[0141] As Figure 7 shown, the enterprise relationship mining device 100 provided by the embodiments of the present application includes a network construction module 101, an enterprise relationship acquisition module 102, an enterprise community acquisition module 103, and an analysis module 104.

[0142] The network construction module 101 collects enterprise-related data of multiple enterprises and constructs an enterprise relationship network according to the enterprise-related data of the multiple enterprises;

[0143] The enterprise relationship acquisition module 102 is used to obtain the enterprise connection relationship set of the multiple enterprises according to the enterprise relationship path finding model based on the genetic algorithm and the enterprise relationship network;

[0144] The enterprise community acquisition module 103 is used to obtain the enterprise adjacent community set of the multiple enterprises according to the community mining model and the enterprise relationship network;

[0145] The analysis module 104 is used to mine and analyze the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set and output them.

[0146] In one embodiment, the network construction module 101 is specifically used for:

[0147] Collect enterprise-related data of multiple enterprises;

[0148] Construct a feature information database for the multiple enterprises according to the unified social credit code of each enterprise and the corresponding enterprise-related data;

[0149] Based on the feature information database, obtain the direct association relationships and indirect association relationships of the multiple enterprises;

[0150] Using the entropy value method, obtain the first weight of the direct association relationship and the second weight of the indirect association relationship;

[0151] Perform normalization processing on the first weight and the second weight to obtain the enterprise relationship network.

[0152] In one embodiment, the enterprise relationship acquisition module 102 is specifically configured to:

[0153] Extract n connected subgraphs of the enterprise relationship network, where n is a positive integer;

[0154] Using the random walk algorithm based on the Markov process, obtain multiple reachable paths between any two enterprise nodes in each connected subgraph;

[0155] Collect all the reachable paths generated in the n connected subgraphs as the initial population, and use the weighted distance of each reachable path as the fitness of each path individual in the initial population;

[0156] Run the genetic algorithm to simulate the evolution process of the initial population through the genetic operations in the genetic algorithm;

[0157] When the genetic algorithm runs to termination, obtain the target population, and select the top k path individuals with the highest fitness in the target population as the enterprise connection relationship set for output, where k is a positive integer.

[0158] In one embodiment, the community mining model includes a first community mining model constructed based on a single clustering algorithm and a second community mining model constructed based on multiple clustering algorithms; the enterprise community acquisition module 103 is specifically configured to:

[0159] According to the feature information database of the multiple enterprises, obtain the enterprise feature vector of each enterprise among the multiple enterprises;

[0160] Input the enterprise feature vector of each enterprise into the first community mining model for enterprise community mining, and obtain the first enterprise neighboring community;

[0161] Input each connected subgraph of the enterprise relationship network into the second community mining model for enterprise community mining, and obtain the second enterprise neighboring community;

[0162] Output the first enterprise adjacent community and the second enterprise adjacent community as the enterprise adjacent community set of the multiple enterprises.

[0163] In one embodiment, the enterprise community acquisition module 103 is further specifically configured to:

[0164] Input each connected subgraph into the second community mining model, apply various clustering algorithms in the second community mining model for community classification, and obtain various community division results of each connected subgraph;

[0165] Evaluate the various community division results of each connected subgraph according to the improved information criterion, and determine one community division result among the various community division results of each connected subgraph as the second enterprise adjacent community according to the evaluation result.

[0166] In one embodiment, the analysis module 104 is specifically configured to:

[0167] Based on the enterprise connection relationship set and the enterprise adjacent community set, obtain the path connection relationship and community relationship between the preset number of target enterprises;

[0168] According to the path connection relationship and the community relationship, obtain the average reach distance of each target enterprise among the preset number of target enterprises;

[0169] Evaluate and output the external influence and importance of each target enterprise according to the average reach distance of each target enterprise.

[0170] In one embodiment, the analysis module 104 is further specifically configured to:

[0171] According to the path connection relationship and the community relationship, obtain the average reachable distance of each target enterprise among the preset number of target enterprises;

[0172] Evaluate and output the degree of influence and risk degree of each target enterprise according to the average reachable distance of each target enterprise. In an application, each unit in the above device can be a software program module, can also be implemented by different logic circuits integrated in the processor or independent physical components connected to the processor, and can also be implemented by multiple distributed processors.

[0173] Embodiment III

[0174] As Figure 8 shown, the embodiment of the present application further provides a terminal device 200, including: at least one processor 201( Figure 8Only one processor is shown), a memory 202, and a computer program 203 stored in the memory 202 and executable on at least one processor 201. When the processor 201 executes the computer program 203, the steps in the above-mentioned method embodiments are implemented.

[0175] In an application, the terminal device may include, but is not limited to, a processor and a memory. Figure 8 These are merely examples of the terminal device and do not constitute a limitation on the terminal device. It may include more or fewer components than those shown in the figure, or combine certain components, or have different components. For example, a human-computer interaction device, input / output devices, network access devices, etc. The network access device may include a communication module for the terminal device to communicate with a user terminal.

[0176] In an application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. For example, the processor may be a timing controller (TCON). The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0177] In an application, in some embodiments, the memory may be an internal storage unit of the terminal device. For example, the hard disk or memory of the terminal device. In other embodiments, the memory may also be an external storage device of the terminal device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device. The memory may also include both the internal storage unit and the external storage device of the terminal device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0178] In an application, the communication module can be set as any device capable of directly or indirectly performing long-distance wired or wireless communication with a user terminal according to actual needs. For example, the communication module can provide communication solutions for applications on network devices, including Wireless Local Area Networks (WLANs) (such as Wi-Fi networks), Bluetooth, Zigbee, mobile communication networks, Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc. The communication module can include an antenna, and the antenna can have only one element or can be an antenna array including multiple elements. The communication module can receive electromagnetic waves through the antenna, perform frequency modulation and filtering processing on the electromagnetic wave signals, and send the processed signals to the processor. The communication module can also receive the signals to be sent from the processor, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna for radiation.

[0179] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / modules, since it is based on the same concept as the method embodiments of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0180] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. Each functional module in the embodiment can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the specific names of each functional module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details will not be elaborated here.

[0181] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0182] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in each of the above method embodiments.

[0183] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in each of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0184] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0185] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0186] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.

[0187] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0188] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An enterprise relationship mining method, characterized in that, The enterprise relationship mining method includes: Collecting enterprise-related data of multiple enterprises, and constructing an enterprise relationship network based on the enterprise-related data of the multiple enterprises; Obtaining an enterprise connection relationship set of the multiple enterprises according to an enterprise relationship path finding model based on a genetic algorithm and the enterprise relationship network; Obtaining an enterprise adjacent community set of the multiple enterprises according to a community mining model and the enterprise relationship network; Mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set, and outputting the results; The step of mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set, and outputting the results, includes: Based on the enterprise connection relationship set and the enterprise adjacent community set, obtaining the path connection relationship and the community relationship between the preset number of target enterprises; According to the path connection relationship and the community relationship, obtaining the average reach distance of each target enterprise among the preset number of target enterprises; the average reach distance refers to the average of the reach distances from one target enterprise node to all other enterprise nodes; Evaluating and outputting the external influence and importance of each target enterprise according to the average reach distance of each target enterprise.

2. The enterprise relationship mining method according to claim 1, characterized in that, The step of collecting enterprise-related data of multiple enterprises and constructing an enterprise relationship network based on the enterprise-related data of the multiple enterprises includes: Collecting enterprise-related data of multiple enterprises; Constructing a feature information database of the multiple enterprises according to the unified social credit code of each enterprise and the corresponding enterprise-related data; Based on the feature information database, obtaining the direct association relationship and the indirect association relationship of the multiple enterprises; Using the entropy method to obtain the first weight of the direct association relationship and the second weight of the indirect association relationship; Performing a normalization process on the first weight and the second weight to obtain the enterprise relationship network.

3. The enterprise relationship mining method according to claim 2, wherein The step of obtaining an enterprise connection relationship set of the multiple enterprises according to an enterprise relationship path finding model based on a genetic algorithm and the enterprise relationship network includes: Extracting n connected subgraphs of the enterprise relationship network, where n is a positive integer; Using a random walk algorithm based on a Markov process to obtain multiple reachable paths between any two enterprise nodes in each connected subgraph; Collecting all the reachable paths generated in the n connected subgraphs as an initial population, and taking the weighted distance of each reachable path as the fitness of each path individual in the initial population; Running a genetic algorithm to simulate the evolution process of the initial population through genetic operations in the genetic algorithm; When the genetic algorithm runs to termination, obtaining a target population, and selecting the top k path individuals with the highest fitness in the target population as the enterprise connection relationship set for output, where k is a positive integer.

4. The enterprise relationship mining method according to claim 3, characterized in that, The community mining model includes a first community mining model constructed based on a single clustering algorithm and a second community mining model constructed based on multiple clustering algorithms; The step of obtaining an enterprise adjacent community set of the multiple enterprises according to a community mining model and the enterprise relationship network includes: Obtain the enterprise feature vectors of each enterprise among the multiple enterprises according to the feature information database of the multiple enterprises; Input the enterprise feature vectors of each enterprise into the first community mining model for mining enterprise communities, and obtain the first enterprise adjacent communities; Input each connected subgraph of the enterprise relationship network into the second community mining model for mining enterprise communities, and obtain the second enterprise adjacent communities; Output the first enterprise adjacent communities and the second enterprise adjacent communities as the enterprise adjacent community set of the multiple enterprises.

5. The enterprise relationship mining method according to claim 4, characterized in that The step of inputting each connected subgraph of the enterprise relationship network into the second community mining model for mining enterprise communities and obtaining the second enterprise adjacent communities includes: Input each connected subgraph into the second community mining model, apply various clustering algorithms in the second community mining model for community classification, and obtain various community division results of each connected subgraph; Evaluate the various community division results of each connected subgraph according to the improved information criterion, and according to the evaluation results, determine one of the various community division results of each connected subgraph as the second enterprise adjacent community.

6. The enterprise relationship mining method according to claim 1, wherein After obtaining the path connection relationship and community relationship between the preset number of target enterprises based on the enterprise connection relationship set and the enterprise adjacent community set, it further includes: Obtain the average reach distance of each target enterprise among the preset number of target enterprises according to the path connection relationship and the community relationship; Evaluate and output the influence degree and risk degree of each target enterprise according to the average reach distance of each target enterprise.

7. An enterprise relationship mining device, characterized in that, The enterprise relationship mining device includes: A network construction module that collects enterprise-related data of multiple enterprises and constructs an enterprise relationship network according to the enterprise-related data of the multiple enterprises; An enterprise relationship acquisition module for obtaining the enterprise connection relationship set of the multiple enterprises according to the enterprise relationship path finding model based on the genetic algorithm and the enterprise relationship network; An enterprise community acquisition module for obtaining the enterprise adjacent community set of the multiple enterprises according to the community mining model and the enterprise relationship network; An analysis module for mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set and outputting; the step of mining and analyzing the relationships between a preset number of target enterprises among the multiple enterprises according to the enterprise connection relationship set and the enterprise adjacent community set and outputting includes: obtaining the path connection relationship and community relationship between the preset number of target enterprises based on the enterprise connection relationship set and the enterprise adjacent community set; obtaining the average reach distance of each target enterprise among the preset number of target enterprises according to the path connection relationship and the community relationship; the average reach distance refers to the average of the reach distances from a target enterprise node to all other enterprise nodes; evaluating and outputting the external influence and importance of each target enterprise according to the average reach distance of each target enterprise.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the enterprise relationship mining method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the enterprise relationship mining method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Optimal path planning algorithm for travel time reliability

    CN110633850A

  • Data recommendation method and device

    CN117648482A

  • Method and device for determining health degree of industrial chain

    CN117709803A