Network structure evolution method and system of emerging industry cluster
By refining data collection and classification methods for industrial clusters, the method improves the accuracy of network structure evolution prediction by integrating sub-networks into a hierarchical network, addressing the challenges of complexity and dynamism in industrial cluster predictions.
Patent Information
- Application Number
- CN202510597052.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-15
AI Technical Summary
It is difficult for the existing technology to accurately predict the evolution of network structure of emerging industry clusters, resulting in low accuracy of prediction results.
Collect information about industrial cluster enterprises, obtain effective data through cleaning processing, divide enterprise relationship types to build sub-networks, determine sub-network association data, connect to form target hierarchical networks, and use pre-trained network structure evolution prediction model to make predictions.
Through refined relationship division and network construction, precise mathematical modeling of enterprise diversified relationships is achieved, and the accuracy and comprehensiveness of network structure evolution prediction are improved.
Smart Images

Figure CN120321133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of industrial cluster structures, and specifically relates to a method and system for the network structure evolution of emerging industrial clusters. Background Art
[0002] With the rapid development of globalization and informatization, emerging industrial clusters have become an important force in promoting regional economic growth and technological innovation. However, due to the complexity and dynamics of industrial clusters, the evolution of their network structures is often difficult to predict, posing great challenges to policy-making and corporate strategies.
[0003] In the prior art, the evolution trend of the network structure is usually predicted by collecting and analyzing the key data of industrial clusters and selecting a model based on the key data.
[0004] However, with the development of industrial clusters, the data obtained becomes increasingly complex and diverse, resulting in incomplete key data analyzed, thus leading to low accuracy of the prediction results of the network structure evolution. Summary of the Invention
[0005] In order to improve the prediction accuracy of the network structure evolution results, this application provides a method, a system, an electronic device, and a storage medium for the network structure evolution of emerging industrial clusters.
[0006] The first aspect of this application provides a method for the network structure evolution of emerging industrial clusters, specifically including: Collect the enterprise information of each enterprise in the industrial cluster, where the industrial cluster is a collection of multiple enterprises in the same field; Clean and process the enterprise information to obtain valid data; According to the valid data, divide the relationships between the enterprises into different types, and construct corresponding sub-networks for each enterprise according to the different types; Determine the associated data between the sub-networks, and connect the sub-networks according to the associated data to obtain a target hierarchical network; Extract the key data in the target hierarchical network, and based on the key data, perform prediction through a pre-trained network structure evolution prediction model to obtain the network structure evolution result of the industrial cluster.
[0007] By adopting the above technical solution, after collecting the enterprise information of each enterprise in the industrial cluster, cleaning and processing the enterprise information to obtain valid data, then according to the valid data, classifying the relationships between enterprises into different types, constructing corresponding sub-networks according to different types, determining the associated data between sub-networks, and connecting the sub-networks according to the associated data to obtain the target hierarchical network, extracting key data from the target hierarchical network, and substituting the key data into the pre-trained network structure evolution prediction model to complete the prediction of network structure evolution. Through refined relationship classification and network construction, accurate mathematical modeling of enterprise multiple relationships is realized. The hierarchical fusion design integrates multi-dimensional sub-networks while retaining the hierarchical topological features of the industrial cluster, establishes a cross-level feature recognition system for the subsequent extraction of key data, improves the comprehensiveness of key data, and thus improves the accuracy of network structure evolution prediction.
[0008] Optionally, the collection of the enterprise information of each enterprise in the industrial cluster specifically includes: Obtaining the basic registration information and geographical location data of each enterprise in the industrial cluster; Invoking the government affairs database interface to obtain the historical change records and policy subsidy data of the enterprise; Performing entity recognition on the enterprise official website text and extracting unstructured data; Taking the basic registration information, the geographical location data, the historical change records, the policy subsidy data, and the unstructured data as the enterprise information.
[0009] By adopting the above technical solution, by integrating the basic registration information and the authoritative data of the government affairs database, a data chain for the entire life cycle of the enterprise is constructed. The entity recognition technology is used to extract unstructured data from the enterprise official website text, effectively supplementing the key dimensions not covered in the administrative database, thereby improving the comprehensiveness of enterprise information.
[0010] Optionally, the cleaning and processing of the enterprise information to obtain valid data specifically includes: Based on a preset rule library, verifying and correcting the error data to obtain corrected data; Performing duplicate removal processing on the corrected data to obtain deduplicated data; Using a generative adversarial network to fill in the missing values in the deduplicated data to obtain valid data.
[0011] By adopting the above technical solution, the automated verification mechanism based on the rule library can intelligently identify and correct logical errors and format exceptions in the data, ensuring data quality. Through multi-level duplicate removal processing, redundant information is effectively eliminated, ensuring data uniqueness. Using a generative adversarial network for missing value filling provides a reliable data basis for subsequent network construction.
[0012] Optionally, based on the valid data, the relationships between the enterprises are divided into different types, and corresponding sub-networks are constructed for each of the enterprises according to the different types, specifically including: Obtain multi-dimensional features corresponding to each enterprise from the valid data; Calculate the feature similarity between the features of each dimension; Based on the feature similarity, divide the relationships between the enterprises into different types; Take the divided types as edges and the enterprises as nodes, and combine the valid data to construct sub-networks corresponding to each of the enterprises.
[0013] By adopting the above technical solution, based on the in-depth mining of enterprise multi-dimensional features, potential associations between enterprises can be comprehensively captured. Through the intelligent calculation of feature similarity, the deviation of subjective classification is overcome. Based on the dynamic classification mechanism of similarity, automatic identification and classification of enterprise relationships are realized, making the network structure modeling more refined.
[0014] Optionally, determine the association data between the sub-networks, and connect the sub-networks according to the association data to obtain a target hierarchical network, specifically including: Analyze the association data between the sub-networks based on the association rule mining technology; Use the association data to establish a connection rule library; Adopt a network fusion algorithm and obtain a target hierarchical network based on the connection rule library and the sub-networks.
[0015] By adopting the above technical solution, through the intelligent network fusion technology, the organic integration of multi-dimensional relationships of industrial clusters is realized, the seamless connection of multi-level network structures is realized, and the multi-dimensional interaction characteristics of the cluster structure are accurately presented.
[0016] Optionally, extract the key data in the target hierarchical network, specifically including: Obtain the node data of each node and the edge data of each edge from the target hierarchical network; Analyze the node data to obtain the time-series dynamic fluctuation characteristics, where the time-series dynamic fluctuation characteristics are the characteristics of the node data fluctuating over time; In combination with a preset industry relationship database, analyze the edge data to obtain associated graph data; Based on the intelligent fusion center, fuse the time-series dynamic fluctuation characteristics and the associated graph data to generate the key data.
[0017] By adopting the above technical solutions, through the extraction of temporal dynamic fluctuation features, the evolution law of node attributes can be accurately captured. Based on the industry relationship database, the semantic understanding of network connection relationships is realized, and association graph data is obtained. Based on the multi-modal data integration mechanism of the intelligent fusion center, the temporal dynamic fluctuation features and the association graph data are fused to obtain key data. In the fusion process, not only the information in the time and space dimensions is integrated, but also the dynamic changes of nodes and the complex associations between nodes are comprehensively considered, thereby improving the comprehensiveness of the key data.
[0018] Optionally, combining with the preset industry relationship database, parsing the edge data to obtain association graph data specifically includes: Extract explicit features and implicit features from the edge data. The explicit features include the direct cooperation relationship between enterprises, and the implicit features include the potential competition relationship between indirectly related enterprises; Convert the explicit features into explicit labels, which are used to describe the direct business connections between enterprises; Convert the implicit features into implicit labels through the industry knowledge base and association rule matching. The industry knowledge base is constructed based on the industry white papers in the field to which the industrial cluster belongs, and the association rules are obtained through the valid data. The implicit labels are used to describe the indirect competition or complementary relationship between enterprises; According to the weights corresponding to the explicit labels and the implicit labels respectively, combine the explicit labels and the implicit labels to generate initial text labels; Perform semantic verification on the initial text labels to obtain text labels; Bind the nodes in the target hierarchical network to the text labels to obtain association graph data.
[0019] By adopting the above technical solutions, based on the dual extraction mechanism of explicit and implicit features, the comprehensive recognition of multi-dimensional relationships between enterprises is realized. Through the intelligent conversion method combining the industry knowledge base and association rules, the abstract implicit features are converted into interpretable semantic labels, significantly improving the comprehensibility of network relationships. Adopting the technical route of dynamic weight fusion and semantic verification ensures the accuracy and consistency of the generated labels. Through the intelligent binding of labels and network nodes, an association knowledge graph with both structural features and semantic information is constructed.
[0020] In the second aspect of the present application, a network structure evolution system for emerging industrial clusters is provided, specifically including: A data collection module, used to collect enterprise information of each enterprise in the industrial cluster, where the industrial cluster is a collection of multiple enterprises in the same field; A data processing module, used to clean and process the enterprise information to obtain valid data; A relationship division and network construction module, configured to divide the relationships between the enterprises into different types according to the valid data, and construct corresponding sub-networks for the enterprises according to the different types; A hierarchical network fusion module, configured to determine the associated data between the sub-networks, and connect the sub-networks according to the associated data to obtain a target hierarchical network; A network evolution prediction module, configured to extract key data from the target hierarchical network, and perform prediction through a pre-trained network structure evolution prediction model according to the key data to obtain the network structure evolution result of the industrial cluster.
[0021] By adopting the above technical solution, after the electronic device collects the enterprise information of each enterprise in the industrial cluster, the enterprise information is cleaned to obtain valid data. Then, according to the valid data, the relationships between the enterprises are divided into different types, and corresponding sub-networks are constructed according to the different types. Next, the associated data between the sub-networks is determined, and the sub-networks are connected according to the associated data to obtain a target hierarchical network. Key data is extracted from the target hierarchical network, and the key data is substituted into the pre-trained network structure evolution prediction model to complete the prediction of the network structure evolution. Through refined relationship division and network construction, accurate mathematical modeling of the multiple relationships of enterprises is realized. The hierarchical fusion design integrates multi-dimensional sub-networks while retaining the hierarchical topological characteristics of the industrial cluster, establishes a cross-level feature recognition system for the extraction of subsequent key data, improves the comprehensiveness of the key data, and thus improves the accuracy of the prediction of the network structure evolution.
[0022] Through the refined design of relationship division and network construction, accurate representation of different types of enterprise relationships is ensured. The design of hierarchical network fusion realizes the organic integration of multi-dimensional sub-networks, completely retains the hierarchical characteristics of the industrial cluster. The prediction of the network structure evolution combined with the pre-trained model enables the system to have the ability to deeply mine the dynamic evolution laws of complex networks, thereby improving the accuracy rate of the network structure evolution result.
[0023] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.
[0024] In the fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed. Description of the Drawings
[0025] Figure 1It is a schematic structural diagram of a method for the evolution of the network structure of an emerging industrial cluster disclosed in an embodiment of the present application; Figure 2 It is a schematic flowchart of a method for the evolution of the network structure of an emerging industrial cluster disclosed in an embodiment of the present application; Figure 3 It is Figure 2 A sub-step flowchart of step S201; Figure 4 It is Figure 2 A sub-step flowchart of step S202; Figure 5 It is Figure 2 A sub-step flowchart of step S203; Figure 6 It is Figure 2 A sub-step flowchart of step S204; Figure 7 It is a flowchart of the steps for extracting key data in the target hierarchical network; Figure 8 It is Figure 7 A sub-step flowchart of step S703; Figure 9 It is a schematic diagram of the modules of a system for the evolution of the network structure of an emerging industrial cluster disclosed in an embodiment of the present application; Figure 10 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application.
[0026] Explanation of reference numerals: 10, system for the evolution of the network structure of an emerging industrial cluster; 11, data acquisition module; 12, data processing module; 13, relationship division and network construction module; 14, hierarchical network fusion module; 15, network evolution prediction module; 901, processor; 902, communication bus; 903, user interface; 904, network interface; 905, memory. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0028] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.
[0029] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0030] Figure 1 Exemplary system architecture 100 showing an embodiment of a method for evolving the network structure of an emerging industrial cluster or a system for evolving the network structure of an emerging industrial cluster to which the present application can be applied.
[0031] As Figure 1 shown, the system architecture 10 may include a terminal device 11, a network 12, and a server 13. The network 12 is used to provide a medium for a communication link between the terminal device 11 and the server 12. The network 12 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0032] A user can use the terminal device 11 to interact with the server 13 through the network 12 to receive or send data, etc.
[0033] The terminal device 11 can be hardware or software. When the terminal device 11 is hardware, it can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc. When the terminal device 11 is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0034] The server is a background server for processing the data displayed on the terminal device 11. The background server can analyze and process the received data, and can feedback the processing result (such as the recognition result) to the terminal device.
[0035] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above is only illustrative. According to the implementation requirements, there may be any number of terminal devices, networks and servers. In particular, when the target data does not need to be acquired remotely, the above system architecture may not include a network, but only include terminal devices or servers.
[0037] This embodiment discloses a method for evolving the network structure of an emerging industry cluster. Figure 2 is a flow chart of a method for evolving a network structure of an emerging industry cluster disclosed in an embodiment of the present application, such as Figure 2 As shown, including step S201 to step S204, the above steps are as follows: S201: Collect enterprise information of each enterprise in an industrial cluster, where an industrial cluster is a collection of multiple enterprises in the same field.
[0038] Among them, the industrial cluster in the embodiment of the present application is a collection of multiple enterprises in the same field, such as a pharmaceutical industry cluster, a new energy vehicle industry cluster, an electronic information industry cluster and other different types of industrial clusters.
[0039] Enterprise information is the original data set obtained after collecting relevant data of industrial clusters, including structured data that conforms to the format specifications and unstructured data that does not conform to the format specifications.
[0040] Specifically, in the embodiment of the present application, structured and unstructured data such as industrial and commercial registration information, supply chain disclosure data, intellectual property related records, etc. are obtained through a commercial data platform and integrated into enterprise information.
[0041] For example, when collecting enterprise information of the new energy vehicle industry cluster, the data collection module is connected to the commercial data platform API to obtain data such as the company's registered capital, shareholder structure, etc., while crawling data such as intellectual property related records in the Chinese patent database, and integrating the above-obtained data into enterprise information.
[0042] Reference Figure 3 , Figure 3 The embodiment of this application provides Figure 2 A schematic flow chart of a sub-step of step S201 in the above step is as follows: S301: Obtain the basic registration information and geographical location data of each enterprise in the industrial cluster.
[0043] Specifically, in the embodiments of this application, through the API interface of the business data platform, batch obtain the basic information of the enterprises in the industrial cluster, including fields such as enterprise name, unified social credit code, establishment date, business scope, etc. At the same time, call the reverse geocoding API of Amap or Baidu Map to convert the enterprise registration address text into longitude and latitude coordinates.
[0044] For example, when calling the "List of Tesla Suppliers in Lingang, Shanghai" interface, directly obtain structured fields such as "Enterprise Name: Shanghai Lingang Yapu Auto Parts Co., Ltd.", "Unified Social Credit Code: 91310115MA1HXXXXXX", "Establishment Date: 2023-05-18", "Business Scope: Manufacturing of new energy vehicle battery box assemblies". At the same time, trigger the reverse geocoding service of Amap to automatically convert the registration address "No. 2855, Canghai Road, Pudong New Area, Shanghai" into longitude and latitude coordinates (121.9012°E, 30.9123°N).
[0045] S302: Call the government affairs database interface to obtain the historical change records and policy subsidy data of the enterprise.
[0046] Specifically, in the embodiments of this application, based on protocol authentication, access the government affairs data interfaces such as the National Enterprise Credit Information Publicity System and the Industrial Policy Department of the Ministry of Industry and Information Technology to dynamically obtain the equity structure, historical change records and policy subsidy details of the enterprise. Government affairs data is usually returned in nested JSON format, and keyword fields (such as shareholder shareholding paths, subsidy application amounts) need to be extracted through JSONPath syntax.
[0047] For example, based on the OAuth2.0 protocol, access the API of the National Enterprise Credit Information Publicity System, send an HTTPS request carrying the enterprise's unified social credit code "91310115MA1HXXXXXX" to the standard interface of the National Enterprise Credit Information Publicity System, and the obtained nested JSON data contains multi-layer equity structure information. Screen important controlling shareholders with a shareholding ratio exceeding 30% through JSONPath expressions, such as the controlling party Yapu Group (with a 51% shareholding) and Guotou Chuangxin (with a 34% shareholding) and other key shareholders. When synchronously calling the subsidy interface of the Industrial Policy Department of the Ministry of Industry and Information Technology, for the "Research and Development of High-Energy Density Power Battery Boxes" subsidy project obtained by this enterprise in 2024, accurately extract the subsidy amount of 12 million yuan.
[0048] S303: Perform entity recognition on the enterprise official website text and extract unstructured data.
[0049] Specifically, in this embodiment, for text data such as corporate official websites and press releases, natural language processing technology is used to achieve entity recognition, and further dependency syntax analysis technology is utilized to extract cooperation relationships from unstructured text.
[0050] For example, taking the press release issued on the official website of Shanghai Lingang Yapu Auto Parts Co., Ltd. as an example, the original text description is: "Reached a strategic cooperation with CATL to jointly develop a new generation of CTP battery box integration technology, and at the same time established a lightweight material laboratory in conjunction with Shanghai Jiao Tong University." Entity recognition is performed on the text in combination with the context, and key entities such as CATL (cooperating party), Shanghai Jiao Tong University (cooperating institution), CTP battery box integration technology (technical field), and lightweight material laboratory (cooperation carrier) are marked. Subsequently, the sentence structure is parsed through dependency syntax analysis technology to identify the core of the verb-object relationship of "reached → strategic cooperation" and the action-object association of "developed → technology", and finally structured cooperation relationship data is generated: Yapu Company and CATL form a R & D cooperation in CTP battery technology and jointly build a laboratory with Shanghai Jiao Tong University to conduct material research.
[0051] S304: Use the basic registration information, geographical location data, historical change records, policy subsidy data, and unstructured data as enterprise information.
[0052] Specifically, in this embodiment, entity alignment is performed on the basic registration information, geographical location data, historical change records, policy subsidy data, and unstructured data. For example, taking the unified social credit code as the same entity, the obtained basic registration information, geographical location data, historical change records, policy subsidy data, and unstructured data are associated with the same entity. The executions of the above S301 to S303 are not in sequence.
[0053] S202: Clean the enterprise information to obtain valid data.
[0054] In the embodiment of this application, the valid data is the enterprise information after eliminating the ambiguity of enterprise entities.
[0055] The cleaning process includes, but is not limited to, cleaning of incorrect data, deduplication of duplicate data, and filling of missing data in the enterprise information. In the embodiment of this application, it is mainly to unify the ambiguous enterprise information, and the ambiguity caused by enterprise abbreviations and aliases in the enterprise information can be eliminated through a preset entity parsing rule library according to the matching rules in the rule library to obtain valid data.
[0056] For example, Shanghai Lingang Yapu Auto Parts Co., Ltd. may appear under aliases such as Yapu Lingang Company and Lingang Yapu New Energy in different data sources. The core word extraction rules in the preset rule library (e.g., "Shanghai Lingang Yapu Auto Parts Co., Ltd." can be simplified to "Yapu Lingang"), the regional abbreviation rules ("Shanghai" is converted to "Hu"), and the industrial and commercial registration name parsing rules (extracting the territorial information in parentheses) are used for cleaning. When variants such as "Hu Yapu New Energy" are detected, core word matching is automatically triggered, and triple verification is carried out in combination with the unified social credit code primary key (91310115MA1HXXXXXX) and the geographical coding of the registered address (e.g., 2855 Canghai Road corresponds to the longitude and latitude 121.9012°E, 30.9123°N). If the verification passes, it is the same entity, eliminating the incorrect information caused by entity ambiguity in the enterprise information and obtaining valid data.
[0057] Refer to Figure 4 , Figure 4 is a schematic diagram of a sub-step process of step S202 provided in an embodiment of the present application. The above steps are as follows: Figure 2 In the above steps: S401: Based on the preset rule library, verify and correct the error data to obtain corrected data.
[0058] Among them, the preset data rule library includes a data source credibility weight table and a verification logic rule. The verification logic rule includes that the enterprise name needs to contain the core word, the change of the registered address needs to meet a reasonable range, and the business scope needs to match the patented technology. The error data is the ambiguous data generated under the same entity in the enterprise information. The corrected data is the enterprise information after eliminating the ambiguous data generated under the same entity.
[0059] Specifically, in the embodiment of the present application, when verifying the enterprise information based on the preset rule library, when ambiguous data appears under the same entity in the enterprise information, based on the data source credibility, the high-credibility data source is used to correct the ambiguous data to obtain corrected data. For example, based on the preset rule library, it is detected that the address data under a certain entity is ambiguous, 2855 Canghai Road (source is the government affairs database) and 123 Feizhou Road (from the enterprise official website). Based on the higher credibility of the government affairs database source, the address data is changed to 2855 Canghai Road.
[0060] S402: Remove duplicates from the corrected data to obtain deduplicated data.
[0061] Among them, the deduplicated data is to eliminate the redundant data in the corrected data. The redundant data includes data aggregation redundancy, business process split redundancy, attribute-level data redundancy, etc. The business rules include rules such as contract number consistency, amount logic verification, time continuity, and technical parameter matching.
[0062] In the embodiment of the present application, the elimination of business process splitting redundancy is mainly taken as an example. Specifically, the splitting keywords (such as the contract number suffix A / B) in the business text in the correction data are parsed, the splitting mode is identified, and the corresponding correction is performed based on the business rules to obtain the deduplicated data.
[0063] For example, by parsing the business text in the revised data, we get contracts CX-2025-033A (amount 120 million) and CX-2025-033B (amount 30 million). Based on business rules, they meet the same project number, the amount logic check, the timestamp must be within the total contract period, and the technical indicators are consistent. Therefore, they are merged into a complete contract CX-2025-033 (amount 150 million).
[0064] S403: Use a generative adversarial network to fill in the missing values in the deduplicated data to obtain valid data.
[0065] Among them, the valid data is the deduplicated data with missing values filled in.
[0066] Specifically, in an embodiment of the present application, missing value filling relies on a conditional generative adversarial network (CGAN) to achieve intelligent completion.
[0067] For example, when Shanghai Lingang Yapu Auto Parts Co., Ltd.'s R&D investment data for Q1 2025 was missing, it relied on the CGAN model for intelligent completion: first, it extracted related features of the same period, including 28 patent authorization data (real-time capture from the API of the State Intellectual Property Office), 120 million yuan of equipment purchases (analysis of the company's ERP system order records) and 3 industry-university-research cooperation projects (through the official website news text to identify the cooperation information of "co-building a solid-state battery laboratory with Shanghai Jiaotong University"). These features constitute the conditional vector input generator, combined with the random noise vector to generate the candidate value interval [130 million, 170 million]. The discriminator is based on historical data distribution (the average R&D investment in Q1 from 2018 to 2024 is 128 million ± 15%) and industry rules (such as "R&D investment / number of patents ≥ 5 million yuan / item") for adversarial training, and finally outputs a filling value of 150 million yuan.
[0068] S203: According to the valid data, the relationships between the enterprises are divided into different types, and corresponding sub-networks are constructed for each enterprise according to the different types.
[0069] Specifically, in the embodiments of the present application, according to the relationships between enterprises in the effective data, three basic types are defined. The material and service flow relationships are classified into supply chain relationship types, the equity investment and merger and acquisition relationships are classified into capital relationship types, and the joint R & D and patent cross-licensing are classified into technical relationship types. Each relationship type independently constructs a sub-network, which includes nodes and edges. The attributes of the nodes contain eigenvectors such as enterprise scale and innovation ability, and the weights of the edges are dynamically adjusted according to the cooperation intensity.
[0070] For example, the multi-dimensional relationship network of Shanghai Lingang Yapu Auto Parts Co., Ltd. is accurately characterized through dynamic data fusion and business rule collaboration. Taking the supply chain relationship sub-network as an example, first, extract the business data of 2025 from the enterprise ERP system: the annual purchase amount of battery cells from CATL reaches 1.2 billion yuan, and the average monthly delivery of battery box modules is 50,000 sets. At the same time, capture the data of the logistics platform and analyze features such as the proportion of suppliers in the Yangtze River Delta region being 83% and the average transportation time being 12 hours. Based on these data, attributes such as production capacity (80 GWh / year), on-time delivery rate (98%), and supply chain localization index (0.87) are assigned to the Yapu company node, forming a "core-satellite" network topology with 23 suppliers such as CATL. The edge weight calculation uses a dynamic composite algorithm and is updated quarterly. For example, in Q1 of 2025, the procurement amount ratio of the CATL edge reaches 67% (1.2 billion / 1.8 billion), and the weight increases by 0.05 due to the on-time delivery rate of 99%; in Q2, the on-time delivery rate drops to 95% due to the congestion of the Sutong Bridge, and the weight is adjusted back by 0.03.
[0071] The construction of the capital relationship sub-network is achieved through equity penetration analysis. Trace the holding level of Yapu company: Yapu Group (51% holding) → Yapu Co., Ltd. (65% holding) → Lingang Yapu. At the same time, integrate the data of a 480 million yuan merger and acquisition event of acquiring a Jiangsu positive electrode material enterprise in 2024. Special indicators such as capital adequacy ratio (18%) and merger and acquisition activity (0.63) are designed in the node attributes. The edge weight calculation integrates multiple factors. For example, the holding edge weight of Yapu Group in Yapu Co., Ltd. = 51% × capital penetration coefficient 0.8 = 0.408, and the edge weight of the merger and acquisition of the Jiangsu enterprise is calculated as 0.77 (480 million / 620 million) according to the ratio of the merger and acquisition amount to the net assets. When Yapu Co., Ltd. increases its shareholding to 70% in 2025, the system updates the weight to 0.433 (51% × 0.85) in real time. This dynamic adjustment mechanism helps the government discover hidden control chains: the State-owned Assets Supervision and Administration Commission actually influences enterprise decisions through three-layer holding, prompting the R & D subsidy coefficient of such enterprises to be adjusted from 1.0 to 1.2 in 2025, and Yapu company thus increases its R & D investment in solid-state batteries by 240 million yuan.
[0072] The construction of the technical relationship sub-network relies on the deep association between patents and R & D data. Crawl the data of the solid-state battery laboratory jointly built by Yapu and Shanghai Jiao Tong University (with an investment of 80 million yuan in 2025), analyze 9 patents generated from the joint R & D, and associate them with the cross-licensing agreement of CTP3.0 patents of CATL (annual license fee of 120 million yuan). The technical node attributes include characteristics such as patent density (28 items / quarter) and the intensity of industry-university-research cooperation (0.67). The edge weight calculation uses a composite index: the edge weight of the cooperation with Shanghai Jiao Tong University = the number of patents 9 × the technical level coefficient 1.3 = 11.7, while the weight of the license edge of CATL = 120 million / 280 million total technical income = 0.43. When a new joint R & D project for sodium-ion batteries is added in Q2 of 2025, the system starts a dynamic increment mechanism, and the edge weight of the school-enterprise cooperation automatically increases by 0.5 per month. This intelligent network evolution enables the government to identify the "Yapu-Jiao Tong-CATL" innovation triangle closed-loop, and through the establishment of a special acceleration channel, the mass production time of solid-state batteries is advanced from 2027 to Q3 of 2026.
[0073] Refer to Figure 5 , Figure 5 is a schematic diagram of a sub-step process of step S203 provided by an embodiment of the present application. The above steps are as follows: Figure 2 In S501: Obtain multi-dimensional features corresponding to each enterprise from the valid data.
[0074] Specifically, in the embodiment of the present application, by analyzing the valid data (such as registration information, patent text, geographical location), combined with the preset industry knowledge base and NLP technology, multi-dimensional features covering business fields, technical directions, geographical distributions, and policy dependencies are extracted.
[0075] For example, in the new energy vehicle industry cluster, the business scope of a power battery enterprise includes "production of high-nickel ternary cathode materials", which is mapped to "upstream - battery materials" through entity recognition; the term "solid-state battery packaging technology" frequently appears in its patent text, and after analysis, it generates the "technical label - solid-state battery"; if the enterprise is located in Shanghai (a policy pilot area) and has received subsidies for 3 consecutive years, then the features of "high policy dependence" and "geographical location - core area of the Yangtze River Delta" are generated.
[0076] S502: Calculate the feature similarity between each dimension feature.
[0077] Specifically, in the embodiment of the present application, different similarity algorithms (such as cosine similarity, Jaccard coefficient, spatial distance calculation) are used for different feature types to quantify the association intensity between enterprises in dimensions such as business, technology, and geography.
[0078] For example, two new energy vehicle companies - Company A (technical tags: 800V high-voltage fast charging, sodium-ion battery) and Company B (technical tags: 800V thermal management system, lithium-sulfur battery), calculated the cosine similarity after quantizing the technical tags. The results showed that the similarity between the two in the "high-voltage system" technology direction was 0.85 (strong correlation), while the similarity in the battery chemical system direction was only 0.2 (weak correlation).
[0079] S503: Classify the relationships between enterprises into different types according to feature similarity.
[0080] Specifically, in an embodiment of the present application, a threshold or a clustering algorithm (such as K-means) is set to map the similarity value to relationship types such as competition, cooperation, and complementarity, and the classification results are calibrated in combination with industry rules.
[0081] For example, for the Yangtze River Delta new energy vehicle cluster, if the market area overlap of two vehicle manufacturers (Company C and Company D) exceeds 70%, the product price ranges are highly overlapped, and the technical similarity is lower than 0.3, they are judged to be in a "direct competition relationship"; and if the technical compatibility of an electric motor supplier (Company E) and multiple vehicle manufacturers (Companies F / G / H) is higher than 0.8, their relationship is defined as a "supply chain cooperation relationship."
[0082] S504: Use the divided types as edges and the enterprises as nodes, and combine with valid data to build a sub-network corresponding to each enterprise.
[0083] Specifically, in the embodiment of the present application, enterprises are used as nodes and relationship types are used as edges. Edge weights are assigned in combination with characteristic data (such as cooperation intensity and competition index) to construct sub-networks such as business, technology, and competition.
[0084] For example, in the business subnetwork, a battery supplier (node S) has established cooperative relations with three vehicle manufacturers, signed a 5-year long-term agreement with vehicle manufacturer T, and its annual supply accounts for 80% of its battery demand, generating a high-weight business edge (weight = 0.8), cooperated with vehicle manufacturer U for 2 years, and its supply accounted for 30%, generating a medium-weight edge (weight = 0.5), and was only in the trial supply stage with start-up car company V, without a long-term agreement, generating a low-weight edge (weight = 0.2), and had no direct business dealings with another battery supplier W, with an edge weight of 0. In the technical cooperation subnetwork, a battery company (node M) and four car companies (nodes N / O / P / Q) formed a high-weight edge (weight = 0.9) due to "joint research and development of solid-state batteries", but had no connection with another battery company (node R) due to low patent overlap and large differences in technical paths; in the competition subnetwork, nodes M and R generated a high-competitive intensity edge (weight = 0.7) because their market share overlap exceeded 60% and their customer groups were similar.
[0085] S204: Determine the association data between each sub-network, and connect each sub-network according to the association data to obtain the target hierarchical network.
[0086] Among them, the association data refers to the key information used to establish cross-layer connections between different sub-networks. In the embodiments of the present application, the association data specifically includes the set of common neighbors of nodes across networks and their overlap degree Nc, and the weight sets W1 and W2 of the edges connected to the common neighbors within each sub-network.
[0087] Specifically, in the embodiments of the present application, a hypergraph model is used to realize the coupling of multiple sub-networks, and cross-network connection rules are established through the analysis of the node overlap degree and edge correlation in the association data. Define the association strength index α = log(N_c + 1) × (W_1 + W_2), where N_c is the number of common neighbors across networks, and W_1 and W_2 are the edge weights of each sub-network. When α exceeds the preset threshold, a cross-layer connection edge is established, and finally a target hierarchical network with a multi-layer topological structure is formed.
[0088] For example, motor enterprise M supplies goods to 3 vehicle factories at the supply chain layer and shares patents with 5 parts enterprises at the technology layer. It is calculated that the association strength α of M in the two-layer network is 2.7 (log(3 + 1) × (0.8 + 0.6)), and a cross-layer connection is established after exceeding the threshold of 2.0. This enables the technical cooperation partners of M to influence its downstream enterprises at the supply chain layer through cross-layer edges.
[0089] Refer to Figure 6 , Figure 6 is a schematic diagram of a sub-step process of step S204 provided by the embodiments of the present application. The above steps are as follows: Figure 2 S601: Analyze the association data between sub-networks based on association rule mining technology. Specifically, by analyzing the co-occurrence rules of enterprise nodes, events or features in different sub-networks (such as business, technology, and competition sub-networks), cross-sub-network association rules are mined. For example, if an enterprise has a strong cooperation relationship in the technology sub-network and supplies goods to multiple enterprises in the business sub-network at the same time, it can be inferred that it has the role of a cross-network hub. At the same time, policy events (such as subsidy adjustments) may synchronously affect the behavior of enterprises in multiple sub-networks, forming event-driven association rules.
[0090] S602: Use the association data to establish a connection rule library.
[0091] Specifically, in the embodiments of the present application, based on the logic in the association rule library (such as "preferential connection of shared nodes" and "event-triggered cross-layer response"), a network fusion strategy is designed. For example, if an enterprise node satisfies the association rule in two sub-networks at the same time, a cross-layer connection edge is generated in its hierarchical network, and the connection strength is adjusted according to the rule weight.
[0092] Specifically, in the embodiments of the present application, based on the logic in the association rule library (such as "preferential connection of shared nodes" and "event-triggered cross-layer response"), a network fusion strategy is designed. For example, if an enterprise node satisfies the association rule in two sub-networks at the same time, a cross-layer connection edge is generated in its hierarchical network, and the connection strength is adjusted according to the rule weight.
[0093] S603: Obtain the target hierarchical network by adopting a network fusion algorithm and based on the connection rule library and the sub-networks.
[0094] Specifically, in the embodiments of the present application, based on the connection rule library, sub-networks such as business, technology, and competition are stacked layer by layer and fused into a unified hierarchical network through cross-layer edge fusion. During the fusion process, the independent structures of the sub-networks are retained, and at the same time, the cross-layer interaction relationships are explicitly expressed.
[0095] For example, an electric vehicle battery swapping technology enterprise X jointly develops battery swapping technology with vehicle enterprise Y in the technology sub-network (weight 0.8), and is connected to government node Z in the policy sub-network due to meeting the subsidy conditions (weight 0.9). According to the "policy-technology synergy rule" in the rule library, X-Y-Z forms a cross-layer triangular structure, revealing the logic of "policy promoting technology implementation"; at the same time, X supplies goods to battery swapping station operator W in the business sub-network (weight 0.7), and through the "technology-supply chain synergy rule", a cross-layer edge X-W (weight 0.6) is generated. Finally, the hierarchical network integrates multi-layer relationships such as technology R & D, business cooperation, and policy intervention, demonstrating the characteristics of enterprise X of "technology-driven + policy-benefited + business expansion".
[0096] S205: Extract key data from the target hierarchical network, and based on the key data, perform prediction through a pre-trained network structure evolution prediction model to obtain the network structure evolution result of the industrial cluster.
[0097] Specifically, in the embodiments of the present application, a spatio-temporal graph neural network (ST-GNN) is constructed as the core architecture of the prediction model. Its temporal convolution module captures the dynamic change patterns of the historical network structure, and the spatial graph attention mechanism learns the cross-layer influence weights between nodes. The model input is key data, which is a snapshot of the hierarchical network of continuous time slices. The network structure evolution result of the industrial cluster is output. A contrastive learning strategy is adopted to enhance the prediction ability of the model for sparse relationships, and a dynamic negative sampling technology is introduced to solve the long-tail distribution problem of enterprise cooperation relationships.
[0098] For example, after inputting the hierarchical network data of the new energy vehicle cluster from 2018 to 2022 into the model, the model successfully predicts the technology cooperation shift of a battery enterprise in 2023. The temporal convolution module identifies that the annual average weight of the technology layer edges of this enterprise has decreased by 12% in the past three years, while the spatial layer correlation has increased by 20%; the graph attention mechanism finds that the potential connection strength with solid-state battery R & D enterprises has reached the threshold.
[0099] Refer to Figure 7 , Figure 7 which is provided by the embodiments of the present application Figure 2Schematic diagram of a sub-step process for extracting key data in the target hierarchical network in step S205, and the above steps are as follows: S701: Obtain the node data of each node and the edge data of each edge from the target hierarchical network.
[0100] Specifically, in the embodiment of the present application, static attributes (such as business type, technology tags) and dynamic attributes (such as scale change, policy response record) of nodes (enterprises) are extracted from the target hierarchical network, as well as the type, weight, and interaction record of edges (relationships between enterprises). For example, in the hierarchical network of the Yangtze River Delta new energy vehicle cluster, a battery enterprise node includes static attributes: technology direction: solid-state battery, policy dependence: high, and dynamic attribute: the annual average growth rate of R & D investment from 2020 to 2023 is 45%; its edge data with vehicle enterprise A is a technology cooperation edge, with a weight of 0.8 and a cooperation period of 3 years.
[0101] S702: Analyze the node data to obtain the time-series dynamic fluctuation characteristics, where the time-series dynamic fluctuation characteristics are the characteristics of the node data fluctuating over time.
[0102] Specifically, in the embodiment of the present application, time-series data of the dynamic attributes of nodes (such as registered capital, market share, R & D investment) is analyzed, fluctuation rules (such as periodicity, trend, abnormal fluctuation) are extracted, and time-series dynamic characteristics are generated. For example, by analyzing the R & D investment data of the battery enterprise from 2018 to 2023, it is found that its annual average growth rate is 40%, but there is a short-term decline (abnormal fluctuation) in 2021 due to rising raw material prices, generating a time-series feature of "high growth trend (main) - short-term fluctuation (secondary)". Another charging pile enterprise has continuously declined in market share after 2022 due to the withdrawal of policy subsidies, generating a time-series feature of "continuous decay".
[0103] S703: Combine a preset industry relationship database and preset association rules to analyze the edge data and obtain association graph data.
[0104] Specifically, in the embodiment of the present application, the edge data is analyzed based on an industry knowledge base (such as a supply chain relationship database, a competition rule database), and the original relationships (such as "technology cooperation", "supply chain") are transformed into association graph data with semantics, including explicit relationships (direct cooperation) and implicit relationships (potential competition). For example, the "technology cooperation edge" between the battery enterprise and vehicle enterprise A is analyzed as "joint R & D relationship (explicit)", and although there is no direct business edge between the two battery enterprises, according to the rule of "default competition among suppliers of the same kind of technology" in the industry knowledge base, an implicit competition edge is generated.
[0105] Refer to Figure 8 , Figure 8 is provided by the embodiment of the present application Figure 7 Schematic diagram of a sub-step process for step S703 in the above, and the above steps are as follows: S801: Extract explicit features and implicit features from edge data. The explicit features include directly associated enterprise cooperation relationships, and the implicit features include indirectly associated enterprise potential competition relationships.
[0106] Specifically, in the embodiments of the present application, the explicit features are directly captured from the edge data. For example, contracts signed between enterprises, publicly disclosed joint patents, or clear supply chain relationships, which are usually recorded in cooperation agreements, transaction records, or government filings; the implicit features need to be derived through data association analysis. For example, calculate the proportion of shared suppliers between enterprises, the overlap of market customers, or identify potential competition relationships through graph algorithms (such as node betweenness centrality). In the new energy vehicle industry cluster, the raw material procurement paths and technology cooperation networks of battery manufacturers are typical explicit data, while the indirect competition caused by enterprises sharing scarce resources (such as lithium mines) or competing for the same customer group belongs to implicit features.
[0107] For example, for the power battery enterprise "CATL", its explicit features include the five-year battery cell supply contract (annual purchase volume of 50 GWh) signed with "NIO" and 15 solid-state battery patents jointly applied for with "CALB"; the implicit features are found through analysis that although CATL has no direct cooperation with "Guoxuan High-Tech", both of them purchase 70% of their lithium raw materials from "Ganfeng Lithium", and the customer overlap in the Yangtze River Delta region reaches 60%. Combining with the industry rule that "sharing core resources + market overlap > 50% is regarded as potential competition", it is deduced that there is an implicit competition relationship between the two.
[0108] S802: Convert the explicit features into explicit labels, and the explicit labels are used to describe the direct business connections between enterprises.
[0109] Specifically, in the embodiments of the present application, the explicit label is a standardized description of the explicit relationship, and the original data needs to be converted into a structured label according to the industry terminology library. For example, convert "annual supply volume of 50 GWh" into "core supplier", mark "15 joint patents" as "in-depth technical cooperation", and attach the relationship direction (such as one-way supply or two-way cooperation) and intensity value (quantified based on transaction scale or cooperation density).
[0110] For example, in the explicit relationship between CATL and NIO, the 50 GWh battery cell supply contract is converted into the label "Supply chain: CATL → NIO (core supplier, intensity 0.92)", and the 15 joint patents generate the label "Technical cooperation: CATL ↔ NIO (joint R & D, intensity 0.85)". The intensity value is calculated by the proportion of the contract amount in NIO's total procurement amount (92%) and the influence of the patented technology (85%).
[0111] S803: Convert implicit features into implicit labels through the industry knowledge base and association rule matching. The industry knowledge base is constructed based on industry white papers in the fields to which the industrial clusters belong. The association rules are obtained through valid data. The implicit labels are used to describe the indirect competition or complementary relationships between enterprises.
[0112] Specifically, in the embodiment of the present application, the generation of implicit labels depends on the predefined rules in the industry knowledge base. For example, it is clearly stated in the "White Paper on the New Energy Vehicle Power Battery Industry" that "if two enterprises share ≥ 2 core raw material suppliers and the market area overlap > 40%, they are marked as potential competitors". The association rule engine will match the implicit feature data. For example, it calculates the number of lithium mine suppliers shared by CATL and Guoxuan High-Tech, the market overlap, and outputs the label and confidence level in combination with historical competition case data (such as the probability of price war). CATL and Guoxuan High-Tech share two core lithium suppliers, Ganfeng Lithium and Huayou Cobalt (accounting for 70% of their respective purchases), and the customer overlap in the East China region reaches 65%. According to the industry rules, the "potential competition" label is triggered, and the confidence level is calculated by the formula "0.4 × the number of shared suppliers + 0.6 × the market overlap" to be 0.72 (0.4 × 2 + 0.6 × 0.65 = 0.72). At the same time, due to the complementary patent layout of the two in the technical route of lithium iron phosphate batteries, the additional label "technical complementarity (confidence level 0.55)" is added.
[0113] S804: Generate initial text labels by combining explicit labels and implicit labels according to the weights corresponding to the explicit labels and implicit labels respectively.
[0114] Specifically, in the embodiment of the present application, the weight distribution needs to reflect the business priorities. For example, the weight of the explicit label accounts for 70% (based on the legal effect of the contract), and the implicit label accounts for 30% (based on the predictive nature). The initial text labels are generated by weighted splicing. For example, "core supplier (competition risk)", where the explicit label determines the main description and the implicit label supplements the warning information. The weight calculation can introduce dynamic factors, such as the impact coefficient of policy changes on the competition relationship.
[0115] For example, the weight of the explicit label "core supplier (strength 0.92)" of CATL is 0.92 × 0.7 = 0.644, and the weight of the implicit label "potential competition (confidence level 0.72)" is 0.72 × 0.3 = 0.216. Since the explicit weight is significantly higher than the implicit weight, the initial text label "core cell supplier (with lithium resource competition with Guoxuan High-Tech)" is generated, and at the same time, the technical complementarity label "solid-state battery technology complementarity (confidence level 0.55)" is marked.
[0116] S805: Perform semantic verification on the initial text labels to obtain text labels.
[0117] Specifically, in the embodiments of the present application, semantic verification detects logical contradictions and term standardization through a natural language processing model. For example, it verifies whether "core supplier" conflicts with "competition risk". If there is a conflict, manual review is triggered, and the term library corrects "lithium resource competition" to "potential resource competition" to ensure compliance with the "New Energy Vehicle Industry Terminology Standard". Redundant tags (such as "cell supplier" and "battery supplier") will be merged.
[0118] For example, after verification of the initial tag "core cell supplier (lithium resource competition)", it is detected that the expression of "lithium resource competition" is ambiguous and is corrected to "core cell supplier (lithium resource procurement competition with Guoxuan High-Tech)"; "Supply chain - cell" and "Supply chain - long-term cooperation" are merged into "core cell supplier (long-term cooperation)"; the complete description of the complementary technology label is supplemented: "Lithium iron phosphate patent complementarity (joint R & D potential)".
[0119] S806: Bind the nodes in the target hierarchical network to text tags to obtain associated graph data.
[0120] Specifically, in the embodiments of the present application, the verified tags are bound to the corresponding enterprise nodes and stored as graph database attributes to support multi-dimensional query and visualization.
[0121] For example, in the new energy vehicle cluster associated graph, the CATL node is displayed. Business layer: "Core cell supplier → NIO (intensity 0.92)", "Secondary supplier → XPeng (intensity 0.6)"; Competition layer: "Potential resource competition ←→ Guoxuan High-Tech (confidence 0.72)", "Market share competition ←→ BYD (confidence 0.68)"; Technology layer: "Leader in solid-state batteries (45 patents)", "Lithium iron phosphate complementarity ←→ Guoxuan High-Tech (confidence 0.55)".
[0122] S704: Based on the intelligent fusion center, fuse the time-series dynamic fluctuation characteristics and associated graph data to generate key data.
[0123] Specifically, in the embodiments of the present application, through the fusion center (such as the time-series - graph joint model), the time-series dynamic characteristics (the changes of the enterprise itself) are combined with the associated graph data (the interactions between enterprises) to generate key data for prediction. For example, the time-series characteristics of a certain battery enterprise show that "R & D investment continues to grow", and its associated graph shows that "technologically bound to 3 leading automobile enterprises". After fusion, the key data is generated: "The technological leading advantage is strengthened, and the probability of supply chain cooperation expansion in the next 2 years > 70%". On the contrary, an enterprise with declining R & D investment and a dense competition edge is marked as "high risk of falling behind in technology and may be squeezed out of the core network".
[0124] Reference Figure 9, this application also provides a network structure evolution system 10 for emerging industrial clusters, specifically including: A data collection module 11, which is used to collect the enterprise information of each enterprise in the industrial cluster, and the industrial cluster is a collection of multiple enterprises in the same field; A data processing module 12, which is used to clean the enterprise information to obtain valid data; A relationship division and network construction module 13, which is used to divide the relationships between the enterprises into different types according to the valid data, and construct corresponding sub-networks for each enterprise according to the different types; A hierarchical network fusion module 14, which is used to determine the associated data between the sub-networks, connect the sub-networks according to the associated data, and obtain a target hierarchical network; A network evolution prediction module 15, which is used to extract key data from the target hierarchical network, and perform prediction through a pre-trained network structure evolution prediction model according to the key data to obtain the network structure evolution result of the industrial cluster.
[0125] Optionally, the information collection module 11 is further used to obtain the basic registration information and geographical location data of each enterprise in the industrial cluster; call the government affairs database interface to obtain the historical change records and policy subsidy data of the enterprise; perform entity recognition on the enterprise official website text and extract unstructured data; use the basic registration information, the geographical location data, the historical change records, the policy subsidy data, and the unstructured data as the enterprise information.
[0126] Optionally, the data processing module 12 is further used to verify and correct the error data based on a preset rule library to obtain corrected data; perform deduplication processing on the corrected data to obtain deduplicated data; use a generative adversarial network to fill in the missing values in the deduplicated data to obtain valid data.
[0127] Optionally, the relationship division and network construction module 13 is further used to obtain multi-dimensional features corresponding to each enterprise from the valid data; calculate the feature similarity between the multi-dimensional features; divide the relationships between the enterprises into different types according to the feature similarity; use the divided types as edges and the enterprises as nodes, and combine the valid data to construct the corresponding sub-networks for each enterprise.
[0128] Optionally, the hierarchical network fusion module 14 is further used to analyze the associated data between the sub-networks based on the association rule mining technology; establish a connection rule library using the associated data; use a network fusion algorithm and based on the connection rule library and the sub-networks to obtain a target hierarchical network.
[0129] Optionally, the network evolution prediction module 15 is further configured to obtain node data of each node and edge data of each edge from the target hierarchical network; analyze the node data to obtain a time-series dynamic fluctuation feature, where the time-series dynamic fluctuation feature is a feature of the node data fluctuating over time; combine a preset industry relationship database and preset association rules to parse the edge data to obtain association graph data; and based on the intelligent fusion center, fuse the time-series dynamic fluctuation feature and the association graph data to generate the key data.
[0130] Optionally, the network evolution prediction module 15 is further configured to extract explicit features and implicit features from the edge data, where the explicit features include directly associated enterprise cooperation relationships, and the implicit features include indirectly associated enterprise potential competition relationships; extract explicit features and implicit features from the edge data, where the explicit features include directly associated enterprise cooperation relationships, and the implicit features include indirectly associated enterprise potential competition relationships; convert the implicit features into implicit labels through a industry knowledge base and association rule matching, where the industry knowledge base is constructed based on industry white papers in the field to which the industrial cluster belongs, and the association rules are obtained through the valid data, and the implicit labels are used to describe the indirect competition or complementary relationships between enterprises; generate initial text labels by combining the explicit labels and the implicit labels according to the weights corresponding to the explicit labels and the implicit labels respectively; perform semantic verification on the initial text labels to obtain text labels; and bind the nodes in the target hierarchical network to the text labels to obtain association graph data.
[0131] It should be noted that when the device provided in the above embodiment implements its functions, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0132] This embodiment also discloses an electronic device 900, referring to Figure 10 , the electronic device may include: at least one processor 901, at least one communication bus 902, a user interface 903, a network interface 904, and at least one memory 905.
[0133] Among them, the communication bus 902 is used to realize the connection and communication between these components.
[0134] Among them, the user interface 903 may include a display screen (Display) and a camera (Camera). Optionally, the user interface may further include a standard wired interface and a wireless interface.
[0135] Among them, the network interface 904 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface).
[0136] Among them, the processor 901 may include one or more processing cores. The processor connects various parts within the entire server using various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling data stored in the memory. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.
[0137] Among them, the memory 905 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory may also be at least one storage device located far from the aforementioned processor. As shown in the figure, the memory, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for the network structure evolution of emerging industrial clusters.
[0138] In Figure 10In the electronic device shown, the user interface is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor can be used to call the application program stored in the memory for the evolution of the network structure of emerging industrial clusters. When executed by one or more processors, the electronic device is enabled to execute the method as described in one or more of the above embodiments.
[0139] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0140] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0141] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0142] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0144] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0145] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will easily think of other implementation manners of the present disclosure after considering the disclosure of the specification. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for the evolution of the network structure of an emerging industrial cluster, characterized in that Applied to an electronic device, the method includes: Collecting enterprise information of each enterprise in an industrial cluster, where the industrial cluster is a collection of multiple enterprises in the same field; Cleaning the enterprise information to obtain valid data; According to the valid data, classifying the relationships between the enterprises into different types, and constructing corresponding sub-networks for each enterprise according to the different types; Determining the associated data between the sub-networks, and connecting the sub-networks according to the associated data to obtain a target hierarchical network; Extracting key data from the target hierarchical network, and performing prediction according to the key data through a pre-trained network structure evolution prediction model to obtain the network structure evolution result of the industrial cluster.
2. The method according to claim 1, characterized in that, The collecting enterprise information of each enterprise in the industrial cluster specifically includes: Obtaining the basic registration information and geographical location data of each enterprise in the industrial cluster; Invoking the government affairs database interface to obtain the historical change records and policy subsidy data of the enterprise; Performing entity recognition on the enterprise official website text and extracting unstructured data; Taking the basic registration information, the geographical location data, the historical change records, the policy subsidy data, and the unstructured data as the enterprise information.
3. The method according to claim 1, wherein The cleaning the enterprise information to obtain valid data specifically includes: Based on a preset rule library, verifying and correcting the error data to obtain corrected data; Performing deduplication processing on the corrected data to obtain deduplicated data; Using a generative adversarial network to fill in the missing values in the deduplicated data to obtain valid data.
4. The method according to claim 1, characterized in that, The according to the valid data, classifying the relationships between the enterprises into different types, and constructing corresponding sub-networks for each enterprise according to the different types specifically includes: Obtaining multi-dimensional features corresponding to each enterprise from the valid data; Calculating the feature similarity between the multi-dimensional features; According to the feature similarity, classifying the relationships between the enterprises into different types; Taking the classified types as edges and the enterprises as nodes, and combining the valid data to construct corresponding sub-networks for each enterprise.
5. The method according to claim 1, wherein The determining the associated data between the sub-networks, and connecting the sub-networks according to the associated data to obtain a target hierarchical network specifically includes: Analyzing the associated data between the sub-networks based on the association rule mining technology; Using the associated data to establish a connection rule library; Adopting a network fusion algorithm and based on the connection rule library and the sub-networks to obtain a target hierarchical network.
6. The method according to claim 4, characterized in that, The extracting key data from the target hierarchical network specifically includes: Obtaining the node data of each node and the edge data of each edge from the target hierarchical network; Analyzing the node data to obtain a time-series dynamic fluctuation feature, where the time-series dynamic fluctuation feature is the feature of the node data fluctuating over time; Combining a preset industry relationship database and preset association rules to parse the edge data to obtain associated graph data; Based on an intelligent fusion center, fusing the time-series dynamic fluctuation feature and the associated graph data to generate the key data.
7. The method according to claim 6, characterized in that, The combining a preset industry relationship database to parse the edge data to obtain associated graph data specifically includes: Extract explicit features and implicit features from the edge data, where the explicit features include directly associated enterprise cooperation relationships, and the implicit features include indirectly associated potential enterprise competition relationships; Extract explicit features and implicit features from the edge data, where the explicit features include directly associated enterprise cooperation relationships, and the implicit features include indirectly associated potential enterprise competition relationships; Convert the implicit features into implicit labels through industry knowledge bases and association rule matching. The industry knowledge bases are constructed based on industry white papers in the field to which the industrial cluster belongs, and the association rules are obtained through the valid data. The implicit labels are used to describe the indirect competition or complementary relationships between enterprises; Generate initial text labels by combining the explicit labels and the implicit labels according to the weights corresponding to the explicit labels and the implicit labels respectively; Perform semantic verification on the initial text labels to obtain text labels; Bind the nodes in the target hierarchical network to the text labels to obtain associated graph data.
8. A network structure evolution system for an emerging industrial cluster, wherein Comprising: A data collection module for collecting enterprise information of each enterprise in the industrial cluster, where the industrial cluster is a collection of multiple enterprises in the same field; A data processing module for cleaning the enterprise information to obtain valid data; A relationship division and network construction module for dividing the relationships between the enterprises into different types according to the valid data, and constructing corresponding sub-networks for each of the enterprises according to the different types; A hierarchical network fusion module for determining the associated data between the sub-networks, and connecting the sub-networks according to the associated data to obtain a target hierarchical network; A network evolution prediction module for extracting key data in the target hierarchical network, and predicting according to the key data through a pre-trained network structure evolution prediction model to obtain a network structure evolution result of the industrial cluster.
9. An electronic device, characterized in that, Comprising a processor, a memory, a user interface, and a network interface. The memory is used for storing instructions. Both the user interface and the network interface are used for communicating with other devices. The processor is used for executing the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-7 is executed.
Citation Information
Cited By
Scientific research field knowledge graph construction method, system and equipment and storage medium
CN121599067A