Industrial gathering area boundary identification method based on industry field
By constructing industry statistical maps and enterprise datasets, and combining density clustering and community detection algorithms, the boundaries of industrial clusters are identified, solving the accuracy problem of industrial cluster boundary identification in existing technologies, and realizing accurate identification and micro-spatial visualization in specific industry fields.
Patent Information
- Application Number
- CN202511121889.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies struggle to accurately identify the boundaries of industrial clusters in specific sectors, particularly in the integration and analysis of macro-level industrial chain logic maps with micro-level geographical space, making it difficult to effectively capture the spatial forms of emerging industries.
By constructing industry statistical maps and industry enterprise datasets, we calculate industry agglomeration and industry connectivity, use density clustering and community detection algorithms to identify the boundaries of industry agglomeration areas, and combine neural networks and geographic information systems for micro-spatial visualization.
It enables precise identification of the boundaries of industrial clusters in specific industry sectors, making up for the lack of microscopic spatial visualization and analysis methods for internal industrial connections in existing technologies, and improving the accuracy and effectiveness of identification.
Smart Images

Figure CN121256313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of industrial cluster boundary identification, and in particular to a method for industrial cluster boundary identification based on an industry sector. Background Technology
[0002] Accurately identifying industrial clusters—the spatial carriers of specific industries—and clearly defining their geographical boundaries is a key foundation for effectively optimizing industrial layout, allocating factor resources, and evaluating policy effectiveness.
[0003] However, from the perspective of statistically identifying specific industry sectors, existing methods face severe challenges in identifying their spatial agglomeration boundaries: First, there is a disconnect between the current "standard classifications" such as the National Economic Industry Classification and the "industrial chain" based on division of labor and cooperation and with inherent economic connections. The competitiveness of industrial clusters often stems from "horizontal" industrial chain collaboration. Second, "industrial chain maps" at the national and regional levels are usually logical and non-spatial. Existing methods lack effective means to "land" the macro-level logical map of the industrial chain in micro-geographical space and integrate it with the location and connection data of specific enterprises to meet the "micro-level needs" of local governments and park managers for spatial implementation. Third, major countries around the world are currently using the development of emerging industries and the cultivation of new productive forces as core strategies to seize the commanding heights of future economic and technological development and promote high-quality development. Emerging industries often exhibit characteristics of cross-border technological integration, rapid business model innovation, and industrial chain restructuring. Cluster identification methods based on traditional industry classifications and historical experience are difficult to adapt quickly and accurately capture the spatial forms of these emerging and complex industrial chains. Summary of the Invention
[0004] In order to overcome the shortcomings of the existing technology, the purpose of this invention is to propose a method for identifying the boundaries of industrial clusters based on industry sectors, which can effectively analyze the characteristics of industrial clusters in specific industry sectors and achieve accurate identification of the boundaries of industrial clusters.
[0005] To achieve the objectives of this invention, the following technical solution is adopted: A method for identifying the boundaries of industrial clusters based on industry sectors, the method comprising the following steps: Obtain industry classification data for the target region, and construct an industry statistical map based on the industry classification data for the target region; Acquire enterprise data for the target region, and construct an industry-specific enterprise dataset based on the industry statistical map and the enterprise data; Based on the enterprise dataset in the aforementioned industry sector, calculate the industrial agglomeration degree and industrial linkage degree of the target region; Based on the calculated industrial agglomeration degree and industrial linkage degree, the scope of industrial agglomeration areas in the target region is identified.
[0006] In the above technical solution, by constructing industry statistical maps and industry enterprise datasets, the industrial data of the target region can be effectively classified and labeled, and then the industrial agglomeration degree and industrial linkage degree of the target region can be calculated. Based on the calculated industrial agglomeration degree and industrial linkage degree, the industrial agglomeration characteristics of specific industry sectors can be effectively analyzed, and the boundary range of industrial agglomeration areas can be accurately identified. This makes up for the lack of methods for visualizing the micro-space of specific industrial chain logic maps and the lack of methods for considering internal industrial linkages within the industry.
[0007] Furthermore, the process of constructing an industry statistical map based on the industry classification data of the target region includes: Based on the acquired industry classification data of the target region, the pre-set industry classification model is trained to obtain a trained industry classification model. The industry classification model extracts upstream, midstream, and downstream data of the industrial chain based on the industry classification data of the target region, and constructs an industry statistical map based on the upstream, midstream, and downstream data of the industrial chain. The upstream data of the industrial chain refers to data on the research, design and production of raw materials or key components; the midstream data of the industrial chain refers to data on product production, processing, assembly and manufacturing; and the downstream data of the industrial chain refers to data on sales, logistics, after-sales service and end consumers.
[0008] Furthermore, the industry classification model, based on the upstream, midstream, and downstream data of the industrial chain, constructs an industry statistical map using a directory mapping method and a keyword semantic expansion method.
[0009] Furthermore, the process of constructing an industry-specific enterprise dataset based on the industry statistical map and the enterprise data includes: Based on the industry statistical map, the enterprise data is cleaned to identify enterprises that do not belong to the industry statistical map, thus obtaining an industry-specific enterprise dataset. Enterprise point coordinate data is obtained based on the enterprise data, and an industry chain enterprise point dataset is constructed based on the industry neighbor enterprise dataset and the enterprise point coordinate data. The industrial chain enterprise point dataset includes upstream industrial chain data, midstream industrial chain data, and downstream industrial chain data.
[0010] Furthermore, the process of calculating the industrial agglomeration degree of the target region includes: Based on the aforementioned industry chain enterprise point dataset, all enterprises are ranked by registered capital, and a core enterprise threshold is set based on the ranking data. Extract the enterprises that reach the core enterprise threshold in the upstream industry data, midstream industry data, and downstream industry data as core enterprises, and obtain several core enterprises; Based on each of the above-mentioned core enterprises and a number of ordinary enterprises that are density-connected, construct a sub-industrial agglomeration area based on the industrial agglomeration degree, obtain several sub-industrial agglomeration areas based on the industrial agglomeration degree, calculate the boundary range of each sub-industrial agglomeration area based on the industrial agglomeration degree, and calculate the boundary range of the industrial agglomeration area based on the industrial agglomeration degree based on the boundary ranges of the several sub-industrial agglomeration areas based on the industrial agglomeration degree; Among them, the core enterprise refers to an enterprise that reaches the core enterprise threshold; the ordinary enterprise refers to an enterprise that does not reach the core enterprise; the density connection means reaching the preset density connection threshold, and both are non-core points but are within the neighborhood of a certain point.
[0011] Further, use the density-based spatial clustering method to calculate the boundary range of each sub-industrial agglomeration area based on the industrial agglomeration degree. The process includes: Traverse all the enterprise point data sets of the industrial chain , randomly select an enterprise point , according to the preset density connection threshold, expand outward according to the condition of density reachability. In the neighborhood formed by a preset radius, if the number of points m < Minpts, then mark the enterprise point as a noise point; if m ≥ Minpts, then mark the enterprise point as a core enterprise point, and jointly construct a sub-industrial agglomeration area with the ordinary enterprise points that are density-connected to the core enterprise point, and obtain several sub-industrial agglomeration areas based on the industrial agglomeration degree; Extract several enterprise points on the periphery of each sub-industrial agglomeration area based on the industrial agglomeration degree , form the range of the sub-industrial agglomeration area based on the industrial agglomeration degree , and based on the range of each sub-industrial agglomeration area based on the industrial agglomeration degree , calculate and obtain the range of the industrial agglomeration area based on the industrial agglomeration degree ; Among them, Minpts represents the minimum number of sample points of the core enterprise; the noise point takes a certain point as the center of the circle and forms a neighborhood with a preset radius. The number of sample points m in the neighborhood < Minpts. At the same time, this point also does not belong to the neighborhood formed by other points as the center of the circle, then this point is a noise point.
[0012] Further, calculate the silhouette coefficient S of each enterprise point based on the industrial agglomeration area range B. The process includes: According to the range of the sub-industrial agglomeration area , calculate the enterprise point With the same sub-industrial cluster Let a be the average distance to other enterprise locations within the same area; and calculate the distance to that enterprise location. Let b be the average distance between the enterprise locations and other sub-industry clusters. Then the profile coefficient of the enterprise location is S = (ba) / max(a,b). Here, max(a,b) represents taking the maximum value of a and b respectively.
[0013] Furthermore, the process of calculating the industrial linkages of the target region includes: A directed network model is constructed based on the aforementioned industry chain enterprise point dataset; The Infomap algorithm for community discovery is used to identify the community structure within the directed network model, resulting in several sub-industry clusters. Calculate the extent of each of the sub-industrial clusters, and calculate the extent of the industrial cluster based on the extent of the sub-industrial clusters.
[0014] Furthermore, the process of constructing a directed network model based on the aforementioned industry chain enterprise point dataset includes: Using the acquired enterprises as nodes in the network and the customer-supplier relationships between enterprises as edges, a directed network model is constructed, expressed as:
[0015] Where X represents the dataset of enterprise points in the industry chain, and E represents the set of edges; The process of identifying the community structure within the directed network model and obtaining several sub-industry clusters includes: Suppose we use n codewords to represent any enterprise in the industry chain enterprise point dataset X. Probability of occurrence during community walks Given n states, the average encoding length is no less than that of any enterprise point in the industry chain enterprise point dataset X. Its own entropy is expressed as:
[0016] For all businesses in the community Perform Huffman coding to represent any path in the network through a combination of codes, as shown in the expression:
[0017] in, This represents the expected average coding length of the paths taken during random walks within and between communities; This represents the probability of a random walk leaving one community and entering another. Entropy represents the probability of a random walk moving between different communities; This represents the probability that a random walk exists within a certain community; when When the change is less than the preset value, the scope of the closely connected sub-industry clusters is obtained. Based on the scope of several sub-industry clusters Industrial cluster area .
[0018] Furthermore, a fishing net model is constructed based on a neural network architecture, and the calculated silhouette coefficient S of each enterprise point is compared with the expected value of the average coding length. The values are assigned to the fishing net model and converted into raster data; The weighted raster data overlay analysis method is used to extract the value of each pixel in the fishing net model, and the weights of industrial agglomeration and industrial linkage are calculated based on the entropy method. The expression is as follows: B´ 1i =B 1i ×W i D´ 1j = D 1j ×W j Among them, B 1i B represents the scope of the sub-industrial clusters formed based on the degree of industrial agglomeration. i The value of the l-th cell within, D 1j B represents the scope of sub-industry clusters formed based on industrial connectivity. j The value of the l-th cell within, W i W represents the weight used to calculate the degree of industrial agglomeration based on the entropy method. j B' represents the weight used to calculate the degree of industrial linkage based on the entropy method. 1i With D´ 1j All of these represent the updated cell values; the term "cell" in this section refers to any grid in the fishing net model.
[0019] The natural discontinuity classification method in ArcGIS was used to update B'. 1i With D´ 1j The scope of the industrial clusters in the target region is obtained by performing hierarchical extraction.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a method for identifying the boundaries of industrial clusters in a specific industry sector. It constructs an industry statistical map and a dataset of enterprises within the industry sector from a supply chain perspective. This method effectively classifies and labels the industrial data of the target region, and then calculates the degree of industrial agglomeration and industrial connectivity in the target region. Based on the calculated degree of industrial agglomeration and industrial connectivity, it can effectively analyze the characteristics of industrial clusters in a specific industry sector, achieving accurate identification of the boundaries of industrial clusters. This method compensates for the lack of methods for visually representing the micro-space of a specific supply chain logic map within the industry, as well as the lack of methods that consider internal industry connections. Attached Figure Description
[0021] Figure 1 A flowchart illustrating the steps of an industry cluster boundary identification method based on an embodiment of this application; Figure 2 A schematic diagram of the directed network model and community discovery process provided in the embodiments of this application; Figure 3 A schematic diagram of the smart home appliance industry chain modules provided in this application embodiment; Figure 4 A schematic diagram of the constructed smart home appliance industry chain provided in the embodiments of this application; Figure 5 This application provides a distribution map of upstream, midstream, and downstream enterprises in the smart home appliance industry. Figure 6 Distribution map of core smart home appliance enterprises provided in the embodiments of this application; Figure 7 A schematic diagram illustrating the analysis results of the density-based spatial clustering method provided in this application embodiment; Figure 8 This is a schematic diagram illustrating the scope of smart home appliance industry clusters based on enterprise clustering characteristics, provided in an embodiment of this application. Figure 9 This is a network diagram based on enterprise connections provided in an embodiment of this application; Figure 10 This is a schematic diagram of community discovery results provided in an embodiment of this application; Figure 11 This application provides a schematic diagram illustrating the clustering scope of the smart home appliance industry based on enterprise contact features in its embodiments. Figure 12 This is a schematic diagram showing the scope of the smart home appliance industry cluster provided in this application embodiment. Detailed Implementation
[0022] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0024] Example 1: This embodiment provides a method for identifying the boundaries of industrial clusters based on a specific industry sector. (See also...) Figure 1 The method includes the following steps: Step S1: Obtain industry classification data for the target region, and construct an industry statistical map based on the industry classification data for the target region; Step S2: Obtain enterprise data for the target region, and construct an industry-specific enterprise dataset based on the industry statistical map and the enterprise data; Step S3: Based on the enterprise dataset of the aforementioned industry sector, calculate the industrial agglomeration degree and industrial linkage degree of the target region; Step S4: Based on the calculated industrial agglomeration degree and industrial linkage degree, identify the scope of industrial agglomeration areas in the target region.
[0025] In step S1, the process of constructing an industry statistical map based on the industry classification data of the target region includes: Based on the acquired industry classification data of the target region, the preset industry classification model is trained to obtain a trained industry classification model, wherein the industry classification model is a BERT model. The industry classification model extracts upstream, midstream, and downstream data of the industrial chain based on the industry classification data of the target region, and constructs an industry statistical map based on the upstream, midstream, and downstream data of the industrial chain. The upstream data of the industrial chain refers to data on the research, design and production of raw materials or key components; the midstream data of the industrial chain refers to data on product production, processing, assembly and manufacturing; and the downstream data of the industrial chain refers to data on sales, logistics, after-sales service and end consumers.
[0026] Furthermore, the industry classification model, based on the upstream, midstream, and downstream data of the industrial chain, constructs an industry statistical map using a directory mapping method and a keyword semantic expansion method.
[0027] Specifically, data on the names and codes of the target region based on the "National Economic Industry Classification Standard (GB / T4754-2017)" is obtained, and the statistical categories and codes of the upstream, midstream and downstream industries are extracted by integrating the statistical caliber catalog of strategic emerging industries and the industrial chain caliber.
[0028] Supply Chain Module Extraction: Based on Porter's value chain model, corporate activities are divided into primary activities (direct value creation) and support activities (indirect assistance). The overall structural relationship of the supply chain is analyzed. The upstream is "raw material or key component R&D, design and production"; the midstream is "product production, processing, assembly and manufacturing"; and the downstream is "sales, logistics, after-sales service and end consumer". The upstream, midstream and downstream relationships and key links of the supply chain are constructed, and specific modules of the upstream, midstream and downstream of the professional industry are extracted.
[0029] Industry statistical map construction: Input the strategic emerging industry statistical catalog published by the national or provincial government and the specific products of the upstream, midstream and downstream professional industries output by sub-module 1 into the BERT model. Using a combination of catalog mapping and keyword semantic expansion, extract the sub-category names and industry codes corresponding to the upstream, midstream and downstream of the industry from the 1381 sub-categories of the "National Economic Industry Classification Standard (GB / T4754-2017)". Combine and summarize to form an industry statistical map that integrates government and industry chain standards.
[0030] Directory mapping. ① Extract the statistical catalog of strategic emerging industries published by national or provincial governments, and match the standardized format using regular expressions. ② Directly extract sub-category codes. ③ Map keywords for entries without codes, with a matching degree = max(|directory keywords|GB / T 4754-2017 description| / |directory keywords|), and accepting matches with a matching threshold > 0.85. ④ Manual logical verification, including fuzzy matches where a single directory term matches multiple sub-categories and the similarity is between 0.75 and 0.85.
[0031] Keyword semantic expansion. The core terms of the specific modules in the upstream, midstream, and downstream of the industry output by submodule 1 are expanded with synonyms, and related subcategories are retrieved in GB / T 4754-2017.
[0032] In step S2, the process of constructing an industry-specific enterprise dataset based on the industry statistical map and the enterprise data includes: Based on the industry statistical map, the enterprise data is cleaned to remove enterprises that do not belong to the industry statistical map, thus obtaining an industry-specific enterprise dataset. Enterprise point coordinate data is obtained based on the enterprise data, and an industrial chain enterprise point dataset is constructed based on the industry sector enterprise dataset and the enterprise point coordinate data. The industrial chain enterprise point dataset includes upstream industrial chain data, midstream industrial chain data, and downstream industrial chain data.
[0033] Specifically, industrial agglomeration is the process by which enterprises gather in geographical space to reduce transaction costs and generate economic connections such as cooperation between different links in the industrial chain and supply chain. The degree of industrial agglomeration and the degree of industrial connectivity are important factors affecting the development of industrial clusters. Industry-specific enterprise data acquisition and cleaning: Based on the Longdun Enterprise Big Data Database, enterprise data is acquired, and enterprise registration location information, enterprise registered capital, enterprise business scope, enterprise customer and supplier information, and industry code are extracted. Based on the industry statistical map formed in sub-module two, enterprises not included in the industry statistical map are cleaned to form an industry-specific enterprise dataset. Geocoding tools parse enterprise location information: Filter out the enterprise's registered address, and use Baidu Maps coding tool to convert the enterprise's registered address into the geographical coordinates of the enterprise's location, forming an enterprise point coordinate dataset; Enterprise Dataset Construction: Using ArcGIS software, enterprise coordinates are converted into spatial enterprise points, forming enterprise point .shp files. Using enterprise names as search terms, an enterprise geographic information database X (i.e., industry chain enterprise point dataset) is formed.
[0034] In step S3, the process of calculating the industrial agglomeration degree of the target region includes: Based on the aforementioned industry chain enterprise point dataset, all enterprises are ranked by registered capital, and a core enterprise threshold is set based on the ranking data. Enterprises that reach the core enterprise threshold from upstream, midstream, and downstream industry data are identified as core enterprises, resulting in a number of core enterprises. Based on each core enterprise and several densely connected ordinary enterprises, a sub-industrial cluster based on industrial agglomeration degree is constructed, resulting in several sub-industrial clusters based on industrial agglomeration degree. The boundary range of each sub-industrial cluster based on industrial agglomeration degree is calculated, and the boundary range of the industrial cluster based on industrial agglomeration degree is calculated based on the boundary range of the several sub-industrial clusters based on industrial agglomeration degree. Among them, the core enterprises refer to the enterprises that reach the core enterprise threshold; the ordinary enterprises refer to the enterprises that do not reach the core enterprises; the density connection means reaching the preset density connection threshold, and both are non-core points, but are both within the neighborhood of a certain point.
[0035] Furthermore, the Density-Based Spatial Clustering of Application with Noise (DBSCAN) method under the improved noisy application background is used to calculate the boundary range of each industrial agglomeration area. The Minpts no longer only represents the minimum number of samples (i.e., the number of enterprises) in a sub-industrial agglomeration area based on the industrial agglomeration degree. The core enterprises are the core carriers that aggregate the upstream, midstream, and downstream of the industrial chain. Replace Minpts with the number of core enterprises (definition of core enterprises: all enterprises in the research area are arranged in reverse order according to the registered capital amount of the enterprises, and the top 100 enterprises are defined as core enterprises). Thus, several sub-industrial agglomeration areas based on the industrial agglomeration degree are obtained. When determining the geographical scope of the industrial agglomeration area, extract the enterprises located on the boundary within the sub-industrial agglomeration area based on the industrial agglomeration degree as geographical boundary points, and use the ArcGIS tool to connect the boundary points and draw the area. These areas are the identified industrial agglomeration area scopes. The process includes: Traverse all the enterprise point data sets of the industrial chain , randomly select an enterprise point , according to the preset density connection threshold, expand outward according to the condition of density reachability. In the neighborhood formed by a preset radius, if the number of points m < Minpts, then mark the enterprise point as a noise point; if m ≥ Minpts, then mark the enterprise point as a core enterprise point, and jointly construct a sub-industrial agglomeration area with the ordinary enterprise points that are density-connected to the core enterprise point to obtain several sub-industrial agglomeration areas based on the industrial agglomeration degree; Extract several enterprise points on the periphery of each sub-industrial agglomeration area based on the industrial agglomeration degree , form the scope of the sub-industrial agglomeration area based on the industrial agglomeration degree , and based on the scope of each sub-industrial agglomeration area based on the industrial agglomeration degree , calculate and obtain the scope of the industrial agglomeration area based on the industrial agglomeration degree ; Among them, Minpts represents the number of points of the minimum sample points of core enterprises; the noise point takes a certain point as the center of the circle and forms a neighborhood with a preset radius. The number of sample points m within the neighborhood < Minpts. At the same time, this point also does not belong to the neighborhood formed by other points as the center of the circle, then this point is a noise point.
[0036] Furthermore, the location of each enterprise is calculated based on the industrial cluster area B, which is determined by the degree of industrial agglomeration. The profile coefficient S is obtained through a process that includes: Based on the scope of sub-industrial clusters according to industrial agglomeration degree Calculate enterprise points With the same sub-industrial cluster based on industrial agglomeration degree Let a be the average distance to other enterprise locations within the same area; and calculate the distance to that enterprise location. Let b be the average distance between the enterprise points and other sub-industrial clusters based on industrial agglomeration degree. Then the profile coefficient of the enterprise point is S=(ba) / max(a,b). Here, max(a,b) represents taking the maximum value of a and b respectively.
[0037] In a preferred embodiment, step S3, the process of calculating the industrial connectivity of the target region, includes: A directed network model is constructed based on the aforementioned industry chain enterprise point dataset; The Infomap algorithm for community discovery is used to identify the community structure within the directed network model, resulting in several sub-industry clusters based on industry connectivity. Calculate the extent of each of the sub-industrial clusters based on industrial connectivity, and calculate the extent of the industrial cluster based on industrial connectivity based on the extent of several of the sub-industrial clusters based on industrial connectivity.
[0038] Specifically, the process of constructing a directed network model based on the aforementioned industry chain enterprise point dataset includes: See Figure 2 (In the diagram, ac represents the network construction process based on the enterprise customer-supplier relationship, and de represents the community discovery process.) The acquired enterprises are the nodes in the network, and the customer-supplier relationships between enterprises are the edges of the network. A directed network model is constructed, expressed as:
[0039] Where X represents the dataset of enterprise points in the industry chain, and E represents the set of edges; The process of identifying the community structure within the directed network model and obtaining several sub-industry clusters based on industry connectivity includes: Based on the community discovery Infomap algorithm, the community structure within the constructed directed network of enterprise connections is identified. Several smaller communities are defined as sub-industry clusters based on industry connectivity, specifically: Community discovery: Let n codewords be used to represent any enterprise in the supply chain enterprise point dataset X. Probability of occurrence during community walks Given n states, the average encoding length is no less than that of any enterprise point in the industry chain enterprise point dataset X. Its own entropy is expressed as:
[0040] To characterize the random walk process, Huffman coding is applied to all nodes. Each node is composed of the Huffman code of its community and the Huffman codes of specific nodes within that community. Huffman coding is used to describe the number and frequency of visits a node makes during random walks in the network, allowing any path in the network to be represented by a combination of codes. Therefore, if the minimum code length is optimized as the objective function, the community structure partitioning problem is transformed into an information flow path coding compression problem.
[0041] For all businesses in the community Huffman coding is performed to represent any path in the network through a combination of codes. According to information entropy theory, the mapping equation is as follows:
[0042] in, This represents the expected average coding length of the paths taken during random walks within and between communities; This represents the probability of a random walk leaving one community and entering another. Entropy represents the probability of a random walk moving between different communities; This represents the probability that a random walk exists within a certain community; Repeat the above calculation process until... When the change is less than the preset value, the scope of closely connected sub-industry clusters based on industrial linkages is obtained. Based on the scope of several sub-industry clusters based on industrial linkages Industrial clusters based on industrial connectivity are obtained. .
[0043] In this embodiment, a 500m×500m fishing net model is constructed based on a neural network architecture. The calculated contour coefficient S of each enterprise point is compared with the expected value of the average coding length. The values are assigned to the fishing net model and converted into raster data; The weighted raster data overlay analysis method is used to extract the value of each pixel in the fishing net model, and the weights of industrial agglomeration and industrial linkage are calculated based on the entropy method. The expression is as follows: B´ 1i =B 1i ×W i D´ 1j = D 1j ×Wj Among them, B 1i B represents the scope of the sub-industrial clusters formed based on the degree of industrial agglomeration. i The value of the l-th cell within, D 1j B represents the scope of sub-industry clusters formed based on industrial connectivity. j The value of the l-th cell within, W i W represents the weight used to calculate the degree of industrial agglomeration based on the entropy method. j B' represents the weight used to calculate the degree of industrial linkage based on the entropy method. 1i With D´ 1j All represent the updated cell values; the cell represents any grid in the fishing net model.
[0044] The natural discontinuity classification method in ArcGIS was used to update B'. 1i With D´ 1j The scope of the industrial clusters in the target region is obtained by performing hierarchical extraction.
[0045] In this embodiment, the present invention aims to focus on a specific industry sector, utilize common enterprise big data, and automatically extract industrial clusters by considering the spatial agglomeration and connections of enterprises through the structural characteristics of the upstream, midstream, and downstream of the industrial chain. It adopts interdisciplinary technologies combining economics, geography, and urban planning, and constructs industry statistical maps and industry sector enterprise datasets to effectively classify and label the industrial data of the target region. Then, it calculates the industrial agglomeration degree and industrial connectivity degree of the target region. Based on the calculated industrial agglomeration degree and industrial connectivity degree, it can effectively analyze the industrial agglomeration characteristics of a specific industry sector, achieve accurate identification of the boundary range of industrial clusters, and make up for the lack of methods for micro-spatial visualization of specific industrial chain logic maps and methods for considering internal industrial connections within the industry.
[0046] Example 2: This embodiment, based on the method described in Embodiment 1, takes the identification of the Guangdong Provincial Smart Home Appliance Industry Cluster as an example and specifically includes: Step 1. Extraction of smart home appliance industry chain modules: See Figure 3(1) The upstream of the smart home appliance industry covers hardware such as electronic components and middleware, as well as software such as system platforms, which are mainly based on the shift towards intelligentization. Specifically, it includes: ① Electronic components: sensors, chips, printed circuits, motors, microcontrollers; ② Middleware: communication modules, intelligent controllers, display modules; ③ Software: system platforms, solutions, Internet of Things technology, human-computer interaction. (2) The midstream of the smart home appliance industry covers the manufacturing, packaging and finished products of different types of home appliances. Specifically, it includes smart air conditioners, smart washing machines, smart refrigerators, smart TVs, smart speakers, cameras, kitchen and bathroom small appliances, home small appliances, care small appliances, as well as quality testing, pilot production, industrial design, etc. (3) The downstream of the smart home appliance industry covers wholesale, retail and after-sales service, etc. Specifically, it includes e-commerce platforms, home appliance alliances, home decoration companies, after-sales service, and replacement service.
[0047] Step 2. Directory Mapping: Extract keywords from the statistical directory of Guangdong Province's strategic pillar industry cluster of smart home appliances. Match them in the 1381 subcategories of the "National Economic Industry Classification Standard (GB / T4754-2017)" using the "direct extraction" or "keyword mapping" method to extract the subcategory names and codes.
[0048] Step 3. Keyword semantic expansion: Expand the core terms of specific modules in the upstream, midstream and downstream of the smart home appliance industry chain with synonyms, and conduct related searches in the 1381 subcategories of the "National Economic Industry Classification Standard (GB / T4754-2017)" to extract the subcategory names and codes corresponding to the upstream, midstream and downstream of the industry.
[0049] Step 4. Match the industry statistical map; merge the search results from Step 2 and Step 3, classify and merge them into the corresponding sub-category names and codes of the upstream, midstream, and downstream of the smart home appliance industry chain, and match to form a smart home appliance industry statistical map, see... Figure 4 .
[0050] Step 5. Enterprise Data Acquisition and Processing: First, enterprise data is acquired based on the Longdun Enterprise Big Data Database, extracting enterprise registered address information, registered capital, business scope, customer and supplier information to form an enterprise database. Second, using a BERT pre-trained model, the extracted business scope terms are input and, referring to the "National Economic Industry Classification," enterprises are classified according to their respective industries. Furthermore, using Baidu Maps encoding tools, the enterprise registered addresses are converted into geographical coordinates of their locations and visualized in ArcGIS, forming an upstream, midstream, and downstream enterprise chain. (See...) Figure 5 .
[0051] Step 6. Core Enterprise Identification: Based on the upstream, midstream, and downstream enterprises in the industrial chain identified in Step 5, all enterprises within the study area are ranked in descending order of their registered capital. The top 100 enterprises are defined as core enterprises. See [link / reference]. Figure 6 .
[0052] Step 7. Identifying the industrial cluster area based on industry agglomeration degree based on enterprise clustering characteristics: Import the enterprise point shapefile data into ArcGIS and perform cluster analysis. Start density-based spatial clustering analysis, select the number of core enterprises and the total number of enterprises, and select a search radius of 4000m to form n neighborhoods. Within a neighborhood, if the number of points m < 5, it will be marked as a noise point; otherwise, the point is a core point. The largest set of objects connected to the core point density forms a sub-industrial cluster area based on industry agglomeration degree. The sub-industrial cluster area formed is the industrial cluster area based on industry agglomeration degree. Extract the enterprise points outside the cluster area as geographic boundary points. Use GIS tools to connect the boundary points and draw the area. These areas are the identified industrial cluster areas based on industry agglomeration degree based on enterprise clustering characteristics. See Figure 7 and Figure 8 .
[0053] Step 8. Construct a directed network based on enterprise relationships: Using the acquired enterprises as nodes in the network, and the "customer-supplier relationship" between enterprises as edges, construct a directed network with the direction of "customer enterprise - supplier enterprise". See [link to relevant documentation]. Figure 9 .
[0054] Step 9. Community Discovery: Identifying the industrial cluster boundaries based on industry connectivity. Treat each node as a small community. Assuming nodes move randomly, calculate the expected average encoding length of the paths taken during random walks within and between communities, considering both node and community merging scenarios. Merge to minimize the number of nodes and communities. Repeat this calculation until the change is minimized, resulting in multiple communities. Each community represents an industrial cluster based on industry connectivity. Extract the coordinates of companies outside the industrial clusters and connect them to obtain the cluster boundaries identified based on company connectivity features. See [link to relevant documentation]. Figure 10 and 11 .
[0055] Step 10. Comprehensive definition of the scope of industrial clusters: An industrial cluster not only refers to a spatial agglomeration of a large number of enterprises, but also reflects the close connection between upstream and downstream industries; this invention identifies industrial clusters by combining the characteristics of enterprise agglomeration and enterprise connection. The specific steps include: (1) extracting the scope of industrial clusters obtained in steps 5 and 7, and using ArcGIS for spatial visualization; (2) using ArcGIS spatial analysis tools, calculating the weights of the industrial agglomeration degree obtained in step 7 and the industrial connection degree obtained in step 9 based on the entropy method, and using the natural discontinuity classification method in ArcGIS to update the B' 1i With D´ 1jBy performing hierarchical extraction, the scope of industrial clusters in the target region is obtained, see... Figure 12 .
[0056] Compared to existing technologies, this invention aims to focus on specific industry sectors, utilize common enterprise big data, and automatically extract industrial clusters by considering the spatial agglomeration and connections of enterprises through the structural characteristics of the upstream, midstream, and downstream of the industrial chain. It adopts interdisciplinary technologies that combine economics, geography, and urban planning to effectively analyze the industrial agglomeration characteristics of specific industry sectors, achieve accurate identification of the boundaries of industrial clusters, and make up for the lack of methods for visually representing the micro-space of specific industrial chain logic maps and the lack of methods for considering internal connections within industries.
[0057] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for identifying the boundaries of industrial clusters based on a specific industry sector, characterized in that, The method includes the following steps: Obtain industry classification data for the target region, and construct an industry statistical map based on the industry classification data for the target region; Acquire enterprise data for the target region, and construct an industry-specific enterprise dataset based on the industry statistical map and the enterprise data; Based on the enterprise dataset in the aforementioned industry sector, calculate the industrial agglomeration degree and industrial linkage degree of the target region; Based on the calculated industrial agglomeration degree and industrial linkage degree, the scope of industrial agglomeration areas in the target region is identified.
2. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 1, characterized in that, The process of constructing an industry statistical map based on the industry classification data of the target region includes: Based on the acquired industry classification data of the target region, the pre-set industry classification model is trained to obtain a trained industry classification model. The industry classification model extracts upstream, midstream, and downstream data of the industrial chain based on the industry classification data of the target region, and constructs an industry statistical map based on the upstream, midstream, and downstream data of the industrial chain. The upstream data of the industrial chain refers to data on the research, design and production of raw materials or key components; the midstream data of the industrial chain refers to data on product production, processing, assembly and manufacturing; and the downstream data of the industrial chain refers to data on sales, logistics, after-sales service and end consumers.
3. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 2, characterized in that, The industry classification model is based on the upstream, midstream, and downstream data of the industrial chain, and uses the directory mapping method and keyword semantic expansion method to construct an industry statistical map.
4. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 2, characterized in that, The process of constructing an industry-specific enterprise dataset based on the industry statistical map and the enterprise data includes: Based on the industry statistical map, the enterprise data is cleaned to remove enterprises that do not belong to the industry statistical map, thus obtaining an industry-specific enterprise dataset. Enterprise point coordinate data is obtained based on the enterprise data, and an industrial chain enterprise point dataset is constructed based on the industry sector enterprise dataset and the enterprise point coordinate data. The industrial chain enterprise point dataset includes upstream industrial chain data, midstream industrial chain data, and downstream industrial chain data.
5. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 4, characterized in that, The process of calculating the industrial agglomeration degree of a target region includes: Based on the aforementioned industry chain enterprise point dataset, all enterprises are ranked by registered capital, and a core enterprise threshold is set based on the ranking data. Enterprises that reach the core enterprise threshold from upstream, midstream, and downstream data of the industrial chain are identified as core enterprises, resulting in a number of core enterprises. Based on each core enterprise and several densely connected ordinary enterprises, a sub-industrial cluster based on industrial agglomeration degree is constructed, resulting in several sub-industrial clusters based on industrial agglomeration degree. The boundary range of each sub-industrial cluster based on industrial agglomeration degree is calculated, and the boundary range of the industrial cluster based on industrial agglomeration degree is calculated based on the boundary range of the several sub-industrial clusters based on industrial agglomeration degree. Here, the core enterprise refers to an enterprise that has reached the core enterprise threshold; the ordinary enterprise refers to an enterprise that has not reached the core enterprise threshold; the density connection refers to a point that has reached the preset density connection threshold, where both are non-core points but are in the neighborhood of a certain point.
6. The method for identifying the boundary of an industrial cluster based on an industry sector, as described in claim 5, is characterized in that... The boundary range of each sub-industrial cluster based on industry agglomeration degree is calculated using a density-based spatial clustering method. The process includes: Traverse all enterprise point data sets in the industrial chain , randomly select an enterprise point , according to the preset density connection threshold, expand outward according to the condition of density reachability. In the neighborhood formed by the preset radius, if the number of points m < Minpts, then mark the enterprise point as a noise point; if m ≥ Minpts, then mark the enterprise point as a core enterprise point, and jointly construct a sub-industrial agglomeration area with the ordinary enterprise points that are density-connected to the core enterprise point, and obtain several sub-industrial agglomeration areas based on the industrial agglomeration degree; Extract several enterprise locations on the periphery of each of the aforementioned sub-industrial clusters based on industrial agglomeration degree. To form sub-industrial clusters based on industrial agglomeration degree And based on the scope of each sub-industrial cluster based on industrial agglomeration degree. The range of industrial clusters based on industrial agglomeration degree is calculated. ; Minpts represents the minimum number of sample points for the core enterprise.
7. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 6, characterized in that, Calculate the location of each enterprise based on the industrial cluster area B, which is determined by the degree of industrial agglomeration. The profile coefficient S is obtained through a process that includes: Based on the scope of sub-industrial clusters according to industrial agglomeration degree Calculate enterprise points With the same sub-industrial cluster based on industrial agglomeration degree Let a be the average distance to other enterprise locations within the same area; and calculate the distance to that enterprise location. Let b be the average distance between the enterprise points and other sub-industrial clusters based on industrial agglomeration degree. Then the profile coefficient of the enterprise point is S=(ba) / max(a,b). Here, max(a,b) represents taking the maximum value of a and b respectively.
8. The method for identifying the boundaries of industrial clusters based on industry sectors according to claim 4, characterized in that, The process of calculating the industrial linkages of a target region includes: A directed network model is constructed based on the aforementioned industry chain enterprise point dataset; The Infomap algorithm for community discovery is used to identify the community structure within the directed network model, resulting in several sub-industry clusters based on industrial connectivity. Calculate the extent of each of the sub-industrial clusters based on industrial connectivity, and calculate the extent of the industrial cluster based on the extent of several of the sub-industrial clusters based on industrial connectivity.
9. The method for identifying the boundary of an industrial cluster based on an industry sector as described in claim 8, characterized in that, The process of constructing a directed network model based on the aforementioned industry chain enterprise point dataset includes: Using the acquired enterprises as nodes in the network, and the customer-supplier relationships between enterprises as edges, a directed network model is constructed, expressed as: Where X represents the dataset of enterprise points in the industry chain, and E represents the set of edges; The process of identifying the community structure within the directed network model and obtaining several sub-industry clusters based on industry connectivity includes: Suppose we use n codewords to represent any enterprise in the industry chain enterprise point dataset X. Probability of occurrence during community walks Given n states, the average encoding length is no less than that of any enterprise point in the industry chain enterprise point dataset X. Its own entropy is expressed as: For all businesses in the community Perform Huffman coding to represent any path in the network through a combination of codes, as shown in the expression: in, This represents the expected average coding length of the paths taken during random walks within and between communities; This represents the probability of a random walk leaving one community and entering another. Entropy represents the probability of a random walk moving between communities; This represents the probability that a random walk exists within a certain community; when When the change is less than the preset value, the scope of the closely connected sub-industry clusters is obtained. Based on several sub-industrial clusters with high industrial linkages Industrial clusters based on industrial connectivity were obtained. .
10. The method for identifying the boundary of an industrial cluster based on any one of claims 5-8, characterized in that, Based on the calculated industrial agglomeration degree and industrial linkage degree, the process of identifying the scope of industrial agglomeration areas in a target region includes: A fishing net model is constructed based on a neural network architecture, and the calculated contour coefficient S of each enterprise point is compared with the expected value of the average coding length. The values are assigned to the fishing net model and converted into raster data; The weighted raster data overlay analysis method is used to extract the value of each pixel in the fishing net model, and the weights of industrial agglomeration and industrial linkage are calculated based on the entropy method. The expression is as follows: B´ 1i =B 1i ×W i D´ 1j = D 1j ×W j Among them, B 1i B represents the scope of the sub-industrial clusters formed based on the degree of industrial agglomeration. i The value of the l-th cell within, D 1j B represents the scope of sub-industry clusters formed based on industrial connectivity. j The value of the l-th cell within, W i W represents the weight used to calculate the degree of industrial agglomeration based on the entropy method. j B' represents the weight used to calculate the degree of industrial linkage based on the entropy method. 1i With D´ 1j Each represents the updated cell value, where the cell represents any grid in the fishing net model; The natural discontinuity classification method in ArcGIS was used to update B'. 1i With D´ 1j The scope of the industrial clusters in the target region is obtained by performing hierarchical extraction.
Citation Information
Patent Citations
Industrial cluster identification method and device, storage medium and electronic equipment
CN112949914A
Regional cultural heritage spatial pattern identification method and device and storage medium
CN117971996A
Method and device for determining investment attracting enterprises in industrial park, equipment and medium
CN119621969A