Manufacturing cluster space identification method based on multi-source heterogeneous industrial big data
By identifying manufacturing cluster spaces through multi-source heterogeneous industrial big data, and classifying and correcting boundaries by combining employment population density and truck density, the problem of unreasonable boundary delineation in existing technologies has been solved, achieving more accurate cluster space identification and urban planning support.
Patent Information
- Application Number
- CN202610581288.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-29
- Publication Date
- 2026-08-25
AI Technical Summary
Existing methods for identifying manufacturing clusters rely on a single data source, leading to unreasonable boundary delineation, an inability to accurately reflect the spatial decay gradient of economic activity intensity within industrial clusters, and an inability to finely distinguish the spatial morphology of different types of clusters.
By employing multi-source heterogeneous industrial big data, we identify employment land plots to form contiguous employment areas, classify them based on employment population density and truck density, determine the type of cluster area, and use multiple data sources for boundary verification and correction to ensure the accuracy and refinement of the boundaries.
It improves the reliability and authenticity of spatial identification of manufacturing clusters, enabling more accurate reflection of the scope and type of economic activities of industrial clusters, and supporting urban planning and industrial upgrading.
Smart Images

Figure CN122634239A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of the intersection of geographic information and industrial planning, and in particular to a method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data. Background Technology
[0002] With the deepening of the global industrial chain division of labor, the high concentration of manufacturing in specific geographical spaces has become a key form of enhancing regional competitiveness. Accurately identifying and delineating the spatial boundaries of manufacturing clusters (or industrial agglomeration areas, employment centers) plays a vital foundational supporting role in carrying out industrial planning, infrastructure development, revitalization of existing land use, and environmental supervision.
[0003] In related technologies, the identification and delineation of the spatial scope of manufacturing clusters mainly adopts methods such as direct demarcation based on administrative jurisdiction boundaries, employment-oriented demarcation, or demarcation based on enterprise directories. These demarcation methods either only consider static land use attributes or only consider single personnel flows, resulting in low reliability and accuracy of the data. Furthermore, they lack a refined demarcation mechanism for the periphery of the cluster, and the boundaries are often drawn in a one-size-fits-all manner, failing to reflect the spatial attenuation gradient of the economic activity intensity of the industrial cluster.
[0004] Therefore, how to provide a manufacturing cluster spatial identification method that can integrate multi-source heterogeneous industrial big data, adaptively identify cluster types, and refine the extraction of spatial boundaries has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This invention provides a spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data, which solves the defects of existing technologies such as single data, unreasonable spatial delineation, and difficulty in truly reflecting the spatial attenuation gradient of the intensity of economic activities in industrial clusters.
[0006] This invention provides a method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data, comprising the following steps:
[0007] Identify employment land parcels based on the land parcel database and select parcels that meet preset conditions;
[0008] By spatially connecting the plots of land that meet the preset conditions, a contiguous employment area can be formed;
[0009] Multiple employment clusters are filtered according to a preset strategy to obtain employment clusters that meet the preset strategy; and the first employment population density and the first truck density of the employment clusters are obtained.
[0010] Based on the first employment population density and the first truck density, the employment clusters are classified to determine the cluster types; thus, technology-intensive clusters and labor-intensive clusters are obtained.
[0011] In response to the fact that the cluster type is a technology-intensive cluster, the second employment population density and enterprise density corresponding to the land parcels on the edge of the cluster are obtained, and the first initial boundary of the technology-intensive cluster is determined based on the second employment population density interval and the enterprise density interval.
[0012] In response to the fact that the cluster type is a labor-intensive cluster, the second truck density corresponding to the plots in the edge area of the cluster is obtained, as well as whether there are independent enterprises occupying land in the edge area of the cluster;
[0013] In response to the existence of an independent land-occupying enterprise in the edge area, a second initial boundary of the labor-intensive cluster area is determined based on the land boundary of the independent land-occupying enterprise; otherwise, the second initial boundary is determined based on the second truck density.
[0014] Based on the spatial attribute data of each plot in the edge region of the first initial boundary, the first initial boundary is checked and corrected to obtain the first determined boundary.
[0015] Based on the spatial attribute data of each plot in the edge region of the second initial boundary, the second initial boundary is checked and corrected to obtain the second determined boundary.
[0016] The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data provided by the present invention identifies employment land plots based on a land plot database and selects plots that meet preset conditions, specifically including the following steps:
[0017] In the pre-configured land database, land parcels that meet the criteria are selected based on whether they are commercial, residential, industrial, or warehousing land.
[0018] The spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data provided by the present invention spatially connects the plots of land that meet the preset conditions, including the following specific steps:
[0019] Based on a predetermined distance extended outward from the boundary of each employment land plot, the plots are spatially connected to form an employment cluster.
[0020] According to the spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data provided by the present invention, the preset distance is 13 meters to 22 meters.
[0021] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, multiple contiguous employment areas are screened according to a preset strategy to obtain employment clusters that meet the preset strategy. Specifically, the method includes the following steps:
[0022] Multiple employment clusters are filtered based on the following criteria: employment area greater than or equal to 30 hectares, employment cluster located within a pre-configured industrial geofence data set, and employment population density greater than 0.5 million people / square kilometer. The resulting clusters are considered employment agglomeration areas.
[0023] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, both the first truck density and the second truck density are obtained based on satellite positioning trajectory data through a stop point identification algorithm; specifically including:
[0024] The spatiotemporal trajectories of all mobile terminals entering the employment cluster area are obtained; target terminals that meet the preset characteristics of freight vehicle movement parameters are filtered; the stopping points of the target terminals in the employment cluster area are identified; and the first truck density is generated based on the ratio of the number of stopping points to the land area of the employment cluster area.
[0025] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, the first employment population density and the second employment population density are obtained based on mobile phone signaling data of operators. The identification rule is: within the statistical period, if the cumulative number of days that a user terminal generates residence signaling in the plot exceeds a preset threshold during a preset time period on weekdays, then the user is marked as a regular residence user of the plot.
[0026] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, the determination of the first defined boundary specifically includes the following methods: determining the first defined boundary based on the occupancy ratio of each plot in the edge region of the first initial boundary; or determining the first defined boundary based on real-time remote sensing image data of each plot in the edge region of the first initial boundary; or determining the first defined boundary based on on-site survey and annotation data of each plot in the edge region of the first initial boundary.
[0027] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, the determination of the second determined boundary specifically includes the following methods: determining the second determined boundary based on the occupancy ratio of each plot in the edge area of the second initial boundary; or determining the second determined boundary based on real-time remote sensing image data of each plot in the edge area of the second initial boundary; or determining the second determined boundary based on on-site survey and annotation data of each plot in the edge area of the second initial boundary.
[0028] According to the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention, the employment cluster area is classified and the cluster type is determined based on the first employment population density and the first truck density. Specifically, the method includes the following steps:
[0029] Based on the first employment population density and the first truck density, the employment cluster area is initially classified; wherein, if the first employment population density of the employment cluster area is greater than a first preset density threshold and the first truck density is less than or equal to the first preset truck threshold, it is initially classified as a technology-intensive cluster area; if the first truck density of the employment cluster area is greater than the first preset truck threshold, it is initially classified as a labor-intensive cluster area.
[0030] Obtain auxiliary data on the industry attributes of the employment cluster; the auxiliary data on industry attributes includes at least one of the following: industry category distribution characteristics extracted based on enterprise business registration information, average plot ratio or building density extracted based on land database, and the total number of regularly resident users in the employment cluster.
[0031] The preliminary classification results are verified and corrected based on the industry attribute auxiliary data to determine the final cluster type.
[0032] This invention provides a spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data. This method systematically integrates land use data, mobile phone signaling data, truck trajectory data, and remote sensing imagery data. This allows the identified cluster boundaries to reflect both the physical land use attributes and the actual intensity of human and logistical activities. Compared to traditional single-data source methods, the identification results of this invention significantly improve the consistency with the actual industrial activity range, effectively enhancing the reliability and accuracy of cluster delineation, improving the rationality of the delineation, and providing reliable boundary data for urban planning. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0034] Figure 1 This is a flowchart illustrating the spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data provided by the present invention.
[0035] Figure 2 This is a logical block diagram of the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data provided by the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0037] In related technologies, the following methods are typically used to identify and delineate the spatial extent of manufacturing clusters:
[0038] The boundary substitution method based on administrative jurisdiction directly uses the approved red line map of the development zone, high-tech zone, or industrial park as the cluster boundary. This is a static, policy-based division method. On the one hand, the land within the approved scope may not have been actually developed, leaving a large amount of vacant land; on the other hand, the spillover effect of industrial clusters often exceeds the administrative red line, forming a large-scale cluster of supporting enterprises in adjacent urban villages or towns, and the administrative boundary cannot perceive this spillover of economic activity.
[0039] The manual visual interpretation method based on remote sensing imagery relies on planners manually delineating cluster outlines by interpreting visual features such as blue-roofed factory buildings, large open-air storage yards, or orderly road networks using high-resolution satellite imagery. This method is highly subjective, inefficient, and cannot be updated frequently.
[0040] The point density analysis method based on a single enterprise directory uses business registration data or tax registration addresses to plot the latitude and longitude of enterprises on a map, generating a heat map through kernel density analysis, and using a certain contour line as the cluster boundary. In this approach, many manufacturing enterprises are registered at virtual addresses in development zone management committee buildings or tax havens, while their actual production workshops are located on industrial sites tens of kilometers away. Simply relying on enterprise registration coordinates causes the cluster center to drift towards administrative office areas, rather than the actual production operation area. Furthermore, this method cannot distinguish the spatial boundary differences between technology-intensive R&D enterprises and labor-intensive processing plants.
[0041] The aforementioned technologies all suffer from drawbacks such as limited data, coarse boundary definitions, and lack of classification, necessitating a more accurate method for identifying and delineating the spatial boundaries of industrial clusters.
[0042] In related technologies, the delineation of manufacturing cluster boundaries relies on subjective experience or a single administrative jurisdiction, leading to an inaccurate reflection of the true economic activity scope of industrial clusters and an inability to distinguish the spatial forms of different cluster types. This application provides a spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data. By spatially connecting multiple scattered employment land parcels to form contiguous areas, and then filtering these contiguous areas according to a preset strategy, employment clusters are identified. The cluster types are classified based on the highest employment population density and the highest truck density within each cluster, and then boundary delineation is achieved using corresponding methods based on the classified types. In other words, this application first selects a large area and classifies the clusters accordingly. Then, specific boundary delineation methods are selected for different types of clusters, achieving a concrete division and definition of the edge boundaries of different clusters. This makes the overall division more realistic and the reliability of the data more high.
[0043] The spatial identification method for manufacturing clusters based on multi-source heterogeneous industrial big data provided by this invention belongs to the interdisciplinary field of geographic information science and industrial planning. Its application scenarios are extremely wide, and it can be deployed in various computing platforms that require industrial spatial monitoring, planning evaluation, and existing land use management. The following non-exhaustive examples illustrate the technical solution of this invention in conjunction with specific application scenarios.
[0044] The method described in this application can be deployed on a national land and space information platform, and the system can automatically perform the following tasks:
[0045] Automatic boundary extraction: The system automatically acquires the latest land parcel database, mobile phone signaling data, and truck trajectory data annually or quarterly to generate the actual industrial cluster boundaries for each development zone in batches. Based on the identified first or second defined boundaries, it accurately calculates core evaluation indicators such as land use intensity, employment population density, and logistics activity within the cluster, eliminating subjective biases caused by manual reporting.
[0046] Idle land identification: The cluster boundary generated by this invention is spatially overlaid with the approval red line of the development zone. Contiguous plots of land located within the approval red line but not included in the cluster boundary by this invention for a long period (e.g., two consecutive quarters) are automatically marked as "suspected inefficient land use" or "approved but not supplied land" by the system and pushed to regulatory authorities for on-site verification.
[0047] It can also be used for assessing the redevelopment potential and spatial identification of inefficient industrial land in cities. With the advancement of urban renewal, local governments urgently need to accurately identify "inefficient land" plots from existing industrial land that are scattered, have outdated production capacity, are actually vacant, or have been converted to non-industrial uses, in order to serve as key targets for demolition, relocation, or industrial upgrading. This approach changes the traditional passive work mode that relies on manual investigation of inefficient land, enabling proactive discovery, dynamic early warning, and precise profiling based on big data, providing clear geographical target areas for urban renewal and industrial upgrading.
[0048] It can also be used for constructing spatial maps of industrial and supply chains and analyzing strengthening and supplementing these chains. Specifically, based on identifying macro-level employment clusters using the method of this invention, industry codes from enterprise registration or patent data can be further overlaid. For example, the enterprise density in the identification algorithm can be replaced with the enterprise density of a specific industry (such as "integrated circuit manufacturing"), thereby cutting out exclusive spatial clusters of specific industrial chains from the macro-level clusters. Utilizing the classification capabilities of this application for technology-intensive and labor-intensive clusters, the system analyzes whether the R&D and design stages (technology-intensive) and production and assembly stages (labor-intensive) within the same industrial chain exhibit a spatially adjacent "industry-city integration" layout, or a pendulum-like separation where R&D is located in the central urban area and production is in the suburbs. If excessive separation is identified, the system can quantitatively assess the resulting additional commuting and freight transport carbon emissions.
[0049] This method overlays data on the boundaries of technology-intensive clusters with those of commercial and residential land. If, within a predetermined distance (e.g., 500 meters) beyond the boundary of a large technology-intensive park, the proportion of identified commercial and residential land plots is extremely low, the system automatically marks the cluster as having a problem of imbalance between jobs and residence or a serious lack of supporting living facilities, providing a quantitative basis for subsequent industry-city integration planning. In other words, this method can extend the identification of macro-level industrial cluster boundaries down to the spatial organization level of the industrial chain, providing local governments with a high-precision spatial analysis tool for accurately mapping the spatial distribution of the industrial chain and identifying supply chain breaks and supporting shortcomings.
[0050] In specific settings, such as Figure 1 , Figure 2 As shown, this embodiment provides a method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data, including the following steps:
[0051] Step S10: Identify employment land parcels based on the land parcel database and select parcels that meet the preset conditions.
[0052] Specifically, the system has a pre-configured local land parcel database. Upon system initialization, the system loads this pre-configured database. This database can be a shapefile or GeoJSON format geographic information vector file, with each parcel polygon containing attribute fields such as land use type, area, and parcel boundaries. The system iterates through all parcels, extracting land patches closely related to manufacturing employment activities based on preset land use type filtering criteria, forming a set of employment-related land parcels.
[0053] In the specific implementation of the screening process, commercial land, residential land, industrial land, and warehousing land are used as screening criteria in a pre-configured land database to select eligible plots. If only industrial land is used as the screening criterion for employment land, it will systematically miss a large number of supporting commercial facilities (such as park canteens, convenience stores, and financial institution branches), supporting residential facilities (such as shift workers' dormitories and employee apartments), and ancillary warehousing facilities within the manufacturing cluster. Although these land uses are not classified as industrial land in the land survey, they actually bear the indispensable functions of accommodating the cluster's employment population and supporting logistics. This embodiment, by including commercial, residential, warehousing, and industrial land in the screening scope, constructs a lower-level data layer that better reflects the real ecosystem of the manufacturing cluster, avoiding the fragmented cluster spatial identification problem caused by the singular nature of land use. This allows the subsequently calculated employment population density and truck density to more comprehensively reflect the area's overall carrying capacity and economic activity intensity.
[0054] Specifically, the pre-configured land parcel database stores a unique identifier and land use type code for each parcel. During filtering, the system reads the "Land Use Type" or "Land Category Code" field from the parcel attribute table. The mapping rules for filtering conditions are as follows, depending on the data standards:
[0055] If the "Classification of Urban Land Use and Standards for Planning and Construction Land" is adopted, the screening criteria are: commercial land (Category B), residential land (Category R, especially the staff dormitory part of R2 second-class residential land), industrial land (Category M), and warehousing land (Category W).
[0056] If the "Third National Land Survey Classification" is adopted, the screening criteria are: commercial and service land (05), residential land (07), industrial and mining land (06), and warehousing land (08).
[0057] The system uses standard SQL query statements or GIS attribute filtering tools to mark the land use type codes that fall into the above set as meeting the preset conditions, and extracts their geometric polygon data for the next processing stage.
[0058] Step S20: Spatially connect the plots of land that meet the preset conditions to form a contiguous employment area.
[0059] Specifically, the system can extend outward by a preset distance based on the boundary of each employment land plot, so that the plots are spatially connected to form an employment cluster.
[0060] In real-world geographical environments, two physically adjacent factory sites, functionally belonging to the same industrial chain, are often separated by an internal road, patrol route, or narrow public passageway. If topological union operations are performed solely based on the original site boundaries, these road separations can fragment what should be a unified cluster into multiple independent pieces, severely impacting the accuracy of subsequent area calculations and density statistics for cluster selection. This embodiment introduces a GIS buffer expansion algorithm to micro-scale outward expansion of the sites before spatial connection, simulating the actual state of pedestrian and vehicle crossing of side roads for daily production activities. This approach eliminates interference from internal side roads while effectively preserving the role of large obstacles such as main urban roads and rivers as natural boundaries for the clusters.
[0061] In other words, after selecting scattered plots that meet the screening criteria, the system calls the buffer function in the GIS spatial analysis engine to generate an equidistant line area with a preset width along the boundary normal direction for each employment land plot. Next, spatial overlay analysis (Union or Intersect) is performed on all plots with generated buffers, merging (Dissolve) the polygons of plots with overlapping or intersecting buffers into a single polygon feature, forming a preliminary contiguous employment area. The purpose of this step is to eliminate the interference of narrow physical barriers such as urban side roads, internal patrol roads, and rivers on the continuity of the plots.
[0062] In executing this step, the system calls the Buffer function interface in the spatial analysis engine. For an input employment land plot polygon, the system calculates its boundary line and generates an equidistant line region with constant bandwidth outward along the boundary normal direction. This bandwidth value is the preset distance. Subsequently, the system performs a spatial union operation on the buffer layers generated for all plots, identifying all buffer polygons with spatial overlap or shared edges, and merging the original plots corresponding to interconnected buffers under the same employment contiguous area identifier. Finally, the system performs aggregation and dissolve on the original plots within each contiguous area, generating a unified outer envelope polygon for that contiguous area.
[0063] The preset distance of the buffer zone directly determines which barriers are crossed and which are retained. In this embodiment, after extensive field sample testing and statistical analysis, a specific value range of 13 meters to 22 meters was determined to be the optimal range. The technical significance of this distance is that it is just larger than the planned right-of-way width of typical internal roads in industrial parks or urban branch roads (usually 12 to 15 meters), but significantly smaller than the width of urban secondary arterial roads (usually 25 to 40 meters) and main roads.
[0064] Specifically, in a typical application scenario, if the road between two plots is a 10-meter-wide internal patrol road, and the preset distance is set to 16 meters, each plot will extend outward by 16 meters (totaling 32 meters). The buffer zone will form a large overlapping area above the center of the road, thus triggering the plot merger. Conversely, if the road between the two plots is a 60-meter-wide urban expressway, even if each side extends outward by 15 meters, there will still be a 32-meter gap in the middle. The buffer zones will not intersect, and the plots will remain separated.
[0065] In a preferred embodiment, the preset distance is set to 15 meters or 20 meters, which can take into account the road cross-sectional characteristics of most new and old industrial areas in cities. The system can store this value as a configurable parameter in a configuration file, allowing users to fine-tune it according to the road planning standards of different cities.
[0066] Step S30: Select multiple contiguous employment areas according to a preset strategy to obtain employment clusters that meet the preset strategy; and obtain the first employment population density and the first truck density of the employment clusters.
[0067] The first two steps enable the fragmented plots of land to be connected into contiguous areas. In this step, clustered areas are obtained through further filtering, and the data of these clustered areas is used for subsequent type classification.
[0068] Further screening of multiple contiguous employment areas involves the following steps: These areas are screened based on the following criteria: an employment area greater than or equal to 30 hectares, location within a pre-configured industrial geofence dataset, and an employment population density greater than 5,000 people per square kilometer. The resulting areas are designated as employment clusters. For clusters formed by spatially connecting multiple employment land parcels, using all contiguous employment areas as analysis objects without further screening would significantly increase computational costs and result in a large amount of noisy data in the output clusters that lack genuine industrial agglomeration effects. This embodiment establishes a triple-constraint filter based on area, spatial location, and activity density, forming a funnel-shaped screening mechanism to ensure that the ultimately retained employment clusters possess significant scale effects, policy location rationality, and actual economic activity intensity.
[0069] The specific screening process is as follows:
[0070] Firstly, area screening. The system calculates the ellipsoidal area (in hectares) of the contiguous employment area polygon. Only areas with an area value greater than or equal to 30 hectares pass this screening.
[0071] Secondly, geofence filtering. The pre-configured industrial geofence dataset refers to the approved boundary vector data of various development zones, high-tech zones, and industrial clusters, stored in spatial database tables or GeoJSON format. The system uses spatial relationship operation functions (such as ST_Intersects or ST_Within) to determine whether the geometric center point of the contiguous employment area falls inside any geofence polygon; or, to determine whether the overlap area between the contiguous area and the geofence exceeds a preset ratio (such as 70%). Contiguous areas that meet the conditions pass this filtering. This eliminates industrial enclaves located in ecological protection zones, basic farmland protection zones, or purely residential areas.
[0072] Third, employment population density screening. The system retrieves the number of regularly active users calculated based on mobile signaling data, divides it by the area of the contiguous zone, and obtains the employment population density value. Only when the employment population density is greater than 0.5 million people / square kilometer (i.e., 50 people / hectare) does the contiguous zone pass this screening. This threshold effectively excludes invalid areas that have factory buildings but are actually vacant or only have a few security guards stationed there. In other words, for a contiguous zone with an area of exactly 30 hectares, as long as the number of regularly active users identified reaches more than 1,500, this condition is met.
[0073] In practice, the system iterates through all the employment clusters generated in the first phase and applies a pre-defined triple-constraint filter for selection: it calculates the ellipsoidal area of the clusters, eliminating scattered clusters with an area less than 30 hectares; it determines whether the geometric center point of the cluster falls within a pre-configured industrial geofence dataset (i.e., the approved boundary vector map of national and provincial development zones); and it calculates the density of regularly resident users within the clusters, eliminating low-activity areas with a density below 0.5 million people per square kilometer. Clusters that pass the selection process are marked as employment clusters.
[0074] For each employment cluster, the system obtains its first employment population density and first truck density. The first employment population density is calculated based on mobile phone signaling data from operators by identifying the number of users who regularly stay for more than 18 days per month during the weekday period from 10:00 to 16:00; the first truck density is calculated based on truck trajectory data by identifying the number of stops of freight vehicles within the plots of land in the cluster.
[0075] The first truck density and the second truck density are both obtained based on satellite positioning trajectory data and through a stop point identification algorithm. Specifically, this includes: acquiring the spatiotemporal trajectories of all mobile terminals entering the employment cluster area; screening target terminals that meet the preset characteristics of freight vehicle movement parameters; identifying the stop points of the target terminals in the employment cluster area; and generating the first truck density based on the ratio of the number of stop points to the land area of the employment cluster area.
[0076] Specifically, the algorithm for calculating truck density can be obtained using existing algorithms in related technologies. The following is an explanation of the specific algorithm for calculating truck density:
[0077] Step S301, Data Acquisition and Preprocessing: The system acquires anonymized mobile communication signaling data or satellite navigation positioning trajectory data from the data service interface. Each data record contains at least: anonymized device ID, timestamp, latitude and longitude coordinates, and instantaneous speed. The system cleans the raw data, removing drift points, duplicate points, and timestamp anomalies.
[0078] Step S302, Freight Vehicle Terminal Identification: There are significant statistical differences in the movement behavior patterns of ordinary passenger vehicles and freight vehicles. This embodiment constructs a target terminal screening model based on a rule engine or machine learning classifier. Preset movement parameter characteristics of freight vehicles include, but are not limited to, at least one of the following: the average daily mileage is significantly higher than that of ordinary commuter vehicles (e.g., >200 km / day); the proportion of driving time during nighttime (22:00-06:00) to the total driving time is higher than a preset value; after road network matching, the trajectory points frequently fall on national highways, provincial highways, expressways, and urban freight corridors, rather than urban micro-circulation roads; the average speed curve of a single trip exhibits typical intercity transportation characteristics.
[0079] Step S303, Dwelling Point Identification: For the marked target terminal, the system performs spatial intersection analysis between its trajectory point sequence and the polygon of the employment cluster area. When multiple consecutive trajectory points of a terminal are located inside the same polygon, and the instantaneous speed of these trajectory points is continuously less than a preset stationary speed threshold (e.g., <5km / h), and the duration of this state exceeds a preset dwell time threshold, a valid dwelling point event is determined to have occurred. Unlike the short-term stop identification of ride-hailing vehicles and taxis in conventional traffic surveys (usually with a threshold of 3-15 minutes), this embodiment, targeting the characteristics of loading and unloading operations in manufacturing logistics, preferably sets the dwell time threshold to 60 to 180 minutes. This longer time window can effectively eliminate interference events such as temporary stops by passing vehicles, commuter bus pick-up and drop-off, and ride-hailing delivery, accurately capturing real loading, unloading, and waiting operations.
[0080] Step S304, Density Aggregation Calculation: The system aggregates dwell point events according to a preset statistical period (e.g., one calendar month). The total number N of dwell points occurring within the polygonal area of the target employment cluster within the statistical period is counted. Then, the first truck density = N / cluster land area. To calculate the second truck density, the statistical scope is limited to the corresponding edge area plots; the remaining steps are the same.
[0081] Step S40: Based on the first employment population density and the first truck density, classify the employment clusters to determine the cluster types; thus, obtain technology-intensive clusters and labor-intensive clusters. Classifying the clusters allows for better boundary determination. The specific methods for obtaining employment population density and truck density are detailed below.
[0082] The specific implementation includes the following steps: Firstly classifying employment clusters based on a first employment population density and a first truck density; wherein, if the first employment population density of an employment cluster is greater than a first preset density threshold, and the first truck density is less than or equal to the first preset truck threshold, it is initially classified as a technology-intensive cluster; if the first truck density of an employment cluster is greater than the first preset truck threshold, it is initially classified as a labor-intensive cluster; acquiring auxiliary data on the industry attributes of the employment clusters; the auxiliary data on industry attributes includes at least one of the following: industry category distribution characteristics extracted based on enterprise registration information, average plot ratio or building density extracted based on a land database, and the total number of regularly resident users within the employment clusters; verifying and correcting the preliminary classification results based on the auxiliary data on industry attributes to determine the final cluster type.
[0083] In the fields of industrial economics and urban planning, there are significant differences in the physical spatial representation between technology-intensive manufacturing (such as electronics and information, biomedicine, and precision instruments) and labor-intensive manufacturing (such as textiles and apparel, metal products, and furniture manufacturing). The former is typically characterized by high-rise, high-density factory buildings or R&D buildings, a large flow of commuters, and extremely low frequency of freight vehicle traffic; the latter is characterized by single-story, large-span workshops, relatively sparse workforce, and extremely high frequency of heavy truck loading and unloading activities.
[0084] This embodiment introduces a verification and correction mechanism for auxiliary data on industry attributes, building upon the initial density classification. By integrating multi-dimensional auxiliary data such as enterprise registration information, building morphology characteristics, and employment population size, the initial classification results are reconfirmed or intelligently corrected, significantly improving the accuracy and robustness of the classification and ensuring that the final output cluster type truly reflects the essence of the region's economic activities.
[0085] Specifically, the specific values of the first preset density threshold and the first preset truck threshold need to be calibrated in conjunction with the regional industrial characteristics. In a typical application example of the present invention, after statistical analysis of known types of development zone samples, preferably: the first preset density threshold (employment population density) is set to 10,000 people / square kilometer (i.e., 100 people / hectare), and the first preset truck threshold (truck density) is set to 30,000 truck trips / square kilometer / month (i.e., 300 truck trips / hectare / month, approximately 10 trucks stopping per hectare per day).
[0086] During operation, the system compares the first employment population density with the first preset density threshold (e.g., 10,000 people / square kilometer) and the first truck density with the first preset truck threshold (e.g., 30,000 truck trips / square kilometer / month). Based on the comparison results, the cluster is classified as a technology-intensive cluster or a labor-intensive cluster.
[0087] Specifically, when making classification decisions:
[0088] Scenario A (Technology-Intensive Cluster Determination): If the average monthly primary employment population density of a certain cluster is 15,000 people / km² (greater than 1.0), and the average monthly primary freight truck density is 10,000 truck trips / km² / month (less than 3.0), the system determines that the area has a high density of employed people, but very few large freight vehicles enter or leave. Its spatial behavior characteristics match the typical profile of a research and development center, software park, headquarters base, or precision manufacturing workshop. Therefore, the output type label is "technology-intensive cluster."
[0089] Scenario B (Labor-intensive cluster determination): If the average monthly primary employment population density of a certain cluster is 0.4 million people / km² (less than 1.0), but the average monthly primary freight truck density is 5.0 million truck trips / km² / month (greater than 3.0), the system determines that the population density in this area is relatively sparse, but freight logistics are extremely busy. The spatial behavior characteristics match the typical profile of large-scale equipment manufacturing, metal processing, textile printing and dyeing, or logistics warehousing bases. Therefore, the output type label is "labor-intensive cluster".
[0090] Scenario C (Comprehensive Judgment): If the average monthly primary employment population density of a certain cluster is 18,000 people / km² (greater than 1.0), and the average monthly primary truck density is 40,000 truck trips / km² / month (greater than 3.0), according to the decision logic of this embodiment, because it meets the condition that the primary truck density is greater than the first preset truck threshold, it is judged as a labor-intensive cluster. The technical rationale for this rule is that the influence weight of high-intensity truck logistics activities on the spatial boundary morphology of the park (such as the need for larger turning radii, more loading and unloading areas, and more fragmented edge morphology) is set higher than the influence of high employment population density in this algorithm. The frequent entry and exit of a large number of trucks will force the park boundary to extend more along the traffic lines and have a more irregular shape, making the labor-intensive truck density demarcation strategy more accurate.
[0091] To avoid potential misjudgments in the initial classification, the system further retrieves auxiliary data on industry attributes for verification. This auxiliary data on industry attributes includes at least one of the following:
[0092] The system extracts industry category distribution characteristics (productivity factors and industry types) based on enterprise registration information. Specifically, it retrieves the business registration information of all registered enterprises within the geographical area of the employment cluster through a data interface and extracts their codes from the National Industrial Classification of Economic Activities (GB / T 4754). The system has a pre-set code library for technology-intensive industries (such as C39 Computer, Communication and Other Electronic Equipment Manufacturing, C27 Pharmaceutical Manufacturing, C40 Instrument and Meter Manufacturing, etc.) and a code library for labor-intensive industries (such as C13 Agricultural and Sideline Food Processing, C17 Textile Industry, C18 Textile and Apparel Industry, C21 Furniture Manufacturing, etc.). The system calculates the percentage of enterprises belonging to technology-intensive industries (P_tech) and the percentage of enterprises belonging to labor-intensive industries (P_labor) within the cluster. If an enterprise is initially classified as technology-intensive, but P_labor is significantly higher than P_tech (e.g., P_labor > 60%), a verification and correction process is triggered.
[0093] The average plot ratio or building density (land type) is extracted from the land parcel database. Specifically, the system retrieves the plot ratio or building footprint area field from the land parcel database to calculate the average plot ratio or average building density of all parcels within the employment cluster area. Generally, the average plot ratio of technology-intensive parks (especially multi-story standard factory buildings and R&D buildings) is usually between 1.2 and 2.5, and the building density is between 30% and 45%; while the average plot ratio of labor-intensive parks (especially single-story large-span workshops and open-air storage yards) is usually between 0.5 and 1.0, and the building density is between 20% and 35%. If the initial classification is technology-intensive, but the average plot ratio of the cluster is lower than 1.0, it indicates that the actual building form is more inclined towards single-story factory buildings, which may trigger a verification and correction.
[0094] The system calculates the total number of regularly resident users within an employment cluster (based on the size of the employed population). Specifically, it counts the total number of regularly resident users within the cluster after deduplication. For clusters of similar size, the total number of resident users in technology-intensive parks (with high population density) is typically significantly higher than that in labor-intensive parks. The system can preset a threshold for the total number of resident users after area normalization. If a cluster is initially classified as labor-intensive, but the total number of resident users is abnormally high (e.g., exceeding the typical value for technology-intensive parks of the same size), a verification and correction may be triggered.
[0095] The system pre-sets weighting coefficients for each type of auxiliary data. In one embodiment, the weight of industry category characteristics is set to 0.5, the weight of plot ratio characteristics is set to 0.3, and the weight of total number of resident users is set to 0.2. The system calculates the consistency score between the preliminary classification result and the results indicated by the auxiliary data. If the consistency score is lower than a preset threshold (e.g., 0.6), the system will issue a classification conflict warning and adjust the classification result according to preset correction rules. For example, if the preliminary classification is technology-intensive, but industry data shows that 80% of the enterprises in the area are labor-intensive textile and apparel enterprises, and the average plot ratio is only 0.8, the system will automatically correct the final classification to a labor-intensive cluster area. Conversely, if the preliminary classification is labor-intensive, but industry data shows that the area is a large electronic equipment assembly plant (a technology-intensive industry), and the total number of regularly resident users is abnormally high, the system can correct it to a technology-intensive cluster area.
[0096] Understandably, through the aforementioned quantitative threshold segmentation and clear decision-making logic, this method transforms the traditionally difficult-to-quantify concept of industry type into two spatial behavioral characteristic indicators that can be automatically calculated and objectively compared. This enables automated, non-intrusive, and highly up-to-date classification and identification of manufacturing cluster types, providing a scientific and reliable basis for differentiated boundary delineation strategies.
[0097] The following section provides practical application examples of the manufacturing cluster spatial identification method based on multi-source heterogeneous industrial big data.
[0098] Classification table of specific cluster areas:
[0099]
[0100] As can be seen from the table above, each major category is further subdivided. Specifically, the classification (i.e., the subcategories under technology-intensive and labor-intensive industries) is based primarily on four dimensions: production factors, industry type, land use type, and population size.
[0101] Based on the screening and classification results of steps S10-S40, the system adopts a differentiated edge determination mechanism to achieve the final determination of the edges.
[0102] Step S50: In response to the cluster type being a technology-intensive cluster, obtain the second employment population density and enterprise density corresponding to the land parcels on the edge of the cluster, and determine the first initial boundary of the technology-intensive cluster based on the second employment population density interval and the enterprise density interval.
[0103] When a cluster passes the screening and is identified as a technology-intensive cluster, the first initial boundary is determined using step S50. That is, the system extracts the edge area plots of the cluster. Edge area plots are defined as plots adjacent to non-employment land outside the cluster or located within a preset width buffer zone inside the convex hull boundary of the cluster. For these edge plots, the system obtains their second employment population density (calculated according to the same rules as the first employment population density, but the statistical scope is limited to the edge plots) and enterprise density (i.e., the number of registered corporate entities per unit area). The system presets second employment population density intervals (e.g., divided into first, second, and third categories) and enterprise density intervals, and determines whether edge plots should be included in the cluster core area by looking up tables or using weighted scoring. If the second employment population density of an edge plot falls into the first category interval and the enterprise density falls into the second category or higher interval, the plot is retained; otherwise, the plot is removed from the initial boundary, thus forming the first initial boundary.
[0104] The following table is the pre-estimation calculation table:
[0105]
[0106] In the table above, within technology-intensive clusters, areas with an employment population density of 20,000-130,000 people / km² and an enterprise density of 10-130 enterprises / km² are classified as Category I (high-end employment population density); areas with an employment population density of 10,000-80,000 people / km² and an enterprise density of less than 100 enterprises / km² are classified as Category II (medium-end employment population density); and areas with an employment population density of 10,000-12,000 people / km² and an enterprise density of less than 30 enterprises / km² are classified as Category III (low-end employment population density).
[0107] Step S51: Based on the spatial attribute data of each plot in the edge region of the first initial boundary, the first initial boundary is checked and corrected to obtain the first determined boundary. After step S50, an initial boundary can be obtained for technology-intensive clusters. Since the initial boundary generated based on statistical data may not completely match the actual physical features, the system introduces the spatial attribute data of the edge plots for secondary verification and obtains the corrected determined boundary.
[0108] The spatial attribute data includes: the proportion of the building footprint area within the plot to the total area of the plot, real-time high-resolution remote sensing image data, or manual on-site survey and annotation data.
[0109] In specific embodiments, the first defined boundary is determined based on the occupancy ratio of each plot in the first initial boundary edge region, or based on real-time remote sensing image data of each plot in the first initial boundary edge region, or based on on-site survey and annotation data of each plot in the first initial boundary edge region.
[0110] In the above embodiments, the three verification and correction methods can be selected and executed by the system based on data availability, or combined and executed in a preset priority order:
[0111] Method 1: Verification based on occupancy ratio. The system retrieves the building footprint vector data of each plot in the edge area. It calculates the ratio of the sum of the building footprint areas to the total land area of each edge plot, i.e., the building density. If the initial boundary passes through an edge plot A, and the calculation shows that the building density of plot A is lower than a preset threshold (e.g., <15%), the system determines that the plot is mainly an open storage yard, internal green space, vacant land awaiting construction, or road plaza, rather than the area where core production buildings are located. Accordingly, the system will perform an inward contraction operation on the initial boundary, snapping the boundary line to the nearest adjacent plot boundary with a building density higher than the threshold.
[0112] Method 2: Verification based on remote sensing imagery. The system retrieves the latest high-resolution satellite remote sensing imagery or aerial orthophotos of the edge area (such as WorldView-3, Gaofen-2, and Google Earth imagery) via an application programming interface (API). The system can use pre-trained deep learning semantic segmentation models (such as U-Net and DeepLabV3+) to classify land features in the imagery, identifying land feature types such as factory roofs (usually regular blue or gray rectangles), container yards (a mottled array of rectangles), water bodies, farmland, and bare soil. If the plots crossed by the first initial boundary are classified as farmland, water bodies, or large areas of bare soil in the image interpretation results, the system will perform cropping and edge smoothing on the initial vector boundary based on the texture edges or plot boundaries of these natural features, excluding the natural features from the cluster boundary.
[0113] Method 3: Verification based on on-site survey and annotation data. The system provides a human-computer interaction interface, pushing an electronic map or remote sensing base map overlaid with the first initial boundary to the mobile terminal of the field verification personnel. Verification personnel carry their terminals to the edge plots. If they find a location within the initial boundary but isolated by a physical wall, a relocated business, or a change in function to non-industrial use, they can delineate the exclusion area or mark corrections on the terminal interface. The system receives this on-site survey and annotation data via a wireless network and updates, trims, or expands the initial boundary accordingly.
[0114] In this embodiment, whether using occupancy ratio, remote sensing imagery, or on-site labeled data, the positioning error of the density algorithm can be effectively eliminated, greatly improving the engineering usability and timeliness of the final output first defined boundary in actual business scenarios such as land spatial planning and existing land use verification.
[0115] Step S60: In response to the cluster type being a labor-intensive cluster, obtain the second truck density corresponding to the land parcels in the edge area of the cluster, and whether there are independent land-occupying enterprises in the edge area of the cluster; In response to the existence of independent land-occupying enterprises in the edge area land parcels, determine the second initial boundary of the labor-intensive cluster based on the land boundary of the independent land-occupying enterprise; otherwise, determine the second initial boundary based on the second truck density. Similar to determining it as a technology-intensive cluster, the specific details are as follows:
[0116] The system also extracts edge plots within the cluster. For these edge plots, the system obtains their second truck density. Simultaneously, the system queries the independent land-occupying enterprise attribute field in the plot database to determine if an independent land-occupying enterprise exists on the edge plot. The criteria for determining an independent land-occupying enterprise include: the land use is industrial or warehousing; only a single legal entity's registered address exists within the plot; and the plot boundary has a physical wall. If an independent land-occupying enterprise exists on an edge plot, the system directly reads the corresponding land boundary polygon and uses it as the boundary of that local area; otherwise, the system determines the boundary based on the attenuation gradient of the second truck density. That is, it checks plots outward from the core of the cluster. When the second truck density of an edge plot is lower than a preset attenuation threshold (e.g., 30% of the core area density), the boundary of that plot closer to the core area is cut to form a second initial boundary.
[0117] Step S61: Based on the spatial attribute data of each plot in the edge region of the second initial boundary, the second initial boundary is checked and corrected to obtain the second determined boundary.
[0118] For labor-intensive clusters, the second initial boundary is determined based on a combination of the second truck density attenuation gradient and the boundaries of independently occupied enterprise land parcels. However, truck trajectory data also suffers from positioning drift issues, and large truck loading and unloading activities sometimes extend to the side of public roads outside the industrial park or to temporarily rented village collective land. Whether these areas should be included in the cluster scope for administrative management and spatial statistics needs to be determined in conjunction with the on-site physical morphology. This embodiment applies the aforementioned three types of spatial attribute data verification mechanisms equally to the second initial boundary of labor-intensive clusters, ensuring that the final boundary outputs of different types of clusters achieve the same accuracy standards and on-site fit.
[0119] Specifically, considering the characteristics of labor-intensive clusters, the verification and correction logic has the following targeted adjustments:
[0120] Adjustments to the occupancy ratio verification: Large open-air storage yards or parking lots often exist on the edges of labor-intensive industrial parks. While their building density may be low, they are the core spaces for logistics activities. Therefore, in this approach, the system not only calculates the building footprint ratio but also the hardened surface ratio—that is, the proportion of non-permeable surfaces such as cement and asphalt identified through remote sensing image spectral analysis of the plot's area. If the building density is low but the hardened surface ratio is high (e.g., >50%), combined with the characteristic of a high second-order truck density, the system can still determine that the plot remains within the boundary.
[0121] Adjustments to remote sensing image verification: For labor-intensive industrial parks, specific ground features such as "large trucks," "containers," and "gantry cranes" have been added to the semantic segmentation model's recognition category library. If a high density of these logistics elements is detected within an edge plot, even if the plot is separated from the core area by a small road, the system can still determine that it functionally belongs to the cluster and appropriately expand the boundary to include the plot.
[0122] Adjustments to on-site survey markings: Field inspectors will focus on the actual usage status of peripheral plots, such as whether they are being used as long-term leased logistics transit centers or whether they are connected to the core factory area by internal access roads (even if there are no formal roads). Inspectors can take photos of the site using mobile devices and upload them as a basis for boundary correction.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data, characterized in that, Includes the following steps: Identify employment land parcels based on the land parcel database and select parcels that meet preset conditions; By spatially connecting the plots of land that meet the preset conditions, a contiguous employment area can be formed; Multiple employment clusters are filtered according to a preset strategy to obtain employment clusters that meet the preset strategy; and the first employment population density and the first truck density of the employment clusters are obtained. Based on the first employment population density and the first truck density, the employment clusters are classified to determine the cluster type; It will develop both technology-intensive and labor-intensive clusters. In response to the fact that the cluster type is a technology-intensive cluster, the second employment population density and enterprise density corresponding to the land parcels on the edge of the cluster are obtained, and the first initial boundary of the technology-intensive cluster is determined based on the second employment population density interval and the enterprise density interval. In response to the fact that the cluster type is a labor-intensive cluster, the second truck density corresponding to the plots in the edge area of the cluster is obtained, as well as whether there are independent enterprises occupying land in the edge area of the cluster; In response to the existence of an independent land-occupying enterprise in the edge area, a second initial boundary of the labor-intensive cluster area is determined based on the land boundary of the independent land-occupying enterprise; otherwise, the second initial boundary is determined based on the second truck density. Based on the spatial attribute data of each plot in the edge region of the first initial boundary, the first initial boundary is checked and corrected to obtain the first determined boundary. Based on the spatial attribute data of each plot in the edge region of the second initial boundary, the second initial boundary is checked and corrected to obtain the second determined boundary.
2. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, Identifying employment land parcels based on a land parcel database and selecting parcels that meet preset conditions includes the following steps: In the pre-configured land database, land parcels that meet the criteria are selected based on whether they are commercial, residential, industrial, or warehousing land.
3. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, The spatial connection of the land parcels that meet the preset conditions includes the following specific steps: Based on a predetermined distance extended outward from the boundary of each employment land plot, the plots are spatially connected to form an employment cluster.
4. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 3, characterized in that, The preset distance is 13 meters to 22 meters.
5. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, The multiple employment clusters are filtered according to a preset strategy to obtain employment clusters that meet the preset strategy. The specific steps include: Multiple employment clusters are filtered based on the following criteria: employment area greater than or equal to 30 hectares, employment cluster located within a pre-configured industrial geofence data set, and employment population density greater than 0.5 million people / square kilometer. The resulting clusters are considered employment agglomeration areas.
6. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, Both the first truck density and the second truck density are obtained based on satellite positioning trajectory data through a stop point identification algorithm; specifically including: The spatiotemporal trajectories of all mobile terminals entering the employment cluster area are obtained; target terminals that meet the preset characteristics of freight vehicle movement parameters are filtered; the stopping points of the target terminals in the employment cluster area are identified; and the first truck density is generated based on the ratio of the number of stopping points to the land area of the employment cluster area.
7. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, The first employment population density and the second employment population density are obtained based on mobile phone signaling data from operators. The identification rule is as follows: within the statistical period, if the cumulative number of days that a user terminal generates residence signaling in the plot exceeds a preset threshold during a preset time period on weekdays, then the user is marked as a regular residence user in the plot.
8. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, The determination of the first defined boundary specifically includes the following methods: determining the first defined boundary based on the occupancy ratio of each plot in the edge area of the first initial boundary; or determining the first defined boundary based on real-time remote sensing image data of each plot in the edge area of the first initial boundary; or determining the first defined boundary based on on-site survey and annotation data of each plot in the edge area of the first initial boundary.
9. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, The determination of the second defined boundary specifically includes the following methods: determining the second defined boundary based on the occupancy ratio of each plot in the edge area of the second initial boundary; or determining the second defined boundary based on real-time remote sensing image data of each plot in the edge area of the second initial boundary; or determining the second defined boundary based on on-site survey and annotation data of each plot in the edge area of the second initial boundary.
10. The method for spatial identification of manufacturing clusters based on multi-source heterogeneous industrial big data according to claim 1, characterized in that, Based on the first employment population density and the first truck density, the employment clusters are classified to determine the cluster type, specifically including the following steps: Based on the first employment population density and the first truck density, the employment cluster area is initially classified; wherein, if the first employment population density of the employment cluster area is greater than a first preset density threshold and the first truck density is less than or equal to the first preset truck threshold, it is initially classified as a technology-intensive cluster area; if the first truck density of the employment cluster area is greater than the first preset truck threshold, it is initially classified as a labor-intensive cluster area. Obtain auxiliary data on the industry attributes of the employment cluster; the auxiliary data on industry attributes includes at least one of the following: industry category distribution characteristics extracted based on enterprise business registration information, average plot ratio or building density extracted based on land database, and the total number of regularly resident users in the employment cluster. The preliminary classification results are verified and corrected based on the industry attribute auxiliary data to determine the final cluster type.