A Method for Managing Enterprise Distribution Information on the Inward and Outward Migration of an Industrial Platform

By standardizing the processing and classification aggregation of multi-source data, regional migration distribution maps and classification trend indicators are generated, and the problem of inactivity and incomplete data in the existing technology is solved, and efficient and accurate management and decision-making support for the migration of industrial platform enterprises is achieved.

CN120123331BActive Publication Date: 2025-07-01HANGZHOU ZHILUO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510618743.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-01
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing technology has problems of lagging response and distribution judgment deviation in the management of enterprise migration and migration information in industrial platforms. This is mainly due to the single source of data and the reliance on enterprises to submit information by themselves, resulting in the information being in real time and incomplete.

Method used

By formatting standard processing of multi-source basic data sets, unify the field structure, eliminate structure conflict fields, and deduplication operations based on data timestamps to form an integrated data set. Then, the enterprise migration time field, enterprise location field and enterprise category field are extracted, time segmentation and spatial mapping are performed, regional migration distribution map is generated, and the migration behavior is classified and aggregated by enterprise category, migration activity is evaluated, and classification trend indicators are generated.

Benefits of technology

It has realized structured, multi-dimensional and standardized management of the distribution information of industrial platform enterprises, improved the timeliness and accuracy of data updates, enhanced the efficiency and integrity of enterprise migration information acquisition, and provided more timely and geographically accurate decision-making support for policy formulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123331B_ABST
    Figure CN120123331B_ABST
Patent Text Reader

Abstract

The present invention provides a method for managing enterprise distribution information for industrial platform relocation, which relates to the technical field of data processing. The method includes: performing formatted standard processing on a multi-source basic data set, unifying the field structures of each data source, eliminating fields with structural conflicts, and performing duplicate removal operations based on data timestamps to obtain an integrated data set; extracting enterprise migration time fields, enterprise location fields, and enterprise category fields one by one; segmenting the data in the integrated data set by time to generate a periodic migration data set; mapping the original location and the target location to a standard coordinate system according to the periodic migration data set and the enterprise location field to generate a regional migration distribution map; classifying and aggregating enterprise migration behaviors by enterprise category, and evaluating the migration activity of different categories of enterprises in the target area within a specific period to generate a classification trend index for enterprise relocation behaviors. The present invention improves the autonomy and accuracy of enterprise distribution information management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for managing the enterprise distribution information of the industrial platform's in and out migrations. Background Art

[0002] In the prior art, for the management of the in and out migration information of enterprises within the industrial platform, it usually relies on manual entry or semi-automated information collection systems. The platform managers obtain enterprise migration data through methods such as enterprise reporting, registration information updates, or regular surveys, and enter it into the information management system. Most of these systems adopt a structured data storage method based on databases, using fields to identify the basic information of enterprises, migration time, original location, target location, etc., and then combining visualization tools to display the changes in enterprise distribution to assist in decision-making analysis. Some systems also introduce Geographic Information System (GIS) to enhance the intuitive display ability of spatial distribution.

[0003] However, in scenarios involving the evaluation of investment promotion projects, the prior art may have problems of response lag and distribution judgment deviation. For example, when a certain industrial park counts the in-migration trend of a certain type of high-tech enterprise through the platform, due to the single data source and only relying on the information reported by the enterprises themselves, if the enterprises do not report in time or the data entry is delayed, the system may misjudge the time node or regional scope of the concentrated in-migration of this type of enterprise, thus affecting the rational allocation of resources. In this case, when formulating support policies or allocating infrastructure, it is easy to make wrong decisions based on incomplete information, exposing the technical defects of the existing system lacking real-time and multi-source data integration capabilities. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for managing the enterprise distribution information of the industrial platform's in and out migrations, aiming to solve the problems mentioned in the background art.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] A method for managing the enterprise distribution information of the industrial platform's in and out migrations, the method comprising:

[0007] Performing formatted standard processing on the multi-source basic data set, unifying the field structures of each data source, removing the fields with structural conflicts, and performing a deduplication operation based on the data timestamp to obtain an integrated data set;

[0008] Extracting the enterprise migration time field, enterprise location field, and enterprise category field item by item according to the integrated data set;

[0009] Segmenting the data in the integrated data set according to the enterprise migration time field to generate a periodic migration data set;

[0010] Based on the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map;

[0011] Based on the regional migration distribution map, periodic migration dataset and enterprise category field, enterprise migration behaviors are classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out.

[0012] Preferably, the multi-source basic data sets are formatted in a standard manner, the field structures of the data sources are unified, the fields with structural conflicts are removed, and deduplication is performed based on the data timestamp to obtain an integrated data set, including:

[0013] According to the semantic relevance of the field names in the multi-source basic data set, the field names of the multi-source basic data set are normalized and converted to obtain a first data set;

[0014] Identify structural conflicts of repeated field contents in the first data set, retain field data with high source priority, and remove content conflict fields to obtain a second data set;

[0015] According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set.

[0016] Preferably, according to the periodic migration data set and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including:

[0017] According to the enterprise location field, the information of the original location and the target location of the enterprise is extracted from the integrated data set, and the unstructured address data is converted into structured coordinate pairs to form a location coordinate pair set;

[0018] According to the enterprise identification field in the periodic migration data set, the enterprise migration records in each period are matched with the location coordinate pair set to generate a periodic location mapping data set, wherein the periodic location mapping data set carries both time and space information;

[0019] According to the periodic location mapping data set, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path data set;

[0020] A directed weighted graph is constructed based on the geographic path dataset, with nodes representing regional identifiers and edges representing migration directions and numbers, to generate a regional migration distribution map.

[0021] Preferably, based on the regional migration distribution map, the periodic migration data set and the enterprise category field, the enterprise migration behavior is classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out, including:

[0022] According to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated;

[0023] According to the set of migration records, the number of migrants in and out of each target area is aggregated and counted for each category of migration subset, and a data table with a three-dimensional structure of category-region-period is constructed. Based on this data table, the migration density function is used to calculate the migration activity of each category in each area in different periods to form a migration activity sequence;

[0024] Based on the data differences of consecutive periods in the migration activity sequence, the cycle change rate is calculated, and the indicator volatility calculation formula based on standard deviation normalization is used to obtain the classification trend indicators of enterprise migration behavior in various categories and regional dimensions.

[0025] Preferably, according to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain the first data set, including:

[0026] Extract all field names from multi-source basic data sets to form a field name set;

[0027] Based on the field name set and the preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a recommended set of standard field names;

[0028] According to the recommended set of standard field names, for field names that are judged to be semantically consistent, the field names in the multi-source basic data set are replaced with the corresponding standard field names by word meaning replacement to generate a standard field set;

[0029] The standard field set is arranged in the order set by the standard field structure template to generate a first data set.

[0030] Preferably, structural conflicts are identified for repeated field contents in the first data set, field data with high source priority is retained, and content conflict fields are removed to obtain a second data set, including:

[0031] Extract the original source of each field according to the field source information in the first dataset, and based on a preset credibility level, assign a field source priority to form field priority data;

[0032] According to the field priority data, perform a conflict comparison on the fields in the first dataset that have semantic consistency but conflicting sources, retain the field data with a higher priority, and eliminate the field values with a lower priority or inconsistent content to generate conflict cleaning result data;

[0033] Use the conflict cleaning result data as the output to obtain the second dataset.

[0034] Preferably, according to the timestamp information corresponding to each data record in the second dataset, combined with the enterprise identification field, identify approximate duplicate records and delete redundant items to obtain an integrated dataset, including:

[0035] Extract the timestamp field and the enterprise identification field from the second dataset, and combine them into a unique identification key to generate record index data;

[0036] Based on the record index data and within a preset time difference tolerance interval, determine whether a data record meets the duplicate condition, where the duplicate condition is that the time difference is within the preset time difference tolerance interval and the enterprise identification is the same;

[0037] For the data records that meet the duplicate condition, retain the optimal record in the second dataset based on data timeliness, delete the remaining records, and output the integrated dataset.

[0038] Preferably, according to the periodic position mapping dataset, map each group of position coordinate pairs to a unified standard geographic coordinate system, and construct an enterprise migration path vector based on this to form a geographic path dataset, including:

[0039] According to the original position and target position coordinate pairs included in each enterprise migration record in the periodic position mapping dataset, convert them from the original coordinate system to spatial point positions in a unified standard geographic coordinate system to form a set of standardized position pairs;

[0040] Based on the spatial discretization principle of grid coding, perform coding processing on each pair of coordinates in the set of standardized position pairs, and convert the coordinate values into indexable regional grid codes;

[0041] According to the enterprise identification field and the time period field, represent the migration direction of each enterprise within a specific period as a path vector from the starting code to the ending code, and attach the migration quantity and enterprise category attribute information to each path vector to construct an enterprise migration path dataset with attributes;

[0042] Organize all enterprise migration path vectors by period to generate a geographic path dataset with a time hierarchy.

[0043] Preferably, a directed weighted graph is constructed based on the geographical path dataset, where nodes represent regional identifiers, and edges represent the migration direction and the number of migrations, generating a regional migration distribution map, including:

[0044] According to the starting grid code and the ending grid code carried by each migration path vector in the geographical path dataset, they are respectively assigned to the standard administrative regions as the node identifiers in the graph, forming a set of regional nodes;

[0045] Perform a merging operation on multiple path vectors that belong to the same time period and have the same starting region and ending region, calculate the total sum of the number of migrations corresponding to the migration path as the edge weight of the path in the graph;

[0046] Using the set of regional nodes as the nodes in the graph and the start and end region pairs after path aggregation as the directed edges in the graph, construct a directed connection relationship between the nodes, and use the total sum of the number of migrations as the weighted value of the edge to form a weighted directed graph structure;

[0047] Output the directed weighted graph in the form of a map as the regional migration distribution map.

[0048] Preferably, a migration density function is used to calculate the migration activity of each category in each region in different periods, forming a migration activity sequence, including:

[0049] According to the data table of the category-region-period three-dimensional structure, extract the number of enterprises moving into and out of each target region by each category of enterprises in each time period, and use them as the measurement values of the migration behavior of the category of enterprises in the region in the period;

[0050] For each target region, extract the corresponding spatial area parameter or industrial capacity parameter as the benchmark factor for density normalization;

[0051] Divide the measurement values of the migration behavior of each category by the corresponding benchmark factor to obtain the normalized migration activity of each category in each target region and each time period;

[0052] Arrange the normalized migration activities in sequence according to the time period to form a migration activity sequence.

[0053] The above solutions of the present invention at least include the following beneficial effects:

[0054] Through the implementation of this method, the structured, multi-dimensional, and standardized management of the enterprise relocation distribution information in the industrial platform is realized. Compared with the passive data collection mechanism that relies on enterprise reporting and manual registration in the existing technology, this method breaks the limitation of single information source by accessing multiple data sources and establishing a unified data fusion and analysis process. Through the integration and processing of multi-source basic data, not only the data coverage range is expanded, but also the timeliness and accuracy of data update are significantly enhanced, and the acquisition efficiency and integrity of enterprise migration information are improved.

[0055] In addition, a standardized information structure mainly composed of "time field" and "location field" is constructed in this method. Combining the periodic time segmentation mechanism and the standard mapping of spatial coordinates, the structured modeling of enterprise migration behavior in the two dimensions of time and space is realized. Different from the traditional method that relies on field table display and simple GIS visualization, the regional migration distribution map generated by this solution has clear time series stratification and coordinate unity, and can accurately reflect the migration path and trend, providing more timely and geographically accurate decision-making support for policymakers.

[0056] Furthermore, this method also introduces the enterprise category field as a classification analysis dimension to conduct fine-grained aggregation of enterprise migration behavior, and quantifies and reflects the migration changes of different categories of enterprises in the target area by calculating the migration activity and trend indicators. This classification trend indicator has triple attributes of space, time, and industry, and can be used to identify the agglomeration trend, outward migration risk, and periodic fluctuations of specific industries, so as to support the refined management of investment promotion, industrial guidance, and regional infrastructure investment.

[0057] In summary, this method has the following significant beneficial effects: enhancing the diversity and fusion ability of data sources, improving the update efficiency and accuracy of enterprise migration information; improving the information expression ability for enterprise behavior through standardized processing and spatio-temporal modeling; providing accurate and quantifiable basis for regional management and policy-making through trend quantification and classification evaluation, and fundamentally solving the technical bottlenecks such as response lag, information misjudgment, and single source existing in the prior art. This method is applicable to various application scenarios such as real-time monitoring, situation analysis, and predictive decision-making in a large-scale industrial platform environment. Brief Description of the Drawings

[0058] Figure 1 is a flowchart of a method for managing the enterprise distribution information of relocation in and out of an industrial platform provided by an embodiment of the present invention. Detailed Embodiments

[0059] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0060] As Figure 1 shown, an embodiment of the present invention proposes a method for managing enterprise distribution information for industrial platform relocation in and out, the method comprising:

[0061] S100. Obtain enterprise basic data from multiple data sources to form a multi-source basic data set;

[0062] S200. Perform formatting standard processing on the multi-source basic data set, unify the field structures of each data source, eliminate fields with structural conflicts, and perform duplicate removal operations based on data timestamps to obtain an integrated data set;

[0063] S300. According to the integrated data set, extract enterprise migration time fields, enterprise location fields, and enterprise category fields item by item;

[0064] S400. According to the enterprise migration time fields, segment the data in the integrated data set by time to generate a periodic migration data set;

[0065] S500. According to the periodic migration data set and the enterprise location fields, map the original location and the target location to a standard coordinate system to generate a regional migration distribution map;

[0066] S600. According to the regional migration distribution map, the periodic migration data set, and the enterprise category fields, classify and aggregate enterprise migration behaviors by enterprise category, and evaluate the migration activity of different categories of enterprises in the target area within a specific period to generate a classification trend index of enterprise relocation in and out behaviors.

[0067] In the embodiment of the present invention, by constructing a complete set of structured management mechanisms for enterprise relocation in and out distribution information, efficient coordination of multi-source data fusion, time segmentation, spatial mapping, and classification evaluation is achieved. First, by obtaining enterprise basic data from multiple data sources, a multi-source basic data set with a wide coverage and diverse data types is formed, avoiding the sample bias problem caused by information silos. The collected data content includes, but is not limited to, enterprise registration information, migration reporting data, and statistical platform data, which are characterized by heterogeneous sources and inconsistent structures.

[0068] In the process of formatting and standardizing the multi-source basic data set, by unifying the field structures of each data source, alignment at the semantic level is achieved, making the subsequent data structures compatible. By eliminating fields with structural conflicts, information conflicts or incorrect merges caused by inconsistent field definitions can be effectively avoided, thus improving the accuracy in the data structure cleaning stage. Further, based on the deduplication operation using data timestamps, time intersections and repetitions of records of the same enterprise in multi-source data are effectively identified and excluded, forming an integrated data set with clear structure and accurate content, ensuring the consistency and uniqueness of information in the subsequent processing.

[0069] After the integrated data set is formed, the enterprise migration time field, enterprise location field and enterprise category field can be extracted item by item to realize the conversion of the data structure to business semantic features, laying a foundation for subsequent processing. By performing time segmentation on the enterprise migration time field to form a periodic migration data set, not only is the overall migration behavior divided into several stages with chronological order, but also trend analysis and time series modeling are supported.

[0070] Further, through the combination of the periodic migration data set and the enterprise location field, the original location and the target location are mapped using spatial standardization technology, realizing the conversion of data from address representation to a standard coordinate system. On this basis, the generated regional migration distribution map has spatial continuity and visualization capabilities, enhancing the platform's monitoring ability of the dynamic distribution of enterprises within the region.

[0071] Finally, based on the regional migration distribution map, the periodic migration data set and the enterprise category field, the enterprise migration behaviors are classified and aggregated, and the migration activity is calculated. This processing not only reflects the ability of dimensionality elevation analysis of data from "enterprise behavior" to "category behavior", but also conducts cross-evaluation through the two axes of time period and spatial region, thereby outputting classification trend indicators of enterprise in-migration and out-migration behaviors. This trend indicator can clearly reveal the agglomeration trend, out-migration risk and fluctuation of a certain category of enterprises in a certain region, providing a predictive decision-making basis for industrial layout, investment promotion policy formulation and infrastructure investment.

[0072] To sum up, through unifying data structures, precise time segmentation, spatial coordinate standardization and category aggregation analysis, this method forms a full-process information management ability for enterprise migration behaviors, with high integrity, high accuracy and good spatio-temporal visualization effects.

[0073] Among them, obtaining the enterprise basic data of multiple data sources to form a multi-source basic data set specifically includes:

[0074] In the enterprise information management and analysis scenario, the collection of basic enterprise data is a prerequisite for identifying in- and out-migration behaviors. "Basic enterprise data" refers to structured or semi-structured data covering enterprise identity, registered address, business category, registration time, legal representative, contact information, migration records, and other information. Traditionally, enterprise information data often comes from a single registration platform, such as an industrial and commercial registration database, an intra-park online reporting platform, or a management system that relies on active reporting by enterprises. These methods all have problems such as limited information dimensions, delayed updates, and insufficient integrity.

[0075] To overcome the above limitations, this method constructs a data access module to collect enterprise basic data from multiple heterogeneous data sources to form a unified data set. The data sources may include but are not limited to the following categories:

[0076] Government information systems (such as business registration, tax filing, land use filing, etc.);

[0077] Third-party enterprise database platforms (such as credit information disclosure systems, business data providers);

[0078] The park builds its own reporting platform;

[0079] Public corporate information collected by web crawlers;

[0080] Data sharing interface for industry associations and chambers of commerce, etc.

[0081] During the collection process, data can be pulled by combining API data interfaces, batch data import, database mirroring synchronization, etc. There may be differences in the enterprise data fields in each data source, such as inconsistent field names, field orders, field types, or even some fields are missing or expressed differently. In this method, all the collected original enterprise information data are organized and compiled into a "multi-source basic data set". This data set has not yet dealt with field conflicts and duplications in the formation stage. Its core goal is to provide basic input for subsequent field normalization, conflict cleaning, and deduplication operations.

[0082] In this way, we avoid reliance on individual enterprise declarations, broaden the breadth of data sources, significantly improve the coverage and update frequency of enterprise basic information, and lay a data foundation for the comprehensive analysis of enterprise migration behavior.

[0083] In a preferred embodiment of the present invention, a multi-source basic data set is formatted in a standard manner, the field structures of various data sources are unified, the structural conflict fields are removed, and deduplication is performed based on the data timestamp to obtain an integrated data set, including:

[0084] Normalize and transform the field names in the multi-source basic dataset according to the semantic relevance of the field names in the multi-source basic dataset to obtain the first dataset;

[0085] Identify structural conflicts in the duplicate field contents in the first dataset, retain the field data with a higher source priority, and eliminate the fields with conflicting contents to obtain the second dataset;

[0086] According to the timestamp information corresponding to each data record in the second dataset, combined with the enterprise identification field, identify approximate duplicate records and delete redundant items to obtain the integrated dataset.

[0087] In the embodiment of the present invention, in the multi-source data processing stage, in the face of multiple data sources with different structures and fields, first, through the analysis method of field semantic relevance, normalize and transform the field names. After extracting all the field names in the multi-source basic dataset, form a field name set, and then perform semantic comparison on these field names in combination with a preset domain vocabulary comparison table to identify fields with the same meaning but different names. This process can not only effectively identify synonymous fields with different names such as "registered capital" and "total capital", but also identify common variant fields such as "enterprise address" and "registered address".

[0088] Replace the original field names with the recommended standard field names, unify the field naming system, and construct a dataset with consistent field semantics. After the field naming is unified, arrange the field order according to the preset field structure template, thereby constructing the first dataset with a standardized field order. This dataset has achieved consistent field semantics and unified structural arrangement, and has the operability for subsequent structural conflict identification.

[0089] Next, further perform structural conflict identification and processing on the situation where there are duplicate field contents in the first dataset. By analyzing the field source information, assign the priority of the field source, and combined with the judgment of field content consistency, eliminate the field values with a lower source priority or inconsistent content, thereby generating the data after structural cleaning. For the situation where there are fields with the same semantics but different values, such as the conflict between "industry" being "manufacturing" and "machinery manufacturing", the field value with accurate semantic expression will be retained according to the set priority, improving the consistency and authority of the field content.

[0090] In addition, perform duplicate removal processing on the records according to the combined identification key of the timestamp and the enterprise identification field of each data record in the second dataset. In this process, judge whether it constitutes a duplicate based on the time difference tolerance interval and the enterprise identification consistency. For the records that meet the duplicate conditions, retain the record item with the most recent time and eliminate the redundant items to ensure that the finally formed data content has time uniqueness and data currency, and improve the data processing quality.

[0091] This process plays a core role in the entire multi-source data standardization processing chain, laying a reliable data foundation for subsequent data mining, behavior analysis, and indicator modeling. The integrated dataset processed by this method achieves a high degree of consistency in multiple aspects such as field structure, field value accuracy, and data uniqueness, providing stable support for the tracking and analysis of enterprise relocation behaviors in the industrial platform.

[0092] In a preferred embodiment of the present invention, according to the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including:

[0093] According to the enterprise location field, extract the information of the original location and the target location of the enterprise from the integrated dataset, and convert the unstructured address data into structured coordinate pairs to form a set of location coordinate pairs;

[0094] According to the enterprise identification field in the periodic migration dataset, match the enterprise migration records in each period with the set of location coordinate pairs to generate a periodic location mapping dataset, and the periodic location mapping dataset carries dual information of time and space;

[0095] According to the periodic location mapping dataset, map each set of location coordinate pairs to a unified standard geographic coordinate system, and construct an enterprise migration path vector based on this to form a geographic path dataset;

[0096] Construct a directed weighted graph based on the geographic path dataset, where the nodes represent regional identifiers, and the edges represent the migration direction and the number of migrations to generate a regional migration distribution map.

[0097] In the embodiment of the present invention, through joint processing based on the periodic migration dataset and the enterprise location field, a complete spatial expression path from address information to geographic coding and then to a directed weighted graph is constructed. First, extract the original location and target location data of the enterprise from the integrated dataset, and extract information based on the enterprise location field. Since the original data often contains unstructured addresses (such as "No. XX, XX Road, Chaoyang District, Beijing"), an address parsing mechanism is used to convert it into a standardized coordinate pair, and a structured coordinate pair set of "original location - target location" is constructed.

[0098] Subsequently, through the enterprise identification field in the periodic migration dataset, accurately match the migration records of each enterprise in a specific period with the set of location coordinate pairs. Each matching result not only binds the migration direction of the enterprise but also carries the specific period information when the migration occurs, thus forming a periodic location mapping dataset with dual dimensions of time and space.

[0099] To achieve unified expression across regions, each pair of coordinates in the periodic location mapping dataset is further converted into points in a unified standard geographic coordinate system. This conversion uses geographic projection conversion rules to handle the mapping between the original coordinates and the target coordinates, avoiding spatial errors caused by different coordinate systems in the multi-source address resolution results. Subsequently, the spatial discretization principle of grid coding is used to convert each pair of standardized coordinates into indexable grid codes, ensuring controllable spatial granularity and efficient grouping and retrieval capabilities.

[0100] On this basis, according to the enterprise identification field and the time period field, the migration behavior of each enterprise within a specific period is encoded into a path vector of "starting point code → ending point code", and attribute values such as the migration quantity and enterprise category are embedded in each path, forming an enterprise migration path dataset with attributes. This process realizes the dimensionality elevation processing of migration information from data records to analyzable paths.

[0101] Finally, all path vectors are sorted according to the period to generate a geographic path dataset at the time level, making the data no longer just a pile of isolated events, but a structured spatio-temporal path system. Through this path system, the platform can clearly display the spatial flow trends of enterprises in different periods on the map, effectively supporting decision-making scenarios such as subsequent regional development assessment and resource scheduling optimization.

[0102] In summary, this embodiment realizes the full-process standardization processing of address information → coordinate data → grid path → directed vector, enhancing the spatial expression ability while enhancing the analyzability and visualization ability of enterprise migration behavior at the geographical level.

[0103] Among them, according to the enterprise location field, the information of the original location and the target location of the enterprise is extracted from the integrated dataset, and the unstructured address data is converted into structured coordinate pairs to form a set of location coordinate pairs, specifically including:

[0104] In the actual enterprise basic information, the "original location" and "target location" are usually recorded in text form, which belongs to typical unstructured data. For example, data such as "No. 27, Zhongguancun Avenue, Haidian District, Beijing" and "Building A3, South Area of Science Park, Nanshan District, Shenzhen" are easy for humans to identify, but difficult for computers to directly perform geographical calculations or spatial analyses. Therefore, in order to realize the structured modeling of enterprise migration behavior in the spatial dimension, such unstructured address texts must be converted into computable geographical coordinate data.

[0105] This method uses address resolution technology (Geocoding) to achieve the process of address to coordinate conversion. Address resolution is a technology that uses a geographical information database and natural language processing means to parse address texts in Chinese or other languages into standard coordinate points. In actual implementation, the address resolution can be achieved through the following two types of methods:

[0106] Access commercial map service APIs, such as Baidu Map, Tencent Location Service, Amap, etc., submit address text requests, and obtain the returned longitude and latitude coordinate results;

[0107] Build its own geographical word library and administrative division dictionary, and through the word segmentation matching and fuzzy comparison mechanism, perform the conversion from local address fields to coordinate points.

[0108] During the address parsing process, to improve the parsing accuracy and coverage rate, the following auxiliary means can also be introduced:

[0109] Preprocess the enterprise address fields, including removing redundant words (such as "office building of Co., Ltd.") and unifying the administrative region expressions (such as replacing "Chaoyang District, Beijing" with "Chaoyang District");

[0110] Set a fault tolerance mechanism for the address text with parsing failures, such as setting the principle of "falling back to the district-level coordinates when the street-level fails", to ensure the availability of spatial data.

[0111] Once the address parsing is successful, a set of structured coordinates can be generated for the original location and target location fields of each enterprise respectively, that is, (original longitude, original latitude) and (target longitude, target latitude). Combine these coordinates into a set of structured data, called "location coordinate pairs".

[0112] Finally, using the "enterprise identifier" as the primary key and combining the original time field, add the location coordinate pairs in each enterprise migration record to the periodic location mapping dataset to form a coordinate data structure with time tags and spatial paths. This set of coordinate pairs not only supports subsequent spatial mapping and atlas generation, but also has clear geographical significance and can be directly visualized as paths on the map.

[0113] Through this step, the conversion of enterprise location fields from unstructured expressions to structured spatial data is effectively realized, endowing enterprise migration data with clear spatial computability and spatial aggregation capabilities, and providing a strong geographical basis for regional analysis, path mapping, and trend analysis.

[0114] In a preferred embodiment of the present invention, according to the regional migration distribution atlas, the periodic migration dataset, and the enterprise category field, classify and aggregate the enterprise migration behaviors by enterprise category, and evaluate the migration activity of different categories of enterprises in the target region within a specific period, and generate classification trend indicators for the enterprise in-migration and out-migration behaviors, including:

[0115] According to the enterprise category field, divide the enterprise migration records in the periodic migration dataset into multiple category migration subsets by category, and combine the original location and target location recorded in the regional migration distribution atlas to generate a set of migration record pairs divided by category and region;

[0116] According to the migration records of the set, for each category of migration subsets, aggregate and count the number of migrated-in and migrated-out in each target region, construct a data table with a three-dimensional structure of category-region-period, and based on this data table, use the migration density function to calculate the migration activity of each category in each region in different periods, forming a migration activity sequence;

[0117] According to the data differences in consecutive periods in the migration activity sequence, calculate the period change rate, and use the index volatility calculation formula normalized based on the standard deviation to obtain the classification trend index of the enterprise migration-in and migration-out behavior in each category and region dimension.

[0118] In the embodiment of the present invention, by focusing on the combined processing of the periodic migration data set, the enterprise category field, and the regional migration distribution map, the refined classification and trend quantification of the enterprise migration behavior are realized. First, through the enterprise category field, the migration records in the periodic migration data set are classified by category. After the classification is completed, combined with the original location and target location carried by each migration path in the regional migration distribution map, the migration data of each category of enterprises are combined according to the migrated-in and migrated-out regions to generate a set of migration record pairs divided by the two dimensions of category and region.

[0119] Based on this set, aggregate and count the number of migrated-in and migrated-out of each category of enterprises in each target region, and organize them structurally according to the period, so as to construct a data table with a three-dimensional structure of category-region-period. This data table not only standardizes the enterprise behavior data, but also provides data support for time trend analysis through the method of period stratification.

[0120] Based on the three-dimensional structure data table, use the migration density function to model and analyze the enterprise activity. In each region and each period, normalize the number of migrated-in or migrated-out of a certain category of enterprises with the regional space resources (such as area, capacity, etc.) to obtain the normalized migration density. This processing method avoids the direct interference of the enterprise migration quantity scale difference on the evaluation result, making the migration behaviors between different categories and different regions comparable on a unified basis.

[0121] Next, taking the time period as the reference axis, calculate the change trend of the migration density under consecutive periods, extract the trend characteristics through the change rate of adjacent periods, and further calculate the volatility intensity index normalized by the standard deviation. This index is used to measure the volatility and directionality of the enterprise migration activity in multiple periods, so as to reflect whether there are phenomena such as stable agglomeration, active fluctuation, or outward migration risk of a certain category of enterprises in a certain region.

[0122] The finally generated classification trend indicator has the three-dimensional fusion characteristics of periodicity, category, and spatiality, and can reveal the dynamic trends of enterprise migration behavior in terms of industrial structure evolution, regional attractiveness change, and policy response effect. This indicator not only shows the clustering or dispersion characteristics of enterprise categories at the spatial distribution level but also captures the industry change rules in the time series, thus providing a strong quantitative decision-making basis for the government's investment promotion, industrial planning formulation, and infrastructure allocation.

[0123] Therefore, through this embodiment, an enterprise behavior evaluation mechanism based on classification aggregation and trend quantification is established, forming a complete closed-loop from data division, indicator construction to trend output, with high stability, high adaptability, and reference value for policies.

[0124] Among them, according to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated, specifically including:

[0125] In the industrial platform, the migration behavior of enterprises usually has significant category differences. For example, the migration paths, cycles, and destination tendencies of manufacturing enterprises may be completely different from those of the financial service industry. Therefore, if we hope to deeply evaluate the migration trend of specific category enterprises, using only the overall enterprise sample as the analysis object will not be able to depict the fine-grained changes, and it is necessary to further classify and disassemble the migration behavior.

[0126] For this reason, this method initially groups the migration data based on the enterprise category field carried by each enterprise migration record in the periodic migration dataset. The enterprise category field usually comes from the "industry affiliation" field in the industrial and commercial registration information or the platform's custom classification field, such as "manufacturing", "information service", "new energy", "medical device", etc. This grouping process can be achieved through string matching, industry code mapping (such as GB / T 4754), or a custom label system.

[0127] After classification, the migration dataset is disassembled into several category migration subsets within each time period. Each subset only contains the migration records within the same enterprise category. On this basis, in order to combine the migration behavior with the spatial flow situation, the original location and target location coding information contained in the regional migration distribution map is further cited to attach a "starting region identifier" and an "ending region identifier" to each record.

[0128] In this way, each enterprise migration record will have the following structural information: enterprise category, migration cycle, original region, target region, and migration quantity. Based on this, a new data set is generated, called the "set of migration record pairs classified by category and region". This set supports two-dimensional aggregation (category × region), which can not only count the number of enterprises of a certain category moving into a certain region but also reflect the spatial flow of the enterprises of this category moving out.

[0129] The above operations achieve the dimensionality increase from the "enterprise record level" to the "regional path level" at the data structure level and realize the cross-integration of "category behavior" and "regional relationship" in the analysis logic, constructing a basic data system for subsequent classification trend modeling.

[0130] Among them, the calculation formula for the index volatility is:

[0131] ,

[0132] is the volatility of the index within the category in the region , that is, the classification trend index;

[0133] is the total number of time periods, such as quarters or months;

[0134] is the change in activity, indicating the difference in activity between the current period and the previous period, and the calculation method is as follows: , is the normalized migration activity of category in the region and period ;

[0135] is the average activity of all periods of category in the region , and the calculation method is: ;

[0136] is the cycle weight coefficient, indicating the fluctuation importance weight of cycle , which can be defined as: , which is the sigmoid function structure; among them,

[0137] is the spatial migration slope, indicating the offset rate of the migration behavior in space within the period, and the calculation method is: , is the category in the region and period The average migration path length within is the cycle length, which is used to standardize the time span. is the adjustment coefficient set by the system, which is used to control the steepness of the weight curve.

[0138] In this formula, in the periodic migration dataset, each category The enterprises in the target area at different time periods constitute the normalized migration activity values shown, and then, based on this, the change in activity between two consecutive periods is extracted. To eliminate the influence brought by the difference in the absolute value of activity between different categories, the system uses the average activity of each category in this area as the normalization denominator, normalizes and scales the activity change value, and squares and averages the results of all periods to measure the intensity of the trend change. In this process, to reflect the influence weight of different periods in the overall fluctuation trend, the system introduces a weight coefficient which is determined by the spatial migration slope of the enterprise within the cycle and reflects the spatial intensity of the migration path within this cycle. Finally, the result output by this formula is defined as the classification trend index of this category of enterprises in this area. The larger the value, the stronger the fluctuation amplitude and the more unstable the trend; the smaller the value, the more stable or continuous the enterprise behavior is in the time series.

[0139] The introduction and calculation of this main formula have the following significant technical advantages and beneficial effects:

[0140] Realize the quantitative expression of trend behavior. Compared with the traditional statistical method that can only output the static comparison of the number of in-migrations and out-migrations, this formula can dynamically depict the behavior fluctuations of enterprises in the cycle dimension by extracting the continuous differences in the time series and normalizing them, solving the problem of overly rough trend judgment in the existing methods.

[0141] Enhance the comparability across categories and regions. The normalization part eliminates the interference brought by the difference in enterprise base through the average activity factor so that the migration activities of different types of enterprises such as manufacturing and service industries can be analyzed for trends on the same scale. This method is especially suitable for horizontal benchmarking and performance evaluation among multiple industrial parks and different investment promotion regions.

[0142] Introduce spatial behavior factors for weighted enhancement. Traditional migration trend assessments often ignore the differences in the length of the moving paths of enterprises in space and only consider the quantity changes. However, this method calculates the spatial migration slope by introducing the path length and cycle span , and construct weights through the Sigmoid function Enhances the sensitivity to behaviors with drastic spatial changes, making the trend indicator better reflect the reconstruction intensity of the enterprise group in terms of physical distribution. Supports behavior prediction and early warning mechanisms. Indicator volatility Can be used as a dynamic monitoring variable in. If the trend indicator of a certain category of enterprises in a certain region continuously exceeds the historical threshold, it can automatically trigger analysis and early warning or resource scheduling suggestions, improving the platform's perception ability and response efficiency to industrial dynamic changes.

[0143] Has good formula extensibility and calculation adaptability. Since all variables are derived from the data sets clearly defined in the claims (such as periodic migration data sets, enterprise category fields, geographical path data sets), this formula can be directly integrated into the existing system architecture without introducing external parameters; at the same time, its form supports the future expansion of more weight dimensions (such as economic indicators, energy consumption constraints, etc.).

[0144] In a preferred embodiment of the present invention, according to the semantic relevance of the field names in the multi-source basic data set, the field names of the multi-source basic data set are normalized and transformed to obtain a first data set, including:

[0145] Extract all field names in the multi-source basic data set to form a field name set;

[0146] According to the field name set, combined with a preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a standard field name recommendation set;

[0147] According to the standard field name recommendation set, for the field names determined to be semantically consistent, use the word meaning replacement method to replace the field names in the multi-source basic data set with the corresponding standard field names to generate a standard field set;

[0148] Arrange the fields in the standard field set in the order set by the standard field structure template to generate a first data set.

[0149] In the embodiment of the present invention, by focusing on realizing the standardization of the multi-source basic data set at the field structure level, the semantic barriers between data from different sources are broken through. In the actually collected multi-source data, due to the heterogeneity of information sources, different data tables often use different naming methods for the same attribute field. For example, some sources use "registered capital", some write "capital value", and some call it "fund scale". This situation where the semantics are similar but the names are not unified seriously interferes with subsequent data fusion operations.

[0150] To achieve a consistent expression of field names, first, all field names in the multi-source basic dataset are extracted to form a field name set. This set not only contains the original field names but also serves as the target set for semantic normalization processing. During the processing, through a preset domain vocabulary comparison table, each field in the field name set is semantically judged one by one to identify those field pairs with similar meanings but different names, and further determine the recommended standard field names. For example, "Company establishment time", "Registration time", and "Establishment date" are uniformly replaced with "Enterprise establishment time".

[0151] Next, according to the recommended set of standard field names, for all field names identified as semantically consistent, a unified replacement operation is performed in the multi-source basic dataset. Through this process, each field name in the original dataset is replaced with the corresponding standard field name, forming a standard field set, achieving semantic normalization and naming consistency of the fields. This operation not only enhances the unity of the field structure but also eliminates the risk of data merging conflicts caused by field duplication or ambiguity.

[0152] In addition, to ensure the unity of the field arrangement in terms of structure, the standard field set is rearranged according to the order set by the preset field structure template. The field structure template is a benchmark for field arrangement based on industry standards or business modeling rules, ensuring that after the data sources are uniformly named, the order of the fields also remains consistent. Through the dual processing of unified field names and unified order, the finally generated first dataset has a highly consistent field structure.

[0153] This embodiment significantly improves the consistency and stability of data structure processing, provides a solid structural foundation for subsequent operations such as structure conflict identification, field content fusion, and duplicate record deduplication, and greatly reduces the workload of manual cleaning and data comparison, improving the processing efficiency.

[0154] Among them, the preset domain vocabulary comparison table is a semantic mapping resource constructed by manual collation and automatic expansion, specifically used to unify field names with similar semantics but different names into standard field names. Its construction principles mainly include the following two methods:

[0155] Manual definition method: According to industry data standards (such as national standards GB / T22240, GB / T4754, etc.), industry databases (such as enterprise industrial and commercial registration systems), and historical project experiences, high-frequency fields and their common synonymous names are sorted out to establish manual mapping relationships.

[0156] Semantic calculation and expansion method: Using natural language processing technologies, such as word vector models, edit distance matching, and word sense disambiguation algorithms, semantic similarity calculations are performed on the existing field name set to automatically discover candidate mapping relationships, which are supplemented into the comparison table after manual review.

[0157] The comparison table is saved in the structure of "standard field name → multiple candidate field names", for example:

[0158] "Registered capital" → ["Capital value", "Fund scale", "Registered fund"]

[0159] "Establishment time" → ["Registration time", "Company establishment date", "Enterprise founding date"]

[0160] During the field normalization process, the system can match whether the field name appears in this table one by one. If a certain candidate name is hit, it will be replaced with the corresponding standard field name. This comparison table, as the basic configuration file before system deployment, can be customized and maintained according to industries, projects or usage scenarios.

[0161] Among them, the standard field structure template refers to the structured specification used to constrain the arrangement order of data table fields and the content of the field set, so as to unify the expression methods of different data sources at the field level, field order and field existence, ensure that the normalized data has consistent structural semantics, and facilitate the subsequent processing module to call and display.

[0162] Its basic composition includes:

[0163] Field set definition: Clearly define which standard fields are included, such as "Enterprise name", "Enterprise type", "Legal representative", "Registered capital", "Original address", "Target address", "Industry classification", etc.;

[0164] Field order convention: To avoid data processing anomalies caused by different field arrangements, the template specifies the sequence number of each field, so as to achieve unified sorting;

[0165] Field attribute description: Set attributes such as type (such as string, date, floating point number), length limit, and whether it can be empty for each field, providing a basis for subsequent data verification and interface mapping.

[0166] The template can be implemented in the form of JSON, XML or database table structure, and used as the output format specification after field normalization. After the field names are unified, the system arranges the fields in the order specified in this template, fills in the missing fields with null values, discards or marks the redundant fields, and finally outputs the first data set with consistent field structures.

[0167] In a preferred embodiment of the present invention, structural conflicts of duplicate field contents in the first data set are identified, the field data with a higher source priority is retained, and the fields with conflicting contents are removed to obtain the second data set, including:

[0168] Extract the original source of each field according to the field source information in the first dataset, and based on a preset credibility level, assign a field source priority to form field priority data;

[0169] According to the field priority data, perform a conflict comparison on the fields in the first dataset that have the same semantics but conflicting sources, retain the field data with a higher priority, and eliminate the field values with a lower priority or inconsistent content to generate conflict cleaning result data;

[0170] Use the conflict cleaning result data as the output to obtain the second dataset.

[0171] In the embodiment of the present invention, it mainly deals with the content-level conflict problems that may still exist after the field structures in the first dataset are unified, that is, the situation where "the field structures are the same but the field values are different". In practical applications, different data sources may record different values for the same field. For example, for the field of "legal representative", one data source may record it as "Zhang San", while another data source may record it as "Zhang Sanfeng". If merged directly without discrimination, it will cause data inconsistency and even incorrect use.

[0172] Therefore, based on the first dataset, first extract the source information of each field, including the source data source identifier, collection time, credibility level, etc. According to the set credibility evaluation rules, assign priorities to different sources. For example, government filing data is superior to web scraping data, and manually reviewed data is superior to automatically filled data. Thus, field source priority data is generated to provide a weight reference for conflict processing.

[0173] Subsequently, identify the records with the same field name but different field values, and perform conflict comparison on these records. By comparing the source priorities, retain the data from the source with a higher priority, and at the same time eliminate the field values with a lower priority or inconsistent semantics to generate the structural conflict cleaning result data. If the field value content is the same, it is determined as a redundant field, and the remaining duplicate values are deleted after retaining one copy. By processing in this way, not only the problem of inconsistent field values is solved, but also the de-duplication at the field value level is completed.

[0174] Finally, through the double cleaning operations of conflict identification and redundant field elimination, a second dataset with clear structure, consistent field values, and complete semantics is obtained. This dataset has high comparability and fusion, providing data guarantee for the accurate management and trustworthy modeling of enterprise data.

[0175] This embodiment effectively improves the consistency and accuracy of field content, avoids interference with the analysis results due to structural conflicts or redundant fields, and provides a reliable data source for constructing a highly trustworthy enterprise information portrait.

[0176] Among them, the main construction methods of the preset credibility level include:

[0177] Source classification dimension: The data sources are classified into "government official", "third-party commercial", "platform filling", "crawler scraping", etc.

[0178] Timeliness dimension: Considering the latest update time of the data, the closer to the current date, the higher the credibility.

[0179] Manual review flag: Determine whether the data has been manually verified or business verified. If the review has been passed, the credibility is improved.

[0180] Each source is assigned a level label, such as "high (level 1)", "medium (level 2)", "low (level 3)". When dealing with field conflicts, if there are value conflicts in multiple fields, the data with a higher level is retained, and the data with a lower level is deleted or ignored. For example, when there is a conflict in the field value of "legal representative", if one source is the industrial and commercial registration system (high level) and the other is the enterprise website (medium level), the former is retained.

[0181] This mechanism ensures that the data selection in the merging and cleaning processes favors more reliable sources by introducing a credibility ranking strategy, improving the overall quality and consistency of the integrated data.

[0182] In a preferred embodiment of the present invention, according to the timestamp information corresponding to each data record in the second dataset, combined with the enterprise identification field, approximate duplicate records are identified, and redundant items are deleted to obtain an integrated dataset, including:

[0183] Extract the timestamp field and the enterprise identification field from the second dataset, and combine them into a unique identification key to generate record index data;

[0184] Based on the record index data and within a preset time difference tolerance interval, determine whether the data record meets the duplicate condition, where the duplicate condition is that the time difference is within the preset time difference tolerance interval and the enterprise identification is the same;

[0185] For the data records that meet the duplicate condition, retain the optimal record in the second dataset based on data timeliness, delete the remaining records, and output the integrated dataset.

[0186] In the embodiment of the present invention, attention is paid to the possible record-level duplicate problems in the second dataset, especially the time-proximity duplicate records caused by synchronization delays and different update frequencies between different data sources. Such records usually have the same field structure and enterprise identification, but there are slight differences in data content or timestamps. If deduplication is not performed in a timely manner, it will directly affect the accuracy and timeliness of data analysis.

[0187] First, extract the timestamp field and the enterprise identifier field contained in each data record from the second dataset, and combine the two to form a unique identification key. This identification key is used to construct a data index set, which can quickly retrieve data records with the same enterprise identifier in the index structure.

[0188] Based on the index, for all data records under each enterprise identifier, execute the duplicate judgment logic based on the time difference tolerance interval. By setting an adjustable time difference threshold, such as "48 hours" or "3 days", determine whether the records are close enough in time. If the enterprise identifiers are the same and the time difference is less than the threshold, it is judged as a duplicate record, constituting a duplicate condition.

[0189] For records that meet the duplicate condition, further execute the strategy of retaining the optimal record. Specifically, give priority to retaining the data record with the latest timestamp and delete the remaining redundant items, so as to ensure that the retained data has the highest timeliness and the most complete field content. All the cleaned results are summarized into a new integrated dataset.

[0190] In this embodiment, through the enterprise identifier and time joint key, accurate identification and deletion of approximate records are achieved, effectively avoiding the redundant accumulation problem caused by multi-source synchronization delay. At the same time, the latestness and uniqueness of the final dataset in the time dimension are ensured, significantly improving the data quality and application value.

[0191] Among them, the preset time difference tolerance interval, that is, when comparing the timestamp fields of records, set a maximum allowable error range to determine whether it belongs to "approximate duplication".

[0192] This interval can be set based on the actual business scenario, such as:

[0193] For a system with a daily update frequency, the time difference tolerance interval can be set to 1 day;

[0194] For a monthly update or near-real-time system, the tolerance interval can be set to 7 days or longer;

[0195] For the historical record data imported in batches, the tolerance interval can be set according to the record coverage period.

[0196] When comparing two records with the same enterprise identifier, the system will calculate the difference between the timestamps of the two records. If the difference is less than or equal to the tolerance interval, the system considers these two records to be approximately duplicate and enters the deduplication process, retaining the record with higher timeliness.

[0197] This tolerance interval can be dynamically set as a system configuration item or can be configured differently for different data sources according to the characteristics of the data source, providing a more flexible policy support for enterprise-level data deduplication and data uniqueness guarantee.

[0198] In a preferred embodiment of the present invention, according to the periodic position mapping data set, each set of position coordinate pairs is mapped to a unified standard geographic coordinate system, and an enterprise migration path vector is constructed based on this to form a geographic path data set, including:

[0199] According to the original position and target position coordinate pairs included in each enterprise migration record in the periodic position mapping data set, convert them from the original coordinate system to spatial points in the unified standard geographic coordinate system to form a standardized position pair set;

[0200] Based on the spatial discretization principle of grid coding, perform coding processing on each pair of coordinates in the standardized position pair set, and convert the coordinate values into indexable regional grid codes;

[0201] According to the enterprise identification field and time period field, represent the migration direction of each enterprise within a specific period as a path vector from the starting point code to the ending point code, and attach the migration quantity and enterprise category attribute information to each path vector to construct an enterprise migration path data set with attributes;

[0202] Organize all enterprise migration path vectors by period to generate a geographic path data set with a time hierarchy.

[0203] In the embodiment of the present invention, after processing the periodic position mapping data set, in order to further enhance the spatial expression ability of enterprise migration behavior, this embodiment maps each set of coordinate pairs of the original position and the target position to a unified standard geographic coordinate system, and constructs an enterprise migration path vector based on this to form a geographic path data set. First, extract the original position and target position coordinate information included in the enterprise migration records from the periodic position mapping data set, and convert them from the original acquisition coordinate system (which may be local or non-uniform standard) to a unified standard geographic coordinate system (such as WGS84 or GCJ02). This standardized conversion eliminates the spatial offset problem caused by coordinate differences and ensures the visualization processing of all position points under a unified reference system.

[0204] Subsequently, to enhance the spatial grouping and indexing efficiency, process the standardized coordinate pairs based on the spatial discretization principle of grid coding. By mapping the coordinate values to uniquely identifiable grid codes (such as GeoHash, Quadtree, or hexagonal coding systems), the spatial discretized expression of the migration path is realized, and operations such as fast geographic range filtering and regional level aggregation are supported.

[0205] Based on the above coding results, further using the enterprise identification field and the time period field as index keys, represent the migration path of each enterprise within a specific period as a path vector of "starting point coding → ending point coding", and embed additional information such as the migration quantity and enterprise category into the path attributes to form migration vector data with attributes. These attribute information not only supports spatial structure mapping but also provides operable dimensions for subsequent multi-dimensional analysis (such as industry and migration intensity).

[0206] All migration path vectors are then organized and layered according to the time period to construct a geographical path dataset with a time hierarchy. This dataset not only has the accuracy of the spatial dimension but also retains the time sequence characteristics, and has multiple application values such as path tracking, trend judgment, and classification aggregation. On the map platform, dynamic path animations, periodic migration heat displays, and visual expressions of enterprise migration flow maps can be realized, greatly enhancing the platform's perception ability and presentation efficiency of regional enterprise activities.

[0207] Through this embodiment, not only are discrete data records converted into a structured migration vector network, but also a high degree of integration is achieved in terms of spatial expression, attribute extension, and time series management, significantly enhancing the usability, visualization, and structured management capabilities of enterprise migration information.

[0208] In a preferred embodiment of the present invention, a directed weighted graph is constructed based on the geographical path dataset, where the nodes represent regional identifiers, and the edges represent the migration direction and the migration quantity, generating a regional migration distribution map, including:

[0209] According to the starting grid coding and the ending grid coding carried by each migration path vector in the geographical path dataset, assign them to the standard administrative regions respectively as the node identifiers in the graph to form a regional node set;

[0210] Perform a merging operation on multiple path vectors belonging to the same time period and having the same starting region and ending region, calculate the total sum of the migration quantities corresponding to this migration path as the edge weight of this path in the graph;

[0211] Using the regional node set as the nodes in the graph and the start and end region pairs after path aggregation as the directed edges in the graph, construct a directed connection relationship between the nodes, and use the total sum of the migration quantities as the weighted value of the edge to form a weighted directed graph structure;

[0212] Output the directed weighted graph in the form of a map as the regional migration distribution map.

[0213] In the embodiment of the present invention, after forming the enterprise migration path vector, in order to further refine the macro flow direction of enterprise migration between regions, this embodiment constructs the geographical path dataset into a structured directed weighted graph to represent the migration relationship between regions. First, extract the grid coding information of the starting point and the ending point included in each path in the geographical path dataset, and map the grid coding to the corresponding standard administrative region according to the preset regional aggregation mapping rule. This mapping relationship is completed according to the regional boundary division table, supporting accurate matching at the levels of cities, districts and counties, and even industrial parks, so as to form a regional node set with administrative units as node identifiers.

[0214] In the process of constructing the graph structure, aggregate and count the path vectors with the same starting region and ending region in each time period, and calculate the total sum of their migration quantities as the weight of the edge of this path in the graph. This path aggregation operation not only eliminates the redundant information at the individual enterprise dimension, but also effectively extracts the overall migration intensity between regions, forming a directed edge with "region pair" as the connection relationship and "total migration volume" as the weight.

[0215] Construct the edges of the graph with the aggregated starting and ending region pairs, and form a complete directed weighted graph structure with all region identifiers as the nodes of the graph. Among them, the nodes represent geographical regions, the edges represent the migration direction, and the weights represent the number of enterprises migrating from one region to another region within a specific period, with a clear meaning of spatial flow. The change in the number and direction of the edges in the graph can dynamically reflect the enterprise flow relationship between regions, so as to identify the in-migration aggregation areas, out-migration high-incidence areas and areas with intensive migration path frequencies.

[0216] Finally, output the graph structure in the form of a regional migration distribution atlas, supporting atlas display methods, including force-directed layout, geographical layout, heat path flow chart, etc., and can clearly present the main paths and hot spots of the enterprise migration network on the interface. This atlas expression form not only enhances the visualization effect, but also supports the insight into the industrial agglomeration trend, the effect of policy attractiveness and the rationality of spatial resource allocation from the macro level.

[0217] This embodiment not only realizes the dimensionality-raising expression from enterprise-level path vectors to regional-level network structures, but also greatly improves the interpretability and decision-making support value of regional migration patterns.

[0218] In a preferred embodiment of the present invention, a migration density function is used to calculate the migration activity of each category in each region in different periods, forming a migration activity sequence, including:

[0219] According to the data table of the category-region-period three-dimensional structure, extract the number of enterprises moving into and out of each target region for each category of enterprises in each time period, and use them as the measurement values of the migration behavior of the enterprises in that category in that region and that period respectively;

[0220] For each target region, extract the corresponding spatial area parameter or industrial capacity parameter as the benchmark factor for density normalization;

[0221] Divide the measurement values of the migration behavior of each category by the corresponding benchmark factor to obtain the normalized migration activity of each category in each target region and each time period;

[0222] Arrange the normalized migration activities in sequence according to the time period to form a migration activity sequence.

[0223] In the embodiment of the present invention, after completing the enterprise category division and regional migration path modeling, this embodiment further realizes the quantification of the classification trend of the enterprise moving-in and moving-out behaviors. The core means is to use the migration density function to perform normalization calculations on the activities of various categories of enterprises in different regions and periods, form a migration activity sequence, and extract its trend change characteristics.

[0224] First, according to the data table of the category-region-period three-dimensional structure, extract the number of enterprises moving into and out of each category of enterprises in each time period and each target region one by one, and use them as the measurement values of the migration behavior. This processing ensures the spatial accuracy and periodic stability of the data, and eliminates the deviation caused by inconsistent time statistical granularity.

[0225] In order to achieve horizontal comparability under different proportional regions and industries, a spatial normalization parameter is introduced, that is, the spatial area or its industrial capacity of each target region is extracted as the benchmark factor for density calculation. The spatial area is mainly used for normalization under the physical scale, while the industrial capacity is used for the evaluation of the bearing capacity dimension, and the two are switched according to the platform configuration.

[0226] Divide the measurement values of the migration behavior by the normalization factor of the corresponding region to obtain the migration density value, thereby forming the migration activity of each category of enterprises in different periods and regions. In chronological order, combine these activity values into a migration activity sequence to form a data set with a time series structure. This sequence has trend visibility, which is convenient for observing the change trajectory of the activity intensity of various categories of enterprises in a certain region.

[0227] On this basis, a difference analysis is carried out on the activity changes between consecutive periods, the change rate between adjacent periods is calculated, and a volatility index is generated by means of standard deviation normalization. This index reflects the stability and change trend of activity, and the larger the value, the stronger the volatility of the migration in and out behavior of enterprises in this category in this region. As the final evaluation result, the trend index can be used to judge whether a certain category of enterprises shows a state of continuous agglomeration, outflow or periodic fluctuation.

[0228] Through the organic combination of density normalization and trend index extraction, this embodiment realizes the unified quantitative expression of enterprise behavior in the time, space and classification dimensions, greatly improving the evaluation accuracy of enterprise classification trends and the policy guidance ability.

[0229] Among them, the spatial area parameter refers to: the physical geographical area corresponding to a certain target area, usually in "square kilometers" or "square meters" as the unit. This parameter is applicable to regional division objects such as administrative regions, functional areas, industrial parks, etc. delimited by geographical boundaries.

[0230] Its specific sources include but are not limited to:

[0231] Administrative division area data in government statistical yearbooks;

[0232] Regional vector boundaries in Geographic Information System (GIS);

[0233] Calculation results of regional fences in map platforms;

[0234] Land planning areas filed by regional construction units.

[0235] In actual operation, the system can preset the area value of each target area, and divide the number of enterprises moving in or out by this area to obtain the "migration density per unit area". For example: in a certain area, if 8 enterprises move in and the area of the region is 2 square kilometers, the average migration density is 4 enterprises per square kilometer. This processing can effectively avoid misjudging areas with "large area but sparse enterprise distribution" as active in migration.

[0236] The spatial area parameter is applicable to horizontal comparison between different administrative regions and functional areas. Especially when conducting geographical heat maps and visual displays of migration density, it provides quantitative support for the maps.

[0237] Among them, the industrial capacity parameter refers to: an estimated value of the number of enterprises or the working population that a certain area can carry, reflecting the industrial resource carrying capacity or construction density of this area. Compared with the geographical area parameter, this parameter is more applicable to functional regional units such as industrial functional areas, industrial parks, and industrial buildings.

[0238] The main ways of its construction include:

[0239] The upper limit of the number of enterprises or the design capacity set in the regional planning document;

[0240] In the historical operation data, the statistical average of the maximum number of enterprises in the region;

[0241] The conversion result of the ratio between the leasable area of the building / plant and the average land area per enterprise;

[0242] In the absence of a specific plan, the system can estimate by using the amplification ratio of the average number of enterprises moved in during the past three years.

[0243] For example: The design capacity of a high-tech industrial park is 100 enterprises, and currently 30 enterprises have moved in, then the current activity density of this region is 30%. If the capacity of another region is 200 enterprises and 50 enterprises have moved in, then its activity is 25%. Although more enterprises have moved into the second region, the density is lower, indicating that the industrial attraction efficiency of the first region is higher.

[0244] The industrial capacity parameter is particularly suitable for evaluating indicators such as the degree of industrial agglomeration, investment promotion saturation, and regional heat saturation risk, and is also convenient for comparing the operation status and potential of the same type of parks.

[0245] Among them, this method automatically selects the applicable normalization parameter according to different scenarios:

[0246] If the region is mainly divided by geographical boundaries (such as cities, districts, counties, streets), then the spatial area parameter is preferably used;

[0247] If the region is mainly divided by industrial functions (such as science and technology parks, industrial parks, building clusters), then the industrial capacity parameter is preferably used;

[0248] In some special scenarios, such as when the regional area is missing but there is a clear industrial capacity, the capacity parameter can be defaultly used instead;

[0249] The parameter can be configured and managed by the system background, and supports manual correction or external data import.

[0250] By introducing the "spatial area parameter or industrial capacity parameter" for normalization processing, the system effectively avoids the evaluation imbalance problem caused by differences in regional size or construction density, thereby enhancing the comparability and fairness of "migration activity" and "trend indicators" across regions and categories.

[0251] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for managing enterprise distribution information of enterprises migrating in and out of an industrial platform, characterized in that: The method comprises: Format the multi-source basic data sets in a standardized manner, unify the field structures of each data source, remove the fields with structural conflicts, and perform deduplication operations based on data timestamps to obtain an integrated data set; According to the integrated data set, the enterprise migration time field, enterprise location field and enterprise category field are extracted one by one; According to the enterprise migration time field, the data in the integrated data set is segmented into time segments to generate a periodic migration data set; Based on the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map; Based on the regional migration distribution map, periodic migration dataset and enterprise category field, the enterprise migration behavior is classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise in-migration and out-migration behavior; Based on the periodic migration dataset and enterprise location fields, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including: According to the enterprise location field, information about the original location and target location of the enterprise is extracted from the integrated data set, and the unstructured address data is converted into structured coordinate pairs to form a location coordinate pair set; According to the enterprise identification field in the periodic migration data set, the enterprise migration records in each period are matched with the location coordinate pair set to generate a periodic location mapping data set, wherein the periodic location mapping data set carries both time and space information; According to the periodic location mapping data set, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path data set; A directed weighted graph is constructed based on the geographic path dataset, where nodes represent regional identifiers and edges represent migration directions and migration quantities, generating a regional migration distribution map. Based on the regional migration distribution map, periodic migration dataset and enterprise category field, the enterprise migration behavior is classified and aggregated by enterprise category, and the migration activity of different types of enterprises in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out, including: According to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated; According to the set of migration records, the number of migrants in and out of each target area is aggregated and counted for each category of migration subset, and a data table with a three-dimensional structure of category-region-period is constructed. Based on this data table, the migration density function is used to calculate the migration activity of each category in each area in different periods to form a migration activity sequence; Based on the data differences of consecutive periods in the migration activity sequence, the cycle change rate is calculated, and the indicator volatility calculation formula based on standard deviation normalization is used to obtain the classification trend indicators of enterprise migration behavior in various categories and regional dimensions.

2. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 1, characterized in that: Format the multi-source basic data sets according to standard processing, unify the field structures of each data source, remove the fields with structural conflicts, and perform deduplication operations based on data timestamps to obtain an integrated data set, including: According to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain a first data set; Identify structural conflicts of repeated field contents in the first data set, retain field data with high source priority, and remove content conflict fields to obtain a second data set; According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set.

3. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 1, characterized in that: According to the semantic relevance of the field names in the multi-source basic data set, the field names of the multi-source basic data set are normalized and converted to obtain a first data set, including: Extract all field names from multi-source basic data sets to form a field name set; Based on the field name set and the preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a recommended set of standard field names; According to the recommended set of standard field names, for field names that are judged to be semantically consistent, the field names in the multi-source basic data set are replaced with the corresponding standard field names by word meaning replacement to generate a standard field set; The standard field set is arranged in the order set by the standard field structure template to generate a first data set.

4. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 3, characterized in that: The structure conflicts of repeated field contents in the first data set are identified, field data with high source priority is retained, and content conflict fields are removed to obtain a second data set, including: Extracting the original source of each field according to the field source information in the first data set, and assigning field source priorities based on a preset credibility level to form field priority data; According to the field priority data, a conflict comparison is performed on the fields in the first data set that have consistent semantics but conflicting sources, the field data with higher priority is retained, and the field values ​​with lower priority or inconsistent content are removed to generate conflict cleaning result data; The conflict clearing result data is taken as output to obtain a second data set.

5. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 4, characterized in that: According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set, including: Extracting a timestamp field and an enterprise identification field from the second data set, and combining them into a unique identification key to generate record index data; According to the record index data, based on the preset time difference tolerance interval, determine whether the data record meets the duplication condition, the duplication condition time difference is within the preset time difference tolerance interval and the enterprise identification is consistent; For data records that meet the duplication condition, the best record in the second data set is retained based on data timeliness, the remaining records are deleted, and the integrated data set is output.

6. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 1, characterized in that: According to the periodic location mapping dataset, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path dataset, including: According to the coordinate pairs of the original position and the target position contained in each enterprise migration record in the periodic location mapping data set, the original coordinate system is converted into spatial points in a unified standard geographic coordinate system to form a standardized location pair set; Based on the spatial discretization principle of grid coding, each pair of coordinates in the standardized position pair set is encoded and the coordinate value is converted into an indexable regional grid code; According to the enterprise identification field and the time period field, the migration direction of each enterprise in a specific period is represented as a path vector from the starting point code to the end point code, and the migration quantity and enterprise category attribute information are added to each path vector to construct an enterprise migration path dataset with attributes; All enterprise migration path vectors are organized into periods to generate a geographic path dataset with a time hierarchy.

7. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 6, characterized in that: A directed weighted graph is constructed based on the geographic path dataset. The nodes represent the regional identifiers, and the edges represent the migration direction and number of migrations. A regional migration distribution map is generated, including: According to the starting grid code and the end grid code carried by each migration path vector in the geographic path dataset, they are respectively assigned to standard administrative regions as node identifiers in the graph to form a regional node set; Merge multiple path vectors that belong to the same time period and have the same starting and ending areas, and calculate the sum of the migration quantities corresponding to the migration path as the edge weight of the path in the graph; The regional node set is used as the node in the graph, and the start and end region pairs after path aggregation are used as the directed edges in the graph. The directed connection relationship between the nodes is constructed, and the sum of the migration quantity is used as the weighted value of the edge to form a weighted directed graph structure. The directed weighted graph is output in the form of a graph as a regional migration distribution graph.

8. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 1, characterized in that: The migration density function is used to calculate the migration activity of each category in each region in different periods to form a migration activity sequence, including: According to the data table of the three-dimensional structure of category-region-period, the number of enterprises of each category moving in and out of each target area in each time period is extracted, and used as the measurement value of the migration behavior of enterprises of this category in this area in this period; For each target area, extract the corresponding spatial area parameter or industrial capacity parameter as the benchmark factor for density normalization; Divide the migration behavior measurement value of each category by the corresponding benchmark factor to obtain the normalized migration activity of each category in each target area and each time period; The normalized migration activity is arranged in order of time period to form a migration activity sequence.

Citation Information

Patent Citations

  • Method and device for evaluating contact demand compactness among functional areas of city

    CN110297875A

  • Public bicycle station dynamic planning method and system for Df-PBS system

    CN111489064A