Enterprise distribution information management method for migration-in and migration-out of industrial platform
Through standardized processing and classification aggregation of multi-source data sets, the response lag problem of enterprise migration information management in the industrial platform is solved, efficient and accurate enterprise distribution monitoring and trend analysis are achieved, and refined management and decision-making are supported.
Patent Information
- Application Number
- CN202510618743.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing technology has lag in response and distribution judgment deviations in the management of enterprise migration information in industrial platforms, resulting in resource allocation errors and lack of real-time and multi-source data integration capabilities.
By formatting the multi-source basic data set, unify the field structure, eliminate conflicting fields, deduplicate based on timestamps, generate integrated data sets, map them to standard coordinate systems, build regional migration distribution maps, and classify and aggregate according to enterprise categories, evaluate migration activity, and generate classified trend indicators.
It realizes structured and multi-dimensional management of enterprise migration information, enhances data coverage and accuracy, provides timeliness and geographical accuracy decision-making support, and supports refined management.
Smart Images

Figure CN120123331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for managing the enterprise distribution information of the relocation of an industrial platform. Background Art
[0002] In the prior art, for the management of the relocation information of enterprises within an industrial platform, it usually relies on manual entry or semi-automated information collection systems. Platform managers obtain enterprise migration data through methods such as enterprise reporting, registration information updates, or regular surveys, and enter it into the information management system. Most of these systems adopt a structured data storage method based on a database, using fields to identify the basic information of enterprises, migration time, original location, target location, etc., and then combining visualization tools to display the changes in enterprise distribution to assist in decision-making analysis. Some systems also introduce a Geographic Information System (GIS) to enhance the intuitive display ability of spatial distribution.
[0003] However, in scenarios involving the evaluation of investment promotion projects, the prior art may have problems of response lag and distribution judgment deviation. For example, when a certain industrial park counts the migration trend of a certain type of high-tech enterprise through the platform, due to the single data source and only relying on the information reported by enterprises themselves, if an enterprise fails to report in a timely manner or there is a delay in data entry, the system may misjudge the time node or regional scope of the concentrated migration of this type of enterprise, thus affecting the rational allocation of resources. In this case, when formulating support policies or allocating infrastructure, it is easy to make incorrect decisions based on incomplete information, exposing the technical defects of the existing system lacking real-time and multi-source data integration capabilities. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for managing the enterprise distribution information of the relocation of an industrial platform, aiming to solve the problems mentioned in the background art.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A method for managing the enterprise distribution information of the relocation of an industrial platform, the method comprising:
[0007] Performing formatting standard processing on a multi-source basic data set, unifying the field structures of each data source, removing structurally conflicting fields, and performing a de-duplication operation based on the data timestamp to obtain an integrated data set;
[0008] Extracting enterprise migration time fields, enterprise location fields, and enterprise category fields item by item according to the integrated data set;
[0009] Segmenting the data in the integrated data set according to the enterprise migration time field to generate a periodic migration data set;
[0010] Based on the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map;
[0011] Based on the regional migration distribution map, periodic migration dataset and enterprise category field, enterprise migration behaviors are classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out.
[0012] Preferably, the multi-source basic data sets are formatted in a standard manner, the field structures of the data sources are unified, the fields with structural conflicts are removed, and deduplication is performed based on the data timestamp to obtain an integrated data set, including:
[0013] According to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain a first data set;
[0014] Identify structural conflicts of repeated field contents in the first data set, retain field data with high source priority, and remove content conflict fields to obtain a second data set;
[0015] According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set.
[0016] Preferably, according to the periodic migration data set and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including:
[0017] According to the enterprise location field, the information of the original location and the target location of the enterprise is extracted from the integrated data set, and the unstructured address data is converted into structured coordinate pairs to form a location coordinate pair set;
[0018] According to the enterprise identification field in the periodic migration data set, the enterprise migration records in each period are matched with the location coordinate pair set to generate a periodic location mapping data set, wherein the periodic location mapping data set carries both time and space information;
[0019] According to the periodic location mapping data set, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path data set;
[0020] A directed weighted graph is constructed based on the geographic path dataset, with nodes representing regional identifiers and edges representing migration directions and numbers, to generate a regional migration distribution map.
[0021] Preferably, based on the regional migration distribution map, the periodic migration data set and the enterprise category field, the enterprise migration behavior is classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out, including:
[0022] According to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated;
[0023] According to the set of migration records, the number of migrants in and out of each target area is aggregated and counted for each category of migration subset, and a data table with a three-dimensional structure of category-region-period is constructed. Based on this data table, the migration density function is used to calculate the migration activity of each category in each area in different periods to form a migration activity sequence;
[0024] Based on the data differences of consecutive periods in the migration activity sequence, the cycle change rate is calculated, and the indicator volatility calculation formula based on standard deviation normalization is used to obtain the classification trend indicators of enterprise migration behavior in various categories and regional dimensions.
[0025] Preferably, according to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain the first data set, including:
[0026] Extract all field names from multi-source basic data sets to form a field name set;
[0027] Based on the field name set and the preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a recommended set of standard field names;
[0028] According to the recommended set of standard field names, for field names that are judged to be semantically consistent, the field names in the multi-source basic data set are replaced with the corresponding standard field names by word meaning replacement to generate a standard field set;
[0029] The standard field set is arranged in the order set by the standard field structure template to generate a first data set.
[0030] Preferably, structural conflicts are identified for repeated field contents in the first data set, field data with high source priority is retained, and content conflict fields are removed to obtain a second data set, including:
[0031] Extract the original source of each field according to the field source information in the first dataset, and based on a preset credibility level, assign field source priorities to form field priority data;
[0032] According to the field priority data, perform conflict comparison on the fields in the first dataset that have semantic consistency but source conflicts, retain the field data with higher priority, and eliminate the field values with lower priority or inconsistent content to generate conflict cleaning result data;
[0033] Take the conflict cleaning result data as the output to obtain the second dataset.
[0034] Preferably, according to the timestamp information corresponding to each data record in the second dataset, combined with the enterprise identification field, identify approximate duplicate records and delete redundant items to obtain an integrated dataset, including:
[0035] Extract the timestamp field and the enterprise identification field from the second dataset, and combine them into a unique identification key to generate record index data;
[0036] According to the record index data, based on a preset time difference tolerance interval, determine whether the data record meets the duplicate condition, where the duplicate condition is that the time difference is within the preset time difference tolerance interval and the enterprise identification is the same;
[0037] For the data records that meet the duplicate condition, retain the optimal record in the second dataset based on data timeliness, delete the remaining records, and output the integrated dataset.
[0038] Preferably, according to the periodic position mapping dataset, map each group of position coordinate pairs to a unified standard geographic coordinate system, and construct an enterprise migration path vector based on this to form a geographic path dataset, including:
[0039] According to the original position and target position coordinate pairs included in each enterprise migration record in the periodic position mapping dataset, convert them from the original coordinate system to spatial point positions in a unified standard geographic coordinate system to form a set of standardized position pairs;
[0040] Based on the spatial discretization principle of grid coding, perform coding processing on each pair of coordinates in the set of standardized position pairs to convert the coordinate values into indexable regional grid codes;
[0041] According to the enterprise identification field and the time period field, represent the migration direction of each enterprise within a specific period as a path vector from the start code to the end code, and attach the migration quantity and enterprise category attribute information to each path vector to construct an enterprise migration path dataset with attributes;
[0042] Organize all enterprise migration path vectors by period to generate a geographic path dataset with a time hierarchy.
[0043] Preferably, a directed weighted graph is constructed based on the geographical path dataset, where the nodes represent regional identifiers, and the edges represent the migration direction and the number of migrations, generating a regional migration distribution map, including:
[0044] According to the starting grid code and the ending grid code carried by each migration path vector in the geographical path dataset, they are respectively assigned to the standard administrative regions as the node identifiers in the graph, forming a regional node set;
[0045] For multiple path vectors belonging to the same time period and having the same starting region and ending region, a merging operation is performed, and the total sum of the number of migrations corresponding to the migration path is calculated as the edge weight of the path in the graph;
[0046] Using the regional node set as the nodes in the graph and the start and end region pairs after path aggregation as the directed edges in the graph, a directed connection relationship between the nodes is constructed, and the total sum of the number of migrations is used as the weighted value of the edge, forming a weighted directed graph structure;
[0047] The directed weighted graph is output in the form of a map as the regional migration distribution map.
[0048] Preferably, a migration density function is used to calculate the migration activity of each category in each region in different periods, forming a migration activity sequence, including:
[0049] According to the data table of the category-region-period three-dimensional structure, the number of migrations into and out of each target region by each category of enterprises in each time period is extracted, and they are respectively used as the measurement values of the migration behavior of the category of enterprises in the region in the period;
[0050] For each target region, the corresponding spatial area parameter or industrial capacity parameter is extracted as the benchmark factor for density normalization;
[0051] The measurement values of the migration behavior of each category are divided by the corresponding benchmark factor to obtain the normalized migration activity of each category in each target region and each time period;
[0052] The normalized migration activities are arranged in sequence according to the time period to form a migration activity sequence.
[0053] The above solution of the present invention has at least the following beneficial effects:
[0054] Through the implementation of this method, the structured, multi-dimensional, and standardized management of the enterprise relocation distribution information in the industrial platform is realized. Compared with the passive data collection mechanism that relies on enterprise reporting and manual registration in the existing technology, this method breaks the limitation of single information source by accessing multiple data sources and establishing a unified data fusion and analysis process. Through the integration and processing of multi-source basic data, not only the data coverage range is expanded, but also the timeliness and accuracy of data updates are significantly enhanced, improving the acquisition efficiency and integrity of enterprise migration information.
[0055] In addition, a standardized information structure mainly composed of "time field" and "location field" is constructed in this method. Combining the periodic time segmentation mechanism and the standard mapping of spatial coordinates, the structured modeling of enterprise migration behavior in the two dimensions of time and space is realized. Different from the traditional method that relies on field table display and simple GIS visualization, the regional migration distribution map generated by this solution has clear time series stratification and coordinate unity, and can accurately reflect the migration path and trend, providing more timely and geographically accurate decision-making support for policymakers.
[0056] Furthermore, this method also introduces the enterprise category field as a classification analysis dimension to conduct fine-grained aggregation of enterprise migration behavior, and quantifies and reflects the migration changes of different categories of enterprises in the target area by calculating the migration activity and trend indicators. This classification trend indicator has triple attributes of space, time, and industry, and can be used to identify the agglomeration trend, outward migration risk, and periodic fluctuations of specific industries, thus supporting the refined management of investment promotion, industrial guidance, and regional infrastructure investment.
[0057] In summary, this method has the following significant beneficial effects: enhancing the diversity and fusion ability of data sources, improving the update efficiency and accuracy of enterprise migration information; improving the information expression ability for enterprise behavior through standardized processing and spatio-temporal modeling; providing accurate and quantifiable basis for regional management and policy-making through trend quantification and classification evaluation, fundamentally solving the technical bottlenecks such as response lag, information misjudgment, and single source existing in the existing technology. This method is applicable to various application scenarios such as real-time monitoring, situation analysis, and predictive decision-making in a large-scale industrial platform environment. Brief Description of the Drawings
[0058] Figure 1 It is a flowchart of a method for managing the distribution information of enterprises relocating in and out of an industrial platform provided by an embodiment of the present invention. Detailed Embodiment
[0059] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0060] As Figure 1 shown, an embodiment of the present invention provides a method for managing enterprise distribution information for moving in and out of an industrial platform. The method includes:
[0061] S100. Obtain enterprise basic data from multiple data sources to form a multi-source basic data set;
[0062] S200. Perform formatting standard processing on the multi-source basic data set, unify the field structures of each data source, eliminate fields with structural conflicts, and perform deduplication operations based on data timestamps to obtain an integrated data set;
[0063] S300. Extract enterprise migration time fields, enterprise location fields, and enterprise category fields one by one according to the integrated data set;
[0064] S400. Segment the data in the integrated data set according to the enterprise migration time field to generate a periodic migration data set;
[0065] S500. Map the original location and the target location to a standard coordinate system according to the periodic migration data set and the enterprise location field to generate a regional migration distribution map;
[0066] S600. Classify and aggregate enterprise migration behaviors by enterprise category according to the regional migration distribution map, the periodic migration data set, and the enterprise category field, and evaluate the migration activity of different categories of enterprises in the target area within a specific period to generate a classification trend index of enterprise moving in and out behaviors.
[0067] In the embodiment of the present invention, by constructing a complete set of structured management mechanisms for enterprise moving in and out distribution information, efficient coordination of multi-source data fusion, time segmentation, spatial mapping, and classification evaluation is achieved. First, by obtaining enterprise basic data from multiple data sources, a multi-source basic data set with a wide coverage and diverse data types is formed, avoiding the sample bias problem caused by information silos. The collected data content includes, but is not limited to, enterprise registration information, migration reporting data, and statistical platform data, which are characterized by heterogeneous sources and inconsistent structures.
[0068] In the process of formatting and standardizing the multi-source basic data set, by unifying the field structures of each data source, alignment at the semantic level is achieved, making the subsequent data structures compatible. By eliminating fields with structural conflicts, information conflicts or incorrect merges caused by inconsistent field definitions can be effectively avoided, thereby improving the accuracy in the data structure cleaning stage. Further, based on the deduplication operation using data timestamps, time intersections and duplicates of records of the same enterprise in multi-source data are effectively identified and excluded, forming an integrated data set with clear structure and accurate content, ensuring the consistency and uniqueness of information in the subsequent processing.
[0069] After the integrated data set is formed, the enterprise migration time field, enterprise location field and enterprise category field can be extracted item by item to realize the conversion of the data structure to business semantic features, laying a foundation for subsequent processing. By performing time segmentation on the enterprise migration time field to form a periodic migration data set, not only is the overall migration behavior divided into several stages with time sequence, but also trend analysis and time series modeling are supported.
[0070] Further, through the combination of the periodic migration data set and the enterprise location field, the original location and the target location are mapped using spatial standardization technology, realizing the conversion of data from address representation to a standard coordinate system. On this basis, the generated regional migration distribution map has spatial continuity and visualization capabilities, enhancing the platform's monitoring ability of the dynamic distribution of enterprises within the region.
[0071] Finally, based on the regional migration distribution map, the periodic migration data set and the enterprise category field, the enterprise migration behaviors are classified and aggregated, and the migration activity is calculated. This processing not only reflects the ability of dimensionality reduction analysis of data from "enterprise behavior" to "category behavior", but also conducts cross-evaluation through the two axes of time period and spatial region, thereby outputting classification trend indicators of enterprise in-migration and out-migration behaviors. This trend indicator can clearly reveal the agglomeration trend, out-migration risk and fluctuation of a certain category of enterprises in a certain region, providing a predictive decision-making basis for industrial layout, formulation of investment promotion policies and infrastructure investment.
[0072] In summary, through unifying data structures, precise time segmentation, spatial coordinate standardization and category aggregation analysis, this method forms a full-process information management ability for enterprise migration behaviors, with high integrity, high accuracy and good spatio-temporal visualization effects.
[0073] Among them, obtaining the enterprise basic data of multiple data sources to form a multi-source basic data set specifically includes:
[0074] In the enterprise information management and analysis scenario, the collection of basic enterprise data is a prerequisite for identifying in- and out-migration behaviors. "Basic enterprise data" refers to structured or semi-structured data covering enterprise identity, registered address, business category, registration time, legal representative, contact information, migration records, and other information. Traditionally, enterprise information data often comes from a single registration platform, such as an industrial and commercial registration database, an intra-park online reporting platform, or a management system that relies on active reporting by enterprises. These methods all have problems such as limited information dimensions, delayed updates, and insufficient integrity.
[0075] To overcome the above limitations, this method constructs a data access module to collect enterprise basic data from multiple heterogeneous data sources to form a unified data set. The data sources may include but are not limited to the following categories:
[0076] Government information systems (such as business registration, tax filing, land use filing, etc.);
[0077] Third-party enterprise database platforms (such as credit information disclosure systems, business data providers);
[0078] The park builds its own reporting platform;
[0079] Public corporate information collected by web crawlers;
[0080] Data sharing interface for industry associations and chambers of commerce, etc.
[0081] During the collection process, data can be pulled by combining API data interfaces, batch data import, database mirroring synchronization, etc. There may be differences in the enterprise data fields in each data source, such as inconsistent field names, field orders, field types, or even some fields are missing or expressed differently. In this method, all the collected original enterprise information data are organized and compiled into a "multi-source basic data set". This data set has not yet dealt with field conflicts and duplications in the formation stage. Its core goal is to provide basic input for subsequent field normalization, conflict cleaning, and deduplication operations.
[0082] In this way, we avoid reliance on individual enterprise declarations, broaden the breadth of data sources, significantly improve the coverage and update frequency of enterprise basic information, and lay a data foundation for the comprehensive analysis of enterprise migration behavior.
[0083] In a preferred embodiment of the present invention, a multi-source basic data set is formatted in a standard manner, the field structures of various data sources are unified, the structural conflict fields are removed, and deduplication is performed based on the data timestamp to obtain an integrated data set, including:
[0084] According to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain a first data set;
[0085] Identify structural conflicts of repeated field contents in the first data set, retain field data with high source priority, and remove content conflict fields to obtain a second data set;
[0086] According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set.
[0087] In an embodiment of the present invention, in the multi-source data processing stage, facing multiple data sources with different structures and fields, the field names are first normalized and converted through the analysis method of field semantic relevance. After extracting all the field names in the multi-source basic data set, a field name set is formed, and then these field names are semantically compared in combination with the preset domain vocabulary comparison table to identify those fields with the same meaning but different names. This process can not only effectively identify synonymous fields such as "registered capital" and "total capital", but also common variant fields such as "company address" and "registered address".
[0088] By replacing the original field names with the recommended standard field names, the field naming system is unified, and a data set with consistent field semantics is constructed. After the unified field naming is completed, the field order is arranged according to the preset field structure template, thereby constructing the first data set with standardized field order. This data set has achieved consistent field semantics and unified structural arrangement, and is operable for subsequent structural conflict identification.
[0089] Next, we further identify and process the structure conflicts in the first data set for the cases where there are repeated field contents. By analyzing the field source information, assigning the priority of the field source, and combining the field content consistency judgment, we remove the field values with low source priority or inconsistent content, thereby generating the structure-cleaned data. For the case where there are fields with the same semantics but different values, such as the conflict between "manufacturing" and "machine manufacturing" in "industry", the field value with accurate semantic expression will be retained according to the set priority, which improves the consistency and authority of the field content.
[0090] In addition, the records are deduplicated based on the joint identification key of the timestamp and enterprise identification field of each data record in the second data set. In this process, whether it constitutes a duplication is determined based on the time difference tolerance interval and the consistency of the enterprise identification. For records that meet the duplication conditions, the most recent record items are retained and redundant items are eliminated to ensure that the final data content has time uniqueness and data update, thereby improving the quality of data processing.
[0091] This process plays a core role in the entire multi-source data standardization processing chain, laying a reliable data foundation for subsequent data mining, behavior analysis, and indicator modeling. The integrated dataset processed by this method achieves a high degree of consistency in multiple aspects such as field structure, field value accuracy, and data uniqueness, providing stable support for the tracking and analysis of enterprise relocation behaviors in the industrial platform.
[0092] In a preferred embodiment of the present invention, according to the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including:
[0093] According to the enterprise location field, extract the information of the enterprise's original location and target location from the integrated dataset, and convert the unstructured address data into structured coordinate pairs to form a set of location coordinate pairs;
[0094] According to the enterprise identification field in the periodic migration dataset, match the enterprise migration records in each period with the set of location coordinate pairs to generate a periodic location mapping dataset, and the periodic location mapping dataset carries dual information of time and space;
[0095] According to the periodic location mapping dataset, map each set of location coordinate pairs to a unified standard geographic coordinate system, and construct an enterprise migration path vector based on this to form a geographic path dataset;
[0096] Construct a directed weighted graph based on the geographic path dataset, where the nodes represent regional identifiers, and the edges represent the migration direction and the number of migrations to generate a regional migration distribution map.
[0097] In the embodiment of the present invention, through the joint processing based on the periodic migration dataset and the enterprise location field, a complete spatial expression path from address information to geographic coding and then to a directed weighted graph is constructed. First, extract the original location and target location data of the enterprise from the integrated dataset, and extract information based on the enterprise location field. Since the original data often contains unstructured addresses (such as "No. XX Road, Chaoyang District, Beijing"), an address parsing mechanism is used to convert it into a standardized coordinate pair, and a structured coordinate pair set of "original location - target location" is constructed.
[0098] Subsequently, through the enterprise identification field in the periodic migration dataset, accurately match the migration records of each enterprise in a specific period with the set of location coordinate pairs. Each matching result not only binds the migration direction of the enterprise but also carries the specific period information when the migration occurs, thus forming a periodic location mapping dataset with dual dimensions of time and space.
[0099] To achieve unified expression across regions, each pair of coordinates in the periodic location mapping dataset is further converted into points in a unified standard geographic coordinate system. This conversion uses geographic projection conversion rules to handle the mapping between the original coordinates and the target coordinates, avoiding spatial errors caused by different coordinate systems in the multi-source address resolution results. Subsequently, the spatial discretization principle of grid coding is used to convert each pair of standardized coordinates into indexable grid codes, ensuring controllable spatial granularity and efficient grouping and retrieval capabilities.
[0100] On this basis, according to the enterprise identification field and the time period field, the migration behavior of each enterprise within a specific period is encoded into a path vector of "starting point code → ending point code", and attribute values such as the migration quantity and enterprise category are embedded in each path, forming an enterprise migration path dataset with attributes. This process realizes the dimensionality elevation processing of migration information from data records to analyzable paths.
[0101] Finally, all path vectors are sorted by period to generate a geographic path dataset at the time level, making the data no longer just a pile of isolated events, but a structured spatio-temporal path system. Through this path system, the platform can clearly display the spatial flow trends of enterprises in different periods on the map, effectively supporting decision-making scenarios such as subsequent regional development assessment and resource scheduling optimization.
[0102] In summary, this embodiment realizes the full-process standardization processing of address information → coordinate data → grid path → directed vector, enhancing the spatial expression ability while enhancing the analyzability and visualization ability of enterprise migration behavior at the geographical level.
[0103] Among them, according to the enterprise location field, the information of the original location and the target location of the enterprise is extracted from the integrated dataset, and the unstructured address data is converted into structured coordinate pairs to form a set of location coordinate pairs, specifically including:
[0104] In the actual enterprise basic information, the "original location" and "target location" are usually recorded in text form, which belongs to typical unstructured data. For example, data such as "No. 27, Zhongguancun Street, Haidian District, Beijing" and "Building A3, South Area of Science and Technology Park, Nanshan District, Shenzhen" are easy for humans to identify, but difficult for computers to directly perform geographical calculations or spatial analyses. Therefore, in order to realize the structured modeling of enterprise migration behavior in the spatial dimension, such unstructured address texts must be converted into computable geographical coordinate data.
[0105] This method uses address resolution technology (Geocoding) to implement the process of address-to-coordinate conversion. Address resolution is a technology that uses a geographical information database and natural language processing means to parse address texts in Chinese or other languages into standard coordinate points. In actual implementation, the address resolution can be realized through the following two types of methods:
[0106] Access commercial map service APIs such as Baidu Map, Tencent Location Service, Amap, etc., submit address text requests, and obtain the returned longitude and latitude coordinate results;
[0107] Build its own geographical word library and administrative division dictionary, and through the word segmentation matching and fuzzy comparison mechanism, perform the conversion from local address fields to coordinate points.
[0108] During the address parsing process, in order to improve the parsing accuracy and coverage rate, the following auxiliary means can also be introduced:
[0109] Preprocess the enterprise address fields, including removing redundant words (such as "office building of Co., Ltd.") and unifying the administrative region expressions (such as replacing "Chaoyang District, Beijing" with "Chaoyang District");
[0110] Set a fault tolerance mechanism for the address text with parsing failures, such as setting the principle of "falling back to the district-level coordinates in case of street-level failure", to ensure the availability of spatial data.
[0111] Once the address parsing is successful, a set of structured coordinates can be generated for the original location and target location fields of each enterprise, namely (original longitude, original latitude) and (target longitude, target latitude). Combine these coordinates into a set of structured data, called "location coordinate pairs".
[0112] Finally, using the "enterprise identifier" as the primary key and combining the original time field, add the location coordinate pairs in each enterprise migration record to the periodic location mapping dataset to form a coordinate data structure with time tags and spatial paths. This set of coordinate pairs not only supports subsequent spatial mapping and graph generation, but also has clear geographical meanings and can be directly visualized as paths on the map.
[0113] Through this step, the conversion of enterprise location fields from unstructured expressions to structured spatial data is effectively realized, endowing enterprise migration data with clear spatial computability and spatial aggregation capabilities, and providing a strong geographical basis for regional analysis, path mapping, and trend analysis.
[0114] In a preferred embodiment of the present invention, according to the regional migration distribution map, the periodic migration dataset, and the enterprise category field, classify and aggregate the enterprise migration behaviors by enterprise category, and evaluate the migration activity of different categories of enterprises in the target region within a specific period, and generate classification trend indicators for the enterprise in-migration and out-migration behaviors, including:
[0115] According to the enterprise category field, divide the enterprise migration records in the periodic migration dataset into multiple category migration subsets by category, and combine the original location and target location recorded in the regional migration distribution map to generate a set of migration record pairs divided by category and region;
[0116] For each category migration subset in the set according to the migration records, aggregate and count the number of migrated-in and migrated-out in each target region, construct a data table with a three-dimensional structure of category-region-period, and based on this data table, calculate the migration activity of each category in each region in different periods using the migration density function to form a migration activity sequence.
[0117] According to the data differences in consecutive periods in the migration activity sequence, calculate the period change rate, and use the index volatility calculation formula normalized based on the standard deviation to obtain the classification trend index of the enterprise migration-in and migration-out behavior in each category and region dimension.
[0118] In the embodiment of the present invention, by focusing on combining the periodic migration data set, the enterprise category field, and the regional migration distribution map, the refined classification and trend quantification of the enterprise migration behavior are realized. First, through the enterprise category field, the migration records in the periodic migration data set are classified by category. After the classification is completed, combined with the original location and target location carried by each migration path in the regional migration distribution map, the migration data of each category of enterprises are combined according to the migrated-in and migrated-out regions to generate a set of migration record pairs divided by the two dimensions of category and region.
[0119] Based on this set, aggregate and count the number of migrated-in and migrated-out of each category of enterprises in each target region, and organize them structurally according to the period, so as to construct a data table with a three-dimensional structure of category-region-period. This data table not only standardizes the enterprise behavior data, but also provides data support for time trend analysis through the method of period stratification.
[0120] Based on the three-dimensional structure data table, use the migration density function to model and analyze the enterprise activity. In each region and each period, normalize the number of migrated-in or migrated-out of a certain category of enterprises with the regional space resources (such as area, capacity, etc.) to obtain the normalized migration density. This processing method avoids the direct interference of the enterprise migration quantity scale difference on the evaluation result, so that the migration behaviors between different categories and different regions can be compared on a unified basis.
[0121] Next, taking the time period as the reference axis, calculate the change trend of the migration density under consecutive periods, extract the trend characteristics through the change rate of adjacent periods, and further calculate the volatility intensity index normalized by the standard deviation. This index is used to measure the volatility and directionality of the enterprise migration activity in multiple periods, so as to reflect whether there are phenomena such as stable aggregation, active fluctuation, or outward migration risk of a certain category of enterprises in a certain region.
[0122] The finally generated classification trend indicator has the three-dimensional fusion characteristics of periodicity, category, and spatiality, and can reveal the dynamic trends of enterprise migration behavior in terms of industrial structure evolution, regional attraction change, and policy response effect. Such an indicator not only shows the clustering or dispersion characteristics of enterprise categories at the spatial distribution level but also captures the industry change rules in the time series, thus providing a strong quantitative decision-making basis for the government's investment promotion, industrial planning formulation, and infrastructure allocation.
[0123] Therefore, through this embodiment, an enterprise behavior evaluation mechanism based on classification aggregation and trend quantification is established, forming a complete closed-loop from data division, indicator construction to trend output, with high stability, high adaptability, and reference value for policies.
[0124] Among them, according to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated, specifically including:
[0125] In the industrial platform, the migration behavior of enterprises usually has significant category differences. For example, the migration paths, cycles, and destination tendencies of manufacturing enterprises may be completely different from those of the financial service industry. Therefore, if we hope to deeply evaluate the migration trend of specific category enterprises, using only the overall enterprise sample as the analysis object will not be able to depict the fine-grained changes, and it is necessary to further classify and disassemble the migration behavior.
[0126] For this reason, this method preliminarily groups the migration data based on the enterprise category field carried by each enterprise migration record in the periodic migration dataset. The enterprise category field usually comes from the "industry" field in the industrial and commercial registration information or the platform custom classification field, such as "manufacturing", "information service", "new energy", "medical device", etc. This grouping process can be achieved through string matching, industry code mapping (such as GB / T 4754), or a custom label system.
[0127] After classification, the migration dataset is disassembled into several category migration subsets within each time period. Each subset only contains the migration records within the same enterprise category. On this basis, in order to combine the migration behavior with the spatial flow situation, the original location and target location coding information included in the regional migration distribution map is further referenced to attach a "starting area identifier" and an "ending area identifier" to each record.
[0128] In this way, each enterprise migration record will have the following structural information: enterprise category, migration cycle, original region, target region, and migration quantity. Based on this, a new data set is generated, called the "set of migration record pairs classified by category and region". This set supports two-dimensional aggregation (category × region), which can not only count the number of enterprises of a certain category moving into a certain region but also reflect the spatial flow of the enterprises of this category moving out.
[0129] The above operations achieve the dimensionality increase from the "enterprise record level" to the "regional path level" at the data structure level and realize the cross-integration of "category behavior" and "regional relationship" in the analysis logic, constructing a basic data system for subsequent classification trend modeling.
[0130] Among them, the calculation formula for the index volatility is: , is the volatility of the index within the category in the region , that is, the classification trend index; is the total number of time periods, such as quarters or months; is the change in activity, representing the difference in activity between the current period and the previous period, and the calculation method is as follows: , is the normalized migration activity of the category in the region and the period ; is the average activity of all periods of the category in the region , and the calculation method is: ; is the period weight coefficient, representing the fluctuation importance weight of the period , and can be defined as: , which is the sigmoid function structure; among them, is the spatial migration slope, representing the offset rate of the migration behavior in space within the period, and the calculation method is: , is the average migration path length of the category in the region and the period , is the period length, used to standardize the time span, is the adjustment coefficient set by the system, used to control the steepness of the weight curve.
[0131] In this formula, for each category in the periodic migration dataset enterprises in the target area at different time periods the normalized migration activity values shown constitute , and then based on this, the change in activity between two consecutive periods is extracted. To eliminate the influence brought by the difference in the absolute value of activity between different categories, the system uses the average activity of each category in this area as the normalization denominator, normalizes and scales the activity change value, and squares and averages the results of all periods to measure the intensity of trend changes. In this process, to reflect the influence weight of different periods in the overall fluctuation trend, the system introduces a weight coefficient , which is determined by the spatial migration slope of the enterprise within the period and reflects the spatial intensity of the migration path within this period. Finally, the result output by this formula is defined as the classification trend indicator of enterprises of this category in this area. The larger the value, the stronger the fluctuation amplitude and the more unstable the trend; the smaller the value, the more stable or continuous the enterprise behavior is in the time series.
[0132] The introduction and calculation of this main formula have the following significant technical advantages and beneficial effects:
[0133] Realize the quantitative expression of trend behavior. Compared with traditional statistical methods that can only output the static comparison of the number of in-migrations and out-migrations, this formula can dynamically depict the behavior fluctuations of enterprises in the cycle dimension by extracting the continuous differences in the time series and normalizing them, solving the problem of overly rough trend judgment in the existing methods.
[0134] Enhance the comparability across categories and regions. The normalization part eliminates the interference brought by the difference in enterprise base through the average activity factor , enabling the migration activities of different types of enterprises such as manufacturing and service industries to be analyzed for trends on the same scale. This method is particularly suitable for horizontal benchmarking and performance evaluation among multiple industrial parks and different investment promotion regions.
[0135] Introduce spatial behavior factors for weighted enhancement. Traditional migration trend assessments often ignore the difference in the length of the spatial movement path of enterprises and only consider the quantity changes. However, this method calculates the spatial migration slope by introducing the path length and cycle span, and constructs a weight through the Sigmoid function to enhance the sensitivity to spatially intense behavior changes, making the trend indicator better reflect the reconstruction intensity of the enterprise group in the physical distribution. Support behavior prediction and early warning mechanisms. The index volatility It can be used as a dynamic monitoring variable in []. If the trend index of a certain type of enterprise in a certain region continuously exceeds the historical threshold, it can automatically trigger analysis warnings or resource scheduling suggestions, improving the platform's perception ability and response efficiency to industrial dynamic changes.
[0136] It has good formula extensibility and calculation adaptability. Since all variables are derived from the data sets clearly defined in the claims (such as periodic migration data sets, enterprise category fields, geographical path data sets), this formula can be directly integrated into the existing system architecture without introducing external parameters; at the same time, its form supports the future expansion of more weight dimensions (such as economic indicators, energy consumption constraints, etc.).
[0137] In a preferred embodiment of the present invention, according to the semantic relevance of the field names in the multi-source basic data set, the field names of the multi-source basic data set are normalized and transformed to obtain a first data set, including:
[0138] Extract all field names in the multi-source basic data set to form a field name set;
[0139] According to the field name set, combined with a preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a standard field name recommendation set;
[0140] According to the standard field name recommendation set, for the field names determined to be semantically consistent, use the word meaning replacement method to replace the field names in the multi-source basic data set with the corresponding standard field names to generate a standard field set;
[0141] Arrange the fields in the standard field set in the order set by the standard field structure template to generate a first data set.
[0142] In the embodiment of the present invention, by focusing on realizing the standardization of the multi-source basic data set at the field structure level, the semantic barrier between data from different sources is broken through. In the actually collected multi-source data, due to the heterogeneity of information sources, different data tables often use different naming methods for the same attribute field. For example, some sources use "registered capital", some write "capital value", and some are called "fund scale". This situation where the semantics are similar but the names are not unified seriously interferes with subsequent data fusion operations.
[0143] In order to achieve consistent expression of field names, all field names in the multi-source basic data set are first extracted to form a field name set. This set not only contains the original field names, but also serves as the target set for semantic standardization. During the processing, the fields in the field name set are semantically judged one by one through the preset domain vocabulary comparison table, and those field pairs with similar meanings but different names are identified, and the recommended standard field names are further determined. For example, "company establishment time", "registration time" and "establishment date" are uniformly replaced with "enterprise establishment time".
[0144] Next, based on the recommended set of standard field names, a unified replacement operation is performed in the multi-source basic data set for all field names that are identified as semantically consistent. Through this process, each field name in the original data set is replaced with the corresponding standard field name, forming a standard field set, achieving field semantic normalization and naming consistency. This operation not only enhances the uniformity of the field structure, but also eliminates the risk of data merging conflicts caused by field duplication or ambiguity.
[0145] In addition, to ensure the structural uniformity of the field arrangement, the standard field set was rearranged in the order set by the preset field structure template. The field structure template is a field arrangement benchmark based on industry standards or business modeling rules, which ensures that the order of fields in each data source remains consistent after being uniformly named. Through the dual processing of field name unification and order unification, the first data set finally generated has a highly consistent field structure.
[0146] This embodiment significantly improves the consistency and stability of data structure processing, provides a solid structural foundation for subsequent operations such as structural conflict identification, field content fusion, and duplicate record removal, and greatly reduces the workload of manual cleaning and data comparison, thereby improving processing efficiency.
[0147] The preset domain vocabulary comparison table is a semantic mapping resource that is manually organized and automatically expanded. It is specifically used to unify fields with similar semantics but different names into standard field names. Its construction principles mainly include the following two methods:
[0148] Manual definition method: Based on industry data standards (such as national standards GB / T22240, GB / T4754, etc.), industry databases (such as business registration systems), and historical project experience, high-frequency fields and their common synonyms are sorted out and manual mapping relationships are established.
[0149] Semantic computing expansion method: Use natural language processing technology, such as word vector model, edit distance matching, and word sense disambiguation algorithm, to calculate the semantic similarity of the existing field name set, automatically discover candidate mapping relationships, and add them to the comparison table after manual review.
[0150] The comparison table is saved in the structure of "standard field name → multiple candidate field names", for example:
[0151] "Registered capital" → ["Capital value", "Fund scale", "Registered funds"]
[0152] "Establishment time" → ["Registration time", "Company establishment date", "Enterprise founding date"]
[0153] During the field normalization process, the system can match whether the field name appears in this table one by one. If a certain candidate name is hit, it will be replaced with the corresponding standard field name. This comparison table serves as the basic configuration file before system deployment and can be customized and maintained according to industries, projects or usage scenarios.
[0154] Among them, the standard field structure template refers to the structured specification used to constrain the arrangement order of data table fields and the content of the field set, so as to unify the expression methods of different data sources at the field level, field order and field existence, ensure that the normalized data has consistent structural semantics, and facilitate the subsequent processing module to call and display.
[0155] Its basic composition includes:
[0156] Field set definition: Clearly define which standard fields are included, such as "Enterprise name", "Enterprise type", "Legal representative", "Registered capital", "Original address", "Target address", "Industry classification", etc.;
[0157] Field order convention: To avoid data processing anomalies caused by different field arrangements, the template specifies the ordinal position of each field, so as to achieve unified sorting;
[0158] Field attribute description: Set attributes such as type (such as string, date, floating point number), length limit, and whether it can be empty for each field, providing a basis for subsequent data verification and interface mapping.
[0159] The template can be implemented in the form of JSON, XML or database table structure, and serves as the output format specification after field normalization. After the field names are unified, the system arranges the fields in the order specified in this template, fills in the missing fields with null values, discards or marks the redundant fields, and finally outputs the first data set with consistent field structures.
[0160] In a preferred embodiment of the present invention, structural conflicts of duplicate field contents in the first data set are identified, the field data with a higher source priority is retained, and the content conflict fields are removed to obtain the second data set, including:
[0161] Extract the original source of each field according to the field source information in the first dataset, and based on a preset credibility level, assign a field source priority to form field priority data;
[0162] According to the field priority data, perform a conflict comparison on the fields in the first dataset that have the same semantics but conflicting sources, retain the field data with a higher priority, and eliminate the field values with a lower priority or inconsistent content to generate conflict cleaning result data;
[0163] Use the conflict cleaning result data as the output to obtain the second dataset.
[0164] In the embodiments of the present invention, it mainly deals with the content-level conflict problems that may still exist after the field structures in the first dataset are unified, that is, the situation where "the field structures are the same but the field values are different". In practical applications, different data sources may record different values for the same field. For example, for the field of "legal representative", one data source may record it as "Zhang San", while another data source may record it as "Zhang Sanfeng". If merged directly without discrimination, it will cause data inconsistency and even incorrect use.
[0165] Therefore, based on the first dataset, first extract the source information of each field, including the source data source identifier, collection time, credibility level, etc. According to the set credibility evaluation rules, assign priorities to different sources. For example, government filing data is superior to web scraping data, and manually reviewed data is superior to automatically filled data. Thus, field source priority data is generated to provide a weight reference for conflict processing.
[0166] Subsequently, identify the records with the same field name but different field values, and perform a conflict comparison on these records. By comparing the source priorities, retain the data from the source with a higher priority, and at the same time eliminate the field values with a lower priority or inconsistent semantics to generate structure conflict cleaning result data. If the field value content is the same, it is determined as a redundant field, and the remaining duplicate values are deleted after retaining one copy. By processing in this way, not only the problem of inconsistent field values is solved, but also the de-duplication at the field value level is completed.
[0167] Finally, through the double cleaning operations of conflict identification and redundant field elimination, a second dataset with clear structure, consistent field values, and complete semantics is obtained. This dataset has high comparability and fusion, providing data guarantee for the accurate management and trustworthy modeling of enterprise data.
[0168] This embodiment effectively improves the consistency and accuracy of field content, avoids interference with the analysis results due to structure conflicts or redundant fields, and provides a reliable data source for constructing a highly trustworthy enterprise information portrait.
[0169] Among them, the main construction methods of the preset credibility level include:
[0170] Source classification dimension: The data sources are classified into "government official", "third-party commercial", "platform filling", "crawler scraping", etc.;
[0171] Timeliness dimension: Considering the latest update time of the data, the closer to the current date, the higher the credibility;
[0172] Manual review flag: Determine whether the data has been manually verified or business verified. If the review has passed, the credibility is improved.
[0173] Each source is assigned a level label, such as "high (level 1)", "medium (level 2)", "low (level 3)". When dealing with field conflicts, if there are value conflicts in multiple fields, the data with a higher level is retained, and the data with a lower level is deleted or ignored. For example, when there is a conflict in the "legal representative" field value, if one source is the industrial and commercial registration system (high level) and the other is the enterprise website (medium level), the former is retained.
[0174] This mechanism ensures that the data selection in the merging and cleaning processes favors more reliable sources by introducing a credibility ranking strategy, improving the overall quality and consistency of the integrated data.
[0175] In a preferred embodiment of the present invention, according to the timestamp information corresponding to each data record in the second dataset, combined with the enterprise identification field, approximate duplicate records are identified, and redundant items are deleted to obtain an integrated dataset, including:
[0176] Extract the timestamp field and the enterprise identification field from the second dataset, and combine them into a unique identification key to generate record index data;
[0177] Based on the record index data, within a preset time difference tolerance interval, determine whether the data record meets the duplicate condition, where the duplicate condition is that the time difference is within the preset time difference tolerance interval and the enterprise identification is the same;
[0178] For the data records that meet the duplicate condition, retain the optimal record in the second dataset based on data timeliness, delete the remaining records, and output the integrated dataset.
[0179] In the embodiment of the present invention, attention is paid to the record-level duplicate problems that may still exist in the second dataset, especially the time-proximity duplicate records caused by synchronization delays and different update frequencies between different data sources. Such records usually have the same field structure and enterprise identification, but there are slight differences in data content or timestamps. If deduplication is not performed in a timely manner, it will directly affect the accuracy and timeliness of data analysis.
[0180] First, extract the timestamp field and the enterprise identifier field contained in each data record from the second dataset, and combine the two to form a unique identification key. This identification key is used to construct a data index set, which can quickly retrieve data records with the same enterprise identifier in the index structure.
[0181] Based on the index, for all data records under each enterprise identifier, execute the duplicate judgment logic based on the time difference tolerance interval. By setting an adjustable time difference threshold, such as "48 hours" or "3 days", determine whether the records are close enough in time. If the enterprise identifiers are the same and the time difference is less than the threshold, it is judged as a duplicate record, constituting a duplicate condition.
[0182] For records that meet the duplicate condition, further execute the strategy of retaining the optimal record. Specifically, give priority to retaining the data record with the latest timestamp and delete the remaining redundant items, so as to ensure that the retained data has the highest timeliness and the most complete field content. All the cleaned results are summarized into a new integrated dataset.
[0183] In this embodiment, through the enterprise identifier and the time joint key, accurate identification and deletion of approximate records are realized, effectively avoiding the problem of redundant accumulation caused by multi-source synchronization delay. At the same time, the latestness and uniqueness of the final dataset in the time dimension are ensured, significantly improving the data quality and application value.
[0184] Among them, the preset time difference tolerance interval, that is, a maximum allowable error range is set when comparing the timestamp fields of records, is used to judge whether it belongs to "approximate duplication".
[0185] This interval can be set based on the actual business scenario, such as:
[0186] For a system with a daily update frequency, the time difference tolerance interval can be set to 1 day;
[0187] For a monthly update or near-real-time system, the tolerance interval can be set to 7 days or longer;
[0188] For the historical record data imported in batches, the tolerance interval can be set according to the record coverage period.
[0189] When comparing two records with the same enterprise identifier, the system will calculate the difference between the timestamps of the two records. If the difference is less than or equal to the tolerance interval, the system considers these two records to be approximately duplicate and enters the deduplication process, retaining the record with higher timeliness.
[0190] This tolerance interval can be dynamically set as a system configuration item, or can be configured differently for different data sources according to the characteristics of the data source, providing a more flexible strategy support for enterprise-level data deduplication and data uniqueness guarantee.
[0191] In a preferred embodiment of the present invention, according to the periodic position mapping data set, each set of position coordinate pairs is mapped to a unified standard geographic coordinate system, and an enterprise migration path vector is constructed based on this to form a geographic path data set, including:
[0192] According to the original position and target position coordinate pairs included in each enterprise migration record in the periodic position mapping data set, convert them from the original coordinate system to spatial points in the unified standard geographic coordinate system to form a standardized position pair set;
[0193] Based on the spatial discretization principle of grid coding, encode each pair of coordinates in the standardized position pair set, and convert the coordinate values into indexable regional grid codes;
[0194] According to the enterprise identification field and time period field, represent the migration direction of each enterprise within a specific period as a path vector from the starting code to the ending code, and attach the migration quantity and enterprise category attribute information to each path vector to construct an enterprise migration path data set with attributes;
[0195] Organize all enterprise migration path vectors by period to generate a geographic path data set with a time hierarchy.
[0196] In the embodiment of the present invention, after processing the periodic position mapping data set, in order to further enhance the spatial expression ability of enterprise migration behavior, this embodiment maps each set of original position and target position coordinate pairs to a unified standard geographic coordinate system, and constructs an enterprise migration path vector based on this to form a geographic path data set. First, extract the original position and target position coordinate information included in the enterprise migration records from the periodic position mapping data set, and convert them from the original acquisition coordinate system (which may be local or non-uniform standard) to a unified standard geographic coordinate system (such as WGS84 or GCJ02). This standardized conversion eliminates the spatial offset problem caused by coordinate differences and ensures the visualization processing of all position points under a unified reference system.
[0197] Subsequently, to enhance the spatial grouping and indexing efficiency, process the standardized coordinate pairs based on the spatial discretization principle of grid coding. By mapping the coordinate values to uniquely identifiable grid codes (such as GeoHash, Quadtree, or hexagonal coding systems), the spatial discretized expression of the migration path is realized, and operations such as fast geographic range filtering and regional level aggregation are supported.
[0198] Based on the above encoding results, further using the enterprise identification field and the time period field as index keys, represent the migration path of each enterprise within a specific period as a path vector of "starting point encoding → ending point encoding", and embed additional information such as the migration quantity and enterprise category into the path attributes to form migration vector data with attributes. These attribute information not only supports spatial structure mapping but also provides operable dimensions for subsequent multi-dimensional analysis (such as industry and migration intensity).
[0199] All migration path vectors are then organized and layered according to the time period to construct a geographical path dataset with a time hierarchy. This dataset not only has the accuracy of the spatial dimension but also retains the time sequence characteristics, and has multiple application values such as path tracking, trend judgment, and classification aggregation. On the map platform, dynamic path animations, periodic migration heat displays, and visual expressions of enterprise migration flow maps can be realized, greatly enhancing the platform's perception ability and presentation efficiency of regional enterprise activities.
[0200] Through this embodiment, not only are discrete data records converted into a structured migration vector network, but also a high degree of integration is achieved in terms of spatial expression, attribute extension, and time series management, significantly enhancing the usability, visualization, and structured management capabilities of enterprise migration information.
[0201] In a preferred embodiment of the present invention, a directed weighted graph is constructed based on the geographical path dataset, where the nodes represent regional identifiers, and the edges represent the migration direction and the migration quantity, generating a regional migration distribution map, including:
[0202] According to the starting grid encoding and the ending grid encoding carried by each migration path vector in the geographical path dataset, assign them to the standard administrative regions respectively as the node identifiers in the graph to form a regional node set;
[0203] Perform a merging operation on multiple path vectors that belong to the same time period and have the same starting region and ending region, calculate the total sum of the migration quantities corresponding to this migration path, and use it as the edge weight of this path in the graph;
[0204] Using the regional node set as the nodes in the graph and the start and end region pairs after path aggregation as the directed edges in the graph, construct a directed connection relationship between the nodes, and use the total sum of the migration quantities as the weighted value of the edge to form a weighted directed graph structure;
[0205] Output the directed weighted graph in the form of a map as the regional migration distribution map.
[0206] In the embodiments of the present invention, after forming the enterprise migration path vectors, in order to further refine the macro flow direction of enterprise migration between regions, this embodiment constructs the geographical path dataset into a structured directed weighted graph to represent the migration relationship between regions. First, extract the grid coding information of the starting point and the ending point included in each path in the geographical path dataset, and map the grid coding to the corresponding standard administrative region according to the preset regional aggregation mapping rules. This mapping relationship is completed based on the regional boundary division table, supporting precise matching at the levels of cities, districts and counties, and even industrial parks, so as to form a set of regional nodes with administrative units as node identifiers.
[0207] In the process of constructing the graph structure, aggregate and statistically analyze the path vectors with the same starting region and ending region in each time period, and calculate the total sum of their migration quantities as the weight of the edge of this path in the graph. This path aggregation operation not only eliminates the redundant information in the enterprise individual dimension, but also effectively extracts the overall migration intensity between regions, forming a directed edge with "region pair" as the connection relationship and "total migration volume" as the weight.
[0208] Construct the edges of the graph with the aggregated starting and ending region pairs, and form a complete directed weighted graph structure with all region identifiers as the nodes of the graph. Among them, the nodes represent geographical regions, the edges represent the migration direction, and the weights represent the number of enterprises migrating from one region to another region within a specific period, with clear spatial flow meaning. The quantity and direction changes of the edges in the graph can dynamically reflect the enterprise flow relationship between regions, so as to identify the in-migration aggregation areas, out-migration high-incidence areas, and areas with intensive migration path frequencies.
[0209] Finally, output the graph structure in the form of a regional migration distribution map, supporting map display methods, including force-directed layout, geographical layout, heat path flow map, etc., and can clearly present the main paths and hot spots of the enterprise migration network on the interface. This map expression form not only enhances the visualization effect, but also supports the insight into the industrial agglomeration trend, the effect of policy attractiveness, and the rationality of spatial resource allocation from the macro level.
[0210] This embodiment not only realizes the dimensionality-raising expression from enterprise-level path vectors to regional-level network structures, but also greatly improves the interpretability and decision-making support value of regional migration patterns.
[0211] In a preferred embodiment of the present invention, a migration density function is used to calculate the migration activity of each category in each region in different periods, forming a migration activity sequence, including:
[0212] According to the data table of the category-region-period three-dimensional structure, extract the number of enterprises moving into and out of each target region by each category of enterprises in each time period, and use them as the measurement values of the migration behavior of this category of enterprises in this region and this period respectively;
[0213] For each target region, extract the corresponding spatial area parameter or industrial capacity parameter as the reference factor for density normalization;
[0214] Divide the measurement values of the migration behavior of each category by the corresponding reference factor to obtain the normalized migration activity of each category in each target region and each time period;
[0215] Arrange the normalized migration activities in sequence according to the time period to form a migration activity sequence.
[0216] In the embodiment of the present invention, after completing the classification of enterprise categories and the modeling of regional migration paths, this embodiment further realizes the quantification of the classification trend of the in-and-out behavior of enterprises. The core method is to use the migration density function to perform normalization calculations on the activities of each category of enterprises in different regions and periods, form a migration activity sequence, and extract its trend change characteristics.
[0217] First, according to the data table of the category-region-period three-dimensional structure, extract the number of enterprises moving into and out of each category of enterprises in each time period and each target region one by one, and use them as the measurement values of the migration behavior. This processing ensures the spatial accuracy and periodic stability of the data, and eliminates the deviation caused by inconsistent time statistical granularity.
[0218] In order to achieve horizontal comparability in different proportion regions and industries, a spatial normalization parameter is introduced, that is, the spatial area or its industrial capacity of each target region is extracted as the reference factor for density calculation. The spatial area is mainly used for normalization under the physical scale, while the industrial capacity is used for the evaluation of the carrying capacity dimension, and the two are switched according to the platform configuration.
[0219] Divide the measurement values of the migration behavior by the normalization factor of the corresponding region to obtain the migration density value, so as to form the migration activity of each category of enterprises in different periods and regions. In the order of time, combine these activity values into a migration activity sequence to form a data set with a time series structure. This sequence has trend visibility, which is convenient for observing the change trajectory of the activity intensity of each category of enterprises in a certain region.
[0220] On this basis, a differential analysis is conducted on the activity changes between consecutive periods, the change rate between adjacent periods is calculated, and a volatility index is generated by means of standard deviation normalization. This index reflects the stability and change trend of activity, and the larger the value, the stronger the volatility of the in-migration and out-migration behaviors of enterprises in this category in this region. As the final evaluation result, the trend index can be used to judge whether a certain category of enterprises shows a state of continuous agglomeration, outflow, or periodic fluctuation.
[0221] Through the organic combination of density normalization and trend index extraction, this embodiment realizes the unified quantitative expression of enterprise behavior in the time, space, and classification dimensions, greatly improving the evaluation accuracy of enterprise classification trends and the policy guidance ability.
[0222] Among them, the spatial area parameter refers to: the physical geographical area corresponding to a certain target area, usually in units of "square kilometers" or "square meters". This parameter is applicable to regional division objects such as administrative regions, functional areas, and industrial parks delimited by geographical boundaries.
[0223] Its specific sources include but are not limited to:
[0224] Administrative division area data in government statistical yearbooks;
[0225] Regional vector boundaries in Geographic Information System (GIS);
[0226] Calculation results of regional fences in map platforms;
[0227] Land planning areas filed by regional construction units.
[0228] In actual operation, the system can preset the area value of each target area, and divide the number of enterprises moving in or out by this area to obtain the "migration density per unit area". For example: in a certain area, if 8 enterprises move in and the area of the region is 2 square kilometers, the average in-migration density is 4 enterprises per square kilometer. This processing can effectively avoid misjudging areas with "large area but sparse enterprise distribution" as active in-migration areas.
[0229] The spatial area parameter is applicable to horizontal comparisons between different administrative regions and functional areas. Especially when conducting geographical heat maps and visual displays of migration density, it provides quantitative support for the maps.
[0230] Among them, the industrial capacity parameter refers to: an estimated value of the number of enterprises or the working population that a certain area can carry, reflecting the industrial resource carrying capacity or construction density of this area. Compared with the geographical area parameter, this parameter is more applicable to functional regional units such as industrial functional areas, industrial parks, and industrial buildings.
[0231] The main ways of its construction include:
[0232] The upper limit of the number of enterprises or the designed capacity set in the regional planning document;
[0233] In the historical operation data, the statistical average of the maximum number of enterprises in the region;
[0234] The conversion result of the ratio between the rentable area of the building / plant and the average land area per enterprise;
[0235] In the absence of a specific plan, the system can estimate by using the amplification ratio of the average number of enterprises moved in during the past three years.
[0236] For example: The designed capacity of a high-tech industrial park is 100 enterprises, and currently 30 enterprises have moved in, then the current activity density of this region is 30%. If the capacity of another region is 200 enterprises and 50 enterprises have moved in, then its activity is 25%. Although more enterprises have moved into the second region, the density is lower, indicating that the first region has a higher industrial attraction efficiency.
[0237] The industrial capacity parameter is particularly suitable for evaluating indicators such as the degree of industrial agglomeration, investment promotion saturation, and regional heat saturation risk, and is also convenient for comparing the operation status and potential of the same type of parks.
[0238] Among them, this method automatically selects the applicable normalization parameter according to different scenarios:
[0239] If the region is mainly divided by geographical boundaries (such as cities, districts and counties, streets), then the spatial area parameter is preferably used;
[0240] If the region is mainly divided by industrial functions (such as science and technology parks, industrial parks, building clusters), then the industrial capacity parameter is preferably used;
[0241] In some special scenarios, such as when the regional area is missing but there is a clear industrial capacity, the capacity parameter can be defaultly used instead;
[0242] The parameter can be configured and managed by the system background, and supports manual correction or external data import.
[0243] By introducing the "spatial area parameter or industrial capacity parameter" for normalization processing, the system effectively avoids the evaluation imbalance problem caused by differences in regional size or construction density, thereby enhancing the comparability and fairness of "migration activity" and "trend indicators" across regions and categories.
[0244] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for managing enterprise distribution information of enterprises migrating in and out of an industrial platform, characterized in that: The method comprises: Perform standard formatting on multi-source basic data sets, unify the field structures of each data source, remove structural conflicting fields, and perform deduplication operations based on data timestamps to obtain an integrated data set; According to the integrated data set, the enterprise migration time field, enterprise location field and enterprise category field are extracted one by one; According to the enterprise migration time field, the data in the integrated data set is segmented into time segments to generate a periodic migration data set; Based on the periodic migration dataset and the enterprise location field, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map; Based on the regional migration distribution map, periodic migration dataset and enterprise category field, enterprise migration behaviors are classified and aggregated by enterprise category, and the migration activity of enterprises of different categories in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out.
2. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 1, characterized in that: Format the multi-source basic data sets according to standard processing, unify the field structures of each data source, remove the fields with structural conflicts, and perform deduplication operations based on data timestamps to obtain an integrated data set, including: According to the semantic relevance of the field names in the multi-source basic data sets, the field names of the multi-source basic data sets are normalized and converted to obtain a first data set; Identify structural conflicts of repeated field contents in the first data set, retain field data with high source priority, and remove content conflict fields to obtain a second data set; According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set.
3. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 1, characterized in that: Based on the periodic migration dataset and enterprise location fields, the original location and the target location are mapped to a standard coordinate system to generate a regional migration distribution map, including: According to the enterprise location field, the information of the original location and the target location of the enterprise is extracted from the integrated data set, and the unstructured address data is converted into structured coordinate pairs to form a location coordinate pair set; According to the enterprise identification field in the periodic migration data set, the enterprise migration records in each period are matched with the location coordinate pair set to generate a periodic location mapping data set, wherein the periodic location mapping data set carries both time and space information; According to the periodic location mapping data set, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path data set; A directed weighted graph is constructed based on the geographic path dataset, with nodes representing regional identifiers and edges representing migration directions and numbers, to generate a regional migration distribution map.
4. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 1, characterized in that: Based on the regional migration distribution map, periodic migration dataset and enterprise category field, the enterprise migration behavior is classified and aggregated by enterprise category, and the migration activity of different types of enterprises in the target area within a specific period is evaluated to generate classification trend indicators of enterprise migration in and out, including: According to the enterprise category field, the enterprise migration records in the periodic migration dataset are divided into multiple category migration subsets by category, and combined with the original location and target location recorded in the regional migration distribution map, a set of migration record pairs divided by category and region is generated; According to the set of migration records, the number of migrants in and out of each target area is aggregated and counted for each category of migration subset, and a data table with a three-dimensional structure of category-region-period is constructed. Based on this data table, the migration density function is used to calculate the migration activity of each category in each area in different periods to form a migration activity sequence; Based on the data differences of consecutive periods in the migration activity sequence, the cycle change rate is calculated, and the indicator volatility calculation formula based on standard deviation normalization is used to obtain the classification trend indicators of enterprise migration behavior in various categories and regional dimensions.
5. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 2, characterized in that: According to the semantic relevance of the field names in the multi-source basic data set, the field names of the multi-source basic data set are normalized and converted to obtain a first data set, including: Extract all field names from multi-source basic data sets to form a field name set; Based on the field name set and the preset domain vocabulary comparison table, determine whether there are fields with the same meaning but different names in different data sources, and determine the recommended standard field names to form a recommended set of standard field names; According to the recommended set of standard field names, for field names that are judged to be semantically consistent, the field names in the multi-source basic data set are replaced with the corresponding standard field names by word meaning replacement to generate a standard field set; The standard field set is arranged in the order set by the standard field structure template to generate a first data set.
6. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 5, characterized in that: The structure conflicts of repeated field contents in the first data set are identified, field data with high source priority is retained, and content conflict fields are removed to obtain a second data set, including: Extracting the original source of each field according to the field source information in the first data set, and assigning field source priorities based on a preset credibility level to form field priority data; According to the field priority data, a conflict comparison is performed on the fields in the first data set that have consistent semantics but conflicting sources, the field data with higher priority is retained, and the field values with lower priority or inconsistent content are removed to generate conflict cleaning result data; The conflict clearing result data is taken as output to obtain a second data set.
7. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 6, characterized in that: According to the timestamp information corresponding to each data record in the second data set, combined with the enterprise identification field, nearly duplicate records are identified, and redundant items are deleted to obtain an integrated data set, including: Extracting a timestamp field and an enterprise identification field from the second data set, and combining them into a unique identification key to generate record index data; According to the record index data, based on the preset time difference tolerance interval, determine whether the data record meets the duplication condition, the duplication condition time difference is within the preset time difference tolerance interval and the enterprise identification is consistent; For data records that meet the duplication condition, the best record in the second data set is retained based on data timeliness, the remaining records are deleted, and the integrated data set is output.
8. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 3, characterized in that: According to the periodic location mapping dataset, each set of location coordinate pairs is mapped to a unified standard geographic coordinate system, and the enterprise migration path vector is constructed to form a geographic path dataset, including: According to the coordinate pairs of the original position and the target position contained in each enterprise migration record in the periodic location mapping data set, the original coordinate system is converted into spatial points in a unified standard geographic coordinate system to form a standardized location pair set; Based on the spatial discretization principle of grid coding, each pair of coordinates in the standardized position pair set is encoded and the coordinate value is converted into an indexable regional grid code; According to the enterprise identification field and the time period field, the migration direction of each enterprise in a specific period is represented as a path vector from the starting point code to the end point code, and the migration quantity and enterprise category attribute information are added to each path vector to construct an enterprise migration path dataset with attributes; All enterprise migration path vectors are organized into periods to generate a geographic path dataset with a time hierarchy.
9. A method for managing enterprise distribution information of migration in and out of an industrial platform according to claim 8, characterized in that: A directed weighted graph is constructed based on the geographic path dataset. The nodes represent the regional identifiers, and the edges represent the migration direction and number of migrations. A regional migration distribution map is generated, including: According to the starting grid code and the end grid code carried by each migration path vector in the geographic path dataset, they are respectively assigned to standard administrative regions as node identifiers in the graph to form a regional node set; Merge multiple path vectors that belong to the same time period and have the same starting and ending areas, and calculate the sum of the migration quantities corresponding to the migration path as the edge weight of the path in the graph; The regional node set is used as the node in the graph, and the start and end region pairs after path aggregation are used as the directed edges in the graph. The directed connection relationship between the nodes is constructed, and the sum of the migration quantity is used as the weighted value of the edge to form a weighted directed graph structure. The directed weighted graph is output in the form of a graph as a regional migration distribution graph.
10. The method for managing enterprise distribution information of an industry platform migration in and out according to claim 4, characterized in that: The migration density function is used to calculate the migration activity of each category in each region in different periods to form a migration activity sequence, including: According to the data table of the three-dimensional structure of category-region-period, the number of enterprises of each category moving in and out of each target area in each time period is extracted, and used as the measurement value of the migration behavior of enterprises of this category in this area in this period; For each target area, extract the corresponding spatial area parameter or industrial capacity parameter as the benchmark factor for density normalization; Divide the migration behavior measurement value of each category by the corresponding benchmark factor to obtain the normalized migration activity of each category in each target area and each time period; The normalized migration activity is arranged in order of time period to form a migration activity sequence.
Citation Information
Patent Citations
Method and device for evaluating contact demand compactness among functional areas of city
CN110297875A
Public bicycle station dynamic planning method and system for Df-PBS system
CN111489064A
Urban operation-oriented enterprise migration big data monitoring and early warning method and device
CN115630732A
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Cited By
Enterprise operation data real-time analysis system based on mobile BI
CN120448842A
Company internal control data automatic management method based on intelligent contract
CN120523864A
Method and device for integrating test data in multiple test environments
CN121501582A