A multi-source data dynamic compilation method and system of a cross-border trade and economic regional development index
Patent Information
- Application Number
- CN202610956944.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明的目的在于针对现有技术的不足,提供一种跨境经贸区域发展指数的多源数据动态编制方法及系统,以解决多源数据融合难、权重静态、更新滞后的技术问题
1. 多源异构数据融合:通过适配器架构,有效整合了海关、港口、物流、统计四类异构数据,大幅提升了指数的覆盖维度与准确性;
Smart Images

Figure CN122840410A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing and index compilation technology, specifically to a method and system for dynamic compilation of multi-source data for cross-border economic and trade regional development index. Background Technology
[0002] With the rapid development of cross-border trade and connectivity, the need for quantitative monitoring of regional development trends is becoming increasingly urgent. Traditional methods for compiling regional development indices typically rely on a single source of statistical data and employ expert scoring or static weighting for calculation.
[0003] However, existing technologies have the following drawbacks: 1. The data source is singular and cannot integrate heterogeneous data from multiple dimensions such as customs, ports, and logistics, resulting in narrow index coverage and low accuracy; 2. The weights are set statically and cannot be dynamically adjusted according to changes in the data over time, making it difficult to reflect the dynamic fluctuations of the market; 3. The workflow has a low degree of automation, relies on manual processing, has a long update cycle, and cannot achieve real-time monitoring and early warning; Therefore, there is an urgent need for an index compilation scheme that can integrate multi-source heterogeneous data, dynamically adjust weights, and operate automatically. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for dynamically compiling a cross-border economic and trade regional development index using multi-source data, thereby solving the technical problems of difficulty in multi-source data integration, static weights, and delayed updates.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for dynamically compiling a cross-border economic and trade regional development index using multi-source data, comprising: S1. Through a multi-source interface adapter, connect to the customs clearance system, port operation system, logistics park platform, and statistical database respectively to obtain heterogeneous raw data and convert it into a unified structured format; S2. Perform outlier removal, missing value imputation, cross-source consistency verification, and standardization on the structured data to generate a standardized dataset; S3. Establish a three-level indicator system. Based on the standardized dataset, use a time-adaptive weighting algorithm to calculate the dynamic weight of each indicator. S4. Following a bottom-up approach, aggregate and calculate the indicators at all levels to generate a comprehensive index for cross-border economic and trade regional development; S5. Based on the SARIMA model, perform time-series prediction on the comprehensive index and generate an early warning signal according to a preset threshold; S6. Timed triggering of full-link incremental data collection and recalculation enables automatic rolling updates of the index.
[0006] Furthermore, in S3, the time-adaptive weighting algorithm includes: Calculate the coefficient of variation (CV) and the recent trend strength (T) of each indicator within a preset sliding window; The coefficient of variation and trend strength are normalized. The dynamic weights are obtained by fusion based on the balance coefficient α: w_i=α·CV_norm_i+(1−α)·T_norm_i, where α∈[0.3,0.7].
[0007] Further, in S4, the hybrid polymerization calculation includes: The third-level sub-indicators are linearly weighted and aggregated into second-level indicators, and the second-level indicators are linearly weighted and aggregated into first-level indicators. The primary indicators are aggregated using multiplication to generate a comprehensive index: Index=(I_trade)^w1×(I_port)^w2×(I_logistics)^w3. In the formula, I_trade is the primary indicator of international trade volume, I_port is the primary indicator of port hub capacity, I_logistics is the primary indicator of international logistics channels, w1, w2, and w3 are the weights of the corresponding primary indicators, and w1+w2+w3=1.
[0008] Correspondingly, the present invention also provides a system for implementing the above method, comprising: - Multi-source data acquisition module: responsible for connecting to heterogeneous data sources and completing data acquisition and format conversion; - Data cleaning and fusion module: performs anomaly handling, missing data imputation, and data standardization; - Adaptive weight calculation module: Runs sequential weight algorithm and outputs dynamic weight parameters; - Index aggregation calculation module: Executes a hybrid aggregation model to generate a comprehensive development index; - Prediction and early warning module: Implements time-series prediction and threshold-based early warning; - Scheduling and Update Module: Responsible for scheduling timed tasks and triggering incremental updates. Beneficial effects
[0009] Compared with the prior art, the present invention has the following beneficial effects: 1. Multi-source heterogeneous data fusion: Through the adapter architecture, four types of heterogeneous data, namely customs, ports, logistics and statistics, are effectively integrated, which greatly improves the coverage dimensions and accuracy of the index; 2. Dynamic Adaptive Weights: A time-series adaptive weighting algorithm based on the coefficient of variation and trend strength is proposed, which solves the lag problem of traditional static weights and can reflect the dynamic changes of data in real time; 3. Automated operation: It realizes full-link automation from data collection to index release, supports incremental updates, shortens the update cycle to T+1, and has predictive and early warning capabilities, providing real-time support for decision-making. Attached Figure Description
[0010] Figure 1 is a block diagram of the overall system architecture of the present invention; Figure 2 is a flowchart of the overall index compilation method of the present invention. Detailed Implementation
[0011] The present invention will now be described in detail through specific embodiments.
[0012] 1. Multi-source heterogeneous data acquisition and standardization In this embodiment, the system connects to four types of heterogeneous data sources through a multi-source interface adapter: - Customs clearance system: Retrieves customs declaration data via RESTful API; - Port Operations System: Obtains real-time data such as container throughput via MQTT subscription; - Logistics park platform: Obtains waybill data via WebService; - Statistical database: Obtain macroeconomic statistical data through open APIs; The system converts the raw data in different formats into a structured format through field mapping, thus solving the problem of data heterogeneity across systems.
[0013] 2. Data Cleaning and Fusion The system cleans the collected data: - Outlier handling: The box plot IQR algorithm is used to identify and remove outlier data that deviates too much; - Missing value handling: Multiple imputation is used for data with low missing rate, and the mean of the same period is used to fill data with high missing rate; - Consistency check: Checks the logical consistency between customs export value and port throughput. If the deviation exceeds 20%, it is marked for review. - Standardization: Unify currency, time zone and unit of measurement to generate standardized datasets.
[0014] 3. Adaptive weight calculation After establishing the three-level indicator system, the system calculates the dynamic weights: - Using a 12-month sliding window, calculate the coefficient of variation (CV) for each indicator (σ / μ) to measure data stability; - Calculate the trend strength T over the past 3 periods to measure the timeliness of the data; - Normalize CV and T, w_i=α·CV_norm_i+(1−α)·T_norm_i. In this example, the balance coefficient α=0.5, that is, w_i=0.5*CV_norm_i+0.5*T_norm_i.
[0015] 4. Exponential Aggregation The system adopts a hybrid aggregation model: - The weighted summation of the tertiary indicators yields the secondary indicators; - The weighted summation of the secondary indicators yields the primary indicators (trade, ports, logistics); - The primary indicators are aggregated through multiplication to obtain the comprehensive index: The three primary indicators are international trade volume, port hub capacity, and international logistics channels, with corresponding symbols I_trade, I_port, and I_logistics, respectively. The comprehensive index is calculated as: Index = I_trade^0.4 * I_port^0.35 * I_logistics^0.25, where w1=0.4, w2=0.35, and w3=0.25, satisfying w1+w2+w3=1. In practical applications, the weights are dynamically adjusted in real time according to the algorithm.
[0016] 5. Forecasting and Early Warning The system uses the SARIMA model to train the index sequence, automatically optimizes parameters, and predicts the trend over the next 6 months. When the index declines by more than 3%, 5%, and 8% month-on-month, it triggers yellow, orange, and red alerts, respectively.
[0017] 6. Automatic updates The system uses a scheduling engine to trigger incremental data collection and full-process recalculation on a daily basis, enabling automatic rolling updates of the index.
[0018] The above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made within the scope of the claims of the present invention should be included in the scope of the present invention.
Claims
1. A method for dynamically compiling a cross-border economic and trade regional development index using multi-source data, characterized in that, include: S1. Through a multi-source interface adapter, connect to the customs clearance system, port operation system, logistics park platform, and statistical database respectively to obtain heterogeneous raw data and convert it into a unified structured format; S2. Perform outlier removal, missing value imputation, cross-source consistency verification, and standardization on the structured data to generate a standardized dataset; S3. Establish a three-level indicator system. Based on the standardized dataset, use a time-adaptive weighting algorithm to calculate the dynamic weight of each indicator. S4. Following a bottom-up approach, aggregate and calculate the indicators at all levels to generate a comprehensive index for cross-border economic and trade regional development; S5. Based on the SARIMA model, perform time-series prediction on the comprehensive index and generate an early warning signal according to a preset threshold; S6. Timed triggering of full-link incremental data collection and recalculation enables automatic rolling updates of the index.
2. The method according to claim 1, characterized in that, In S3, the time-adaptive weighting algorithm includes: Calculate the coefficient of variation (CV) and the recent trend strength (T) of each indicator within a preset sliding window; The coefficient of variation and trend strength are normalized. The dynamic weights are obtained by fusion based on the balance coefficient α: w_i=α・CV_norm_i+(1−α)・T_norm_i, where α∈[0.3,0.7], and the sum of the weights of the same level is 1 after normalization.
3. The method according to claim 1, characterized in that, In S4, the hybrid polymerization calculation includes: The third-level sub-indicators are linearly weighted and aggregated into second-level indicators, and the second-level indicators are linearly weighted and aggregated into first-level indicators. The primary indicators are aggregated using multiplication to generate a comprehensive index: Index=(I_trade)^w1×(I_port)^w2×(I_logistics)^w3. In the formula, I_trade is the primary indicator of international trade volume, I_port is the primary indicator of port hub capacity, I_logistics is the primary indicator of international logistics channels, w1, w2, and w3 are the weights of the corresponding primary indicators, and w1+w2+w3=1.
4. A system for implementing the method according to any one of claims 1-3, characterized in that, include: The system includes a multi-source data acquisition module, a data cleaning and fusion module, an adaptive weight calculation module, an index aggregation calculation module, a prediction and early warning module, and a scheduling and update module.