Digital economic statistical system and method

Through the digital economic statistics system, the problems of multi-source heterogeneous data processing and indicator calculation in traditional economic statistics have been solved, efficient data integration and real-time analysis have been achieved, and accurate regional economic portraits and decision-making support have been provided.

CN120687726AInactive Publication Date: 2025-09-23刘珊
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510782568.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional economic statistics methods have difficulty processing multi-source heterogeneous data and are unable to achieve efficient data fusion, resulting in inaccurate calculation of economic indicators, inability to dynamically reflect industry changes, insufficient regional economic analysis, and inability to provide scientific decision-making support.

Method used

A digital economic statistics system is adopted, including a data acquisition module, a data preprocessing module, a dynamic weight allocation module, a real-time economic indicator calculation module and a regional economic portrait generation module. Through timestamp alignment, industry knowledge graph, dynamic weight allocation, GIS technology and LSTM prediction methods, multi-source heterogeneous data processing and real-time economic indicator calculation are realized.

Benefits of technology

It improves the integrity and timeliness of economic statistical data, dynamically adjusts weights to conform to actual economic conditions, generates accurate regional economic portraits, and provides scientific support for economic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687726A_ABST
    Figure CN120687726A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of economic statistics, and particularly discloses a digital economic statistics system and method.The digital economic statistics system is characterized in that a data acquisition module acquires multi-source heterogeneous data, and a preprocessing module processes the data through a timestamp alignment algorithm and an industry knowledge graph; the dynamic weight distribution module integrates multiple factors to determine an industry weight; the real-time economic index calculation module defines indexes such as a digital consumption index and an industrial toughness index and performs real-time calculation; and the regional economy portrait generation module realizes regional economy visualization and prediction based on GIS and LSTM. According to the method, efficient processing of multi-source data and accurate calculation of economic indexes are realized, a regional economic portrait can be effectively generated, and comprehensive and real-time data support is provided for economic analysis and decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of economic statistics, and in particular to a digital economic statistics system and method. Background Art

[0002] With the rapid development of the global economy and the deep integration of digitalization, the complexity and diversity of economic activities have become increasingly prominent, placing unprecedented demands on the accuracy, timeliness, and comprehensiveness of economic statistics. Traditional economic statistics rely primarily on structured data, manually collecting, entering, and analyzing data such as corporate financial statements and registration information provided by tax, industry and commerce departments to calculate economic indicators and assess the economic situation. This approach has significant limitations in data collection, making it difficult to efficiently integrate cross-departmental data. For example, tax payment data from tax authorities, business registration data from industry and commerce departments, and import and export data from customs departments face numerous obstacles in the data integration process due to the varying standards and formats of these data, making it difficult to quickly form a comprehensive and coherent picture of economic data.

[0003] With the booming internet economy, a wealth of economic data, such as semi-structured e-commerce platform transaction logs, logistics company shipping documents, and unstructured news reports and social media comments, has emerged. However, traditional statistical methods are extremely limited in their ability to process this type of data. For example, e-commerce transaction logs, which record information such as transaction time, amount, and user reviews, can intuitively reflect the dynamic trends of the consumer market. Economic discussions and public opinion on social media can, to a certain extent, predict potential market demand. However, due to a lack of effective data parsing and extraction technologies, traditional statistics cannot incorporate this valuable data into economic statistics. This significantly reduces the timeliness and comprehensiveness of the data, making it difficult to accurately capture the real-time dynamics of economic operations.

[0004] When calculating economic indicators, traditional methods often use fixed weights, often based on a single factor such as an industry's historical scale or output value. These methods fail to promptly reflect the dynamic fluctuations of an industry or the impact of policy changes on the economy. For example, during the rapid rise of emerging industries, traditional fixed-weight statistical methods fail to promptly highlight their crucial role in economic development, making it difficult to accurately assess an industry's resilience to risk and the degree of digitalization in its consumption structure. When the state introduces industrial support policies, traditional statistics are unable to quickly incorporate the impact of policy factors on related industries into weight calculations, leading to deviations between economic statistics and actual economic development.

[0005] In the field of regional economic analysis, traditional methods rely primarily on static data reports and simple charts. These methods fail to intuitively present the geographical distribution of industries, making it difficult to effectively predict regional economic trends or accurately identify key industries and areas of weakness in regional economic development. For example, when analyzing a region's industrial layout, traditional methods fail to clearly demonstrate the geographical agglomeration effects and synergistic development relationships among various industries, hindering local governments from formulating scientifically sound industrial planning and economic development strategies.

[0006] To sum up, in the digital age, traditional economic statistical methods can no longer meet the needs of accurate economic analysis and scientific decision-making. There is an urgent need for a digital economic statistical system and method that can efficiently process multi-source heterogeneous data, dynamically calculate economic indicators, and comprehensively characterize the regional economy. Summary of the Invention

[0007] The purpose of this invention is to provide a digital economic statistics system and method to solve the problems existing in traditional economic statistics in data processing, indicator calculation and regional economic analysis, realize the efficient processing of multi-source heterogeneous economic data, dynamic and accurate calculation of economic indicators and comprehensive analysis of regional economy, and provide scientific and accurate data support for economic decision-making.

[0008] To achieve the above-mentioned purpose, the technical solution provided by the present invention is as follows: a digital economic statistics system, comprising a data acquisition module, a data preprocessing module, a dynamic weight allocation module, a real-time economic indicator calculation module, and a regional economic portrait generation module;

[0009] The data collection module is used to obtain structured data such as corporate financial statements and import and export data from tax, industry and commerce, and customs departments; parse e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; and extract economic keywords from news reports and social media through NLP to obtain unstructured data.

[0010] The data preprocessing module uses a timestamp alignment algorithm to unify the time granularity of different data sources; based on the industry knowledge graph, it corrects the semantic ambiguity in the data;

[0011] The dynamic weight allocation module assigns initial weights to various industries based on the national economic industry classification; dynamically adjusts weights by calculating the covariance between industry stock index volatility and GDP growth; extracts keywords based on national industrial policy documents and increases policy weights for relevant industries; and uses a combination of entropy weighting and hierarchical analysis to derive final weights, ensuring the total weight is 100% and avoiding dominance by a single factor.

[0012] The real-time economic indicator calculation module defines a digital consumption index indicator: based on the mobile payment transaction volume, the number of online education users and the advertising revenue of the short video platform, the degree of digitalization of the consumption structure is calculated;

[0013] Define the industry resilience index indicators: Combine the company's debt-to-asset ratio, R&D investment intensity, and supply chain diversification to assess the industry's risk resistance;

[0014] The pre-processed data is classified by industry; a dynamic weight allocation model is applied to calculate the weighted value of each industry; and economic indicators are generated through real-time aggregation using a distributed computing framework;

[0015] The regional economic portrait generation module is based on GIS technology, mapping economic indicators to geographic grids to generate industrial distribution heat maps; using LSTM to predict regional economic trends in the next three months; calculating the deviation of regional economic indicators from the national average, and identifying advantageous industries and weak areas.

[0016] Furthermore, the specific method of using the timestamp alignment algorithm to unify the time granularity of different data sources is as follows:

[0017] For the n data sources obtained, their time granularities are: T1, T2, ..., T n , select the smallest time granularity T min =min(T1,T2,…,T n ) as a unified benchmark granularity; for the i-th data source, define the time conversion function f i (t), convert its original timestamp t into the timestamp t′ at the base time granularity:

[0018] For the timestamp t of the same event in different data sources i1 and t j1 , calculate the time deviation Δt=|f i (t i1 )-f j (t j1 )|; If Δt>∈, where ∈ is the set time deviation threshold; then according to the reliability of the data and the continuity of the time series, the linear interpolation method is used for adjustment, as follows: for two adjacent timestamps t k and t k+1 and its corresponding data value x k and x k+1 , the linear interpolation is:

[0019]

[0020] Furthermore, based on the industry knowledge graph, the specific method for correcting semantic ambiguity in data is as follows:

[0021] Construct an industry knowledge graph G = (V, E) consisting of nodes and edges, where nodes V represent industry terms and entities, and edges E represent semantic relationships between nodes. Each edge is assigned a weight w. ij , representing node v i and v j the closeness of the semantic relationship between them;

[0022] When an ambiguous term a appears in the data, the node set N related to a is searched in the knowledge graph. a , calculate each relevant node n∈N a Semantic relevance S(n) to other term nodes in the data context: S(n) = ∑ m∈C w nm ; Where C is the set of term nodes involved in the data context; select the semantic interpretation corresponding to the node with the largest semantic association S(n) as the correct semantics of the term in the current data to complete the correction of the data.

[0023] Furthermore, the initial weight allocation method is as follows:

[0024] According to the national economic industry classification, the economic industry is divided into m categories; an initial weight w is assigned to each industry i i0 When considering the economic scale and employment contribution of the industry, we can make a comprehensive consideration, where i = 1, 2, ..., m; the economic scale of industry i is S i , the number of employed people is E i The total economic scale of all industries is The total number of employed people is The initial weight calculation formula is:

[0025] Wherein, α and β are adjustment coefficients, and α+β=1.

[0026] Furthermore, the method of dynamically adjusting weights is as follows: the stock index of industry i in time period t is I it , its volatility σ it It is obtained by calculating the standard deviation of the stock index during the time period, that is:

[0027] Where n is the number of samples in the time period t, is the mean of the stock index of industry i during the time period;

[0028] At the same time, the GDP growth rate during this time period is g t , then the covariance Cov between the stock index volatility of industry i and GDP growth rate is i for:

[0029] in, is the mean volatility of the stock index of industry i during the time period, is the average GDP growth rate during this time period;

[0030] According to the covariance Cov i For the initial weight w i0 Perform dynamic adjustments to obtain the adjusted weight w i1 :

[0031] Among them, γ is the adjustment coefficient, which is used to control the influence of covariance on weight adjustment; max(|Cov1|,|Cov2|,…,|Cov m |) represents the maximum absolute value of the covariance of all industries, which is used to normalize the adjustment amplitude.

[0032] Furthermore, the specific method of increasing policy weight is as follows: analyze national industrial policy documents through natural language processing, extract the keyword set K, and for each industry i, calculate its correlation r with the keyword set K i ; The number of keywords contained in the policy document is N, and the number of keywords related to industry i is n i , then the correlation calculation formula is:

[0033] According to the correlation r i Add policy weight w to industry i i2 :w i2 =δ×r i ; Among them, δ is the policy weight coefficient, which is used to adjust the impact of policy factors on the weight.

[0034] Furthermore, the specific calculation method of the final weight is as follows:

[0035] Use the entropy weight method to calculate the objective weight of each weight o i ; Let x i1 =w i0 ,x i2 =w i1 ,x i3 =w i2 , then the entropy value H j for:

[0036] in,

[0037] Then we get the entropy weight e j :

[0038] Use the analytic hierarchy process to determine the subjective weight of each weighti ;

[0039] Comprehensive objective weight i and subjective weight s i , get the final weight w i :w i =θ×o i +(1-θ)×s i ; where θ is the balance coefficient, 0≤θ≤1, and

[0040] The present invention also provides a digital economic statistics method, which is implemented based on the above system and includes the following steps:

[0041] S1. Multi-source heterogeneous data collection and preprocessing: Obtain structured data on corporate financial statements and imports and exports from tax, industry and commerce, and customs departments; Analyze e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; Use NLP to extract economic keywords from news reports and social media to obtain unstructured data; Use timestamp alignment algorithms to align the time granularity of different data sources; Correct semantic ambiguity in the data based on industry knowledge graphs;

[0042] S2. Dynamic Weight Allocation Model: Assign initial weights to each industry based on the national economic industry classification; dynamically adjust weights by calculating the covariance between industry stock index volatility and GDP growth; extract keywords based on national industrial policy documents and increase policy weights for related industries; and use a combination of entropy weighting and the analytic hierarchy process to derive final weights, ensuring the total weight is 100% and avoiding dominance by a single factor.

[0043] S3. Real-time economic indicator calculation engine: Define the digital consumption index indicator: Calculate the degree of digitalization of the consumption structure based on mobile payment transaction volume, number of online education users, and advertising revenue of short video platforms; Define the industry resilience index indicator: Combine corporate asset-liability ratio, R&D investment intensity, and supply chain diversification to assess the industry's risk resistance; Classify pre-processed data by industry; Apply a dynamic weight allocation model to calculate the weighted value of each industry; Generate economic indicators through real-time aggregation through a distributed computing framework;

[0044] S4. Generation of regional economic profiles: Map economic indicators to geographic grids based on GIS to generate a heat map of industrial distribution; use LSTM to predict regional economic trends over the next three months; calculate the deviation between regional economic indicators and the national average to identify advantageous industries and areas of weakness.

[0045] The advantages of the present invention compared with the prior art are: the present invention realizes the comprehensive collection of structured, semi-structured and unstructured data, and improves the integrity and richness of economic statistical data.

[0046] The present invention comprehensively considers economic scale, employment contribution, market fluctuations and policy factors, and dynamically adjusts industry weights to make the statistical results more in line with actual economic conditions.

[0047] The present invention generates economic indicators through real-time aggregation via a distributed computing framework, thereby improving the timeliness and real-time performance of statistics.

[0048] Based on GIS technology and multiple algorithms, this invention generates industrial distribution heat maps, predicts economic trends, identifies areas of strength and weakness, and provides accurate decision-making support for regional economic development.

[0049] The present invention adopts a timestamp alignment algorithm and a semantic ambiguity correction method to improve the consistency and accuracy of data, and the anomaly detection mechanism ensures the reliability of economic indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a system block diagram of a digital economic statistics system of the present invention.

[0051] Figure 2 It is a flow chart of a digital economic statistics method of the present invention.

[0052] Figure 3 It is a flowchart of dynamic weight allocation. DETAILED DESCRIPTION

[0053] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention.

[0054] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0055] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0056] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0057] The following is a further detailed description of a digital economic statistics system and method of the present invention with reference to the accompanying drawings.

[0058] Combined with attachment Figure 1-3 , the present invention is introduced in detail.

[0059] A digital economic statistics system, including a data acquisition module, a data preprocessing module, a dynamic weight distribution module, a real-time economic indicator calculation module, and a regional economic profile generation module;

[0060] The data collection module is used to obtain structured data such as corporate financial statements and import and export data from tax, industry and commerce, and customs departments; parse e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; and extract economic keywords from news reports and social media through NLP to obtain unstructured data.

[0061] The data preprocessing module uses a timestamp alignment algorithm to unify the time granularity of different data sources. The specific method is as follows:

[0062] For the n data sources obtained, their time granularities are: T1, T2, ..., T n , select the smallest time granularity T min =min(T1,T2,…,T n ) as a unified benchmark granularity; for the i-th data source, define the time conversion function f i (t), convert its original timestamp t into the timestamp t′ at the base time granularity:

[0063] For the timestamp t of the same event in different data sources i1 and t j1 , calculate the time deviation Δt=|f i (t i1 )-f j (t j1 )|; If Δt>∈, where ∈ is the set time deviation threshold; then according to the reliability of the data and the continuity of the time series, the linear interpolation method is used for adjustment, as follows: for two adjacent timestamps t k and t k+1 and its corresponding data value x k and x k+1 , the linear interpolation is:

[0064]

[0065] Based on the industry knowledge graph, the semantic ambiguity in the data is corrected. The specific method is as follows:

[0066] Construct an industry knowledge graph G = (V, E) consisting of nodes and edges, where nodes V represent industry terms and entities, and edges E represent semantic relationships between nodes. Each edge is assigned a weight w. ij , representing node v i and v j the closeness of the semantic relationship between them;

[0067] When an ambiguous term a appears in the data, the node set N related to a is searched in the knowledge graph. a , calculate each relevant node n∈N a Semantic relevance S(n) to other term nodes in the data context: S(n) = ∑ m∈C w nm ; Where C is the set of term nodes involved in the data context; select the semantic interpretation corresponding to the node with the largest semantic association S(n) as the correct semantics of the term in the current data to complete the correction of the data.

[0068] The dynamic weight allocation module allocates initial weights to each industry based on the national economic industry classification. The specific allocation method is as follows:

[0069] According to the national economic industry classification, the economic industry is divided into m categories; an initial weight w is assigned to each industry i i0 When considering the economic scale and employment contribution of the industry, we can make a comprehensive consideration, where i = 1, 2, ..., m; the economic scale of industry i is S i , the number of employed people is E i The total economic scale of all industries is The total number of employed people is The initial weight calculation formula is:

[0070] Wherein, α and β are adjustment coefficients, and α+β=1.

[0071] By calculating the covariance between the volatility of industry stock index and GDP growth rate, the weight is adjusted dynamically. The specific method is as follows: the stock index of industry i in time period t is I it , its volatility σ it It is obtained by calculating the standard deviation of the stock index during the time period, that is:

[0072] Where n is the number of samples in the time period t, is the mean of the stock index of industry i during the time period;

[0073] At the same time, the GDP growth rate during this time period is g t , then the covariance Cov between the stock index volatility of industry i and GDP growth rate is i for:

[0074] in, is the mean volatility of the stock index of industry i during the time period, is the average GDP growth rate during this time period;

[0075] According to the covariance Cov i For the initial weight w i0 Perform dynamic adjustments to obtain the adjusted weight w i1 :

[0076] Among them, γ is the adjustment coefficient, which is used to control the influence of covariance on weight adjustment; max(|Cov1|,|Cov2|,…,|Cov m |) represents the maximum absolute value of the covariance of all industries, which is used to normalize the adjustment amplitude.

[0077] According to the national industrial policy documents, keywords are extracted to increase the policy weight for related industries. The specific method is as follows: Analyze the national industrial policy documents through natural language processing, extract the keyword set K, and for each industry i, calculate its correlation r with the keyword set K i ; The number of keywords contained in the policy document is N, and the number of keywords related to industry i is n i , then the correlation calculation formula is:

[0078] According to the correlation r i Add policy weight w to industry i i2 :w i2 =δ×r i ; Among them, δ is the policy weight coefficient, which is used to adjust the impact of policy factors on the weight.

[0079] The final weights are obtained by combining the entropy weight method with the hierarchical analysis method to ensure that the total weight is 100% and to avoid the dominance of a single factor. The specific calculation method of the final weight is as follows:

[0080] Use the entropy weight method to calculate the objective weight of each weight o i ; Let x i1 =w i0 ,x i2 =w i1 ,x i3 =w i2 , then the entropy value H j for:

[0081] in,

[0082] Then we get the entropy weight e j :

[0083] Use the analytic hierarchy process to determine the subjective weight of each weight i ;

[0084] Comprehensive objective weight iand subjective weight s i , get the final weight w i :w i =θ×o i +(1-θ)×s i ; where θ is the balance coefficient, 0≤θ≤1, and

[0085] The real-time economic indicator calculation module defines a digital consumption index indicator: based on the mobile payment transaction volume, the number of online education users, and the advertising revenue of short video platforms, the degree of digitalization of the consumption structure is calculated. The calculation method of the digital consumption index indicator (DCI) is as follows:

[0086] Based on mobile payment transaction amount T mp , the number of online education users U oe and short video platform advertising revenue R sp , introducing the consumption growth trend factor λ t and consumption structure change coefficient μ i , construct the calculation formula of digital consumption index indicator:

[0087]

[0088] in, are the benchmark values ​​of the corresponding indicators; ω1+ω2+ω3=1 is the weight coefficient of each indicator; λ t1 ,λ t2 ,λ t3 It is a consumption growth trend factor based on time series, which is calculated by predicting future consumption growth trends through the ARIMA model; μ1, μ2, and μ3 are consumption structure change coefficients, which are dynamically adjusted according to changes in market share of consumer categories, reflecting changes in the degree of digitalization of consumption structure in different dimensions.

[0089] Defining the Industry Resilience Index: The industry's risk resilience is assessed by combining corporate debt-to-asset ratios, R&D investment intensity, and supply chain diversification. The specific method for calculating the Industry Resilience Index (PRI) is as follows:

[0090] Combined with the enterprise's debt-to-asset ratio L ar , R&D investment intensity I ri and the degree of supply chain diversification D sc , introducing risk buffer factors and innovation potential factor ψ i , construct the calculation formula of the industry resilience index:

[0091]

[0092] in, are the industry average thresholds of the corresponding indicators respectively; σ1+σ2+σ3=1 is the weight coefficient of each indicator; The risk buffer factor is derived based on factors such as the company's cash flow reserves and debt maturity structure, reflecting the company's ability to withstand financial risks; i1 , ψ i2 It is an innovation potential factor, which is dynamically adjusted based on indicators such as the number of patent applications and the conversion rate of technological achievements to reflect the industry's innovation and development potential and the supply chain's resilience to risks.

[0093] The pre-processed data is classified by industry. In addition to the National Economic Industry Classification (GB / T4754), the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is also introduced to analyze the text descriptions in the data to mine the potential industry-related features of the data. Suppose the data set contains n data samples. For the jth keyword w in the i-th sample, ij , its TF-IDF value is calculated as follows: TF-IDF ij =TF ij ×IDF j Among them, word frequency Indicates keyword w ij Frequency of occurrence in the i-th sample; inverse document frequency n is the total number of samples j To contain the keyword w ij Based on the matching degree between the TF-IDF value and the industry characteristic vocabulary, the industry classification of the data sample is rechecked and adjusted to improve the classification accuracy.

[0094] Apply the dynamic weight distribution model to calculate the weighted value of each industry and introduce the data credibility factor τ i This factor is derived from a comprehensive assessment of factors such as the authority of the data source, data integrity, and data update frequency. i The calculation formula is: V i =w i ×∑ k x ik ×τ i ; Among them, w i is the final weight of the industry, x ik is the kth economic data indicator value of the i-th industry.

[0095] Generate economic indicators in real time through real-time aggregation using a distributed computing framework; Spark Streaming is used as the distributed computing framework, and the sliding window time is set to Δt and the window sliding step is δ t In each window, the weighted values ​​of each industry are aggregated in real time to calculate the regional economic indicators: E z =∑i V i ;

[0096] At the same time, an anomaly detection mechanism is introduced based on the mean of historical data and standard deviation σ z , calculate the Z-score value of the real-time economic indicator:

[0097]

[0098] When the absolute value of the Z-score is greater than the set threshold ε, an abnormal alarm is triggered, prompting you to check and correct the data or calculation process to ensure the accuracy and reliability of economic indicators.

[0099] The regional economic portrait generation module is based on GIS technology, which maps economic indicators to geographic grids and generates industrial distribution heat maps; for each geographic grid g ij , its economic indicator value is E ij , convert it into thermal value H ij The calculation formula is as follows:

[0100]

[0101] Among them, N ij is the geographic grid g ij The neighborhood grid set, ω kl is the neighborhood grid g kl With the current grid g ij The economic correlation weight is obtained by calculating the similarity of the main industries in the two grids; ρ is the thermal diffusion coefficient, which is used to adjust the impact of industrial agglomeration on the surrounding areas, and its value range is [0,1]. Finally, according to the calculated thermal value H ij Generate an industrial distribution heat map so that the heat map can more intuitively and accurately show the agglomeration and diffusion of industries in geographical space.

[0102] Use LSTM to predict regional economic trends in the next three months; calculate the deviation between regional economic indicators and the national average to identify advantageous industries and weak areas; when calculating the deviation between regional economic indicators and the national average, combine the industry dynamic weight w i Perform weighted calculations to more accurately identify advantageous industries and weak areas. Suppose the economic indicator value of the i-th industry in the region is E zi , the national average value of industry i is The calculation formula for the weighted deviation D of regional economic indicators is: Industries are ranked based on the calculated deviation D. When D > 0 and the deviation is large, the industry is considered a dominant industry; when D < 0 and the deviation is large in absolute value, the industry is considered a weak sector, thus providing a more targeted reference for the formulation of regional economic development strategies.

[0103] The present invention also provides a digital economic statistics method, which is implemented based on the above system and includes the following steps:

[0104] S1. Multi-source heterogeneous data collection and preprocessing: Obtain structured data on corporate financial statements and imports and exports from tax, industry and commerce, and customs departments; Analyze e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; Use NLP to extract economic keywords from news reports and social media to obtain unstructured data; Use timestamp alignment algorithms to align the time granularity of different data sources; Correct semantic ambiguity in the data based on industry knowledge graphs;

[0105] S2. Dynamic Weight Allocation Model: Assign initial weights to each industry based on the national economic industry classification; dynamically adjust weights by calculating the covariance between industry stock index volatility and GDP growth; extract keywords based on national industrial policy documents and increase policy weights for related industries; and use a combination of entropy weighting and the analytic hierarchy process to derive final weights, ensuring the total weight is 100% and avoiding dominance by a single factor.

[0106] S3. Real-time economic indicator calculation engine: Define the digital consumption index indicator: Calculate the degree of digitalization of the consumption structure based on mobile payment transaction volume, number of online education users, and advertising revenue of short video platforms; Define the industry resilience index indicator: Combine corporate asset-liability ratio, R&D investment intensity, and supply chain diversification to assess the industry's risk resistance; Classify pre-processed data by industry; Apply a dynamic weight allocation model to calculate the weighted value of each industry; Generate economic indicators through real-time aggregation through a distributed computing framework;

[0107] S4. Generation of regional economic profiles: Map economic indicators to geographic grids based on GIS to generate a heat map of industrial distribution; use LSTM to predict regional economic trends over the next three months; calculate the deviation between regional economic indicators and the national average to identify advantageous industries and areas of weakness.

[0108] The specific implementation process of a digital economic statistics system and method of the present invention is as follows:

[0109] 1. Multi-source heterogeneous data collection and preprocessing

[0110] 1. Data Collection

[0111] The data collection module obtains structured monthly financial statement data from the province's manufacturing and service industries, covering revenue, profit, and tax payments, from the tax authorities; structured data on business registrations and changes from the industrial and commercial authorities; and structured monthly import and export data from the customs authorities. Simultaneously, it analyzes semi-structured data from daily transaction logs of major local e-commerce platforms over the past six months, including product categories, amounts, and transaction times; and obtains semi-structured data on shipping documents from logistics companies. Furthermore, using natural language processing (NLP) technology, it extracts unstructured data on economic keywords, such as industry trends and policy impacts, from local news media reports and discussions on related topics on social media platforms over the past three months.

[0112] 2. Data preprocessing

[0113] Timestamp alignment: In the acquired data sources, the time granularity of tax data is monthly (T1), the time granularity of e-commerce transaction logs is daily (T2), and the time granularity of news report data is hourly (T3). Determine the minimum time granularity T min =T3 (hours) as the unified benchmark granularity. For tax data (the i-th data source, i=1), define a time conversion function f1(t) to convert monthly data into the timestamp t′ corresponding to each hour of the month. For example, a company's April tax data is converted into the timestamps corresponding to 0:00 on April 1, 1:00 on April 30, and so on.

[0114] If a company’s transaction timestamp in the e-commerce transaction log is t i1 (10:00 on April 15) and the timestamp of the economic events related to the enterprise mentioned in the news report j1 (12:00 on April 15), the time deviation Δt is calculated as |f i (t i1 )-f j (t j1 )|=2 hours, set the time deviation threshold∈=1 hour, Δt>∈. The enterprise k )'s transaction amount x k =1000 yuan, 11 o'clock (t k+1 )'s transaction amount x k+1 = 1500 yuan, and the transaction amount at 10 o'clock (t) is calculated by linear interpolation as x = 1000 + (10-9) / (11-9) × (1500-1000) = 1250 yuan, completing the timestamp alignment adjustment.

[0115] Semantic ambiguity correction: Construct an industry knowledge graph G = (V, E) that includes industry terms and corporate entities such as manufacturing and service industries. When the ambiguous term "cloud computing" appears in the data, search for the relevant node set N in the knowledge graph. a, including nodes such as "cloud computing technology" and "cloud computing services." The set of nodes C involving terms such as "big data" and "server" in the data context calculates the semantic relevance S(n) between each related node and the nodes in C. For example, the sum of the edge weights of the "cloud computing technology" node and the "big data" and "server" nodes yields S(cloud computing technology). Similarly, the sum of the edge weights of the "cloud computing service" node yields S(cloud computing service). The semantic interpretation corresponding to the node with the greatest semantic relevance is selected. If S(cloud computing service) > S(cloud computing technology), the semantic meaning of "cloud computing" in the current data is determined to be "cloud computing service," completing the semantic ambiguity correction.

[0116] 2. Dynamic Weight Allocation Model

[0117] 1. Initial weight distribution

[0118] According to the national economic industry classification, the province's economic industries are divided into m = 5 categories, namely manufacturing, service, finance, agriculture, and construction. Taking the manufacturing industry (industry i = 1) as an example, its economic scale S1 = 200 billion yuan and the number of employees E1 = 500,000 people; the total economic scale of all industries is billion yuan, total number of employed people 10,000 people, adjustment coefficient α=0.6, β=0.4. Then the initial weight of manufacturing industry is w i0 =0.6×2000 / 5000+0.4×50 / 150≈0.24+0.13=0.37.

[0119] 2. Dynamically adjust weights

[0120] In the past year (time period t), the manufacturing stock index I it , the sample size n = 252 trading days, calculate its volatility Get σ it =0.15; GDP growth rate of the province during the same period g t , the covariance between the volatility of the manufacturing stock index and the GDP growth rate is calculated The maximum absolute value of the covariance of all industries max(|Cov1|,|Cov2|,…,|Cov m |)=0.12, adjustment coefficient γ=0.5. Then the adjusted manufacturing weight w i1 =0.37×(1+0.5×0.08 / 0.12)≈0.37×(1+0.33)=0.49.

[0121] 3. Increase policy weight

[0122] The country has issued a policy document to support the development of the manufacturing industry. Through natural language processing, a keyword set K of 20 keywords was extracted, of which the number of keywords related to the manufacturing industry was n. i= 8. The number of policy document keywords N = 20, then the correlation between manufacturing industry and keyword set K is r i =8 / 20=0.4, policy weight coefficient δ=0.2, increasing the policy weight w for the manufacturing industry i2 =0.2×0.4=0.08.

[0123] 4. Final weight calculation

[0124] Use entropy weight method to calculate objective weight o i , let x i1 =w i0 =0.37,x i2 =w i1 =0.49,x i3 =w i2 =0.08, calculate the entropy value H j and entropy weight e j . Using the analytic hierarchy process to determine the subjective weights i , the balance coefficient θ=0.6, and the final weight of the manufacturing industry w is obtained i =0.6×o i +0.4×s i , ensuring that the final sum of weights of each industry is 1.

[0125] 3. Real-time Economic Indicator Calculation Engine

[0126] 1. Indicator definition and calculation

[0127] Digital Consumption Index (DCI): The province's mobile payment transaction volume T mp =10 billion yuan, base value 100 million yuan; number of online education users U oe =300,000 people, baseline value 10,000 people; advertising revenue of short video platform R sp =2 billion yuan, base value 100 million yuan. Weight coefficients ω1=0.4, ω2=0.3, ω3=0.3. Using the ARIMA model to predict the consumption growth trend factor λ t1 =1.1,λ t2 =1.05,λ t3 =1.1; Based on the changes in the market share of consumer categories, the consumption structure change coefficients μ1=1.2, μ2=1.1, μ3=1.05. Then:

[0128] DCI=0.4×100 / 80×1.1×1.2+0.3×30 / 25×1.05×1.1+0.3×20 / 15×1.1×

[0129] 1.05≈0.792+0.416+0.462=1.67

[0130] Industry Resilience Index (PRI): The debt-to-asset ratio of a manufacturing enterprise L ar =0.6, industry average threshold R&D investment intensity I ri =0.05, industry average threshold Supply chain diversification degree D sc =0.8, industry average threshold Weight coefficients σ1 = 0.3, σ2 = 0.4, σ3 = 0.3. Evaluate risk buffer factors based on the company's cash flow reserves, etc. Determine the innovation potential factor ψ based on the number of patent applications, etc. i1 =1.2,ψ i2 =1.1; then:

[0131] PRI=0.3×(1-0.6 / 0.7)×1.1+0.4×0.05 / 0.04×1.2+0.3×0.8 / 0.7×1.1≈

[0132] 0.051+0.6+0.394=1.045

[0133] 2. Data Industry Classification

[0134] For the pre-processed data, in addition to the national economic industry classification, the TF-IDF algorithm is introduced. In a data sample, the keyword "intelligent machine tool" has a word frequency of TF ij =n ij / ∑ k n ik , assuming that the number of times “intelligent machine tool” appears in the sample is n ij =3, the total number of occurrences of all keywords in the sample∑ k n ik =20, then TF ij =3 / 20=0.15; the total number of samples n=100, including the number of samples of “intelligent machine tools” n j =10, inverse document frequency IDF j =ln(100 / (1+10))≈2.3, TF-IDF ij = 0.15 × 2.3 = 0.345. Based on the matching degree with the manufacturing characteristic vocabulary, the sample is classified as manufacturing and a second verification and adjustment is performed.

[0135] 3. Calculation of industry weighted value

[0136] Introducing data credibility factor τ i For manufacturing data, τ is evaluated based on the authority of the tax department, data integrity, and update frequency. i =0.9. Final weight of manufacturing industry w i=0.4, the kth economic data indicator value of the industry x ik (such as revenue index value), calculate the weighted value V of each industry i =0.4×∑ k x ik ×0.9.

[0137] 4. Real-time aggregation and generation of economic indicators

[0138] Spark Streaming is used as the distributed computing framework, and the sliding window time Δt = 1 hour and the window sliding step δ t = 15 minutes. In each window, the weighted values ​​of each industry are aggregated in real time to calculate the regional economic index E z =∑ i V i Calculate the average of regional economic indicators based on historical data Standard deviation σ z =50, a real-time economic indicator E is calculated z =650, calculate Z-score = (650-500) / 50 = 3, set threshold ∈ = 2, |Z-score| > ∈, trigger an abnormal alarm, and prompt to check the data or calculation process.

[0139] 4. Generation of Regional Economic Profiles

[0140] 1. Drawing of industrial distribution heat map

[0141] A geographic grid g ij Economic indicator value E ij =200, its neighborhood grid N ij There are 3 grids, economic correlation weight ω kl are 0.7, 0.6, and 0.5 respectively, and the thermal diffusion coefficient ρ = 0.6. Then the thermal value is:

[0142]

[0143] According to the calculated thermal value H ij Generate an industrial distribution heat map to visually display the industrial agglomeration and diffusion in the region.

[0144] 2. Regional economic trend forecast

[0145] The LSTM model was used to train the province’s economic data for the past five years to predict regional economic trends in the next three months, providing a forward-looking reference for economic decision-making.

[0146] 3. Deviation calculation

[0147] The province's manufacturing economic index value E zi =220 billion yuan, the national manufacturing average billion yuan, the final weight of manufacturing industry is w i =0.4, calculate the weighted deviation of regional economic indicators:

[0148] If D>0 and the deviation is large, manufacturing is determined to be the province's advantageous industry; similarly, the deviation of other industries is calculated to identify weak areas and provide a basis for the formulation of regional economic development strategies.

[0149] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.

Claims

1. A digital economic statistics system, characterized by: It includes data acquisition module, data preprocessing module, dynamic weight allocation module, real-time economic indicator calculation module and regional economic portrait generation module; The data collection module is used to obtain structured data such as corporate financial statements and import and export data from tax, industry and commerce, and customs departments; parse e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; and extract economic keywords from news reports and social media through NLP to obtain unstructured data. The data preprocessing module uses a timestamp alignment algorithm to unify the time granularity of different data sources; based on the industry knowledge graph, it corrects the semantic ambiguity in the data; The dynamic weight allocation module allocates initial weights to various industries according to the national economic industry classification; By calculating the covariance between the volatility of industry stock indices and GDP growth rates, we dynamically adjust weights. Based on national industrial policy documents, we extract keywords and increase policy weights for related industries. The entropy weight method is combined with the hierarchical analysis method to obtain the final weight, ensuring that the total weight is 100% and avoiding the dominance of a single factor; The real-time economic indicator calculation module defines a digital consumption index indicator: based on the mobile payment transaction volume, the number of online education users and the advertising revenue of the short video platform, the degree of digitalization of the consumption structure is calculated; Define the industry resilience index indicators: Combine the company's debt-to-asset ratio, R&D investment intensity, and supply chain diversification to assess the industry's risk resistance; The pre-processed data is classified by industry; a dynamic weight allocation model is applied to calculate the weighted value of each industry; and economic indicators are generated through real-time aggregation using a distributed computing framework; The regional economic portrait generation module is based on GIS technology, mapping economic indicators to geographic grids to generate industrial distribution heat maps; using LSTM to predict regional economic trends in the next three months; calculating the deviation of regional economic indicators from the national average, and identifying advantageous industries and weak areas.

2. A digital economic statistics system according to claim 1, characterized in that: The specific method of using the timestamp alignment algorithm to unify the time granularity of different data sources is as follows: For the n data sources obtained, their time granularities are: T1, T2, ..., T n , select the smallest time granularity T min =min(T1,T2,…,T n ) as a unified benchmark granularity; for the i-th data source, define the time conversion function f i (t), convert its original timestamp t into the timestamp t′ at the base time granularity: For the timestamp t of the same event in different data sources i1 and t j1 , calculate the time deviation Δt=|f i (t i1 )-f j (t j1 )|; If Δt>∈, where ∈ is the set time deviation threshold; then according to the reliability of the data and the continuity of the time series, the linear interpolation method is used for adjustment, as follows: for two adjacent timestamps t k and t k+1 and its corresponding data value x k and x k+1 , the linear interpolation is:

3. A digital economic statistics system according to claim 2, characterized in that: Based on the industry knowledge graph, the specific method for correcting semantic ambiguity in data is as follows: Construct an industry knowledge graph G = (V, E) consisting of nodes and edges, where nodes V represent industry terms and entities, and edges E represent semantic relationships between nodes. Each edge is assigned a weight w. ij , representing node v i and v j the closeness of the semantic relationship between them; When an ambiguous term a appears in the data, the node set N related to a is searched in the knowledge graph. a , calculate each relevant node n∈N a Semantic relevance S(n) to other term nodes in the data context: S(n) = ∑ m∈C w nm ; Where C is the set of term nodes involved in the data context; select the semantic interpretation corresponding to the node with the largest semantic association S(n) as the correct semantics of the term in the current data to complete the correction of the data.

4. A digital economic statistics system according to claim 3, characterized in that: The initial weight distribution method is as follows: According to the national economic industry classification, the economic industry is divided into m categories; an initial weight w is assigned to each industry i i0 When considering the economic scale and employment contribution of the industry, we can make a comprehensive consideration, where i = 1, 2, ..., m; the economic scale of industry i is S i , the number of employed people is E i The total economic scale of all industries is The total number of employed people is The initial weight calculation formula is: Wherein, α and β are adjustment coefficients, and α+β=1.

5. A digital economic statistics system according to claim 4, characterized in that: The method of dynamically adjusting weights is as follows: the stock index of industry i in time period t is I it , its volatility σ it It is obtained by calculating the standard deviation of the stock index during the time period, that is: Where n is the number of samples in the time period t, is the mean of the stock index of industry i during the time period; At the same time, the GDP growth rate during this time period is g t , then the covariance Cov between the stock index volatility of industry i and GDP growth rate is i for: in, is the mean volatility of the stock index of industry i during the time period, is the average GDP growth rate during this time period; According to the covariance Cov i For the initial weight w i0 Perform dynamic adjustments to obtain the adjusted weight w i1 : Among them, γ is the adjustment coefficient, which is used to control the influence of covariance on weight adjustment; max(|Cov1|,|Cov2|,…,|Cov m |) represents the maximum absolute value of the covariance of all industries, which is used to normalize the adjustment amplitude.

6. A digital economic statistics system according to claim 5, characterized in that: The specific method of increasing policy weight is as follows: Analyze national industrial policy documents through natural language processing, extract keyword set K, and for each industry i, calculate its correlation r with keyword set K. i ; The number of keywords contained in the policy document is N, and the number of keywords related to industry i is n i , then the correlation calculation formula is: According to the correlation r i Add policy weight w to industry i i2 :w i2 =δ×r i ; Among them, δ is the policy weight coefficient, which is used to adjust the impact of policy factors on the weight.

7. A digital economic statistics system according to claim 6, characterized in that: The specific calculation method of the final weight is as follows: the objective weight of each weight is calculated using the entropy weight method. i ; Let x i1 =w i0 ,x i2 =w i1 ,x i3 =w i2 , then the entropy value H j for: in, Then we get the entropy weight e j : Use the analytic hierarchy process to determine the subjective weight of each weight i ; Comprehensive objective weight i and subjective weight s i , get the final weight w i :w i =θ×o i +(1-θ)×s i ; where θ is the balance coefficient, 0≤θ≤1, and 8. A digital economic statistics method, implemented based on the digital economic statistics system according to any one of claims 1 to 7, characterized in that: The following steps are involved: S1. Multi-source heterogeneous data collection and preprocessing: Obtain structured data on corporate financial statements and imports and exports from tax, industry and commerce, and customs departments; Analyze e-commerce platform transaction logs and logistics company shipping documents to obtain semi-structured data; Use NLP to extract economic keywords from news reports and social media to obtain unstructured data; Use timestamp alignment algorithms to align the time granularity of different data sources; based on industry knowledge graphs, correct semantic ambiguity in the data; S2. Dynamic weight allocation model: assign initial weights to various industries based on the national economic industry classification; By calculating the covariance between the volatility of industry stock indices and GDP growth rates, we dynamically adjust weights. Based on national industrial policy documents, we extract keywords and increase policy weights for related industries. The entropy weight method is combined with the hierarchical analysis method to obtain the final weight, ensuring that the total weight is 100% and avoiding the dominance of a single factor; S3. Real-time economic indicator calculation engine: Define the digital consumption index indicator: Calculate the degree of digitalization of the consumption structure based on mobile payment transaction volume, number of online education users, and advertising revenue of short video platforms; Define the industry resilience index indicator: Combine enterprise asset-liability ratio, R&D investment intensity, and supply chain diversification to assess the industry's risk resistance; The pre-processed data is classified by industry; a dynamic weight allocation model is applied to calculate the weighted value of each industry; and economic indicators are generated through real-time aggregation using a distributed computing framework; S4. Generation of regional economic profiles: Map economic indicators to geographic grids based on GIS to generate a heat map of industrial distribution; use LSTM to predict regional economic trends over the next three months; calculate the deviation between regional economic indicators and the national average to identify advantageous industries and areas of weakness.