Distributed statistical data analysis optimization system and method

Through distributed statistical data analysis and optimization systems and methods, problems such as inaccurate missing value processing in economic data processing and dependence on a single traditional filter are solved, and efficient and accurate data analysis and optimized resource utilization are achieved, ensuring data security and privacy.

CN120012135APending Publication Date: 2025-05-16SHENZHEN ZHIXIN DATA TECHNOLOGY SERVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510148490.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems such as inaccurate missing value processing when processing economic data, noisy filtering relies on a single traditional filter and requires manual intervention, inflexible data allocation strategies, waste of resources and insufficient system security.

Method used

A distributed statistical data analysis optimization system and method is proposed, including data acquisition, preprocessing, data sharding and task allocation, data analysis and encrypted storage. Economic data is obtained through crawlers, missing value processing and filtering is performed using clustering algorithms and deep learning models, data sharding is used to slice and task allocation is performed by dynamically adjusting weights, and data is finally encrypted and stored.

Benefits of technology

It realizes fine preprocessing and reasonable allocation of economic data, improves the efficiency and accuracy of data analysis, optimizes system resource utilization, avoids resource waste and overload, and ensures data security and privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012135A_ABST
    Figure CN120012135A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of economic data processing, and discloses a distributed statistical data analysis optimization system and method. Comprising the steps that an economic data set is collected, and the economic data set comprises economic numerical data and an economic time series data set; preprocessing the economic data set to obtain a preprocessed economic data set; performing data fragmentation on the preprocessed economic data set to obtain a fragmented economic data set; performing task allocation processing on the fragmented economic data set to obtain an optimal allocation data set; performing data analysis on the optimal distribution data set to obtain a prediction analysis result; combining the optimal allocation data set and the prediction analysis result into a prediction analysis data set; performing encryption processing on the prediction analysis data set by using an encryption algorithm to obtain an encrypted data set; sending the encrypted data set to a preset database for storage; high-precision processing is carried out on the economic data set, and resource utilization and safety guarantee are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of economic data processing, and more specifically, to a distributed statistical data analysis optimization system and method. Background Art

[0002] The patent application with the publication number CN114840641A discloses a statistical data analysis classification system and method, including a data processing system, wherein the data processing system includes a credibility division unit, a source definition unit, a data preliminary screening unit and an analysis integration unit, wherein the credibility division unit is used to define a trusted end, and set a trusted end data accuracy rating standard, divide a dangerous integrated area and a blacklist, and the credibility division unit is connected with a source definition unit, wherein the source definition unit is used to define directory and sub-directory data, and relates to the field of data processing technology. The statistical data analysis classification system and method divides the information acquisition points corresponding to the data source by credibility authorization, and sets a trusted end data accuracy rating standard, and uses the form of points to encourage different trusted ends to correct each other, thereby screening out low-error information acquisition points, reducing the processing difficulty, and achieving high-accuracy classification of statistical data.

[0003] The economic data of corporate revenue may contain missing values ​​or noisy data due to system errors or network delays during the collection process. However, the existing preprocessing methods for economic data are not accurate enough. For example, when processing missing values, the impact of the different positions of missing data points on the missing value processing results cannot be taken into account. When filtering noise, a single traditional filter is used to remove the noise, and manual intervention is frequently required, which greatly reduces the efficiency of data preprocessing. It is difficult to flexibly adjust the allocation strategy when allocating data, and fixed rules are often used to allocate data. For example, each node is required to process a fixed amount of data, which leads to waste of resources and reduced work efficiency. There is a lack of reasonable security protection when using databases to store data, resulting in insufficient system security.

[0004] In view of this, the present invention proposes a distributed statistical data analysis optimization system and method to solve the above problems. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a distributed statistical data analysis optimization method, comprising: S1. Collect economic data sets, including economic numerical data and economic time series data sets; S2. preprocessing the economic data set to obtain a preprocessed economic data set; S3. Slice the preprocessed economic data set to obtain a slicing economic data set; perform task allocation processing on the slicing economic data set to obtain an optimal allocation data set; S4. Perform data analysis on the optimal allocation data set to obtain a prediction analysis result; merge the optimal allocation data set and the prediction analysis result into a prediction analysis data set; encrypt the prediction analysis data set using an encryption algorithm to obtain an encrypted data set; and send the encrypted data set to a preset database for storage.

[0006] Furthermore, a crawler method is used to obtain economic numerical data of a pre-selected enterprise revenue data platform and economic time series data within a preset time period; the economic numerical data include the average consumption amount of each customer, and the economic time series data include the daily revenue amount; a number is added to any numerical data point in the economic numerical data, and all numerical data points of the economic numerical data are arranged in ascending order according to the size of the number; the preset time period to which the economic time series data belongs is divided into Y time period intervals, and all time series data points in each time period interval are regarded as one economic time series data; each economic time series data is numbered in the order of the time period interval, and all numbered economic time series data are integrated to obtain an economic time series data set; any time series data point in any economic time series data has a timestamp corresponding to it; the economic numerical data and the economic time series data set are integrated to obtain an economic data set.

[0007] Furthermore, the method of preprocessing the economic data set includes: Process the missing values ​​of economic numerical data to obtain complete numerical data; process the missing values ​​of economic time series data sets to obtain complete time series data; filter the complete time series data to obtain filtered time series data; combine the complete numerical data and the filtered time series data to obtain the preprocessed economic data set; The methods for handling missing values ​​of economic numerical data include: Traverse the economic numerical data, if the average consumption amount value in the numerical data point corresponding to the number is missing, then the numerical data point is determined as a missing numerical data point, and the position of the missing numerical data point is determined based on the number; additionally mark the missing numerical data points according to the traversal order; The Euclidean distance between each missing numerical data point and any other numerical data point is calculated, and all numerical data points in the economic numerical data are clustered using a clustering algorithm based on the Euclidean distance; the missing numerical data points corresponding to the additional labels are used as cluster centers, and if the Euclidean distance between any numerical data point and any cluster center is less than a preset Euclidean distance threshold, the numerical data point is classified into the cluster corresponding to the cluster center; if there is a numerical data point that does not belong to any cluster after clustering, the numerical data point is classified into the cluster with the closest Euclidean distance to it; based on the numerical data points in each cluster, the missing numerical data points are fitted with the missing values ​​of the average consumption amount to obtain the fitted value of the average consumption amount for each missing numerical data point; the value of the missing numerical data point is filled with the fitted value of the average consumption amount for each missing numerical data point to obtain complete numerical data.

[0008] Furthermore, the method of fitting the missing value of the average consumption amount for the missing numerical data points includes: The values ​​of all missing numerical data points are preliminarily estimated using the linear interpolation method to obtain the initial value of each missing numerical data point; The location distribution of missing numerical data points is divided into two categories, including one-category location distribution and two-category location distribution; one-category location distribution means that any two missing numerical data points are not adjacent to each other, and the second-category location distribution means that there are two or more missing numerical data points adjacent to each other; when any two missing numerical data points are not adjacent to each other, a first-category objective function is constructed; the average consumption amount fitting value when the missing numerical data points belong to the first-category location distribution is calculated by minimizing the first-category objective function; A type of objective function ;in, Indicates that the additional label is Initial values ​​for missing numerical data points; Indicates that the additional label is The set of all numerical data points in the cluster corresponding to the missing numerical data point; Representing a collection Any numerical data point in ; Indicates that the additional label is Missing numerical data points and sets Any numerical data point in The weight of is a constant; Use the gradient descent method to The initial value of the missing numerical data point is updated until the function value of the first type of objective function no longer decreases, and the average consumption amount fitting value when the missing numerical data point belongs to a type of position distribution is obtained; When there are two or more missing numerical data points adjacent to each other, a second-class objective function is constructed; Second type objective function ;in, Indicates that from the additional label The missing numerical data points are given additional labels The initial value of any missing numerical data point between the missing numerical data points of ; Indicates that from the additional label The missing numerical data points are given additional labels The set of all numerical data points in all clusters formed by the missing numerical data points; Representing a collection Any numerical data point in ; Indicates that from the additional label The missing numerical data points are given additional labels Any missing numerical data point between the missing numerical data points and the set The Euclidean distance of any numerical data point in ; is a constant; By calculating each of the two objective functions The partial derivatives are obtained to obtain a system of simultaneous equations; the system of simultaneous equations is solved to obtain the fitted value of the average consumption amount for each missing numerical data point belonging to the second type of location distribution.

[0009] Furthermore, the method of processing missing values ​​for the economic time series data set includes: Build a deep learning model and use the GAN network model as the basic framework of the deep learning model; collect a complete historical time series data set and use the data set as the validation set of the deep learning model; randomly select Mark the historical time series data points and mask the values ​​of these historical time series data points to obtain a fuzzy time series data set, which is used as the training set for the deep learning model. Construct the original matrix , where the rows of the original matrix represent any economic time series data in the fuzzy time series data set, and the columns of the original matrix represent the timestamps in any economic time series data in the fuzzy time series data set; the fuzzy time series data set is represented by the original matrix, where 0 represents missing time series data points and 1 represents known time series data points; Input the training set, the original matrix and a random noise vector generated based on Gaussian distribution into the generator of the deep learning model; define the attention score, calculate the attention score of each time series data point in each economic time series data, and sum the attention scores of all time series data points to obtain the attention vector; the generator generates the predicted value of the missing time series data point based on the input data and the attention vector, and uses the predicted value to fill the fuzzy time series data set to obtain the preliminary processed data set; the discriminator outputs a discrimination probability by comparing the validation set and the preliminary processed data set, and calculates the loss function of the deep learning model at the same time; repeat the generation and discrimination process until the discrimination probability output by the discriminator is greater than or equal to the preset judgment threshold and the function value of the loss function of the deep learning model no longer decreases, and the trained deep learning model is obtained at this time; The trained deep learning model is used to process missing values ​​in the economic time series data set to obtain complete time series data.

[0010] Furthermore, the method of filtering the complete time series data includes: Decompose the complete time series data into N time series signal components by using wavelet decomposition, preset the time series noise threshold, process the high-frequency noise in each time series signal component by soft threshold method, and obtain preliminary denoised time series data; construct an adaptive filter, and optimize the adaptive filter to obtain the optimal performance adaptive filter; use the optimal performance adaptive filter to denoise the preliminary denoised time series data to obtain filtered time series data; The method of constructing an adaptive filter to perform denoising on the preliminary denoised time series data includes: Initializing performance parameters of the adaptive filter, the performance parameters including a weight vector and a performance index of the adaptive filter; Based on the denoised time series data The value corresponding to the timestamp and the The weight vector corresponding to the timestamp is calculated The filtering error corresponding to the timestamp ; Construct an updated covariance matrix; based on the The filter error corresponding to the timestamp and the updated covariance matrix are used to update the weight vector of the adaptive filter to obtain the The update weight corresponding to the timestamp ; Based on The updated weights corresponding to the timestamps and the denoised time series data The value corresponding to the timestamp is calculated The filtering error corresponding to the timestamp is based on the The filtering error corresponding to the timestamp is calculated as Performance indicators corresponding to timestamps; Repeat the updating of the weight vector and the performance index until the value of the performance index is less than or equal to the preset performance index threshold; fix the parameters at this time to obtain the optimal performance adaptive filter.

[0011] Furthermore, the method of sharding the preprocessed economic data set includes: Based on different types of data in the preprocessed economic data set, different sharding strategies are used to shard the preprocessed economic data set; The complete numerical data in the preprocessed economic data set is sharded using the uniform sharding method; the amount of data in the complete numerical data is counted, and the complete numerical data is divided into U groups based on the amount of data; each group of complete numerical data is written into a numerical task shard, and each numerical task shard is numbered until all the complete numerical data are written into U numerical task shards, and a set of numerical task shards is obtained; The filtered time series data in the preprocessed economic data set are sharded using the time period sharding method; the time period interval to which each filtered time series data point in the filtered time series data belongs is queried, and the filtered time series data is divided into Y groups based on the time period interval; each group of filtered time series data is written into a time series task shard, and each time series task shard is marked using the time period interval until all the filtered time series data are written into Y time series task shards to obtain a time series task shard set; the time series task shard set and the numerical task shard set are combined to obtain a sharded economic data set.

[0012] Furthermore, the method of performing task allocation processing on the shard economic data set includes: Defining a node collection and node parameters, where Represents each node in the node set; node parameters include the load limit of any node ;in, Indicates the upper limit of the computational load of any node in the node set; Indicates the upper limit of memory load of any node in the node set; Query the data complexity and memory usage of each task shard in the shard economic data set; calculate the load of each task shard based on the data complexity and memory usage; The load of each task shard ;in, Indicates the data complexity of any task shard; Indicates the memory usage of any task slice; The weight representing the complexity of the data; The weight representing the memory usage; The weight of data complexity and the weight of memory usage are dynamically adjusted. The calculation formula for dynamic adjustment is: ;in, Represents the mean data complexity of all task shards; The variance representing the complexity of the data; Indicates the variance of memory usage; Represents the sum of the upper limits of computing loads of all nodes; Indicates the total computing load that has been used; ;in, Represents the average memory usage of all task slices; Indicates the sum of the memory load limits of all nodes; Indicates the total memory load that has been used; Based on the load of each task slice, all task slices are arranged in descending order from large to small to obtain a sorted slice data set; a node is allocated to each task slice in the sorted slice data set to obtain an allocated node sequence; the allocated node sequence is traversed from the first node to check the load of the task slice in the second node. If the total load obtained by adding the load of the task slice in the second node to the load of the task slice in the first node is less than or equal to the load upper limit of the node, the task slice in the second node is added to the first node, and the second node is released at the same time; otherwise, the total load of the task slice of the third node and the first node is calculated until the load of the first node reaches the load upper limit or cannot accommodate any of the remaining task slices. At this time, traversal starts from the second node and the process of calculating the total load to fill the node is repeated; the resource utilization of each node is calculated. If there is a node whose resource utilization is lower than the preset utilization threshold, the task slice of the node is added to other nodes until all task slices are allocated and the number of nodes containing task slices reaches the minimum. At this time, the optimal node sequence is obtained, which is the optimal allocated data set.

[0013] Furthermore, the method of performing data analysis on the optimal allocation data set includes: Feature extraction is performed on the optimal allocation data set to obtain the economic numerical data features and economic time series data features of the optimal allocation data set; a machine learning model is constructed based on the economic numerical data features and the economic time series data features to process the optimal allocation data set, and the average consumption amount prediction results and daily revenue amount prediction results are obtained respectively; the average consumption amount prediction results and the daily revenue amount prediction results are combined to obtain the prediction analysis results.

[0014] A distributed statistical data analysis and optimization system, which is used to implement a distributed statistical data analysis and optimization method, comprising: A data collection module is used to collect economic data sets, which include economic numerical data and economic time series data sets; A preprocessing module is used to preprocess the economic data set to obtain a preprocessed economic data set; The allocation module is used to perform data sharding on the preprocessed economic data set to obtain a sharded economic data set; perform task allocation processing on the sharded economic data set to obtain an optimally allocated data set; The analysis module is used to perform data analysis on the optimal allocation data set to obtain a prediction analysis result; merge the optimal allocation data set and the prediction analysis result into a prediction analysis data set; encrypt the prediction analysis data set using an encryption algorithm to obtain an encrypted data set; send the encrypted data set to a preset database for storage; and connect each module via wired and / or wireless means.

[0015] The technical effects and advantages of a distributed statistical data analysis optimization system and method of the present invention are as follows: By collecting statistical data, which includes the average consumption amount data of each customer in the enterprise revenue economic data and the daily revenue amount data of the enterprise; the statistical data are carefully preprocessed, and at the same time, the statistical data are reasonably distributed, the data analysis of the statistical data is completed, and the statistical data is encrypted and stored, realizing the optimization of the statistical data analysis process in the field of enterprise revenue economic data; compared with the existing experience, the preprocessing is completed for different types of data using appropriate methods, which is more accurate than traditional methods in large data sets, and indirectly improves the efficiency of data analysis; considering the system load and storage space, the statistical data are segmented, and an effective task allocation strategy is proposed to optimize the resource utilization of the system and avoid the occurrence of resource waste or overload; a model is built to analyze the statistical data, and the numerical and development trend of the statistical data are predicted. At the same time, the statistical data and the prediction results are encrypted and stored using encryption algorithms to ensure the security and privacy of the statistical data in the big data environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of a distributed statistical data analysis optimization method of the present invention; Figure 2 A schematic diagram of a distributed statistical data analysis and optimization system of the present invention. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Example

[0018] See also Figure 1 As shown, the distributed statistical data analysis optimization method described in this embodiment includes: S1. Collect economic data sets, including economic numerical data and economic time series data sets; S2. preprocessing the economic data set to obtain a preprocessed economic data set; S3. Slice the preprocessed economic data set to obtain a slicing economic data set; perform task allocation processing on the slicing economic data set to obtain an optimal allocation data set; S4. Perform data analysis on the optimal allocation data set to obtain a prediction analysis result; merge the optimal allocation data set and the prediction analysis result into a prediction analysis data set; encrypt the prediction analysis data set using an encryption algorithm to obtain an encrypted data set; and send the encrypted data set to a preset database for storage.

[0019] Use a crawler method (for example, use Python to crawl web pages or data in databases) to obtain economic numerical data of a pre-selected enterprise revenue data platform and economic time series data within a preset time period; the economic numerical data includes the average consumption amount of each customer, and the economic time series data includes the daily revenue amount (the daily revenue amount is different within the preset time period, gradually changes and shows a certain trend, so it is regarded as time series data); add a number to any numerical data point in the economic numerical data, and arrange all numerical data points of the economic numerical data in ascending order according to the size of the number; divide the preset time period to which the economic time series data belongs into Y time period intervals (the value of Y can be selected according to the specific situation), and regard all time series data points in each time period interval as one economic time series data; number each economic time series data in the order of the time period interval, and integrate all numbered economic time series data to obtain an economic time series data set; any time series data point in any economic time series data has a timestamp corresponding to it; integrate the economic numerical data and the economic time series data set to obtain an economic data set.

[0020] Before analyzing the statistical data in the field of corporate revenue economic data, the data needs to be preprocessed, because there may be certain errors in the original data collected initially. For example, economic numerical data is prone to missing values, and economic time series data sets may not only have missing values ​​but also noise signals. If the errors in the original data are not processed, directly using the original data for data analysis will only lead to erroneous results or no analysis results, which will undermine the integrity and accuracy of the data analysis. Therefore, it is necessary to preprocess the original data.

[0021] Ways to preprocess economic data sets include: Process the missing values ​​of economic numerical data to obtain complete numerical data; process the missing values ​​of economic time series data sets to obtain complete time series data; filter the complete time series data to obtain filtered time series data; combine the complete numerical data and filtered time series data to obtain the preprocessed economic data set.

[0022] The methods for handling missing values ​​of economic numerical data include: In a complex data set with a large amount of data, only using traditional missing value processing methods (such as calculating the mean or median to fill the missing values) may cause the calculated missing values ​​to become erroneous data. Therefore, choosing to use clustering algorithms to process missing values ​​for economic numerical data is an effective method. By calculating the Euclidean distance between numerical data points and grouping the numerical data points into different clusters, numerical data points with similar characteristics can be identified, and information can be obtained from the numerical data points belonging to the same cluster to predict the values ​​of missing numerical data points.

[0023] The economic numerical data are traversed. If the average consumption amount value in the numerical data point corresponding to the number is missing, the numerical data point is determined to be a missing numerical data point, and the position of the missing numerical data point is determined based on the number; the missing numerical data points are additionally numbered in the traversal order (the additional number is greater than 1 and less than or equal to the number of missing numerical data points).

[0024] Calculate the Euclidean distance between each missing numerical data point and any other numerical data point, and cluster all numerical data points in the economic numerical data using a clustering algorithm (such as the DBSCAN clustering algorithm) based on the Euclidean distance; use the missing numerical data points corresponding to the additional labels as cluster centers, and if the Euclidean distance between any numerical data point and any cluster center is less than a preset Euclidean distance threshold, classify the numerical data point into the cluster corresponding to the cluster center; if there is a numerical data point that does not belong to any cluster after clustering, assign the numerical data point to the cluster with the closest Euclidean distance to it; fit the missing values ​​of the average consumption amount of the missing numerical data points based on the numerical data points in each cluster to obtain the average consumption amount fitting value of each missing numerical data point; use the average consumption amount fitting value of each missing numerical data point to fill in the value of the missing numerical data point to obtain complete numerical data.

[0025] The methods for fitting the missing values ​​of the average consumption amount for missing numerical data points include: In the process of fitting the missing values ​​of the average consumption amount for missing numerical data points, considering that there are two possible location distributions of the missing numerical data points in the original numerical type data, a different objective function is constructed for each possible location distribution, and the fitted value of the average consumption amount for each missing numerical data point is obtained by minimizing the objective function.

[0026] The linear interpolation method is used to preliminarily estimate the values ​​of all missing numerical data points to obtain the initial value of each missing numerical data point.

[0027] The location distribution of missing numerical data points is divided into two categories, including type I location distribution and type II location distribution; type I location distribution is that any two missing numerical data points are not adjacent to each other, and type II location distribution is that there are two or more missing numerical data points that are adjacent to each other; when any two missing numerical data points are not adjacent to each other, a type I objective function is constructed; the average consumption amount fitting value when the missing numerical data points belong to type I location distribution is calculated by minimizing the type I objective function.

[0028] A type of objective function ;in, Indicates that the additional label is Initial values ​​for missing numerical data points; Indicates that the additional label is The set of all numerical data points in the cluster corresponding to the missing numerical data point; Representing a collection Any numerical data point in ; Indicates that the additional label is Missing numerical data points and sets Any numerical data point in The weight of ,in Indicates that the additional label is Missing numerical data points and sets The Euclidean distance of any numerical data point in the data set; if the Euclidean distance between two points is closer, the similarity between the two points is higher and the weight value is larger; if the Euclidean distance between two points is farther, the similarity between the two points is lower and the weight value is smaller); is a constant ( ).

[0029] Use the gradient descent method to The initial value of the missing numerical data point is updated until the function value of the first-class objective function no longer decreases. At this time, the average consumption amount fitting value when the missing numerical data point belongs to a class of location distribution is obtained.

[0030] When there are multiple adjacent missing numerical data points, the dependency between the missing numerical data points increases, and the numerical data points that can be used for reference decrease, resulting in a lack of information. Inference can only be made based on global characteristics. In addition, adjacent missing numerical data points affect each other, making it difficult to accurately calculate the average consumption amount fitting value of the missing numerical data points. For example ,in, Indicates missing numerical data points; in this case, if you want to calculate the average consumption amount fitting value of the missing numerical data points, you can only use and Make a judgment, but and and The distance between the adjacent missing numerical data points is far, and errors are inevitable in the calculation results; therefore, it is necessary to construct an objective function that can take into account the adjacent missing numerical data points for calculation.

[0031] When there are two or more missing numerical data points adjacent to each other, a second-class objective function is constructed.

[0032] Second type objective function ;in, Indicates that from the additional label The missing numerical data points are given additional labels The initial value of any missing numerical data point between the missing numerical data points of (all missing numerical data points in this interval are adjacent); Indicates that from the additional label The missing numerical data points are given additional labels The set of all numerical data points in all clusters formed by the missing numerical data points; Representing a collection Any numerical data point in ; Indicates that from the additional label The missing numerical data points are given additional labels Any missing numerical data point between the missing numerical data points and the set The Euclidean distance of any numerical data point in ; is a constant ( ).

[0033] By calculating each of the two objective functions The partial derivatives are obtained to obtain a system of simultaneous equations; the system of simultaneous equations is solved to obtain the fitted value of the average consumption amount for each missing numerical data point belonging to the second type of location distribution.

[0034] The location of missing numerical data points is discussed in a classified manner. When calculating the fitting value of the average consumption amount of missing numerical data points, the appropriate objective function can be selected according to the specific situation. This improves the flexibility of missing value fitting, reduces the error of missing value processing, and indirectly improves the accuracy and reliability of subsequent data analysis.

[0035] Ways to handle missing values ​​for economic time series data sets include: The GAN network model uses a generative adversarial mechanism to continuously optimize the generator and ultimately generate data that is closest to the real data. This model is generally used to generate new image data, but it can also predict missing values ​​in other types of data and can adapt to various complex data missing scenarios, such as random missing, continuous missing, or structural missing. Since economic time series data has obvious time correlation, an attention mechanism is added to the GAN network model to accurately capture the dependency between time series data points corresponding to different timestamps, and generate prediction values ​​based on the dependency that conform to the time series characteristics of the current economic time series data, thereby improving the prediction accuracy and efficiency of the model.

[0036] Build a deep learning model and use the GAN network model as the basic framework of the deep learning model; collect a complete historical time series data set and use the data set as the validation set of the deep learning model; randomly select Historical time series data points (including Less than the total number of historical time series data points in the historical time series data set) are marked and the values ​​of these historical time series data points are masked to obtain a fuzzy time series data set, which is used as a training set for the deep learning model.

[0037] Construct the original matrix , where the rows of the original matrix represent any economic time series data in the fuzzy time series data set, and the columns of the original matrix represent the timestamps in any economic time series data in the fuzzy time series data set; the original matrix is ​​used to represent the fuzzy time series data set, where 0 represents missing time series data points (indicating that the value of the time series data point corresponding to the timestamp does not exist), and 1 represents known time series data points; for example, Indicates The timestamp in the time series data is The time series data points of are known time series data points; Indicates The timestamp in the time series data is The time series data points are missing time series data points.

[0038] The training set, the original matrix and a random noise vector generated based on Gaussian distribution are input into the generator of the deep learning model; the attention score (such as the dot product function) is defined, the attention score of each time series data point in each economic time series data is calculated, and the attention scores of all time series data points are summed up to obtain the attention vector; the generator generates the predicted value of the missing time series data point based on the input data (including the training set, the original matrix and the random noise vector) and the attention vector, and uses the predicted value to fill the fuzzy time series data set to obtain the preliminary processed data set (the original matrix helps the generator identify the location of the missing time series data points in the training set, ensuring that only the values ​​of the missing time series data points are generated; the random noise vector provides diversity and Randomness makes the generator's generation results richer; the attention vector can reflect the temporal dependency of the time series, and is used to assist the generator in generating values ​​that conform to the temporal characteristics of the training set; the generator generates the predicted values ​​of the missing time series data points by combining the training set, the attention vector and the random noise vector, and outputs the supplemented data set); the discriminator outputs a discrimination probability by comparing the validation set and the preliminary processed data set, and calculates the loss function of the deep learning model (including the adversarial loss function and the reconstruction loss function of the GAN network model); the generation and discrimination process is repeated until the discrimination probability output by the discriminator is greater than or equal to the preset judgment threshold and the function value of the loss function of the deep learning model no longer decreases, and a trained deep learning model is obtained.

[0039] The trained deep learning model is used to process missing values ​​in the economic time series data set to obtain complete time series data.

[0040] The methods for filtering the complete time series data include: In time series data, noise often appears in the form of high-frequency signals, so wavelet decomposition is used to remove high-frequency noise; in existing denoising algorithms, wavelet decomposition is a powerful noise processing tool, and has developed quite maturely, which can effectively separate useful signals and noise; after removing the high-frequency noise in the complete time series data through wavelet decomposition, an adaptive filter is constructed to further optimize the filtering processing of the data; the combination of wavelet decomposition and adaptive filter can give full play to the advantages of both and obtain more accurate filtered time series data.

[0041] The complete time series data is decomposed into N time series signal components by using wavelet decomposition, a time series noise threshold is preset, and the high-frequency noise in each time series signal component is processed by the soft threshold method (if the absolute value of a time series signal component is less than the preset time series noise threshold, the time series signal component is set to zero; if the absolute value of a time series signal component is greater than the preset time series noise threshold, the time series component is subtracted from the preset time series noise threshold and the absolute value is taken), and preliminary denoised time series data is obtained; an adaptive filter is constructed, and the adaptive filter is optimized to obtain the optimal performance adaptive filter; the optimal performance adaptive filter is used to denoise the preliminary denoised time series data to obtain filtered time series data.

[0042] The method of constructing an adaptive filter to perform denoising on the preliminary denoised time series data includes: Initialize the performance parameters of the adaptive filter, which include a weight vector and a performance index of the adaptive filter.

[0043] Based on the denoised time series data The value corresponding to the timestamp and the The weight vector corresponding to the timestamp is calculated The filtering error corresponding to the timestamp ; The calculation formula of filtering error is: ;in, The first output of the adaptive filter is The value corresponding to the timestamp; Indicates The weight vector corresponding to the timestamp; Indicates the first The value corresponding to the timestamp.

[0044] Use Bayesian estimation to construct and update the covariance matrix; based on the The filter error corresponding to the timestamp and the updated covariance matrix are used to update the weight vector of the adaptive filter to obtain the The update weight corresponding to the timestamp ; The calculation formula for updating the weight vector is: ;in, Indicates The updated covariance matrix corresponding to the timestamp.

[0045] Based on The updated weights corresponding to the timestamps and the denoised time series data The value corresponding to the timestamp is calculated The filtering error corresponding to the timestamp is based on the The filtering error corresponding to the timestamp is calculated as The performance indicator corresponding to the timestamp.

[0046] No. Performance indicators corresponding to timestamps ;in, Indicates The filtering error corresponding to the timestamp.

[0047] Repeat the updating of the weight vector and the performance index until the value of the performance index is less than or equal to the preset performance index threshold; fix the parameters at this time to obtain the optimal performance adaptive filter.

[0048] By dynamically adjusting the performance parameters of the adaptive filter, the filter can automatically optimize its performance according to changes in input data; the flexibility of the filtering process is improved, manual intervention is reduced, and the efficiency of the filtering process is improved.

[0049] The methods for sharding the preprocessed economic data set include: Based on different types of data in the preprocessed economic dataset, different sharding strategies are used to shard the preprocessed economic dataset.

[0050] The complete numerical data in the preprocessed economic data set is sharded using the uniform sharding method; the amount of data in the complete numerical data is counted, and the complete numerical data is divided into U groups based on the data amount; for example, if there are 1000 complete numerical data points in total, every 100 complete numerical data points are divided into a group; each group of complete numerical data is written into a numerical task shard (a numerical task shard refers to a separate file or a database table used to record each group of complete numerical data), and each numerical task shard is numbered until all the complete numerical data are written into U numerical task shards to obtain a set of numerical task shards.

[0051] The filtered time series data in the preprocessed economic data set are sharded using the time period sharding method; the time period interval to which each filtered time series data point in the filtered time series data belongs is queried, and the filtered time series data is divided into Y groups based on the time period interval; for example, the filtered time series data points belonging to the same time period interval are grouped together; each group of filtered time series data is written into a time series task shard, and each time series task shard is marked using the time period interval (marking means that each time series task shard corresponds to a time period interval), until all the filtered time series data are written into Y time series task shards, and a time series task shard set is obtained; the time series task shard set and the numerical task shard set are combined to obtain a sharded economic data set.

[0052] The ways to allocate tasks to shard economic data sets include: Since different task shards may have significant differences in data complexity (the amount of computing resources required for the task shard during execution, such as the number of operations required to execute) and memory usage (the storage space required for the task shard), using only a simple fixed allocation method cannot reflect the actual resource requirements of these task shards, which can easily lead to node resource overload or waste. For a system, computing power and storage space are limited, so an efficient allocation method is needed to allocate tasks to the sharded economic data set, with the goal of completing the allocation of all task shards with the least number of nodes.

[0053] Defining a node collection and node parameters, where Represents each node in the node set; node parameters include the load limit of any node ;in, Indicates the upper limit of the computational load of any node in the node set; Indicates the upper limit of memory load of any node in the node set.

[0054] Query the data complexity and memory usage of each task shard in the shard economic data set; calculate the load of each task shard based on the data complexity and memory usage.

[0055] The load of each task shard ;in, Indicates the data complexity of any task shard; Indicates the memory usage of any task slice; The weight representing the complexity of the data; The weight representing the memory usage.

[0056] The data complexity and memory usage of different types of task shards are different. For example, numerical data generally has higher data complexity but lower memory usage; time series data generally has lower data complexity but higher memory usage. Therefore, it is necessary to dynamically adjust the weight of data complexity and memory usage based on the remaining computing load and memory load in the system after each allocation, so as to accurately measure the actual load of the task shards and ensure that the allocation of task shards is more reasonable.

[0057] The weight of data complexity and the weight of memory usage are dynamically adjusted. The calculation formula for dynamic adjustment is: ;in, Represents the mean data complexity of all task shards; The variance representing the complexity of the data; Indicates the variance of memory usage; Represents the sum of the upper limits of computing loads of all nodes; Indicates the total computing load that has been used.

[0058] ;in, Represents the average memory usage of all task slices; Indicates the sum of the memory load limits of all nodes; Indicates the total memory load that has been used.

[0059] In order to maximize the use of the load of each node, a greedy algorithm is used to allocate task slices to achieve the goal of minimizing the number of nodes; when the amount of data is large and computing resources are limited, the use of a greedy algorithm for allocation has the advantages of high efficiency and low complexity.

[0060] Based on the load of each task slice, all task slices are sorted in descending order from large to small to obtain a sorted slice data set; a node is assigned to each task slice in the sorted slice data set to obtain an assigned node sequence; the assigned node sequence is traversed from the first node to check the load of the task slice in the second node. If the total load obtained by adding the load of the task slice in the second node to the load of the task slice in the first node is less than or equal to the node's load limit, the task slice in the second node is added to the first node and the second node is released at the same time; otherwise, the total load of the task slices of the third node and the first node is calculated until the load of the first node reaches the load limit or cannot accommodate any of the remaining task slices. At this time, traversal is started from the second node, and the process of calculating the total load and filling nodes is repeated; for example, the load limit of each node is 20, the first node already contains a task slice with a load of 11, and the load of the task slice in the second node is 8. Since the total load of the task slices in the two nodes is less than 20, The task slice of the second node is added to the first node, and the second node is released. At this time, the load of the second node is 0. If the load of the task slice in the second node is 10, it cannot be added to the first node. At this time, the load of the task slice of the third node is checked. At this time, the load of the task slice of the third node must be less than or equal to 8, then the task slice can be added to the first node, and the judgment method for all subsequent nodes is similar. When the first node reaches the load limit or cannot accommodate any other task slices, for example, when the load of the first node is 20 or the load is 19, there is no node with a load of 1, then traverse from the second node to determine whether there is a task slice with a suitable load that can be added to the node, and the operation of all other nodes is similar. Calculate the resource utilization of each node. If there is a node whose resource utilization is lower than the preset utilization threshold, add the task slice of the node to other nodes until all task slices are allocated and the number of nodes containing task slices is minimized. At this time, the optimal node sequence is obtained, which is the optimal allocation data set.

[0061] By efficiently allocating tasks to sharded economic data sets, it is able to adapt to the load requirements of diverse task shards, optimize system resource utilization, and improve system work efficiency.

[0062] Methods for data analysis of the optimal allocation data set include: Feature extraction is performed on the optimal allocation data set to obtain the economic numerical data features and economic time series data features of the optimal allocation data set; a machine learning model is constructed based on the economic numerical data features and the economic time series data features to process the optimal allocation data set, and the average consumption amount prediction results and the daily revenue amount prediction results are obtained respectively, and the average consumption amount prediction results and the daily revenue amount prediction results are combined to obtain the prediction analysis results; for example, the time series decomposition algorithm is used to extract features of the filtered time series data in the optimal allocation data set, and the trend features and periodic features of the filtered time series data are obtained within a preset time window to obtain the economic time series data features; the principal component analysis algorithm is used to reduce the dimension of the complete numerical data in the optimal allocation data set and generate principal components, and the principal components are used as the economic numerical data features; a numerical LSTM model and a time series LSTM model are constructed to collect historical economic numerical data, and feature extraction is performed on the historical economic numerical data to obtain The historical economic numerical data features are used as training labels; the historical economic numerical data are used to train the numerical LSTM model to obtain a trained numerical LSTM model, and the economic numerical data features and the complete numerical data in the optimal allocation data set are used as the input data of the trained numerical LSTM model (the economic numerical data features are used to guide the model to generate the prediction results of the complete numerical data), and the average consumption amount prediction results are obtained; historical economic time series data are collected, and the historical economic time series data features obtained by feature extraction of the historical economic time series data are used as training labels; the historical economic time series data are used to train the time series LSTM model to obtain a trained time series LSTM model, and the economic time series data features and the filtered time series data in the optimal allocation data set are used as the input data of the trained numerical LSTM model (the time series data features are used to guide the model to predict the future development trend of the filtered time series data), and the daily revenue amount prediction results are obtained.

[0063] Methods of using encryption algorithms to encrypt predictive analysis data sets include: The data in the prediction analysis data set are uniformly converted into binary format, and the converted prediction analysis data set is encoded to obtain the data set code; an encryption key and a decryption key are generated using an asymmetric encryption algorithm, the data set code is encrypted using the encryption key, the key and the data set code are combined into an encrypted ciphertext, and the encrypted ciphertext and the decryption key are combined to obtain an encrypted data set; after the database receives the encrypted data set, it identifies the correctness of the decryption key, and if it is correct, the encrypted data set is read and stored at the same time.

[0064] This embodiment collects statistical data, which includes the average consumption amount data of each customer in the enterprise's revenue economic data and the daily revenue amount data of the enterprise; performs fine preprocessing on the statistical data, and reasonably distributes the statistical data, completes the data analysis of the statistical data, and encrypts and stores the statistical data, thereby optimizing the statistical data analysis process in the field of enterprise revenue economic data; compared with existing experience, preprocessing is completed for different types of data using appropriate methods, which is more accurate than traditional methods in data sets with huge data volumes, and indirectly improves the efficiency of data analysis; considering the system load and storage space, the statistical data is segmented, and an effective task allocation strategy is proposed, which optimizes the resource utilization of the system and avoids the occurrence of resource waste or overload; a model is built to analyze the statistical data, and the statistical data is predicted in terms of numerical values ​​and development trends. At the same time, the statistical data and the prediction results are encrypted and stored using encryption algorithms, thereby ensuring the security and privacy of the statistical data in a big data environment. Example

[0065] See also Figure 2 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a distributed statistical data analysis and optimization system, including: A data collection module is used to collect economic data sets, which include economic numerical data and economic time series data sets; A preprocessing module is used to preprocess the economic data set to obtain a preprocessed economic data set; The allocation module is used to perform data sharding on the preprocessed economic data set to obtain a sharded economic data set; perform task allocation processing on the sharded economic data set to obtain an optimally allocated data set; The analysis module is used to perform data analysis on the optimal allocation data set to obtain a prediction analysis result; merge the optimal allocation data set and the prediction analysis result into a prediction analysis data set; encrypt the prediction analysis data set using an encryption algorithm to obtain an encrypted data set; send the encrypted data set to a preset database for storage; and connect each module via wired and / or wireless means. Example

[0066] This embodiment discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the operation mode of the distributed statistical data analysis optimization method provided above is implemented.

[0067] Since the electronic device introduced in this embodiment is an electronic device used to implement a distributed statistical data analysis optimization method in the embodiment of this application, based on the distributed statistical data analysis optimization method introduced in the embodiment of this application, the technical personnel of this field can understand the specific implementation of the electronic device of this embodiment and its various variations, so how the electronic device implements the method in the embodiment of this application is not introduced in detail here. As long as the technical personnel of this field implement the electronic device used in the distributed statistical data analysis optimization method in the embodiment of this application, it belongs to the scope of protection of this application.

[0068] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0069] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technical users in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A distributed statistical data analysis optimization method, characterized in that: include: S1. Collect economic data sets, including economic numerical data and economic time series data sets; S2. preprocessing the economic data set to obtain a preprocessed economic data set; S3. Slice the preprocessed economic data set to obtain a slicing economic data set; Perform task allocation processing on the shard economic data set to obtain the optimal allocation data set; S4. Performing data analysis on the optimal allocation data set to obtain a prediction analysis result; merging the optimal allocation data set and the prediction analysis result into a prediction analysis data set; The prediction analysis data set is encrypted using an encryption algorithm to obtain an encrypted data set; the encrypted data set is sent to a preset database for storage.

2. A distributed statistical data analysis optimization method according to claim 1, characterized in that: Use the crawler method to obtain the economic numerical data of the pre-selected enterprise revenue data platform and the economic time series data within the preset time period; the economic numerical data includes the average consumption amount of each customer, and the economic time series data includes the daily revenue amount; add a number to any numerical data point in the economic numerical data, and arrange all numerical data points of the economic numerical data in ascending order according to the size of the number; divide the preset time period to which the economic time series data belongs into Y time period intervals, and regard all time series data points in each time period interval as one piece of economic time series data; number each piece of economic time series data in the order of the time period interval, and integrate all numbered economic time series data to obtain an economic time series data set; Any time series data point in any economic time series data has a timestamp corresponding to it; the economic numerical data and the economic time series data set are integrated to obtain the economic data set.

3. A distributed statistical data analysis optimization method according to claim 2, characterized in that: The method of preprocessing the economic data set includes: Process the missing values ​​of economic numerical data to obtain complete numerical data; process the missing values ​​of economic time series data sets to obtain complete time series data; filter the complete time series data to obtain filtered time series data; combine the complete numerical data and the filtered time series data to obtain the preprocessed economic data set; The methods for handling missing values ​​of economic numerical data include: Traverse the economic numerical data, if the average consumption amount value in the numerical data point corresponding to the number is missing, then the numerical data point is determined as a missing numerical data point, and the position of the missing numerical data point is determined based on the number; additionally mark the missing numerical data points according to the traversal order; The Euclidean distance between each missing numerical data point and any other numerical data point is calculated, and all numerical data points in the economic numerical data are clustered using a clustering algorithm based on the Euclidean distance; the missing numerical data points corresponding to the additional labels are used as cluster centers, and if the Euclidean distance between any numerical data point and any cluster center is less than a preset Euclidean distance threshold, the numerical data point is classified into the cluster corresponding to the cluster center; if there is a numerical data point that does not belong to any cluster after clustering, the numerical data point is classified into the cluster with the closest Euclidean distance to it; based on the numerical data points in each cluster, the missing numerical data points are fitted with the missing values ​​of the average consumption amount to obtain the fitted value of the average consumption amount for each missing numerical data point; the value of the missing numerical data point is filled with the fitted value of the average consumption amount for each missing numerical data point to obtain complete numerical data.

4. A distributed statistical data analysis optimization method according to claim 3, characterized in that: The method of fitting the missing value of the average consumption amount for the missing numerical data points includes: The values ​​of all missing numerical data points are preliminarily estimated using the linear interpolation method to obtain the initial value of each missing numerical data point; The location distribution of missing numerical data points is divided into two categories, including one-category location distribution and two-category location distribution; one-category location distribution means that any two missing numerical data points are not adjacent to each other, and the second-category location distribution means that there are two or more missing numerical data points adjacent to each other; when any two missing numerical data points are not adjacent to each other, a first-category objective function is constructed; the average consumption amount fitting value when the missing numerical data points belong to the first-category location distribution is calculated by minimizing the first-category objective function; A type of objective function ;in, Indicates that the additional label is Initial values ​​for missing numerical data points; Indicates that the additional label is The set of all numerical data points in the cluster corresponding to the missing numerical data point; Representing a collection Any numerical data point in ; Indicates that the additional label is Missing numerical data points and sets Any numerical data point in The weight of is a constant; Use the gradient descent method to The initial value of the missing numerical data point is updated until the function value of the first type of objective function no longer decreases, and the average consumption amount fitting value when the missing numerical data point belongs to a type of position distribution is obtained; When there are two or more missing numerical data points adjacent to each other, a second-class objective function is constructed; Second type objective function ;in, Indicates that from the additional label The missing numerical data points are given additional labels The initial value of any missing numerical data point between the missing numerical data points of ; Indicates that from the additional label The missing numerical data points are given additional labels The set of all numerical data points in all clusters formed by the missing numerical data points; Representing a collection Any numerical data point in ; Indicates that from the additional label The missing numerical data points are given additional labels Any missing numerical data point between the missing numerical data points and the set The Euclidean distance of any numerical data point in ; is a constant; By calculating each of the two objective functions The partial derivatives are obtained to obtain a system of simultaneous equations; the system of simultaneous equations is solved to obtain the fitted value of the average consumption amount for each missing numerical data point belonging to the second type of location distribution.

5. A distributed statistical data analysis optimization method according to claim 4, characterized in that: The method of processing missing values ​​for economic time series data sets includes: Build a deep learning model and use the GAN network model as the basic framework of the deep learning model; collect a complete historical time series data set and use the data set as the validation set of the deep learning model; randomly select Mark the historical time series data points and mask the values ​​of these historical time series data points to obtain a fuzzy time series data set, which is used as the training set for the deep learning model. Construct the original matrix , where the rows of the original matrix represent any economic time series data in the fuzzy time series data set, and the columns of the original matrix represent the timestamps in any economic time series data in the fuzzy time series data set; the fuzzy time series data set is represented by the original matrix, where 0 represents missing time series data points and 1 represents known time series data points; Input the training set, the original matrix and a random noise vector generated based on Gaussian distribution into the generator of the deep learning model; define the attention score, calculate the attention score of each time series data point in each economic time series data, and sum the attention scores of all time series data points to obtain the attention vector; the generator generates the predicted value of the missing time series data point based on the input data and the attention vector, and uses the predicted value to fill the fuzzy time series data set to obtain the preliminary processed data set; the discriminator outputs a discrimination probability by comparing the validation set and the preliminary processed data set, and calculates the loss function of the deep learning model at the same time; repeat the generation and discrimination process until the discrimination probability output by the discriminator is greater than or equal to the preset judgment threshold and the function value of the loss function of the deep learning model no longer decreases, and the trained deep learning model is obtained at this time; The trained deep learning model is used to process missing values ​​in the economic time series data set to obtain complete time series data.

6. A distributed statistical data analysis optimization method according to claim 5, characterized in that: The method of filtering the complete time series data includes: Decompose the complete time series data into N time series signal components by using wavelet decomposition, preset the time series noise threshold, process the high-frequency noise in each time series signal component by soft threshold method, and obtain preliminary denoised time series data; construct an adaptive filter, and optimize the adaptive filter to obtain the optimal performance adaptive filter; use the optimal performance adaptive filter to denoise the preliminary denoised time series data to obtain filtered time series data; The method of constructing an adaptive filter to perform denoising on the preliminary denoised time series data includes: Initializing performance parameters of the adaptive filter, the performance parameters including a weight vector and a performance index of the adaptive filter; Based on the denoised time series data The value corresponding to the timestamp and the The weight vector corresponding to the timestamp is calculated The filtering error corresponding to the timestamp ; Construct an updated covariance matrix; based on the The filter error corresponding to the timestamp and the updated covariance matrix are used to update the weight vector of the adaptive filter to obtain the The update weight corresponding to the timestamp ; Based on The updated weights corresponding to the timestamps and the denoised time series data The value corresponding to the timestamp is calculated The filtering error corresponding to the timestamp is based on the The filtering error corresponding to the timestamp is calculated as Performance indicators corresponding to timestamps; Repeat the updating of the weight vector and the performance index until the value of the performance index is less than or equal to the preset performance index threshold; fix the parameters at this time to obtain the optimal performance adaptive filter.

7. A distributed statistical data analysis optimization method according to claim 6, characterized in that: The method of sharding the preprocessed economic data set includes: Based on different types of data in the preprocessed economic data set, different sharding strategies are used to shard the preprocessed economic data set; The complete numerical data in the preprocessed economic data set is sharded using the uniform sharding method; the amount of data in the complete numerical data is counted, and the complete numerical data is divided into U groups based on the amount of data; each group of complete numerical data is written into a numerical task shard, and each numerical task shard is numbered until all the complete numerical data are written into U numerical task shards, and a set of numerical task shards is obtained; The filtered time series data in the preprocessed economic data set are sharded using the time period sharding method; the time period interval to which each filtered time series data point in the filtered time series data belongs is queried, and the filtered time series data is divided into Y groups based on the time period interval; each group of filtered time series data is written into a time series task shard, and each time series task shard is marked using the time period interval until all the filtered time series data are written into Y time series task shards to obtain a time series task shard set; the time series task shard set and the numerical task shard set are combined to obtain a sharded economic data set.

8. A distributed statistical data analysis optimization method according to claim 7, characterized in that: The method of performing task allocation processing on the shard economic data set includes: Defining a node collection and node parameters, where Represents each node in the node set; node parameters include the load limit of any node ;in, Indicates the upper limit of the computational load of any node in the node set; Indicates the upper limit of memory load of any node in the node set; Query the data complexity and memory usage of each task shard in the shard economic data set; calculate the load of each task shard based on the data complexity and memory usage; The load of each task shard ;in, Indicates the data complexity of any task shard; Indicates the memory usage of any task slice; The weight representing the complexity of the data; The weight representing the memory usage; The weight of data complexity and the weight of memory usage are dynamically adjusted. The calculation formula for dynamic adjustment is: ;in, Represents the mean data complexity of all task shards; Indicates the variance of data complexity; Indicates the variance of memory usage; Represents the sum of the upper limits of computing loads of all nodes; Indicates the total computing load that has been used; ;in, Represents the average memory usage of all task slices; Indicates the sum of the memory load limits of all nodes; Indicates the total memory load that has been used; Based on the load of each task slice, all task slices are arranged in descending order from large to small to obtain a sorted slice data set; a node is allocated to each task slice in the sorted slice data set to obtain an allocated node sequence; the allocated node sequence is traversed from the first node to check the load of the task slice in the second node. If the total load obtained by adding the load of the task slice in the second node to the load of the task slice in the first node is less than or equal to the load upper limit of the node, the task slice in the second node is added to the first node, and the second node is released at the same time; otherwise, the total load of the task slice of the third node and the first node is calculated until the load of the first node reaches the load upper limit or cannot accommodate any of the remaining task slices. At this time, traversal starts from the second node and the process of calculating the total load to fill the node is repeated; the resource utilization of each node is calculated. If there is a node whose resource utilization is lower than the preset utilization threshold, the task slice of the node is added to other nodes until all task slices are allocated and the number of nodes containing task slices reaches the minimum. At this time, the optimal node sequence is obtained, which is the optimal allocated data set.

9. A distributed statistical data analysis optimization method according to claim 8, characterized in that: The method of performing data analysis on the optimal allocation data set includes: Feature extraction is performed on the optimal allocation data set to obtain the economic numerical data features and economic time series data features of the optimal allocation data set; a machine learning model is constructed based on the economic numerical data features and the economic time series data features to process the optimal allocation data set, and the average consumption amount prediction results and daily revenue amount prediction results are obtained respectively; the average consumption amount prediction results and the daily revenue amount prediction results are combined to obtain the prediction analysis results.

10. A distributed statistical data analysis and optimization system, used to implement a distributed statistical data analysis and optimization method according to any one of claims 1 to 9, characterized in that: include: A data collection module is used to collect economic data sets, which include economic numerical data and economic time series data sets; A preprocessing module is used to preprocess the economic data set to obtain a preprocessed economic data set; An allocation module, used for sharding the preprocessed economic data set to obtain a sharded economic data set; Perform task allocation processing on the shard economic data set to obtain the optimal allocation data set; An analysis module is used to perform data analysis on the optimal allocation data set to obtain a prediction analysis result; and to merge the optimal allocation data set and the prediction analysis result into a prediction analysis data set; The prediction analysis data set is encrypted using an encryption algorithm to obtain an encrypted data set; the encrypted data set is sent to a preset database for storage; and each module is connected by wire and / or wireless means.

Citation Information

Patent Citations

  • Statistical data analysis and classification system and method

    CN114840641A