Uncertain modeling method for data hybrid exploration

Through the pre-processing, parallel processing and deep learning model construction of massive heterogeneous data, combined with adaptive sampling strategies, the problem of how to accurately identify key areas in the modeling of massive heterogeneous data is solved, efficient and accurate data analysis and model stability are achieved, and data processing efficiency and model adaptability are improved.

CN120296692APending Publication Date: 2025-07-11CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510231575.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In parallel hybrid modeling of massive heterogeneous data, how to accurately identify and focus key areas, dynamically adjust the model to balance flexibility and stability, uniformly measure the degree of influence of heterogeneous data, and efficiently utilize computing resources to avoid resource waste.

Method used

By pre-processing, parallel processing, deep learning model construction, feature importance evaluation and adaptive sampling of massive heterogeneous data, identifying key areas, dynamically adjusting sampling density, and building a hybrid modeling framework to achieve accurate depiction with the greatest impact on modeling tasks.

Benefits of technology

The efficiency of large-scale heterogeneous data processing is improved, the correlation mode and key characteristics are revealed, the accuracy and adaptability of models are improved, and the efficient utilization of resources and the balance of model stability and flexibility is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296692A_ABST
    Figure CN120296692A_ABST
Patent Text Reader

Abstract

The invention provides an uncertainty modeling method for data hybrid exploration, which comprises the following steps of: based on a preliminary data clustering result, analyzing a potential relationship among data items, calculating a support degree and a confidence degree among different features, mining a hidden mode and an association rule in data, and forming a data association network; according to the data association network, constructing a multi-layer neural network model, learning a nonlinear relationship between data features, and obtaining a deep learning model capable of representing an internal structure of the data; performing feature importance evaluation on the data by using a deep learning model, calculating the contribution degree of each feature to a modeling task, sorting the features according to the contribution degrees, and selecting a feature subset with the highest contribution degree as a key region; and performing adaptive sampling on the selected key area, wherein the sampling density is dynamically adjusted according to data distribution characteristics, sampling points are increased in an important area, sampling points are reduced in a secondary area, and a targeted sampling scheme is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an uncertainty modeling method for data mixing exploration. Background Art

[0002] In the parallel hybrid modeling of massive heterogeneous data, accurately identifying and focusing on key areas has become a major challenge in the field of data science. With the rapid expansion of data scale and the increasing diversity of data types, the hidden patterns and complex associations in massive data require the model to automatically discover potential laws, and the dynamic changes in the focus of modeling tasks over time have put forward a severe test for the real-time adjustment of key area selection strategies. In addition, the importance evaluation standards of different types of data vary significantly. How to uniformly measure the impact of heterogeneous data and consider the efficient use of computing resources while ensuring model accuracy to avoid resource waste have constituted technical barriers that need to be overcome. Adjusting key areas too frequently may lead to model instability. How to strike a balance between flexibility and stability is also a thorny issue. These multi-dimensional technical problems are intertwined, forming an intricate challenge. Accurately identifying key areas and effectively focusing on them has become a technical problem that needs to be solved in the field of big data analysis and modeling. Summary of the invention

[0003] The present invention provides an uncertainty modeling method for data mixing exploration, which mainly includes:

[0004] Preprocessing of massive heterogeneous data, including removing noise and outliers, unifying the format and scale of data from different sources through data standardization methods, and classifying and storing them according to data types and features to obtain structured multidimensional data sets;

[0005] Use the distributed computing framework to process structured multidimensional data sets in parallel, perform preliminary data grouping, and identify data with similar characteristics by calculating the similarity and distance between data points to obtain preliminary data clustering results;

[0006] Based on the preliminary data clustering results, analyze the potential relationships between data items, calculate the support and confidence between different features, mine the hidden patterns and association rules in the data, and form a data association network;

[0007] Based on the data association network, a multi-layer neural network model is constructed to learn the nonlinear relationship between data features and obtain a deep learning model that can represent the intrinsic structure of the data;

[0008] Use deep learning models to evaluate the importance of data features, calculate the contribution of each feature to the modeling task, sort the features according to their contribution, and select the feature subset with the highest contribution as the key area;

[0009] Perform adaptive sampling on selected key regions, including dynamically adjusting the sampling density according to the characteristics of data distribution, increasing sampling points in important regions, and reducing sampling points in secondary regions, to generate a targeted sampling scheme;

[0010] Analyze the data in the key regions based on the sampling scheme, capture the temporal variation characteristics of the data, analyze the interactions between variables, construct a targeted hybrid modeling framework, and achieve accurate characterization and modeling of the key regions that have the greatest impact on the modeling task.

[0011] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0012] The present invention discloses an uncertainty modeling method for data hybrid exploration. By ensuring data quality through data preprocessing, optimizing the organization by classified storage, and constructing a parallel data processing and deep learning model, hidden patterns are automatically discovered, and the regions that have the most significant impact on the modeling task are accurately identified and dynamically adjusted. Through feature importance evaluation and combined with an adaptive sampling strategy, this method not only realizes the efficient utilization of resources, but also takes into account flexibility while ensuring the accuracy and stability of the model. Overall technical effects or functions summary: The method of the present invention significantly improves the processing efficiency and analysis depth of large-scale heterogeneous data, can more accurately reveal the association patterns and key features in the data, enhances the processing efficiency and analysis depth of large-scale heterogeneous data, and provides strong technical support for decision-making support and data mining. In addition, through adaptive sampling and targeted hybrid modeling, the accuracy and adaptability of the model can be effectively improved, and it has a wide range of application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of an uncertainty modeling method for data hybrid exploration according to the present invention.

[0014] Figure 2 It is a schematic diagram of an uncertainty modeling method for data hybrid exploration according to the present invention.

[0015] Figure 3 It is another schematic diagram of an uncertainty modeling method for data hybrid exploration according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] To further understand the content of the present invention, the present invention will be described in detail in combination with the accompanying drawings and embodiments. The following further describes the present application in detail with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that for the sake of description, only the parts related to the invention are shown in the drawings.

[0017] AsFigures 1-3 , a method for uncertainty modeling of data hybrid exploration in this embodiment may specifically include:

[0018] S101. Preprocess the massive heterogeneous data, including removing noise and outliers, unifying the formats and scales of data from different sources through data standardization methods, and classifying and storing the data according to data types and characteristics to obtain a structured multi-dimensional data set.

[0019] Receive massive heterogeneous data, label the data according to the diversity and heterogeneity degree of the data sources to obtain a preliminarily classified heterogeneous data set; for the preliminarily classified heterogeneous data set, filter the noise in the data to obtain a cleaned data set; perform standardization processing on the cleaned data set to unify data of different scales into a specified interval to obtain a standardized data set; group the numerical data according to the data types and characteristics of the standardized data set to obtain a classified and stored data set; construct a multi-dimensional data model, define the dimensions and metrics of the multi-dimensional data model, create a fact table and dimension tables, and establish a hierarchical structure to obtain a structured multi-dimensional data set.

[0020] Specifically, collect and preliminarily classify massive heterogeneous data, and mark the data according to the diversity of data sources and the degree of heterogeneity, including structured data, semi-structured data, and unstructured data. For unstructured text data, successively apply the word frequency statistics and keyword extraction methods for preliminary parsing and structuring. Use distributed storage to store different types of original data into corresponding databases respectively, and establish the association relationship between data through database indexing to obtain a preliminarily classified heterogeneous data set. Perform data cleaning on the preliminarily classified heterogeneous data set, and filter out the noise in the data through the wavelet transform method. Select an appropriate wavelet basis function, calculate the wavelet coefficients, set a threshold for coefficient screening, and reconstruct the signal through inverse transformation. Use statistical methods to identify outliers, and use the median replacement method to process outliers. Use the linear interpolation method to fill in the missing data, and use regularization to smooth the abnormal data to obtain a cleaned data set. Standardize the cleaned data set, and use the minimum-maximum normalization to unify data of different scales into a specified interval. Eliminate the dimensional difference between different data sources through z-score standardization, and the calculation formula is (x - μ) / σ, where x is the original value, μ is the mean, and σ is the standard deviation. Use logarithmic transformation to process skewed distribution data, and use one-hot encoding to perform encoding conversion on discrete variables to obtain a standardized data set with a unified format. Classify and store according to the data type and characteristics of the standardized data set, and use K-means clustering to automatically group numerical data. Perform feature selection on categorical data through chi-square test, and use principal component analysis to reduce the dimension of high-dimensional data. Use a data cube to construct a multi-dimensional data model, define dimensions and measures, create fact tables and dimension tables, and establish a hierarchical structure. Thus, a structured multi-dimensional data set is obtained, realizing the preprocessing, cleaning, standardization, and structured storage of massive heterogeneous data. When collecting and preliminarily classifying massive heterogeneous data, data source identifiers and heterogeneity degree scores can be set. For example, structured data such as database tables can be marked as S1 with a score of 0.8; semi-structured data such as JSON files are marked as S2 with a score of 0.5; unstructured data such as text documents are marked as S3 with a score of 0.2. For S3 type text data, calculate the word frequency through the TF-IDF algorithm, extract keywords, and use the vocabulary with a word frequency greater than 0.01 as features. Use the Hadoop distributed file system to store different types of data, and use HBase to establish an index to achieve fast retrieval and association. During the data cleaning process, select the Daubechies wavelet function for noise filtering. Set the number of wavelet decomposition layers to 3 and the soft threshold to 0.05. Detect outliers in the reconstructed data, calculate the interquartile range IQR, and consider the data outside the range of Q1 - 1.5IQR or Q3 + 1.5IQR as outliers. Use the median replacement method to process outliers. For example, if the median of a certain numerical column is 100, then replace the outliers with 100.For missing data, the linear interpolation algorithm is used for filling. For example, if the two adjacent valid values are 80 and 120 and there is one missing data point in between, it is filled as 100. The L2 regularization method is used to smooth the abnormal data, and the regularization coefficient is set to 0.01. When performing standardization processing, the min-max normalization is used to map the data to the

[0021] [0, 1] interval. For example, if the original data range is [-10, 50], then x' = (x + 10) / 60. In z-score standardization, assuming the mean μ of a certain feature is 10 and the standard deviation σ is 2, the original value 15 becomes 2.5 after standardization. For skewed distribution data, if the skewness of a certain feature is greater than 1, the natural logarithm transformation is applied. When encoding discrete variables, for example, if the gender feature has two values, "male" and "female", they are encoded as [1, 0] and [0, 1]. In the classification storage stage, the K-means algorithm is used to cluster the numerical data, and the value of K is set to 5. When performing feature selection on categorical data, the chi-square statistic threshold is set to 3.84 (corresponding to a significance level of 0.05). When performing principal component analysis for dimensionality reduction, the principal components with a cumulative contribution rate reaching 85% are selected. When constructing the data cube, dimensions such as time, geographical location, and product category are defined, and sales amount, profit margin, etc. are used as measurement indicators. A sales fact table and dimension tables are created, and a hierarchical structure of "country - province - city" is established on the geographical location dimension. Through these processes, a structured multi-dimensional dataset is finally obtained.

[0022] S102. Use the distributed computing framework to perform parallel processing on the structured multi-dimensional dataset, perform preliminary grouping on the data, and identify the data with similar features by calculating the similarity and distance between data points to obtain a preliminary data clustering result.

[0023] Obtain the structured multi-dimensional dataset in the distributed file system, and perform preliminary data grouping according to the dataset using the MapReduce computing framework; for the preliminary grouping result, calculate the comprehensive similarity matrix between data points, and the comprehensive similarity matrix is obtained by calculating the Euclidean distance of numerical features, the cosine similarity of text features, and the Jaccard coefficient of set features; according to the comprehensive similarity matrix, use the K-means clustering algorithm to perform preliminary clustering on the data, and the preliminary clustering is implemented through locality-sensitive hashing; evaluate the preliminary clustering result, use the silhouette coefficient to evaluate the clustering quality, if the silhouette coefficient is less than the preset threshold, increase the number of clusters and re-perform K-means clustering; perform dimensionality reduction processing on the data through the principal component analysis method to obtain a dimensionality-reduced dataset; re-apply the K-means algorithm on the dimensionality-reduced dataset to obtain the final data clustering result.

[0024] Specifically, obtain a structured multi-dimensional dataset from a distributed file system and use the MapReduce framework to perform preliminary grouping on the data. In the Map stage, calculate the hash value according to the data characteristics and allocate similar data to the same computing node. In the Reduce stage, merge the data with the same hash value. Dynamically adjust the grouping method according to the data scale and the number of computing nodes, and set the number of hash buckets to 3 times the number of nodes to ensure load balancing and obtain the preliminary grouping result. For the preliminary grouping result, calculate the similarity and distance between data points. Use the Euclidean distance formula to calculate the distance of numerical features, which is calculated after numerical normalization. For text features, first use TF-IDF to convert the text into a vector representation, and then calculate the similarity between vectors through cosine similarity. Use the Jaccard coefficient to calculate the similarity of set features. Combine multiple distance measurement methods and use the weighted average method to obtain a comprehensive similarity matrix, and the weights are preset according to the feature importance. Based on the comprehensive similarity matrix, use the K-means clustering algorithm to perform preliminary clustering on the data. Improve the clustering efficiency through parallel processing, and use the polling method to allocate data points to different computing nodes. Use the MinHash algorithm to implement locality-sensitive hashing and reduce the search space. Calculate the distance from the data point to the cluster center in parallel on each node, and synchronously update the cluster center through the message passing interface. Set the maximum number of iterations to 100, or stop the iteration when the change in the position of the cluster center is less than 0.001. Obtain the preliminary clustering result. Evaluate and optimize the preliminary clustering result, and use the silhouette coefficient to evaluate the clustering quality. Calculate the silhouette value of each data point and take the average as the overall clustering quality index. If the silhouette coefficient is less than 0.5, increase the number of clusters and re-perform K-means clustering. Reduce the complexity of high-dimensional data through the principal component analysis method, and select the principal components with a cumulative contribution rate reaching 95%. Perform eigenvalue decomposition on the selected principal components to obtain the reduced-dimensional dataset. Reapply the K-means algorithm on the reduced-dimensional dataset to obtain the final data clustering result. When processing a large amount of multi-dimensional datasets, first read the data from the distributed file system. Assume that the data scale is 10TB and contains 1000 feature dimensions. Use MapReduce for preliminary grouping and set 100 computing nodes. In the Map stage, use the MurmurHash algorithm to calculate the hash value of each piece of data, and set the number of hash buckets to 300. For example, for the feature vector [0.5, 0.3, 0.8], the calculated hash value is 7823, which is allocated to bucket 23 (7823 % 300 = 23). In the Reduce stage, merge the data within the same hash bucket to obtain the preliminary grouping result. Calculate the similarity and distance between data points for the grouped result. The Euclidean distance is used for numerical features. For example, the distance between two points [1, 2, 3] and [4, 5, 6] is 5.196.Text features are converted into vectors using TF-IDF. For example, "data mining" is converted into [0.4, 0.6, 0], and "machine learning" is converted into [0, 0.5, 0.5]. The cosine similarity is calculated to be 0.424. The set features use the Jaccard coefficient. For example, the similarity between {a, b, c} and {b, c, d} is 0.5. The comprehensive similarity uses weighted average with weights set as [0.5, 0.3, 0.2] to obtain the final similarity matrix. The K-means algorithm is used for preliminary clustering with the K value set to 50. The data is distributed to 100 computing nodes in a polling manner. The MinHash algorithm is used to implement locality-sensitive hashing, and 100 hash functions are selected. The distances are calculated in parallel on each node. For example, node 1 is responsible for calculating the distances from the first 1000 data points to the cluster centers. The cluster centers are updated synchronously through MPI (Message Passing Interface). The maximum number of iterations is set to 100, and the process stops when the change in the positions of the cluster centers is less than 0.001. The preliminary clustering result is obtained, including 50 cluster centers and the cluster labels of each data point. The clustering result is evaluated and optimized. The silhouette coefficient is calculated. For example, for the data point [1, 2] in cluster A, the nearest other cluster is B. The cohesion degree a = 0.8, the separation degree b = 1.5, and the silhouette value s = (b - a) / max(a, b) = 0.467. The average value of all points is taken as the overall silhouette coefficient. If the silhouette coefficient is less than 0.5, the number of clusters is increased to 60 and re-clustering is performed. Dimension reduction is carried out through principal component analysis, and the first 100 principal components are selected with the cumulative contribution rate reaching 95%. The eigenvalue decomposition is performed on these 100 principal components to obtain the dimension-reduced dataset. K-means is reapplied to the dimension-reduced data to obtain the final clustering result, including 60 cluster centers and the cluster labels of each data point.

[0025] S103. Based on the preliminary data clustering result, analyze the potential relationships between data items, calculate the support and confidence between different features, mine the hidden patterns and association rules in the data, and form a data association network.

[0026] Obtain the data items of each cluster according to the clustering result, use the Apriori algorithm to perform frequent item set mining on the data items to obtain a list of frequent item sets. Calculate the conditional probability between item sets for the list of frequent item sets. If the conditional probability is greater than the preset minimum confidence threshold, generate preliminary association rules. Calculate the lift for the preliminary association rules. If the lift is greater than 1, retain the association rule to obtain the optimized association rules. Construct a data association network according to the optimized association rules, where the item sets in the association rules are used as network nodes and the association rules are used as the edges between nodes. Layout the data association network to obtain a visual data association network diagram.

[0027] Specifically, according to the preliminary data clustering results, frequent itemset mining is performed separately for each cluster. The Apriori algorithm is used to scan the data items within each cluster, and the minimum support threshold is set to 0.05. For the clustering results, the weights of the items are adjusted. The weight of the items near the cluster center is set to 1.2, and the weight of the marginal items is set to 0.8. The occurrence frequencies of the item sets that meet the weighted minimum support are counted to obtain the frequent itemset list. For the obtained frequent itemset list, the conditional probabilities between the item sets are calculated. The confidence formula is P(B|A) = P(A∩B) / P(A), where P(A∩B) is the probability that A and B appear together, and P(A) is the probability that A appears. The minimum confidence threshold is set to 0.6, and the item set pairs that meet the minimum confidence are screened out to generate the preliminary association rules. The preliminary association rules are optimized, and the lift of the rules is calculated. The lift formula is Lift(A→B) = P(B|A) / P(B), which represents the ratio of the occurrence probability of B under the condition of containing A to the general occurrence probability of B. The rules with a lift greater than 1 are retained, indicating positive correlation. The remaining rules are sorted according to the weighted scores of support, confidence, and lift, and the weights are set to 0.3, 0.4, and 0.3 respectively. The top 1000 rules with the highest scores are selected as the final association rules. Based on the final association rules, a data association network is constructed. The item sets in the rules are used as network nodes, and the rules are used as the edges between the nodes. The weight of the edge is set to the weighted score of the rule. The Fruchterman-Reingold algorithm is used to layout the network. For large-scale networks, first use hierarchical clustering to divide the nodes into 50 groups, and then display them layer by layer. Through the interactive zoom function, multi-level browsing of the network is achieved, and a visualized data association network diagram is obtained. When processing the preliminary data clustering results, assume there are 10 clusters, and each cluster contains 1000 data items. The Apriori algorithm is used to perform frequent itemset mining for each cluster, and the minimum support threshold is set to 0.05. For example, in cluster 1, the item set {A,B} appears 60 times, and the support is 0.06, which exceeds the threshold and is retained. The weight of the items within the radius of the cluster center is set to 1.2. For example, item C is 0.3 away from the center, and the weight is 1.2; the marginal item D is 0.8 away from the center, and the weight is 0.8. After weighting, the support of the item set {C,D} is increased from 0.049 to 0.0564, meeting the threshold requirements. Finally, a frequent itemset list containing 5000 item sets is obtained. When calculating the conditional probabilities between the item sets, assume that the item set {A,B} appears 100 times in the dataset, and item set A appears alone 150 times. Then P(B|A) = 100 / 150 = 0.667. The minimum confidence threshold 0.6 is set, and the confidence of {A,B} is 0.667, which is greater than the threshold and is retained as an association rule. For the item set {E,F}, E appears 200 times, and {E,F} co-occurs 80 times. The confidence is 0.4, which is lower than the threshold and is excluded. After screening, 3000 preliminary association rules are obtained.When optimizing association rules, the lift is calculated. If P(F) = 0.1 and P(F|E) = 0.4, then Lift(E→F) = 0.4 / 0.1 = 4, indicating a positive correlation, and this rule is retained. For the rule G→H, if P(H) = 0.5 and P(H|G) = 0.45, then Lift(G→H) = 0.45 / 0.5 = 0.9, indicating a negative correlation, and this rule is removed. The remaining rules are sorted by weighted scores. For example, for the rule I→J with a support of 0.1, a confidence of 0.7, and a lift of 2.5, the score is 0.1×0.3 + 0.7×0.4 + 2.5×0.3 = 0.88. The top 1000 rules with the highest scores are selected as the final rules. When constructing a data association network, the 1000 rules form approximately 500 nodes and 1000 edges. The Fruchterman-Reingold algorithm is used for layout, setting the number of iterations to 1000 and the temperature parameter to decrease from 1 to 0.1. For large-scale networks, the hierarchical clustering algorithm is first used to divide the nodes into 50 groups, setting the clustering threshold to 0.6. In the visualization interface, the 50 group nodes are initially displayed, and clicking on a group node can expand to view the internal structure. Multilevel browsing is achieved through the scroll wheel zoom function, with the zoom ratio ranging from 0.1 to 10. The finally generated interactive data association network diagram supports dynamic adjustment and in-depth exploration.

[0028] S104. According to the data association network, construct a multi-layer neural network model to learn the non-linear relationship between data features and obtain a deep learning model that can represent the internal structure of the data.

[0029] According to the structure of the data association network, obtain the topological structure of the graph convolutional network; wherein, the graph convolutional network includes three graph convolutional layers; preprocess the data set; wherein, the preprocessing includes: performing min-max scaling on numerical features to map the data into the interval [0, 1]; performing Z-score standardization on continuous features to make the data mean 0 and the standard deviation 1; performing one-hot encoding transformation on categorical features. Initialize the parameters of the graph convolutional network, including: initializing the weights of the ReLU activation function layer using the He initialization method; initializing the weights of the tanh activation function layer using the Xavier initialization method; initializing the bias to 0. Start the training of the graph convolutional network, including: setting the loss function to mean squared error; using the Adam optimization algorithm and setting the learning rate; evaluate the graph convolutional network, including: calculating the accuracy, precision, recall, and F1 score using the test set; using 5-fold cross-validation to evaluate the stability of the model performance; if the validation set accuracy is higher than the preset threshold or the maximum number of iterations is reached, stop the training.

[0030] Specifically, according to the structure of the data association network, obtain the topological structure of the graph convolutional network. Take the nodes in the association network as features and the weights of the edges as connection strengths. Set three graph convolutional layers, and the output dimensions of each layer are 128, 64, and 32 respectively. After the graph convolutional layers, add a fully connected layer with 16 neurons. The number of neurons in the output layer is the same as the dimension of the target variable. Select the ReLU function as the activation function to introduce non-linearity. Preprocess the dataset, including data normalization and standardization. Perform min-max scaling on numerical features to map the data into the interval [0,1]. Perform Z-score standardization on continuous features to make the data have a mean of 0 and a standard deviation of 1. Perform one-hot encoding transformation on categorical features. Divide the training set and the validation set with a ratio of 8:2, and use the stratified sampling method to ensure the consistency of the data distribution. Initialize the network parameters. For the layers with the ReLU activation function, use the He initialization method for the weights; for the layers with the tanh activation function, use the Xavier initialization method. Initialize the biases to 0. Set the loss function as the mean squared error, select the Adam optimization algorithm, and set the learning rate to 0.001. Set the batch size to 64 and the maximum number of iterations to 1000 epochs. Adopt an early stopping strategy during training, and stop training if the validation set loss does not decrease for 10 consecutive epochs. Use a learning rate decay strategy to reduce the learning rate to 0.9 times the original value every 100 epochs. Start the model training, evaluate the model performance on the validation set every 50 epochs, and record the loss value and accuracy. Use 5-fold cross-validation to evaluate the stability of the model performance. Optimize the hyperparameters, including the learning rate, batch size, and number of network layers, using grid search. Stop training if the validation set accuracy is higher than the preset threshold of 0.95 or the maximum number of iterations is reached. Save the parameters of the best-performing model to obtain a deep learning model representing the internal structure of the data. Evaluate the final model using the test set and calculate the accuracy, precision, recall, and F1 score. When constructing the graph convolutional network, assume that the data association network contains 1000 nodes and 5000 edges. The first graph convolution converts the 1000-dimensional input features into 128 dimensions, the second into 64 dimensions, and the third into 32 dimensions. Each graph convolution operation uses a normalized version of the adjacency matrix to ensure that information is not amplified or reduced during propagation. The fully connected layer compresses the 32-dimensional features to 16 dimensions, and the final output layer dimension is set to 10, corresponding to 10 target variables. The ReLU activation function is applied after each layer to introduce non-linearity. In the data preprocessing stage, process 1 million data records. The minimum value of the numerical feature A is -50 and the maximum value is 100, and it is mapped into the interval [0,1] through min-max scaling. The mean of the continuous feature B is 10 and the standard deviation is 5. After Z-score standardization, the mean becomes 0 and the standard deviation becomes 1. The categorical feature C has 5 possible values and is converted into a 5-dimensional binary vector through one-hot encoding.Using the stratified sampling method to maintain the proportion of each category, 800,000 pieces of data are divided into the training set, and 200,000 pieces are divided into the validation set. When initializing the network parameters, for the layers with the ReLU activation function, the weight initialization adopts the He initialization, multiplied by sqrt(2 / n_in), where n_in is the input dimension. The initial learning rate of the Adam optimizer is set to 0.001, β1 = 0.9, and β2 = 0.999. The batch size is set to 64, that is, 64 pieces of data are used for each training. The maximum number of iterations is 1000 rounds, and the learning rate is multiplied by 0.9 every 100 rounds. The early stopping strategy sets the threshold to 10 rounds, that is, if the validation set loss does not decrease for 10 consecutive rounds, the training stops. During the model training process, the performance is evaluated on 200,000 pieces of validation data every 50 rounds. Using 5-fold cross-validation, the training data is divided into 5 parts, 4 parts are used for training each time, and 1 part is used for validation, repeating 5 times. Hyperparameters are optimized by grid search, the learning rate range is [0.0001, 0.001, 0.01], the batch size is [32, 64, 128], and the number of network layers is [3, 4, 5]. Stop training when the validation set accuracy reaches 0.95 or after 1000 rounds of iteration. The final model is evaluated on 100,000 pieces of test data, and the accuracy, precision, recall, and F1 score are calculated to comprehensively evaluate the model performance.

[0031] S105. Use the deep learning model to evaluate the feature importance of the data, calculate the contribution of each feature to the modeling task, sort the features according to the contribution size, and select the feature subset with the highest contribution as the key area.

[0032] Obtain the partial derivative of the loss function with respect to the feature, and the partial derivative is calculated for each sample; accumulate the partial derivatives of all samples according to the partial derivative and take the average to obtain the feature importance score. Use the softmax function to process the feature importance score; convert the feature importance score into a contribution percentage through the softmax function, the contribution percentage is between 0 and 1, and the sum of the contributions of all features is 1. Sort the contribution percentages in descending order; set the contribution threshold, and filter out the feature subset with a contribution greater than the contribution threshold to obtain the preliminary key area. Perform 5-fold cross-validation, and the 5-fold cross-validation is used to evaluate the performance of the model using only the preliminary key area; compare the model performance with the model performance using all features, and if the performance degradation does not exceed the preset threshold, retain the feature subset. Perform correlation analysis on the features in the preliminary key area; calculate the Pearson correlation coefficient matrix between features, and the Pearson correlation coefficient matrix is used to represent the correlation degree between features; set the correlation threshold, and perform hierarchical clustering on the features with a correlation coefficient higher than the correlation threshold; select the feature with the highest contribution from each cluster as the representative to obtain the optimized key area feature subset.

[0033] Specifically, according to the structure of the deep learning model, the importance of each feature is calculated for a specific modeling task. The partial derivative of the loss function with respect to the feature is calculated for each sample, and the partial derivatives of all samples are accumulated and averaged. The specific process is to calculate the loss through forward propagation, calculate the gradient through backward propagation, take the absolute value of the gradient for each feature and accumulate it, and finally divide by the number of samples to obtain the importance score of each feature. Based on the feature importance score, the softmax function is used to convert the score into a contribution percentage.

[0034] The softmax function converts the scores of each feature into values between 0 and 1, and the sum of the contribution degrees of all features is 1. The calculation formula is exp(score_i) / sum(exp(score)), where i represents the i-th feature and sum represents the exponential sum of all feature scores. Sort the feature contribution degrees in descending order, set a contribution degree threshold, and filter out the feature subset with contribution degrees greater than the threshold. If the number of features in the subset is less than the preset minimum number of features, select the top N features with the highest contribution degrees as the preliminary key region. Use 5-fold cross-validation to evaluate the performance of the model using only the selected features and compare it with the performance of the model using all features. If the performance degradation does not exceed the preset threshold, retain the feature subset. Conduct a correlation analysis on the features in the preliminary key region, calculate the Pearson correlation coefficient matrix between the features. Set a correlation threshold, and perform hierarchical clustering on the features with correlation coefficients higher than the threshold. Select the feature with the highest contribution degree from each cluster as the representative. Merge the highly correlated features and retain the representative features with higher contribution degrees to finally obtain the optimized key region feature subset. During the feature importance evaluation process, assume that a deep learning model has 100 input features for predicting the sales volume of a certain product. Process 10,000 samples. For each sample, calculate the loss through the forward propagation of the model and then calculate the gradient through the backward propagation. For feature A, the absolute value of the cumulative gradient on all samples is 500, and the importance score is 0.05 after averaging. Repeat this process for all features to obtain the importance scores of 100 features. Use the softmax function to convert the feature importance scores into contribution percentages. For example, the score of feature A is 0.05, and the contribution degree after softmax conversion is 2.1%. The sum of the contribution degrees of all features is strictly equal to 100%. Sort the contribution degrees of 100 features in descending order, set the contribution degree threshold to 1%, and filter out the feature subset with contribution degrees greater than 1%, obtaining 25 features. Since 25 is greater than the preset minimum number of features 20, these 25 features form the preliminary key region. Use 5-fold cross-validation to evaluate the model performance. The average prediction accuracy of the full-feature model is 85%, and the accuracy of the model using only 25 selected features is 83%. The performance degradation is 2.35%, which is less than the preset threshold of 3%, so these 25 features are retained. Calculate the Pearson correlation coefficient matrix for these 25 features, and set the correlation threshold to 0.8. It is found that the correlation coefficient between feature B and feature C is 0.85, exceeding the threshold. Perform hierarchical clustering on the feature pairs with correlation coefficients higher than 0.8 to form 5 feature groups. In the feature group composed of B and C, the contribution degree of B is 1.8% and the contribution degree of C is 1.5%, and B is selected as the representative feature. Process other highly correlated feature groups similarly. Finally, the 25 initial features are compressed into 20 representative features, forming the optimized key region feature subset. These 20 features capture approximately 90% of the information in the original 100 features, significantly reducing the model complexity while maintaining a high prediction performance.

[0035] S106. Perform adaptive sampling on the selected key regions, including dynamically adjusting the sampling density according to the data distribution characteristics, increasing the sampling points in important regions and decreasing the sampling points in secondary regions, and generating a targeted sampling scheme.

[0036] Obtain the probability density function of each feature within the selected key regions to obtain the data distribution analysis result. Calculate the importance weights of the sampling points according to the data distribution analysis result, where the importance weights are mapped from the probability density function. Set the sampling density parameters for important regions and secondary regions. If the sampling point is located in an important region, a high sampling density parameter is used; if the sampling point is located in a secondary region, a low sampling density parameter is used. Generate candidate sampling points that meet the minimum distance threshold requirement based on the sampling density parameters and the preset spatial constraint conditions; judge whether it is necessary to adjust the definition threshold or the sampling density parameters of the important region and the secondary region according to the distribution of the sampled points. If adjustment is needed, update the definition threshold or the sampling density parameters; if no adjustment is needed, continue the sampling process until the preset stop condition is met.

[0037] Specifically, perform data distribution analysis on the selected key regions, and use the kernel density estimation algorithm to calculate the probability density function of each feature. Determine the appropriate grid division granularity through cross-validation methods. Try different granularity values and select the granularity that performs best on the validation set. Generate uniformly distributed sampling points in the feature space and perform preliminary scoring on the sampling points according to the probability density function. Based on the preliminary scoring results, calculate the importance weights of the sampling points. Use the sigmoid function to map the probability density values to the interval (0,1), and the formula is 1 / (1+exp(-x)), where x is the probability density value. Set the weight threshold, mark the sampling points above the threshold as important regions, and mark the sampling points below the threshold as secondary regions. The weight threshold is initially set to 0.5 and is dynamically adjusted according to the sampling results. Set the sampling density parameters for the important regions and secondary regions respectively. Use a higher sampling density for the important regions and a lower sampling density for the secondary regions. Combine the spatial distance constraints and use a KD tree to store the selected sampling points. When adding a new sampling point, check whether the distance between it and the nearest neighbor point is greater than the preset threshold to ensure that the minimum distance between sampling points is not less than this threshold and avoid over-concentration of sampling points. Generate candidate sampling points using the Monte Carlo method according to the sampling density parameters and spatial constraint conditions. Use the rejection sampling algorithm to screen the sampling points that meet the conditions, and the acceptance probability is proportional to the importance weight of the sampling point. Dynamically adjust the number of sampling points until the preset total number of sampling points or sampling coverage rate is reached. Dynamically adjust the definition threshold of the important region and the secondary region, or adjust the sampling density parameters according to the distribution of the sampled points. Repeat this process until the stop condition is met, and finally form a targeted adaptive sampling scheme. When performing adaptive sampling on the key regions of a certain product sales data, first perform kernel density estimation on 10 key features. Use the Gaussian kernel function, and the bandwidth parameter is determined to be 0.15 through cross-validation. The grid division granularity is cross-validated by 5 folds and try

[0038] [10, 20, 50, 100] are equal. Finally, 50 is selected as the optimal granularity. 50^10 uniformly distributed sampling points are generated in the 10-dimensional feature space, and the probability density value of each point is calculated to obtain a preliminary score. The sigmoid function is applied to the preliminary score result for normalization, mapping the probability density value to the interval (0, 1). The original probability density value of 0.8 becomes 0.69 after being transformed by the sigmoid function. The initial weight threshold is set to 0.5. Sampling points higher than 0.5 are marked as important regions, and those lower than 0.5 are marked as secondary regions. Among 100,000 sampling points, approximately 30,000 are marked as important regions. The sampling density parameter for the important region is set to 0.01, and that for the secondary region is set to 0.001. The KD-tree data structure is used to store the selected sampling points, and the minimum distance threshold is set to 0.05. When adding a new sampling point, check its distance from the nearest neighbor point in the KD-tree. If it is less than 0.05, reject the point. The distance between the new sampling point (0.2, 0.3, 0.4) and the nearest neighbor point is 0.03, which is less than the threshold, so it is rejected. The Monte Carlo method is used to generate 100,000 candidate sampling points. For each candidate point, calculate its importance weight, such as 0.75. Generate a uniformly random number in the range [0, 1]. If it is less than 0.75, accept the point. The target number of sampling points is set to 5000. During the sampling process, the weight threshold is dynamically adjusted. When the number of sampling points in the important region reaches 4000, the threshold is increased to 0.6 to increase the sampling of the secondary region. Finally, an adaptive sampling scheme containing 5000 sampling points is generated, where approximately 70% are distributed in the important region and 30% are distributed in the secondary region, realizing the dynamic adjustment of the sampling density.

[0039] S107. Analyze the data in the key region based on the sampling scheme, capture the time-varying characteristics of the data, analyze the interaction between variables, construct a targeted hybrid modeling framework, and achieve the accurate characterization and modeling of the key region that has the greatest impact on the modeling task.

[0040] Obtain the sliding window features from the time series data, calculate the autocorrelation coefficient and partial autocorrelation coefficient through the sliding window features, and judge the stationarity and periodicity of the time series data. If the time series data is non-stationary, perform differencing processing on the time series data until the time series data becomes stationary. Use the principal component analysis method to perform dimensionality reduction processing on the time series features to obtain the dimensionality-reduced feature set. Calculate the Pearson correlation coefficient matrix and mutual information matrix between the features in the dimensionality-reduced feature set, and construct a feature interaction network. Construct a hybrid modeling framework, including prior probability, constraint conditions, and a custom loss function. Perform seasonal decomposition on the periodic features in the feature interaction network to obtain the trend component, seasonal component, and random component. Model the data in the key region based on the hybrid modeling framework, and set a rolling prediction window using the time series cross-validation method.

[0041] Specifically, time series analysis is performed on the data of key regions according to the sampling scheme, and the sliding window method is used to extract time series features. Multiple window sizes are set, such as 1 day, 7 days, and 30 days, and the optimal window size is selected through 5-fold cross-validation. The autocorrelation coefficient and partial autocorrelation coefficient of each feature are calculated to judge the stationarity and periodicity of the data. Differencing is performed on non-stationary sequences until they become stationary. Trend, seasonal, and periodic features are extracted to obtain a description of the time-varying features. Multivariate statistical analysis is performed on the extracted time-varying features and original features, and principal component analysis is used to reduce the feature dimension, retaining the principal components with a cumulative explained variance ratio reaching 95%. The Pearson correlation coefficient matrix and mutual information matrix between features are calculated to capture linear and nonlinear feature relationships. The maximum relevance minimum redundancy algorithm is used to identify significantly correlated feature pairs and construct a feature interaction network. Community detection is performed on the network to identify closely related feature clusters. A targeted hybrid modeling framework is constructed in combination with expert knowledge. Expert knowledge is incorporated into the model by setting prior probabilities, constraints, or custom loss functions. Seasonal decomposition is used for features with strong periodicity to separate out the trend, season, and random components. Polynomial regression is used for nonlinear relationships, and the order is determined through cross-validation. Random forest is used for complex interactions, and the number and depth of trees are dynamically adjusted according to the data scale. Time series prediction is combined with feature regression to construct an overall framework. Based on the constructed hybrid modeling framework, the data of key regions is modeled. The time series cross-validation method is used to evaluate the model performance, and a rolling prediction window is set. The Shapley value is used to calculate the contribution of each sub-model to the final prediction. The sub-models are weighted and fused according to the contribution, and the weights are optimized by gradient descent. Residual analysis is performed on the fused model to identify regions with large prediction errors and conduct targeted optimization. Finally, a hybrid model that accurately depicts the key regions is obtained. When analyzing the sales data of a large retail chain enterprise, time series processing is first performed on the data of key regions. Three sliding windows of 1 day, 7 days, and 30 days are set, and through 5-fold cross-validation, it is found that the 7-day window has the best effect. The autocorrelation coefficient of the sales volume feature is calculated, and a peak value of 0.85 is reached at a lag of 7, indicating an obvious weekly periodicity. First-order differencing is performed on the non-stationary price sequence to obtain a stationary sequence. Trend, seasonal, and periodic features are extracted, and it is found that the sales volume shows a seasonal peak of 250% in December every year. Multivariate analysis is performed on the time features and original features, and the results of principal component analysis show that the cumulative explained variance of the first 10 principal components reaches 96.7%. The Pearson correlation coefficient and mutual information are calculated, and it is found that the correlation coefficient between price and sales volume is -0.6, and the mutual information value is 0.4, indicating a strong nonlinear relationship. The maximum relevance minimum redundancy algorithm is used to screen out 20 pairs of significantly correlated feature pairs and construct a feature interaction network. Through the Louvain algorithm for community detection, 3 closely related feature clusters are identified, corresponding to product attributes, promotional activities, and customer behavior respectively.Construct a hybrid modeling framework by integrating expert knowledge. Set the prior probability of sales volume increase due to promotional activities to 0.8 and use it as a model constraint condition. Perform seasonal decomposition on the weekly sales volume series and extract an annual growth trend of 2.5%. Use a third-order polynomial regression for the non-linear relationship between price and sales volume, and the cross-validation mean squared error is 156. Use a random forest for complex interactions with 20 features, set 500 trees, and the maximum depth is 8. Combine the time series prediction of the ARIMA model with feature regression to construct the overall framework. Use a 60-day rolling window time series cross-validation to evaluate the model performance. Calculate the Shapley value and find that the contribution of the ARIMA sub-model is 0.4, polynomial regression is 0.3, and random forest is 0.3. Optimize the sub-model weights through gradient descent to obtain the optimal weights of 0.35, 0.35, and 0.3. Conduct a residual analysis on the fusion model and find that the prediction error is relatively large during the promotion period, and the mean absolute error is 15% of the daily average sales volume. Optimize the prediction during the promotion period specifically to reduce the error to 10%. Finally, the R value of the hybrid model on the test set reaches 0.92, achieving an accurate characterization and modeling of the key sales areas.

[0042] For those skilled in the art, it is obvious that this application is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of this application, this application can be implemented in other specific forms. Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in this application. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. An uncertainty modeling method for data hybrid exploration, characterized in that The method includes: Preprocessing massive heterogeneous data, including removing noise and outliers, unifying the formats and scales of data from different sources through data standardization methods, and classifying and storing the data according to data types and features to obtain a structured multi-dimensional dataset; Using a distributed computing framework to perform parallel processing on the structured multi-dimensional dataset, initially grouping the data, and identifying data with similar features by calculating the similarity and distance between data points to obtain a preliminary data clustering result; Based on the preliminary data clustering result, analyzing the potential relationships between data items, calculating the support and confidence between different features, mining the hidden patterns and association rules in the data, and forming a data association network; According to the data association network, constructing a multi-layer neural network model, learning the non-linear relationships between data features, and obtaining a deep learning model that can represent the internal structure of the data; Using the deep learning model to evaluate the importance of data features, calculating the contribution of each feature to the modeling task, sorting the features according to the contribution size, and selecting the feature subset with the highest contribution as the key region; Performing adaptive sampling on the selected key region, including dynamically adjusting the sampling density according to the data distribution characteristics, increasing the sampling points in important regions, and reducing the sampling points in secondary regions to generate a targeted sampling scheme; Based on the sampling scheme, analyzing the data in the key region, capturing the time-varying characteristics of the data, analyzing the interaction between variables, constructing a targeted hybrid modeling framework, and achieving accurate characterization and modeling of the key region that has the greatest impact on the modeling task.

2. The method according to claim 1, wherein The preprocessing of the massive heterogeneous data, including removing noise and outliers, unifying the formats and scales of data from different sources through data standardization methods, and classifying and storing the data according to data types and features to obtain a structured multi-dimensional dataset, includes: Receiving massive heterogeneous data, marking the data according to the diversity of data sources and the degree of heterogeneity to obtain a preliminarily classified heterogeneous dataset; Filtering the noise in the data for the preliminarily classified heterogeneous dataset to obtain a cleaned dataset; Performing standardization processing on the cleaned dataset to unify data of different scales into a specified interval to obtain a standardized dataset; Grouping the numerical data according to the data types and features of the standardized dataset to obtain a classified and stored dataset; Constructing a multi-dimensional data model, defining the dimensions and metrics of the multi-dimensional data model, creating fact tables and dimension tables, and establishing a hierarchical structure to obtain a structured multi-dimensional dataset.

3. The method according to claim 1, wherein, The use of a distributed computing framework to perform parallel processing on the structured multi-dimensional dataset, initially grouping the data, and identifying data with similar features by calculating the similarity and distance between data points to obtain a preliminary data clustering result, includes: Obtaining the structured multi-dimensional dataset in the distributed file system and performing preliminary data grouping according to the dataset using the MapReduce computing framework; For the preliminary grouping results, calculate the comprehensive similarity matrix between data points, where the comprehensive similarity matrix is obtained by calculating the Euclidean distance of numerical features, the cosine similarity of text features, and the Jaccard coefficient of set features; According to the comprehensive similarity matrix, use the K-means clustering algorithm to perform preliminary clustering on the data, where the preliminary clustering is achieved through locality-sensitive hashing; Evaluate the preliminary clustering results, use the silhouette coefficient to evaluate the clustering quality. If the silhouette coefficient is less than the preset threshold, increase the number of clusters and re-perform K-means clustering; Perform dimensionality reduction on the data through the principal component analysis method to obtain a dimensionality-reduced data set; Re-apply the K-means algorithm on the dimensionality-reduced data set to obtain the final data clustering results.

4. The method according to claim 1, wherein Based on the preliminary data clustering results, analyze the potential relationships between data items, calculate the support and confidence between different features, mine the hidden patterns and association rules in the data, and form a data association network, including: Obtain the data items of each cluster according to the clustering results, use the Apriori algorithm to perform frequent item set mining on the data items to obtain a list of frequent item sets; Calculate the conditional probability between item sets for the list of frequent item sets. If the conditional probability is greater than the preset minimum confidence threshold, generate preliminary association rules; Calculate the lift for the preliminary association rules. If the lift is greater than 1, retain the association rule to obtain optimized association rules; Construct a data association network according to the optimized association rules, where the item sets in the association rules are used as network nodes and the association rules are used as the edges between nodes; Layout the data association network to obtain a visualized data association network graph.

5. The method according to claim 1, wherein According to the data association network, construct a multi-layer neural network model to learn the non-linear relationships between data features and obtain a deep learning model that can represent the internal structure of the data, including: Obtain the topological structure of the graph convolutional network according to the structure of the data association network; Among them, the graph convolutional network includes three graph convolutional layers; Preprocess the data set; Among them, the preprocessing includes: performing min-max scaling on numerical features and mapping the data into the interval [0, 1]; Performing Z-score standardization on continuous features to make the data mean 0 and the standard deviation 1; Performing one-hot encoding transformation on categorical features; Initialize the parameters of the graph convolutional network, including: initializing the weights of the ReLU activation function layer using the He initialization method; Initializing the weights of the tanh activation function layer using the Xavier initialization method; Initializing the bias to 0; Start the training of the graph convolutional network, including: setting the loss function as the mean squared error; Using the Adam optimization algorithm and setting the learning rate; Evaluate the graph convolutional network, including: calculating the accuracy, precision, recall, and F1 score using the test set; Using 5-fold cross-validation to evaluate the stability of the model performance; If the validation set accuracy is higher than the preset threshold or the maximum number of iterations is reached, stop the training.

6. The method according to claim 1, wherein, Evaluating the feature importance of data using a deep learning model, calculating the contribution degree of each feature to the modeling task, sorting the features according to the contribution degree, and selecting the feature subset with the highest contribution degree as the key region, including: Obtaining the partial derivative of the loss function with respect to the feature, which is calculated for each sample; Accumulating the partial derivatives of all samples according to the partial derivative and taking the average to obtain the feature importance score; Processing the feature importance score using the softmax function; Converting the feature importance score into a contribution percentage through the softmax function, where the contribution percentage is between 0 and 1, and the sum of the contribution degrees of all features is 1; Sorting the contribution percentages in descending order; Setting a contribution threshold, screening out the feature subset with a contribution degree greater than the contribution threshold to obtain a preliminary key region; Performing 5-fold cross-validation, which is used to evaluate the performance of the model using only the preliminary key region; Comparing the performance of the model with the performance of the model using all features. If the performance degradation does not exceed the preset threshold, the feature subset is retained; Performing a correlation analysis on the features in the preliminary key region; Calculating the Pearson correlation coefficient matrix between features, which is used to represent the correlation degree between features; Setting a correlation threshold and performing hierarchical clustering on the features with a correlation coefficient higher than the correlation threshold; Selecting the feature with the highest contribution degree from each cluster as a representative to obtain an optimized key region feature subset.

7. The method according to claim 1, wherein, The adaptive sampling of the selected key region includes dynamically adjusting the sampling density according to the data distribution characteristics, increasing the sampling points in the important region and reducing the sampling points in the secondary region to generate a targeted sampling scheme, including: Obtaining the probability density function of each feature within the selected key region to obtain the data distribution analysis result; Calculating the importance weight of the sampling points according to the data distribution analysis result, where the importance weight is mapped from the probability density function; Setting sampling density parameters for the important region and the secondary region. If the sampling point is in the important region, a high sampling density parameter is used; If the sampling point is in the secondary region, a low sampling density parameter is used; Generating candidate sampling points that meet the minimum distance threshold requirement based on the sampling density parameter and the preset spatial constraint conditions; Judging whether it is necessary to adjust the boundary threshold or the sampling density parameter of the important region and the secondary region according to the distribution of the sampled points. If adjustment is needed, update the boundary threshold or the sampling density parameter; If no adjustment is needed, continue the sampling process until the preset stop condition is met.

8. The method according to claim 1, wherein Analyzing the key region data based on the sampling scheme, capturing the time-varying characteristics of the data, analyzing the interaction between variables, constructing a targeted hybrid modeling framework, and realizing the accurate characterization and modeling of the key region that has the greatest impact on the modeling task, including: Obtaining sliding window features from time series data, calculating the autocorrelation coefficient and the partial autocorrelation coefficient through the sliding window features, and judging the stationarity and periodicity of the time series data; If the time series data is non-stationary, perform differencing on the time series data until the time series data becomes stationary; Use the principal component analysis method to perform dimensionality reduction on the time series features to obtain a reduced-dimensional feature set; Calculate the Pearson correlation coefficient matrix and mutual information matrix between the features in the reduced-dimensional feature set, and construct a feature interaction network; Construct a hybrid modeling framework, including prior probability, constraint conditions, and a custom loss function; Perform seasonal decomposition on the periodic features in the feature interaction network to obtain a trend component, a seasonal component, and a random component; Based on the hybrid modeling framework, model the key region data, and set a rolling prediction window using the time series cross-validation method.