Feature-based data screening and predicting method and system
By building a probability density network in data screening and prediction, and automatically screening high-density sample data, the problems of time-consuming and labor-intensive data screening and inaccurate model prediction in the prior art are solved, and efficient and reliable data processing and prediction results are achieved.
Patent Information
- Application Number
- CN202510351540.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology lacks automation and systematization in data screening and sample selection, and the feature selection and weighting processing are not in-depth enough, resulting in data screening time-consuming and labor-intensive and inconsistent results, complex probability modeling and high computing resources, and difficult to adapt to complex data distribution, affecting the model training effect and prediction accuracy.
By obtaining the initial features of the training set and the prediction set, calculating the correlation degree and filtering the basic features, assigning feature weights to build a probability density network, filtering high-density sample data for prediction, and forming an automated closed-loop process for data screening and model optimization.
The automation of feature selection and probability modeling is realized, the efficiency and accuracy of data screening are improved, and the changes in data distribution are dynamically adapted to the reliability and consistency of the prediction results are ensured, and the computational complexity is reduced.
Smart Images

Figure CN120296373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of battery capacity prediction, and particularly to a method and system for data screening and prediction based on features. Background Art
[0002] In the fields of data analysis and machine learning, data preprocessing and sample selection are key steps for optimizing model performance. However, there are still many defects and deficiencies in the existing technologies in terms of data screening and sample selection. Among them, in terms of data screening, it usually relies on manual experience and domain knowledge, lacking an automated and systematic process, resulting in a time-consuming and laborious data screening process, and being easily affected by subjective factors, making it difficult to ensure the objectivity and consistency of the screening results; in terms of feature selection and weighting processing, it relies on simple statistical indicators, lacking in-depth analysis of feature importance and weights; in terms of probability modeling, it is often too complex to efficiently run on large-scale data sets, resulting in excessive consumption of computing resources; in terms of sample selection, the existing methods usually rely on fixed rules or simple statistics, making it difficult to adapt to complex data distributions and dynamically changing business requirements, resulting in the selected samples may not fully represent the overall characteristics of the data, affecting the training effect and generalization ability of subsequent models;
[0003] In addition, when using a small-scale training set, the model cannot fully capture global information due to the small amount of data used for training, resulting in poor performance when predicting the prediction set. However, not all prediction results are inaccurate. For example, if the sample data in the prediction set has a high similarity to the sample data in the training set, then the prediction model trained based on this training set will surely obtain reliable prediction results when predicting such sample data;
[0004] In view of the above problems, how to achieve automation of feature selection, weighting processing, and probability modeling, and on this basis, effectively eliminate the sample data in the prediction set that the model does not master, so as to obtain reliable prediction results through efficient and accurate sample selection, this method is particularly important. Summary of the Invention
[0005] The main technical problem to be solved by the present invention is how to achieve automation of feature selection, weighting processing, and probability modeling, and on this basis, effectively eliminate the sample data in the prediction set that the model does not master, so as to obtain reliable prediction results through efficient and accurate sample selection.
[0006] According to a first aspect, in one embodiment, a method for data screening and prediction based on features is provided, including:
[0007] Obtain a training set and a prediction set, and each sample data in the training set and the prediction set corresponds to multiple initial features;
[0008] For any initial feature: based on the initial feature corresponding to each sample data in the training set and the true value corresponding to each sample data in the training set, obtain the correlation degree between the initial feature and the true value;
[0009] Based on the correlation degrees corresponding to all initial features, obtain multiple basic features; based on each basic feature corresponding to each sample data in the training set and the correlation degree corresponding to this basic feature, obtain the target feature vector of each sample data in the training set;
[0010] Construct a probability density network based on the target feature vectors corresponding to all sample data in the training set;
[0011] Obtain the target feature vector corresponding to each sample data in the prediction set; based on the obtained probability density network and the target feature vector corresponding to each sample data in the prediction set, obtain the density value corresponding to each sample data in the prediction set;
[0012] Screen the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain a screened prediction set;
[0013] Train a preset model according to the training set and perform prediction on the screened prediction set to obtain a reliable prediction result.
[0014] According to a second aspect, an embodiment provides a system for data screening and prediction based on features, including:
[0015] A data acquisition module, configured to acquire a training set and a prediction set, and each sample data in the training set and the prediction set corresponds to multiple initial features;
[0016] A correlation degree analysis module, configured to, for any initial feature: based on the initial feature corresponding to each sample data in the training set and the true value corresponding to each sample data in the training set, obtain the correlation degree between the initial feature and the true value;
[0017] A target feature vector acquisition module, configured to, based on the correlation degrees corresponding to all initial features, obtain multiple basic features; based on each basic feature corresponding to each sample data in the training set and the correlation degree corresponding to this basic feature, obtain the target feature vector of each sample data in the training set;
[0018] A probability density network acquisition module, configured to construct a probability density network based on the target feature vectors corresponding to all sample data in the training set;
[0019] A density value acquisition module, configured to acquire a target feature vector corresponding to each sample data in the prediction set; and obtain a density value corresponding to each sample data in the prediction set according to the obtained probability density network and the target feature vector corresponding to each sample data in the prediction set.
[0020] A predicted sample screening module, configured to screen the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set, so as to obtain a screened prediction set.
[0021] A reliable prediction result acquisition module, configured to train a preset model according to the training set, and predict the screened prediction set to obtain a reliable prediction result.
[0022] A method and system for data screening and prediction based on features according to the above embodiments. First, feature extraction is performed on each sample data in the entire data set, so that each sample data corresponds to multiple initial features. By analyzing the correlation between different initial features and the true value, features that have an important impact on the target variable (true value) can be more accurately identified, avoiding the limitations of relying on artificial experience and simple statistical indicators in traditional feature selection methods, improving the efficiency and accuracy of feature selection, and thus realizing the automatic screening of initial features to obtain multiple basic features with a strong correlation with the true value. Then, corresponding reference weights are assigned to different basic features according to the magnitude of the correlation degree corresponding to different basic features, obtaining the target feature vector corresponding to each sample data in the training set, and then constructing a probability density network. Different from the complex and computationally resource-consuming probability modeling methods in the prior art, the probability density network can significantly reduce the computational complexity while ensuring the modeling accuracy, improving the modeling efficiency, and thus better adapting to the processing requirements of large-scale data sets. Then, based on the obtained probability density network, combined with the target feature vector corresponding to each sample data in the prediction set, the density value corresponding to each sample data in the prediction set is obtained, and the sample data with the highest density value is screened out from the prediction set according to a preset ratio (i.e., the second preset threshold) to obtain the screened prediction set. By dynamically adapting to the changes in the data distribution, more representative and informative sample data are screened out from the prediction set, significantly improving the flexibility and accuracy of sample selection. And because the sample data retained in the screened prediction set has a high similarity with the sample data in the training set, predicting the screened prediction set can obtain a reliable prediction result. Different from the phenomenon of isolation and lack of coordination in each link in the prior art, by organically combining data preprocessing, feature selection, probability modeling, and sample screening together, a complete closed-loop process of automatic data screening and model optimization is formed, realizing the seamless connection between data processing and model optimization, significantly improving the efficiency and effect of the overall process, and thus effectively removing the sample data not covered by the model in the prediction set on the basis of realizing the automation of feature selection, weighted processing, and probability modeling, obtaining a reliable prediction result, which is more simple and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flowchart of a method for a data screening and prediction method based on features;
[0024] Figure 2 is a system block diagram of a system for a data screening and prediction based on features;
[0025] Figure 3 is a schematic diagram of the performance of the coefficient of determination of the prediction result obtained by the traditional method;
[0026] Figure 4Schematic diagram of the mean absolute error performance of the prediction results obtained by the traditional method;
[0027] Figure 5 Schematic diagram of the coefficient of determination performance of the prediction results obtained by this application;
[0028] Figure 6 Schematic diagram of the mean absolute error performance of the prediction results obtained by this application. Detailed implementation manners
[0029] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific implementation manners. Similar elements in different implementation manners adopt related similar element numbers. In the following implementation manners, many detailed descriptions are provided to enable a better understanding of this application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, and methods. In some cases, some operations related to this application are not shown or described in the specification to avoid the core part of this application being overwhelmed by excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the descriptions in the specification and general technical knowledge in the art.
[0030] In addition, the features, operations, or characteristics described in the specification can be combined in any appropriate manner to form various implementation manners. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean a necessary sequence, unless it is stated that a certain sequence must be followed.
[0031] The serial numbers assigned to the components in this article, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning. And the "connection" and "coupling" mentioned in this application, unless otherwise specified, both include direct and indirect connection (coupling).
[0032] The present invention overcomes the problems in the prior art such as complex data preprocessing processes, lack of systematicness in feature selection, low efficiency of probability modeling, and inflexible sample screening strategies. It can significantly improve the efficiency and accuracy of data screening, and at the same time optimize the rationality of sample selection, providing high-quality data support for subsequent machine learning model training and application. Furthermore, it solves the deficiencies in data preprocessing, feature selection, probability modeling, and sample screening in the prior art, and can achieve more efficient, intelligent, and highly automated data processing and sample screening.
[0033] In the embodiments of the present invention, first, feature extraction is performed on each sample data based on the entire data set to obtain a variety of initial features; according to the degree of association between the variety of initial features corresponding to each sample data in the training set and the true value, multiple features with relatively high relevance are screened out from all the initial features and denoted as basic features, and these basic features will be used as the main reference features for subsequent data screening and analysis; then, corresponding reference weights are assigned to different basic features according to the magnitude of the relevance corresponding to these basic features, so as to screen out sample data in the prediction set that is similar to the sample data features in the training set; since the prediction model is trained based on the training set, and the sample data in the screened prediction set is highly similar to the sample data in the training set, reliable prediction is achieved by predicting the screened prediction set, and a reliable prediction result is obtained.
[0034] Please refer to Figure 1 , in some embodiments, a method for data screening and prediction based on features is provided, which includes the following steps:
[0035] Step S100: Obtain a training set and a prediction set, where each sample data in the training set and the prediction set corresponds to a variety of initial features.
[0036] Obtain a data set composed of sample data, such as cell capacity data, perform data cleaning on the obtained data set, such as outlier detection and missing value processing, and standardize and normalize all sample data; then perform feature extraction on the data set after cleaning is completed. Exemplarily, feature engineering methods can be used to perform feature extraction on the entire data set to obtain a variety of features corresponding to each sample data. In this embodiment, the obtained variety of features are denoted as initial features, that is, at this time, each sample data in the data set corresponds to a variety of initial features, and the number of types of initial features corresponding to each sample data is the same;
[0037] Then, the entire data set is divided into a training set and a prediction set. Exemplarily, the entire data set can be divided according to a preset ratio, such as taking 30% of the entire data set as the training set, and the remaining 70% of the data as the prediction set, where each sample data in the training set also has its corresponding true value.
[0038] Step S110: Obtain a variety of basic features according to the relevance corresponding to all the initial features.
[0039] Since the variety of initial features corresponding to the sample data are all helpful for the prediction target, some features may be irrelevant to the prediction target, or even introduce noise and reduce the model performance. Therefore, in this embodiment, it is necessary to first screen out the features with a strong correlation with the true value from all the initial features, so as to subsequently screen out sample data in the prediction set that is highly similar to the sample data in the training set based on these features;
[0040] In this embodiment, for any initial feature: according to the initial feature corresponding to each sample data in the training set and the true value corresponding to each sample data in the training set, the correlation degree between the initial feature and the true value is obtained;
[0041] For example, for a certain initial feature A, the numerical values corresponding to the initial feature A of each sample data in the training set are obtained. Then, the numerical values corresponding to the initial feature A of all sample data in the training set form a feature value sequence; the true values corresponding to each sample data in the training set are obtained, and the true values corresponding to all sample data in the training set form a true value sequence. Since the Spearman method has better capturing ability for non-linear relationships and better robustness to outliers compared with the traditional Pearson coefficient, and can further weaken the influence of outliers, the Spearman method can be used to calculate the correlation degree between the obtained feature value sequence and the true value sequence; and so on, the correlation degree corresponding to each initial feature is obtained.
[0042] Since the positive and negative correlations of the correlation degree are caused by the co-directionality and reverse directionality of the change trends between variables, and both positive correlation and negative correlation indicate strong correlations between variables, in this embodiment, all initial features are screened based on the absolute value of the correlation degree. Among them, for any initial feature, when the absolute value of the correlation degree corresponding to the initial feature is greater than the first preset threshold, the initial feature is used as a basic feature; according to the correlation degrees corresponding to all initial features and the first prediction threshold, multiple basic features are obtained, and thus multiple features with strong correlations with the true value are screened out from multiple initial features as basic features. Denote the number of types of the obtained basic features as N.
[0043] Exemplarily, the first preset threshold can be set according to the actual situation. Since 0.4 is the dividing line between weak correlation and medium correlation in terms of statistical significance, removing uncorrelated features will have a better impact on the construction of the subsequent probability density network, so the first preset threshold in this embodiment can be set to 0.4.
[0044] Step S120: Obtain the target feature vector of each sample data in the training set according to each basic feature of each sample data in the training set and the correlation degree corresponding to the basic feature.
[0045] Although the obtained multiple basic features all have a high correlation with the true value, the specific correlation degrees between different basic features and the true value are not the same. That is, when judging the similarity between the sample data in the prediction set and the sample data in the training set, different basic features should have different reference degrees. Therefore, in this embodiment, different reference weights are further assigned to each basic feature according to the correlation degree corresponding to each basic feature.
[0046] Exemplarily, the reference weight corresponding to each basic feature can be dynamically generated according to the proportion of the relevance corresponding to each basic feature relative to the sum of the relevances corresponding to all basic features. In this embodiment, first, the absolute values of the relevances corresponding to all basic features are accumulated, and then the ratio between the absolute value of the relevance corresponding to each basic feature and the obtained accumulated sum is calculated, and the obtained ratio is used as the reference weight corresponding to this basic feature. Denote the reference weight corresponding to the i-th basic feature as c i ;
[0047] Then, for each sample data in the training set, calculate the product between the value corresponding to each basic feature of this sample data and the reference weight corresponding to this basic feature respectively. The products corresponding to all basic features of this sample data constitute the target feature vector of this sample data;
[0048] For example, for the i-th basic feature of this sample data, if the value corresponding to this basic feature is s i , then the product corresponding to this basic feature is c i ×s i , then the value on the i-th dimension in the target feature vector corresponding to this sample data is c i ×s i .
[0049] Step S130: Construct a probability density network according to the target feature vectors corresponding to all sample data in the training set.
[0050] Exemplarily, through the kernel density estimation model KDE, construct a probability density network based on the target feature vectors corresponding to all sample data in the training set;
[0051] Among them, the target feature vectors corresponding to all sample data in the training set constitute a feature matrix. Input the obtained feature matrix into the kernel density estimation model KDE to construct a multi-dimensional probability density function. In this embodiment, the obtained multi-dimensional probability density function is called a probability density network, and high-quality sample data is screened based on the obtained probability density network in the follow-up;
[0052] It should be noted that in this embodiment, the kernel function used by the kernel density estimation function can be dynamically selected according to the data distribution characteristics. The kernel function it uses can be a Gaussian kernel, a uniform kernel, a cosine kernel, etc. The kernel function used in this embodiment is a Gaussian kernel; in addition, when the kernel density estimation function selects the bandwidth, it can adaptively adjust the bandwidth size according to the dimension of the target feature vector, and the preferred range is 0.1 - 0.5.
[0053] Step S140: Obtain the target feature vector corresponding to each sample data in the prediction set, and combine the obtained probability density network to obtain the density value corresponding to each sample data in the prediction set.
[0054] Since each sample data in the entire dataset corresponds to multiple initial features, each sample data in the prediction set also corresponds to multiple initial features. Then, basic features are selected from the multiple initial features corresponding to each sample data in the prediction set. At this time, the types of basic features corresponding to each sample data in the prediction set are the same as those corresponding to each sample data in the training set;
[0055] For any sample data in the prediction set, multiply the value corresponding to each basic feature of the sample data by the reference weight corresponding to the basic feature. Here, the reference weight corresponding to the basic feature is the reference weight obtained in step S120. Then, the product corresponding to all the basic features of the sample data constitutes the target feature vector of the sample data; thus, the target feature vector corresponding to each sample data in the prediction set is obtained; and the dimension of the target feature vector is also N;
[0056] Then, input the target feature vector corresponding to each sample data in the prediction set into the probability density network to obtain the density value corresponding to each sample data in the prediction set;
[0057] It should be noted that for the probability density network constructed based on the kernel density estimation model KDE, its output result is the logarithm of the density value. Therefore, in order to avoid floating - point underflow, the density value also needs to be restored through exponential operation. For example, for the j - th sample data in the prediction set, the result obtained after processing the sample data by the probability density network is d j , then the density value corresponding to the sample data is exp(d j ), where exp() represents the exponential function with the natural constant e as the base.
[0058] Step S150: Obtain the filtered prediction set according to the density value corresponding to each sample data in the prediction set.
[0059] Sort in descending order according to the density value corresponding to each sample data in the prediction set. At this time, the sample data with a higher density value indicates that it has a higher similarity to the sample data in the training set; and the higher the density value, the higher the corresponding similarity, and the more reliable the prediction result obtained after predicting the sample data. Therefore, the sample data with a high density value is considered to be of high quality and needs to be retained; while the sample data with a lower density value indicates that its similarity to the sample data in the training set is lower, and it is more likely to be a sample data that the model has not mastered or covered. The higher the possibility that the prediction result obtained after predicting the sample data is unreliable. Therefore, the sample data with a low density value is considered to be of poor quality and needs to be excluded;
[0060] Although retaining only a small amount of high-quality data in the prediction set will yield more reliable prediction results, in practical applications, it is usually necessary to predict a large number of sample data at one time. That is, in practical applications, it is often inclined to retain most of the sample data in the prediction set. Moreover, due to the complexity and diversity of real-world data, in order to meet the prediction requirements in practical applications, it is necessary to retain some sample data with low density values. In this embodiment, the obtained sorting result is screened according to a second preset threshold, and the second preset threshold represents the proportion of sample data to be retained. The preferred range of the second preset threshold in this embodiment is 0.9 to 0.99. This means that 90% to 99% of the sample data in the prediction set should be retained.
[0061] Among them, the number of samples to be predicted is obtained according to the total number of samples included in the prediction set and the second preset threshold. Exemplarily, the second preset threshold is multiplied by the total number of samples included in the prediction set, and the result obtained after rounding down the obtained multiplication result is the amount of sample data / to-be-predicted sample number that the prediction set finally needs to retain. Then, according to the sorting result and the obtained number of samples to be predicted, a screened prediction set is obtained. For example, when the second preset threshold is 0.9, the first 90% of the sample data in the sorting result needs to be retained.
[0062] Step S160: Train a preset model according to the training set and predict the screened prediction set to obtain a reliable prediction result.
[0063] Since the sample data included in the screened prediction set at this time has a high similarity with the sample data in the training set, the preset model obtained after training based on the training set will have a high reliability of the prediction result after using the preset model to predict the screened prediction set. Thus, a reliable prediction result is obtained. Among them, the preset model can be set as a Linear Regression model, a Random Forest model, a LightGBM model, an XGBoost model, etc.
[0064] In addition, the method in this embodiment can be applied not only in the field of battery cell capacity prediction, but also in other fields, such as housing price prediction. Taking the Linear Regression model, the Random Forest model, the LightGBM model, and the XGBoost model as examples, the method in this embodiment is further verified using these four representative models. Among them, the original number of sample data included in the prediction set is 9858, and the second preset threshold is set to 0.95. Then, the number of samples to be predicted included in the screened prediction set is 9365.
[0065] After different models adopt the traditional method and the method in this embodiment, the coefficient of determination R corresponding to the obtained prediction results 2 is verified with the change of the mean absolute error (MAE). The coefficient of determination essentially measures the fitting degree of the model to the data. The higher the coefficient of determination, the better the fitting effect of the model. The mean absolute error is used to measure the average deviation between the predicted value of the model and the true value. The smaller the mean absolute error, the smaller the deviation between the prediction result and the true value, and the more reliable the corresponding prediction result;
[0066] The difference between the traditional method and the method in this embodiment is that after training the model with the training set, the former predicts the unfiltered prediction set to obtain the prediction result, while the latter predicts the filtered prediction set to obtain the prediction result;
[0067] Using the traditional method, that is, predicting the unfiltered prediction set, the performance of the coefficient of determination of the obtained prediction result is as shown in Figure 3 shown, and the performance of the mean absolute error of the obtained prediction result is as shown in Figure 4 shown; Using the method in this embodiment, that is, predicting the filtered prediction set, the performance of the coefficient of determination of the obtained prediction result is as shown in Figure 5 shown, and the performance of the mean absolute error of the obtained prediction result is as shown in Figure 6 shown; Figure 3 and Figure 5 The horizontal axis in corresponds to R 2 Score, indicating the magnitude of the coefficient of determination R 2 , and the vertical axis corresponds to different models Model; Figure 4 and Figure 6 The horizontal axis in corresponds to Mean Absolute Error, indicating the mean absolute error, and the vertical axis corresponds to different models. It can be seen that compared with the traditional method, for different models, the coefficient of determination R corresponding to the obtained prediction results 2 has increased significantly, and the mean absolute error has decreased significantly. This means that the fitting effects of these models have been greatly improved, the deviation between the prediction result and the true value has decreased, and the reliability of the corresponding prediction result has also been improved.
[0068] In this embodiment, first, based on the entire data set, feature extraction is performed on each sample data in the data set, so that each sample data corresponds to multiple initial features. By analyzing the correlation between different initial features and the true value, features that have an important impact on the target variable (true value) can be more accurately identified, avoiding the limitations of relying on manual experience and simple statistical indicators in traditional feature selection methods, improving the efficiency and accuracy of feature selection, and thus realizing the automatic screening of initial features to obtain multiple basic features with a strong correlation with the true value. Then, according to the magnitude of the correlation corresponding to different basic features, corresponding reference weights are assigned to different basic features to obtain the target feature vector corresponding to each sample data in the training set. Furthermore, a probability density network is constructed. Different from the complex and computationally resource-consuming probability modeling methods in the prior art, the probability density network used in this method can significantly reduce the computational complexity while ensuring the modeling accuracy, improve the modeling efficiency, and thus better adapt to the processing requirements of large-scale data sets.
[0069] Then, based on the obtained probability density network, combined with the target feature vector corresponding to each sample data in the prediction set, the density value corresponding to each sample data in the prediction set is obtained, and the sample data with the highest density value is selected from the prediction set according to a preset ratio (i.e., the second preset threshold) to obtain the filtered prediction set. Compared with the traditional sample screening method based on fixed rules or simple statistics, this method can dynamically adapt to the changes in data distribution, thereby screening out more representative and informative sample data, significantly improving the flexibility and accuracy of sample selection. And because the sample data retained in the filtered prediction set has a high similarity with the sample data in the training set, predicting the filtered prediction set can obtain reliable prediction results. Different from the phenomenon of isolation and lack of coordination in each link in the prior art, in this embodiment, data preprocessing, feature selection, probability modeling, and sample screening are organically combined to form a complete closed-loop process of automatic data screening and model optimization, realizing the seamless connection between data processing and model optimization, significantly improving the efficiency and effect of the overall process, and thus effectively removing the sample data not covered by the model in the prediction set on the basis of realizing the automation of feature selection, weighting processing, and probability modeling to obtain reliable prediction results.
[0070] Please refer to Figure 2 , in some embodiments, a system for data screening and prediction based on features is provided, and the system includes the following modules:
[0071] A data acquisition module 200, configured to acquire a training set and a prediction set, and each sample data in the training set and the prediction set corresponds to multiple initial features;
[0072] The relevance analysis module 210 is used to obtain the relevance between any initial feature and the true value according to the corresponding initial feature of each sample data in the training set and the true value corresponding to each sample data in the training set;
[0073] The target feature vector acquisition module 220 is used to obtain multiple basic features according to the relevance corresponding to all initial features; obtain the target feature vector of each sample data in the training set according to each basic feature of each sample data in the training set and the relevance corresponding to the basic feature;
[0074] The probability density network acquisition module 230 is used to construct a probability density network according to the target feature vectors corresponding to all sample data in the training set;
[0075] The density value acquisition module 240 is used to obtain the target feature vector corresponding to each sample data in the prediction set; obtain the density value corresponding to each sample data in the prediction set according to the obtained probability density network and the target feature vector corresponding to each sample data in the prediction set;
[0076] The predicted sample screening module 250 is used to screen the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain the screened prediction set;
[0077] The reliable prediction result acquisition module 260 is used to train a preset model according to the training set and predict the screened prediction set to obtain a reliable prediction result.
[0078] It should be noted that each module in this embodiment corresponds to the method steps in the above method for data screening and prediction based on features, and its specific implementation manners have been specifically described in the above embodiments, and will not be elaborated here.
[0079] In this embodiment, first, in the data acquisition module 200, feature extraction is performed on each sample data in the entire data set, so that each sample data corresponds to multiple initial features. Then, in the relevance analysis module 210, by analyzing the relevance between different initial features and the true value, features that have an important impact on the target variable (true value) can be more accurately identified, avoiding the limitations of relying on manual experience and simple statistical indicators in traditional feature selection methods, improving the efficiency and accuracy of feature selection, and thus realizing the automatic screening of initial features to obtain multiple basic features with a strong correlation with the true value. In the target feature vector acquisition module 220, corresponding reference weights are assigned to different basic features according to the magnitude of the relevance corresponding to different basic features, obtaining the target feature vector corresponding to each sample data in the training set. Based on the target feature vector corresponding to each sample data in the training set, a probability density network is constructed in the probability density network acquisition module 230. Different from the complex and computationally resource-consuming probability modeling methods in the prior art, the probability density network used in this method can significantly reduce the computational complexity while ensuring the modeling accuracy, improving the modeling efficiency, and thus better adapting to the processing requirements of large-scale data sets.
[0080] Subsequently, in the density value acquisition module 240, based on the obtained probability density network and combined with the target feature vector corresponding to each sample data in the prediction set, the density value corresponding to each sample data in the prediction set is obtained. In the predicted sample screening module 250, sample data with the highest density value is screened out from the prediction set according to a preset ratio (i.e., the second preset threshold), obtaining the screened prediction set. Compared with the traditional sample screening method based on fixed rules or simple statistics, this method can dynamically adapt to the changes in data distribution, thereby screening out more representative and informative sample data, significantly improving the flexibility and accuracy of sample selection. In the reliable prediction result acquisition module 260, since the sample data retained in the screened prediction set has a high similarity with the sample data in the training set, predicting the screened prediction set can obtain reliable prediction results. Different from the phenomenon of isolation and lack of coordination in each link in the prior art, this embodiment organically combines data preprocessing, feature selection, probability modeling, and sample screening together to form a complete closed-loop process of automatic data screening and model optimization, realizing the seamless connection between data processing and model optimization, significantly improving the efficiency and effect of the overall process, and thus effectively eliminating the sample data not covered by the model in the prediction set on the basis of realizing the automation of feature selection, weighting processing, and probability modeling, obtaining reliable prediction results.
[0081] Those skilled in the art can understand that all or part of the functions of the various methods in the above embodiments can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions can be realized by a computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, the above all or part of the functions can be realized. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, and saved to the memory of the local device by downloading or copying, or the system of the local device is updated in version. When the processor executes the program in the memory, all or part of the functions in the above embodiments can be realized.
[0082] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art of the present invention, according to the idea of the present invention, several simple deductions, deformations or substitutions can also be made.
Claims
1. A method for data screening and prediction based on features, characterized in that Including: Obtain a training set and a prediction set, where each sample data in the training set and the prediction set corresponds to multiple initial features; For any one of the initial features: according to the initial feature corresponding to each sample data in the training set and the true value corresponding to each sample data in the training set, obtain the correlation degree between the initial feature and the true value; Obtain multiple basic features according to the correlation degrees corresponding to all the initial features; obtain the target feature vector of each sample data in the training set according to each basic feature of each sample data in the training set and the correlation degree corresponding to the basic feature; Construct a probability density network according to the target feature vectors corresponding to all the sample data in the training set; Obtain the target feature vector corresponding to each sample data in the prediction set; according to the obtained probability density network and the target feature vector corresponding to each sample data in the prediction set, obtain the density value corresponding to each sample data in the prediction set; Screen the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain a screened prediction set; Train a preset model according to the training set and make a prediction on the screened prediction set to obtain a reliable prediction result.
2. The method according to claim 1, characterized in that, The obtaining multiple basic features according to the correlation degrees corresponding to all the initial features includes: Obtain multiple basic features according to the correlation degrees corresponding to all the initial features and a first prediction threshold; wherein, for any one of the initial features, when the absolute value of the correlation degree corresponding to the initial feature is greater than the first preset threshold, use the initial feature as a basic feature.
3. The method according to claim 1, characterized in that, The obtaining the target feature vector of each sample data in the training set according to each basic feature of each sample data in the training set and the correlation degree corresponding to the basic feature includes: Normalize the absolute value of the correlation degree corresponding to each basic feature, and use the obtained normalization result as the reference weight corresponding to each basic feature; For any one sample data, calculate the product between the value corresponding to each basic feature of the sample data and the reference weight corresponding to the basic feature respectively, and the products corresponding to all the basic features of the sample data constitute the target feature vector of the sample data.
4. The method according to claim 1, wherein The screening the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain a screened prediction set includes: sort each sample data in the prediction set according to the corresponding density value, and screen the obtained sorting result according to a second preset threshold to obtain a screened prediction set.
5. The method according to claim 4, characterized in that, The screening the obtained sorting result according to a second preset threshold to obtain a screened prediction set includes: obtain the number of samples to be predicted according to the total number of samples included in the prediction set and the second preset threshold, and obtain the screened prediction set according to the sorting result and the obtained number of samples to be predicted.
6. A system for data screening and prediction based on features, characterized in that, Including: A data acquisition module, configured to obtain a training set and a prediction set, where each sample data in the training set and the prediction set corresponds to multiple initial features; A relevance analysis module, which is used for any initial feature: according to the initial feature corresponding to each sample data in the training set and the true value corresponding to each sample data in the training set, obtain the relevance between the initial feature and the true value; A target feature vector acquisition module, which is used to obtain multiple basic features according to the relevance corresponding to all initial features; according to each basic feature of each sample data in the training set and the relevance corresponding to the basic feature, obtain the target feature vector of each sample data in the training set; A probability density network acquisition module, which is used to construct a probability density network according to the target feature vectors corresponding to all sample data in the training set; A density value acquisition module, which is used to obtain the target feature vector corresponding to each sample data in the prediction set; according to the obtained probability density network and the target feature vector corresponding to each sample data in the prediction set, obtain the density value corresponding to each sample data in the prediction set; A predicted sample screening module, which is used to screen the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain a screened prediction set; A reliable prediction result acquisition module, which is used to train a preset model according to the training set and predict the screened prediction set to obtain a reliable prediction result.
7. The system according to claim 6, wherein The obtaining multiple basic features according to the relevance corresponding to all initial features includes: Obtaining multiple basic features according to the relevance corresponding to all initial features and a first prediction threshold; wherein, for any initial feature, when the absolute value of the relevance corresponding to the initial feature is greater than the first preset threshold, the initial feature is used as a basic feature.
8. The system according to claim 6, wherein The obtaining the target feature vector of each sample data in the training set according to each basic feature of each sample data in the training set and the relevance corresponding to the basic feature includes: Normalizing the absolute value of the relevance corresponding to each basic feature, and using the obtained normalization result as the reference weight corresponding to each basic feature; For any sample data, calculate the product between the numerical value corresponding to each basic feature of the sample data and the reference weight corresponding to the basic feature respectively, and the products corresponding to all basic features of the sample data constitute the target feature vector of the sample data.
9. The system according to claim 6, wherein The screening the sample data in the prediction set according to the density value corresponding to each sample data in the prediction set to obtain a screened prediction set includes: sorting each sample data in the prediction set according to the corresponding density value, and screening the obtained sorting result according to a second preset threshold to obtain a screened prediction set.
10. A computer-readable storage medium, characterized in that, The medium stores a computer program, and the computer program can be executed by a processor to implement the method according to any one of claims 1-5.
Citation Information
Cited By
Model training method and related equipment
CN120632469A
Model training method and related device
CN120632469B