Data analysis method and system based on digital intelligent brain
By introducing multi-scale manifold distance and feature importance assessment in digital brain data analysis methods, the problem that the difference in feature contributions of traditional random forest models in multi-source heterogeneous data is not fully considered, and the decision accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN202510685421.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
When processing multi-source heterogeneous data, the traditional random forest model fails to fully consider the differences in the contribution of different characteristics to decision-making, which leads to the decision-making process being easily misled by inefficient characteristics and inaccurate output results.
Using a data analysis method based on the digital brain, a more accurate random forest model is constructed through multi-source data acquisition, standardization, label coding, t-SNE dimensionality reduction, multi-scale manifold distance calculation, reliability evaluation and feature importance calculation.
Through multi-scale processing of sample data and evaluation of feature importance, key features can be more accurately identified, inefficient feature interference can be suppressed, and the decision accuracy and robustness of the model can be improved.
Smart Images

Figure CN120197073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a data analysis method and system based on a digital intelligence brain. Background Art
[0002] In the wave of the big data era, the amount of data has grown exponentially. Its characteristics of diversity, high speed, and massiveness have profoundly changed the decision-making patterns in various fields. Data has become an extremely crucial driving factor in decision-making. Traditional decision-making methods have long relied on the past experience, subjective intuition of decision-makers, and extremely limited data samples. In today's complex and ever-changing business environment and scientific research problems and other scenarios, traditional methods rely only on a small amount of data and experience, making it difficult to comprehensively capture these factors, resulting in a significant reduction in decision-making accuracy, a long decision-making process, and low efficiency. The Chinese patent application document with the publication number CN114549211A discloses a data analysis method, device, equipment, and storage medium based on a random forest. The method includes: receiving an analysis instruction for hanging account data, where the analysis instruction includes hanging account information, and the hanging account information is generated from the attribute information of the hanging account data according to a preset format; parsing the hanging account information according to the preset format to obtain the attribute information of the hanging account data; based on the attribute information of the hanging account data, obtaining the scenario information of the hanging account data and the attribute information of the associated data associated with the hanging account data; obtaining a target random forest model corresponding to the scenario information; and inputting the attribute information of the hanging account data and the attribute information of the associated data into the target random forest model to obtain the reason for the hanging account of the hanging account data.
[0003] However, as a commonly used decision-making model, the random forest also has certain defects. When constructing a random forest model and performing integrated decision-making, traditional random forests lack in-depth consideration of the contribution differences of different features in the decision-making process. When facing multi-source heterogeneous data, such as scenarios that integrate various types of data such as text, images, and numerical values, the model cannot accurately distinguish important features from inefficient features, resulting in the decision-making process being frequently misled by inefficient features, and ultimately the accuracy of the output decision result is poor. Summary of the Invention To solve the problem that traditional random forests do not fully consider the contribution differences of different features to decision-making, and when facing multi-source heterogeneous data, the decision of the random forest model is easily interfered by inefficient features, resulting in inaccurate final output results, the present invention provides a data analysis method and system based on a digital intelligence brain.
[0004] In a first aspect, the present invention provides a data analysis method based on a digital intelligence brain, adopting the following technical solutions: A data analysis method based on a digital intelligence brain, including: obtaining samples containing multiple features in a target field through multi-source data collection channels; standardizing each feature of the samples, performing label encoding based on the standardized features of all samples to obtain a labeled sample data set, and using t-SNE to reduce the dimension of the sample data set to generate several subspaces at different scales; for any two samples in the sample data set, determining the multi-scale manifold distance between the two samples according to the difference in the values of the two samples in each subspace and the average overlap number of all samples in each subspace and the preset neighboring samples of the remaining subspaces; determining multiple neighboring samples of each sample according to the magnitude of the multi-scale manifold distance, and taking the normalized result of the average of the multi-scale manifold distances between each sample and all neighboring samples of each sample as the reliability of each sample; determining the feature importance of each feature according to the reliability, the value of each sample in each feature, and the mean value of all samples in each feature; constructing a random forest model based on the feature importance, training the random forest model using the sample data set, and using the trained random forest model to analyze newly obtained data to be measured.
[0005] The beneficial effects are as follows: Through the standardization, label encoding, and t-SNE dimensionality reduction of the sample data, multi-dimensional information of the data can be captured from multiple scales. The subspaces generated after dimensionality reduction effectively present the complex structure of the data and the correlations between various features, providing higher-quality feature input for subsequent analysis; The measurement method of multi-scale manifold distance is introduced. By calculating the relative positions and overlap degrees of samples in multiple subspaces, the model can more accurately identify the similarities and differences between samples, effectively avoid interference caused by feature noise, and improve the robustness of the model when processing heterogeneous data; By combining the multi-scale manifold distance of each sample and the information of its neighboring samples, the reliability of the sample is calculated, and based on this, the importance of the feature is evaluated by comprehensively considering the performance of each sample in different features and the mean value of all samples, which can more accurately focus on those features that have a greater impact on decision-making, thereby improving the accuracy of the model; On the basis of calculating the feature importance, a random forest model is reconstructed, enabling the model to pay more attention to key features and suppress the interference of inefficient or irrelevant features, thereby improving the overall performance of the random forest model; By standardizing different features and optimizing feature selection, heterogeneous data with diverse sources and complex structures can be effectively processed, providing an efficient data analysis method for processing large-scale and diverse application scenarios.
[0006] Further, the multi-source data collection channel is a relational database.
[0007] Further, the standardization adopts Z-score standardization.
[0008] Further, the method for obtaining the multi-scale manifold distance is as follows: Denote any subspace as the target subspace. For any sample, denote the ratio of the average number of overlapping samples of the sample with the preset nearest neighbor samples in the target subspace and the remaining subspaces to the number of preset nearest neighbor samples as the confidence weight of the sample, and denote the sum of the confidence weights of all samples as the confidence weight of the target subspace; Denote the product of the confidence weight and the difference between the values of the two samples in the target subspace as the manifold distance of the two samples in the target subspace; Denote the standard normalization result of the sum of the manifold distances of the two samples in all subspaces as the multi-scale manifold distance of the two samples.
[0009] The beneficial effects are as follows: By introducing the distance calculation of subspaces and the differences between samples within each subspace, the similarity of samples is comprehensively evaluated at multiple scales, better capturing the relationships between samples in different dimensions, more accurately reflecting the similarity between samples, and avoiding errors in a single scale; By performing weighted summation on the distances of different subspaces, the weights of distance calculation can be adaptively adjusted according to the degree of difference between samples within each subspace, enabling the model to dynamically adjust its attention to each subspace according to data characteristics, thereby more precisely reflecting the true relationship between samples.
[0010] Further, determining multiple nearest neighbor samples for each sample includes: For any sample, sort the multi-scale manifold distances of the sample with all the other samples in ascending order, and select multiple samples as the nearest neighbor samples of the sample based on the ascending order.
[0011] The beneficial effects are as follows: Through the multi-scale manifold distance, the structure of the data at different scales is considered, which can more accurately reflect the local structure of the data, thereby selecting more representative nearest neighbor samples, avoiding errors in a single scale, and being able to capture more potential relationships between samples.
[0012] Further, the normalization is performed using a logistic function.
[0013] Further, the method for obtaining the feature importance is as follows: Denote any feature as the target feature. For any sample, denote the negative correlation mapping of the square of the difference between the value of the sample in the target feature and the mean value of all samples in the target feature as the conformity degree of the sample in the target feature and all samples in the target feature; Denote the product of the conformity degree and the reliability of the sample as the importance of the sample in the target feature; Denote the mean value of the sum of the importances of all samples in the target feature as the feature importance of the target feature.
[0014] The beneficial effects are as follows: By introducing the reliability of samples, it can effectively avoid the distortion of feature importance evaluation caused by poor quality or excessive noise of some samples. The influence of samples with low reliability on the results will be reduced, thereby improving the robustness of the model. Since the exponential function (negative correlation mapping) is used, the contribution of samples with large deviations to feature importance will be more prominent, which helps to identify key samples with large differences from the mean in features, and can identify samples and features that may be important for classification or prediction tasks in feature selection. By evaluating feature importance, it can help select features with large variability among samples and significant contributions to the task, and avoid features that are too trivial or have little information from interfering with the performance of the model.
[0015] Further, the construction of the random forest model includes: using the mapping result of the feature importance of each feature through the function as the selection probability of each feature when constructing the random forest model, and based on the selection probability, constructing the random forest model to obtain the random forest model.
[0016] The beneficial effects are as follows: By considering the importance of features, feature selection is more intelligent, which helps to reduce the influence of redundant or irrelevant features and improve the accuracy of the model. Traditional random forests randomly select features, which may ignore some very important features. By mapping the feature importance to the selection probability, it can ensure that important features have a higher selection probability when constructing decision trees, thereby improving the performance of the model. Through the probability distribution of feature selection, the model can more targeted select features for training, reduce unnecessary calculations, and thus improve the training efficiency.
[0017] Further, the analysis of the newly obtained data to be measured includes: inputting the newly obtained data to be measured into the trained random forest model, obtaining the decision result corresponding to the label of the sample data set, and completing the data analysis based on the digital brain.
[0018] In the second aspect, the present invention provides a data analysis system based on a digital brain, adopting the following technical solutions: The data analysis system based on a digital brain includes: a processor and a memory, and the memory stores computer program instructions, which implement the above-mentioned data analysis method based on a digital brain when the computer program instructions are executed by the processor.
[0019] By adopting the above technical solutions, the above-mentioned data analysis method based on a digital brain is generated into a computer program and stored in the memory to be loaded and executed by the processor, so as to manufacture a terminal device according to the memory and the processor, which is convenient to use.
[0020] The present invention has the following technical effects: (1) By determining the feature importance of each feature, the differences in the contributions of different features to the decision-making can be fully considered. When constructing a random forest model, the interference of inefficient features to the model decision-making can be reduced, avoiding the problem of inaccurate final output results caused by the influence of inefficient features in the case of multi-source heterogeneous data, and improving the accuracy of the model decision-making.
[0021] (2) Using t-SNE to reduce the dimension of the sample dataset and generate several subspaces at different scales, and processing the data from multiple scales can more comprehensively mine the potential structure and information in the data, making the subsequent analysis of the relationship between samples more accurate and in-depth; determining the multi-scale manifold distance according to the values of the samples in different subspaces and the average overlapping number of the preset neighbor samples can more accurately measure the similarity and distance relationship between samples, so as to more reasonably determine the neighbor samples of each sample, providing a more reliable basis for calculating the sample reliability and feature importance subsequently.
[0022] (3) Calculating the reliability of each sample can quantitatively evaluate the quality of the sample, which helps to more reasonably utilize the sample in model training, improve the stability and generalization ability of the model, and avoid the decline of model performance caused by sample quality problems; constructing and training a random forest model based on feature importance enables the model to better adapt to the characteristics of multi-source heterogeneous data, and can more accurately analyze the newly obtained data to be measured, improving the processing ability and adaptability of the model to different types of data.
[0023] (4) Considering the reliability, the value of the sample in the feature, and the mean value of the sample in the feature comprehensively to determine the feature importance, making full use of various aspects of information of the sample, making the evaluation of the feature importance more scientific and reasonable, thus constructing a better random forest model and improving the accuracy of subsequent data analysis. Brief Description of the Drawings
[0024] Figure 1 is the flowchart of the data analysis method based on the digital intelligence brain in the embodiment of the present invention.
[0025] Figure 2 is the schematic diagram of the multi-scale manifold distance in the data analysis method based on the digital intelligence brain in the embodiment of the present invention. Detailed Embodiments
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] An embodiment of the present invention discloses a data analysis method based on a digital intelligence brain, with reference to Figure 1 , including steps S1 - S6: S1: Obtain samples containing multiple features in the target field through multi-source data collection channels.
[0028] It should be noted that the present invention can be widely applied to various fields, such as finance, medical care, transportation, energy, etc. Therefore, it is necessary to process multi-source heterogeneous data, that is, data with different sources and types. Since the random forest model can make decisions well based on multi-source heterogeneous data, therefore, for analysis based on the random forest, in order to facilitate subsequent training of the random forest model, it is necessary to first obtain sample data for training. Directly extract tabular data (for example, user basic information, device operation parameters, etc.) from the multi-source data collection channels to obtain samples.
[0029] Implementers can set the number of samples according to the specific implementation situation, for example, 1000.
[0030] Specifically, the multi-source data collection channel is a relational database (for example, MySQL, Oracle, etc.).
[0031] In another embodiment, the multi-source data collection channel is an application programming interface (for example, financial platform trading API, traffic sensor real-time data interface, etc.) S2: Obtain a sample data set with labels and generate several subspaces of the sample data set.
[0032] It should be noted that according to the target field (for example, financial risk control requires transaction records, medical diagnosis requires electronic medical records, etc.), define data labels (for example, risk level, disease type, etc.) and key features (for example, transaction amount, consultation time, energy consumption peak, etc.).
[0033] Standardize each feature of the samples, perform label encoding based on the standardized features of all samples to obtain a sample data set with labels, and use t-SNE to reduce the dimension of the sample data set to generate several subspaces at different scales. Among them, when using t-SNE for dimensionality reduction, perform independent t-SNE dimensionality reduction operations (exemplarily, ), preset the number of neighboring samples (exemplarily, ), and randomly select the perplexity within a preset range (exemplarily, ) each time, so as to generate subspaces at different scales.
[0034] Implementers can set the label encoding method according to the specific implementation situation, for example, one-hot encoding or label encoding.
[0035] Specifically, the standardization adopts Z-score standardization.
[0036] S3: Determine the multi-scale manifold distance between samples.
[0037] It should be noted that since the feature distributions of samples are in different dimensions and the sample position distributions are complex in the high-dimensional space, for samples with low overall aggregation degree, directly using the distance between each sample and the centroid of the sample set will result in low reliability of the samples. Therefore, as Figure 2 shown, it is necessary to first obtain the multi-scale manifold distance between each sample. In the figure, each origin represents a sample, and the black connection lines are the multi-scale manifold distances between two samples.
[0038] For any two samples in the sample data set, determine the multi-scale manifold distance between the two samples according to the difference in the values of the two samples in each subspace and the average overlap number of the preset neighboring samples of all samples in each subspace and the remaining subspaces.
[0039] Specifically, the multi-scale manifold distance satisfies: ; In the formula, is the multi-scale manifold distance between the th sample and the th sample, and are respectively the values of the th sample and the th sample in the th subspace, is the average overlap number of the preset neighboring samples of the th sample in the th subspace and the remaining subspaces, is the number of preset neighboring samples; is the number of samples in the sample data set, is the number of subspaces, is the standard normalization function, is the absolute value symbol.
[0040] Among them, represents the confidence weight of the th subspace. When the average overlap number of the preset neighboring samples of all samples in the subspace and the samples in other subspaces is larger, the local structure under this subspace is more credible and has a higher impact on the multi-scale manifold distance; represents the distance between the th sample and the th sample in the th subspace. When in different subspaces the When the overall distance between one sample and the th sample is closer, the multi-scale manifold distance between the two samples is smaller.
[0041] S4: Obtain the reliability of each sample.
[0042] It should be noted that in order to avoid the reduction of accuracy in calculating feature importance caused by the feature differences of bad samples (samples that are significantly different from normal samples due to sampling bias or noise interference), it is necessary to obtain the reliability of each sample first.
[0043] According to the magnitude of the multi-scale manifold distance, determine multiple nearest neighbor samples for each sample, and use the normalized result of the mean of the multi-scale manifold distances between each sample and all its nearest neighbor samples as the reliability of each sample.
[0044] Specifically, the determination of multiple nearest neighbor samples for each sample includes: For any sample, sort the multi-scale manifold distances between this sample and all other samples in ascending order, and select multiple samples as the nearest neighbor samples of this sample based on the ascending order.
[0045] Specifically, the normalization is performed using the logistic function ( , ).
[0046] Among them, the logistic function is a common S-shaped curve, and its mathematical expression is: ; is the input value (real number), is the natural exponential function, which is an exponential function with the natural constant as the base (the natural constant is approximately equal to 2.71828). The logistic function effectively solves the mapping problem from manifold distance to reliability through non-linear normalization, and is especially suitable for dealing with local structure differences and global noise interference in high-dimensional data, providing a stable and interpretable weight basis for subsequent feature importance calculation.
[0047] S5: Determine the feature importance of each feature.
[0048] It should be noted that due to the variety of data types in multi-source heterogeneous data, it is easy to have significant differences in individual features of bad samples, resulting in inaccurate calculation of feature importance. By obtaining the reliability of each sample through the above steps, and then based on the features of high-reliability samples, obtain the feature importance of each feature. Finally, establish a random forest model by fitting as many effective features as possible according to the feature importance, so that the integrated decision-making of the random forest model is more accurate.
[0049] Determine the feature importance of each feature according to the reliability, the value of each sample in each feature, and the mean value of all samples in each feature.
[0050] Specifically, the feature importance satisfies: ; In the formula, is the feature importance of the th feature, is the reliability of the th sample, is the value of the th sample in the th feature, is the mean value of all samples in the th feature, is the number of samples in the sample data set, is the natural exponential function.
[0051] Among them, represents the degree of conformity of the th sample in the th feature with all samples in the th feature. When the overall degree of conformity of all samples with high reliability in the th feature with all samples in the th feature is higher, the feature importance of the th feature is higher.
[0052] S6: Based on the feature importance, construct a random forest model, train the random forest model using the sample data set, and use the trained random forest model to analyze the newly obtained data to be measured.
[0053] It should be noted that to improve the decision-making accuracy of the model, when constructing the random forest model, more features with high feature importance need to be integrated.
[0054] Specifically, the construction of the random forest model includes: Use the mapping result of the feature importance of each feature through the function as the selection probability of each feature when constructing the random forest model. Based on the selection probability, construct the random forest model to obtain the random forest model.
[0055] It should be noted that when training the random forest model using the sample data set with labels, the implementer can set various parameters of the random forest model according to the specific implementation situation. For example, the preset number of trees is , and the maximum tree depth is , the minimum number of samples in the leaf nodes is 5. If it is a classification task, the cross-entropy loss function is selected; if it is a regression task, the mean squared error loss function is selected. In the sample dataset samples are used as the training set, and 20% of the samples are used as the test set. The random forest model with improved ensemble decision-making is trained using the training set data. Through model iteration, the training is terminated until the loss function is less than the preset value of 0.8 or the accuracy of the test set is higher than the preset value of 96%, and the trained random forest model is obtained.
[0056] Specifically, the analysis of the newly acquired data to be measured includes: The newly acquired data to be measured is input into the trained random forest model (in the same way as the sample acquisition method), and the decision result corresponding to the label of the sample dataset is obtained (the category for the classification task and the value for the regression task). Based on the model decision result, subsequent tasks in the corresponding field are carried out (for example, hierarchical diagnosis and treatment in medicine, risk interception threshold in finance, etc.), and the data analysis based on the digital intelligence brain is completed.
[0057] The embodiment of the present invention also discloses a data analysis system based on the digital intelligence brain, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the data analysis method based on the digital intelligence brain according to the present invention is implemented.
[0058] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.
[0059] The above are all the preferred embodiments of the present invention. The protection scope of the present invention is not limited hereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A data analysis method based on a digital intelligence brain, characterized in that including: obtaining samples containing multiple features in the target field through multi-source data collection channels; standardizing each feature of the samples, performing label encoding based on the standardized features of all samples to obtain a labeled sample dataset, and using t-SNE to reduce the dimension of the sample dataset to generate several subspaces at different scales; for any two samples in the sample dataset, determining the multi-scale manifold distance between the two samples according to the difference in the values of the two samples in each subspace and the average number of overlapping samples of all samples in each subspace with the preset neighboring samples in the remaining subspaces; determining multiple neighboring samples of each sample according to the magnitude of the multi-scale manifold distance, and taking the normalized result of the mean of the multi-scale manifold distances between each sample and all neighboring samples of each sample as the reliability of each sample; determining the feature importance of each feature according to the reliability, the value of each sample in each feature, and the mean of all samples in each feature; constructing a random forest model based on the feature importance, training the random forest model using the sample dataset, and using the trained random forest model to analyze newly obtained data to be tested.
2. The data analysis method based on the digital intelligence brain according to claim 1, wherein The multi-source data collection channel is a relational database.
3. The data analysis method based on the digital intelligence brain according to claim 1, wherein, The standardization uses Z-score standardization.
4. The data analysis method based on the digital intelligence brain according to claim 1, wherein The method for obtaining the multi-scale manifold distance is as follows: Denote any subspace as the target subspace. For any sample, denote the ratio of the average number of overlapping samples of the sample in the target subspace with the preset neighboring samples in the remaining subspaces to the number of preset neighboring samples as the trust weight of the sample, and denote the sum of the trust weights of all samples as the trust weight of the target subspace; Denote the product of the trust weight and the difference in the values of the two samples in the target subspace as the manifold distance between the two samples in the target subspace; Denote the standard normalization result of the sum of the manifold distances between the two samples in all subspaces as the multi-scale manifold distance between the two samples.
5. The data analysis method based on the digital intelligence brain according to claim 1, wherein Determining multiple neighboring samples of each sample includes: For any sample, sort the multi-scale manifold distances between the sample and all other samples according to their magnitudes, and select multiple samples in ascending order as the neighboring samples of the sample.
6. The data analysis method based on the digital intelligence brain according to claim 1, wherein The normalization uses a logistic function for normalization.
7. The data analysis method based on the digital intelligence brain according to claim 1, wherein The method for obtaining the feature importance is as follows: Denote any feature as the target feature. For any sample, denote the negative correlation mapping of the square of the difference between the value of the sample in the target feature and the mean of all samples in the target feature as the conformity degree of the sample in the target feature with all samples in the target feature; Denote the product of the conformity degree and the reliability of the sample as the importance of the sample in the target feature; Denote the mean of the sum of the importances of all samples in the target feature as the feature importance of the target feature.
8. The data analysis method based on the digital intelligence brain according to claim 1, characterized in that Constructing the random forest model includes: The feature importance of each feature is used as the selection probability of each feature in the construction of the random forest model through the mapping result of the function. Based on the selection probability, a random forest model is constructed to obtain the random forest model.
9. The data analysis method based on the digital intelligence brain according to claim 1, characterized in that Analyzing the newly obtained data to be tested includes: Inputting the newly obtained data to be tested into the trained random forest model, obtaining a decision result corresponding to the label of the sample dataset, and completing the data analysis based on the digital intelligence brain.
10. A data analysis system based on a digital intelligence brain, characterized in that, including: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the data analysis method based on the digital intelligence brain according to any one of claims 1-9.
Citation Information
Patent Citations
Data analysis method and device based on random forest, equipment and storage medium
CN114549211A
Method and system for optimizing classification of random forest based on weighted decision trees
CN107766883A
Laser radar human leg detection method based on multi-scale adaptive random forest
CN111444769A
Non-data watershed runoff simulation method based on domain adaptation and machine learning
CN117973237A
Data analysis method and device based on improved random forest and medium
CN118761480A
Cited By
Engine health management method and system
CN120724160A
An engine health management method and system
CN120724160B