Background event discrimination method based on random forest classification model

By adopting the background event screening method based on the random forest classification model in the liquid flash detection system, the problems of low background reduction efficiency and high construction cost in the prior art are solved, and more accurate measurement and more effective background event screening of lower activity samples are achieved.

CN120067896AActive Publication Date: 2025-05-30NATIONAL INSTITUTE OF METROLOGY CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510538510.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When the prior art reduces the background level of the liquid flash detection system, there are problems of marginal effects and high construction costs. At the same time, the identification ability of the pulse identification method based on signal processing is limited, and it is easy to mistakenly identify real events as background events.

Method used

The background event screening method based on the random forest classification model is adopted, and the random forest classification model is constructed and optimized to distinguish real signals and background events by pre-processing and feature dimensionality reduction of the original data sets of high-activity standard samples and blank samples.

Benefits of technology

Without significantly increasing the system construction cost, the measurement ability of the liquid flash analyzer for lower activity samples is improved, the screening effect of background events is enhanced, and the loss of real events is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067896A_ABST
    Figure CN120067896A_ABST
Patent Text Reader

Abstract

The invention discloses a background event discrimination method based on a random forest classification model, and the method comprises the steps: obtaining an original data set through a high-activity standard sample and a blank sample which are measured by a detection system, and carrying out the preprocessing of the original data set, and obtaining a standard data set; constructing a random forest classification model according to the standard data set, training the random forest classification model, and optimizing the random forest classification model according to a training error; and performing background event discrimination on a to-be-tested sample by using the optimized classification model, performing time sequence relationship analysis on the to-be-tested sample and the standard data set, and outputting a discrimination result. The method not only can improve the precision of background event discrimination, but also has good interpretability, and can be directly applied to a background event discrimination system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of event discrimination, and in particular to a background event discrimination method based on a random forest classification model. Background Art

[0002] During the low-background liquid scintillation measurement process, in order to achieve accurate measurement of lower-activity samples, it is necessary to further reduce the background level of the detection system while maintaining a high detection efficiency as much as possible. Traditional methods for reducing background include using low-background liquid scintillation detectors with better passive shielding performance, and reducing the dark current noise of photomultiplier tubes and background events caused by cosmic rays or environmental γ-rays through time coincidence and anticoincidence techniques. In addition, some pulse discrimination methods based on signal processing are also widely used, such as screening abnormal pulse signals by using the ratio of pulse length gate integration and the ratio of pulse amplitudes of multiple channels, or discriminating background events by counting the number of after-pulses within a certain time after the pulse backend.

[0003] In recent years, the development of high-speed digitizers has made it possible to acquire the full waveforms of fast liquid scintillation signals. During the acquisition process, the full waveform points of each liquid scintillation pulse signal are acquired and the data is sent to a computer via USB. With the powerful computing ability of the computer, machine learning methods can be used for further classification of liquid scintillation pulse signals. Based on methods such as feature dimensionality reduction in machine learning, all features of each pulse can be effectively utilized to achieve a better discrimination effect for background events.

[0004] Existing methods are effective to a certain extent, but they still have limitations. For example, the improvement of the construction materials of the detection system, such as increasing the thickness of the external lead shield, although it can further reduce the background, this improvement has a marginal effect and will significantly increase the construction cost of the system. In addition, the discrimination ability of existing pulse discrimination methods based on signal processing is limited. During the process, when relatively loose parameters are selected, the background discrimination effect is limited, but when more stringent parameters are selected, more real events are misdiscriminated as background events, resulting in the loss of real events, thereby reducing the detection efficiency of the system, and ultimately the ability to measure lower-activity samples cannot be improved well. Summary of the Invention

[0005] The object of the present invention is to provide a background event discrimination method based on a random forest classification model.

[0006] To achieve the above object, the present invention is implemented according to the following technical solutions: The present invention includes the following steps: By measuring high-activity standard samples and blank samples through a detection system, an original data set is obtained, and the original data set is preprocessed to obtain a standard data set; Construct a random forest classification model based on the standard data set, train the random forest classification model, and optimize the random forest classification model according to the training error; including: Perform feature dimensionality reduction on the standard data set through principal component analysis, calculate the feature matrix, perform eigenvalue decomposition on the eigenvalues, and obtain the eigenvalues and eigenvectors. The expression is: Among them, the standardized standard data matrix is Z, and the transpose matrix of the standard data matrix Z is , the eigenvector matrix is , the eigenvector matrix The transpose matrix of is , each column represents an eigenvector, the diagonal matrix is , the elements on the diagonal are eigenvalues, the decomposition matrix is S, and the number of eigenvalues is n; Extract eigenvalues and eigenvectors according to the decomposition matrix, sort the eigenvalues in descending order, select the first k largest eigenvalues and the corresponding eigenvectors, construct a projection matrix with the first k eigenvectors, project the original data onto the principal components, and obtain the dimensionality-reduced data. The expression is: Among them, the k-th eigenvector is , the first eigenvector is , the second eigenvector is , the projection matrix is W, and the training data matrix is Y; Calculate the proportion of explained variance: Among them, the i-th proportion of explained variance is , the i-th eigenvalue is , the number of standard data features is ; Determine the information volume of the dimensionality-reduced data retaining the original data according to the proportion of explained variance, and select the principal component features with an explained variance greater than 95% to establish a random forest-based classification model; Use the optimized classification model to perform background event discrimination on the sample to be tested, perform time series relationship analysis on the sample to be tested and the standard data set, and output the discrimination result.

[0007] Furthermore, the method for training the random forest classification model includes: The dimensionality-reduced standard dataset is split into a training dataset and a validation dataset according to a ratio of 4:1. The classification model is trained using the training dataset, and the performance of the trained classification model is evaluated using the validation dataset. During the training process of the classification model, the prediction results are integrated by constructing multiple decision trees. When constructing each decision tree, a random number of samples are drawn from the training data according to the principle of random sampling with replacement to form a new training set. When constructing the nodes of the decision tree, k features are randomly selected from m features, and the number of selected features is calculated: where the number of selected features is k, the initial number of features is m, and the initial feature with the minimum Gini index is used as the splitting criterion to select the best splitting point. The expression of the Gini index is: where the probability of class i is , the number of classes is , and the Gini index is ; Recursively repeat the selection of the best splitting point for the child nodes until the tree reaches the maximum depth, then stop the recursion. The voting method is used to select the prediction results of the majority of decision trees as the prediction result, and the expression is: where the prediction result of the T-th decision tree is , the prediction result of the first decision tree is , the prediction result of the second decision tree is , and the number of decision trees is T.

[0008] Furthermore, a method for optimizing the random forest classification model according to the training error includes: Obtain the training dataset and the validation dataset for training the random forest classification model, and dynamically adjust the value range of the hyperparameters according to the scale and feature dimension of the training dataset. Given the optimization objective function, the expression is: where the hyperparameter search space is , the objective function of the hyperparameter search space is , the total number of samples in the training data is , the prediction result of the a-th training data is , the true label of the a-th training data is , and the true label of the a-th training data in class i is , the prediction result of the a-th training data on class i is , the number of classes is , the loss function is , the weight coefficients are respectively 、 、 , the F1 score is F, and the accuracy is X; Adopt a Gaussian process regression model based on a Gaussian kernel as a surrogate model. According to the training data set, take the acquisition function as the expectation, and obtain the optimal hyperparameter combination for the improved hyperparameters. The expression is: where the best hyperparameter combination is , the best hyperparameter combination includes the best number of decision trees and the best maximum depth. The hyperparameter search space is H, and the c-th hyperparameter combination is , the j-th hyperparameter combination is , the hyperparameter combination and the hyperparameter combination have a kernel function of , the length scale is b, and the standard deviation of the hyperparameter combination and the hyperparameter combination is ; Based on the training data set, take the weighted sum of the F1 score and the accuracy as the validation threshold for performance evaluation. Select the test data set to obtain a data subset, and calculate the performance evaluation value of the random forest classification model through the data subset. When the performance evaluation value is greater than the validation threshold, update the validation threshold; otherwise, trigger negative feedback. When negative feedback is triggered, add the data subset to the training data set and adjust the hyperparameter search space. The expression is: where the adjusted search space for the number of decision trees is , the adjusted search space for the maximum depth is , the index of the number of decision trees sorted in non-decreasing order is u, the index of the maximum depth sorted in non-decreasing order is a, the number of decision trees is e, and the maximum depth is , the index of the adaptively adjusted number of decision trees is , the index of the adaptively adjusted maximum depth is ; Recalculate the current performance evaluation value. If it is greater than the validation threshold, traverse the test set and output the hyperparameter search space; otherwise, increase the number of data subsets and adjust the hyperparameter search space.

[0009] Further, a method for analyzing the time series relationship between the sample to be measured and the standard data set includes: Sort the sample to be measured and the standard data set according to time, use the sample to be measured as the experimental group, use the standard data set as the control group, and perform an initial sliding window division on the experimental group and the control group; Calculate the mean and variance of the sliding data window data sets of the experimental group and the control group, and perform standardization processing on the window data; Establish a slow feature analysis model for the window, and perform feature mapping and dimensionality reduction on the data of the experimental group and the control group; Perform correlation analysis on the slow feature analysis models of the experimental group and the control group to generate a correlation spectrum diagram and a clustering spectrum diagram, and judge the relevant data blocks according to the correlation spectrum diagram and the clustering spectrum diagram; Convert the window data into a data classification result. In the correlation spectrum diagram, large continuous windows with a correlation greater than 0.759 are steady-state data. Then, re-establish a slow feature analysis model according to the steady-state data to obtain a steady-state slow feature analysis monitoring module; if the correlation is less than or equal to 0.759, set the short sliding window size, sliding step length, and correlation threshold; Calculate the correlation between the short sliding window and the sliding window. If the correlation is greater than the correlation threshold, the short sliding window and the sliding window are correlated, and calculate the window merging length: where the sliding step length is and the correlation between the short sliding window and the sliding window is and the window merging length is ; Merge the first g data of the increased window into the short sliding window, update the short sliding window to obtain the current window, and calculate the correlation between the current window and the sliding window; if the correlation is less than or equal to the correlation threshold, divide the short sliding window into sub-stages, and establish a slow feature analysis model according to the sub-stages; Divide the unmodeled window data according to the initial window length and the initial sliding step length to obtain data blocks. Establish a synchronous breeding tree for the first sub-stage and the last s data of the previous steady-state data, and obtain the correlation relationship rules between the experimental group and the control group according to the synchronous breeding tree; Establish a synchronous breeding tree for the last sub-stage and the last s data of the next steady-state data, and obtain the correlation relationship rules between the experimental group and the control group according to the synchronous breeding tree.

[0010] The beneficial effects of the present invention are: The present invention is a background event discrimination method based on a random forest classification model. Compared with the prior art, the present invention has the following technical effects: Through steps of preprocessing, constructing a random forest classification model, training the random forest classification model, and optimizing the model, in order to improve the measurement ability of the existing low-background liquid scintillation analyzer for lower-activity samples without significantly increasing the system construction cost, it is necessary to seek a more effective background reduction strategy while ensuring the detection efficiency. Through machine learning technology, real signals and background events can be more accurately distinguished, thereby further reducing the background level of the system without significantly increasing the system construction cost and enhancing the measurement ability of the system for lower-activity samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of the steps of a method for discriminating background events based on a random forest classification model of the present invention; Figure 2 It is a comparison graph of the energy spectrum of the background of the method of the embodiment of this specification and the prior art. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The present invention will be further described below through specific embodiments. The illustrative embodiments and explanations of this invention are used to explain the present invention, but do not limit the present invention.

[0013] A method for discriminating background events based on a random forest classification model of the present invention includes the following steps: As Figure 1 shown, in this embodiment, it includes the following steps: By measuring high-activity standard samples and blank samples through a detection system, an original data set is obtained, and the original data set is preprocessed to obtain a standard data set; In actual evaluation, a high-speed digital acquisition instrument with a sampling rate of 500M per second and model DT5730S of a certain company is used as the experimental acquisition instrument; On a low-background liquid scintillation device through a high-speed digital acquisition instrument, a 3H sample with a standard activity of 100 Bq and a blank sample are measured for 10 minutes, obtaining 5000 labeled real event data and 5000 labeled background event data sets. The 10,000 waveform data sets are preprocessed, and the finally preprocessed data is used as the standard data set; among them, the waveform data set: the original data set has a dimension of 10000*360, a total of 10,000 samples, and each sample has 360 features; Perform dimensionality reduction processing on the standard data set, calculate how much proportion of the features can be explained by each principal component after dimensionality reduction, and find that the first 8 features can explain a variance proportion of more than 95%. Finally, the first 8 features after dimensionality reduction are selected to form a new data set, where the new data set has a dimension of 10000*8, a total of 10,000 samples, and each sample has 8 features; Construct a random forest classification model based on the standard data set, train the random forest classification model, and optimize the random forest classification model according to the training error; including: Perform feature dimensionality reduction on the standard data set through principal component analysis, calculate the feature matrix, perform eigenvalue decomposition on the eigenvalues, and obtain the eigenvalues and eigenvectors. The expression is: where the standardized standard data matrix is Z, and the transpose matrix of the standard data matrix Z is , the eigenvector matrix is , the eigenvector matrix The transpose matrix of is , each column represents an eigenvector, the diagonal matrix is , the elements on the diagonal are the eigenvalues, the decomposition matrix is S, and the number of eigenvalues is n; Extract the eigenvalues and eigenvectors according to the decomposition matrix, sort the eigenvalues in descending order, select the first k largest eigenvalues and the corresponding eigenvectors, construct a projection matrix with the first k eigenvectors, project the original data onto the principal components, and obtain the dimensionality-reduced data. The expression is: where the k-th eigenvector is , the first eigenvector is , the second eigenvector is , the projection matrix is W, and the training data matrix is Y; Calculate the proportion of explained variance: where the i-th proportion of explained variance is , the i-th eigenvalue is , the number of standard data features is ; Determine the information volume of the dimensionality-reduced data retaining the original data according to the proportion of explained variance, and select the principal component features with an explained variance greater than 95% to establish a random forest-based classification model; Use the optimized classification model to perform background event discrimination on the sample to be tested, perform time series relationship analysis on the sample to be tested and the standard data set, and output the discrimination result.

[0014] In this embodiment, the method for training the random forest classification model includes: The dimensionality-reduced standard dataset is split into a training dataset and a validation dataset according to a ratio of 4:1. The classification model is trained using the training dataset, and the performance of the trained classification model is evaluated using the validation dataset. During the training process of the classification model, the prediction results are integrated by constructing multiple decision trees. When constructing each decision tree, a random number of samples are randomly drawn from the training data according to the principle of sampling with replacement to form a new training set. When constructing the nodes of the decision tree, k features are randomly selected from m features, and the number of selected features is calculated: where the number of selected features is k, the initial number of features is m, and the initial feature with the minimum Gini index is used as the splitting criterion to select the best split point. The expression for the Gini index is: where the probability of class i is , the number of classes is , and the Gini index is ; Recursively repeat the selection of the best split point for the child nodes until the tree reaches the maximum depth, and then stop the recursion. The voting method is used to select the prediction results of the majority of decision trees as the prediction result, and the expression is: where the prediction result of the T-th decision tree is , the prediction result of the first decision tree is , the prediction result of the second decision tree is , and the number of decision trees is T; In the actual evaluation, the prediction results are integrated by constructing multiple decision trees. When constructing each decision tree, a random number of samples are randomly drawn from the training data with replacement to form a new training set. When constructing each node of the decision tree, 3 features are randomly selected from 8 features, and the best split point is selected according to the selected features and the Gini index as the splitting criterion. The above process is recursively repeated for each child node until the stop condition is met, and the voting method is used to select the prediction results of the majority of decision trees as the prediction result.

[0015] In this embodiment, the method for optimizing the random forest classification model according to the training error includes: Obtain the training dataset and the validation dataset for training the random forest classification model, and dynamically adjust the value range of the hyperparameters according to the scale and feature dimension of the training dataset. Given the optimization objective function, the expression is: where the hyperparameter search space is , the hyperparameter search space has an objective function of , the total number of samples of the training data is , the prediction result of the a-th training data is , the true label of the a-th training data is , the true label of the a-th training data on class i is , the prediction result of the a-th training data on class i is , the number of classes is , the loss function is , the weight coefficients are respectively , , , the F1 score is F and the accuracy is X; Adopt a Gaussian process regression model based on a Gaussian kernel as the surrogate model. According to the training data set, take the acquisition function as the expectation and obtain the optimal hyperparameter combination for the improved hyperparameters. The expression is: where the best hyperparameter combination is , the best hyperparameter combination includes the best number of decision trees and the best maximum depth. The hyperparameter search space is H, and the c-th hyperparameter combination is , the j-th hyperparameter combination is , the hyperparameter combination and the hyperparameter combination have a kernel function of , the length scale is b, and the standard deviation of the hyperparameter combination and the hyperparameter combination is ; Based on the training data set, use the weighted sum of the F1 score and the accuracy as the validation threshold for performance evaluation. Select the test data set to obtain a data subset, and calculate the performance evaluation value of the random forest classification model through the data subset. When the performance evaluation value is greater than the validation threshold, update the validation threshold; otherwise, trigger negative feedback; When negative feedback is triggered, add the data subset to the training data set and adjust the hyperparameter search space. The expression is: where the adjusted search space for the number of decision trees is , and the adjusted search space for the maximum depth is , the index after non-decreasing sorting of the number of decision trees is u, the index after non-decreasing sorting of the maximum depth is a, the number of decision trees is e, and the maximum depth is , the index of the adaptively adjusted number of decision trees is , the index of the adaptively adjusted maximum depth is ; Recalculate the current performance evaluation value. If it is greater than the verification threshold, traverse the test set and output the hyperparameter search space; otherwise, increase the number of data subsets and adjust the hyperparameter search space; In the actual evaluation, then optimize the hyperparameters of the trained random forest classification model, mainly the optimization of the number of decision trees and the number of samples per node from 1 to 5. By trying all different parameter combinations, obtain the best hyperparameters and solidify the model with the best prediction effect on the verification dataset; among them, the number of decision trees ranges from 100 to 1000 with a step size of 100; The solidified random forest classification model is used for the background discrimination of the samples to be tested, and the events identified as background events are excluded.

[0016] In this embodiment, the method for analyzing the time series relationship between the sample to be tested and the standard dataset includes: Sort the sample to be tested and the standard dataset according to time. Take the sample to be tested as the experimental group and the standard dataset as the control group, and perform an initial sliding window division on the experimental group and the control group; Calculate the mean and variance of the sliding data window datasets of the experimental group and the control group, and perform standardization processing on the window data; Establish a slow feature analysis model for the window, and perform feature mapping and dimensionality reduction on the data of the experimental group and the control group; Perform correlation analysis on the slow feature analysis models of the experimental group and the control group, generate correlation spectrograms and clustering maps, and judge the relevant data blocks according to the correlation spectrograms and clustering maps; Convert the window data into data classification results. In the correlation map, the large continuous windows with a correlation greater than 0.759 are steady-state data. Then, re-establish a slow feature analysis model according to the steady-state data to obtain a steady-state slow feature analysis monitoring module; if the correlation is less than or equal to 0.759, set the short sliding window size, sliding step length, and correlation threshold; Calculate the correlation between the short sliding window and the sliding window. If the correlation is greater than the correlation threshold, the short sliding window and the sliding window are correlated, and calculate the window merging length: where the sliding step length is , the correlation between the short sliding window and the sliding window is , the window merging length is ; Merge the first g data of the increasing window into the short sliding window, update the short sliding window to obtain the current window, and calculate the correlation between the current window and the sliding window; if the correlation is less than or equal to the correlation threshold, divide the short sliding window into sub-phases, and establish a slow feature analysis model according to the sub-phases; Divide the unmodeled window data according to the initial window length and the initial sliding step to obtain data blocks. Establish a synchronous breeding tree with the first sub-phase and the last s data of the previous steady-state data, and obtain the correlation rule between the experimental group and the control group according to the synchronous breeding tree; Establish a synchronous breeding tree with the last sub-phase and the last s data of the next steady-state data, and obtain the correlation rule between the experimental group and the control group according to the synchronous breeding tree; In the actual evaluation, the correlation threshold is 0.767.

[0017] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A background event identification method based on a random forest classification model, characterized in that: The following steps are involved: A high-activity standard sample and a blank sample are measured by a detection system to obtain an original data set, and the original data set is preprocessed to obtain a standard data set; Constructing a random forest classification model according to the standard data set, training the random forest classification model, and optimizing the random forest classification model according to the training error; comprising: The standard data set is reduced in dimension through principal component analysis, the characteristic matrix is ​​calculated, and the eigenvalues ​​are decomposed to obtain the eigenvalues ​​and eigenvectors. The expression is: The standardized standard data matrix is ​​Z, and the transposed matrix of the standard data matrix Z is , the eigenvector matrix is , the eigenvector matrix The transposed matrix of , each column represents a eigenvector, and the diagonal matrix is , the elements on the diagonal are eigenvalues, the decomposition matrix is ​​S, and the number of eigenvalues ​​is n; Extract eigenvalues ​​and eigenvectors according to the decomposition matrix, sort the eigenvalues ​​in descending order, select the first k largest eigenvalues ​​and corresponding eigenvectors, use the first k eigenvectors to construct a projection matrix, project the original data onto the principal component, and obtain the reduced-dimensional data. The expression is: The kth eigenvector is , the first eigenvector is , the second eigenvector is , the projection matrix is ​​W, and the training data matrix is ​​Y; Calculate the proportion of variance explained: The i-th explained variance ratio is , the i-th eigenvalue is , the number of standard data features is ; The amount of information retained in the original data by the dimension-reduced data was determined based on the explained variance ratio, and the principal component features with explained variance greater than 95% were selected to establish a classification model based on random forests; The optimized classification model is used to identify background events of the sample to be tested, and the time series relationship analysis is performed on the sample to be tested and the standard data set, and the identification result is output.

2. A background event identification method based on a random forest classification model according to claim 1, characterized in that: The method for training the random forest classification model comprises: The standard data set after dimensionality reduction is split into a 4:1 ratio to obtain a training data set and a validation data set. The training data set is used to train the classification model, and the validation data set is used to evaluate the performance of the trained classification model. In the classification model training process, the prediction results are integrated by building multiple decision trees. When building each decision tree, a random number of samples are extracted from the training data according to the principle of random sampling with replacement to form a new training set. When constructing the nodes of the decision tree, k features are randomly selected from m features, and the number of selected features is calculated: The number of features is selected as k, the number of initial features is m, and the initial feature with the smallest Gini index is used as the splitting criterion to select the best splitting point. The expression of the Gini index is: The probability of category i is , the number of categories is , the Gini index is ; Recursively select the best split point for the child nodes until the tree reaches the maximum depth, then stop the recursion; The voting method is used to select the prediction results of the majority decision trees as the prediction results. The expression is: The prediction result of the Tth decision tree is , the prediction result of the first decision tree is , the prediction result of the second decision tree is , the number of decision trees is T.

3. The background event identification method based on the random forest classification model according to claim 1, characterized in that: The method for optimizing the random forest classification model according to the training error comprises: Obtain the training data set and validation data set for training the random forest classification model, and dynamically adjust the value range of the hyperparameters according to the scale and feature dimension of the training data set; Given the optimization objective function, the expression is: The hyperparameter search space is , the hyperparameter search space The objective function is The total number of samples of training data is , the prediction result of the ath training data is , the true label of the a-th training data is , the true label of the a-th training data in category i is , the prediction result of the a-th training data on category i is , the number of categories is , the loss function is The weight coefficients are , , , F1 score is F, accuracy is X; The Gaussian process regression model based on Gaussian kernel is used as the surrogate model. The acquisition function is taken as the expectation according to the training data set, and the improved hyperparameters obtain the optimal hyperparameter combination, which is expressed as: The best hyperparameter combination is , the best hyperparameter combination includes the optimal number of decision trees and the optimal maximum depth. The hyperparameter search space is H, and the cth hyperparameter combination is , the jth hyperparameter combination is , hyperparameter combination and hyperparameter combinations The kernel function is , length scale is b, hyperparameter combination and hyperparameter combinations The standard deviation of ; Based on the training data set, the weighted sum of the F1 score and accuracy is used as the validation threshold for performance evaluation. The test data set is selected to obtain a data subset. The performance evaluation value of the random forest classification model is calculated through the data subset. When the performance evaluation value is greater than the validation threshold, the validation threshold is updated, otherwise negative feedback is triggered. When negative feedback is triggered, a data subset is added to the training data set and the hyperparameter search space is adjusted. The expression is: The search space for the number of adjusted decision trees is , the adjusted maximum depth search space is , the index of the non-decreasing sorted number of decision trees is u, the index of the non-decreasing sorted maximum depth is a, the number of decision trees is e, and the maximum depth is , the number of decision trees adjusted adaptively is indexed as , the maximum depth index of adaptive adjustment is ; Recalculate the current performance evaluation value. If it is greater than the verification threshold, traverse the test set and output the hyperparameter search space; otherwise, increase the number of data subsets and adjust the hyperparameter search space.

4. The background event identification method based on the random forest classification model according to claim 1, characterized in that: The method for performing time series relationship analysis on the sample to be tested and the standard data set comprises: Sort the samples to be tested and the standard data set by time, use the samples to be tested as the experimental group and the standard data set as the control group, and perform initial sliding window division on the experimental group and the control group; Calculate the mean and variance of the sliding data window data set of the experimental group and the control group, and standardize the window data; A slow feature analysis model was established for the window, and feature mapping and dimensionality reduction were performed on the data of the experimental and control groups; Perform correlation analysis on the slow feature analysis models of the experimental group and the control group, generate correlation spectra and clustering spectra, and determine relevant data blocks based on the correlation spectra and clustering spectra; The window data is converted into data classification results. In the correlation map, the long continuous window with a correlation greater than 0.759 is steady-state data. The slow feature analysis model is re-established based on the steady-state data to obtain a steady-state slow feature analysis monitoring module. If the correlation is less than or equal to 0.759, the short sliding window size, sliding step size and correlation threshold are set. Calculate the correlation between the short sliding window and the sliding window. If the correlation is greater than the correlation threshold, the short sliding window and the sliding window are correlated. Calculate the window merging length: The sliding step length is , the correlation between the short sliding window and the sliding window is , the window merging length is ; The first g data of the added window are merged into the short sliding window, the short sliding window is updated to obtain the current window, and the correlation between the current window and the sliding window is calculated; if the correlation is less than or equal to the correlation threshold, the short sliding window is divided into sub-stages, and a slow feature analysis model is established according to the sub-stages; The unmodeled window data is divided according to the initial window length and the initial sliding step length to obtain data blocks, and a synchronous propagation tree is established for the last s data of the first sub-stage and the previous steady-state data. The correlation rules between the experimental group and the control group are obtained according to the synchronous propagation tree; The last s data of the last sub-stage and the next steady-state data are used to establish a synchronous propagation tree, and the correlation rules between the experimental group and the control group are obtained based on the synchronous propagation tree.

Citation Information

Patent Citations

  • Background coincidence event judging and selecting method, device and equipment and readable storage medium

    CN112102426A

  • A parameter selection optimization method, system and equipment in random forest model training

    CN113591944A

  • Kernel event screening method based on CatBoost model

    CN117892232A

  • Intelligent access control management method and system based on multi-mode identification and Internet of Things technology

    CN118968665A

  • Active signal detection using adaptive identification of a noise floor

    US20190293769A1