Laboratory detection data processing method and system based on machine learning technology
Through dynamic noise reduction and multimodal feature fusion, combined with dual-model collaboration and environmental compensation, the limitations of noise processing and feature extraction in laboratory data analysis systems are solved, and high-precision and robust intelligent analysis are achieved, which improves the reliability and timeliness of detection results.
Patent Information
- Application Number
- CN202510838746.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing laboratory data analysis system relies on fixed thresholds when processing multi-source heterogeneous signals, making it difficult to adapt to the dynamic noise environment. Feature extraction is limited to a single mode, abnormal detection is separated from quantitative regression models, and environmental parameter interference cannot respond in real time, resulting in reduced analysis accuracy and insufficient reliability.
Through dynamic noise reduction, multimodal feature fusion, dual-model collaboration and environmental compensation methods, wavelet transformation, sliding window, SIFT algorithm, graph convolution network and LSTM network are used to build anomaly detection and quantitative regression models to achieve signal quality improvement, feature screening and environmental calibration.
It significantly improves the reliability and timeliness of laboratory test results, improves signal-to-noise ratio, reduces feature redundancy, enhances abnormal detection accuracy and environmental interference resistance.
Smart Images

Figure CN120337110A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing. Specifically, it particularly relates to a method and system for processing laboratory test data based on machine learning technology. Background Art
[0002] In high-end laboratory scenarios such as biomedicine, environmental monitoring, and materials science, the data generated by instruments presents multi-source heterogeneity (such as spectrometer waveforms, chromatograph values, microscopic images, etc.), and is vulnerable to interference from temperature, humidity, and voltage fluctuations. Traditional analysis methods rely on manual experience for segmented processing: first denoising, then feature extraction, and finally single-model prediction, resulting in large information loss, lagging process response, and serious cumulative errors in environmental interference. The industry still needs an intelligent analysis system that can automatically fuse multi-modal data, resist noise in real time, and collaboratively optimize anomaly filtering and environmental compensation.
[0003] Existing laboratory data analysis systems have significant defects: the noise processing of multi-source heterogeneous signals relies on fixed-threshold denoising methods, which are difficult to adapt to dynamic noise environments; feature extraction is limited to a single modality and lacks the ability of multi-modal fusion; anomaly detection and quantitative regression models are split in training, resulting in regression results contaminated by abnormal samples; environmental parameter interference is only processed through static compensation and cannot respond to complex environmental changes in real time, ultimately resulting in a decline in analysis accuracy and insufficient reliability. Summary of the Invention
[0004] (I) Technical Problems to be Solved In view of the problems in the related art, the present invention provides a method for processing laboratory test data based on machine learning technology to overcome the above-mentioned technical problems existing in the existing related technologies.
[0005] (II) Technical Solutions To solve the above technical problems, the present invention is realized through the following technical solutions: S1. Collect multi-source signal data and environmental parameter data of the laboratory; perform dynamic denoising processing on the multi-source signal data of the laboratory to obtain preprocessed multi-source data of the laboratory; S2. Extract frequency-domain features and spatial features from the preprocessed multi-source data of the laboratory, and perform fusion to obtain multi-source features of the laboratory; Perform mutual information feature selection processing on the multi-source features of the laboratory to obtain effective laboratory features; S3. Construct an anomaly detection classification model and a quantitative regression model; train the anomaly detection classification model and the quantitative regression model respectively to obtain a final anomaly detection classification model and a final quantitative regression model; S4. Use the final anomaly detection classification model and the final quantitative regression model to perform dual-model collaborative analysis on the effective laboratory features to obtain an anomaly detection result and a regression prediction value; Eliminate the abnormal data in the regression prediction values according to the anomaly detection results to obtain normal regression prediction values; S5. Obtain the predicted compensation amount according to the environmental parameter data in combination with the LSTM network; obtain the calibrated laboratory data according to the normal regression prediction values in combination with the predicted compensation amount; Through the full-chain innovation of dynamic noise reduction, multi-modal fusion, feature screening, dual-model collaboration, and environmental compensation, the present invention overcomes the technical problems of multi-source data being sensitive to noise, insufficient utilization of features, isolated prediction of models, and cumulative environmental interference, providing a high-precision, strong-robustness, and adaptive intelligent analysis solution for the laboratory, and significantly improving the reliability and timeliness of detection results.
[0006] Preferably, S1 includes the following steps: S11. Obtain the multi-source signal data of the laboratory in real time through the industrial bus or API interface; the multi-source signal data of the laboratory includes structured signal data and unstructured signal data; Synchronously collect the environmental parameter data to obtain the environmental parameter data; S12. Perform dynamic noise reduction processing on the multi-source signal data of the laboratory to obtain the preprocessed multi-source data of the laboratory; The present invention collects the structured data and unstructured data of the laboratory through the industrial bus / API dual track, synchronously obtains environmental parameters such as temperature and humidity, and realizes the unified access of multi-source heterogeneous data; adopts dynamic noise reduction processing, adaptively filters out noise and retains the edge features of effective signals, improves the signal-to-noise ratio, and provides a high-fidelity data basis for subsequent analysis.
[0007] Preferably, S12 includes the following steps: S121. Perform wavelet transform on the multi-source signal data of the laboratory to obtain a wavelet coefficient set; the wavelet coefficients include approximation coefficients and detail coefficients; S122. Calculate the noise standard deviation according to the wavelet coefficient set; S123. Calculate the real-time dynamic threshold according to the noise standard deviation; S124. Process each wavelet coefficient in the wavelet coefficient set through a soft threshold function according to the real-time dynamic threshold to obtain a filtered wavelet coefficient set; S125. Perform signal reconstruction on the wavelet signals in the filtered wavelet coefficient set through wavelet inverse transform processing to obtain the preprocessed multi-source data of the laboratory; The present invention decomposes the signal into approximation / detail coefficients through wavelet transform, calculates the noise standard deviation based on the median of the highest-frequency detail coefficients, dynamically generates a threshold to adapt to the signal length, intelligently shrinks the noise coefficients with a soft threshold function, and reconstructs a high-fidelity signal through wavelet inverse; adaptively filters out noise while retaining the edge features of effective signals, improves the signal-to-noise ratio, and avoids the over-smoothing or under-denoising problems caused by fixed thresholds.
[0008] Preferably, S2 includes the following steps: S21. Calculate the local statistics of the time series signal data in the preprocessed laboratory multi-source data through a sliding window to obtain frequency domain features; Adopt the SIFT algorithm to extract the key point descriptors of the spatial data in the preprocessed laboratory multi-source data to obtain spatial features; The frequency domain features and spatial features together constitute the laboratory multi-source features; S22. Calculate the mutual information score set between the features of the preprocessed laboratory multi-source data and the target variable; Set the retention percentage; retain the effective features in the mutual information score set according to the retention percentage to obtain effective laboratory features; The present invention extracts the local statistics of the time series signal and the FFT frequency domain features through a sliding window, combines the SIFT algorithm to capture the key point descriptors of the spatial data, and realizes the time series-spatial multi-modal feature fusion; calculates the mutual information score between the features and the target variable, and retains the highly correlated features; through in-depth mining of the dynamic law of the signal and the spatial structure information, mutual information screening reduces the feature redundancy and improves the model training efficiency and representation accuracy.
[0009] Preferably, S3 includes the following steps: S31. Based on the graph convolutional network and the attention mechanism, construct an anomaly detection classification model; set the weight parameters of the anomaly detection classification model; Based on XGBoost and LightGBM, construct a quantitative regression model; the quantitative regression model adopts the dual-model weighted fusion strategy of XGBoost and LightGBM; S32. Collect historical laboratory feature data; S33. Input the historical laboratory feature data into the anomaly detection classification model for training; During the training process, combine the random search and cross-validation algorithms to find the optimal weight parameters of the anomaly detection classification model to obtain the optimal solution; use the optimal solution as the weight parameters of the anomaly detection classification model to obtain the final anomaly detection classification model; S34. Input the historical laboratory feature data into the quantitative regression model for training, and optimize the parameters of the quantitative regression model by combining grid search to obtain the final quantitative regression model; The present invention improves the analysis ability through a hybrid model architecture and an intelligent optimization strategy; realizes anomaly detection by accurately capturing data correlation features through a graph convolutional network + attention mechanism, and enhances the robustness of regression prediction through the weighted fusion of XGBoost and LightGBM; the graph convolutional network processes topological data, the attention mechanism focuses on key features, and the combination of XGBoost + LightGBM takes into account both speed and accuracy; combines random search cross-validation to optimize the weights of the classification model and grid search to tune the parameters of the regression model; random search avoids grid calculation explosion, and cross-validation prevents overfitting to ensure the reliability of the optimal weights; directly improves the efficiency of automated analysis of laboratory data, reduces manual intervention errors; significantly improves the sensitivity of anomaly recognition and the accuracy of quantitative prediction, reduces the cost of manual parameter tuning, is especially suitable for complex and high-dimensional laboratory data scenarios, and improves the generalization ability of the model.
[0010] Preferably, during the training process in S33, a random search and cross-validation algorithm are combined to find the optimal weight parameters of the anomaly detection classification model, and obtaining the optimal solution includes the following steps: S331. Set the search space of the weight parameters of the anomaly detection classification model as p, the maximum number of random searches as m, and the number of cross-validation folds as W; S332. Randomly set the weight parameters of the anomaly detection classification model according to the search space p of the weight parameters to obtain an initial weight parameter set; Initialize the optimal weight parameters of the anomaly detection classification model as u and the best performance as v, and randomly select the weight parameters of the anomaly detection classification model from the initial weight parameter set as e; S333. Divide the training data in the historical laboratory feature data into W folds. For each fold w, ; use the training data in all historical laboratory feature data except the w-th fold to train the anomaly detection classification model, and use the training data in the w-th fold of the historical laboratory feature data to evaluate the anomaly detection classification model to obtain a performance index set; Calculate the mean of the performance indexes of all folds to obtain the average performance r, If the average performance r > the best performance v, update the best performance to r and update the optimal weight parameters to e; otherwise, keep the original best performance and optimal weight parameters; S334. Repeat S333. When the maximum number of random searches m is reached, stop the iteration to obtain the optimal solution; The present invention optimizes the anomaly detection model by combining random search and cross-validation; efficiently samples weight combinations in the parameter space to avoid the computational overhead of grid search; divides the data into W folds for multiple training and evaluation to ensure the statistical stability of performance metrics and reduce the risk of overfitting; continuously compares the average performance with the best performance, and finally converges to the optimal weights; balances efficiency and robustness in a limited number of searches, significantly improving the generalization ability of the model, especially suitable for scenarios with limited laboratory data volume and the need to avoid evaluation bias.
[0011] Preferably, the S4 includes the following steps: S41. Input the effective laboratory features into the final anomaly detection classification model and the final quantitative regression model respectively to obtain the anomaly detection results and regression prediction values; S42. Set the anomaly probability threshold; mark the effective laboratory features with an anomaly probability ≥ the anomaly probability threshold in the anomaly detection results as abnormal data; Eliminate the regression prediction values corresponding to the abnormal data in the regression prediction values to obtain the normal regression prediction values; The present invention improves the analysis reliability through the cooperation of dual models; uses the classification model to identify abnormal data (probability ≥ threshold) and accurately mark abnormal samples; eliminates the regression prediction values corresponding to the abnormal data and retains the normal regression prediction values; avoids the interference of abnormal values on the regression results, and improves the accuracy and stability of quantitative prediction, especially suitable for scenarios with fluctuating laboratory data quality.
[0012] Preferably, the S5 includes the following steps: S51. Construct an LSTM network; set the calibration rule; update the LSTM network according to the calibration rule; S52. Input the environmental parameter data into the LSTM network to obtain the prediction compensation amount; S53. Calculate the calibrated laboratory data according to the normal regression prediction values combined with the prediction compensation amount; The present invention realizes dynamic calibration through LSTM time series modeling; uses LSTM to learn the influence of environmental parameters (such as temperature and humidity) on the data and outputs the prediction compensation amount; combines the normal regression prediction values with the compensation amount to generate the calibrated laboratory data; significantly reduces the systematic error caused by environmental interference, improves the data reliability and the accuracy of experimental conclusions, and is applicable to high-precision environment-sensitive scenarios.
[0013] A processing system for laboratory test data based on machine learning technology, used to implement the above-mentioned processing method for laboratory test data based on machine learning technology, including a multi-source data acquisition and dynamic noise reduction module, a multi-modal feature fusion and selection module, a dual-model collaborative training and optimization module, an anomaly-aware regression prediction module, and an environmental compensation dynamic calibration module; The multi-source data acquisition and dynamic noise reduction module is used for real-time acquisition and preprocessing of multi-source signals and environmental parameters in the laboratory; structured and unstructured data are obtained through industrial buses or API interfaces, and environmental parameters such as temperature, humidity, and voltage are synchronously acquired; the discrete wavelet transform is used to decompose the coefficients, the noise standard deviation is estimated based on the median of the highest frequency coefficients, the threshold is dynamically calculated, and then the noise coefficients are shrunk through the soft threshold function and the signal is reconstructed, and finally the multi-source data after noise reduction is output; The multi-modal feature fusion and selection module is used to extract and fuse frequency domain and spatial features from the preprocessed data; for time series signals, local statistics are calculated through a sliding window and FFT frequency domain features are extracted; for spatial data, the SIFT algorithm is used to extract key point descriptors; the fused multi-source features are optimized through mutual information feature selection, the mutual information scores between the features and the target variables are calculated, the high-score features are retained, the redundant features are removed, and the effective laboratory features are output; The dual-model collaborative training and optimization module is used to construct and train an anomaly detection and quantitative regression dual model; based on the graph convolutional network and the attention mechanism, a graph structure is constructed, the weight parameters are optimized by random search cross-validation, random sampling is performed in the parameter space, the performance is evaluated through W-fold cross-validation, and the optimal weights are iteratively updated to obtain the final anomaly detection classification model; by fusing XGBoost and LightGBM, it is trained in two stages: first, XGBoost is used to predict the base value, and then LightGBM is used to learn the residuals; finally, the weighted fusion coefficient is optimized on the validation set through grid search, and the weight that minimizes the RMSE is selected to obtain the final quantitative regression model; The anomaly perception regression prediction module is used to realize dual-model collaborative inference and anomaly filtering; the effective features are respectively input into the final anomaly detection classification model and the final quantitative regression model, and the anomaly probability and the regression prediction value are output; an anomaly probability threshold is set to mark the anomaly samples; the regression prediction values corresponding to the anomaly samples are removed, and the pure normal regression prediction results are output; The environmental compensation dynamic calibration module eliminates environmental interference based on the LSTM network; the environmental parameters are input into the LSTM, and the prediction compensation amount is calculated through the gating mechanism; an incremental learning mechanism is designed to optimize the LSTM; the normal regression prediction value and the compensation amount are superimposed, and the final laboratory data after environmental calibration is output.
[0014] (III) Beneficial effects The present invention has the following beneficial effects: Through the full-chain innovation of dynamic noise reduction, multi-modal fusion, feature screening, dual-model collaboration, and environmental compensation, the present invention overcomes the technical problems of multi-source data being sensitive to noise, insufficient utilization of features, isolated prediction of models, and cumulative environmental interference, provides a high-precision, strong-robustness, and adaptive intelligent analysis solution for the laboratory, and significantly improves the reliability and timeliness of the detection results.
[0015] Through dynamic adaptive noise reduction, the present invention significantly improves the signal quality; by decomposing signal coefficients through discrete wavelet transform, accurately estimating the noise standard deviation based on the median of the highest-frequency detail coefficients, dynamically calculating the threshold and combining with the soft threshold function to intelligently shrink the noise coefficients, while retaining the edge features of the effective signal, thoroughly filtering out the noise, improving the signal-to-noise ratio of the reconstructed signal, and providing a high-fidelity data basis for subsequent analysis.
[0016] Through multi-modal feature fusion, the present invention comprehensively enhances the information representation ability; for time-series signals, local statistics and FFT frequency-domain features are extracted using a sliding window; for spatial data, key-point descriptors are extracted through the SIFT algorithm; the fused multi-source features are screened by mutual information scoring, retaining highly correlated features, effectively reducing feature redundancy, and greatly improving the model training efficiency and representation accuracy.
[0017] Through a dual-model collaboration mechanism, the present invention achieves anomaly immunity; the anomaly detection model combines a graph convolutional network and an attention mechanism, optimizes parameters through random search and W-fold cross-validation, improving the anomaly classification accuracy; quantitative regression adopts a weighted fusion strategy of XGBoost and LightGBM, and after determining the optimal weights through grid search, reduces the prediction error; according to the anomaly detection results, contaminated samples are directly removed, eliminating the interference of outliers on the regression results from the source.
[0018] Through environmental compensation dynamic calibration, the long-term analysis accuracy is guaranteed; the LSTM network dynamically learns the non-linear relationship between environmental parameters and interference amounts through a gating mechanism, the cell state long-term remembers the environmental change pattern, and outputs accurate prediction compensation amounts; an incremental learning mechanism is designed to optimize the LSTM, solving the problem of cumulative errors caused by laboratory environment drift.
[0019] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the invention, and for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic flow chart of the method for processing laboratory test data based on machine learning technology of the present invention; Figure 2 It is a schematic flow chart of dynamic threshold denoising of laboratory multi-source signal data in the method for processing laboratory test data based on machine learning technology of the present invention; Figure 3 This is a schematic flowchart for obtaining normal regression prediction values in the method for processing laboratory test data based on machine learning technology of the present invention; Figure 4 This is a schematic diagram of the modules of the processing system for laboratory test data based on machine learning technology of the present invention. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the invention. Obviously, the described embodiments are only a part of the embodiments of the invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the invention without making creative efforts belong to the scope of protection of the invention.
[0023] In the description of the present invention, it should be understood that the terms "openings", "upper", "lower", "top", "middle", "inner", etc. indicating orientations or positional relationships are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0024] Embodiment 1: Please refer to Figure 1 , Figure 2 , Figure 3 , the present invention discloses a method for processing laboratory test data based on machine learning technology, including the following steps: S1. Collect multi-source signal data and environmental parameter data in the laboratory; perform dynamic noise reduction processing on the multi-source signal data in the laboratory to obtain preprocessed multi-source data in the laboratory; The S1 includes the following steps: S11. Obtain multi-source signal data in the laboratory in real time through an industrial bus (Modbus / CAN) or an API interface; the multi-source signal data in the laboratory includes structured signal data and unstructured signal data; the structured signal data includes detection values (concentration values, absorbance, etc.), and the unstructured signal data includes spectral waveforms, chromatograms, and mass spectra; Synchronously collect environmental parameter signals to obtain environmental parameter signal data; the environmental parameter temperature data, humidity data, and equipment voltage values; S12. Perform dynamic noise reduction processing on the multi-source signal data in the laboratory to obtain preprocessed multi-source data in the laboratory; The S12 includes the following steps: S121. Perform wavelet transform (such as discrete wavelet transform DWT) on the multi-source signal data in the laboratory to obtain a wavelet coefficient set; the wavelet coefficients include approximation coefficients and detail coefficients; the highest-frequency detail coefficients are mainly composed of noise; S122. Calculate the noise standard deviation based on the wavelet coefficient set. The calculation formula is as follows: ; where σ represents the noise standard deviation; ck represents the highest-frequency wavelet coefficient in the wavelet coefficient set; median(ck) is the median of the highest-frequency wavelet coefficient (used to eliminate the influence of signal components); dividing by 0.6745 is because for normally distributed noise, the relationship between MAD and the standard deviation is MAD / 0.6745; S123. Calculate the real-time dynamic threshold based on the noise standard deviation. The calculation formula is as follows: ; where λ represents the real-time dynamic threshold (the wavelet coefficient for distinguishing noise and useful signals), and N represents the signal length of the laboratory multi-source signal data (the threshold increases with the increase of the signal length to adapt to longer signals); S124. Process each wavelet coefficient in the wavelet coefficient set through a soft threshold function according to the real-time dynamic threshold to obtain a filtered wavelet coefficient set. The formula of the soft threshold function is as follows: ; where y represents the wavelet coefficient in the wavelet coefficient set, and η(y,λ) represents the wavelet coefficient in the wavelet coefficient set after being processed by the soft threshold function; the wavelet coefficient greater than λ is subtracted by λ (shrinking towards zero), the wavelet coefficient less than -λ is added by λ (shrinking towards zero), and the coefficient with an absolute value less than or equal to λ is set to zero (regarded as noise); S125. Reconstruct the signal by performing wavelet inverse transform on the wavelet signal in the filtered wavelet coefficient set to obtain preprocessed laboratory multi-source data; S2. Extract the frequency-domain features and spatial features from the preprocessed laboratory multi-source data, and fuse them to obtain laboratory multi-source features; Perform mutual information feature selection processing on the laboratory multi-source features to obtain effective laboratory features; The S2 includes the following steps: S21. Calculate the local statistics (mean, variance, skewness) of the time-series signal data (such as spectral scanning) in the preprocessed laboratory multi-source data through a sliding window (width = 50 sampling points) to obtain frequency-domain features (such as the main frequency amplitude and power spectrum entropy after FFT transform). The sliding window calculation formula is as follows: ; Among them, μt represents the mean of the t-th window, w represents the window length (i.e., the number of sample points of the time series signal data in the preprocessed laboratory multi-source data), and x represents the value of the time series signal at index i; σt2 represents the variance of the t-th window, 1 / (w - 1) represents the denominator for calculating the unbiased estimate of the variance (using w - 1 instead of w can provide a better estimate of the population variance, especially when the sample size w is not large), and (xi - μt)2 represents the square of the deviation calculated for each sample point from the window mean; The SIFT algorithm is used to extract the key point descriptors of the spatial data (such as cell microscopic images) in the preprocessed laboratory multi-source data, obtaining spatial features; The frequency domain features and spatial features together constitute the laboratory multi-source features; S22. Calculate the mutual information score set between the features of the preprocessed laboratory multi-source data and the target variable; the calculation formula is as follows, ; Among them, X represents the feature variable in the features of the preprocessed laboratory multi-source data, and Y represents the target variable (such as class label, predicted value, etc.); I(X, Y) represents the mutual information score between the feature variable X and the target variable Y (i.e., measures the degree of reduction in the uncertainty of Y after knowing X, and vice versa, measuring the statistical dependence between X and Y, the larger the value, the stronger the correlation between X and Y); x and y respectively represent all possible values of X and Y; p(x, y) represents the joint probability that X takes the value x and Y takes the value y; p(x) and p(y) respectively represent the marginal probabilities that X takes the value x and Y takes the value y; Set the retention percentage; retain the effective features in the mutual information score set according to the retention percentage to obtain the effective laboratory features; for example, retain the features with the top 20% of the high mutual information score values and eliminate redundant features; S3. Construct an anomaly detection classification model and a quantitative regression model; train the anomaly detection classification model and the quantitative regression model respectively to obtain the final anomaly detection classification model and the final quantitative regression model; The S3 includes the following steps: S31. Based on the graph convolutional network and the attention mechanism, construct an anomaly detection classification model; set the weight parameters of the anomaly detection classification model; the construction of the anomaly detection classification model includes graph structure construction, graph convolutional layer design, setting the attention mechanism, and designing the output layer; the graph structure construction includes inputting the training sample features, defining the adjacency matrix, and adding self-loops; Based on XGBoost and LightGBM, a quantitative regression model is constructed; the quantitative regression model adopts a dual-model weighted fusion strategy of XGBoost and LightGBM; the XGBoost model is based on gradient boosting trees and optimizes the XGBoost model by controlling the complexity through a regularization term; accelerated based on the histogram algorithm, and adopts the Leaf-wise growth strategy; S32. Collect historical laboratory feature data; S33. Input the historical laboratory feature data into the anomaly detection classification model for training; During the training process, the optimal weight parameters of the anomaly detection classification model are found by combining random search and cross-validation algorithms to obtain the optimal solution; the optimal solution is used as the weight parameters of the anomaly detection classification model to obtain the final anomaly detection classification model; The process of finding the optimal weight parameters of the anomaly detection classification model by combining random search and cross-validation algorithms during the training process in S33 to obtain the optimal solution includes the following steps: S331. Set the search space for the weight parameters of the anomaly detection classification model as p, the maximum number of random searches as m, and the number of cross-validation folds as W; S332. Randomly set the weight parameters of the anomaly detection classification model according to the weight parameter search space p to obtain the initial weight parameter set , where ji represents the i-th initial weight parameter and k represents the total number of initial weight parameters; Initialize the optimal weight parameters of the anomaly detection classification model as u and the best performance as v. Randomly select the weight parameters of the anomaly detection classification model as e from the initial weight parameter set ; Divide the training data in the historical laboratory feature data into W folds. For each fold w, ; Use the training data in all historical laboratory feature data except the w-th fold to train the anomaly detection classification model, and use the training data in the w-th fold of the historical laboratory feature data to evaluate the anomaly detection classification model to obtain the performance index set ; where, tw represents the w-th performance index and W represents the total number of performance indexes; Calculate the mean of the performance indexes of all folds to obtain the average performance r. The calculation formula is as follows, ; If the average performance r > the best performance v, update the best performance to r and update the optimal weight parameters to e; otherwise, keep the original best performance and optimal weight parameters; S334. Repeat S333. When the maximum number of random searches m is reached, stop the iteration to obtain the optimal solution; S34. Input the historical laboratory feature data into a quantitative regression model for training, and optimize the parameters of the quantitative regression model by combining grid search to obtain the final quantitative regression model; Divide the historical laboratory feature data into a training set and a validation set; the training of the quantitative regression model is carried out in stages; in the first stage, use XGBoost to independently train the basic model, set 200 trees, a maximum depth of 7, a learning rate of 0.05, and use a subsample ratio of 0.8 to improve the anti-overfitting ability. After the model training is completed, calculate the prediction residuals of the training set as the target of the second stage; in the second stage, fix the XGBoost parameters, and use LightGBM to specifically learn these residuals, configure 150 leaf nodes, a learning rate of 0.1, and a feature sampling ratio of 0.7 to form a residual compensation model; On the retained validation set, perform grid search on the weight values in the interval [0.5, 0.7] with a step size of 0.01; at each weight value, fuse the predictions of the two models (α×XGB prediction + (1-α)×compensation prediction), and calculate the RMSE of the validation set; select the weight value corresponding to the minimum RMSE as the final fusion coefficient; when predicting, first use XGBoost to generate the basic prediction, then use LightGBM to generate the residual compensation, and finally weight and fuse the output results according to the optimal weight to obtain the final quantitative regression model; S4. Use the final anomaly detection classification model and the final quantitative regression model to perform dual-model collaborative analysis on the effective laboratory features to obtain the anomaly detection results and regression prediction values; Eliminate the abnormal data in the regression prediction values according to the anomaly detection results to obtain the normal regression prediction values; The S4 includes the following steps: S41. Input the effective laboratory features into the final anomaly detection classification model and the final quantitative regression model respectively to obtain the anomaly detection results and regression prediction values; S42. Set the anomaly probability threshold; mark the effective laboratory features with an anomaly probability ≥ anomaly probability threshold in the anomaly detection results as abnormal data; Eliminate the regression prediction values corresponding to the abnormal data in the regression prediction values to obtain the normal regression prediction values; S5. Predict the compensation amount according to the environmental parameter data by combining the LSTM network; obtain the calibrated laboratory data by combining the normal regression prediction values with the predicted compensation amount; The S5 includes the following steps: S51. Construct an LSTM network; set the calibration rules; update the LSTM network according to the calibration rules; For example, when the prediction error of 10 consecutive samples > 5%, automatically trigger incremental learning, and fine-tune the parameters of the last two layers of the LSTM with new data to obtain the updated LSTM network; S52. Input the environmental parameter data into the LSTM network to obtain the predicted compensation amount. The calculation process of the LSTM cell is as follows. ; Among them, it represents the input gate value at time step t (used to control which new information will be written into the cell state), σ represents the Sigmoid activation function (used to map the input to the interval (0, 1)), Wi represents the weight matrix of the input gate, [ht-1, et] represents concatenating the hidden state of the previous time step and the environmental parameters of the current time step into a vector as the input of the LSTM cell, and bi represents the bias vector of the input gate; ft represents the forget gate value at time step t (controlling how much past information needs to be discarded from the cell state), Wf represents the weight matrix of the forget gate, and bf represents the bias vector of the forget gate; ot represents the output gate value at time step t (used to control which parts of the cell state will be output as the hidden state), Wo represents the weight matrix of the output gate, and bo represents the bias vector of the output gate; c1t represents the candidate cell state value at time step t (used to represent the new information that may be added to the cell state at the current time step), tanh represents the hyperbolic tangent activation function (used to map the input to the interval [-1, 1]), Wc represents the weight matrix of the candidate cell state, and bc represents the bias vector of the candidate cell state; ct represents the cell state value at time step t (used to store long-term memory), ct-1 represents the cell state of the previous time step, ft * ct-1 represents how much past cell state information is retained, and it * c1t represents how much new candidate information is added to the cell state; ht represents the hidden state value at time step t (i.e., the predicted compensation amount), tanhct represents applying the tanh function to the current cell state and mapping it to the interval [-1, 1], and ot * tanhct represents the output gate determining which information in the hidden state is important; S53. Calculate the calibrated laboratory data based on the normal regression prediction value combined with the predicted compensation amount. The calculation formula is as follows. ; Among them, f represents the calibrated laboratory data, d represents the normal regression prediction value, and ht represents the predicted compensation amount.
[0025] Example 2: Please refer to Figure 4, A processing system for laboratory test data based on machine learning technology, which is used to implement the above-mentioned processing method for laboratory test data based on machine learning technology, including a multi-source data acquisition and dynamic noise reduction module, a multi-modal feature fusion and selection module, a dual-model collaborative training and optimization module, an anomaly-aware regression prediction module, and an environmental compensation dynamic calibration module; The multi-source data acquisition and dynamic noise reduction module is used for the real-time acquisition and preprocessing of laboratory multi-source signals and environmental parameters; structured data and unstructured data are obtained through an industrial bus or an API interface, and environmental parameters such as temperature, humidity, and voltage are synchronously acquired; the discrete wavelet transform is used to decompose the coefficients, the noise standard deviation is estimated based on the median of the highest frequency coefficients, the threshold is dynamically calculated, and then the noise coefficients are shrunk through a soft threshold function and the signal is reconstructed, and finally the denoised multi-source data is output; The multi-modal feature fusion and selection module is used to extract and fuse frequency domain and spatial features from the preprocessed data; for time series signals, local statistics are calculated through a sliding window and FFT frequency domain features are extracted; for spatial data, the SIFT algorithm is used to extract key point descriptors; the fused multi-source features are optimized through mutual information feature selection, the mutual information score between the features and the target variable is calculated, the high-score features are retained, the redundant features are removed, and the effective laboratory features are output; The dual-model collaborative training and optimization module is used to construct and train an anomaly detection and quantitative regression dual model; based on a graph convolutional network and an attention mechanism, a graph structure is constructed, the weight parameters are optimized by random search cross-validation, random sampling is performed in the parameter space, the performance is evaluated through W-fold cross-validation, and the optimal weights are iteratively updated to obtain the final anomaly detection classification model; by fusing XGBoost and LightGBM, it is trained in two stages: first, XGBoost is used to predict the base value, and then LightGBM is used to learn the residuals; finally, the weighted fusion coefficient is optimized on the validation set through grid search, and the weight that minimizes the RMSE is selected to obtain the final quantitative regression model; The anomaly-aware regression prediction module is used to implement dual-model collaborative inference and anomaly filtering; the effective features are respectively input into the final anomaly detection classification model and the final quantitative regression model, and the anomaly probability and the regression prediction value are output; an anomaly probability threshold is set to mark the anomaly samples; the regression prediction values corresponding to the anomaly samples are removed, and the pure normal regression prediction result is output; The environmental compensation dynamic calibration module eliminates environmental interference based on an LSTM network; the environmental parameters are input into the LSTM, and the prediction compensation amount is calculated through a gating mechanism; an incremental learning mechanism is designed to optimize the LSTM; the normal regression prediction value and the compensation amount are superimposed, and the final laboratory data after environmental calibration is output.
[0026] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0027] The preferred embodiments of the invention disclosed above are only used to help illustrate the invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the invention, so that those skilled in the art can understand and utilize the invention well.
Claims
1. A method for processing laboratory test data based on machine learning technology, characterized in that, It includes the following steps: S1. Collect multi-source signal data and environmental parameter data in the laboratory; Perform dynamic noise reduction processing on the multi-source signal data in the laboratory to obtain preprocessed multi-source data in the laboratory; S2. Extract the frequency-domain features and spatial features from the preprocessed multi-source data in the laboratory, and fuse them to obtain multi-source features in the laboratory; Perform mutual information feature selection processing on the multi-source features in the laboratory to obtain effective laboratory features; S3. Construct an anomaly detection classification model and a quantitative regression model; Train the anomaly detection classification model and the quantitative regression model respectively to obtain the final anomaly detection classification model and the final quantitative regression model; S4. Use the final anomaly detection classification model and the final quantitative regression model to perform dual-model collaborative analysis on the effective laboratory features to obtain anomaly detection results and regression prediction values; Eliminate the abnormal data in the regression prediction values according to the anomaly detection results to obtain normal regression prediction values; S5. According to the environmental parameter data, combine with the LSTM network to obtain a prediction compensation amount; according to the normal regression prediction values and the prediction compensation amount, obtain calibrated laboratory data.
2. The processing method of laboratory test data based on machine learning technology according to claim 1, wherein The S1 includes the following steps: S11. Real-time obtain multi-source signal data in the laboratory through an industrial bus or an API interface; the multi-source signal data in the laboratory includes structured signal data and unstructured signal data; Synchronously collect environmental parameter data to obtain environmental parameter data; S12. Perform dynamic noise reduction processing on the multi-source signal data in the laboratory to obtain preprocessed multi-source data in the laboratory.
3. The processing method of laboratory test data based on machine learning technology according to claim 2, wherein The S12 includes the following steps: S121. Perform wavelet transform on the multi-source signal data in the laboratory to obtain a wavelet coefficient set; the wavelet coefficients include approximation coefficients and detail coefficients; S122. Calculate the noise standard deviation according to the wavelet coefficient set; S123. Calculate the real-time dynamic threshold according to the noise standard deviation; S124. Process each wavelet coefficient in the wavelet coefficient set through a soft threshold function according to the real-time dynamic threshold to obtain a filtered wavelet coefficient set; S125. Perform signal reconstruction on the wavelet signals in the filtered wavelet coefficient set through inverse wavelet transform processing to obtain preprocessed multi-source data in the laboratory.
4. The processing method of laboratory test data based on machine learning technology according to claim 1, characterized in that The S2 includes the following steps: S21. Calculate the local statistics of the time-series signal data in the preprocessed multi-source data in the laboratory through a sliding window to obtain frequency-domain features; Use the SIFT algorithm to extract the key point descriptors of the spatial data in the preprocessed multi-source data in the laboratory to obtain spatial features; The frequency-domain features and the spatial features jointly constitute the multi-source features in the laboratory; S22. Calculate the mutual information score set of the features of the preprocessed multi-source data in the laboratory and the target variable; Set a retention percentage; retain the effective features in the mutual information score set according to the retention percentage to obtain effective laboratory features.
5. The processing method of laboratory test data based on machine learning technology according to claim 1, characterized in that The S3 includes the following steps: S31. Based on the graph convolutional network and the attention mechanism, construct an anomaly detection classification model; set the weight parameters of the anomaly detection classification model; Based on XGBoost and LightGBM, construct a quantitative regression model; the quantitative regression model adopts a dual-model weighted fusion strategy of XGBoost and LightGBM; S32. Collect historical laboratory feature data; S33. Input the historical laboratory feature data into the anomaly detection classification model for training; During the training process, combine random search and cross-validation algorithms to find the optimal weight parameters of the anomaly detection classification model, and obtain the optimal solution; use the optimal solution as the weight parameters of the anomaly detection classification model to obtain the final anomaly detection classification model; S34. Input the historical laboratory feature data into the quantitative regression model for training, and optimize the parameters of the quantitative regression model by combining grid search to obtain the final quantitative regression model.
6. The processing method of laboratory test data based on machine learning technology according to claim 5, characterized in that The step of combining random search and cross-validation algorithms to find the optimal weight parameters of the anomaly detection classification model and obtain the optimal solution during the training process in S33 includes the following steps: S331. Set the search space of the weight parameters of the anomaly detection classification model as p, the maximum number of random searches as m, and the number of cross-validation folds as W; S332. Randomly set the weight parameters of the anomaly detection classification model according to the search space p of the weight parameters to obtain the initial weight parameter set; Initialize the optimal weight parameters of the anomaly detection classification model as u and the best performance as v, and randomly select the weight parameters of the anomaly detection classification model from the initial weight parameter set as e; Divide the training data in the historical laboratory feature data into W folds. For each fold w, ; use the training data in all historical laboratory feature data except the w-th fold to train the anomaly detection classification model, and use the training data in the historical laboratory feature data of the w-th fold to evaluate the anomaly detection classification model to obtain a set of performance metrics; Calculate the mean of the performance indicators for all folds to obtain the average performance r, If the average performance r > the best performance v, update the best performance to r and update the optimal weight parameters to e; otherwise, keep the original best performance and optimal weight parameters; S334. Repeat S333. When the maximum number of random searches m is reached, stop the iteration to obtain the optimal solution.
7. The processing method of laboratory test data based on machine learning technology according to claim 1, characterized in that The S4 includes the following steps: S41. Input the valid laboratory features into the final anomaly detection classification model and the final quantitative regression model respectively to obtain the anomaly detection result and the regression prediction value; S42. Set the anomaly probability threshold; mark the valid laboratory features with an anomaly probability ≥ the anomaly probability threshold in the anomaly detection result as abnormal data; Exclude the regression prediction values corresponding to the abnormal data in the regression prediction values to obtain the normal regression prediction values.
8. The processing method of laboratory test data based on machine learning technology according to claim 1, characterized in that The S5 includes the following steps: S51. Construct an LSTM network; set the calibration rule; update the LSTM network according to the calibration rule; S52. Input the environmental parameter data into the LSTM network to obtain the prediction compensation amount; S53. Calculate and obtain the calibrated laboratory data based on the normal regression prediction values combined with the prediction compensation amount.
9. A processing system for laboratory test data based on machine learning technology, characterized in that, Implement the processing method of laboratory test data based on machine learning technology according to any one of claims 1-8. The system includes a multi-source data acquisition and dynamic noise reduction module, a multi-modal feature fusion and selection module, a dual-model collaborative training and optimization module, an anomaly perception regression prediction module, and an environmental compensation dynamic calibration module; The multi-source data acquisition and dynamic noise reduction module is used for real-time acquisition and preprocessing of laboratory multi-source signals and environmental parameters; obtain structured data and unstructured data through an industrial bus or API interface, and synchronously collect environmental parameters; Decompose the coefficients by discrete wavelet transform, estimate the noise standard deviation based on the median of the highest frequency coefficients, dynamically calculate the threshold, then shrink the noise coefficients through the soft threshold function and reconstruct the signal, and finally output the multi-source data after noise reduction; The multi-modal feature fusion and selection module is used to extract and fuse frequency domain and spatial features from the preprocessed data; for time series signals, calculate local statistics through a sliding window and extract FFT frequency domain features; for spatial data, use the SIFT algorithm to extract key point descriptors; the fused multi-source features are optimized through mutual information feature selection, calculate the mutual information score between the features and the target variable, retain the high-score features, eliminate redundant features, and output effective laboratory features; The dual-model collaborative training and optimization module is used to construct and train an anomaly detection and quantitative regression dual-model; based on the graph convolutional network and attention mechanism, construct a graph structure, use random search cross-validation to optimize the weight parameters, randomly sample in the parameter space, evaluate the performance through W-fold cross-validation, and iteratively update the optimal weights to obtain the final anomaly detection classification model; by fusing XGBoost and LightGBM, train in two stages: first use XGBoost to predict the base value, and then use LightGBM to learn the residuals; finally, optimize the weighted fusion coefficient on the validation set through grid search, select the weight that minimizes RMSE, and obtain the final quantitative regression model; The anomaly-aware regression prediction module is used to achieve dual-model collaborative inference and anomaly filtering; input the effective features into the final anomaly detection classification model and the final quantitative regression model respectively, and output the anomaly probability and the regression prediction value; Set the anomaly probability threshold to mark the abnormal samples; eliminate the regression prediction values corresponding to the abnormal samples and output the pure normal regression prediction results; The environment compensation dynamic calibration module eliminates environmental interference based on the LSTM network; Input the environmental parameters into the LSTM, calculate the prediction compensation amount through the gating mechanism; design an incremental learning mechanism to optimize the LSTM; superimpose the normal regression prediction value and the compensation amount, and output the final laboratory data after environmental calibration.
10. A storage medium, characterized in that, A program is stored thereon, and when the program is executed by a processor, it implements the processing method of laboratory test data based on machine learning technology as described in any one of claims 1-8.
Citation Information
Patent Citations
Qualitative and quantitative combination water quality monitoring method
CN106442420A
Remote sensing monitoring method of waterlogging of winter wheat, based on fusion of satellite-ground multi-source rainfall data
CN108764688A
Odometer calibration method, odometer calibration model training method and related device
CN118274878A
High-precision management method and system for managing service life of probe card
CN118797972A
Cited By
Cascae data management method and device, equipment and storage medium
CN120763488A
Method and device for correcting helium concentration data based on mass spectrometric detection
CN120951271A
A method and device for correcting helium concentration data based on mass spectrometry detection
CN120951271B
Big data analysis-based online multi-terminal interconnection production and education fusion precise management system
CN121213307A
An online multi-terminal interconnection production and teaching integration precise management system based on big data analysis
CN121213307B