Quality control method in sample collection and transportation process based on machine learning

By employing machine learning methods and utilizing tensor decomposition and dual-path deep neural network models, the problems of spatiotemporal structural information loss and lack of dynamic interaction relationships of features during sample collection and transportation were solved, enabling accurate prediction and early warning of sample failure risk and ensuring sample quality.

CN121961333APending Publication Date: 2026-05-01SHANDONG CENT FOR DISEASE CONTROL & PREVENTION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG CENT FOR DISEASE CONTROL & PREVENTION
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as loss of spatiotemporal structural information in environmental feature processing, insufficient distinguishability of feature importance, and lack of modeling ability for dynamic interaction relationships of features during sample collection and transportation, making it difficult to achieve sample quality control.

Method used

A machine learning-based approach is adopted to collect multi-source data by deploying a sensor array. Tensor decomposition and Tucker decomposition are used to preserve the spatiotemporal correlation structure, and a dual-path interactive deep neural network model is constructed. The model is combined with a significant weighting mechanism of information entropy and feature discriminative power to predict the risk of sample failure.

Benefits of technology

It achieves accurate prediction and early warning of sample failure risk, ensures sample quality, fully preserves the inherent coupling relationship between time, space and feature types, enhances the adaptability of key discriminative features, and dynamically adjusts the contribution ratio of environmental information and sample ontology information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961333A_ABST
    Figure CN121961333A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical sample transportation process intelligent monitoring, in particular to a sample collection and transportation process quality control method based on machine learning, and the method specifically comprises the following steps: deploying a sensor array to collect sample collection and transportation full life cycle multi-source data, and combining a detection result to mark a quality label to form a data set; performing tensor decomposition on the data to construct an environment spatial-temporal feature tensor, and performing Tucker decomposition; constructing an ontology feature vector based on the data of the sample, and generating a weighted feature vector through a significance weighting mechanism; constructing a dual-path interactive deep neural network, extracting environment features and fusing the environment features with weighted sample features, and predicting a sample failure risk by means of a channel attention multilayer perceptron; adopting a multi-task joint loss function to calculate total loss; dividing a data set training, verification and test model; and predicting a new sample risk by using the trained model and performing early warning according to a threshold value. According to the method, the sample quality can be accurately controlled, and the quality control efficiency and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring technology for medical sample transportation, and in particular to a quality control method based on machine learning during sample collection and transportation. Background Technology

[0002] The emergence and spread of carbapenem-resistant Enterobacteriaceae (NDM-CRE) with metallo-β-lactamase genes in New Delhi in recent years has become a major threat to global public health. These bacteria exhibit high levels of drug resistance, leading to limited clinical treatment options and high mortality rates. Continuous and precise epidemiological surveillance of specific communities or environments is crucial for effectively tracking the epidemiological sources, transmission routes, and variation patterns of NDM-CRE. In this process, collecting samples from the field and subsequently isolating and identifying strains in the laboratory are fundamental steps in obtaining key scientific data.

[0003] However, quality control during sample collection and transportation is the "last mile" that determines the success or failure of such monitoring projects. Taking this project as an example, the scientific value of multiple time-series samplings from the same community highly depends on the quality stability of the samples throughout the transportation process. Deviations at any stage—collection, packaging, storage, or transportation to the laboratory—such as improper sampling procedures leading to distorted sample representativeness, inconsistent sample labeling, or critical environmental parameters like temperature and time exceeding permissible limits, can directly result in the death of target microorganisms, sample invalidation, or cross-contamination. These quality issues will cause failed isolation and identification, distorted monitoring data, and even misleading conclusions, significantly diminishing the scientific value of the entire research project, wasting valuable public health resources, and potentially delaying accurate assessments of epidemic risks.

[0004] Existing technologies objectively suffer from the following shortcomings: Environmental feature processing typically employs vectorization followed by dimensionality reduction, which disrupts the inherent three-dimensional structure of spatiotemporal data, leading to the loss of information regarding temporal evolution patterns, spatial correlation characteristics, and coupling relationships between features; Sample feature preprocessing often uses standardization and normalization techniques, which only eliminate dimensional differences but cannot distinguish feature importance, resulting in key discriminative features being submerged in a large number of ordinary or noisy features; Multi-source feature fusion often adopts simple splicing or early fusion strategies, lacking the ability to model the dynamic interaction relationships between features and failing to adapt to changes in the relative importance of environmental and sample features under different scenarios; Model optimization relies solely on a loss function oriented towards prediction accuracy, lacking modeling of prior knowledge in the field of sample transportation quality control, thus limiting the physical rationality and generalization ability of the model.

[0005] Therefore, in order to address the serious public health threat and based on the urgent need for high-quality epidemiological surveillance data, this invention proposes a quality control method based on machine learning in the sample collection and transportation process to solve the above problems. Summary of the Invention

[0006] This invention addresses the shortcomings of existing technologies by developing a quality control method for sample collection and transportation based on machine learning. This invention can solve the problem of difficult quality control in sample collection and transportation, achieve accurate prediction and early warning of sample failure risk, and thus ensure sample quality.

[0007] The technical solution of this invention to solve the technical problem is a quality control method for sample collection and transportation based on machine learning, comprising the following steps: S1. Deploy sensor arrays to collect multi-source data associated with medical samples throughout their collection and transportation lifecycle, including environmental data and their own data. After transportation, label each sample with a quality label based on the test results to form a sample dataset. S2. Tensor decomposition is performed on the data in the sample dataset to construct the environmental spatiotemporal feature tensor. Then, Tucker decomposition is performed to decompose it into a core tensor, the product of three factor matrices and a residual tensor. S3. Construct sample ontology feature vectors based on ontology features in its own data, and then adaptively reweight the ontology feature vectors based on the significance weighting mechanism of information entropy and feature discrimination to generate weighted sample feature vectors. S4. Construct a dual-path interactive deep neural network model. First, extract the overall environmental pattern vector and the local environmental anomaly factor from the decomposed environmental spatiotemporal feature tensor. Then, fuse the overall environmental pattern vector, the local environmental anomaly factor, and the saliency-weighted sample feature vector through a gating mechanism. Finally, use a multilayer perceptron with a channel attention mechanism to predict the sample failure risk. S5. Employ a multi-task joint loss function to integrate the weighted exponential cumulative exposure loss, tensor decomposition structure preservation regularization loss, and anomaly perception marginal contrast loss to calculate the total loss of the model. S6. The sample dataset is divided into a training set, a validation set, and a test set. The model is trained using the training set and the validation set, and the parameters of the model with the best performance are saved. The trained model is evaluated using the test set. S7. Predict the failure risk of new samples using the trained model, set a risk threshold, and determine whether to issue an early warning.

[0008] S1 is as follows: The collected environmental data consists of multi-dimensional spatiotemporal monitoring data of the environment in which the sample is located, while the self-data consists of the ontological characteristic data of the individual sample itself. Networkable sensor arrays are deployed along the sample transportation route and at key storage nodes to continuously monitor and record seven types of environmental data, including temperature, humidity, light intensity, atmospheric pressure, wind speed, carbon dioxide concentration, and ultraviolet radiation intensity, forming a raw environmental monitoring sequence covering both time and space dimensions. The data is determined based on the sample type, with one sample type corresponding to one feature type. After the samples are transported and delivered to the laboratory, each sample is labeled with a binary category label based on authoritative quality testing results, with the label category being qualified and invalid. The dataset consists of multiple samples, and each sample in the dataset includes environmental data in the spatiotemporal dimensions, its own data, and corresponding labels.

[0009] S2 is as follows: Tensor decomposition technology is used to preserve the inherent spatiotemporal correlation structure in environmental features. Specifically, the collected data is organized into a third-order tensor. The three dimensions of this tensor correspond to time series, spatial location and feature type, respectively. This third-order tensor represents the spatiotemporal feature tensor of the environment. The environment spatiotemporal feature tensor is decomposed using Tucker decomposition, which extracts and preserves the low-rank spatiotemporal correlation structure by decomposing it into the product of a core tensor and three factor matrices and a residual tensor. The three factor matrices include a time factor matrix, a spatial factor matrix, and a feature factor matrix; the core tensor is used to compress the low-rank interactions and coupling structures between the time, space, and feature dimensions in the environmental features; and the residual tensor is used to capture local details, anomalies, or noise information that are not explained by the low-rank decomposition.

[0010] S3 is as follows: After generating the weighted sample feature vector, the statistical discriminant of each ontology feature among the labels is calculated using the labeled labels. The higher the discriminant, the more important the feature is in distinguishing the sample state. Next, calculate the information entropy of each ontology feature on the training dataset; then, combine the information entropy with the feature discrimination to calculate a comprehensive score for each ontology feature; then, normalize the scores of all ontology features using Softmax to obtain the saliency weight vector; finally, multiply the saliency weight vector element-wise with the sample ontology feature vector to obtain the weighted sample feature vector.

[0011] S4 is as follows: The specific steps for extracting the overall environment pattern vector in S4.1 are as follows: The core vector, temporal factor matrix, and spatial factor matrix in the decomposed environment spatiotemporal feature tensor are vectorized, temporal pattern pooling, and spatial pattern pooling operations are performed respectively. The outputs are concatenated and then dimensionality reduction and nonlinear transformation are performed through a fully connected layer to output the overall environment pattern vector.

[0012] S4.2 The specific steps for calculating local anomaly factors in the environment are as follows: The residual tensor in the decomposed environmental spatiotemporal feature tensor contains local details and potential anomaly information that are not explained by the low-rank mode. Multidimensional feature residuals corresponding to specific spatiotemporal locations are extracted from the residual tensor, and local environmental anomaly factors are calculated.

[0013] S4.3, Dual-path gated interactive feature fusion: The overall environmental pattern vector, local environmental anomaly factors, and weighted sample feature vectors are fused. Through a learnable gating mechanism, the contribution ratio of information from the environmental path and information from the sample ontology path in the final fused feature is dynamically adjusted to obtain the fused feature vector. S4.4, Attention-based multilayer perceptron risk prediction: The fused feature vectors are input into a multilayer perceptron that incorporates a channel attention mechanism, and the output is the probability of sample failure.

[0014] S5 is detailed below: S5.1. Combining the norm of the overall environmental pattern vector with the saliency weights of the sample ontological features, construct a weighted index cumulative exposure loss; S5.2. Tensor decomposition structure is used to preserve regularization loss by imposing nuclear norm constraints on the core tensor and orthogonality constraints on the factor matrix. S5.3 Calculate the marginal contrast loss of anomaly perception based on local environmental anomaly factors, and introduce a dynamically growing margin for the classification task. When the local environmental anomaly factors are high, the model is forced to widen the difference in prediction probabilities between positive and negative samples. S5.4. Weighted sum of the three losses to obtain the total loss function.

[0015] S6 is detailed below: End-to-end training of the dual-path interactive deep neural network model is conducted using the training set, employing an iterative optimization strategy that traverses the entire training data in each round. Within each training batch, forward propagation is performed first, executing the operations in S2 and S3 on the data in the training set before inputting it into the model to obtain the batch sample failure risk prediction probability. Subsequently, backpropagation is performed, combining the sample's true label, the model's predicted probability, and the intermediate tensor from the forward propagation to calculate the total loss function value. Using stochastic gradient descent or its variant, the gradient of the loss function with respect to the model's trainable parameters is calculated, and the parameters are updated to minimize the loss. During training, the model performance is periodically evaluated using an independent validation set. When the preset maximum number of training rounds is reached, or when the validation set performance metrics do not improve or even decline for several consecutive rounds, an early stopping mechanism is triggered to save the optimal parameters and complete the training. Finally, the generalization ability of the trained model is evaluated using a test set.

[0016] S7 is detailed below: The trained model is deployed in the online system for quality control of sample collection and transportation. The environmental and self-data of new samples after collection or during transportation are analyzed and processed to output the predicted probability of failure risk of the new sample in the current time and space. Set a risk threshold. If the probability is below the threshold, mark it as low risk and continue transportation; if it exceeds the threshold, mark it as high risk and trigger an alert.

[0017] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: This invention discloses a quality control method for sample collection and transportation based on machine learning. It proposes a three-dimensional tensor modeling and Tucker decomposition method for environmental spatiotemporal features, which fully preserves the inherent coupling relationship between time, space, and feature type dimensions during feature extraction, avoiding the loss of spatiotemporal structural information caused by traditional vectorization dimensionality reduction. It proposes a sample ontology feature saliency weighting mechanism based on information entropy and inter-class discriminative power, achieving adaptive enhancement of key discriminative features and overcoming the limitation of traditional standardization methods that treat all features equally while ignoring differences in feature contribution. It constructs a dual-path interactive deep neural network architecture, dynamically adjusting the contribution ratio of environmental information and sample ontology information through a gating fusion mechanism, solving the problem that traditional simple feature concatenation methods cannot model cross-modal dynamic interaction relationships. It employs a multi-task joint loss function, integrating three physical priors—weighted exponential cumulative exposure, tensor decomposition structure preservation, and anomaly perception marginal comparison—into the optimization process, overcoming the limitation of traditional cross-entropy loss functions lacking domain knowledge guidance. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0019] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0020] Figure 2 This is a schematic diagram of the operation process of S2.

[0021] Figure 3 This is a schematic diagram of the operation process of S4.

[0022] Figure 4This is a schematic diagram of the S7 operation process. Detailed Implementation

[0023] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0024] Example 1 like Figure 1 As shown, a quality control method for sample collection and transportation based on machine learning includes the following steps: S1. Deploy sensor arrays to collect multi-source data associated with medical samples throughout their collection and transportation lifecycle, including environmental data and their own data. After transportation, label each sample with a quality label based on the test results to form a sample dataset. S2. Tensor decomposition is performed on the data in the sample dataset to construct the environmental spatiotemporal feature tensor. Then, Tucker decomposition is performed to decompose it into a core tensor, the product of three factor matrices and a residual tensor. S3. Construct sample ontology feature vectors based on ontology features in its own data, and then adaptively reweight the ontology feature vectors based on the significance weighting mechanism of information entropy and feature discrimination to generate weighted sample feature vectors. S4. Construct a dual-path interactive deep neural network model. First, extract the overall environmental pattern vector and the local environmental anomaly factor from the decomposed environmental spatiotemporal feature tensor. Then, fuse the overall environmental pattern vector, the local environmental anomaly factor, and the saliency-weighted sample feature vector through a gating mechanism. Finally, use a multilayer perceptron with a channel attention mechanism to predict the sample failure risk. S5. Employ a multi-task joint loss function to integrate the weighted exponential cumulative exposure loss, tensor decomposition structure preservation regularization loss, and anomaly perception marginal contrast loss to calculate the total loss of the model. S6. The sample dataset is divided into a training set, a validation set, and a test set. The model is trained using the training set and the validation set, and the parameters of the model with the best performance are saved. The trained model is evaluated using the test set. S7. Predict the failure risk of new samples using the trained model, set a risk threshold, and determine whether to issue an early warning.

[0025] In a specific implementation, S1 is as follows: To construct a dataset for model training, multi-source data related to the samples throughout their entire lifecycle of collection and transportation were systematically collected. Data collection mainly focused on two aspects: First, there is multi-dimensional spatiotemporal monitoring data of the environment in which the sample is located; second, there is ontological characteristic data of the individual sample itself.

[0026] Specifically, network-connected sensor arrays are deployed along the pre-defined sample transport route and at key storage nodes to continuously monitor and record seven environmental parameters: temperature, humidity, light intensity, atmospheric pressure, wind speed, carbon dioxide concentration, and ultraviolet radiation intensity. This forms a raw environmental monitoring sequence covering both temporal and spatial dimensions. Specifically, temperature (unit: ...) ), humidity (unit: ), light intensity (unit: Atmospheric pressure (unit: atmosphere) Wind speed (unit: ), carbon dioxide concentration (unit: ), ultraviolet radiation intensity (unit: ); At the same time, for each transported sample individual, seven ontological characteristics were recorded.

[0027] The sample status is labeled. After the sample is transported and delivered to the laboratory, each sample is assigned a binary category label of "qualified" or "invalid" based on authoritative quality test results (such as activity retention rate, degree of component degradation, contamination status, etc.).

[0028] By collecting a large amount of such sample data and dividing it into training, validation, and test sets, a complete dataset is constructed for subsequent model training and evaluation.

[0029] In specific implementation methods, such as Figure 2 As shown, S2 is as follows: Environmental spatiotemporal features have a structure with three dimensions: time, space, and feature type. Conventional methods usually expand the data into vectors and then use principal component analysis for dimensionality reduction, which will destroy the inherent correlation and coupling effect between these three dimensions and easily lead to the loss of feature properties related to spatiotemporal change patterns.

[0030] This invention utilizes tensor decomposition technology to preserve the inherent spatiotemporal correlation structure in environmental features. Specifically, the collected environmental spatiotemporal feature data is organized into a third-order tensor, whose three dimensions correspond to time series, spatial location, and feature type, respectively. The environment spatiotemporal feature tensor has dimension 1. ; in, This represents the size of the time dimension, i.e., the total number of sampling points in the time series. For example, sampling hourly data from 24 hours of data would... ; The dimension representing the spatial dimension, i.e., the total number of different spatial monitoring locations or sampling points; Furthermore, the feature dimension has a size of 7, corresponding to 7 preset environmental feature types.

[0031] For the three-dimensional feature tensor of the environment spatiotemporal feature tensor Performing Tucker decomposition decomposes the data into the sum of a core tensor, the product of three factor matrices, and a residual tensor. This extracts and preserves the low-rank spatiotemporal correlation structure. The Tucker decomposition yields: Core Tensor , dimension And satisfy , , It is used to compress the low-rank interactions and coupling structures between the temporal, spatial, and feature dimensions in characterizing environmental features; Time factor matrix , dimension Each column of this matrix is ​​a basis vector in the time dimension, used to extract and represent the main patterns of environmental features changing over time; Space factor matrix , dimension Each column of this matrix is ​​a basis vector in the spatial dimension, used to extract and represent the association patterns between different geographical locations or sampling points; Feature factor matrix , dimension Each column of this matrix is ​​a basis vector along the feature dimension, used to fuse seven environmental features; residual tensor , dimension It is used to capture local details, anomalies, or noise information that are not interpreted by low-rank decomposition.

[0032] in, The rank represents the time dimension. The rank represents the spatial dimension. Represents the rank of the feature dimension.

[0033] In a specific implementation, S3 is as follows: Because sample ontology features not only have different dimensions, but also contribute differently to the prediction of sample failure risk, conventional Z-score normalization methods can only eliminate dimensions, but treat all features equally, failing to enhance the contribution of key features or suppress irrelevant or noisy features, which may lead to low model learning efficiency and limited discrimination performance. Therefore, this invention adaptively reweights the sample ontology feature vectors through a saliency weighting mechanism based on information entropy and feature discriminativeness, in order to enhance key features that contribute significantly to failure risk discrimination. The specific steps are as follows: 1) Constructing the sample ontology feature vector Collect 7-dimensional ontology features for each sample to construct the sample ontology feature vector. ; in, The first feature vector representing the sample ontology The element, i.e., the th element The specific numerical values ​​of the ontological features of the sample; Indicates feature index, ; This represents the vector transpose symbol.

[0034] 2) Calculate the feature discrimination vector Using the labeled "qualified" and "failed" samples in the training set, calculate the statistical discriminant of each ontology feature between the two classes. The higher the discriminant, the more important the feature is in distinguishing the sample states, thus obtaining the feature discriminant vector. ; in, Indicates the first The discriminative power of a sample ontological feature is the ratio of the sum of the absolute differences of the means and the standard deviations between "qualified" and "failed" samples. A higher value indicates that the feature is more important in distinguishing sample states, characterizing the feature's statistical discriminative ability between two classes of samples. The calculation method is expressed as follows: ; This indicates the first "failure" sample in the training set. The arithmetic mean of the features, This indicates the first "qualified" sample in the training set. The arithmetic mean of the features; This indicates the first "failure" sample in the training set. The standard deviation of each feature This indicates the first "qualified" sample in the training set. Features The standard.

[0035] 3) Calculate the feature saliency weights based on information entropy. Calculate the information entropy of each feature over the entire training set to measure the uncertainty or information content of that feature's value; that is, the information entropy of the first feature. Information entropy of each feature ; Then, the information entropy is combined with the feature discrimination vector to calculate a comprehensive score for each feature, i.e., the first feature. The combined score of each feature ; Finally, the score vectors of all features are normalized using Softmax to obtain the significance weight vector. ,Right now ,satisfy and ; in, Indicates the first The significance weight scalar of each feature represents the relative importance of the feature in predicting the risk of sample failure after being calculated and normalized by combining information entropy and inter-class discrimination. This represents the Softmax normalization function, used to transform an input vector into a probability distribution; Indicates the first The total number of bins after binning based on each feature can be determined based on the training set size and feature distribution, for example, by setting... That is, tenths; The index representing the binning. ; Indicates the first The proportion of samples in each bin to the total number of training samples is used to divide continuous features into bins, and then the proportion of samples in each bin to the total number of samples is used as the empirical probability of the value range of that bin. The logarithmic function is represented, with the default base being the natural constant; d represents the feature discrimination vector, i.e. .

[0036] 4) Generate weighted sample feature vectors The saliency weight vector is multiplied element-wise with the sample ontology feature vector to obtain the weighted sample feature vector. , dimension This enhances key features and suppresses secondary features.

[0037] In specific implementation methods, such as Figure 3 As shown, S4 is as follows: Conventional deep neural networks typically input features from different sources by simply concatenating them, neglecting the overall coupling of the spatiotemporal patterns of the environment, local details, and the dynamic interactions between sample features. Furthermore, in sample collection and transportation scenarios, the cumulative and abrupt effects of environmental conditions can significantly impact sample quality, and traditional network structures lack the ability to specifically model these factors. Therefore, this invention constructs a dual-path interactive deep neural network module. First, it extracts the overall environmental pattern vector through temporal pattern pooling, spatial pattern pooling, and core tensor vectorization. Simultaneously, it calculates local environmental anomaly factors from the residual tensor. Then, it fuses the overall environmental pattern vector, local environmental anomaly factors, and saliency-weighted sample feature vectors through a gating mechanism. Finally, it uses a multilayer perceptron with channel attention to predict sample failure risk. The specific steps are as follows: S4.1, Overall Environmental Pattern Extraction and Vectorization The influence of the environment on samples has a cumulative effect and a spatial proximity effect. This invention extracts richer global patterns through temporal pattern pooling, spatial pattern pooling, and feature cross-coding. Specifically: a) Time-mode pooling: For the time factor matrix... Each row (representing a time pattern at a specific point in time) undergoes temporal attention pooling to obtain a... dimensional vector The time base pattern is calculated by weighting the elements and summing them over the entire time period, highlighting the time periods that have a significant impact on sample quality. b) Spatial pattern pooling: For the spatial factor matrix... Each row (representing a spatial pattern of a spatial point) is averaged to obtain a... dimensional vector It represents the overall spatial distribution pattern; in, The temporal pattern vector is extracted from the temporal factor matrix through temporal attention pooling, and has a dimension of [missing information]. We should focus on time periods that have a significant impact on sample quality. The spatial pattern vector is extracted from the spatial factor matrix using average pooling, and has a dimension of [missing information]. It represents the overall spatial distribution pattern.

[0038] The temporal pattern vector, spatial pattern vector, and low-rank core tensor obtained from Tucker decomposition are concatenated, and then dimensionality reduction and nonlinear transformation are performed through a fully connected layer to extract the overall environmental pattern vector that can represent the low-rank coupling relationship between the three dimensions of time, space, and feature type. This vector is represented as follows: In the formula, Represents the overall environmental pattern vector, with dimension . It represents the global, low-rank relational structure of environmental features in the temporal, spatial, and feature dimensions; This represents a vectorization operation used to flatten an input tensor into a one-dimensional column vector along all dimensions. represents the core mapping weight matrix, used to perform linear transformations on the vectorized core tensor, and is a trainable parameter; This represents the core mapping bias vector, which is a bias added after the linear transformation and is a trainable parameter. The dimension of the overall environmental pattern vector is a hyperparameter, preferably set to 128; This represents the LeakyReLU activation function, used to introduce nonlinear modeling capabilities.

[0039] S4.2 Extraction and Factor Calculation of Local Environmental Anomalies The residual tensor contains local details and potential anomaly information not explained by the low-rank mode. To quantify the degree of anomaly in the local environment of each sample, multidimensional feature residuals corresponding to specific spatiotemporal locations are extracted from the residual tensor, and the local anomaly factor of the environment is calculated, expressed as: In the formula, Indicates time index and spatial index The local environmental anomaly factor at a given location is a scalar value. The larger the value, the greater the degree of anomaly in the environmental state at that spatiotemporal point, which deviates from the overall low-rank pattern. This represents a time index, with a value range of [value range missing]. ; This represents a spatial index, with a value range of 100. ; This represents the feature category index, with a value range of [value range missing]. ; Representing the residual tensor In position The element value on, i.e., the first Environmental characteristics in the first The time point, the first The residual at the nth spatial point represents the residual at the nth spatial point. The time point, the first At the spatial location, the first The degree to which environmental characteristics deviate from the overall low-rank coupling model; a positive value indicates that the actual value is higher than the low-rank model prediction, while a negative value indicates the opposite. A larger absolute value may indicate local anomalies, such as sensor transient failures or sudden changes in local microclimate. This indicates that at all spatiotemporal points in the training set, the th The arithmetic mean of the residuals of various environmental characteristics is a statistic pre-calculated based on the training data; This indicates that at all spatiotemporal points in the training set, the th The standard deviation of the residuals of environmental characteristics is a statistic pre-calculated based on the training data; This represents a very small positive number, used to prevent the denominator in a calculation formula from being zero. Examples of its values ​​are shown below. .

[0040] S4.3, Dual-path Gated Interactive Feature Fusion The overall environmental pattern vector, local environmental anomaly factors, and weighted sample feature vectors are fused. Through a learnable gating mechanism, the contribution ratios of information from environmental paths and information from sample ontology paths in the final fused features are dynamically adjusted, as expressed below: In the formula, This represents the fused feature vector, with dimension . It dynamically integrates global environmental patterns, local environmental anomalies, and saliency-weighted sample ontological features to characterize the comprehensive state of a sample under a specific spatiotemporal context, which is then used for the final failure risk assessment. This represents a vector concatenation operation; This represents the dimension of the fused feature vector; it is a hyperparameter and is set to 64. This represents a first small multilayer perceptron, specifically designed for processing environmental information. Its input is a concatenated vector of the overall environmental pattern vector and anomaly factors, and its output dimension is... A two-layer multilayer perceptron structure is preferred. This represents a second small multilayer perceptron, specifically designed for processing weighted sample features. The input is a weighted sample feature vector, and the output dimension is... A two-layer multilayer perceptron structure is preferred. Represents the fusion gate vector, with dimension . Each element has a value between 0 and 1, and each element dynamically controls the fusion ratio of environmental information and sample ontology information on the corresponding feature dimension. A value closer to 1 indicates greater trust in environmental information, while a value closer to 0 indicates greater trust in the sample ontology information. The calculation method is expressed as follows: ; Indicates and A vector of all 1s with the same dimension; This represents the Sigmoid activation function, used to map the output of the gated layer to... interval; This represents element-wise multiplication; The weight matrix of the gated layer is a trainable parameter with dimension 1. ; The bias vector of the gated layer is a trainable parameter with dimension . .

[0041] S4.4 Multilayer Perceptron Risk Prediction Based on Attention Mechanism The fused feature vectors are input into a multilayer perceptron incorporating a channel attention mechanism, and the output is the probability of sample failure. The attention mechanism enables the network to adaptively emphasize the feature dimensions most relevant to failure risk judgment, as shown below: In the formula, This represents the predicted probability of sample failure, and is a scalar value between 0 and 1. The larger the value, the higher the risk that the model determines the sample to be failure. This represents the attention-weighted feature vector, which recalibrates the feature channels of the first hidden layer output vector through an attention mechanism, enhancing features related to failure risk. The calculation method is expressed as follows: , dimension ; This represents the channel attention weight vector. The closer an element is to 1, the more important the information from that feature channel; the closer it is to 0, the less important or suppressed the information from that channel. The calculation method is expressed as follows: , dimension ; The output vector of the first hidden layer is a non-linear transformation of the fused feature vector, with dimension 1. It captures the high-level abstract features of the comprehensive state, and the calculation method is represented as follows. ; This represents the weight matrix of the output layer, with dimension 1. , are trainable parameters used to map hidden layer features to risk probabilities; This represents the bias scalar of the output layer, with dimension . , are trainable parameters; This represents the weight matrix of the hidden layer, with dimension 1. , are trainable parameters used to map attention-weighted features to the hidden space; This represents the bias vector of the hidden layer, with dimension . , are trainable parameters; This represents the weight matrix of the channel attention layer, with dimension 1. , are trainable parameters used to generate attention weights from pooled features; This represents the bias vector of the channel attention layer, with dimension . , are trainable parameters; This represents the dimension of the second hidden layer; it is a hyperparameter and is set to 64. This represents the dimension of the first hidden layer; it is a hyperparameter and is set to 32. This represents the global average pooling operation, which averages all elements of the input vector. This represents the global max pooling operation, which takes the maximum value of all elements in the input vector.

[0042] In a specific implementation, S5 is as follows: Because the conventional binary cross-entropy loss function only utilizes the difference between the true sample label and the model's predicted probability for optimization, it fails to model multiple key physical priors and internal model state information of the samples in the collection and transportation quality control task. This invention employs a multi-task joint loss function, integrating weighted exponential cumulative exposure loss, tensor decomposition structure preservation regularization loss, and anomaly-aware marginal contrast loss. This ensures that the model optimization process not only fits the sample labels but also conforms to the inherent physical laws and data characteristics of the sample quality control task. The specific steps are as follows: S5.1 Calculation of Weighted Index Cumulative Exponent Exponent Loss The risk of sample failure often exhibits a non-linear relationship with the cumulative exposure to adverse environmental conditions. Furthermore, the contribution of different sample ontological features to the risk varies. By combining the norm of the overall environmental pattern vector with the saliency weights of sample ontological features, a weighted index cumulative exposure loss is constructed to model the cumulative environmental effect and enhance the influence of key features on loss calculation. This is expressed as: , In the formula, This represents the weighted index cumulative exposure loss, which is a scalar value. Higher loss weights are assigned to failure samples with high environmental cumulative exposure and qualified samples with high risk characteristics, thereby guiding the model to pay more attention to them. Indicates the total number of samples in the batch; This represents the sample index, with a value range of [value range missing]. ; Indicates the first The true label for each sample is 0, which represents "qualified" and 1, which represents "invalid". The model represents the first The probability of sample failure risk prediction for a single sample is calculated using the same method as the probability of sample failure risk prediction. The calculation method is the same; Indicates the first The overall environmental pattern vector corresponding to each sample has a dimension of . For a single sample, the calculation method is similar to the overall environmental pattern vector. The calculation method is the same; Represents the L2 norm; The intensity coefficient representing cumulative environmental exposure is a hyperparameter greater than 0, used to control the influence of the overall environmental mode vector norm on loss modulation. An example value is 0.01. The intensity coefficient represents the key features of the sample ontology. It is a hyperparameter greater than 0 and is used to control the degree of influence of key sample features on loss adjustment. An example value is 0.1. Indicates the first The first sample The saliency weight scalar of each ontological feature is a saliency weight vector. The i-th element, with a value range between 0 and 1, represents the relative importance of the feature; Indicates the first The first sample ontology feature vector of the nth sample One element; This represents the natural exponential function.

[0043] S5.2 Calculation of Tensor Decomposition Structure Preservation Regularization Loss To ensure that the core tensor and factor matrices obtained from the environmental spatiotemporal feature tensor decomposition maintain their low rank and decoupling between factors during end-to-end training, and to avoid destroying their physical meaning, a tensor decomposition structure-preserving regularization loss is adopted. This is achieved by imposing nuclear norm constraints on the core tensor and orthogonality constraints on the factor matrices to maintain the structural characteristics of the decomposition results, expressed as: , In the formula, represents the tensor decomposition structure-preserving regularization loss, a scalar value used to constrain the intermediate representation of the environmental feature extraction path during training, encouraging it to maintain low rank and structural properties; represents the core tensor obtained from Tucker decomposition, with dimension . ; The nuclear norm of a matrix is ​​represented by the number of atoms in a matrix. Before calculating the nuclear norm, it needs to be treated as a matrix and expanded along the first dimension; express transpose; express transpose; express transpose; The dimension is The identity matrix; The dimension is The identity matrix; The dimension is The identity matrix; Represents the Frobenius norm; The weight coefficients for low-rank regularization of the core tensor are hyperparameters greater than 0, which control the proportion of the kernel norm in the total regularization loss. An example value is 0.1. This represents the weighting coefficient for orthogonality regularization of the factor matrix. It is a hyperparameter greater than 0 that controls the proportion of the orthogonality constraint term in the total regularization loss. An example value is 0.01.

[0044] S5.3, Anomaly Perception Marginal Contrast Loss Calculation During sample transportation, abnormal fluctuations in the local environment are a significant risk factor leading to sample failure. To make the model's predictions more confident in such scenarios, an anomaly perception marginal contrast loss is calculated based on the local environmental anomaly factor. This introduces a dynamically growing margin into the classification task. When the local environmental anomaly factor is high, it forces the model to widen the probability gap between positive and negative samples, thus making more discriminative predictions. This is expressed as: , In the formula, This represents the marginal contrast loss for anomaly perception. It is a scalar value and is designed to improve the model's prediction confidence and discrimination ability in abnormal environmental scenarios. This represents the baseline marginal value, a hyperparameter greater than 0. It defines the minimum classification margin that the model should strive to achieve even under normal environmental conditions, with an example value of 0.3. The anomaly sensitivity coefficient is a hyperparameter greater than 0, used to control the dynamic adjustment strength of local environmental anomalies on the classification margin. An example value is 1.0. Indicates the first The local environmental anomaly factor for each sample, calculated using the same method as in the time index, is... and spatial index Local environmental anomalies The calculation method is the same; This represents the hinge loss function, where the loss is 0 when a sample is correctly classified and its prediction confidence exceeds the dynamic margin, which is the sum of the base margin and the dynamic adjustment part; otherwise, a positive penalty is applied.

[0045] It should be noted that, Item representation will be the true label Mapped to Where qualified samples are mapped to -1 and invalid samples are mapped to +1. The term representation will predict the probability Mapped to The two work together to map the binary classification labels and predicted probabilities onto a symmetric scale, which is used to construct a unified "confidence" metric.

[0046] S5.3 Calculation of Total Loss Function The overall loss function during model training is a weighted sum of the weighted exponential cumulative exposure loss, the tensor decomposition structure preservation regularization loss, and the anomaly perception marginal contrast loss. Through joint optimization, the model simultaneously satisfies data fitting and physical constraints, expressed as: , In the formula, This represents the total loss function, which is a scalar value and is the final optimization objective that the model needs to minimize during the training phase. The fusion weight represents the loss of tensor decomposition structure preservation regularization. It is a hyperparameter greater than 0 and is used to balance the main loss and the structure regularization term. An example value is 0.05. The fusion weight represents the marginal contrast loss for anomaly perception. It is a hyperparameter greater than 0 and is used to balance the main loss and the anomaly perception discrimination term. An example value is 0.2.

[0047] In a specific implementation, S6 is as follows: The dual-path interactive deep neural network model is trained end-to-end using the training dataset. The training process employs an iterative optimization strategy, with each training round (called an epoch) traversing the entire training dataset.

[0048] In each training batch, the forward propagation process is first executed: according to step S2, the spatiotemporal environmental data of each sample in the batch is decomposed using Tucker to obtain the core tensor, temporal factor matrix, spatial factor matrix, feature factor matrix, and residual tensor; according to step S3, the saliency weight vector of the ontological features of the samples in this batch is calculated and a weighted sample feature vector is generated; then, according to the process of step S4, the overall environmental pattern extraction and vectorization, local environmental anomaly extraction and factor calculation, dual-path gated interactive feature fusion, and multilayer perceptron risk prediction based on attention mechanism are performed in sequence to finally obtain the failure risk prediction probability of all samples in this batch.

[0049] Enter the backpropagation process: According to step S5, use the true labels of the samples, the model's predicted probabilities, and the intermediate tensors generated during the forward propagation process (such as the overall environmental pattern vector, core tensor, factor matrix, local environmental anomaly factors, etc.) to calculate the value of the total loss function.

[0050] Then, using stochastic gradient descent or its variants, the gradient of the total loss function with respect to all trainable parameters of the model is calculated, and these parameters are updated to minimize the total loss function.

[0051] During training, model performance is evaluated periodically using independent validation sets, such as calculating accuracy, precision, recall, F1 score, or AUC value.

[0052] The training of the model is stopped when the preset maximum number of training epochs is reached, or when the performance metrics evaluated on the validation set no longer improve or even begin to decline in multiple consecutive training epochs. In this case, the early stopping mechanism can be triggered to save the optimal model parameters, prevent overfitting, and complete the model training.

[0053] In specific implementation methods, such as Figure 4 As shown, S7 is as follows: The trained and saved dual-path interactive deep neural network model is deployed in an online system for sample collection and transportation quality control to achieve real-time risk warning and decision support.

[0054] In practical applications, for any new sample after collection or during transportation, the system collects its associated spatiotemporal environmental monitoring data (sequences of 7 environmental parameters along the transportation path) and the sample's ontological feature data (7-dimensional features) in real time. Specifically, First, following step S2, the environmental data of the new sample is constructed into an environmental spatiotemporal feature tensor and Tucker decomposition is performed. At the same time, following step S3, the feature discrimination vector calculated and fixed from all training data during the training phase and the information entropy correlation statistic are used to calculate the significance weights of the ontological features of the new sample and generate a weighted sample feature vector. Then, the core tensor, factor matrix, residual tensor, and weighted sample feature vector obtained from the decomposition are input into the pre-trained model described in step S4. The model performs forward propagation calculations and finally outputs the predicted probability of failure risk of the new sample under the current spatiotemporal conditions. Based on this predicted probability, one or more risk thresholds can be set, for example, a threshold of 0.7. If the predicted probability is lower than the threshold, the current risk of the sample is considered controllable and marked as "low risk," and the normal transportation process continues. If the predicted probability exceeds the threshold, an early warning is immediately triggered and the sample is marked as "high risk." The system can automatically remind transportation managers or quality control personnel through a visual interface, SMS, or application notifications. In this way, proactive, intelligent, and prediction-based quality control of the sample collection and transportation process is achieved, effectively reducing the risk of sample failure.

[0055] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the invention. Based on the technical solutions of the invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the invention.

Claims

1. A quality control method for sample collection and transportation based on machine learning, characterized in that, Includes the following steps: S1. Deploy sensor arrays to collect multi-source data associated with medical samples throughout their collection and transportation lifecycle, including environmental data and their own data. After transportation, label each sample with a quality label based on the test results to form a sample dataset. S2. Tensor decomposition is performed on the data in the sample dataset to construct the environmental spatiotemporal feature tensor. Then, Tucker decomposition is performed to decompose it into a core tensor, the product of three factor matrices and a residual tensor. S3. Construct sample ontology feature vectors based on ontology features in its own data, and then adaptively reweight the ontology feature vectors based on the significance weighting mechanism of information entropy and feature discrimination to generate weighted sample feature vectors. S4. Construct a dual-path interactive deep neural network model. First, extract the overall environmental pattern vector and the local environmental anomaly factor from the decomposed environmental spatiotemporal feature tensor. Then, fuse the overall environmental pattern vector, the local environmental anomaly factor, and the saliency-weighted sample feature vector through a gating mechanism. Finally, use a multilayer perceptron with a channel attention mechanism to predict the sample failure risk. S5. Employ a multi-task joint loss function to integrate the weighted exponential cumulative exposure loss, tensor decomposition structure preservation regularization loss, and anomaly perception marginal contrast loss to calculate the total loss of the model. S6. The sample dataset is divided into a training set, a validation set, and a test set. The model is trained using the training set and the validation set, and the parameters of the model with the best performance are saved. The trained model is evaluated using the test set. S7. Predict the failure risk of new samples using the trained model, set a risk threshold, and determine whether to issue an early warning.

2. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S1 is as follows: The collected environmental data consists of multi-dimensional spatiotemporal monitoring data of the environment in which the sample is located, while the self-data consists of the ontological characteristic data of the individual sample itself. Networkable sensor arrays are deployed along the sample transportation route and at key storage nodes to continuously monitor and record seven types of environmental data, including temperature, humidity, light intensity, atmospheric pressure, wind speed, carbon dioxide concentration, and ultraviolet radiation intensity, forming a raw environmental monitoring sequence covering both time and space dimensions. The data is determined based on the sample type, with one sample type corresponding to one feature type. After the samples are transported and delivered to the laboratory, each sample is labeled with a binary category label based on authoritative quality testing results, with the label category being qualified and invalid. The dataset consists of multiple samples, and each sample in the dataset includes environmental data in the spatiotemporal dimensions, its own data, and corresponding labels.

3. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S2 is as follows: Tensor decomposition technology is used to preserve the inherent spatiotemporal correlation structure in environmental features. Specifically, the collected data is organized into a third-order tensor. The three dimensions of this tensor correspond to time series, spatial location and feature type, respectively. This third-order tensor represents the spatiotemporal feature tensor of the environment. The environment spatiotemporal feature tensor is decomposed using Tucker decomposition, which extracts and preserves the low-rank spatiotemporal correlation structure by decomposing it into the product of a core tensor and three factor matrices and a residual tensor. The three factor matrices include a time factor matrix, a spatial factor matrix, and a feature factor matrix; the core tensor is used to compress the low-rank interactions and coupling structures between the time, space, and feature dimensions in the environmental features; and the residual tensor is used to capture local details, anomalies, or noise information that are not explained by the low-rank decomposition.

4. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S3 is as follows: After generating the weighted sample feature vector, the statistical discriminant of each ontology feature among the labels is calculated using the labeled labels. The higher the discriminant, the more important the feature is in distinguishing the sample state. Next, calculate the information entropy of each ontology feature on the training dataset; then, combine the information entropy with the feature discrimination to calculate a comprehensive score for each ontology feature; then, normalize the scores of all ontology features using Softmax to obtain the saliency weight vector; finally, multiply the saliency weight vector element-wise with the sample ontology feature vector to obtain the weighted sample feature vector.

5. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that: The specific steps for extracting the overall environment pattern vector in S4.1 are as follows: The core vector, temporal factor matrix, and spatial factor matrix in the decomposed environment spatiotemporal feature tensor are vectorized, temporal pattern pooling, and spatial pattern pooling operations are performed respectively. The outputs are concatenated and then dimensionality reduction and nonlinear transformation are performed through a fully connected layer to output the overall environment pattern vector.

6. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that: S4.2 The specific steps for calculating local anomaly factors in the environment are as follows: The residual tensor in the decomposed environmental spatiotemporal feature tensor contains local details and potential anomaly information that are not explained by the low-rank mode. Multidimensional feature residuals corresponding to specific spatiotemporal locations are extracted from the residual tensor, and local environmental anomaly factors are calculated.

7. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that: S4.3, Dual-path gated interactive feature fusion: The overall environmental pattern vector, local environmental anomaly factors, and weighted sample feature vectors are fused. Through a learnable gating mechanism, the contribution ratio of information from the environmental path and information from the sample ontology path in the final fused feature is dynamically adjusted to obtain the fused feature vector. S4.4, Attention-based multilayer perceptron risk prediction: The fused feature vectors are input into a multilayer perceptron that incorporates a channel attention mechanism, and the output is the probability of sample failure.

8. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S5 is detailed below: S5.

1. Combining the norm of the overall environmental pattern vector with the saliency weights of the sample ontological features, construct a weighted index cumulative exposure loss; S5.

2. Tensor decomposition structure is used to preserve regularization loss by imposing nuclear norm constraints on the core tensor and orthogonality constraints on the factor matrix. S5.3 Calculate the marginal contrast loss of anomaly perception based on local environmental anomaly factors, and introduce a dynamically growing margin for the classification task. When the local environmental anomaly factors are high, the model is forced to widen the difference in prediction probabilities between positive and negative samples. S5.

4. Weighted sum of the three losses to obtain the total loss function.

9. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S6 Specifically as follows: End-to-end training of the dual-path interactive deep neural network model is conducted using the training set, employing an iterative optimization strategy that traverses the entire training data in each round. Within each training batch, forward propagation is performed first, executing the operations in S2 and S3 on the data in the training set before inputting it into the model to obtain the batch sample failure risk prediction probability. Subsequently, backpropagation is performed, combining the sample's true label, the model's predicted probability, and the intermediate tensor from the forward propagation to calculate the total loss function value. Using stochastic gradient descent or its variant, the gradient of the loss function with respect to the model's trainable parameters is calculated, and the parameters are updated to minimize the loss. During training, the model performance is periodically evaluated using an independent validation set. When the preset maximum number of training rounds is reached, or when the validation set performance metrics do not improve or even decline for several consecutive rounds, an early stopping mechanism is triggered to save the optimal parameters and complete the training. Finally, the generalization ability of the trained model is evaluated using a test set.

10. The quality control method for sample collection and transportation based on machine learning according to claim 1, characterized in that, S7 is detailed below: The trained model is deployed in the online system for quality control of sample collection and transportation. The environmental and self-data of new samples after collection or during transportation are analyzed and processed to output the predicted probability of failure risk of the new sample in the current time and space. Set a risk threshold; if the probability is below the threshold, mark it as low risk and continue transportation. If the threshold is exceeded, a high-risk indicator is marked, triggering an alert.