Food illegal additives ai assisted mass spectrometry qualitative and quantitative detection method

By employing technologies such as gas-phase assisted desorption chemical ionization mass spectrometry and deep semantic matching networks, the complexity and matrix effect problems in the detection of illegal food additives have been solved, enabling rapid and accurate quantitative and qualitative detection, and improving the ability to identify novel additives and the versatility of detection.

CN122432648APending Publication Date: 2026-07-21GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
Filing Date
2026-06-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies for detecting illegal food additives are complex, time-consuming, have incompatible data across instruments, lack the ability to identify unknown illegal additives, and the matrix effect significantly affects quantitative accuracy.

Method used

Non-destructive detection was performed using a gas-phase assisted desorption chemical ionization mass spectrometer. By combining spectral normalization and feature discretization, a deep semantic matching network was used to match unknown mass spectrometry signals. An adversarial domain adaptive quantitative prediction model was constructed and a multi-attention mechanism was introduced. Graph neural networks were then used for structural analysis.

Benefits of technology

It enables rapid and accurate detection of illegal food additives, improves the ability to discover novel additives and the quantitative accuracy in complex matrices, and ensures the transferability and universality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432648A_ABST
    Figure CN122432648A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of food safety detection, and particularly discloses an AI-assisted mass spectrometry qualitative and quantitative detection method for food illegal additives, which comprises mass spectrometry fingerprint acquisition, fingerprint preprocessing and feature extraction, illegal additive qualitative screening, quantitative prediction model construction, unknown illegal additive discovery, and detection result output and tracing. The scheme is used for non-destructive and rapid detection of food samples, maps mass spectrometry data to a unified feature space through fingerprint normalization and feature discretization; a deep semantic matching network based on contrast learning is used to realize intelligent matching of unknown mass spectrometry signals, a graph neural network is combined for structure analysis, and the discovery capability for new illegal additives is improved; an adversarial domain adaptation technology is used to eliminate the distribution difference between different food matrices, a quantitative prediction model is constructed, a multi-attention mechanism is combined to adaptively extract feature peaks related to the concentration of illegal additives, and the quantitative accuracy and generalization capability in a complex matrix are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of food safety testing technology, specifically to an AI-assisted mass spectrometry qualitative and quantitative detection method for illegal food additives. Background Technology

[0002] Illegal food additives refer to substances added during food production and processing without national approval or exceeding permitted limits, such as melamine, Sudan Red, formaldehyde, and borax. These substances pose serious threats to human health; therefore, achieving rapid, accurate, and non-destructive detection of illegal additives is of great significance for food safety supervision.

[0003] Currently, commonly used methods for detecting illegal food additives mainly include high-performance liquid chromatography (HPLC), gas chromatography-mass spectrometry (GC-MS), and liquid chromatography-mass spectrometry (LC-MS). While these methods offer high sensitivity and accuracy, they suffer from the following technical limitations: Traditional detection methods typically require complex sample pretreatment processes, including extraction, purification, and derivatization. These processes are cumbersome and time-consuming, making them unsuitable for rapid on-site screening. The pretreatment process may result in the loss of the target substance or the introduction of interfering substances. Furthermore, mass spectrometry data collected by different instruments can vary, affecting the accuracy of the detection results. Existing mass spectrometry detection methods lack the ability to identify unknown illegal additives. When illegal additives are not on the pre-set detection list, traditional methods struggle to detect and identify unknown substances, leading to a risk of missed detections. Moreover, the complex composition of food matrices and the significant differences in matrix interference between different samples necessitate the establishment of separate standard curves for each matrix using traditional quantitative methods, resulting in a large workload and poor applicability. Summary of the Invention

[0004] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides an AI-assisted mass spectrometry method for the qualitative and quantitative detection of illegal food additives. Addressing the problems of complex and time-consuming preprocessing and cross-instrument data incompatibility in traditional detection methods, this solution employs a gas-phase assisted desorption-chemical ionization mass spectrometry (GC-CMS) device to achieve non-destructive and rapid detection of food samples. Through spectral normalization and feature discretization, mass spectrometry data collected from different instruments are mapped to a unified feature space, ensuring model transferability and method universality. To address the insufficient ability to identify unknown illegal additives, this solution uses a deep semantic matching network based on contrastive learning to achieve intelligent matching of unknown mass spectrometry signals, combined with graph neural networks for structural analysis, improving the ability to discover novel illegal additives. To address the significant matrix effects and poor quantitative accuracy, this solution employs adversarial domain adaptation technology to eliminate distribution differences between different food matrices, constructing a universal quantitative prediction model. Combined with a multi-attention mechanism, it adaptively extracts feature peaks related to the concentration of illegal additives, significantly improving quantitative accuracy and generalization ability in complex matrices.

[0005] The technical solution adopted in this invention is as follows: This invention provides an AI-assisted mass spectrometry qualitative and quantitative detection method for illegal food additives, which includes the following steps:

[0006] Step S1: Mass spectrometry fingerprint acquisition. In-situ, non-destructive mass spectrometry analysis of food samples is performed using a gas phase-assisted desorption / chemical ionization mass spectrometer to obtain the original mass spectrometry fingerprint.

[0007] Step S2: Spectrum preprocessing and feature extraction. After baseline correction, peak alignment and noise reduction preprocessing of the original mass spectrometry fingerprint spectrum, mass spectrometry feature peaks are extracted, and sample feature vectors are constructed based on the extracted mass spectrometry feature peaks.

[0008] Step S3: Qualitative screening of illegal additives. Input the sample feature vector into the trained deep semantic matching network and compare it with the standard feature vector in the illegal additive mass spectrometry reference database. Output a list of suspected illegal additives and matching confidence, as well as food samples that did not match any known illegal additives.

[0009] Step S4: Quantitative prediction model construction. An illegal additive concentration prediction model is constructed using adversarial domain adaptation and multi-attention mechanism to predict the content of the screened illegal additives and output the concentration prediction value.

[0010] Step S5: Discovery of unknown illegal additives. For food samples that cannot be matched with known illegal additives, a graph neural network is used to analyze the structure of fragment ion spectra and infer the structural category of unknown compounds.

[0011] Step S6: Output the detection results. Combine the qualitative and quantitative results with the structural category inference results of the unknown compounds to generate a detection report.

[0012] Further, in step S2, the map preprocessing and feature extraction specifically include the following steps:

[0013] Step S21: Baseline correction. The original mass spectrometry fingerprint spectrum is corrected using an asymmetric least squares baseline correction algorithm to eliminate background noise and baseline drift, and the corrected spectrum is obtained.

[0014] Step S22: Peak alignment. The dynamic time warping algorithm is used to align the mass axis of the corrected spectrum. The signals on the mass axis of the corrected spectrum of different food samples are arranged in ascending order of mass-charge ratio. The mass spectrum peaks are aligned to a unified mass coordinate. The resulting signal intensity value sequence is the mass sequence, and the aligned spectrum is generated.

[0015] Step S23: Peak detection and feature extraction. Continuous wavelet transform is used to identify mass spectrometry characteristic peaks from the aligned spectra. The parameters of each characteristic peak include mass-to-charge ratio, peak height, and peak area. Each food sample after peak detection is represented as a set of characteristic peaks, using the following formula: ; ; ; ;

[0016] In the formula, Indicates the index of the detected feature peak. Indicates the first One characteristic peak, Indicates the first The mass-charge ratio of each characteristic peak This represents a single scan point within the characteristic peak. Indicates the first The mass-to-charge ratio of each scan point Indicates the first Signal strength at each scan point Indicates the first The height of each characteristic peak Indicates the first The area of ​​each characteristic peak, This indicates the mass-to-charge ratio interval between adjacent scan points. Represents the set of characteristic peaks of a food sample. This represents the total number of detected characteristic peaks;

[0017] Step S24: Sample feature vector construction. Discretize the set of feature peaks according to a fixed mass-charge ratio interval. Sum the areas of the feature peaks in each interval to obtain a sample feature vector of fixed dimensions. The component of the mass-charge ratio interval into which no feature peak falls is 0.

[0018] Furthermore, in step S3, the qualitative screening of illegal additives specifically includes the following steps:

[0019] Step S31: Obtain the mass spectrometry reference database of illegal additives and store the standard feature vector of each compound;

[0020] Step S32: Construct a deep semantic matching network, which includes two sub-networks: a sample encoder and a reference encoder. The two sub-networks share weight parameters and are used to calculate the matching score between the sample feature vector and the standard feature vector. The specific steps are as follows:

[0021] Step S321: The sample encoder maps the sample feature vector to the embedding space and uses a three-layer fully connected network to implement the encoding function. Each layer is followed by batch normalization and ReLU activation function to generate the sample encoded feature vector.

[0022] Step S322: The reference encoder maps the standard feature vector to the embedding space to generate the standard encoded feature vector;

[0023] Step S323: Calculate the matching score between the sample feature vector and the standard feature vector using cosine similarity;

[0024] Step S33: Train the deep semantic matching network. The training data is divided into positive sample pairs and negative sample pairs: positive sample pairs are pairs of sample feature vectors and their standard feature vectors for the same illegal additive at different concentrations and in different matrices; negative sample pairs are pairs of sample feature vectors for different illegal additives. Set the learning rate, batch size, and number of training epochs. Use the Adam optimizer and contrastive loss function to train the network. The formulas used are as follows: ;

[0025] In the formula, Indicates comparative loss, Represents the sample feature vector. Indicates the matching score. Represents the characteristic vector of the sample Standard feature vectors of the same type Standard feature vectors representing different classes Indicates the temperature coefficient. This represents the natural exponential function. Represent the natural logarithm function;

[0026] Step S34: Qualitative screening. Use the trained deep semantic matching network to calculate the matching score between the sample feature vector and the standard feature vector. Set a similarity threshold. Sort the matching scores of the sample feature vector and all standard feature vectors from high to low. Output a list of illegal additives with matching scores exceeding the similarity threshold and the matching confidence. For food samples that do not match any known illegal additives, proceed to step S5.

[0027] Further, in step S4, the construction of the quantitative prediction model specifically includes the following steps:

[0028] Step S41: Collect mass spectrometry data of different concentrations of illegal additives added to different food matrices as training sets. Each training sample in the training set contains a sample feature vector and the corresponding true concentration value.

[0029] Step S42: Construct a quantitative prediction model based on adversarial domain adaptation, comprising three modules: feature extractor, concentration predictor, and domain discriminator, with the following structure:

[0030] Step S421: The feature extractor consists of two fully connected layers. The first layer maps the input sample feature vector to 512 dimensions, and the second layer maps it to 64-dimensional deep features. Each layer is followed by batch normalization and ReLU activation function.

[0031] Step S422: The concentration predictor consists of two fully connected layers. The first layer maps 64-dimensional deep features to 32-dimensional features, and the second layer maps to 1-dimensional concentration prediction values. The output layer uses the softplus activation function to ensure that the concentration prediction value is positive.

[0032] Step S423: The domain discriminator determines which food matrix domain the 64-dimensional deep feature originates from. This consists of two fully connected layers: the first layer maps the 64-dimensional deep feature to 32 dimensions, and the second layer maps it to... Dimensional output, This represents the number of food matrix types, and the output layer uses the softmax activation function to obtain the domain label probability vector;

[0033] Step S43: Introduce a multi-head attention module into the feature extractor, assigning different weights to different quality ranges of the sample feature vector. The multi-head attention module generates a weighted feature vector, which replaces the sample feature vector and is input into the feature extractor. The formula used is as follows: ; ;

[0034] In the formula, Indicates the number of attention heads. Indicates the index of the attention head. Indicates the first The first attention head to the first The weights of each quality interval This represents a vector of learnable attention parameters. Indicates the first Characteristic values ​​of a quality interval Indicates the first The offset of each attention head, [⋅;⋅] represents the vector concatenation operation. Represents the weighted eigenvector;

[0035] Step S44: Construct the total loss function, including concentration prediction loss, domain discrimination loss, and domain confusion loss. By minimizing the concentration prediction loss and adversarial training, the feature extractor learns concentration-related features independent of the food matrix. The formula used is as follows: ; ; ; ;

[0036] In the formula, , and These represent the concentration prediction loss, domain discrimination loss, and domain confusion loss, respectively. Represents the total loss function. This represents the total number of training samples in the training set. Indicates the index of the training sample. Indicates the first Concentration prediction values ​​for each training sample. This represents the actual concentration value. Indicate the domain discriminative loss weights, Represents the cross-entropy loss function. Representation domain discriminator, Indicates feature extractor, Indicates the first The weighted feature vector of each training sample. Indicates the first The true domain labels of each training sample Indicates the weight of the resistance loss. Indicates a uniformly distributed label;

[0037] Step S45: The model training adopts a two-stage optimization strategy: In the first stage, the parameters of the domain discriminator are fixed, the parameters of the feature extractor and the concentration predictor are updated, and the concentration prediction loss is minimized; In the second stage, the parameters of the feature extractor and the concentration predictor are fixed, the parameters of the domain discriminator are updated, the domain discrimination loss is maximized, and the domain confusion loss is minimized at the same time; The alternating iteration continues until convergence. After training is completed, the concentration prediction value is directly output from the sample feature vector of the unknown food through the feature extractor and the concentration predictor.

[0038] Furthermore, in step S5, the discovery of the unknown illegal additive specifically includes the following steps:

[0039] Step S51: For food samples that do not match known illegal additives, extract the abnormal characteristic peak set that does not match the known illegal additives. Use the tandem mass spectrometry mode of the mass spectrometer to collect the secondary mass spectrometry fragment ion spectrum of the unknown compound in the collision energy range of 15~45eV by using the collision-induced dissociation method on the abnormal characteristic peak set.

[0040] Step S52: Construct a fragment ion structure analytical model based on a graph neural network. The fragment ion spectrum of the unknown compound is modeled as an undirected graph, including a node set and an edge set. Each node in the node set corresponds to a fragment ion, and node features include mass-to-charge ratio, relative abundance, and neutral mass loss. Each edge in the edge set connects two fragment ions that may have a parent-child relationship, and edge features include mass difference and fragmentation type probability. Iteratively update the node embedding graph neural network. The node features are linearly transformed to obtain the initial embedding vector. The readout function is used to calculate the full graph representation vector after graph convolution, using the following formula: ; ; ;

[0041] In the formula, Represents an undirected graph. Represents a set of nodes. Denotes the set of edges. Represents a node. , and They represent the first Layer, First Layer and first Nodes after layer graph convolution Embedded vector, Represents nodes The set of adjacent nodes, Indicates the connection node and The edge feature vectors, , , Indicates the first Layer-learnable weight matrix, Represents the linear rectification activation function. Indicates the number of convolutional layers. Represents the full graph representation vector;

[0042] Step S53: Construct a classifier using the softmax function. Input the full image representation vector and output the structural category of the unknown compound. Simultaneously output the predicted molecular weight and the range of the mass-charge ratio. The range of the mass-charge ratio is ±0.2 Da of the predicted molecular weight. The formula used is as follows: ;

[0043] In the formula, This indicates that the unknown compound belongs to the first... The probability of class structure category, and Indicates the first The class's weight vector and bias, Indicates the total number of structure categories;

[0044] Step S54: Perform targeted high-performance liquid chromatography-high-resolution mass spectrometry to verify the structure category and confirm the precise structure and concentration of the unknown compound; after verification as a newly discovered illegal additive, add its sample feature vector and related information to the illegal additive mass spectrometry reference database, and record the first discovery time and food matrix source of the compound for subsequent automated screening.

[0045] The beneficial effects achieved by the present invention using the above solution are as follows:

[0046] (1) In view of the problems of complex pretreatment, long time consumption and incompatibility of cross-instrument data in traditional detection methods, this scheme adopts gas phase-assisted desorption chemical ionization mass spectrometry to realize non-destructive rapid detection of food samples. Through spectrum normalization and feature discretization, mass spectrometry data collected by different instruments are mapped to a unified feature space to ensure the transferability of the model and the universality of the method.

[0047] (2) To address the problem of insufficient ability to identify unknown illegal additives, this solution adopts a deep semantic matching network based on contrastive learning to achieve intelligent matching of unknown mass spectrometry signals, and combines it with graph neural network for structural analysis to improve the ability to discover new illegal additives.

[0048] (3) To address the issues of significant matrix effects and poor quantitative accuracy, this scheme employs adversarial domain adaptation technology to eliminate distribution differences between different food matrices, constructs a general quantitative prediction model, and combines multi-attention mechanism to adaptively extract characteristic peaks related to the concentration of illegal additives, thereby significantly improving the quantitative accuracy and generalization ability in complex matrices. Attached Figure Description

[0049] Figure 1 This is a schematic diagram of the process for an AI-assisted mass spectrometry qualitative and quantitative detection method for illegal food additives proposed in this invention.

[0050] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0051] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0052] Example 1, see Figure 1 This invention provides an AI-assisted mass spectrometry qualitative and quantitative detection method for illegal food additives, which includes the following steps:

[0053] Step S1: Mass spectrometry fingerprint acquisition. In-situ, non-destructive mass spectrometry analysis of food samples is performed using a gas phase-assisted desorption / chemical ionization mass spectrometer to obtain the original mass spectrometry fingerprint.

[0054] Step S2: Spectrum preprocessing and feature extraction. After baseline correction, peak alignment and noise reduction preprocessing of the original mass spectrometry fingerprint spectrum, mass spectrometry feature peaks are extracted, and sample feature vectors are constructed based on the extracted mass spectrometry feature peaks.

[0055] Step S3: Qualitative screening of illegal additives. Input the sample feature vector into the trained deep semantic matching network and compare it with the standard feature vector in the illegal additive mass spectrometry reference database. Output a list of suspected illegal additives and matching confidence, as well as food samples that did not match any known illegal additives.

[0056] Step S4: Quantitative prediction model construction. An illegal additive concentration prediction model is constructed using adversarial domain adaptation and multi-attention mechanism to predict the content of the screened illegal additives and output the concentration prediction value.

[0057] Step S5: Discovery of unknown illegal additives. For food samples that cannot be matched with known illegal additives, a graph neural network is used to analyze the structure of fragment ion spectra and infer the structural category of unknown compounds.

[0058] Step S6: Output the detection results. Combine the qualitative and quantitative results with the structural category inference results of the unknown compounds to generate a detection report.

[0059] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, the mass spectrometry fingerprint acquisition specifically includes the following steps:

[0060] Step S11: Construct a gas-phase assisted desorption chemical ionization mass spectrometry device. The device includes a carrier gas control system, a humidification unit, a sample chamber, and a mass spectrometer. Nitrogen is used as the carrier gas, and the flow rate is controlled at 25 mL / min through a precision gas flow valve. The nitrogen flows through the humidification unit containing deionized water, and the water vapor volume content is maintained at 70% to assist in the desorption and ionization of compounds on the sample surface.

[0061] Step S12: Place a 5.00g food sample directly into the sample chamber without any grinding, extraction, or derivatization pretreatment. The temperature of the sample chamber is controlled at 25℃, and the relative humidity is maintained at 60%.

[0062] Step S13: Set the mass spectrometry analysis parameters, adopt positive ion detection mode, capillary voltage of 4.5kV, ion transmission tube temperature of 250℃, mass spectrometry scanning range of m / z 50~1000, scanning speed of 0.5 seconds / s, collect three consecutive samples for each food sample and take the average spectrum as the original mass spectrometry fingerprint spectrum of the food sample, and the total acquisition time for each food sample shall not exceed 1 minute.

[0063] Example 3, see Figure 1 This embodiment is based on the above embodiment. In step S2, the map preprocessing and feature extraction specifically include the following steps:

[0064] Step S21: Baseline correction. The original mass spectrometry fingerprint spectrum is corrected using an asymmetric least squares baseline correction algorithm to eliminate background noise and baseline drift, resulting in the corrected spectrum. The formula used is as follows: ; ; ; ;

[0065] In the formula, This represents the signal intensity sequence of the original mass spectrometry fingerprint. This indicates the number of mass spectrometry scan points, with a value of 9500. This represents the estimated baseline signal sequence. This represents the optimal baseline sequence. This represents the function that minimizes the parameters. Indicates the index of the scan point. Indicates the first The weighting coefficient of each scan point, when If the value is 0.5, then use 10; This represents the smoothing parameter, with a value of 10. 6 , This indicates the corrected spectrum;

[0066] Step S22: Peak Alignment. A dynamic time warping algorithm is used to align the mass axis of the corrected spectra. The signals on the mass axis of the corrected spectra of different food samples are arranged in ascending order of mass-charge ratio. The mass spectral peaks are aligned to a unified mass coordinate system, resulting in a signal intensity value sequence as the mass sequence. The aligned spectra are generated using the following formula: ;

[0067] In the formula, Indicates the first two quality sequences One mass point and the front The cumulative distance between mass points and These represent the reference spectrum and the spectrum to be aligned at the [number]th [number]th [number]. The and the first Signal strength at each mass point;

[0068] Step S23: Peak detection and feature extraction. Continuous wavelet transform is used to identify mass spectrometry characteristic peaks from the aligned spectra. The parameters of each characteristic peak include mass-to-charge ratio, peak height, and peak area. Each food sample after peak detection is represented as a set of characteristic peaks, using the following formula: ; ; ; ;

[0069] In the formula, Indicates the index of the detected feature peak. Indicates the first One characteristic peak, Indicates the first The mass-charge ratio of each characteristic peak This represents a single scan point within the characteristic peak. Indicates the first The mass-to-charge ratio of each scan point Indicates the first Signal strength at each scan point Indicates the first The height of each characteristic peak Indicates the first The area of ​​each characteristic peak, This indicates the mass-to-charge ratio interval between adjacent scan points. Represents the set of characteristic peaks of a food sample. This represents the total number of detected characteristic peaks;

[0070] Step S24: Sample feature vector construction. The set of feature peaks is discretized according to a fixed mass-charge ratio interval, with each interval having a width of 0.1 Da. The areas of the feature peaks within each interval are summed to obtain a sample feature vector with fixed dimensions. The number of intervals with mass-to-charge ratio is denoted as Let's take 9500. The formula used is as follows: ;

[0071] In the formula, Indices representing quality ranges Represents the sample feature vector The One portion, Indicates the first For each quality interval, there is no characteristic peak falling within that interval. The value is 0.

[0072] By performing the aforementioned operations, this solution addresses the problems of complex and time-consuming pretreatment and cross-instrument data incompatibility in traditional detection methods. It employs a gas-phase assisted desorption chemical ionization mass spectrometry device to achieve non-destructive and rapid detection of food samples. Through spectral normalization and feature discretization, mass spectrometry data collected by different instruments are mapped to a unified feature space, ensuring the transferability of the model and the universality of the method.

[0073] Example 4, see Figure 1 This embodiment is based on the above embodiment. In step S3, the qualitative screening of illegal additives specifically includes the following steps:

[0074] Step S31: Obtain the mass spectrometry reference database of illegal additives and store the standard feature vector of each compound;

[0075] Step S32: Construct a deep semantic matching network, which includes two sub-networks: a sample encoder and a reference encoder. The two sub-networks share weight parameters and are used to calculate the matching score between the sample feature vector and the standard feature vector. The specific steps are as follows:

[0076] Step S321: The sample encoder maps the sample feature vector to the embedding space, and uses a three-layer fully connected network to implement the encoding function. The number of neurons in each layer is 512, 256, and 128, respectively. Each layer is followed by batch normalization and ReLU activation function. The formula used is as follows: ;

[0077] In the formula, Represents the sample feature vector. Represents the sample encoded feature vector. Represents the encoding function. This represents the trainable parameters of the sample encoder. Let be the set of real numbers. This represents the dimension of the embedding vector, with a value of 128.

[0078] Step S322: The reference encoder maps the standard feature vectors to the embedding space using the following formula: ;

[0079] In the formula, Represents the standard eigenvector. Represents the standard encoded feature vector;

[0080] Step S323: Calculate the matching score between the sample encoded feature vector and the standard encoded feature vector using cosine similarity. The formula used is as follows: ;

[0081] In the formula, This represents the matching score between the sample encoded feature vector and the standard encoded feature vector;

[0082] Step S33: Train the deep semantic matching network. The training data is divided into positive sample pairs and negative sample pairs: positive sample pairs are pairs of sample feature vectors and their standard feature vectors for the same illegal additive at different concentrations and in different matrices; negative sample pairs are pairs of sample feature vectors for different illegal additives. The learning rate is initialized to 0.001, the batch size is 64, and the training is performed for 50 epochs. The Adam optimizer and contrastive loss function are used to train the network. The formulas used are as follows: ;

[0083] In the formula, Indicates comparative loss, Represents the characteristic vector of the sample Standard feature vectors of the same type Standard feature vectors representing different classes This represents the temperature coefficient, with a value of 0.1. This represents the natural exponential function. Represent the natural logarithm function;

[0084] Step S34: Qualitative screening. Using a trained deep semantic matching network, calculate the matching score between the sample feature vector and the standard feature vector. Set a similarity threshold of 0.75. Sort the matching scores of the sample feature vector with all standard feature vectors from high to low. Output a list of illegal additives with matching scores exceeding the similarity threshold and their matching confidence scores. For food samples that do not match any known illegal additives, proceed to step S5, which is as follows: like A score ≥0.85 is considered a high match confidence level and is output directly. If 0.75≤ <0.85, which is judged as low matching confidence, and proceeds to step S4 for quantitative confirmation; like If the value is less than 0.75, it is determined that no known illegal additives were matched, and the process proceeds to step S5 to discover unknown illegal additives.

[0085] Example 5, see Figure 1This embodiment is based on the above embodiment. In step S31, the illegal additive mass spectrometry reference database contains standard information for a total of 156 known illegal additives, including melamine, Sudan I-IV, Rhodamine B, Acid Orange II, Basic Orange II, Sodium formaldehyde sulfoxylate, borax, sodium formaldehyde sulfoxylate, malachite green, crystal violet, ractopamine, clenbuterol, salbutamol, tetracyclines, and sulfonamides. The data stored for each illegal additive includes: Name, chemical formula, CAS number, and molecular weight of the compound; Standard mass spectra acquired in positive ion mode include the mass-charge ratio and relative abundance of the parent ion and no less than 5 characteristic fragment ions. Mass spectrometry information for different adduct ion forms; Reference values ​​for the limits of detection and limits of quantitation of this compound in different food matrices.

[0086] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S4, the construction of the quantitative prediction model specifically includes the following steps:

[0087] Step S41: Collect mass spectrometry data of different concentrations of illegal additives added to different food matrices as training sets. Each illegal additive is added at 6 concentration levels in 5 typical food matrices: the 5 typical food matrices include milk, beverages, meat products, condiments, and grains and oils, and the 6 concentration levels are 0.1, 0.5, 1, 5, 10, and 50 mg / kg, respectively; each concentration level is repeated 3 times, resulting in a total of 156×5×6×3=14,040 training samples. Each training sample contains a sample feature vector and the corresponding true concentration value.

[0088] Step S42: Construct a quantitative prediction model based on adversarial domain adaptation, including three modules: feature extractor, concentration predictor, and domain discriminator.

[0089] Step S421: The feature extractor consists of two fully connected layers. The first layer maps the input sample feature vector to 512 dimensions, and the second layer maps it to 64-dimensional deep features. Each layer is followed by batch normalization and the ReLU activation function, using the following formula: ;

[0090] In the formula, Indicates feature extractor, Indicates the parameters of the feature extractor. Indicates the dimension of deep features;

[0091] Step S422: The concentration predictor consists of two fully connected layers. The first layer maps the 64-dimensional depth features to 32 dimensions, and the second layer maps them to 1-dimensional concentration predictions. The output layer uses the softplus activation function to ensure that the concentration predictions are positive. The formula used is as follows: ;

[0092] In the formula, Indicates concentration predictor, This represents the parameters of the concentration predictor. This represents the predicted concentration value;

[0093] Step S423: The domain discriminator determines which food matrix domain the 64-dimensional deep feature originates from. This consists of two fully connected layers: the first layer maps the 64-dimensional deep feature to 32 dimensions, and the second layer maps it to... Dimensional output, The output layer uses the softmax activation function to obtain the domain label probability vector, and the formula used is as follows: ;

[0094] In the formula, Representation domain discriminator, The parameters of the domain discriminator, Represents the domain label probability vector;

[0095] Step S43: Introduce a multi-head attention module into the feature extractor, assigning different weights to different quality ranges of the sample feature vector. The multi-head attention module generates a weighted feature vector, which replaces the sample feature vector and is input into the feature extractor. The formula used is as follows: ; ;

[0096] In the formula, This indicates the number of attention heads, with a value of 8. Indicates the index of the attention head. Indicates the first The first attention head to the first The weights of each quality interval This represents a vector of learnable attention parameters. Indicates the first Characteristic values ​​of a quality interval Indicates the first The offset of each attention head, with values ​​corresponding to different widths of the quality range being {0, 10, 20, 50, 100, 200, 500, 1000}, and [⋅;⋅] representing vector concatenation operations. Represents the weighted eigenvector;

[0097] Step S44: Construct the total loss function, including concentration prediction loss, domain discrimination loss, and domain confusion loss. By minimizing the concentration prediction loss and adversarial training, the feature extractor learns concentration-related features independent of the food matrix. The formula used is as follows: ; ; ; ;

[0098] In the formula, , and These represent the concentration prediction loss, domain discrimination loss, and domain confusion loss, respectively. Represents the total loss function. This represents the total number of training samples in the training set. Indicates the index of the training sample. Indicates the first Concentration prediction values ​​for each training sample. This represents the actual concentration value. The domain discrimination loss weight is represented by a value of 0.1. Represents the cross-entropy loss function. Indicates the first The ground truth domain labels for each training sample are represented by a one-hot vector of length 5. This represents the adversarial loss weight, with a value of 0.05. To represent uniformly distributed labels, use a vector of length 5, where each element is 1 / 5;

[0099] Step S45: The model training adopts a two-stage optimization strategy: In the first stage, the parameters of the domain discriminator are fixed, and the parameters of the feature extractor and concentration predictor are updated to minimize the concentration prediction loss; in the second stage, the parameters of the feature extractor and concentration predictor are fixed, and the parameters of the domain discriminator are updated to maximize the domain discrimination loss while minimizing the domain confusion loss; the alternating iteration is performed for 100 rounds, with 10 batches trained in the first stage and 5 batches trained in the second stage in each round; after training, the concentration prediction value is directly output from the sample feature vector of the unknown food through the feature extractor and concentration predictor.

[0100] By performing the aforementioned operations, this solution addresses the issues of significant matrix effects and poor quantitative accuracy by employing adversarial domain adaptation technology to eliminate distribution differences between different food matrices, constructing a universal quantitative prediction model, and combining multi-attention mechanism to adaptively extract characteristic peaks related to the concentration of illegal additives, thereby significantly improving the quantitative accuracy and generalization ability in complex matrices.

[0101] Example 7, see Figure 1 This embodiment is based on the above embodiment. In step S5, the discovery of the unknown illegal additive specifically includes the following steps:

[0102] Step S51: For food samples that do not match known illegal additives, extract the set of abnormal characteristic peaks that do not match the known illegal additives. Using the tandem mass spectrometry mode of the mass spectrometer, use collision-induced dissociation to collect the secondary mass spectrometry fragment ion spectrum of the unknown compound in the collision energy range of 15~45 eV. The formula used is as follows: ;

[0103] In the formula, Represents a set of anomalous characteristic peaks. This represents the set of characteristic peak masses of all known compounds in the database, including the parent ion and characteristic fragment ion masses of each compound, with a tolerance of ±0.2 Da;

[0104] Step S52: Construct a fragment ion structure analysis model based on a graph neural network. The fragment ion spectrum of the unknown compound is modeled as an undirected graph, including a node set and an edge set. Each node in the node set corresponds to a fragment ion. Node features include mass-charge ratio, relative abundance, and neutral loss mass. The mass-charge ratio is normalized to the [0,1] interval. The relative abundance is the proportion relative to the base peak, ranging from 0 to 1. The neutral loss mass is the normalized mass difference relative to the parent ion. Each edge in the edge set connects two fragment ions that may have a parent-child relationship. Edge features include mass difference and breakage type probability. The mass difference is normalized, and the breakage type probability is the chemical bond breakage type distribution predicted by the pre-trained model. Iteratively update the node embedding graph neural network. The node features undergo linear transformation to obtain an initial embedding vector of dimension 64. The readout function is used to calculate the full graph representation vector after graph convolution. The formula used is as follows: ; ; ;

[0105] In the formula, Represents an undirected graph. Represents a set of nodes. Denotes the set of edges. Represents a node. , and They represent the first Layer, First Layer and first Nodes after layer graph convolution Embedded vector, Represents nodes The set of adjacent nodes, Indicates the connection node and The edge feature vectors have a dimension of 32. , , Indicates the first Layer-learnable weight matrix, Represents the linear rectification activation function. Indicates the number of convolutional layers. Represents the full graph representation vector;

[0106] Step S53: A classifier is constructed using the softmax function. The input is the full-image representation vector, and the output is the structural category of the unknown compound. There are 15 categories: azo dyes, triphenylmethane dyes, aromatic amines, sulfonamides, quinolones, tetracyclines, β-receptor agonists, nitrofurans, hormones, bisphenols, phthalates, organophosphates, carbamates, benzimidazoles, and others. Simultaneously, the predicted molecular formula and mass-charge ratio range are inferred through precise mass matching and elemental composition rules. The mass-charge ratio range is taken as ±0.2 Da of the predicted molecular weight, using the following formula: ;

[0107] In the formula, This indicates that the unknown compound belongs to the first... The probability of class structure category, and Indicates the first The class's weight vector and bias, This represents the total number of structure categories, with a value of 15.

[0108] Step S54: Perform targeted high-performance liquid chromatography-high-resolution mass spectrometry to verify the structure category and confirm the precise structure and concentration of the unknown compound; after verification as a newly discovered illegal additive, add its mass spectrometry feature vector and related information to the illegal additive mass spectrometry reference database, and record the first discovery time and food matrix source of the compound.

[0109] By performing the aforementioned operations, this solution addresses the problem of insufficient ability to identify unknown illegal additives. It employs a deep semantic matching network based on contrastive learning to achieve intelligent matching of unknown mass spectrometry signals, combined with graph neural networks for structural analysis, thereby improving the ability to detect novel illegal additives.

[0110] Example 8, see Figure 1 This embodiment is based on the above embodiment. In step S6, the detection result is output, which specifically includes the following steps:

[0111] Step S61: For each food sample tested, generate a test report. The report content includes:

[0112] Unique sample identifier, testing time, and testing personnel;

[0113] Qualitative screening results: Names of detected illegal additives, matching confidence levels, and a list of corresponding characteristic peaks;

[0114] Quantitative prediction results: predicted concentration values ​​and 95% confidence intervals for each detected illegal additive;

[0115] Unknown compound discovery tips: If there are mismatched abnormal mass spectrometry signals, indicate the structural category of the suspected unknown substance and suggest verification methods;

[0116] Overall assessment conclusion: Pass / Fail / Requires re-inspection;

[0117] Step S62: Write the detection results into the blockchain evidence storage system, and generate a unique hash value for each detection result for tamper-proof verification;

[0118] Step S63: Regularly summarize and analyze the test data, statistically analyze the detection rate, concentration distribution, and associated food categories of various illegal additives, and generate a food safety risk warning report; for newly confirmed illegal additives, after expert review, their mass spectrometry characteristic data are synchronously updated to the cloud reference database, and all networked testing terminals can obtain the updated database in real time, realizing dynamic expansion of the database and online incremental learning of the model.

[0119] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0120] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0121] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A qualitative and quantitative detection method for illegal food additives using mass spectrometry, characterized in that: The method includes the following steps: Step S1: Mass spectrometry fingerprint acquisition. In-situ, non-destructive mass spectrometry analysis of food samples is performed using a gas phase-assisted desorption / chemical ionization mass spectrometer to obtain the original mass spectrometry fingerprint. Step S2: Spectrum preprocessing and feature extraction. After baseline correction, peak alignment and noise reduction preprocessing of the original mass spectrometry fingerprint spectrum, mass spectrometry feature peaks are extracted, and sample feature vectors are constructed based on the extracted mass spectrometry feature peaks. Step S3: Qualitative screening of illegal additives. Input the sample feature vector into the trained deep semantic matching network and compare it with the standard feature vector in the illegal additive mass spectrometry reference database. Output a list of suspected illegal additives and matching confidence, as well as food samples that did not match any known illegal additives. Step S4: Quantitative prediction model construction. An illegal additive concentration prediction model is constructed using adversarial domain adaptation and multi-attention mechanism to predict the content of the screened illegal additives and output the concentration prediction value. Step S5: Discovery of unknown illegal additives. For food samples that cannot be matched with known illegal additives, a graph neural network is used to analyze the structure of fragment ion spectra and infer the structural category of unknown compounds. Step S6: Output the detection results. Combine the qualitative and quantitative results with the structural category inference results of the unknown compounds to generate a detection report.

2. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 1, characterized in that: In step S2, the map preprocessing and feature extraction specifically include the following steps: Step S21: Baseline correction. The original mass spectrometry fingerprint spectrum is corrected using an asymmetric least squares baseline correction algorithm to eliminate background noise and baseline drift, and the corrected spectrum is obtained. Step S22: Peak alignment. The dynamic time warping algorithm is used to align the mass axis of the corrected spectrum. The signals on the mass axis of the corrected spectrum of different food samples are arranged in ascending order of mass-charge ratio. The mass spectrum peaks are aligned to a unified mass coordinate. The resulting signal intensity value sequence is the mass sequence, and the aligned spectrum is generated. Step S23: Peak detection and feature extraction. The continuous wavelet transform method is used to identify mass spectrometry characteristic peaks from the aligned spectrum. The parameters of each characteristic peak include mass-charge ratio, peak height and peak area. Each food sample after peak detection is represented as a set of characteristic peaks. Step S24: Sample feature vector construction. Discretize the set of feature peaks according to a fixed mass-charge ratio range, sum the feature peak areas in each range, and obtain a sample feature vector with fixed dimensions.

3. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 1, characterized in that: In step S3, the qualitative screening of illegal additives specifically includes the following steps: Step S31: Obtain the mass spectrometry reference database of illegal additives and store the standard feature vector of each compound; Step S32: Construct a deep semantic matching network, which includes two sub-networks: a sample encoder and a reference encoder. The two sub-networks share weight parameters and are used to calculate the matching score between the sample feature vector and the standard feature vector. Step S33: Train a deep semantic matching network. The training data is divided into positive sample pairs and negative sample pairs. Positive sample pairs are the pairing of the sample feature vectors of the same illegal additive in different concentrations and matrices with their standard feature vectors. Negative sample pairs are the pairing of sample feature vectors of different illegal additives. Set the learning rate, batch size and training epochs. Use the Adam optimizer and contrastive loss function to train the network. Step S34: Qualitative screening. Use the trained deep semantic matching network to calculate the matching score between the sample feature vector and the standard feature vector. Set a similarity threshold. Sort the matching scores of the sample feature vector and all standard feature vectors from high to low. Output a list of illegal additives with matching scores exceeding the similarity threshold and the matching confidence. For food samples that do not match any known illegal additives, proceed to step S5.

4. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 3, characterized in that: In step S32, constructing the deep semantic matching network includes the following steps: Step S321: The sample encoder maps the sample feature vector to the embedding space and uses a three-layer fully connected network to implement the encoding function. Each layer is followed by batch normalization and ReLU activation function to generate the sample encoded feature vector. Step S322: The reference encoder maps the standard feature vector to the embedding space to generate the standard encoded feature vector; Step S323: Calculate the matching score between the sample feature vector and the standard feature vector using cosine similarity.

5. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 1, characterized in that: In step S4, the quantitative prediction model is constructed, specifically including the following steps: Step S41: Collect mass spectrometry data of different concentrations of illegal additives added to different food matrices as training sets. Each training sample in the training set contains a sample feature vector and the corresponding true concentration value. Step S42: Construct a quantitative prediction model based on adversarial domain adaptation, including three modules: feature extractor, concentration predictor, and domain discriminator; Step S43: Introduce a multi-head attention module into the feature extractor, assign different weights to different quality ranges of the sample feature vector, generate a weighted feature vector through the multi-head attention module, and input it into the feature extractor instead of the sample feature vector; Step S44: Construct the total loss function, which includes concentration prediction loss, domain discrimination loss and domain confusion loss. By minimizing the concentration prediction loss and adversarial training, the feature extractor learns concentration-related features that are independent of the food matrix. Step S45: The model training adopts a two-stage optimization strategy: In the first stage, the parameters of the domain discriminator are fixed, the parameters of the feature extractor and the concentration predictor are updated, and the concentration prediction loss is minimized; In the second stage, the parameters of the feature extractor and the concentration predictor are fixed, the parameters of the domain discriminator are updated, the domain discrimination loss is maximized, and the domain confusion loss is minimized at the same time; The alternating iteration continues until convergence. After training is completed, the concentration prediction value is directly output from the sample feature vector of the unknown food through the feature extractor and the concentration predictor.

6. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 5, characterized in that: In step S42, the quantitative prediction model has the following structure: Step S421: The feature extractor consists of two fully connected layers. The first layer maps the input sample feature vector to 512 dimensions, and the second layer maps it to 64-dimensional deep features. Each layer is followed by batch normalization and ReLU activation function. Step S422: The concentration predictor consists of two fully connected layers. The first layer maps 64-dimensional deep features to 32-dimensional features, and the second layer maps to 1-dimensional concentration prediction values. The output layer uses the softplus activation function to ensure that the concentration prediction value is positive. Step S423: The domain discriminator determines which food matrix domain the 64-dimensional deep feature originates from. This consists of two fully connected layers: the first layer maps the 64-dimensional deep feature to 32 dimensions, and the second layer maps it to... Dimensional output, This represents the number of food matrix types. The output layer uses the softmax activation function to obtain the domain label probability vector.

7. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 1, characterized in that: In step S5, the discovery of the unknown illegal additive specifically includes the following steps: Step S51: For food samples that do not match known illegal additives, extract the abnormal characteristic peak set that does not match the known illegal additives. Use the tandem mass spectrometry mode of the mass spectrometer to collect the secondary mass spectrometry fragment ion spectrum of the unknown compound in the collision energy range of 15~45eV by using the collision-induced dissociation method on the abnormal characteristic peak set. Step S52: Construct a fragment ion structure analytical model based on graph neural network. Model the fragment ion spectrum of the unknown compound as an undirected graph, including a set of nodes and a set of edges. Iteratively update the node embedding graph neural network. The node features are linearly transformed to obtain the initial embedding vector. The readout function is used to calculate the full graph representation vector after graph convolution. Step S53: Construct a classifier using the softmax function, input the full image representation vector, output the structural category to which the unknown compound belongs, and output the predicted molecular weight and mass-charge ratio range; Step S54: Perform targeted high-performance liquid chromatography-high-resolution mass spectrometry to verify the structure category and confirm the precise structure and concentration of the unknown compound.

8. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 7, characterized in that: In step S52, in the graph neural network, each node in the node set corresponds to a fragment ion, and the node features include mass-to-charge ratio, relative abundance, and neutral loss mass; each edge in the edge set connects two fragment ions that may have a parent-child relationship, and the edge features include mass difference and fracture type probability.

9. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 3, characterized in that: In step S33, the contrast loss function uses the following formula: ; In the formula, Indicates comparative loss, Represents the sample feature vector. Indicates the matching score. Represents the characteristic vector of the sample Standard feature vectors of the same type Standard feature vectors representing different classes Indicates the temperature coefficient. This represents the natural exponential function. This represents the natural logarithm function.

10. The method for qualitative and quantitative detection of illegal food additives using AI-assisted mass spectrometry according to claim 5, characterized in that: In step S44, the total loss function is constructed using the following formula: ; ; ; ; In the formula, , and These represent the concentration prediction loss, domain discrimination loss, and domain confusion loss, respectively. Represents the total loss function. This represents the total number of training samples in the training set. Indicates the index of the training sample. Indicates the first Concentration prediction values ​​for each training sample. This represents the actual concentration value. Indicate the domain discriminative loss weights, Represents the cross-entropy loss function. Representation domain discriminator, Indicates feature extractor, Indicates the first The weighted feature vector of each training sample. Indicates the first The true domain labels of each training sample Indicates the weight of the resistance loss. This indicates a uniformly distributed label.