Tomato root rot pathogen volatile biomarker detection system and method based on GC-IMS technology
Through the detection system of volatile biomarker for tomato root rot pathogenic bacteria based on GC-IMS technology, the limitations of early diagnosis and precise prevention and control of tomato root rot in the existing technology have been solved, and efficient and accurate disease monitoring and prevention and control management have been achieved.
Patent Information
- Application Number
- CN202510229282.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology has limitations in the early diagnosis and precise prevention and control of tomato root rot, including time-consuming detection, non-real-time diagnosis, difficulty in realizing real-time monitoring in the field and intelligent prevention and control decisions.
The volatile biomarker detection system for tomato root rot pathogenic bacteria based on GC-IMS technology is adopted to achieve early warning and precise prevention and control through non-invasive sampling, multimodal denoising, intelligent feature extraction and closed-loop prevention and control recommendation.
It significantly improves the early diagnosis efficiency of tomato root rot, realizes the full process of closed-loop management from detection to prevention and treatment, and has the advantages of high sensitivity, non-invasive sampling, intelligent data processing and accurate diagnosis.
Smart Images

Figure CN120199336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural disease early warning, and particularly relates to a detection system and method for volatile biomarkers of the pathogenic bacteria of tomato root rot based on GC-IMS technology. Background Art
[0002] Tomato root rot is one of the important diseases threatening agricultural production. Its pathogenic bacteria infect the roots, causing the plants to wilt and even die. Traditional detection methods rely on laboratory culture or molecular biology techniques, which have limitations such as long time consumption, destructive sampling, and difficulty in early diagnosis. Volatile organic compounds (VOCs), as metabolites of plant-pathogen interactions, have the potential for non-invasive collection and real-time monitoring. However, the efficient extraction and accurate analysis of trace VOCs in complex environmental samples still face challenges. Gas chromatography-ion mobility spectrometry (GC-IMS) technology has significant advantages in gas detection due to its high resolution and sensitivity. However, its application in agricultural disease diagnosis is not yet mature, especially lacking systematic solutions in aspects such as soil gas sampling around the roots, complex noise suppression, characteristic biomarker extraction, and intelligent analysis.
[0003] The Chinese patent application with the publication number CN116413357A discloses a method for screening characteristic VOCs of grape mildew based on GC-IMS technology combined with chemometrics. The method steps are as follows: Select grapes with uniform size, full grains, and pedicels retained, and inoculate Alternaria alternata in the way of cutting wounds of 1mm×1mm. On the 1st, 3rd, 5th, and 7th days after inoculation, 3 samples are randomly selected from each group for GC-IMS detection. Chemometric and other analysis methods are used to extract features from the GC-IMS data, and the number of characteristic parameters is optimized through cross-validation. Finally, a partial least squares model between the characteristic VOCs of grape mildew and different infection days is established; the partial least squares model is used for the monitoring of grape mildew; by combining GC-IMS technology with chemometrics to characterize and monitor the changes of VOCs in mildewed grape fruits at different infection days, without complex pretreatment, high sensitivity, and the goal of monitoring grape mildew at different infection days can be achieved, providing a new monitoring method for grape diseases.
[0004] In the prior art, the processing of GC-IMS data mostly relies on traditional algorithms, with insufficient ability to suppress low-frequency noise and residual interference in ion mobility spectrometry diagrams, and limited recognition accuracy of the feature selection and classification model for complex pathogenic bacteria markers, making it difficult to meet the requirements of field real-time diagnosis. In addition, the intelligent level of agricultural detection systems in terms of data standardization verification and the linkage of prevention and control decisions is relatively low. To address the above problems, there is an urgent need to develop a GC-IMS detection system that integrates non-invasive sampling, multi-modal denoising, intelligent feature extraction, and closed-loop prevention and control recommendations to achieve early warning and precise prevention and control of tomato root rot and provide technical support for smart agriculture. Summary of the Invention
[0005] The object of the present invention is to propose a detection system and method for volatile biomarkers of pathogenic bacteria causing tomato root rot based on GC-IMS technology in view of the problems existing in the background technology.
[0006] The technical solution of the present invention: A method for detecting volatile biomarkers of pathogenic bacteria causing tomato root rot based on GC-IMS technology includes the following specific implementation steps:
[0007] S1. Collect samples of volatile chemical substances dissolved in the air in the soil;
[0008] S2. Generate an ion mobility spectrometry diagram based on gas chromatography-ion mobility spectrometry (GC-IMS) technology;
[0009] S3. Denoise based on a multi-modal denoising method. Perform multi-level time-frequency analysis on the ion mobility spectrometry diagram signal through wavelet packet transform, decompose and remove noise, and use an adaptive filter to dynamically adjust coefficients to suppress noise and restore effective signals, weight important frequency components, suppress residual noise, restore the time-domain signal through inverse wavelet packet transform, and generate a standard verification symbol through a hash function and quadratic residue operation;
[0010] S4. Conduct standardization verification through the standard verification symbol, extract peak feature parameters using the Gaussian fitting algorithm, screen key features through mutual information analysis, construct a multi-scale convolutional neural network and an adaptive attention mechanism to capture local and global features, and finally classify through a fully connected layer and output the disease diagnosis result;
[0011] S5. Automatically generate a diagnostic report based on the diagnosis result and provide corresponding prevention and control suggestions according to the disease recognition result;
[0012] S6. Visually display the diagnosis result and prevention and control suggestions, and upload all diagnostic data, disease feature information, and treatment suggestions to the cloud for storage and update.
[0013] Preferably, the denoising process of denoising based on the multi-modal denoising method is as follows:
[0014] S21. Decompose the ion mobility spectrometry signal using wavelet packet transform:
[0015]
[0016] In the formula, f(t) represents the ion mobility spectrometry signal; a n represents the coefficient after wavelet packet decomposition; ψ n (t) represents the corresponding wavelet function;
[0017] S22. Automatically adjust the filter coefficients based on the statistical characteristics of the noise through an adaptive filter for dynamic denoising:
[0018]
[0019]
[0020] w k (t + 1) = w k (t) + μ·e(t)·x(t - k);
[0021] In the formula, e(t) represents the error signal, that is, the difference between the filter output and the target signal; d(t) represents the target signal, which is the signal after wavelet packet transform denoising; represents the filter output, the estimated signal; w k (t) represents the filter coefficient; μ represents the step size, controlling the update speed of the filter coefficient; x(t - k) represents the input signal, that is, the input data of the filter, and t - k is the historical data of the signal;
[0022] S23. Remove the residual noise from the frequency-domain signal, that is, by weighting the important frequency components in the signal, so that the noise components are more effectively suppressed:
[0023] S enhanced (f) = S(f)·w(f);
[0024] In the formula, S(f) represents the signal of the ion mobility spectrometry in the frequency domain, that is, the denoised signal w(f) represents the weighting function; S enhanced (f) represents the enhanced frequency-domain signal;
[0025] S24. The signal after frequency-domain enhancement is reconstructed in the time domain to restore the time-domain representation of the original ion mobility spectrometry. Use the inverse wavelet packet transform IWPT to convert the enhanced frequency-domain signal back to the time domain:
[0026]
[0027] Wherein, I(t) represents the time-domain signal after denoising, which is output after being restored by the inverse wavelet packet transform; a n represents the enhanced coefficient, representing the signal components of each frequency band.
[0028] Preferably, the generation process of the standard verification symbol is as follows:
[0029] S31. Calculate the intermediate parameter s = α × L(H(data)) mod n;
[0030] Wherein, α is a pre-compiled generation factor, α = [lcm(p - 1, q - 1)] - 1 mod n; p and q are predefined large prime numbers, p ≡ 3 mod 8, p ≡ 7 mod 8; n is a pre-compiled encoding factor, n = p × q; data is the binary form of the denoised ion mobility spectrogram; H() represents a pre-compiled hash function; L() is a pre-compiled generation function, L(x) = (x - 1) / n; and x ≡ 1 mod n; lcm() represents the least common multiple function;
[0031] S32. Calculate the first-order verification symbol S1:
[0032] If [s / n] = 1, then S1 = 0;
[0033] If [s / n] = -1, then S1 = 1;
[0034] S33. Calculate the second-order verification symbol S2:
[0035]
[0036] Wherein, Q n represents the set of quadratic residues modulo n; is the set of non-quadratic residues;
[0037] S34. Calculate the auxiliary generation factor d = (n - p - q + 5) / 8, and calculate the third-order verification symbol S3 accordingly:
[0038]
[0039] S35. Calculate the fourth-order verification symbol S4:
[0040]
[0041] S36. Generate the standard verification symbol Sv = {S1, S2, S3, S4}.
[0042] Preferably, the verification process for standard verification using the standard verification symbol is as follows:
[0043] S41. Convert the denoised ion mobility spectrometry graph I(t) into a binary string data'.
[0044] S42. Calculate the first-order review parameter AP1 = H(data');
[0045] where H() represents a pre-compiled hash function;
[0046] S43. Calculate the second-order review parameter AP2:
[0047]
[0048] where n is a pre-compiled encoding factor, n = p × q; p and q are predefined large prime numbers, p ≡ 3 mod 8, p ≡ 7 mod 8;
[0049] S44. If AP1 = AP2, the denoised ion mobility spectrometry graph I(t) meets the standard; otherwise, give an alarm immediately.
[0050] Preferably, the screening process for key features is as follows:
[0051] S51. Adopt the Gaussian Fitting algorithm to fit each peak and extract the characteristic parameters of the peak:
[0052]
[0053] In the formula, I(t) represents the intensity of the peak, that is, the intensity in the ion mobility spectrometry; A represents the height of the peak; μ represents the center position of the peak, that is, the migration time; σ represents the width of the peak;
[0054] S52. From the extracted peak values, use mutual information analysis to evaluate the contribution of each feature, and remove redundant or irrelevant features through feature selection:
[0055]
[0056] In the formula, I(X,Y) represents the mutual information; P(x,y) represents the joint probability distribution of feature x and target variable y; P(x) and P(y) represent the marginal probabilities of feature x and target variable y respectively;
[0057] S53. Sort the calculated mutual information values of each feature from high to low, set a threshold θ t , and select the features with mutual information greater than the threshold θ t .
[0058] Preferably, the structure of the multi-scale convolutional neural network is:
[0059] Input layer: The selected features are used as the input;
[0060] Convolution Layer 1: A 3×3 convolutional kernel is used to extract local features through a local perception area;
[0061] Convolution Layer 2: A 5×5 convolutional kernel is used to extract global features;
[0062] Adaptive Attention Mechanism Layer AM: Automatically focuses on the feature regions related to diseases in the ion mobility spectrometry map;
[0063] Adaptive Attention Mechanism Formula:
[0064]
[0065] In the formula, A(x, y) represents the attention weight; W a represents the learned attention matrix; b a represents the bias term; f(x, y) represents the output feature map of the convolutional layer; exp() represents the exponential function;
[0066] Fully Connected Layer FC: Flattens the features processed by the convolutional layer and the attention mechanism and passes them to the fully connected layer for classification:
[0067] y = F(W fc ·flatten(f) + b fc );
[0068] In the formula, y represents the output of the fully connected layer; W fc represents the weight matrix of the fully connected layer; flatten(f) represents flattening the output f of the convolutional layer into a one-dimensional vector; b fc represents the bias term of the fully connected layer; F() represents the activation function;
[0069] Output Layer: Uses the softmax activation function to output the probability classification result of root rot:
[0070]
[0071] In the formula, P(y = z k |X) represents the predicted probability of class c k ; z k represents the logits value output by the network, that is, the original score of class c k ; X represents the class, that is, "diseased" and "disease-free"; C represents the number of classes.
[0072] Preferably, the volatile chemical substance sample is collected based on a miniaturized non-invasive sampling device, and the invasive sampling device locates the gas around the tomato roots through an intelligent probe to capture the volatile chemical substance sample dissolved in the air in the soil.
[0073] Preferably, the generation process of the ion mobility spectrometry is as follows: The volatile chemical substance sample is separated by a chromatographic column to separate various volatiles. The separated gases are ionized by an ion mobility spectrometer to form an ion mobility spectrometry.
[0074] The technical solution of the present invention: A detection system for volatile biomarkers of tomato root rot pathogens based on GC-IMS technology, which is used to execute the above-mentioned detection method for volatile biomarkers of tomato root rot pathogens based on GC-IMS technology, includes:
[0075] A sample collection module for collecting volatile organic compound samples;
[0076] A volatile organic compound detection module for detecting volatile organic compounds (VOCs) in the sample by using GC-IMS technology, vaporizing and separating the sample, and migrating and separating various gas molecules to finally obtain an ion mobility spectrometry;
[0077] A data processing module for preprocessing the ion mobility spectrometry;
[0078] A data analysis module for automatically outputting the detection result according to the ion mobility spectrometry;
[0079] An intelligent prevention and control recommendation module for providing prevention and control plan recommendations based on historical data and real-time detection results;
[0080] A disease prevention and control visualization module for visually displaying the ion mobility spectrometry, real-time detection results, and the intelligent recommended prevention and control plans.
[0081] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:
[0082] The present invention designs a detection system and method for volatile biomarkers of tomato root rot pathogens based on GC-IMS technology, which not only improves the early diagnosis efficiency of tomato root rot, but also constructs a full-process closed-loop system from detection to prevention and control, with outstanding practical value and industry promotion potential:
[0083] (1) High sensitivity and early disease identification ability: By using GC-IMS technology to detect volatile organic compounds (VOCs) in tomato roots, combined with multi-modal denoising methods (wavelet packet transform, adaptive filter, and frequency domain enhancement), the signal-to-noise ratio of the ion mobility spectrometry is significantly improved, and high-sensitivity capture of pathogen biomarkers is achieved;
[0084] (2) Non-invasive sampling and non-destructive detection: Use a miniaturized intelligent probe to accurately locate the gas around the root system, and collect samples of VOCs dissolved in the soil in a non-invasive manner to avoid mechanical damage or chemical contamination to the plant root system;
[0085] (3) Intelligent data processing and precise diagnosis: Based on a multi-scale convolutional neural network and an adaptive attention mechanism, extract local and global features from complex ion mobility spectrometry maps, dynamically focus on the signal regions related to pathogenic bacteria, and optimize feature selection by combining the mutual information and Gaussian fitting algorithms;
[0086] (4) Data standardization and integrity guarantee: Introduce a standard verification symbol generation mechanism, and perform multiple verifications on the denoised spectral data through modular arithmetic, hash functions, and quadratic residue theory to ensure the standardization and anti-tampering of the denoised spectral data, effectively resist external interference or malicious tampering, and improve the robustness of the system in a complex agricultural environment;
[0087] (5) Closed-loop intelligent prevention and control and decision support: Integrate an intelligent prevention and control recommendation module, and automatically generate multi-dimensional prevention and control plans (such as soil improvement, chemical spraying, intelligent irrigation) based on real-time detection results and historical databases to achieve closed-loop management from detection to prevention and control;
[0088] (6) High-efficiency visualization and cloud collaboration: Convert ion mobility spectrometry maps, diagnostic results, and prevention and control suggestions into intuitive charts through a disease prevention and control visualization module to lower the user's interpretation threshold and assist in quick decision-making. At the same time, all data is uploaded to the cloud in real time to support multi-terminal access and historical data backtracking, providing data support for regional disease trend analysis and prevention and control strategy optimization. Description of the Drawings
[0089] Figure 1 It is a system architecture diagram of a detection system for volatile biomarkers of pathogenic bacteria causing tomato root rot based on GC-IMS technology proposed by the present invention;
[0090] Figure 2 It is a method flow diagram of a method for detecting volatile biomarkers of pathogenic bacteria causing tomato root rot based on GC-IMS technology proposed by the present invention. Detailed Embodiments
[0091] Example 1, as Figure 1 shown, a detection system for volatile biomarkers of pathogenic bacteria causing tomato root rot based on GC-IMS technology proposed by the present invention includes: a sample collection module, a volatile organic compound detection module, a data processing module, a data analysis module, an intelligent prevention and control recommendation module, and a disease prevention and control visualization module.
[0092] The sample collection module collects samples of volatile organic compounds;
[0093] The volatile organic compound detection module uses GC-IMS technology to detect volatile organic compounds (VOCs) in samples. The samples are vaporized and separated, and various gas molecules are migrated and separated, and finally an ion mobility spectrum is obtained;
[0094] The data processing module preprocesses the ion mobility spectrum;
[0095] The data analysis module introduces an AI-based pattern recognition algorithm and a deep learning model to construct a disease diagnosis model. According to the characteristic spectra of volatile markers, it automatically identifies the signature volatiles of different pathogenic bacteria (including but not limited to root rot), and compares them with the database to output the detection results;
[0096] The intelligent prevention and control recommendation module provides multi-dimensional prevention and control plan recommendations based on historical data and real-time detection results;
[0097] The disease prevention and control visualization module visually displays the ion mobility spectrum, real-time detection results, and the intelligent recommended prevention and control plans.
[0098] Example 2, as Figure 2 shown, a method for detecting volatile biomarkers of pathogenic bacteria of tomato root rot based on GC-IMS technology proposed by the present invention is applied to a detection system for volatile biomarkers of pathogenic bacteria of tomato root rot based on GC-IMS technology proposed in Example 1. The specific implementation steps are as follows:
[0099] S1. The sample collection module uses a miniaturized non-invasive sampling device to accurately locate the gas around the tomato root system through an intelligent probe, avoiding any damage to the root system, capturing the sample of volatile chemical substances dissolved in the air in the soil, and then transmitting the collected soil volatile organic compounds to the volatile organic compound detection module.
[0100] S2. The volatile organic compound detection module inputs the received soil volatile organic compounds into the GC-IMC analysis system:
[0101] In the gas chromatography (GC) system, the sample is first separated by the chromatographic column to separate various volatiles, and the separated various gases are analyzed by the ion mobility spectrometry (IMS) module;
[0102] The separated volatile organic compounds are ionized in the IMS module to form an ion mobility spectrum. The ion migration time of each volatile compound is related to its molecular structure, forming a unique spectrum, which is transmitted to the data analysis module in real time.
[0103] S3. The data processing module designs a multi-modal denoising method by combining wavelet packet transform (WPT), adaptive filter, and frequency domain enhancement method. It is carefully optimized in both the frequency domain and the time domain, which can effectively remove different types of noise while retaining important signal features, ensuring the accuracy of subsequent analysis and feature recognition. Specifically:
[0104] S31. Use wavelet packet transform (WPT) to decompose the ion mobility spectrometry signal. It is a tool for simultaneous time-frequency analysis at multiple levels, which can decompose each frequency band of the signal more finely, making it easier to distinguish and remove noise in some frequency bands. Then, through layer-by-layer decomposition, noise components of different frequencies can be captured, and denoising is achieved by filtering low-frequency noise signals. Specifically:
[0105]
[0106] In the formula, f(t) represents the ion mobility spectrometry signal; a n represents the coefficient after wavelet packet decomposition; ψ n (t) represents the corresponding wavelet function;
[0107] Accordingly: After the wavelet packet transform decomposes the signal into different frequency bands, the a n coefficient is obtained, which represents the signal components contained in each frequency band. In the subsequent steps, noise will be identified and processed based on these frequency bands;
[0108] S32. To further reduce low-frequency noise and restore the effective part of the signal, an adaptive filter is designed, which automatically adjusts its filter coefficients based on the statistical characteristics of the noise, thereby achieving dynamic denoising. That is, in the frequency domain, the noise component is estimated by minimizing the mean square error (MSE), and the noise component is separated from the signal component. Specifically:
[0109]
[0110]
[0111] w k (t + 1) = w k (t) + μ · e(t) · x(t - k);
[0112] In the formula, e(t) represents the error signal, that is, the difference between the filter output and the target signal; d(t) represents the target signal (i.e., the signal after wavelet packet transform denoising); represents the filter output, the estimated signal; w k (t) represents the filter coefficient; μ represents the step size, which controls the update speed of the filter coefficient; x(t - k) represents the input signal, that is, the delayed signal sample, the input data of the filter, and t - k is the historical data of the signal;
[0113] Accordingly, the weight coefficient w k (t) of the filter is adjusted through adaptive learning to achieve noise suppression, and the output is the denoised signal, which will be used as the input for the next frequency-domain enhancement;
[0114] S33. After the adaptive filter denoises the ion mobility spectrogram, there may still be some unobvious residual noise or data flaws. Combining with frequency-domain enhancement technology, the residual noise is further removed by further enhancing the frequency-domain signal, that is, by weighting the important frequency components in the signal, so that the noise components are more effectively suppressed. Specifically:
[0115] S enhanced (f) = S(f)·w(f);
[0116] In the formula, S(f) represents the signal of the ion mobility spectrogram in the frequency domain, that is, the denoised signal w(f) represents the weighting function; S enhanced (f) represents the enhanced frequency-domain signal;
[0117] Accordingly, the input signal S(f) is the signal obtained through the adaptive filter is enhanced in the frequency domain, and the weighting coefficient w(f) adjusts the signal to suppress the noise components in the frequency domain. The output S enhanced (f) will be used for time-domain reconstruction;
[0118] S34. The signal after frequency-domain enhancement will be reconstructed in the time domain to restore the time-domain representation of the original ion mobility spectrogram. The enhanced frequency-domain signal is re-converted back to the time domain using the inverse wavelet packet transform (IWPT). Specifically:
[0119]
[0120] In the formula, I(t) represents the denoised time-domain signal, which is output after being restored through the inverse transform of the wavelet packet transform; a n represents the enhanced coefficient, representing the signal components of each frequency band; ψ n (t) represents the wavelet packet basis function, that is, the signal components in the time domain;
[0121] Accordingly, the signal is re-converted back to the time domain through the inverse wavelet packet transform, and the wavelet packet coefficient a n is obtained from the signal processed by the adaptive filter and the signal S enhanced (f) after frequency-domain enhancement. The finally output I(t) is the denoised signal, ready for subsequent feature extraction and analysis;
[0122] S35. Generate a standard verification symbol for the denoised ion mobility spectrometry graph I(t), and the generation process is as follows:
[0123] S3501. Calculate the intermediate parameter s = α × L(H(data)) mod n;
[0124] Where α is a precompiled generation factor, α = [lcm(p - 1, q - 1)] - 1 mod n; p and q are predefined large prime numbers, p ≡ 3 mod 8, p ≡ 7 mod 8; n is a precompiled encoding factor, n = p × q; data is the binary form of the denoised ion mobility spectrometry graph; H() represents a precompiled hash function; L() is a precompiled generation function, L(x) = (x - 1) / n; and x ≡ 1 mod n; lcm() represents the least common multiple function;
[0125] S3502. Calculate the first-order verification symbol S1:
[0126] If [s / n] = 1, then S1 = 0;
[0127] If [s / n] = -1, then S1 = 1;
[0128] S3503. Calculate the second-order verification symbol S2:
[0129]
[0130] Where Q n represents the set of quadratic residues modulo n; is the set of non - quadratic residues;
[0131] S3504. Calculate the auxiliary generation factor d = (n - p - q + 5) / 8, and calculate the third - order verification symbol S3 based on this;
[0132]
[0133] S3505. Calculate the fourth - order verification symbol S4:
[0134]
[0135] S3506. Generate the standard verification symbol Sv = {S1, S2, S3, S4};
[0136] S36. Transmit {the standard verification symbol Sv = {S1, S2, S3, S4}, the denoised ion mobility spectrometry graph I(t)} to the data analysis module.
[0137] S4. The data analysis module constructs a multi-scale convolutional neural network and combines an attention mechanism to achieve accurate diagnosis of tomato root rot. It extracts disease-related features from complex ion mobility spectra and classifies diseases based on these features, providing accurate and real-time disease detection results. The specific implementation process is as follows:
[0138] S41. Receive {standard verification symbol Sv = {S1, S2, S3, S4}, denoised ion mobility spectrum I(t)}, extract the standard verification symbol Sv = {S1, S2, S3, S4} and the denoised ion mobility spectrum I(t), and verify whether the received denoised ion mobility spectrum I(t) meets the standard according to this. The verification process is as follows:
[0139] S4101. Convert the received denoised ion mobility spectrum I(t) into a binary string data'.
[0140] S4102. Calculate the first-order verification parameter AP1 = H(data').
[0141] Where H() represents a precompiled hash function.
[0142] S4103. Calculate the second-order verification parameter AP2:
[0143]
[0144] Where n is a precompiled coding factor, n = p × q; p and q are predefined large prime numbers, p ≡ 3 mod 8, p ≡ 7 mod 8.
[0145] S4104. If AP1 = AP2, the received denoised ion mobility spectrum I(t) meets the standard; otherwise, an alarm is given immediately.
[0146] S42. There are multiple peaks in the ion mobility spectrum, and each peak corresponds to a different chemical substance. The Gaussian Fitting algorithm is used to fit each peak and extract the characteristic parameters of the peak (including but not limited to peak position, peak height, full width at half maximum), specifically:
[0147]
[0148] In the formula, I(t) represents the intensity of the peak, that is, the intensity in the ion mobility spectrum; A represents the height of the peak; μ represents the center position of the peak (i.e., the migration time); σ represents the width of the peak, reflecting the speed of ion migration.
[0149] S43. From the extracted peaks, select the most representative features related to the characteristics of tomato root rot. Use mutual information analysis to evaluate the contribution of each feature. Through feature selection, redundant or irrelevant features can be removed, thereby improving the efficiency and accuracy of the subsequent model. Specifically:
[0150]
[0151] In the formula, I(X,Y) represents mutual information; P(x,y) represents the joint probability distribution of feature x and target variable y; P(x) and P(y) represent the marginal probabilities of feature x and target variable y respectively;
[0152] Accordingly: Sort the calculated mutual information values of each feature in descending order. The larger the mutual information of a feature, the stronger the dependence relationship between the feature and the target variable, and the more valuable information it can provide. That is, the sorted features will be arranged in descending order of their importance. Set a threshold θ t , and select the features with mutual information greater than the threshold θ t ;
[0153] S44. Build a multi-scale convolutional neural network (Multi-Scale CNN). By using different convolutional kernel sizes, capture local and global features in the ion mobility spectrogram, and add an adaptive attention mechanism (Adaptive Attention Mechanism) to dynamically focus on the key regions in the spectrogram, thereby enhancing the sensitivity to disease characteristics and transmitting the diagnosis results to the intelligent prevention and control recommendation module. The model architecture is as follows:
[0154] (1) Input layer: The selected features are used as the input;
[0155] (2) Convolutional layer 1 (C1): Use multiple convolutional kernels of size 3×3 to extract basic features through the local perception area;
[0156] (3) Convolutional layer 2 (C2): Use multiple convolutional kernels of size 5×5 to further extract higher-level features and increase the receptive field;
[0157] The convolution operation formula is:
[0158]
[0159] In the formula, f(x,y) represents the convolution output; W i,j represents the weight of the convolutional kernel; I(x+I,y+j) represents the pixel value of the input image; b represents the bias term; K represents the size of the convolutional kernel;
[0160] (4) Adaptive Attention Mechanism Layer (AM): Automatically focuses on the feature regions related to diseases in the ion mobility spectrometry graph;
[0161] Adaptive attention mechanism formula:
[0162]
[0163] In the formula, A(x, y) represents the attention weight; W a represents the learned attention matrix; b a represents the bias term; f(x, y) represents the output feature map of the convolutional layer; exp() represents the exponential function, which is used to calculate the attention value at each position. The exponential operation makes the features with larger values have higher weights for attention allocation, further focusing on the key regions;
[0164] (5) Fully Connected Layer (FC): Flattens the features processed by the convolutional layer and the attention mechanism and passes them to the fully connected layer for classification:
[0165] y = F(W fc ·flatten(f) + b fc );
[0166] In the formula, y represents the output of the fully connected layer; W fc represents the weight matrix of the fully connected layer, which is responsible for weighted summation of the input features and determines the importance of the features; flatten(f) represents flattening the output f of the convolutional layer into a one-dimensional vector; b fc represents the bias term of the fully connected layer; F() represents the activation function;
[0167] (6) Output Layer: Uses the softmax activation function to output the probability classification result of root rot:
[0168]
[0169] In the formula, P(y = z k |X) represents the predicted probability of class c k ; z k represents the logits value output by the network, that is, the original score of class c k ; C represents the number of classes, indicating how many classes there are in total in the classification task. In this implementation, C = 2; X represents the classes, that is, "diseased" and "disease-free".
[0170] S5. Based on the above analysis results, the intelligent prevention and control recommendation module automatically generates a diagnostic report, and based on historical data, provides corresponding prevention and control suggestions according to the results of disease identification, including but not limited to soil improvement and recommendation of disease control agents;
[0171] For example, if characteristic volatiles of high-concentration tomato root rot are detected, local soil treatment, spraying of root control agents, or an intelligent irrigation plan is recommended.
[0172] S6. The disease control visualization module visually displays the diagnosis results and control suggestions, and all diagnostic data, disease characteristic information, and treatment suggestions will be uploaded to the cloud for storage and update.
[0173] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.
Claims
1. A method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology, characterized in that: The specific implementation steps include the following: S1. Collect samples of volatile chemicals dissolved in the air from the soil; S2, generating an ion mobility spectrum based on gas chromatography-ion mobility spectrometry GC-IMS technology; S3, denoising based on multimodal denoising method, multi-level time-frequency analysis of ion mobility spectrum signal by wavelet packet transform, decomposition and noise removal, and dynamic adjustment of coefficients by adaptive filter, suppressing noise and restoring effective signal, weighting important frequency components, suppressing residual noise, restoring time domain signal by inverse wavelet packet transform, and generating standard audit symbol by hash function and quadratic residue operation; S4. Conduct standard audits through standard audit symbols, extract peak feature parameters using Gaussian fitting algorithm, select key features through mutual information analysis, build multi-scale convolutional neural networks and adaptive attention mechanisms to capture local and global features, and finally classify and output disease diagnosis results through the fully connected layer; S5. Based on the diagnosis results, a diagnosis report is automatically generated, and corresponding prevention and control suggestions are provided according to the results of disease identification; S6. The diagnosis results and prevention and control recommendations will be visualized, and all diagnostic data, disease characteristic information and treatment recommendations will be uploaded to the cloud for storage and updating.
2. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 1, characterized in that: The denoising process based on the multimodal denoising method is as follows: S21. Decompose the ion mobility spectrum signal using wavelet packet transform: Where f(t) represents the ion mobility spectrometer signal; a n represents the coefficient after wavelet packet decomposition; ψ n (t) represents the corresponding wavelet function; S22, through adaptive filter, automatically adjust the filter coefficient based on the statistical characteristics of noise to perform dynamic denoising: w k (t+1)=w k (t)+μ·e(t)·x(t-k); Where, e(t) represents the error signal, that is, the difference between the filter output and the target signal; d(t) represents the target signal, that is, the signal after denoising by wavelet packet transform; represents the filter output, the estimated signal; w k (t) represents the filter coefficient; μ represents the step size, which controls the speed of updating the filter coefficient; x(tk) represents the input signal, that is, the input data of the filter, and tk is the historical data of the signal; S23, by removing the residual noise from the frequency domain signal, that is, by weighting the important frequency components in the signal, the noise components are suppressed more effectively: S enhanced (f)=S(f)·w(f); In the formula, S(f) represents the signal of the ion mobility spectrum in the frequency domain, that is, the signal after denoising w(f) represents the weighting function; S enhanced (f) represents the enhanced frequency domain signal; S24. The signal after frequency domain enhancement is reconstructed in the time domain to restore the original time domain representation of the ion mobility spectrum. The enhanced frequency domain signal is converted back to the time domain using the inverse wavelet packet transform (IWPT): Where I(t) represents the denoised time domain signal, which is restored and output after inverse transformation of wavelet packet transform; a n It represents the enhanced coefficient, which represents the signal component of each frequency band.
3. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 1, characterized in that: The process of generating a standard audit symbol is as follows: S31, calculating the intermediate parameter s=α×L(H(data))mod n; Wherein, α is the precompiled generation factor, α = [lcm(p-1,q-1)]-1mod n; p and q are predefined large prime numbers, p≡3mod 8, p≡7mod 8; n is the precompiled coding factor, n = p×q; data is the binary form of the denoised ion migration spectrum; H() represents the precompiled hash function; L() is the precompiled generation function, L(x) = (x-1) / n; And x≡1mod n; lcm() represents the least common multiple function; S32. Calculate the first-order audit symbol S1: If [s / n] = 1, then S1 = 0; If [s / n] = -1, then S1 = 1; S33. Calculate the second-order audit symbol S2: Among them, Q n represents the set of quadratic residues modulo n; is a non-quadratic residue set; S34, calculate the auxiliary generation factor d = (np-q+5) / 8, and calculate the third-order audit symbol S3 based on it: S35. Calculate the fourth-order audit symbol S4: S36. Generate standard audit symbol Sv={S1, S2, S3, S4}.
4. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 3, characterized in that: The audit process for standard auditing through standard audit symbols is as follows: S41, converting the denoised ion migration spectrum I(t) into a binary string data'; S42, calculate the first-order audit parameter AP1 = H (data'); Where H() represents the precompiled hash function; S43. Calculate the second-order audit parameter AP2: Wherein, n is a pre-compiled coding factor, n=p×q; p and q are pre-defined large prime numbers, p≡3mod8, p≡7mod8; S44. If AP1=AP2, the ion migration spectrum I(t) after denoising meets the standard; otherwise, an alarm is immediately issued.
5. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 1, characterized in that: The screening process for key features is as follows: S51. Use Gaussian Fitting algorithm to fit each peak and extract the characteristic parameters of the peak: In the formula, I(t) represents the intensity of the peak, that is, the intensity in the ion mobility spectrum; A represents the height of the peak; μ represents the center position of the peak, that is, the migration time; σ represents the width of the peak; S52. From the extracted peaks, mutual information analysis is used to evaluate the contribution of each feature, and redundant or irrelevant features are removed through feature selection: Where I(X,Y) represents the mutual information; P(x,y) represents the joint probability distribution of feature x and target variable y; P(x) and P(y) represent the marginal probabilities of feature x and target variable y, respectively; S53, sort the calculated mutual information value of each feature in descending order, and set the threshold θ t , select the mutual information greater than the threshold θ t characteristics.
6. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 5, characterized in that: The structure of the multi-scale convolutional neural network is: Input layer: The selected features are used as input; Convolutional layer 1: Use a convolution kernel of size 3×3 to extract local features through the local perception area; Convolutional layer 2: Use a convolution kernel of size 5×5 to extract global features; Adaptive attention mechanism layer AM: automatically focuses on the feature areas related to the disease; Adaptive attention mechanism formula: Where A(x,y) represents the attention weight; W a represents the learned attention matrix; b a represents the bias term; f(x,y) represents the output feature map of the convolutional layer; exp() represents the exponential function; Fully connected layer FC: Flattens the features processed by the convolutional layer and the attention mechanism and passes them to the fully connected layer for classification: y=F(W fc ·flatten(f)+b fc ); Where y represents the output of the fully connected layer; W fc represents the weight matrix of the fully connected layer; flatten(f) means flattening the output of the convolutional layer into a one-dimensional vector; b fc represents the bias term of the fully connected layer; F() represents the activation function; Output layer: Use the softmax activation function to output the probability classification results of root rot: In the formula, P(y=z k |X) indicates category c k The predicted probability of k Represents the logits value of the network output, that is, category c k The original score of; X represents the category, that is, "disease" and "no disease"; C represents the number of categories.
7. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 1, characterized in that: Volatile chemical samples were collected based on a miniaturized non-invasive sampling device. The invasive sampling device used a smart probe to locate the gas around the tomato roots to capture samples of volatile chemicals dissolved in the air in the soil.
8. The method for detecting volatile biomarkers of tomato root rot pathogens based on GC-IMS technology according to claim 1, characterized in that: The process of generating an ion mobility spectrum is as follows: a volatile chemical sample is separated by a chromatographic column to separate various volatile substances, and the separated gases are ionized by an ion mobility spectrum to form an ion mobility spectrum.
9. A volatile biomarker detection system for tomato root rot pathogens based on GC-IMS technology, which is used to perform a volatile biomarker detection method for tomato root rot pathogens based on GC-IMS technology according to any one of claims 1 to 8, characterized in that: include: A sample collection module, used for collecting volatile organic compound samples; The volatile organic compound detection module is used to detect volatile organic compounds (VOCs) in samples using GC-IMS technology, gasify and separate the samples, and migrate and separate various gas molecules to finally obtain an ion migration spectrum; A data processing module, used for preprocessing the ion mobility spectrum; Data analysis module, used to automatically output the test results based on the ion migration spectrum; Intelligent prevention and control recommendation module, used to provide prevention and control plan recommendations based on historical data and real-time detection results; The disease control visualization module is used to visualize ion migration spectra, real-time detection results, and intelligently recommended control plans.
Citation Information
Patent Citations
Method for screening grape mildew characteristic VOCs based on GC-IMS technology in combination with chemometrics
CN116413357A