Fruit Mycotoxin Contamination Early Warning System Based on High-Throughput Sequencing and Machine Learning

By adopting a fruit mycotoxin pollution warning system based on high-throughput sequencing and machine learning in fruit mycotoxin detection, the problem of insufficient efficiency and accuracy of traditional detection methods is solved, and timely early warning and effective control of mycotoxin pollution is achieved.

CN118711663BActive Publication Date: 2025-06-10FRUIT TREE INST OF CHINESE ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410653674.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-06-10
Estimated Expiration
2044-05-24

AI Technical Summary

Technical Problem

Traditional fruit mycotoxin detection methods have problems such as cumbersome pre-processing, low detection throughput, insufficient detection limit and high false positive rate, which is difficult to meet the needs of food safety supervision. At the same time, in the analysis of high-throughput sequencing data, low calculation efficiency, difficult parameter tuning, insufficient comparison accuracy and lack of noise processing mechanisms, it is difficult to achieve real-time analysis and accurate identification of trace and novel toxin genes.

Method used

A fruit mycotoxin pollution warning system based on high-throughput sequencing and machine learning is adopted, including sample preprocessing module, sequencing data analysis module, intelligent risk assessment module, early warning information release module and system self-learning optimization module. The system realizes accurate detection and risk assessment of mycotoxin genes through literature coding comparison algorithms, semantic feature extraction algorithms and adaptive clustering algorithms, and constructs a pollution risk grading warning model through a random forest algorithm.

Benefits of technology

It significantly improves detection efficiency and accuracy, achieves timely warning and effective control of mycotoxin pollution, overcomes the limitations of traditional methods, and improves detection sensitivity and warning timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118711663B_ABST
    Figure CN118711663B_ABST
Patent Text Reader

Abstract

A fruit mycotoxin pollution early warning system based on high-throughput sequencing and machine learning, including: a sample pretreatment module, a sequencing data analysis module, an intelligent risk assessment module, an early warning information release module, and a system self-learning optimization module. The sequencing data analysis module is used to perform high-throughput sequencing on the library to obtain sequencing reads data, and uses a literature coding alignment algorithm and a semantic feature extraction algorithm to accurately detect mycotoxin-related toxin-producing genes; the intelligent risk assessment module is used to determine the risk level of sample toxin pollution according to the abundance and harmfulness assessment of the toxin-producing genes detected by the sequencing data analysis module; the early warning information release module is used to automatically generate an early warning report and publish and share it through the Web end; the system self-learning optimization module is used to perform incremental learning and iterative optimization on the coding alignment algorithm and the semantic feature extraction algorithm in the data analysis module according to the new data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fruit mycotoxin pollution warning systems, and more specifically, to a fruit mycotoxin pollution warning system based on high-throughput sequencing and machine learning. Background Art

[0002] Traditional methods for detecting fruit mycotoxins mainly include thin-layer chromatography (TLC), high-performance liquid chromatography (HPLC), gas chromatography-mass spectrometry (GC-MS), enzyme-linked immunosorbent assay (ELISA), etc. These methods are difficult to meet the increasingly strict food safety supervision requirements in terms of specificity, sensitivity, etc. Taking HPLC as an example, it mainly has the following limitations:

[0003] (1) The sample pretreatment is cumbersome. The sample needs to go through complex steps such as extraction, purification, and concentration before injection analysis, which is time-consuming and laborious. Moreover, different matrices and toxin types require different pretreatment schemes, with poor generality.

[0004] (2) The detection throughput is low. HPLC belongs to single-channel detection, and only one sample can be analyzed each time, making it difficult to achieve large-scale screening and unable to meet the rapid response requirements for mycotoxin outbreaks.

[0005] (3) The detection limit is insufficient. The detection sensitivity of HPLC is limited by factors such as injection volume and matrix effect, and its ability to detect trace toxins is limited, with a risk of missed detection.

[0006] (4) The false positive rate is high. Interfering components in the matrix may produce chromatographic peaks similar to the target toxin, resulting in false positive results and affecting the detection accuracy.

[0007] Entering the 21st century, a disruptive technology in the biological field - high-throughput sequencing - is being introduced into food safety detection. However, how to process massive sequencing data and accurately identify trace and novel toxin genes still faces huge challenges. Some scholars have tried to directly apply traditional short sequence alignment algorithms such as BLAST to the analysis of toxin gene sequencing data, but there are the following deficiencies:

[0008] (1) The computational efficiency is low. BLAST adopts an exhaustive local alignment strategy, and the time complexity is as high as O(mn) (where m is the reads length and n is the reference sequence length), making it difficult to meet the real-time analysis requirements under the scale of big data.

[0009] (2) Parameter tuning is difficult. The selection of threshold parameters in BLAST has a great impact on the results, but there is no adaptive adjustment mechanism for different data characteristics, and the subjectivity is large.

[0010] (3)Insufficient alignment accuracy. BLAST only considers the local similarity of sequences and ignores the global semantic information, making it difficult to accurately distinguish highly similar toxin gene sequences.

[0011] (4)Lack of noise processing mechanism. Real sequencing data often contains a large number of low-quality and irrelevant reads. Directly using them for alignment will result in redundant calculations and false positives, reducing the analysis efficiency and accuracy.

[0012] It can be seen that traditional analysis methods and algorithms can no longer meet the needs of high-throughput sequencing data analysis, and innovative technological breakthroughs are urgently needed. In view of the above technical problems, the present invention proposes an integrated and intelligent sequencing data analysis solution, and constructs a complete early warning system for mycotoxin contamination, aiming to contain toxin contamination from the source. Summary of the Invention

[0013] The purpose of the present invention is to provide an early warning system for mycotoxin contamination of fruits and vegetables that integrates sequencing analysis, machine learning, and self-learning optimization in view of the deficiencies of the prior art, so as to significantly improve the detection efficiency and accuracy, and achieve timely early warning and effective control of mycotoxin contamination.

[0014] The present invention provides an early warning system for mycotoxin contamination of fruits and vegetables based on high-throughput sequencing and machine learning. 1. An early warning system for mycotoxin contamination of fruits and vegetables based on high-throughput sequencing and machine learning, characterized by comprising: a sample pretreatment module, a sequencing data analysis module, an intelligent risk assessment module, a warning information release module, and a system self-learning optimization module;

[0015] The sample pretreatment module is used for crushing, homogenizing, DNA extraction, and library construction of fruit and vegetable samples;

[0016] The sequencing data analysis module is used for performing high-throughput sequencing on the library to obtain sequencing reads data, and realizing accurate detection of mycotoxin genes by using a literature coding alignment algorithm and a semantic feature extraction algorithm;

[0017] The intelligent risk assessment module is used for evaluating the sample contamination risk level according to the mycotoxin gene abundance and harmfulness detected by the sequencing data analysis module;

[0018] The warning information release module is used for automatically generating a warning report and publishing and sharing it through the Web end;

[0019] The system self-learning optimization module is used for performing incremental learning and iterative optimization on the literature coding alignment algorithm and the semantic feature extraction algorithm in the sequencing data analysis module according to new data.

[0020] Specifically, the literature coding alignment algorithm includes the following steps:

[0021] First, construct a database encoding toxin genes, and use the SimHash algorithm to convert the toxin gene sequences into 64-bit binary encoding vectors;

[0022] Then, calculate the Hamming distance H between the sequencing reads and the encoding vectors in the database d :

[0023]

[0024] where read i and ref i represent the i-th bit of the sequencing reads and the reference encoding vector respectively, is the exclusive OR operator,

[0025] Next, set an adaptive threshold

[0026]

[0027] where α and β are balance factors, n is the number of reads in the current batch, is the mean of H d value,

[0028] Finally, screen the suspected toxin gene reads: Determine the reads with H d less than as the suspected toxin gene reads.

[0029] Specifically, the semantic feature extraction algorithm includes the following steps:

[0030] First, convert the suspected toxin gene reads into a one-hot encoding matrix R ∈ R l×4 , and reduce the dimension through the embedding matrix W e ∈ R 4×d to obtain a dense vector X ∈ R l×d : X = R · W e

[0031] where l is the read length and d is the embedding vector dimension;

[0032] Then, use a multi-scale convolutional kernel W c ∈ R k×d to extract the local features of the reads and perform a non-linear transformation through the ReLU function σ:

[0033]

[0034] where is the convolution operation, b c is the bias term

[0035] Then, the multi-scale convolution features are adaptively integrated through the attention mechanism to obtain the semantic representation vector of reads.

[0036]

[0037] Among them, α k is the attention weight of the k-th convolution feature:

[0038]

[0039] is a learnable query vector;

[0040] Design a loss function that fuses the similarity and positional relationship of reads to optimize the model parameters:

[0041]

[0042] Among them, P and N are the sets of positive and negative samples respectively, d ij is the distance between readsi and j on the genome, and λ is the balance coefficient;

[0043] Finally, perform adaptive DBSCAN clustering on the reads semantic representation vector S to obtain a high-confidence toxin gene reads cluster.

[0044] Specifically, the process of the adaptive DBSCAN clustering includes:

[0045] First, calculate the Euclidean distance d i between any two semantic representation vectors S j and S ij = ||S i - S j || 2; Then, construct a similarity graph between reads based on the shared nearest neighbor similarity. The similarity sim ij is calculated as follows:

[0046]

[0047] Among them, Γ(S i ) represents the set of k nearest neighbors of S i ;

[0048] Finally, apply the DBSCAN clustering algorithm on the similarity graph to automatically determine the number of clusters and obtain the toxin gene reads cluster.

[0049] Specifically, after obtaining the reads clusters of the toxin genes, a cluster confidence scoring function is further used to evaluate the quality of each cluster, and the high-confidence clusters are used for subsequent analysis. The cluster confidence score score i The calculation formula is:

[0050]

[0051] where L i is the minimum inter-cluster distance of cluster i, reflecting the inter-cluster separation:

[0052]

[0053] μ i and μ j are the centroid vectors of cluster i and cluster j respectively; H i is the average intra-cluster distance of cluster i, reflecting the intra-cluster compactness:

[0054] Specifically, the intelligent risk assessment module adopts the random forest algorithm, and constructs a sample mycotoxin contamination risk grading and early warning model with the types, abundances, and hazards of the toxin genes detected by the sequencing data analysis module as features.

[0055] Specifically, the early warning information release module includes: (1) a report generation unit, which is used to automatically generate a structured pollution early warning report according to the early warning results of the intelligent risk assessment module;

[0056] (2) an information release unit, which is used to publish the early warning report through a Web service to realize the real-time sharing of early warning information.

[0057] Specifically, the system self-learning and optimization module includes: (1) an incremental learning unit, which is used to regularly incorporate newly collected fruit samples into the data set and perform incremental training on the literature coding comparison algorithm and semantic feature extraction algorithm of the sequencing data analysis module; (2) a parameter tuning unit, which is used to continuously optimize the hyperparameters of the algorithm based on new data to improve the algorithm performance.

[0058] Specifically, the system further includes:

[0059] (1) a data storage module, which is used to persistently store and manage the sample sequencing data, metadata, model parameters, etc. using a distributed file system;

[0060] (2) a high-performance computing module, which is used to accelerate computationally intensive tasks such as sequencing data analysis and model training using a parallel computing framework.

[0061] The present invention has the following beneficial effects:

[0062] 1. Synergistic effect of multiple technology modules. This system adopts a modular design, decouples and encapsulates key links such as sample pretreatment, sequencing analysis, risk assessment, information release, and self-learning optimization, and realizes interconnection through standardized interfaces, enabling each module to evolve independently and operate collaboratively. At the same time, the system innovatively integrates the cutting-edge technology of high-throughput sequencing with machine learning algorithms to form a technical synergistic mechanism driven by massive data and empowered by intelligent algorithms, which not only excavates the deep value of metagenomic data but also overcomes the inherent limitations of traditional toxin detection methods, greatly improving the sensitivity of detection and the timeliness of early warning.

[0063] 2. Information superposition of multi-dimensional features. The intelligent risk assessment module ingeniously integrates multi-source heterogeneous features such as toxin gene abundance, harmfulness, and environmental meteorology, and combines these features through a random forest algorithm for high-dimensional combination to explore their internal relationships and achieve accurate hierarchical early warning of pollution risks. This multi-dimensional feature superposition mechanism overcomes the one-sidedness of single-feature early warning and provides a more comprehensive and reliable risk analysis perspective. For example, even if a high-abundance toxin gene is detected in some samples, if its toxicity is low and the environmental conditions are not conducive to toxin expression, the system can determine its risk level as medium or low based on this to avoid over-early warning.

[0064] 3. Iterative complementarity of active learning and incremental learning. The self-learning optimization module integrates two machine learning paradigms of active learning and incremental learning, which are iteratively complementary during the operation of the system to continuously optimize the model performance and improve the early warning effect. Active learning focuses on the manual annotation of a small number of valuable samples to specifically improve the model's early warning ability for complex and unknown pollution events; incremental learning continuously trains the model on general big data to make it adapt to the dynamic evolution of the mycotoxin prevalence spectrum. The complementary combination of the two learning paradigms enables the system to continuously learn and evolve while dealing with known and common pollution, and calmly respond to new and sudden pollution events.

[0065] 4. Resolution of the contradiction between distributed storage and parallel computing. Massive sequencing data poses a huge challenge to storage and computing resources. The traditional centralized storage and serial computing paradigms are difficult to meet the system requirements under the premise of controllable time and cost, and there are problems of storage bottlenecks and computing bottlenecks. The data storage module adopts the HDFS distributed file system to store TB / PB-level sequencing data distributedly in a cluster of inexpensive servers, and the theoretical storage capacity can be infinitely expanded; the high-performance computing module adopts the Spark in-memory computing framework to split and parallelize computing tasks such as machine learning and send them to cluster nodes for concurrent execution, significantly improving the computing speed. The organic combination of this distributed storage and parallel computing technology fundamentally solves the storage and computing bottleneck problems brought by big data.

[0066] 5. The sequencing data analysis module is the core of the present invention. It innovatively integrates multiple algorithms such as literature coding alignment, semantic feature extraction, and adaptive clustering to form an efficient, accurate, and robust sequencing data mining solution. First, the literature coding alignment unit uses the SimHash locality-sensitive hashing algorithm to map known toxin gene sequences into binary codes, and realizes the preliminary screening of reads by calculating the Hamming distance between the codes. This algorithm cleverly utilizes the principle of locality sensitivity, that is, similar sequences have a small Hamming distance after hash mapping. At the same time, the binary coding greatly compresses the data volume, enabling the subsequent distance calculation to be completed within a very low spatio-temporal complexity of O(n), thus achieving an efficiency improvement in high-throughput reads alignment. In addition, this unit also introduces a threshold adaptive strategy, dynamically adjusting the distance threshold according to the mean and variance of the reads Hamming distance, taking into account both sensitivity and specificity in improving the alignment, reflecting the adaptability and robustness of the algorithm. Second, the semantic feature extraction unit innovatively introduces a convolutional neural network into the feature representation learning of sequencing reads. Through successive abstractions of the embedding layer, convolutional layer, and attention layer, it automatically learns the deep semantic features of reads. Traditional sequencing data analysis methods usually rely on shallow features such as the k-mer frequency of reads, which are difficult to characterize complex and subtle sequence differences. The convolutional neural network can effectively capture the key sequence patterns of reads through the local receptive field mechanism and achieve a multi-granularity understanding of the semantics of read segments through multi-scale convolutional kernels. At the same time, the attention mechanism endows the model with the ability to automatically adjust the focus according to the context, further enhancing the discriminability of semantic features. In addition, this unit also cleverly introduces an immune negative sample similarity penalty term into the loss function to minimize the similarity of reads with inconsistent labels, further widening the inter-class differences, which conforms to the biological mechanism of sequence analysis. It can be seen that the semantic feature extraction unit synergistically enhances efficiency through multiple mechanisms, greatly improving the accuracy and generalization of feature representation. Finally, in the high-confidence toxin gene reads identification link, the adaptive clustering unit improves the traditional DBSCAN algorithm by introducing shared nearest neighbor similarity, overcoming the preference of the original algorithm for spherical clusters and enabling it to discover cluster structures of arbitrary shapes. At the same time, the construction of the SNN graph suppresses the influence of isolated noise points to a certain extent, making the clustering process more robust. In addition, this unit also designs a cluster confidence scoring mechanism to evaluate the clustering quality from two dimensions of cluster compactness and separation, making full use of the rich information of the clustering structure and further improving the reliability of toxin gene identification. These improvements are organically combined to form an effective cluster optimization scheme. In summary, the sequencing data analysis module is based on the cutting-edge technology of high-throughput sequencing, deeply integrating multiple innovative algorithms such as literature coding alignment, semantic feature extraction, and adaptive clustering to form a set of technical combination punches that complement each other and synergistically enhance efficiency.The coding comparison unit realizes the efficient and accurate preliminary screening of reads of toxin genes. The semantic feature extraction unit realizes the automatic learning of deep discriminative features. The adaptive clustering unit realizes the intelligent processing of noise data and clusters of arbitrary shapes. The three major algorithms are closely linked, and the processing details are refined. Through data mining from multiple granularities and perspectives, the information value of sequencing data is maximally amplified, laying a solid data foundation for downstream intelligent early warning. Therefore, this module is of milestone significance in improving the sensitivity, accuracy, and usability of mycotoxin detection, representing the current technological forefront and development direction in this field.

[0067] In summary, the present invention realizes the synergistic effect among multiple modules and the information superposition of multi-dimensional features in the system architecture, and realizes the iterative complementarity of active learning and incremental learning and the seamless integration of distributed storage and parallel computing in the algorithm design. Overall, it forms an intelligent solution with advanced technology, complete functions, and excellent performance. The cross-integration of these innovative technologies produces an early warning effect far beyond conventional methods, providing solid technical support for ensuring the quality and safety of fruits, safeguarding the rights and interests of consumers, opening up new paths and providing new ideas for food safety supervision and risk governance, and having broad application prospects. Brief Description of the Drawings

[0068] Figure 1 It is a schematic diagram of the framework of a fruit mycotoxin contamination early warning system based on high-throughput sequencing and machine learning.

[0069] Figure 2 It is a schematic diagram of the framework of the sequencing data analysis module.

[0070] Figure 3 It is a flowchart of the method of the present invention. Detailed Embodiments

[0071] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the drawings and preferred embodiments to detail the specific embodiments, structures, features, and effects of the fruit mycotoxin contamination early warning system based on high-throughput sequencing and machine learning proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0073] The following specifically describes the specific solution of the fruit mycotoxin contamination early warning system based on high-throughput sequencing and machine learning provided by the present invention with reference to the drawings.

[0074] Example 1

[0075] Please refer to Figure 1-2 , the present invention provides a warning system for fruit mycotoxin contamination based on high-throughput sequencing and machine learning, including: a sample pretreatment module 1, a sequencing data analysis module 2, an intelligent risk assessment module 3, a warning information release module 4, and a system self-learning optimization module 5;

[0076] The sample pretreatment module 1 is used to process fruit samples to obtain a library to be sequenced. The following process is specifically adopted: First, use a Retsch Grindomix GM 200 homogenizer to crush the fruit samples, with the parameter settings of 10,000 rpm and a crushing time of 60 s; then add PBS buffer and use an IKA Ultra-Turrax T25 homogenizer for homogenization treatment, with the parameter settings of 13,500 rpm and a homogenization time of 120 s; subsequently, use a QIAGEN DNeasy Plant Mini Kit to extract the sample DNA; finally, use an Illumina TruSeq DNA PCR-Free library construction kit to construct a sequencing library, and the fragment size is selected to be 300-500 bp. The above series of sample pretreatment operations maximally break the fruit tissue at the physical level, fully lyse the cells at the chemical level, highly pure extract DNA at the molecular level, and construct a high-quality sequencing library, laying a foundation for subsequent high-throughput sequencing experiments.

[0077] The sequencing data analysis module 2 is used to perform high-throughput sequencing on the library constructed by the sample pretreatment module to obtain raw sequencing data, and call a literature coding comparison unit 21 and a semantic feature extraction unit 22 to achieve accurate detection of mycotoxin gene reads. The literature coding comparison unit 21 adopts a literature coding comparison algorithm, and the semantic feature extraction unit 22 adopts a semantic feature extraction algorithm. The module process is as follows:

[0078] The present invention designs an innovative rapid comparison algorithm for toxin genes in the literature coding comparison unit 21: First, download known mycotoxin gene sequences from public nucleic acid databases such as NCBI and EMBL through web crawler technology, and use the SimHash locality-sensitive hashing algorithm to convert them into 64-bit binary coding vectors. The above process can be completed offline to form a toxin gene coding database, greatly improving the comparison efficiency. When processing sequencing data, this unit calculates the Hamming distance between each read and the coding vector in the database in real time, and the formula is as follows:

[0079]

[0080] where, read iand ref i respectively represent the i-th bit of the sequencing reads and the reference coding vector, where ⊕ is the exclusive OR (XOR) operator.

[0081] To adaptively determine the optimal Hamming distance threshold, this unit introduces a threshold dynamic adjustment strategy:

[0082]

[0083] where α and β are balance factors, and n is the number of reads in the current batch is the mean of H d . Experiments show that when α takes values from 0.6 to 0.8 and β takes values from 1.2 to 1.5, a robust threshold estimation effect can be obtained. The literature coding comparison unit 21 finally determines the reads with H d less than as suspected toxin gene-related reads and outputs them to the downstream analysis unit.

[0084] To further improve the accuracy of reads screening, the semantic feature extraction unit 22 innovatively introduces a semantic segmentation convolutional neural network to perform feature representation learning on suspected toxin gene reads from the sequence semantic level. The core process is as follows:

[0085] (1) Convert the suspected toxin gene reads into a sparse matrix R ∈ R l×4 , and right-multiply it by a randomly initialized dense embedding matrix W e ∈ R 4×d to achieve feature dimensionality reduction and obtain a low-dimensional dense feature matrix X ∈ R l×d . The dimension d of the embedding vector is usually set to 64 - 512 to balance the time and space complexity.

[0086] (2) Use a multi-scale convolutional kernel to perform a convolution operation on X to extract the local features of the reads, and introduce a non-linear transformation through the ReLU activation function σ. The formula is as follows:

[0087]

[0088] In the above formula, is the convolution operator, and b c is the bias term. The size k of the convolutional kernel usually takes odd numbers, such as 3, 5, 7, etc., to cover semantic segments of different sizes.

[0089] (3) Adopt an attention mechanism to adaptively fuse multi-scale convolution features, assign different importance to different semantic segments, and obtain the semantic representation vector of the reads after aggregation. The formula is as follows:

[0090]

[0091] where α k is the attention weight of the k-th convolutional feature, obtained by taking the inner product of the convolutional feature map C k and the learnable query vector and performing softmax normalization:

[0092]

[0093] (4) Introduce a new loss function that, while supervising the read segment similarity, minimizes the semantic similarity between immune negative samples and penalizes the clustering tendency of reads with inconsistent labels in the vector space. The formula is as follows:

[0094]

[0095] The first and second terms of the formula are the log-likelihoods of positive and negative samples respectively. The third term is a regularization term based on the semantic distance of immune negative samples, and the coefficient λ controls the penalty strength. d ij is the distance between readsi and j on the genome. Continuously optimize the model parameters through gradient backpropagation to finally obtain highly discriminative semantic features.

[0096] (5) Input the semantic feature vector S of reads into the clustering unit, and use the adaptive DBSCAN algorithm to divide the high-confidence toxin gene reads clusters. The process is as follows: Based on the semantic representation vector, calculate the Euclidean distance d ij = ||S i - S j || 2 ;

[0097] (2) Construct a read similarity graph. To overcome the defect that the traditional DBSCAN algorithm is highly sensitive to the distance threshold ∈, this unit improves the calculation strategy and uses the shared nearest neighbor similarity to define the edge weight sim ij , and the formula is as follows:

[0098]

[0099] where Γ(S i ) represents the set of k nearest neighbors of S i . The selection of the number of neighbors k is related to the data distribution, and the empirical value range is 5 - 20. sim ij considers both local similarity and global distribution, making the clustering process more robust.

[0100] (3) Apply the DBSCAN algorithm to the similarity graph to automatically identify the cluster structure and obtain the clusters of toxin gene reads. To further improve the recognition confidence, the present invention designs a cluster confidence scoring function to quantitatively evaluate the qualities of each cluster, and preferably selects high-confidence clusters for downstream analysis. The formula is as follows:

[0101]

[0102] where, L i is the minimum inter-cluster distance of cluster i, reflecting the inter-cluster separation degree, and the calculation formula is:

[0103] L i = min j≠i ||μ i - μ j || 2

[0104] μ i and μ j are the centroid vectors of cluster i and cluster j respectively.

[0105] H i is the average intra-cluster distance of cluster i, reflecting the intra-cluster compactness, and the calculation formula is:

[0106]

[0107] In summary, through a series of innovative algorithms such as the above-mentioned literature coding comparison, semantic feature extraction, and adaptive clustering screening, the sequencing data analysis module can efficiently and accurately detect trace and novel mycotoxin genes in the sample, providing reliable data support for downstream pollution risk assessment.

[0108] To achieve the quantitative risk warning of mycotoxin pollution, the intelligent risk assessment module 3 adopts advanced machine learning algorithms to fully mine the characteristic information of the sequencing analysis results and construct a pollution grading warning model. The process is as follows:

[0109] (1) Extract the mycotoxin gene expression abundance characteristics. Use tools such as salmon to quantify the number of toxin gene reads of each sample into relative abundance (TPM) values to characterize the sample pollution degree;

[0110] (2) Extract the mycotoxin hazard characteristics. Extract the toxicity parameters (such as LD50, TD50, etc.) of each detected toxin from the toxin database (such as EFSA, FDA, etc.) to characterize its hazard degree;

[0111] (3) Extract the environmental meteorological characteristics. Obtain data such as temperature, humidity, and rainfall at the sampling point from the meteorological department to characterize the influence of external conditions on mycotoxin pollution;

[0112] (4) The random forest classification algorithm is adopted, with the above-mentioned multi-dimensional heterogeneous features as the input, to train a warning model for the risk classification (high, medium, low) of mycotoxin contamination in the training samples. To balance the sparsity of the training samples and the problem of class imbalance, the SMOTE method is used for data augmentation and balancing.

[0113] Through the above process, the intelligent risk assessment module 3 can real-time warn of the mycotoxin contamination risk level according to the sample test results, providing a quantitative basis for the safety assessment, warning, and traceability of fruits. Based on this module, government regulatory agencies can timely take targeted control measures to effectively prevent food safety risks.

[0114] To facilitate the government and the public to timely obtain the warning information on mycotoxin contamination in fruits, the warning information release module 4 provides an automated information release function, including:

[0115] (1) Report generation unit: According to the warning level output by the intelligent risk assessment module, an early warning report is automatically generated according to a customized template, and the content includes sample information, detected mycotoxin types, contamination levels, traceability analysis, prevention and control suggestions, etc.;

[0116] (2) Information release unit: Generate a unique electronic identifier for the warning report, and release it to the public in a multi-terminal and mobile-first manner through web services, and at the same time seamlessly connect with the government food safety supervision platform. The public and government agencies can query and subscribe to warning information at any time through browsers, mobile apps, etc.

[0117] Through the above automated and structured warning information release process, this system can significantly shorten the warning response time and improve the dissemination range and accessibility of warning information.

[0118] To continuously improve the efficiency and accuracy of the detection and warning of this system, the system self-learning and optimization module 5 introduces an online learning and active learning mechanism to realize the automatic tuning of model parameters and the incremental update of the knowledge base, including:

[0119] (1) Incremental learning unit: This unit continuously monitors the latest scientific research literature in the field of mycotoxin research through natural language processing technology, and calls the named entity recognition and relationship extraction algorithms to automatically extract the new mycotoxin gene sequences reported in the literature. After the extracted new toxin gene sequences are manually reviewed, they are automatically added to the toxin gene coding database to realize the dynamic update of the knowledge base. At the same time, the newly incorporated sample data is regularly supplemented to the training set, and the model parameters of the sequencing data analysis module and the intelligent risk assessment module are updated and optimized through incremental learning, so as to adapt to the dynamic changes of the mycotoxin prevalence spectrum.

[0120] (2) Active learning unit: Aiming at the complexity and unknownness of mycotoxin contamination, this unit designs an active learning strategy based on uncertainty. It regularly screens out the samples with the most information gain for model discrimination from a massive sample pool, guides experts to prioritize manual annotation and experimental verification, and promptly incorporates the newly obtained annotated samples into the training set to update the model. The expert feedback and manual annotation required by the active learning unit are provided by a third-party authoritative testing agency to ensure data quality. This unit can significantly reduce the manpower and material resources required for model training and iterative optimization, and continuously improve the system's ability to handle new and complex pollution events.

[0121] In addition, the system also sets up a data storage module 6 and a high-performance computing module 7 to provide support for large-scale data storage, computing, and sharing:

[0122] (1) The data storage module 6 adopts the HDFS distributed file system to uniformly and highly fault-tolerantly persistently store and manage sample sequencing data, process metadata, model parameters, etc. The data storage capacity can be elastically expanded;

[0123] (2) The high-performance computing module 7 adopts the Spark in-memory computing framework to use distributed nodes to parallel accelerate computationally intensive tasks such as sequencing data analysis, feature engineering, and machine learning training, significantly improving the system's computing efficiency and throughput capacity.

[0124] In summary, the fruit mycotoxin contamination early warning system constructed by the present invention takes the sequencing data analysis module as the core, integrates a number of innovative technologies such as semantic feature extraction, adaptive clustering, machine learning early warning, and self-learning optimization, and forms a set of high-throughput, intelligent, and full-process mycotoxin contamination prevention and control solutions. Compared with traditional detection methods, this system has obvious advantages in terms of detection sensitivity, analysis efficiency, early warning timeliness, etc. At the same time, the system adopts a Web-based and service-oriented architecture, with good openness, scalability, and interoperability.

[0125] The mycotoxin contamination early warning system of the present invention adopts a modular and distributed software and hardware architecture and can be flexibly deployed in various computing environments. The hardware structure of the system mainly includes data acquisition devices, high-performance computing clusters, distributed storage clusters, and Web servers, etc.

[0126] Among them, the data acquisition devices mainly refer to laboratory instruments used for sample preparation and high-throughput sequencing, including homogenizers, centrifuges, library construction instruments, sequencers, etc. These instruments exchange data with the sample pretreatment module 1 and the data storage module 6 of the system through the laboratory information management system (LIMS).

[0127] The high-performance computing cluster consists of multiple computing nodes, and each node is equipped with computing resources such as high-performance CPUs, GPUs, and large-capacity memories. Each computing node communicates through a high-speed interconnection network such as InfiniBand. Each algorithm in the sequencing data analysis module runs on this cluster. Among them, the literature coding comparison unit 21, the semantic feature extraction unit 22, and the adaptive clustering unit respectively correspond to different computing nodes in the cluster. Each computing node works collaboratively to jointly complete the distributed processing and analysis of sequencing data. To achieve efficient task scheduling and load balancing, the system adopts big data processing frameworks such as Hadoop and Spark, and uses parallel programming models such as MapReduce to automatically decompose and distribute computing tasks.

[0128] The distributed storage cluster consists of multiple storage nodes, and each node is equipped with a large-capacity disk array. Each storage node jointly provides PB-level massive data storage capabilities through a network file system (such as HDFS). Sequencing raw data, metadata, feature data, model parameters, etc. are all persistently stored in this cluster. The data storage module uniformly manages and schedules the above data, and adopts mechanisms such as hot and cold tiering, automatic fault tolerance, and load balancing to ensure the high availability and high-performance access of the data.

[0129] The Web server deploys a warning information publishing module, uses Web development frameworks such as PHP and Node.js, and provides a Web interaction interface for users. The server adopts load balancing technology and supports concurrent access by multiple users.

[0130] In addition, this system is also equipped with a dedicated management server, which deploys an intelligent risk assessment module 3 and a system self-learning optimization module 5 to coordinate the calls and data flows between various modules, and implement management functions such as monitoring, scheduling, and optimization of the entire system. Among them, the intelligent risk assessment module 3 uses machine learning frameworks such as TensorFlow to call GPU resources to train and infer machine learning models; the system self-learning optimization module 5 uses deep learning frameworks such as PyTorch and uses incremental learning and active learning strategies to optimize various indicators of the system.

[0131] The mapping relationships between the functional modules of this system and the hardware resources are as follows: The sample pretreatment module 1 depends on the data acquisition device and the distributed storage cluster to complete the acquisition, storage, and management of data; the sequencing data analysis module 2 depends on the high-performance computing cluster and the distributed storage cluster, and uses algorithms such as literature coding comparison, semantic feature extraction, and adaptive clustering to realize the intelligent analysis of high-throughput sequencing data; both the intelligent risk assessment module 3 and the system self-learning optimization module 5 run on the management server, depending on the machine learning framework and the deep learning framework respectively, to realize the training of the early warning model and the self-optimization of the system performance; the early warning information publishing module 4 runs on the Web server to provide users with interactive and friendly early warning information query and statistical analysis functions.

[0132] According to the task requirements, each module dynamically applies for the required computing and storage resources, and persists the intermediate results and metadata to the distributed storage cluster. By decoupling the software modules from the hardware resources, this system realizes flexible resource sharing and on-demand allocation, improving the overall resource utilization rate of the system. At the same time, thanks to the high scalability of the distributed architecture, this system can be easily expanded and upgraded to meet the growing data processing requirements.

[0133] It should be noted that although the present invention provides a relatively complete software and hardware system solution, the specific configuration of the system can be appropriately tailored and adjusted according to the actual application scenarios and business requirements. Those skilled in the art can, inspired by the present invention, design various different software and hardware architectures and deployment solutions in order to further optimize the system performance and enhance the applicability and economy of the system.

[0134] Example 2

[0135] Please refer to Figure 3 , the present invention also provides a method for warning fruit mycotoxin contamination based on high-throughput sequencing and machine learning, including the following steps:

[0136] Step 1. Sample pretreatment: Using the sample pretreatment module 1, perform a series of pretreatment operations such as physical crushing, chemical lysis, and molecular purification on the fruit samples to be tested, and construct a high-quality sequencing library;

[0137] Step 2. Sequencing data analysis: Using the sequencing data analysis module 2, perform high-throughput sequencing on the library to obtain the original sequencing data; and screen the suspected toxin gene reads through the innovative literature coding comparison algorithm, then learn the deep semantic representation of the reads through the semantic feature extraction algorithm, and finally, through adaptive DBSCAN clustering and cluster confidence evaluation, identify the high-confidence fungal toxin gene read clusters;

[0138] Step 3. Intelligent risk assessment: Using the intelligent risk assessment module 3, based on the expression abundance, harmfulness, and environmental meteorological characteristics of mycotoxin genes detected in Step 2, the risk level of mycotoxin contamination in the sample is real-time warned through the random forest classification algorithm;

[0139] Step 4. Early warning information release: Using the early warning information release module 4, a structured early warning report is automatically generated and released to achieve the timely transmission and sharing of early warning information to government regulatory departments and the public;

[0140] Step 5. System self-learning optimization: Using the system self-learning optimization module 5, the system adapts to the dynamic changes of the mycotoxin prevalence spectrum through incremental learning and active learning mechanisms, and continuously improves the efficiency and accuracy of the full-process detection and early warning.

[0141] The innovation points of the present invention are mainly reflected in:

[0142] (1) Aiming at the problem of low sequence alignment efficiency, the SimHash locality-sensitive hashing algorithm and threshold adaptive strategy are innovatively introduced to achieve the optimal balance between alignment sensitivity and computational efficiency, realizing high-throughput toxin gene screening.

[0143] (2) Aiming at the problem of shallow feature representation, the semantic segmentation convolutional neural network is first introduced into the characterization learning of sequencing reads, and deep discriminative features contained in the reads are automatically extracted through means such as embedding layers, multi-scale convolutions, and attention mechanisms.

[0144] (3) Aiming at the problem of insufficient accuracy in cluster recognition, the traditional DBSCAN clustering algorithm is improved, and an adaptive clustering strategy based on shared nearest neighbor similarity is proposed, which can flexibly discover cluster structures of arbitrary shapes, suppress noise interference, and improve clustering quality.

[0145] (4) At the system level, cutting-edge technologies such as multi-dimensional heterogeneous feature fusion, self-learning optimization, and distributed computing are innovatively integrated and optimized to construct an all-process and integrated intelligent early warning system for mycotoxin contamination based on high-throughput sequencing and machine learning.

[0146] In summary, the present invention is based on the urgent needs and technical challenges faced in the current food safety field, conducts systematic and original technological innovations, and forms a complete and excellent mycotoxin contamination prevention and control solution. This system is expected to fill a number of technical gaps in related fields at home and abroad, and has very important practical significance and broad application prospects for improving the level of food safety guarantee in China, safeguarding the vital interests of consumers, and promoting the sustainable development of the agricultural economy.

[0147] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fruit mycotoxin contamination early warning system based on high-throughput sequencing and machine learning, characterized in that: include: Sample pre-processing module, sequencing data analysis module, intelligent risk assessment module, early warning information release module and system self-learning optimization module; The sample pre-processing module is used to crush, homogenize, extract DNA and construct a library for the fruit samples; The sequencing data analysis module is used to perform high-throughput sequencing on the library to obtain sequencing reads data, and to accurately detect mycotoxin genes using a literature coding comparison algorithm and a semantic feature extraction algorithm; The intelligent risk assessment module is used to assess the sample contamination risk level according to the abundance and harmfulness of the mycotoxin genes detected by the sequencing data analysis module; The warning information publishing module is used to automatically generate warning reports and publish and share them through the Web terminal; The system self-learning optimization module is used to perform incremental learning and iterative optimization on the literature coding comparison algorithm and semantic feature extraction algorithm in the sequencing data analysis module according to the newly added data; The document coding comparison algorithm comprises the following steps: Firstly, a toxin gene coding database was constructed, and the toxin gene sequence was converted into a 64-bit binary coding vector using the SimHash algorithm; Then, the Hamming distance H between the sequencing reads and the encoding vector in the database is calculated. d : Among them, read i and ref i represents the i-th bit of the sequencing reads and the reference encoding vector respectively, ⊕ is the XOR operator, Next, set the adaptive threshold Among them, α and β are balance factors, n is the number of reads in the current batch, H d Mean, Finally, screen the suspected toxin gene reads: d Less than The reads were judged as suspected toxin gene reads.

2. The system according to claim 1, characterized in that The semantic feature extraction algorithm comprises the following steps: First, the suspected toxin gene reads are converted into a one-hot encoding matrix R∈R l×4 , and through the embedding matrix W e ∈R 4×d Dimensionality reduction to obtain dense vector X∈R l×d :X=R·W e Among them, l is the reads length, d is the embedding vector dimension; Then, a multi-scale convolution kernel W is used c ∈R k×d Extract the local features of reads and perform nonlinear transformation through the ReLU function σ: in, is the convolution operation, b c is the bias term Then, the multi-scale convolution features are adaptively integrated through the attention mechanism to obtain the semantic representation vector of reads. : Among them, α k is the attention weight of the kth convolutional feature: is a learnable query vector; Design a loss function that integrates read similarity and position relationship to optimize model parameters: Among them, P and N are the positive and negative sample sets respectively, d ij is the distance between readsi and j on the genome, and λ is the balance coefficient; Finally, adaptive DBSCAN clustering is performed on the reads semantic representation vector S to obtain high-confidence toxin gene reads clusters.

3. The system according to claim 2, characterized in that The process of the adaptive DBSCAN clustering includes: First, calculate any two semantic representation vectors S i and S j The Euclidean distance d ij =∥S i -S j ∥ 2; Then, a similarity graph between reads is constructed based on the shared nearest neighbor similarity. ij The calculation formula is: Among them, Γ(S i ) indicates S i The k nearest neighbor set of ; Finally, the DBSCAN clustering algorithm was applied on the similarity graph to automatically determine the number of clusters and obtain the toxin gene reads clusters.

4. The system according to claim 3, characterized in that After obtaining the toxin gene reads cluster, the cluster confidence score function is further used to evaluate the quality of each cluster, and the high confidence cluster is used for subsequent analysis. i The calculation formula is: Among them, L i is the minimum distance between clusters i, reflecting the separation between clusters: μ i and μ j are the centroid vectors of cluster i and cluster j respectively; H i is the average intra-cluster distance of cluster i, reflecting the compactness within the cluster:

5. The system according to claim 1, characterized in that ,The intelligent risk assessment module adopts the random forest algorithm and ,takes the toxin gene types, abundance and harmfulness detected by the ,sequencing data analysis module as features to construct a ,graded warning model for the risk of mycotoxin contamination in samples.

6. The system according to claim 1, characterized in that , the warning information release module includes: (1) a report generation unit, used to automatically generate a structured pollution warning report based on the warning results of the intelligent risk assessment module; (2) An information publishing unit, used to publish the warning report through Web services to achieve real-time sharing of warning information.

7. The system according to claim 1, characterized in that , the system self-learning optimization module includes: (1) an incremental learning unit, used to regularly incorporate newly collected fruit samples into the data set, and perform incremental training on the literature coding comparison algorithm and semantic feature extraction algorithm of the sequencing data analysis module; (2) A parameter tuning unit, which is used to continuously optimize the hyperparameters of the algorithm based on new data to improve algorithm performance.

8. The system according to any one of claims 1 to 7, characterized in that , the system further comprises: (1) a data storage module, used to persistently store and manage the sample sequencing data, metadata, and model parameters using a distributed file system; (2) A high-performance computing module for accelerating the computationally intensive tasks of sequencing data analysis and model training using a parallel computing framework.

Citation Information

Patent Citations

  • Microbial identification and analysis system and device based on next-generation high-throughput sequencing

    CN111276185A

  • Analysis method for high-throughput prediction of phage hosts based on next-generation sequencing technology

    CN115662516A