Plant volatile oil detection error traceability analysis system and method based on machine learning
By using a machine learning system to collect, test, and process plant volatile oil samples, combined with error source analysis, the problems of low analysis efficiency and insufficient consistency in existing technologies have been solved, achieving efficient and accurate quality evaluation and error source tracing.
Patent Information
- Application Number
- CN202511390677.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-06
AI Technical Summary
In existing technologies, the detection and evaluation of plant volatile oils largely rely on manual analysis, resulting in low analytical efficiency and insufficient consistency.
A machine learning-based error tracing and analysis system for plant volatile oil detection was adopted, which includes modules for sample collection, component detection, data preprocessing, machine learning analysis, and error tracing. The system achieves quality evaluation and error identification through a joint prediction model, and performs correction and quality evaluation by combining the error contribution.
It improves the efficiency and consistency of plant volatile oil quality analysis, realizes the quantitative expression of errors and the accuracy of quality results, and outputs intuitive quality evaluation and error traceability reports.
Smart Images

Figure CN121275929A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant volatile oil detection technology, specifically to a machine learning-based error tracing and analysis system and method for plant volatile oil detection. Background Technology
[0002] Plant volatile oils are a class of natural organic compounds with distinctive odors and biological activities, synthesized and released by plants through metabolic pathways. They have wide applications in medicine, fragrance, and daily chemical industries. The volatile oils contained in plants of the *Litsea* genus are rich and complex in composition, often exhibiting differences between samples from different parts, origins, and processing conditions. Detecting and analyzing these components, and conducting quality evaluation accordingly, is a crucial aspect of plant resource utilization and product development. In recent years, with the development of artificial intelligence, machine learning-based detection and analysis methods have been introduced into plant volatile oil research to improve detection accuracy and intelligent processing of results. Currently, the detection and evaluation of plant volatile oils largely rely on researchers performing component analysis, quality evaluation, or error analysis separately.
[0003] However, current technologies separate the identification of quality results and error sources, resulting in low analysis efficiency and insufficient consistency. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a machine learning-based error tracing and analysis system and method for detecting plant volatile oils, solving the problems of low analysis efficiency and insufficient consistency.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a machine learning-based error tracing and analysis system for plant volatile oil detection, comprising:
[0006] The sample collection module collects samples of plants in the genus Litsea and records sampling information and experimental conditions;
[0007] The component detection module performs chemical component analysis on the sample to obtain the raw detection data of Litsea cubeba volatile oil;
[0008] The data preprocessing module performs signal processing and data standardization on the raw detection data to form standardized data that can be used for modeling.
[0009] The machine learning analysis module builds and runs models based on standardized data, and outputs quality evaluation results and detection error identification results for Litsea cubeba volatile oil.
[0010] The error tracing module analyzes the sources of error by detecting error identification results and determines the contribution of each error source.
[0011] The quality evaluation module corrects the quality evaluation results by combining the contribution of error sources, and outputs a comprehensive quality evaluation and error tracing report for Litsea cubeba volatile oil.
[0012] Preferably, the sample collection includes:
[0013] Collect samples from different parts of plants in the genus Litsea, including leaves, fruits, or branches, and number and label the samples with their origin.
[0014] Record the time, location, ecological environment conditions, processing methods, and experimental conditions of the sample.
[0015] Preferably, the component detection includes:
[0016] The chemical components of the Litsea genus plant samples were detected by high performance liquid chromatography and gas chromatography-mass spectrometry, and the detection signals were obtained.
[0017] The detection signal is converted into raw detection data of Litsea cubeba volatile oil and stored.
[0018] Preferably, the data preprocessing includes:
[0019] Signal processing is performed on the raw detection data, including baseline correction, noise removal, and peak identification;
[0020] The processed data is then aligned by retention time and normalized by intensity to form standardized data.
[0021] Preferably, the machine learning analysis includes:
[0022] A joint prediction model is established based on standardized data to simultaneously output quality evaluation results and detection error identification results.
[0023] The standardized data of the sample to be tested are input into the joint prediction model to obtain the corresponding quality evaluation results and detection error identification results.
[0024] The results are accompanied by confidence level information, and a verification or retesting mark is given when the confidence level is lower than a preset threshold;
[0025] Preferably, the machine learning analysis module further includes abnormal sample identification and cluster analysis of standardized data to distinguish between normal samples and abnormal samples.
[0026] Preferably, the error tracing includes:
[0027] The residuals between the predicted and measured values are calculated and compared with the results of the reference samples.
[0028] Statistical analysis methods are used to classify the errors generated during the detection process and to calculate the contribution of each type of error.
[0029] Preferably, the contribution of each type of error is calculated by decomposing the residual between the predicted and measured values and using weight allocation and regression analysis.
[0030] Preferably, the quality evaluation module constructs a comprehensive quality index by weighting and combining the quality evaluation results with the various error contribution rates calculated by the error tracing module, and classifies the quality of Litsea cubeba volatile oil based on the comprehensive quality index.
[0031] A machine learning-based method for error tracing and analysis in plant volatile oil detection, the method comprising:
[0032] S1. Collect samples of Litsea genus plants and record sampling information and experimental conditions;
[0033] S2. Chemical composition analysis of the sample was performed to obtain the raw analysis data of Litsea cubeba volatile oil;
[0034] S3. Perform signal processing and data standardization on the raw detection data to form standardized data that can be used for modeling;
[0035] S4. Establish and run a machine learning model based on standardized data to output the quality evaluation results and detection error identification results of Litsea cubeba volatile oil;
[0036] S5. Conduct source analysis on the detection error identification results, determine the different types of errors generated during the detection process, and calculate their contribution.
[0037] S6. Combine the quality evaluation results with the error contribution rate, correct the quality evaluation results, and generate a quality evaluation and error traceability report for Litsea cubeba volatile oil based on the corrected results.
[0038] This invention provides a machine learning-based error tracing and analysis system for plant volatile oil detection. It has the following beneficial effects:
[0039] 1. This invention incorporates quality evaluation and error identification into the same machine learning framework, enabling the simultaneous output of two types of results. Furthermore, through a multi-task joint optimization strategy, it simultaneously learns quality and error characteristics, thereby improving the efficiency and consistency of Litsea cubeba volatile oil quality analysis.
[0040] 2. This invention categorizes detection errors after prediction by using residual calculation, reference sample comparison, and statistical analysis. Based on residual decomposition and weight allocation, it calculates the contribution of different error sources, thus clarifying the sources of error and achieving a quantitative expression of the impact of error, providing a basis for result interpretation and correction.
[0041] 3. This invention combines quality evaluation results with various error contribution factors in a weighted manner, and classifies quality levels based on an index, thereby achieving coupled evaluation of quality and error factors, outputting intuitive grading conclusions and reports, and forming a complete quality evaluation system that integrates testing, traceability and evaluation. Attached Figure Description
[0042] Figure 1 This is an architecture diagram of the machine learning-based error tracing and analysis system for plant volatile oil detection according to the present invention.
[0043] Figure 2 This is a flowchart of the machine learning-based error tracing analysis method for plant volatile oil detection according to the present invention. Detailed Implementation
[0044] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Please see the appendix Figure 1 This invention provides a machine learning-based error tracing and analysis system for plant volatile oil detection, comprising:
[0046] The sample collection module collects samples of plants in the genus Litsea and records sampling information and experimental conditions;
[0047] Furthermore, sample collection includes:
[0048] Collect samples from different parts of plants in the genus Litsea, including leaves, fruits, or branches, and number and label the samples with their origin.
[0049] Record the time, location, ecological environment conditions, processing methods, and experimental conditions of the samples.
[0050] Specifically, samples are taken from the leaves, fruits, or branches of the target Litsea genus plants. The samples are cut and packaged using disposable or sterilized sampling tools in accordance with aseptic or clean procedures. Immediately after collection, a unique number, such as a barcode or QR code, is generated for each sample. The source information, such as the sampling site, plant number, or plot number, is simultaneously marked on the container and electronic record. At the same time, sealing and preliminary status recording are completed to ensure that the sample and information correspond one by one and maintain traceability.
[0051] During sampling, the collection time, such as timestamp, collection location, and post-collection processing methods and experimental conditions, such as drying temperature, drying time, and drying method, are recorded simultaneously. The participants and the numbers of the instruments or equipment used are also recorded and bound to the sample number to form structured metadata, which is used for condition control and source tracing during quality evaluation and error tracing.
[0052] The component detection module performs chemical component analysis on the sample to obtain the raw detection data of Litsea cubeba volatile oil;
[0053] Furthermore, component analysis includes:
[0054] High performance liquid chromatography and gas chromatography-mass spectrometry were used to detect the chemical components of Litsea genus plant samples and obtain the detection signals.
[0055] The detection signal is converted into raw detection data of Litsea cubeba volatile oil and stored.
[0056] Specifically, the processed Litsea cubeba plant samples were analyzed by high performance liquid chromatography and gas chromatography-mass spectrometry to separate and identify various compounds in the volatile oil, such as identifying major components like citral and linalool and determining their content, thereby comprehensively obtaining the chemical composition of Litsea cubeba volatile oil and ensuring the accuracy and completeness of the test results.
[0057] The signals output by the testing instrument are converted into digital raw test data and stored in a database, corresponding to the sample number and sampling information. This avoids errors from manual recording, ensures data traceability, and provides reliable input for subsequent data preprocessing and model analysis.
[0058] The data preprocessing module performs signal processing and data standardization on the raw detection data to form standardized data that can be used for modeling.
[0059] Further data preprocessing includes:
[0060] Signal processing is performed on the raw detection data, including baseline correction, noise removal, and peak identification;
[0061] The processed data is then aligned by retention time and normalized by intensity to form standardized data.
[0062] Specifically, baseline correction, noise removal, and peak identification are performed on the raw detection data. The algorithm automatically corrects the baseline drift of the chromatographic signal, eliminates random noise, and accurately identifies the position and size of the peaks, ensuring that the detection signal is more stable and clear, and obtaining effective data that can truly reflect the information of volatile oil components.
[0063] After signal processing, the data is retained for time alignment and intensity normalization to ensure that data obtained from different batches and under different detection conditions are consistent in terms of time scale and signal intensity. This eliminates the influence of batch differences and instrument fluctuations, generates standardized data in a unified format, and provides reliable input for subsequent machine learning modeling.
[0064] The machine learning analysis module builds and runs models based on standardized data, and outputs quality evaluation results and detection error identification results for Litsea cubeba volatile oil.
[0065] Furthermore, machine learning analysis includes:
[0066] A joint prediction model is established based on standardized data to simultaneously output quality evaluation results and detection error identification results;
[0067] The standardized data of the sample to be tested are input into the joint prediction model to obtain the corresponding quality evaluation results and detection error identification results.
[0068] The results are accompanied by confidence level information, and a verification or retesting mark is given when the confidence level is lower than a preset threshold;
[0069] The machine learning analysis module also includes abnormal sample identification and cluster analysis of standardized data to distinguish between normal and abnormal samples.
[0070] Specifically, a joint prediction model is established based on standardized data, capable of simultaneously outputting quality evaluation results and detection error identification results. During training, the feature vector of each sample is used as input, and the corresponding quality grade label and error category label are used as supervision signals. A multi-task joint optimization strategy is adopted to enable the same model to learn the discrimination patterns related to quality and the discrimination patterns related to error, so as to cover two types of tasks with one modeling, reduce the bias caused by inconsistency between models, and thus improve prediction efficiency and consistency under the same data distribution.
[0071] The standardized data of the sample to be tested is input into the trained joint prediction model. The model forward inference simultaneously provides the quality evaluation result and the detection error identification result of the sample. The two types of results and the sample number are written into the system cache or database for subsequent traceability and report generation. The feature vector is automatically mapped into a classification output that can be directly used for decision-making, realizing fast and automated batch judgment.
[0072] Then, the probability scores given by the model for each candidate category are read, and the maximum value is taken as the confidence level of the sample on the corresponding task. It is compared with the preset threshold. If it is lower than the threshold, it is marked as needing to be reviewed or retested to enter the manual or repeated detection process, thereby avoiding low confidence conclusions from directly entering the quality evaluation stage.
[0073] Anomaly detection can be performed using Mahalanobis distance, calculated as follows:
[0074] Where X is the standardized feature vector of the sample, μ is the mean vector of the training or calibration set, Σ is the covariance matrix of the set, and D... M The statistical distance from the sample to the population distribution is defined as the distance to the sample. Samples with a distance greater than the threshold τ are considered abnormal.
[0075] Cluster analysis can employ a partitioning method based on distortion minimization, with the objective function being:
[0076]
[0077] Where K is the number of clusters, C i Let μ be the sample set of the i-th cluster. i Let J be the centroid vector of the i-th cluster. By alternately performing sample allocation and centroid update until J converges, clustering is completed. In addition to model inference, quality control at the data distribution level is provided to promptly identify distribution shifts or mixed samples and prevent them from interfering with subsequent evaluation and tracing.
[0078] The error tracing module analyzes the sources of error by detecting error identification results and determines the contribution of each error source.
[0079] Furthermore, error tracing includes:
[0080] The residuals between the predicted and measured values are calculated and compared with the results of the reference samples.
[0081] Statistical analysis methods are used to classify the errors generated during the detection process and to calculate the contribution of each type of error.
[0082] The contribution of various errors is calculated by decomposing the residuals between predicted and measured values and using weighting and regression analysis.
[0083] Specifically, after obtaining the detection error identification results, the magnitude of the error is first measured by calculating the residual between the predicted value and the measured value, and then compared with the results of the reference sample. The predicted value is denoted as... The measured value is denoted as y, and the residual calculation formula is:
[0084] r represents the residual vector. By comparing the detection results with those of the reference sample, it can be determined whether the residual comes from instrument drift, sample processing, or differences in environmental conditions. Subsequent tracing provides quantitative error information, making the existence and magnitude of the error clearly reflected.
[0085] After obtaining the residual distribution, the errors generated during the detection process are classified through statistical analysis. For example, by using analysis of variance or discriminant analysis, the residual characteristics caused by different sources are clustered and classified, thereby classifying the errors into equipment-related errors, sample processing-related errors, personnel operation-related errors, environmental condition-related errors, and errors caused by inherent differences in the samples. This transforms complex error information into structured categories, making the sources of errors clear and interpretable.
[0086] After error classification, the contribution of each type of error is calculated, which involves decomposing the residuals between the predicted and measured values and assigning weights accordingly.
[0087] Where r is the total residual, r k Let α represent the residual component of the k-th type of error. k The contribution weight of this type of error is determined by regression analysis or least squares method, where m is the total number of error categories. k By obtaining numerical values, we can obtain the quantitative contribution of different error sources to the total error, so that the error is not only limited to the classification level, but also presents its impact in numerical results, providing a basis for subsequent quality evaluation and correction.
[0088] The quality evaluation module corrects the quality evaluation results by combining the contribution of error sources, and outputs a comprehensive quality evaluation and error tracing report for Litsea cubeba volatile oil.
[0089] Furthermore, the quality evaluation module constructs a comprehensive quality index by weighting and combining the quality evaluation results with the various error contribution rates calculated by the error tracing module, and then classifies the quality of Litsea cubeba volatile oil based on the comprehensive quality index.
[0090] Specifically, after obtaining the quality evaluation results, the system will correct the quality evaluation results by combining the various error contribution rates output by the error tracing module. The quality evaluation results are recorded as follows: The contribution of each type of error is denoted as α. k The corrected quality results are expressed as follows:
[0091] Among them, Q ' The corrected quality evaluation result is shown below, where m is the number of error categories. Based on the above, deviations caused by different error sources are eliminated, making the corrected result closer to the true quality of the sample, ensuring the objectivity of the evaluation result, and improving the reliability of the test result.
[0092] After correcting the quality results, a comprehensive quality index is constructed by weighting the corrected quality evaluation results with various error contribution rates.
[0093] Where LQI represents the overall quality index, w1 is the weight of the quality result, and w k+1 By assigning weights to the contribution of various errors, the evaluation considers both the impact of quality and error, thus quantifying and comparing the quality of different samples.
[0094] Furthermore, the quality of Litsea cubeba volatile oil is graded based on a comprehensive quality index. The grading method uses interval division; for example, when LQI ≥ θ1, it is judged as Grade 1 quality, when θ2 ≤ LQI < θ1, it is judged as Grade 2 quality, and so on. θ1 and θ2 are preset quality grading thresholds, thereby transforming complex detection and correction results into intuitive grading conclusions and automatically generating reports containing quality grading and error traceability, providing a clear basis for the development, utilization, standard setting, and variety selection of Litsea cubeba volatile oil.
[0095] Please see the appendix Figure 2 A machine learning-based method for error tracing analysis in plant volatile oil detection, comprising the following methods:
[0096] S1. Collect samples of Litsea genus plants and record sampling information and experimental conditions;
[0097] S2. Chemical composition analysis of the sample was performed to obtain the raw analysis data of Litsea cubeba volatile oil;
[0098] S3. Perform signal processing and data standardization on the raw detection data to form standardized data that can be used for modeling;
[0099] S4. Establish and run a machine learning model based on standardized data to output the quality evaluation results and detection error identification results of Litsea cubeba volatile oil;
[0100] S5. Conduct source analysis on the detection error identification results, determine the different types of errors generated during the detection process, and calculate their contribution.
[0101] S6. Combine the quality evaluation results with the error contribution rate, correct the quality evaluation results, and generate a quality evaluation and error traceability report for Litsea cubeba volatile oil based on the corrected results.
[0102] Specifically, the process begins with standardized collection and information recording of Litsea cubeba plant samples. Then, chromatographic and mass spectrometric techniques are used to detect chemical components and obtain raw data on Litsea cubeba volatile oil. Subsequently, the data is standardized through baseline correction, noise reduction, alignment, and normalization to form inputs for model analysis. Based on this, a machine learning model is built and run to evaluate sample quality and identify detection errors. Furthermore, residual calculation and statistical analysis methods are used to classify error sources and calculate their contribution. Finally, the quality evaluation results are weighted and combined with the error contribution to correct the quality results, generating a comprehensive report containing quality grading and error tracing conclusions. This achieves an integrated process for Litsea cubeba volatile oil detection and error tracking.
[0103] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A machine learning based plant volatile oil detection error root cause analysis system, characterized in that, The system comprises: a sample collection module for collecting Litsea samples and recording sampling information and experimental conditions; a component detection module for detecting the chemical components of the samples to obtain original detection data of Litsea volatile oil; a data preprocessing module for signal processing and data standardization of the original detection data to form standardized data for modeling; a machine learning analysis module for establishing and running a model based on the standardized data to output quality evaluation results and detection error identification results of Litsea volatile oil; an error tracing module for analyzing error sources and determining the contribution of each error source based on the detection error identification results; a quality evaluation module for correcting the quality evaluation results in combination with the contribution of the error sources and comprehensively outputting a quality evaluation and error tracing report of Litsea volatile oil.
2. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The sample collection comprises: collecting samples of different parts of Litsea plants, including leaves, fruits or branches, and numbering and labeling the sources of the samples; recording the collection time, collection location, ecological environment conditions and processing methods and experimental conditions after collection of the samples.
3. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The component detection comprises: detecting the chemical components of the Litsea samples by high-performance liquid chromatography and gas chromatography-mass spectrometry to obtain detection signals; converting the detection signals into original detection data of Litsea volatile oil and storing them.
4. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The data preprocessing comprises: signal processing of the original detection data, including baseline correction, noise removal and peak identification; retention time alignment and intensity normalization of the processed data to form standardized data.
5. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The machine learning analysis comprises: establishing a joint prediction model for simultaneously outputting quality evaluation results and detection error identification results based on the standardized data; inputting the standardized data of the samples to be tested into the joint prediction model to obtain the corresponding quality evaluation results and detection error identification results; attaching confidence information to the results and giving a review or retest mark when the confidence is lower than a preset threshold.
6. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The machine learning analysis module further comprises abnormal sample identification and cluster analysis of the standardized data for distinguishing normal samples from abnormal samples.
7. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The error tracing comprises: calculating the residual error between the predicted value and the measured value and comparing it with the reference sample results; classifying the errors generated in the detection process by statistical analysis method and calculating the contribution of each type of error.
8. The machine learning based plant volatile oil detection error root cause analysis system of claim 7, wherein, The contribution of each type of error is calculated by decomposing the residual error between the predicted value and the measured value and using weight distribution and regression analysis method.
9. The machine learning based plant volatile oil detection error root cause analysis system of claim 1, wherein, The quality evaluation module constructs a comprehensive quality index by weighting and combining the quality evaluation results and the contribution of each type of error calculated by the error tracing module, and grades the quality of Litsea volatile oil based on the comprehensive quality index.
10. A method for error provenance analysis of plant volatile oil detection based on machine learning, characterized in that, The method for the machine learning-based plant volatile oil detection error tracing analysis system of any one of claims 1-9 comprises: S1, collecting Litsea samples and recording sampling information and experimental conditions; S2, detecting the chemical components of the samples to obtain original detection data of Litsea volatile oil; S3, signal processing and data standardization of the original detection data to form standardized data for modeling; S4, establishing and running a model based on the standardized data to output quality evaluation results and detection error identification results of Litsea volatile oil; S5, analyzing error sources and determining the contribution of each error source based on the detection error identification results; S6, correcting the quality evaluation results in combination with the contribution of the error sources and comprehensively outputting a quality evaluation and error tracing report of Litsea volatile oil. S4, based on the standardized data, a machine learning model is established and run to output quality evaluation results and detection error identification results of the Litsea cubeba volatile oil; S5, traceability analysis is performed on the detection error identification results to determine different types of errors generated in the detection process and calculate their contribution degrees; S6, the quality evaluation results are combined with the error contribution degrees, the quality evaluation results are corrected, and a quality evaluation and error traceability report of the Litsea cubeba volatile oil is generated based on the corrected results.