Grain and oil matrix standard substance candidate raw material accurate screening method based on characteristic fingerprint spectrum
By employing characteristic fingerprinting technology and a dynamic sampling decision system, the problems of subjectivity and insufficient sampling representativeness in the screening of candidate raw materials for standard substances have been solved, enabling precise screening and efficient preparation for complex polluted environments.
Patent Information
- Application Number
- CN202511623839.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-07
AI Technical Summary
In existing technologies, the screening of candidate raw materials for standard substances relies on the experience of researchers and limited targeted detection techniques. This results in problems such as strong subjectivity, blind spots in detection techniques, and insufficient sample representativeness, which leads to the prepared standard substances failing to accurately cover the concentration range of pollutants in actual samples.
A feature fingerprint-based approach is adopted, and a sampling point weight allocation model is established using remote sensing and historical soil data. Combined with multi-dimensional pollutant feature analysis, sparse partial least squares discriminant analysis, and multi-objective optimization algorithm, a pollutant fingerprint map and a three-dimensional pollution hotspot map are constructed to realize a dynamic sampling decision system.
It improves the accuracy and efficiency of raw material screening for standard substances, enabling it to more accurately reflect the actual conditions under complex pollution environments and enhancing the success rate and efficiency of standard substance preparation.
Smart Images

Figure CN121583313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental analysis and testing technology, and more specifically, to a method and system for precise screening of candidate raw materials for grain and oil matrix standard substances based on characteristic fingerprinting. Background Technology
[0002] Currently, the screening of candidate raw materials for reference materials mainly relies on researchers' experience and limited targeted detection technologies. Specifically, it typically involves simple random sampling in suspected contaminated areas, followed by laboratory analysis of the types and levels of pollutants in the samples to determine whether they meet the requirements for reference material development. This approach has significant drawbacks: First, raw material screening is highly subjective, lacking objective and quantitative indicators reflecting the spatial heterogeneity of pollutant distribution as decision support, leading to uncertainty about whether the pollutant concentration gradients of the screened raw materials meet the requirements for reference material development. Second, detection technologies have blind spots; traditional targeted detection is like "seeing the leopard through a tube," only confirming pollutants on a pre-set list and failing to discover and identify other unknown or novel pollutants and their combinations that may exist in the sample. This results in a limited variety of prepared reference materials, unable to meet the needs of complex pollution monitoring. Finally, sampling representativeness is severely insufficient. Due to the high spatial heterogeneity of pollutant distribution in environmental media (such as soil and crops), traditional sparse sampling schemes are prone to missing local pollution hotspots or underestimating the degree of pollution, making it impossible for the final prepared reference materials to accurately cover the concentration range in actual samples, affecting their accuracy and reliability as a measurement benchmark.
[0003] Therefore, it is necessary to provide a precise screening method for candidate raw materials of grain and oil matrix standard materials based on characteristic fingerprint spectrum to solve the problems of current standard material candidate material screening relying on researchers' experience and limited targeted detection technology, strong subjectivity in raw material screening, blind spots in detection technology, and serious lack of sampling representativeness. Summary of the Invention
[0004] In view of this, the present invention proposes a precise screening method for candidate raw materials of grain and oil matrix standard materials based on feature fingerprint spectrum, which aims to solve the problems of current standard material candidate raw material screening relying on researchers' experience and limited targeted detection technology, strong subjectivity in raw material screening, blind spots in detection technology, and serious lack of sampling representativeness.
[0005] This invention proposes a method for precise screening of candidate raw materials for grain and oil matrix standard substances based on characteristic fingerprint spectroscopy, comprising: A sampling point weight allocation model was established based on remote sensing and historical soil data to determine the sampling points; Grain samples were collected at the sampling points, and multi-dimensional pollutant characteristic analysis was performed on the grain samples to obtain a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins, and pesticide residues, as well as microbial community structure data; wherein, the multi-dimensional pollutant characteristic analysis includes non-targeted screening analysis, fungal toxin metabolomics analysis, and microbial community analysis; The initial data matrix is constructed by integrating the list of compounds for non-targeted screening, concentration data of heavy metals, toxins, and pesticide residues for targeted quantification, and microbial community structure data. The initial data matrix is then screened based on sparse partial least squares discriminant analysis and Moran's index to construct pollutant fingerprint profiles and generate a three-dimensional pollution hotspot map. A pollutant magnitude gradient database is established based on a 3D pollution hotspot map, and the optimal solution set is iteratively selected from the sampling points using a multi-objective optimization algorithm.
[0006] Furthermore, the sampling point weight allocation model established based on remote sensing and historical soil data, when determining sampling points, includes: The target farmland area is scanned by air using optical and radar sensors to obtain spectral image data of the entire target farmland area. In GIS software, the target farmland area is divided into regular grids; The acquired remote sensing data was overlaid with historical soil pollution survey data for analysis. The entropy weight method was used to calculate the weights of multiple environmental factors within each grid, and a sampling point weight allocation model was constructed.
[0007] Furthermore, the process of overlaying and analyzing the acquired remote sensing data with historical soil pollution survey data, and calculating the weights of multiple environmental factors within each grid using the entropy weight method to construct a sampling point weight allocation model includes: Construct an m×n decision matrix X = (xij)m×n; where xij represents the original measurement value of the j-th environmental factor in the i-th grid; m represents the total number of grids; and n represents the number of environmental factors. Environmental factors are normalized; among them, positive indicators of environmental factors are normalized: Yij=[xij-min(xj)] / [max(xj)-min(xj)]; In the above formula, Yij represents the normalized environmental factor; xij represents the original measurement value of the j-th environmental factor in the i-th grid; min(xj) represents the minimum value among all m grids for the j-th environmental factor; and max(xj) represents the maximum value among all m grids for the j-th environmental factor. Negative indicators among environmental factors: Yij=[max(xj)-xij] / [max(xj)-min(xj)]; In the above formula, Yij represents the normalized environmental factor; max(xj) represents the maximum value among all m grids for the j-th environmental factor; xij represents the original measurement value of the j-th environmental factor in the ith grid; and min(xj) represents the minimum value among all m grids for the j-th environmental factor. The information entropy ej of the j-th environmental factor is calculated using the following formula: ; ; In the above formula, ej represents the information entropy of the j-th environmental factor; k = 1 / ln(m) > 0; m represents the total number of grids; Yij represents the normalized environmental factor; The weight wj of the j-th environmental factor is calculated using the following formula: ; In the above formula, wj represents the weight of the j-th environmental factor; ej represents the information entropy of the j-th environmental factor; and n represents the number of environmental factors. The comprehensive pollution risk score Si for each grid is calculated using the following formula: ; In the above formula, Si represents the comprehensive pollution risk score of the i-th grid; n represents the number of environmental factors; wj represents the weight of the j-th environmental factor; Yij represents the normalized environmental factor; where Si∈[0,1]; The total number of sampling points is preset to N. The theoretical number of sampling points qi for each grid is calculated using the following formula: qi = Si × N; Round qi down to obtain the basic allocation quota ai; The difference R between the total number of sampling points and the total number of sampling points is calculated using the following formula: ; In the above formula, R represents the difference between the total number of sampling points and the total number of sampling points, that is, the remaining sampling point quota that needs to be redistributed; m represents the total number of grids; ai represents the basic allocation quota. Sort all grids in descending order of their corresponding qi fractional part (qi-ai). Allocate the remaining R slots to the R grids with the largest fractional parts in turn. Add 1 sampling point to the base allocation slot ai of each selected grid to obtain the actual number of sampling points ni for each grid.
[0008] Furthermore, the process of overlaying and analyzing the acquired remote sensing data with historical soil pollution survey data, calculating the weights of multiple environmental factors within each grid using the entropy weight method, and constructing a sampling point weight allocation model includes: Based on the final determined number of actual sampling points ni in each grid, the target farmland area is divided into three sampling priorities: If the actual number of sampling points ni ≥ 2, the grid is marked as a first-level priority area; if the actual number of sampling points ni = 1, the grid is marked as a second-level priority area; if the actual number of sampling points ni = 0, the grid is marked as a temporarily unsampled area.
[0009] Furthermore, when collecting grain samples at the sampling point and performing multi-dimensional pollutant characteristic analysis on the grain samples to obtain a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins, pesticide residues, and microbial community structure data, the process includes: Grain samples were collected at the sampling points and analyzed in depth using three methods: Using a high-resolution mass spectrometer, a full scan was performed, and the detected compound peaks were preliminarily identified by comparing them with public databases, thus identifying unknown pollutants or metabolites. Liquid chromatography-tandem mass spectrometry was used to perform metabolomics analysis of fungal toxins, quantifying major fungal toxins and their metabolites, including cryptic toxins; DNA was extracted from grain samples, and the ITS2 region of fungi was amplified. High-throughput sequencing was performed using the Illumina NovaSeq platform, and microbial community analysis was conducted. Through bioinformatics analysis, the types and relative abundance of toxin-producing fungi in the samples were identified. After three-way in-depth analysis, three types of core data were obtained for each sampling point: a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins, and pesticide residues, and microbial community structure data.
[0010] Furthermore, the main mycotoxins include: Aflatoxins, ochratoxins, fumonisins, deoxynivalenols, zearalenones, latrinones, patulin, cyclopiranic acid, penicillic acid, trichothecenes, and emerging Fusarium toxins.
[0011] Furthermore, when integrating the list of compounds for non-targeted screening, the concentration data of targeted quantitative heavy metals, toxins, and pesticide residues, and the microbial community structure data to form the initial data matrix, it includes: The non-targeted screening compound list, the concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and the microbial community structure data were merged according to sample ID to obtain an initial data matrix A×b; where A represents the total number of variables in the non-targeted screening compound list, the concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and the microbial community structure data, and b represents the number of grain samples.
[0012] Furthermore, the process of filtering the initial data matrix based on sparse partial least squares discriminant analysis and Moran's index, constructing a pollutant fingerprint map, and generating a three-dimensional pollution hotspot map includes: Based on the limits and internal control thresholds of GB 2761, GB 2762, and GB 2763, grain samples are divided into three categories: Clean sample G1: The concentrations of all detected contaminants are below the detection limit; Medium pollution G2: The concentration of any one of the detected pollutants is greater than or equal to the detection limit, and the concentration of any one of the detected pollutants is less than twice the limit. High pollution G3: The concentration of any one of the detected pollutants is greater than or equal to twice the limit; The initial A×b data matrix and the sample classification labels (G1, G2, G3) are input together into the sparse partial least squares discriminant analysis model; The five-fold cross-validation method was used to run the test a preset number of times to optimize the number of latent variables and the number of variables retained for each latent variable in the sparse partial least squares discriminant analysis model, and to determine the optimal parameter combination of the sparse partial least squares discriminant analysis model. Based on the trained sparse partial least squares discriminant analysis model, the variable projection importance score of each variable is calculated; a threshold for the variable projection importance score is set, and combined with the statistical significance level controlled by the false discovery rate, variables with variable projection importance scores greater than the variable projection importance score threshold and corrected p-values less than the significance level are selected from the total number of A variables to form a subset of key feature variables for constructing the pollutant feature fingerprint map; Spatial correlation is performed between the concentration or abundance data of each variable in the subset of key feature variables and the geographic coordinates of their corresponding sampling points. Based on the correlated data, the Moran index of the target farmland area is calculated to determine the overall spatial clustering of pollutants, and the local Moran index of the local space formed by each sampling point and its adjacent sampling points is calculated. The spatial clustering type is identified by the local Moran index; wherein, the spatial clustering type includes high-high clustering, low-low clustering, high-low anomaly and low-high anomaly. High-high concentration areas are defined as spatial hotspots of pollutants; using the concentration or abundance data of the subset of key feature variables as Z values, combined with their geographic coordinates, a continuously distributed three-dimensional pollution hotspot map is generated using the Kriging spatial interpolation method.
[0013] Furthermore, the process of establishing a pollutant magnitude gradient database based on a three-dimensional pollution hotspot map includes: Based on the three-dimensional pollution hotspot map and the corresponding subset of key feature variables, a relational database is constructed. The database includes unique identifiers for sampling points, geographical coordinates of sampling points, spatial hotspot partition identifiers to which sampling points belong, and concentration or abundance data of each variable in the subset of key feature variables.
[0014] Furthermore, when using a multi-objective optimization algorithm to iteratively select the optimal solution set among the sampling points, the process includes: Maximizing the coverage of pollutant concentration gradient range and minimizing the number of sampling points are set as two objectives that need to be optimized simultaneously. The multi-objective optimization algorithm is run on the pollutant quantity gradient database to iteratively select the optimal solution set, i.e., a series of sampling schemes, from all sampling points.
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: Firstly, this invention introduces the concept of "interaction networks" in ecosystems into the screening of raw materials for standard substances. By constructing pollutant-microorganism association networks through multi-omics data, it not only identifies individual pollutants but also reveals their synergistic and antagonistic coexistence patterns. This allows the standard substances prepared accordingly to more realistically reflect the actual situation under complex pollution environments, elevating the characterization of raw materials for real pollution scenarios from the one-sidedness of traditional methods to a systematic level. Secondly, this invention uniquely and seamlessly couples remote sensing optics and radar imagery, high-resolution mass spectrometry, and spatial geostatistics, forming a "spectral-mass spectrometry-spatial" three-in-one fingerprinting technology. This achieves a full-chain analysis from macroscopic distribution detection to microscopic molecular identification and spatial pattern quantification, improving the spatial representativeness of sampling points from discrete point inference to precise characterization of continuous space. Finally, this invention establishes a dynamic sampling decision system based on machine learning models, transforming raw material screening from a static, passive "screening" to a dynamic, proactive "design," thereby improving the accuracy of standard substance raw material screening and increasing the success rate and efficiency of standard substance preparation. Attached Figure Description
[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart of a method for precise screening of candidate raw materials for grain and oil matrix standard substances based on feature fingerprinting provided in an embodiment of the present invention; Figure 2 A three-dimensional visualization map provided for embodiments of the present invention. Detailed Implementation
[0017] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] In some embodiments of this application, see Figure 1 As shown, this embodiment provides a method for precise screening of candidate raw materials for grain and oil matrix standard substances based on feature fingerprinting, including the following steps: S100. Establish a sampling point weight allocation model based on remote sensing and historical soil data to determine the sampling points; S200. Collect grain samples at the sampling point and perform multi-dimensional pollutant characteristic analysis on the grain samples to obtain a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and microbial community structure data; wherein, the multi-dimensional pollutant characteristic analysis includes non-targeted screening analysis, fungal toxin metabolomics analysis and microbial community analysis. S300, integrating non-targeted screening compound list, targeted quantitative heavy metal, toxin and pesticide residue concentration data and microbial community structure data to form an initial data matrix, and using sparse partial least squares discriminant analysis and Moran index to screen the initial data matrix, construct pollutant fingerprint spectrum and generate a three-dimensional pollution hotspot map. S400. Based on the three-dimensional pollution hotspot map, establish a pollutant value gradient database, and use a multi-objective optimization algorithm to iteratively select the optimal solution set from the sampling points.
[0019] Understandably, this invention is the first to introduce the concept of "interaction networks" in ecosystems into the screening of raw materials for standard substances. By constructing pollutant-microorganism association networks through multi-omics data, it not only identifies individual pollutants but also reveals their synergistic and antagonistic coexistence patterns. This allows the standard substances prepared accordingly to more realistically reflect the actual situation under complex pollution environments, elevating the characterization of raw materials for real pollution scenarios from the one-sidedness of traditional methods to a systematic level. Secondly, this invention uniquely and seamlessly couples remote sensing, high-resolution mass spectrometry, and spatial geostatistics, forming a "spectral-mass spectrometry-spatial" three-in-one fingerprinting technology. This achieves a full-chain analysis from macroscopic distribution detection to microscopic molecular identification and spatial pattern quantification, improving the spatial representativeness of sampling points from the inference of discrete locations to the precise characterization of continuous space. Finally, this invention establishes a dynamic sampling decision system based on machine learning models, transforming raw material screening from a static, passive "screening" to a dynamic, proactive "design," thereby improving the accuracy of standard substance raw material screening and increasing the success rate and efficiency of standard substance preparation.
[0020] In some embodiments of this application, the step of establishing a sampling point weight allocation model based on remote sensing and historical soil data to determine sampling points includes: The target farmland area is scanned by air using optical and radar sensors to obtain spectral image data of the entire target farmland area. In GIS software, the target farmland area is divided into a regular grid of 100m × 100m; The acquired remote sensing data was overlaid with historical soil pollution survey data (such as heavy metal content, pH value, organic matter content, etc.), and the weights of various environmental factors (such as specific spectral reflectance characteristics and soil pollutant concentration) in each grid were calculated using the entropy weight method to construct a sampling point weight allocation model.
[0021] In some embodiments of this application, the process of overlaying and analyzing the acquired remote sensing data with historical soil pollution survey data, calculating the weights of multiple environmental factors within each grid using the entropy weight method, and constructing a sampling point weight allocation model includes: Construct an m×n decision matrix X = (xij)m×n; where xij represents the original measurement value of the j-th environmental factor in the i-th grid; m represents the total number of grids; and n represents the number of environmental factors. Environmental factors were normalized; among them, positive indicators of environmental factors (such as soil Cd content and mean 555nm reflectance R555, the higher the value, the more dangerous the environment) were normalized. Yij=[xij-min(xj)] / [max(xj)-min(xj)]; In the above formula, Yij represents the normalized environmental factor; xij represents the original measurement value of the j-th environmental factor in the i-th grid; min(xj) represents the minimum value among all m grids for the j-th environmental factor; and max(xj) represents the maximum value among all m grids for the j-th environmental factor. For negative indicators among environmental factors (such as soil pH, where higher soil pH indicates stronger alkalinity and decreased soil Cd activity, meaning higher soil pH is better): Yij=[max(xj)-xij] / [max(xj)-min(xj)]; In the above formula, Yij represents the normalized environmental factor; max(xj) represents the maximum value among all m grids for the j-th environmental factor; xij represents the original measurement value of the j-th environmental factor in the ith grid; and min(xj) represents the minimum value among all m grids for the j-th environmental factor. The information entropy ej of the j-th environmental factor is calculated using the following formula: ; ; In the above formula, ej represents the information entropy of the j-th environmental factor; k = 1 / ln(m) > 0; m represents the total number of grids; Yij represents the normalized environmental factor; The weight wj of the j-th environmental factor is calculated using the following formula: ; In the above formula, wj represents the weight of the j-th environmental factor; ej represents the information entropy of the j-th environmental factor; and n represents the number of environmental factors. The comprehensive pollution risk score Si for each grid is calculated using the following formula: ; In the above formula, Si represents the comprehensive pollution risk score of the i-th grid; n represents the number of environmental factors; wj represents the weight of the j-th environmental factor; Yij represents the normalized environmental factor; where Si∈[0,1]; The total number of sampling points is preset to N. The theoretical number of sampling points qi for each grid is calculated using the following formula: qi = Si × N; Round qi down to obtain the basic allocation quota ai; The difference R between the total number of sampling points and the total number of sampling points is calculated using the following formula: ; In the above formula, R represents the difference between the total number of sampling points and the total number of sampling points, that is, the remaining sampling point quota that needs to be redistributed; m represents the total number of grids; ai represents the basic allocation quota. Sort all grids in descending order of their corresponding qi fractional part (qi-ai). Allocate the remaining R slots to the R grids with the largest fractional parts in turn. Add 1 sampling point to the base allocation slot ai of each selected grid to obtain the actual number of sampling points ni for each grid.
[0022] In some embodiments of this application, the process of overlaying and analyzing the acquired remote sensing data with historical soil pollution survey data, calculating the weights of multiple environmental factors within each grid using the entropy weight method, and constructing a sampling point weight allocation model includes: Based on the final determined number of actual sampling points ni in each grid, the target farmland area is divided into three sampling priorities: If the actual number of sampling points ni ≥ 2, the grid is marked as a first-level priority area; if the actual number of sampling points ni = 1, the grid is marked as a second-level priority area; if the actual number of sampling points ni = 0, the grid is marked as a temporarily unsampled area.
[0023] In some embodiments of this application, when collecting grain samples at the sampling point and performing multi-dimensional pollutant characteristic analysis on the grain samples to obtain a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and microbial community structure data, the process includes: Grain samples were collected at the sampling points and analyzed in depth using three methods: Using a high-resolution mass spectrometer, a full scan was performed, and the detected compound peaks were preliminarily identified by comparing them with public databases, thus identifying unknown pollutants or metabolites. Liquid chromatography-tandem mass spectrometry was used to perform metabolomics analysis of fungal toxins, quantifying major fungal toxins and their metabolites, including cryptic toxins; DNA was extracted from grain samples, and the ITS2 region of fungi was amplified. High-throughput sequencing was performed using the Illumina NovaSeq platform, and microbial community analysis was conducted (using QIIME2 software). Through bioinformatics analysis, the types and relative abundance of toxin-producing fungi in the samples were identified. After three-way in-depth analysis, three types of core data were obtained for each sampling point: a list of compounds for non-targeted screening, concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and microbial community structure data.
[0024] In some embodiments of this application, the major mycotoxin includes: Aflatoxins, ochratoxins, fumonisins, deoxynivalenols, zearalenones, latrinones, patulin, cyclopiranic acid, penicillic acid, trichothecenes, and emerging Fusarium toxins.
[0025] Specifically, 1. Aflatoxins: AFB1, AFB2, AFG1, AFG2, AFM1; 2. Ochratoxins: OTA, OTB, OTC; 3. Fumonisins: FB1, FB2, FB3; 4. Trichothecenes-Type B: DON, 3-ADON, 15-ADON, DON-3G (caused); 5. Trichothecenes-Type B: ... A): T-2, HT-2, NEO, DAS; 6. Zearalenones: ZEN, α-ZEL, β-ZEL, ZAN, ZEN-14G (cryptic); 7. Alternariatoxins: AOH, AME, ALT, TeA; 8. Patulin: PAT; 9. Cyclopiazonic acid: CPA; 10. Mycophenolic acid: MPA; 11. Satratoxins (mainly found in moist stored grains): SAT-G; 12. Emerging Fusariumtoxins: Beauvericin, Enniatin A1, B, B1.
[0026] In some embodiments of this application, the integration of the list of compounds for non-targeted screening, the concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and microbial community structure data to form an initial data matrix includes: The non-targeted screening compound list, the concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and the microbial community structure data were merged according to sample ID to obtain an initial data matrix A×b; where A represents the total number of variables in the non-targeted screening compound list, the concentration data of targeted quantitative heavy metals, toxins and pesticide residues, and the microbial community structure data, and b represents the number of grain samples.
[0027] In some embodiments of this application, the step of filtering the initial data matrix based on sparse partial least squares discriminant analysis and Moran's index, constructing a pollutant fingerprint map, and generating a three-dimensional pollution hotspot map includes: Based on the limits and internal control thresholds of GB 2761, GB 2762, and GB 2763, grain samples are divided into three categories: Clean sample G1: The concentrations of all detected contaminants are below the detection limit; Medium pollution G2: The concentration of any one of the detected pollutants is greater than or equal to the detection limit, and the concentration of any one of the detected pollutants is less than twice the limit. High pollution G3: The concentration of any one of the detected pollutants is greater than or equal to twice the limit; The initial A×b data matrix and the sample classification labels (G1, G2, G3) are input together into the sparse partial least squares discriminant analysis model; The five-fold cross-validation method was used to run the test a preset number of times to optimize the number of latent variables and the number of variables retained for each latent variable in the sparse partial least squares discriminant analysis model, and to determine the optimal parameter combination of the sparse partial least squares discriminant analysis model. Based on the trained sparse partial least squares discriminant analysis model, the variable projection importance score of each variable is calculated; a threshold for the variable projection importance score is set, and combined with the statistical significance level controlled by the false discovery rate, variables with variable projection importance scores greater than the variable projection importance score threshold and corrected p-values less than the significance level are selected from the total number of A variables to form a subset of key feature variables for constructing the pollutant feature fingerprint map; Spatial correlation is performed between the concentration or abundance data of each variable in the subset of key feature variables and the geographic coordinates of their corresponding sampling points. Based on the correlated data, the Moran index of the target farmland area is calculated to determine the overall spatial clustering of pollutants, and the local Moran index of the local space formed by each sampling point and its adjacent sampling points is calculated. The spatial clustering type is identified by the local Moran index; wherein, the spatial clustering type includes high-high clustering, low-low clustering, high-low anomaly and low-high anomaly. High-high concentration areas are defined as spatial hotspots of pollutants; using the concentration or abundance data of the subset of key feature variables as Z values, combined with their geographic coordinates, a continuously distributed three-dimensional pollution hotspot map is generated using the Kriging spatial interpolation method.
[0028] In some embodiments of this application, the process of establishing a pollutant magnitude gradient database based on a three-dimensional pollution hotspot map includes: Based on the three-dimensional pollution hotspot map and the corresponding subset of key feature variables, a relational database is constructed. The database includes unique identifiers for sampling points, geographical coordinates of sampling points, spatial hotspot partition identifiers to which sampling points belong, and concentration or abundance data of each variable in the subset of key feature variables.
[0029] In some embodiments of this application, the step of iteratively selecting the optimal solution set from the sampling points using a multi-objective optimization algorithm includes: Maximizing the coverage of pollutant concentration gradient range and minimizing the number of sampling points are set as two objectives that need to be optimized simultaneously. The multi-objective optimization algorithm is run on the pollutant quantity gradient database to iteratively select the optimal solution set, i.e., a series of sampling schemes, from all sampling points.
[0030] Application examples: 1. Preparation of cadmium-deoxynivalenol (DON) synergistic contamination standard material in wheat-producing areas of the Huang-Huai-Hai Plain Technology Implementation (1) Grid sampling based on Geographic Information System (GIS) A 5,000-mu wheat field was selected at the border of Shandong and Henan provinces and divided into 100m×100m grids (500 grids in total). After integrating remote sensing data with soil survey data, 72 high-risk grids (weight score > 0.8) were selected. (2) Table 1 shows the results of multi-omics detection. Table 1
[0031] As shown in Table 1, in the high-risk fields identified by GIS grid mapping, cadmium is present in a highly bioavailable form (60% as Cd). 2+ The cadmium-Fusarium-DON co-contamination was simultaneously enriched with DON (including 15–22% of cryptic DON-3G), and the abundance of Fusarium could explain 73% of the DON variation, confirming that the "cadmium-Fusarium-DON" co-contamination can serve as a typical target for the preparation of standard materials in the Huang-Huai-Hai wheat region.
[0032] (3) Fingerprint mapping Fingerprint maps were constructed using sparse partial least squares discriminant analysis (sPLS-DA), and the stability of the features was verified by principal component analysis (PCA).
[0033] (4) Intelligent sampling decision The multi-objective genetic algorithm (NSGA-II) was used to select 12 sampling points from 72 candidate points, covering 4 pollution combination modes.
[0034] Application Example 2: Development of Standard Reference Materials for Multiple Toxins in Northeast China's Corn-Producing Area Solve the challenge of synergistic monitoring of cryptic toxins (such as zearalenone-14-O-β-glucoside, ZEN-14G) and novel mycotoxins (such as Alternaria alternata, AAL).
[0035] (1) Spatial heterogeneity analysis Moran's index showed that aflatoxin (AFB1) exhibited a significant clustered distribution (I=0.41, p<0.001). Hotspot areas account for 12% of the total area, but contribute 68% of the toxin load of the entire region; (2) Dynamic sampling verification: Table 2 shows the dynamic sampling verification results of the traditional method and the method in this invention.
[0036] Table 2
[0037] As shown in Table 2, after adopting the dynamic sampling method of the present invention, the non-compliance rate of the three rounds of verification was consistently ≤8%, which is about 24% lower than the traditional method on average. This significantly reduces the risk of missed detection of hidden / novel toxins such as ZEN-14G and AAL, and is suitable for the accurate development of multi-toxin standard materials in the corn-producing areas of Northeast China.
[0038] Application Example 3: Standard Material for Pesticide-Heavy Metal Complex Pollution in Rice in the Yangtze River Basin - Achieving Simultaneous Calibration of Pesticide Metabolites (such as Chlorpyrifos-3,5,6-TCP) and Heavy Metals (Inorganic Arsenic, As(III) / As(V)) Forms.
[0039] (1) Integration of multi-omics data: Table 3 shows the integration of multi-omics sampling data.
[0040] Table 3
[0041] Table 3 shows that pesticide-heavy metal complex pollution in rice in the Yangtze River Basin is characterized by high detection and high variability: chlorpyrifos and its metabolite 3,5,6-TCP are commonly found in the same pollutants, and the spatial variation (CV 92%) of the metabolites is much higher than that of the parent compound. Inorganic As(III) is detected in 100% of the samples, with a concentration variation of 65%. This confirms that simultaneous calibration of the parent compound, metabolites and As forms is necessary to meet the accuracy requirements of the standard reference materials for complex pollution in this region.
[0042] (2) See Figure 2 , Figure 2 The map is presented in a 3D visualization format, where the red bars represent chlorpyrifos and the blue bars represent As(III). Depend on Figure 2 (Due to the inclusion of pollution data, the coordinate axes are omitted in this schematic diagram for illustrative purposes only.) It can be seen that the overlapping areas of multiple pollutants can be visually queried, which serves as a guide for the precise sampling sites of standard substances. At the same time, it also serves as a reminder for daily monitoring to increase the frequency of sampling in such areas.
[0043] (3) Quantity gradient design Key gradient concentration points were identified through database association analysis (Table 4 shows the standard substance gradient design): Table 4
[0044] As shown in Table 4, chlorpyrifos and As(III) showed a synergistic increase in each gradient (R). 2 The concentration of chlorpyrifos (>0.99) in the industrial zone and its surrounding area (L5) is close to or reaches the national limits (chlorpyrifos 2 mg / kg, inorganic arsenic 0.35 mg / kg), highlighting the dual risk of exceeding the limits for both pesticides and heavy metals. Therefore, L5 should be used as the upper limit anchor point for the standard material. Using L1, L3, and L5, the minimum concentration difference can be ≥15 times, meeting the dual requirements of "representative matrix + concentration range" for the standard material. The three points correspond to clearly defined regional types and can be directly used to prepare candidates for the three levels of standard materials: "background - transition - pollution", achieving synchronous value transfer of "parent material + metabolite + As form" in three dimensions.
[0045] Through the implementation of the above application examples, the characteristic values of the candidate raw materials for the collected standard substances are consistent, and the concentration levels deviate from the design values by less than 5%, which provides technical support for the accurate acquisition of raw materials for standard substances and reduces the development cost of standard substances.
[0046] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0048] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0049] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for accurately screening grain and oil matrix standard material candidate raw materials based on characteristic fingerprint spectrum, characterized in that, The application relates to a method for constructing a three-dimensional pollution hotspot map based on a pollution fingerprint spectrum. The method comprises the following steps: a sampling point weight distribution model is established based on aerospace remote sensing and historical soil data, and sampling points are determined; grain samples are collected at the sampling points, and multi-dimensional pollutant characteristic analysis is performed on the grain samples to obtain a non-target screening compound list, concentration data of target quantitative heavy metal, toxin and pesticide residue and microbial community structure data; wherein the multi-dimensional pollutant characteristic analysis comprises non-target screening analysis, fungal toxin metabolomics analysis and microbial community analysis; the non-target screening compound list, the concentration data of target quantitative heavy metal, toxin and pesticide residue and the microbial community structure data are integrated to form an initial data matrix, data screening is performed on the initial data matrix based on a sparse partial least squares discriminant analysis method and a Moran index, a pollutant fingerprint spectrum is constructed, and a three-dimensional pollution hotspot map is generated based on the spectrum; 2. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 1, characterized in that, a pollutant value gradient database is established based on the three-dimensional pollution hotspot map, and a multi-objective optimization algorithm is used to iteratively screen optimal solution sets in the sampling points. When the sampling point weight distribution model is established based on aerospace remote sensing optical and radar images and historical soil data, the following steps are included: optical and radar sensors are used to perform flight scanning on a target farmland region to obtain image data of the entire target farmland region; the target farmland region is divided into regular grids in GIS software; 3. The method according to claim 2, wherein, the obtained remote sensing data are superimposed and analyzed with historical soil pollution census data, an entropy weight method is used to calculate the weights of various environmental factors in each grid, and a sampling point weight distribution model is constructed. When the obtained remote sensing data are superimposed and analyzed with historical soil pollution census data, the weights of various environmental factors in each grid are calculated by using an entropy weight method, and a sampling point weight distribution model is constructed, the following steps are included: an m*n order decision matrix X=(xij)m*n is constructed; wherein xij represents the original measurement value of the jth environmental factor of the ith grid; m represents the total number of grids; and n represents the number of environmental factors; normalization is performed on the environmental factors; wherein for positive indicators in the environmental factors: Yij=[xij-min(xj)] / [max(xj)-min(xj)]; in the formula, Yij represents the normalized environmental factor; xij represents the original measurement value of the jth environmental factor of the ith grid; min(xj) represents the minimum value of the jth environmental factor among all m grids; and max(xj) represents the maximum value of the jth environmental factor among all m grids; The information entropy ej of the jth environmental factor is calculated by the following formula: ; ; in the above formula, ej represents the information entropy of the jth environmental factor; k = 1 / ln(m) > 0; m represents the total number of grids; Yij represents the normalized environmental factor, and if åYij = 0, then Pij = 1 / m; The weight wj of the jth environmental factor is calculated by the following formula: ; in the formula, wj represents the weight of the jth environmental factor; ej represents the information entropy of the jth environmental factor; n represents the number of environmental factors, and if ∑(1-ej)=0, then wj=1 / n; The comprehensive pollution risk score Si of each grid is calculated by the following formula: ; in the formula, Si represents the comprehensive pollution risk score of the i-th grid; n represents the number of environmental factors; wj represents the weight of the j-th environmental factor; Yij represents the normalized environmental factor; and Si ∈ [0, 1]. for negative indicators in the environmental factors: Yij=[max(xj)-xij] / [max(xj)-min(xj)]; in the formula, Yij represents the normalized environmental factor; max(xj) represents the maximum value of the jth environmental factor among all m grids; xij represents the original measurement value of the jth environmental factor of the ith grid; and min(xj) represents the minimum value of the jth environmental factor among all m grids; the total number of preset sampling points is N, and the theoretical sampling point number basis qi of each grid is calculated by the following formula: qi=Si*N; Downward rounding qi, get the basic allocation quota ai; The difference R between the total number of sampling points and the total number of sampling points is calculated by the following formula: ; in the above formula, R represents the difference between the total number of sampling points and the total number of sampling points, i.e. the remaining sampling point quota that needs to be redistributed; m represents the total number of grids; ai represents the basic allocation quota; All grids are sorted according to the decimal part (qi-ai) of their corresponding qi from large to small, and the remaining R quotas are allocated to the first R grids with the largest decimal part in turn. Each selected grid increases 1 sampling point on its basic allocation quota ai to obtain the actual sampling point number ni of each grid.
4. The feature fingerprint-based grain oil base standard substance candidate raw material precision screening method according to claim 3, characterized in that, The acquired remote sensing data is superimposed and analyzed with historical soil pollution census data, the weight of each grid is calculated by entropy weight method, and a sampling point weight distribution model is constructed, including: According to the final determined actual sampling point number ni of each grid, the target farmland area is divided into three sampling priority levels: If the actual sampling point number ni of the grid is greater than or equal to 2, the grid is marked as a first priority area; if the actual sampling point number ni of the grid is equal to 1, the grid is marked as a second priority area; if the actual sampling point number ni of the grid is equal to 0, the grid is marked as a temporary sampling area.
5. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 4, characterized in that, When the grain seed sample is collected at the sampling point, the multi-dimensional pollutant characteristic analysis of the grain seed sample is performed to obtain the non-target screening compound list, the concentration data of targeted heavy metal, toxin and pesticide residue, and the microbial community structure data, including: Grain seed samples are collected at the sampling point, and deep analysis is performed in three ways: A high-resolution mass spectrometer is used for full scan, and the detected compound peaks are preliminarily identified by comparing public databases to identify unknown pollutants or metabolites; Liquid chromatography-tandem mass spectrometry is used for fungal toxin metabolomics analysis to quantitatively analyze main fungal toxins and their metabolites, including hidden toxins; DNA is extracted from the grain seed sample, the ITS2 region of fungi is amplified, high-throughput sequencing is performed on the Illumina NovaSeq platform, and microbial community analysis is performed. Through bioinformatics analysis, the types and relative abundance of toxin-producing fungi in the sample are determined. After three-way deep analysis, three types of core data of each sampling point are obtained: non-target screening compound list, concentration data of targeted heavy metal, toxin and pesticide residue, and microbial community structure data.
6. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 5, characterized in that, The main fungal toxins include: Aflatoxins, ochratoxins, fumonisins, deoxynivalenols, nivalenols, zearalenones, alternariols, patulins, cyclopiazonic acids, penicillic acids, trichothecenes, and emerging fusarium toxins.
7. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 6, characterized in that, When the non-target screening compound list, the concentration data of targeted heavy metal, toxin and pesticide residue, and the microbial community structure data are integrated to form an initial data matrix, including: The non-target screening compound list, the concentration data of targeted heavy metal, toxin and pesticide residue, and the microbial community structure data are merged according to the sample ID to obtain an Axb initial data matrix; wherein A represents the total number of variables of the non-target screening compound list, the concentration data of targeted heavy metal, toxin and pesticide residue, and the microbial community structure data, and b represents the number of grain seed samples.
8. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 7, characterized in that, The sparse partial least squares discriminant analysis method and the Moran index are used for data screening of an initial data matrix, and a pollutant fingerprint spectrum is constructed. According to the limit and internal control threshold values in GB 2761, GB 2762 and GB 2763, the grain kernel samples are divided into three categories: Clean samples G1: the concentrations of all detected pollutants are less than the detection limit; Medium pollution G2: the concentration of any detected pollutant is greater than or equal to the detection limit, and the concentration of any detected pollutant is less than twice the limit; High pollution G3: the concentration of any detected pollutant is greater than or equal to twice the limit; The A x b initial data matrix and the sample classification label (G1, G2, G3) are input into the sparse partial least squares discriminant analysis model. The five-fold cross-validation method is used to run the preset number of times to optimize the number of latent variables and the number of variables retained by each latent variable of the sparse partial least squares discriminant analysis model, and determine the optimal parameter combination of the sparse partial least squares discriminant analysis model. Based on the trained sparse partial least squares discriminant analysis model, the variable projection importance score of each variable is calculated; a variable projection importance score threshold is set, and combined with the statistically significant level controlled by the false discovery rate, variables with a variable projection importance score greater than the variable projection importance score threshold and a corrected p value less than the significant level are selected from the total number of A variables to form a key feature variable subset for constructing a pollutant characteristic fingerprint spectrum. The concentration or abundance data of each variable in the key feature variable subset is spatially associated with the corresponding sampling point geographic coordinates. Based on the associated data, the Moran index of the target farmland area is calculated to judge the overall aggregation of the spatial distribution of pollutants, and the local Moran index of each sampling point and its adjacent sampling points is calculated to identify the spatial aggregation type; wherein the spatial aggregation type includes high-high aggregation, low-low aggregation, high-low anomaly and low-high anomaly. The high-high aggregation area is defined as a spatial hotspot of pollutants; the concentration or abundance data of the key feature variable subset is used as the Z value, combined with the geographic coordinates, and the Kriging spatial interpolation method is used to generate a continuous distribution of three-dimensional pollution hotspot map.
9. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 8, characterized in that, When the three-dimensional pollution hotspot map is used to establish a pollutant value gradient database, the following steps are included: Based on the three-dimensional pollution hotspot map and the corresponding key feature variable subset, a relational database is constructed. The database includes the unique identifier of the sampling point, the geographic coordinates of the sampling point, the spatial hotspot partition identifier to which the sampling point belongs, and the concentration or abundance data of each variable in the key feature variable subset.
10. The feature fingerprint-based grain oil matrix reference material candidate raw material precision screening method according to claim 9, characterized in that, When the multi-objective optimization algorithm is used to iteratively screen the optimal solution set from the sampling points, the following steps are included: Maximizing the coverage of the pollutant concentration gradient range and minimizing the number of sampling points are set as two objectives that need to be optimized simultaneously. The multi-objective optimization algorithm is run in the pollutant value gradient database to iteratively screen the optimal solution set from all sampling points, i.e., a series of sampling schemes.
Citation Information
Patent Citations
A fingerprint database construction method and device for lake and reservoir water pollution tracing
CN109711674A
Tracing method and system for new pollutants in various environmental media based on machine learning
CN120387598A
Soybean producing area soil micro-ecological risk assessment method and system based on combined pollution
CN120706889A
Coastal wetland soil sample point quality evaluation method and system based on multi-source environmental data
CN120744705A
Soil detection system and detection method
CN120820697A