A multi-modal fusion complex matrix dish safety detection method and system
By employing a multimodal fusion detection method that combines global images, microscopic morphology, and diffuse reflectance spectra, and selecting detection strategies and performing dynamic detection, the problem of low accuracy in Chinese food safety detection has been solved, enabling efficient analysis of pollutant concentrations even under complex food composition conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BINZHOU POLYTECHNIC
- Filing Date
- 2026-02-06
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies are insufficient for effectively detecting the safety of Chinese cuisine and similar dishes, especially given the low accuracy of rapid testing due to the use of large amounts of raw materials, oils, and seasonings.
A multimodal fusion detection method is adopted, which combines global image, micro-morphology and diffuse reflectance spectrum. The detection strategy is selected by obtaining matrix feature vector, and dynamic detection is performed. High-frequency continuous sampling is performed using electrochemical and optical subsystems. The signal is separated and reconstructed by AI module, and the pollutant concentration is output.
It effectively solves the problem of low accuracy in rapid detection caused by the complexity of food ingredients, and can capture and eliminate non-ideal interference signals caused by food matrix, providing accurate pollutant concentration detection results.
Smart Images

Figure CN122200630A_ABST
Abstract
Description
Technical Field
[0001] This application relates to a multimodal fusion method and system for food safety testing of complex matrix products. Background Technology
[0002] Food safety, as a major issue concerning people's livelihood, has received increasing attention. Currently, it is mainly ensured from two aspects: one is the control of raw materials, mainly to avoid the use of various illegal drugs; the other is the monitoring of finished products after processing. For industrialized food products, due to the uniformity of management, food safety can be controlled through random sampling. However, for Chinese cuisine and other types of catering, due to the wide variety of raw materials, oils, seasonings, etc. used, there is a need to develop more effective testing methods for the safety of this type of food. Summary of the Invention
[0003] To address the aforementioned issues, this application proposes a multimodal fusion method for the safety detection of complex matrix foods, comprising the following steps: The system acquires the global image, microstructure, and diffuse reflectance spectrum of the food, and analyzes the matrix feature vector V. matrix ; Based on the matrix feature vector V matrix Select a detection strategy; The original reaction kinetic curves were obtained by performing kinetic detection according to the detection strategy. For the matrix feature vector V matrix The original reaction kinetic curves are analyzed to obtain the analytical results. This application selects a detection strategy based on global images, microscopic morphology, and diffuse reflectance spectra, and then performs kinetic detection to obtain the detection results, effectively solving the problem of low accuracy in rapid detection caused by the complexity of food components.
[0004] Preferably, the global image is obtained in the following manner: The container holding the food is placed in a dark box. A top diffused lighting circuit is set up in the dark box to create a uniform white light environment to eliminate local reflections. A macroscopic RGB texture image of the food is acquired using a CMOS sensor. The microstructure was obtained in the following manner: Switch the top diffuse illumination circuit to a low-angle side grazing mode, so that the light is incident at an angle close to the sample surface, and use the "shadow enhancement principle" to highlight the micro-undulations and liquid film features of the surface. The diffuse reflectance spectrum was obtained as follows: The 45° / 0° optical geometric scanning unit is activated, and the full-color light source is incident at a 45-degree angle. The MEMS micro grating spectrometer receives the diffuse reflection light signal in the vertical direction at a 0-degree angle. This optical path design effectively filters out specular reflection light generated by liquids or smooth surfaces through the principle of geometric optics, ensuring that the collected light is a diffuse reflection spectrum of 380nm-780nm carrying internal chemical information.
[0005] Preferably, the biological processing attribute category of the dish is determined by the morphological features and color distribution features of the whole-domain image. ID The biological processing attribute categories include pre-cooked food categories, vegetable and fruit categories, and fungi categories. Based on microscopic morphology and biological processing attributes, the following methods were used: For pre-prepared dishes, a connected component analysis algorithm was applied to statistically analyze the distribution density of highly reflective areas and quantify the oil film coverage index. index ; For fruits and vegetables and fungi, a texture roughness analysis algorithm is applied to identify soil-adhered particles or wrinkling features caused by oxidative browning, quantifying the surface cleanliness / freshness index, Browning. level ; Diffuse reflectance spectra were processed based on biological processing properties as follows: For prepared vegetables, the 400nm-420nm wavelength band was calculated to quantify the intensity of the brown background interference; for fruits and vegetables, the absorbance in the 640nm-660nm wavelength band was calculated by integration to quantify the chlorophyll density. density For fungal species, detect abnormal reflection peaks around 380nm to rule out fluorescent whitening agents. The V matrix =[Category ID Oil index Chlorophyll density Browning level ,……].
[0006] Preferably, based on the matrix feature vector V matrix The detection strategy should be selected as follows: Set up a multi-dimensional rule base; According to V matrix The detection strategy is obtained by matching a multi-dimensional rule base. The multi-dimensional rule base includes channel arbitration strategies and pre-processing interaction logic strategies: The channel arbitration strategy includes optical / electrochemical channel switching and adaptive correction of enrichment parameters: For high-pigment samples (such as dark vegetables), the optical / electrochemical channel switching logic kernel forcibly disables the optical colorimetric channel that is interfered with by pigments, and instead activates the electrochemical detection channel, and loads the enzyme electrode parameter configuration for pesticide residues, thereby avoiding absorbance background noise from a physical principle level; For samples that need to detect heavy metals, in order to cope with the dissolution requirements of low-concentration ions, the system automatically overwrites the control script of the electrochemical instrument; specifically, it lowers the enrichment voltage to a more sensitive -1.2V and dynamically extends the enrichment time from the default 60s to 120s to improve the signal-to-noise ratio; The preprocessing interaction logic strategy includes abnormal matrix response and human-machine collaborative instructions: the abnormal matrix response is for some extreme matrices that exceed the sensor's direct injection tolerance (such as V). matrix (For pre-cooked dishes displaying extremely high oil film index), the system determines that relying solely on algorithmic noise reduction may fail; the human-machine collaborative instruction refers to triggering the "human-machine intervention" mechanism to generate structured UI interaction instructions; this instruction is not a simple error report, but a specific pre-processing operation guide (e.g., "Heavy oil matrix detected, it is recommended to add ethyl acetate for shaking extraction, let stand, and then aspirate the lower clear liquid for injection"), which is equivalent to dynamically inserting a virtual technical guidance link into the detection process.
[0007] Preferably, the original reaction kinetic curve is obtained by performing kinetic detection according to the detection strategy in the following manner: Construct a dual-mode detection circuit: It includes an electrochemical subsystem and an optical subsystem. The electrochemical subsystem is based on a closed-loop potentiostat architecture, uses a 16-bit high-precision DAC to generate a scanning voltage, and drives a three-electrode system through an operational amplifier. The receiving end uses a low-noise transimpedance amplifier (TIA) to linearly convert the weak reaction current into a voltage signal, which is then digitized by a 24-bit ADC. The optical subsystem constructs a miniature colorimetric cavity and integrates a constant current source-driven specific wavelength narrowband LED (such as 412nm / 520nm) and a high-sensitivity silicon photodiode for real-time monitoring of the transmittance changes of the colorimetric reaction. Dynamic reaction activation: Instruction parsing: The microprocessor (MCU) receives the "control instruction packet" from stage two and parses out the specific detection mode and driving parameters (such as scan voltage range and illumination duration); Execution logic: Electrochemical mode: The MCU controls the DAC to output a dynamically changing scan waveform (such as the step voltage in square wave voltammetry), establishing a non-constant potential environment on the working electrode surface to induce the target analyte (such as heavy metal ions) to undergo a redox reaction; Optical mode: The MCU controls the constant current source to light up the LED, establishing a stable photometric measurement environment and initiating the enzyme inhibition reaction or chemical color development process; High-frequency continuous sampling of the entire process kinetic signal: Sampling mechanism: Unlike the "endpoint method" of traditional instruments that read a single value after the reaction ends (e.g., at the 3rd minute), this system activates a high-frequency timer to synchronously trigger data reading at a frequency of 10Hz (i.e., 0.1 seconds / time); Data capture: Time-series recording: The system fully records the current response (electrochemical) from the initial potential scan to the termination potential, or the absorbance change (optical) at every moment from the start to the end of the colorimetric reaction; Feature retention: This "cinematic" recording method not only retains the signal peaks or slopes representing the actual contaminant concentration, but also captures the non-ideal interference signals caused by the food matrix intact—for example, random sharp burrs generated by plant fibers colliding with the electrode, or baseline drift and fluctuations caused by pigments / oil films.
[0008] Input the original reaction kinetic curve C raw : C raw ={(t0,v0),(t1,v1),……,(t n ,v n )} t n Represents a relative time point or sequence index for high-frequency sampling; v n Represents t n The physical response values collected at any given time are as follows: In electrochemical mode: representing the current response value, that is, the weak current flowing between the working electrode and the counter electrode when a specific scan voltage is applied; In optical mode: represents absorbance or transmittance, which is the value of the light intensity received by the photodiode after photoelectric conversion and logarithmic processing; This includes n consecutive sampling points (e.g., 1800 data points generated by a 180-second reaction). The original reaction kinetic curves obtained in this application, which mix real chemical signals with complex matrix noise, constitute all the original materials for the subsequent AI module to perform "signal separation and reconstruction".
[0009] Preferably, for the matrix feature vector V matrix The original reaction kinetic curves were analyzed and processed to obtain the analytical results as follows: Dual-stream coding and spatial mapping of heterogeneous data: Dimension alignment: due to the input V matrix It is low-dimensional discrete data, while the original dynamic curve C raw Since it is high-dimensional time-series data, the two cannot be directly computed; the system first maps them to a latent feature space of the same dimension through two independent encoder branches; Signal stream coding: Utilizing a one-dimensional convolutional neural network (1D-CNN) as a feature extractor, sliding and scanning C along the time axis. raw Capture the local shape of the waveform and generate the signal feature sequence H. signal ; Matrix Flow Coding: Using Multilayer Perceptron (MLP) to encode V matrix Perform a dimensionality-up mapping to transform it into a global context vector H that implicitly contains specific matrix noise patterns. context ; Noise locking and suppression based on cross-attention mechanism: Query construction: The system constructs the matrix context vector H context Convert to a query vector (Query, Q); Key-value matching: The system matches the signal feature sequence H signal Convert Q into a key vector (Key, K) and a value vector (Value, V); the algorithm calculates the dot product of Q and K to generate the attention score matrix; Inverse suppression logic: In the scoring matrix, waveform segments with extremely high matching degrees to the visual prior Q are assigned high weights; the model then performs weighted fusion and residual connections accordingly, specifically suppressing these high-scoring interference terms, preserving and enhancing the features of the unmatched real chemical signals, and generating the fused feature H. fused ; Signal reconstruction and quantitative inversion: Temporal decoding: Utilizing a one-dimensional deconvolutional decoder to fuse features H fused Restore to the temporal space; this step outputs the reconstructed pure curve C. clean — Its baseline drift was corrected to zero, random spikes were smoothed, and only specific peaks related to pollutant concentration were retained; Numerical Regression: The parallel regression head directly reads the purity features, performs linear regression calculations through a fully connected layer, and outputs the final predicted pollutant concentration value Y. pred This process is subject to a dual loss function and a morphological constraint L. MSE+ Precision Constraint L MAE The supervision ensured the synchronous accuracy of waveform reconstruction and numerical calculation; Summary output: Quantitative results: Precise pollutant concentration values Y pred ; Qualitative evidence: A visual pure reaction curve C clean This is used to compare the curve with the original curve in the next stage, demonstrating the reliability of the detection to the user.
[0010] Preferred, C raw The mathematical form is a one-dimensional tensor of length L; the signal feature sequence Hsignal Do it in the following way: A one-dimensional convolutional neural network (1D-CNN) is used as a feature extractor; it contains multiple layers of convolutional kernels that slide along the time axis to capture the rising slope, peak curvature, and high-frequency glitches of the waveform; the weight parameters of the convolutional layers are set to W. CNN The bias is b cnn After convolution and the ReLU activation function, the original time-series data is transformed into a feature sequence H. signal : ; Where L' is the length of the downsampled sequence, which represents the time segment of the waveform, and d is the feature embedding dimension; The global context vector H context Do it in the following way: The matrix vector is obtained through linear transformation and nonlinear activation. ; Here, d is consistent with the dimension of the signal coding branch.
[0011] Preferably, the noise locking and suppression based on the cross-attention mechanism is performed in the following manner: Using three learnable linear projection matrices, W Q W K W V The generated encoded features are converted into query vectors (Q), key vectors (K), and value vectors (V) required by the attention mechanism. Generation of query vector Q: using matrix context vector H context Generate Q; ; Q represents a judgment based on the current visual perception; Generation of key vector K and value vector V: using signal feature sequence H signal Generate K and V; ; K is the index label of each local waveform segment in the dynamic curve, and V is the information content actually carried by these segments; Calculate the dot product of Q and K, and scale it to obtain the attention score matrix A; ; Softmax is a normalization exponential function used to transform an input numerical vector into a probability distribution vector. K TThis refers to the transpose of K; d k This refers to the feature dimension of K; A full-domain scanning process is performed, calculating the similarity between the "visual prior" and "each waveform segment." If the visual module identifies a sample with extremely high oil content, then the Q vector will contain information about the "oil film noise pattern." At this point, any segment in the curve that matches the "random spike" characteristic, i.e., noise caused by oil, will have a very high matching degree between its corresponding K value and Q, thus obtaining a high score in matrix A. Conversely, the actual pollutant signal waveform will have a lower score due to feature mismatch. This step achieves accurate localization of interference signals. A high score means that in the attention scoring matrix A, the value at that position is significantly higher than the uniform distribution probability value in the same dimension or significantly higher than the weight of other non-interfering segments; this means that the features of the identified waveform segment are highly similar to Q. A low score refers to a value in the attention scoring matrix A that is close to zero or significantly lower than the weight of the main interference item; this means that the characteristics of the waveform segment do not match the matrix noise pattern and are judged as real chemical signals or non-specific background noise. The value vector V is weighted using the scoring matrix A, and residual connections are introduced to generate the fusion feature H. fused ; In this step, the model employs a reverse suppression strategy: features with high attention scores are weighted or filtered out in subsequent feature reorganization; features with low attention scores are retained and enhanced; thus completing the qualitative change from "noisy features" to "clean features". It also includes the processes of signal reconstruction and concentration inversion: Using a one-dimensional deconvolutional neural network, Transposed Conv1D, as a decoder, the fused features H are... fused The data is restored from the feature space back to the temporal space; the decoder gradually recovers the length of the data through upsampling operations, and finally outputs a reconstructed, clean response curve C. clean ; Compared to the original input curve, this reconstructed curve has significant features: baseline drift is corrected to zero, random spikes caused by plant fibers are smoothed, background absorbance caused by pigments is subtracted, and only specific peaks directly related to pollutant concentration are retained. To directly output the detection results, a regression head is connected in parallel; it consists of a global average pooling layer and a fully connected layer, directly reading the peak height and peak area characteristics of the purity curve, and outputting the final pollutant concentration value Y through linear regression calculation. pred .
[0012] Preferably, the system's built-in memory contains a database of food safety testing standards, including maximum residue limits (T) for various pesticide residues, heavy metals, and illegal additives. limit When the final pollutant concentration value Y is output... pred Then, numerical comparisons are performed: If Y pred <T limit The system determined the test result to be "qualified"; If Y pred ≥T limit The system determines the test result as "unqualified" and automatically calculates the excess multiple.
[0013] On the other hand, a multimodal fusion-based food safety detection system based on complex matrices is also disclosed, comprising the following modules: The data acquisition module is used to acquire the global image, microstructure, and diffuse reflectance spectrum of the dish, and analyze it to obtain the matrix feature vector V. matrix ; The selection module is used to select based on the matrix feature vector V. matrix Select a detection strategy; The kinetic detection module is used to perform kinetic detection according to the detection strategy to obtain the original reaction kinetic curve; Analysis module, used for analyzing matrix feature vector V matrix The original reaction kinetic curves were analyzed to obtain the analytical results.
[0014] This application can bring the following beneficial effects: 1. This application selects a detection strategy based on global image, micro-morphology and diffuse reflectance spectrum, and then performs dynamic detection to obtain the detection result, which effectively solves the problem of low accuracy of rapid detection caused by the complexity of food ingredients.
[0015] 2. This application can capture non-ideal interference signals caused by the food matrix intact—for example, random sharp burrs generated by plant fibers colliding with the electrode, or baseline drift and fluctuations caused by pigments / oil films.
[0016] 3. The original reaction kinetic curves obtained in this application are a mixture of real chemical signals and complex matrix noise, which constitute all the original materials for the subsequent AI module to perform "signal separation and reconstruction". Attached Figure Description
[0017] The accompanying drawings, which are provided to further understand this application and constitute a part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.
[0018] In the attached diagram: Figure 1 This is a flowchart illustrating the process of this application.
[0019] Figure 2 This is a schematic diagram of the modules in this application.
[0020] Figure 3 This is a schematic diagram of a dual-mode detection circuit. Detailed Implementation
[0021] To clearly illustrate the technical features of this solution, the following detailed description, in conjunction with the accompanying drawings, will be provided.
[0022] like Figure 1 As shown, a multimodal fusion method for food safety testing of complex matrix products includes the following steps: S1. Obtain the global image, microstructure, and diffuse reflectance spectrum of the dish, and analyze to obtain the matrix feature vector V. matrix ; The global image is obtained in the following manner: The container holding the food is placed in a dark box. A top diffused lighting circuit is set up in the dark box to create a uniform white light environment to eliminate local reflections. A macroscopic RGB texture image of the food is acquired using a CMOS sensor. The microstructure was obtained in the following manner: Switch the top diffuse illumination circuit to a low-angle side grazing mode, so that the light is incident at an angle close to the sample surface, and use the "shadow enhancement principle" to highlight the micro-undulations and liquid film features of the surface. The diffuse reflectance spectrum was obtained as follows: The 45° / 0° optical geometric scanning unit is activated, and the full-color light source is incident at a 45-degree angle. The MEMS micro grating spectrometer receives the diffuse reflection light signal in the vertical direction at a 0-degree angle. This optical path design effectively filters out specular reflection light generated by liquids or smooth surfaces through the principle of geometric optics, ensuring that the collected light is a diffuse reflection spectrum of 380nm-780nm carrying internal chemical information.
[0023] By analyzing the morphological and color distribution features of the global image, the biological processing attribute category of the dish is determined. ID The biological processing attribute categories include pre-cooked food categories, vegetable and fruit categories, and fungi categories. Based on microscopic morphology and biological processing attributes, the following methods were used: For pre-prepared dishes, a connected component analysis algorithm was applied to statistically analyze the distribution density of highly reflective areas and quantify the oil film coverage index. indexFor fruits and vegetables and fungi, a texture roughness analysis algorithm is applied to identify soil-attached particles or wrinkling features caused by oxidative browning, quantifying the surface cleanliness / freshness index, Browning. level ; Diffuse reflectance spectra were processed based on biological processing properties as follows: For prepared vegetables, the 400nm-420nm wavelength band was calculated to quantify the intensity of the brown background interference; for fruits and vegetables, the absorbance in the 640nm-660nm wavelength band was calculated by integration to quantify the chlorophyll density. density For fungal species, detect abnormal reflection peaks around 380nm to rule out fluorescent whitening agents. The V matrix =[Category ID Oil index Chlorophyll density Browning level ,……].
[0024] S2. Based on the matrix eigenvector V matrix Select a detection strategy; Based on the matrix feature vector V matrix The detection strategy should be selected as follows: Set up a multi-dimensional rule base; According to V matrix The detection strategy is obtained by matching a multi-dimensional rule base. The multi-dimensional rule base includes channel arbitration strategies and pre-processing interaction logic strategies: The channel arbitration strategy includes optical / electrochemical channel switching and adaptive correction of enrichment parameters: For high-pigment samples (such as dark vegetables), the optical / electrochemical channel switching logic kernel forcibly disables the optical colorimetric channel that is interfered with by pigments, and instead activates the electrochemical detection channel, and loads the enzyme electrode parameter configuration for pesticide residues, thereby avoiding absorbance background noise from a physical principle level; For samples that need to detect heavy metals, in order to cope with the dissolution requirements of low-concentration ions, the system automatically overwrites the control script of the electrochemical instrument; specifically, it lowers the enrichment voltage to a more sensitive -1.2V and dynamically extends the enrichment time from the default 60s to 120s to improve the signal-to-noise ratio; The preprocessing interaction logic strategy includes abnormal matrix response and human-machine collaborative instructions: the abnormal matrix response is for some extreme matrices that exceed the sensor's direct injection tolerance (such as V). matrix(For pre-cooked dishes displaying extremely high oil film index), the system determines that relying solely on algorithmic noise reduction may fail; the human-machine collaborative instruction refers to triggering the "human-machine intervention" mechanism to generate structured UI interaction instructions; this instruction is not a simple error report, but a specific pre-processing operation guide (e.g., "Heavy oil matrix detected, it is recommended to add ethyl acetate for shaking extraction, let stand, and then aspirate the lower clear liquid for injection"), which is equivalent to dynamically inserting a virtual technical guidance link into the detection process.
[0025] S3. Obtain the original reaction kinetic curve by performing kinetic detection according to the detection strategy; The original reaction kinetic curves were obtained by performing kinetic detection according to the detection strategy as follows: like Figure 3 As shown, a dual-mode detection circuit is constructed: It includes an electrochemical subsystem and an optical subsystem. The electrochemical subsystem is based on a closed-loop potentiostat architecture, uses a 16-bit high-precision DAC to generate a scanning voltage, and drives a three-electrode system through an operational amplifier. The receiving end uses a low-noise transimpedance amplifier (TIA) to linearly convert the weak reaction current into a voltage signal, which is then digitized by a 24-bit ADC. The optical subsystem constructs a miniature colorimetric cavity and integrates a constant current source-driven specific wavelength narrowband LED (such as 412nm / 520nm) and a high-sensitivity silicon photodiode for real-time monitoring of the transmittance changes of the colorimetric reaction. Dynamic reaction activation: Instruction parsing: The microprocessor (MCU) receives the "control instruction packet" from stage two and parses out the specific detection mode and driving parameters (such as scan voltage range and illumination duration); Execution logic: Electrochemical mode: The MCU controls the DAC to output a dynamically changing scan waveform (such as the step voltage in square wave voltammetry), establishing a non-constant potential environment on the working electrode surface to induce the target analyte (such as heavy metal ions) to undergo a redox reaction; Optical mode: The MCU controls the constant current source to light up the LED, establishing a stable photometric measurement environment and initiating the enzyme inhibition reaction or chemical color development process; High-frequency continuous sampling of the entire process kinetic signal: Sampling mechanism: Unlike the "endpoint method" of traditional instruments that read a single value after the reaction ends (e.g., at the 3rd minute), this system activates a high-frequency timer to synchronously trigger data reading at a frequency of 10Hz (i.e., 0.1 seconds / time); Data capture: Time-series recording: The system fully records the current response (electrochemical) from the initial potential scan to the termination potential, or the absorbance change (optical) at every moment from the start to the end of the colorimetric reaction; Feature retention: This "cinematic" recording method not only retains the signal peaks or slopes representing the actual contaminant concentration, but also captures the non-ideal interference signals caused by the food matrix intact—for example, random sharp burrs generated by plant fibers colliding with the electrode, or baseline drift and fluctuations caused by pigments / oil films.
[0026] Input the original reaction kinetic curve C raw : C raw ={(t0,v0),(t1,v1),……,(t n ,v n )} t n Represents the relative time point or sequence index of high-frequency sampling; Acquisition method: Recorded by a high-frequency timer built into the microprocessor at a frequency of 10Hz.
[0027] v n Represents t n The physical response values collected at any given time are as follows: In electrochemical mode: representing the current response value (Current, A), which is the weak current flowing between the working electrode and the counter electrode when a specific scan voltage is applied; In optical mode: it represents the absorbance (OD) or transmittance, which is the value of the light intensity received by the photodiode after photoelectric conversion and logarithmic processing; This includes n consecutive sampling points (e.g., a 180-second reaction generates 1800 data points).
[0028] S4 for matrix feature vector V matrix The original reaction kinetic curves were analyzed to obtain the analytical results.
[0029] For the matrix feature vector V matrix The original reaction kinetic curves were analyzed and processed to obtain the analytical results as follows: Dual-stream coding and spatial mapping of heterogeneous data: Dimension alignment: due to the input V matrix It is low-dimensional discrete data, while the original dynamic curve C raw Since it is high-dimensional time-series data, the two cannot be directly computed; the system first maps them to a latent feature space of the same dimension through two independent encoder branches. Signal stream coding: Utilizing a one-dimensional convolutional neural network (1D-CNN) as a feature extractor, sliding and scanning C along the time axis. raw Capture the local shape of the waveform and generate the signal feature sequence H. signal ; Matrix Flow Coding: Using Multilayer Perceptron (MLP) to encode V matrix Perform a dimensionality-up mapping to transform it into a global context vector H that implicitly contains specific matrix noise patterns. context ; Noise locking and suppression based on cross-attention mechanism: Query construction: The system constructs the matrix context vector H context Convert to a query vector (Query, Q); Key-value matching: The system matches the signal feature sequence H signal Convert Q into a key vector (Key, K) and a value vector (Value, V); the algorithm calculates the dot product of Q and K to generate the attention score matrix; Inverse suppression logic: In the scoring matrix, waveform segments with extremely high matching degrees to the visual prior Q are assigned high weights; the model then performs weighted fusion and residual connections accordingly, specifically suppressing these high-scoring interference terms, preserving and enhancing the features of the unmatched real chemical signals, and generating the fused feature H. fused ; Signal reconstruction and quantitative inversion: Temporal decoding: Utilizing a one-dimensional deconvolutional decoder to fuse features H fused Restore to the temporal space; this step outputs the reconstructed pure curve C. clean — Its baseline drift was corrected to zero, random spikes were smoothed, and only specific peaks related to pollutant concentration were retained; Numerical Regression: The parallel regression head directly reads the purity features, performs linear regression calculations through a fully connected layer, and outputs the final predicted pollutant concentration value Y. pred This process is subject to a dual loss function and a morphological constraint L. MSE+ Precision Constraint L MAE The supervision ensured the synchronous accuracy of waveform reconstruction and numerical calculation; Summary output: Quantitative results: Precise pollutant concentration values Y pred ; Qualitative evidence: A visual pure reaction curve C clean This is used to compare the curve with the original curve in the next stage, demonstrating the reliability of the detection to the user.
[0030] C raw The mathematical form is a one-dimensional tensor of length L; the signal feature sequence H signal Do it in the following way: A one-dimensional convolutional neural network (1D-CNN) is used as a feature extractor; it contains multiple layers of convolutional kernels that slide along the time axis to capture the rising slope, peak curvature, and high-frequency glitches of the waveform; the weight parameters of the convolutional layers are set to W. CNN The bias is b cnn After convolution and the ReLU activation function, the original time-series data is transformed into a feature sequence H. signal : ; Where L' is the length of the downsampled sequence, which represents the time segment of the waveform, and d is the feature embedding dimension; The global context vector H context Do it in the following way: The matrix vector is obtained through linear transformation and nonlinear activation. ; Here, d is consistent with the dimension of the signal coding branch.
[0031] The noise locking and suppression based on the cross-attention mechanism is performed in the following manner: Using three learnable linear projection matrices, W Q W K W V The generated encoded features are converted into query vectors (Q), key vectors (K), and value vectors (V) required by the attention mechanism. Generation of query vector Q: using matrix context vector H context Generate Q; ; Q represents a judgment based on the current visual perception; Generation of key vector K and value vector V: using signal feature sequence H signal Generate K and V; ; K is the index label of each local waveform segment in the dynamic curve, and V is the information content actually carried by these segments; Calculate the dot product of Q and K, and scale it to obtain the attention score matrix A; ; Softmax is a normalization exponential function used to transform an input numerical vector into a probability distribution vector (all values are between 0 and 1, and their sum is 1). In this application, its physical meaning is to calculate the "relevance weight," which determines the extent to which the AI model should focus on a particular waveform segment.
[0032] K T This refers to the transpose of K; essentially, it calculates the similarity (dot product) between Q and K.
[0033] d kThis refers to the feature dimension of K; it is usually a constant (such as 64 or 128) to prevent the dot product result from becoming too large, causing the Softmax function to enter the "saturation region" (gradient vanishing), thus ensuring the numerical stability of model training.
[0034] A full-domain scanning process is performed, calculating the similarity between the "visual prior" and "each waveform segment." If the visual module identifies a sample with extremely high oil content, then the Q vector will contain information about the "oil film noise pattern." At this point, any segment in the curve that matches the "random spike" characteristic, i.e., noise caused by oil, will have a very high matching degree between its corresponding K value and Q, thus obtaining a high score in matrix A. Conversely, the actual pollutant signal waveform will have a lower score due to feature mismatch. This step achieves accurate localization of interference signals. A high score means that in the attention scoring matrix A, the value at that position is significantly higher than the uniform distribution probability value in the same dimension or significantly higher than the weight of other non-interfering segments; this means that the features of the identified waveform segment are highly similar to Q. A low score refers to a value in the attention scoring matrix A that is close to zero or significantly lower than the weight of the main interference item; this means that the characteristics of the waveform segment do not match the matrix noise pattern and are judged as real chemical signals or non-specific background noise. The value vector V is weighted using the scoring matrix A, and residual connections are introduced to generate the fusion feature H. fused ; In this step, the model employs a reverse suppression strategy: features with high attention scores are weighted or filtered out in subsequent feature reorganization; features with low attention scores are retained and enhanced; thus completing the qualitative change from "noisy features" to "clean features". It also includes the processes of signal reconstruction and concentration inversion: Using a one-dimensional deconvolutional neural network, Transposed Conv1D, as a decoder, the fused features H are... fused The data is restored from the feature space back to the temporal space; the decoder gradually recovers the length of the data through upsampling operations, and finally outputs a reconstructed, clean response curve C. clean ; Compared to the original input curve, this reconstructed curve has significant features: baseline drift is corrected to zero, random spikes caused by plant fibers are smoothed, background absorbance caused by pigments is subtracted, and only specific peaks directly related to pollutant concentration are retained. To directly output the detection results, a regression head is connected in parallel; it consists of a global average pooling layer and a fully connected layer, directly reading the peak height and peak area characteristics of the purity curve, and outputting the final pollutant concentration value Y through linear regression calculation.pred .
[0035] The system's built-in memory contains a database of food safety testing standards, including maximum residue limits (T) for various pesticide residues, heavy metals, and illegal additives. limit When the final pollutant concentration value Y is output... pred Then, numerical comparisons are performed: If Y pred <T limit The system determined the test result to be "qualified"; If Y pred ≥T limit The system determines the test result as "unqualified" and automatically calculates the excess multiple.
[0036] In the second embodiment, as Figure 2 As shown, a multimodal fusion-based food safety detection system for complex matrices includes the following modules: The data acquisition module 201 is used to acquire the global image, microstructure, and diffuse reflectance spectrum of the dish, and analyze it to obtain the matrix feature vector V. matrix ; Selection module 202 is used to select based on the matrix feature vector V matrix Select a detection strategy; The kinetic detection module 203 is used to perform kinetic detection according to the detection strategy to obtain the original reaction kinetic curve; Analysis module 204 is used for analyzing the matrix feature vector V matrix The original reaction kinetic curves were analyzed to obtain the analytical results.
[0037] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A multimodal fusion method for food safety testing in complex matrices, characterized in that: Includes the following steps: The system acquires the global image, microstructure, and diffuse reflectance spectrum of the food, and analyzes the matrix feature vector V. matrix ; Based on the matrix feature vector V matrix Select a detection strategy; The original reaction kinetic curves were obtained by performing kinetic detection according to the detection strategy. For the matrix feature vector V matrix The original reaction kinetic curves were analyzed to obtain the analytical results.
2. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 1, characterized in that: The global image is obtained in the following manner: The container holding the food is placed in a dark box. A top diffused lighting circuit is set up in the dark box to create a uniform white light environment to eliminate local reflections. A macroscopic RGB texture image of the food is acquired using a CMOS sensor. The microstructure was obtained in the following manner: Switch the top diffuse illumination circuit to a low-angle side grazing mode, where the light is incident at an angle close to the sample surface, using the "shadow enhancement principle" to highlight the surface's micro-undulations and liquid film features; The diffuse reflectance spectrum was obtained as follows: The 45° / 0° optical geometric scanning unit is activated, and the full-color light source is incident at a 45-degree angle. The MEMS micro grating spectrometer receives the diffuse reflection light signal in the vertical direction at a 0-degree angle. This optical path design effectively filters out specular reflection light generated by liquids or smooth surfaces through the principle of geometric optics, ensuring that the collected light is a diffuse reflection spectrum of 380nm-780nm carrying internal chemical information.
3. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 2, characterized in that: By analyzing the morphological and color distribution features of the global image, the biological processing attribute category of the dish is determined. ID The biological processing attribute categories include pre-cooked food categories, vegetable and fruit categories, and fungi categories. Based on microscopic morphology and biological processing attributes, the following methods were used: For pre-prepared dishes, a connected component analysis algorithm was applied to statistically analyze the distribution density of highly reflective areas and quantify the oil film coverage index. index ; For fruits and vegetables and fungi, a texture roughness analysis algorithm is applied to identify soil-adhered particles or wrinkling features caused by oxidative browning, quantifying the surface cleanliness / freshness index, Browning. level ; Diffuse reflectance spectra were processed based on biological processing properties as follows: For prepared vegetables, the 400nm-420nm wavelength band was calculated to quantify the intensity of the brown background interference; for fruits and vegetables, the absorbance in the 640nm-660nm wavelength band was calculated by integration to quantify the chlorophyll density. density For fungal species, detect abnormal reflection peaks around 380nm to rule out fluorescent whitening agents. The said V matrix =[Category ID , Oil index , Chlorophyll density , Browning level 。 4. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 1, characterized in that: Based on the matrix feature vector V matrix The detection strategy should be selected as follows: Set up a multi-dimensional rule base; According to V matrix The detection strategy is obtained by matching a multi-dimensional rule base. The multi-dimensional rule base includes channel arbitration strategies and pre-processing interaction logic strategies: The channel arbitration strategy includes optical / electrochemical channel switching and adaptive correction of enrichment parameters: For samples with high pigment content, the optical / electrochemical channel switching logic kernel forcibly disables the optical colorimetric channel that is interfered with by pigments, and instead activates the electrochemical detection channel, and loads the enzyme electrode parameter configuration for pesticide residues, thereby avoiding absorbance background noise from a physical principle perspective; For samples that need to detect heavy metals, in order to cope with the dissolution requirements of low concentration ions, the system automatically overwrites the control script of the electrochemical instrument; specifically, the enrichment voltage is lowered to a more sensitive -1.2V, and the enrichment time is dynamically extended from the default 60s to 120s to improve the signal-to-noise ratio; The preprocessing interaction logic strategy includes abnormal matrix response and human-machine collaboration instructions: the abnormal matrix response is for some extreme matrices that exceed the sensor's tolerance for direct injection, and the system determines that relying solely on algorithm denoising may fail; the human-machine collaboration instructions refer to triggering the "human-machine intervention" mechanism to generate structured UI interaction instructions.
5. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 1, characterized in that: The original reaction kinetic curves were obtained by performing kinetic detection according to the detection strategy as follows: Construct a dual-mode detection circuit: It includes an electrochemical subsystem and an optical subsystem. The electrochemical subsystem is based on a closed-loop potentiostat architecture, uses a 16-bit high-precision DAC to generate a scanning voltage, and drives a three-electrode system through an operational amplifier. The receiving end uses a low-noise transimpedance amplifier to linearly convert the weak reaction current into a voltage signal, which is then digitized by a 24-bit ADC. The optical subsystem constructs a miniature colorimetric cavity and integrates a specific wavelength narrowband LED driven by a constant current source and a high-sensitivity silicon photodiode for real-time monitoring of the transmittance changes of the colorimetric reaction. Dynamic reaction activation: Instruction parsing: The microprocessor (MCU) receives the "control instruction packet" from stage two and parses out the specific detection mode and driving parameters; Execution logic: Electrochemical mode: The MCU controls the DAC to output a dynamically changing scanning waveform, establishing a non-constant potential environment on the working electrode surface to induce a redox reaction in the target analyte; Optical mode: The MCU controls the constant current source to light up the LED, establishing a stable photometric measurement environment and initiating an enzyme inhibition reaction or a chemical color development process; High-frequency continuous sampling of dynamic signals throughout the entire process: Sampling mechanism: Data reading is synchronously triggered at a frequency of 10Hz; Data capture: Time-series recording: The system fully records the current response from the initial potential scan to the final potential, or the absorbance change at every moment from the start to the end of the color reaction; Feature retention: This "cinematic" recording method not only retains the signal peaks or slopes representing the true concentration of pollutants, but also captures non-ideal interference signals caused by the food matrix; Input the original reaction kinetic curve C raw : C raw ={(t0,v0),(t1,v1),……,(t n ,v n )} t n Represents a relative time point or sequence index for high-frequency sampling; v n Represents t n The physical response values collected at any given time are as follows: In electrochemical mode: representing the current response value, that is, the weak current flowing between the working electrode and the counter electrode when a specific scan voltage is applied; In optical mode: represents absorbance or transmittance, which is the value of the light intensity received by the photodiode after photoelectric conversion and logarithmic processing; It contains n consecutive sampling points.
6. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 1, characterized in that: For the matrix feature vector V matrix The original reaction kinetic curves were analyzed and processed to obtain the analytical results as follows: Dual-stream coding and spatial mapping of heterogeneous data: Dimension alignment: due to the input V matrix It is low-dimensional discrete data, while the original dynamic curve C raw Since it is high-dimensional time-series data, the two cannot be directly computed; the system first maps them to a latent feature space of the same dimension through two independent encoder branches. Signal stream coding: Utilizing a one-dimensional convolutional neural network (1D-CNN) as a feature extractor, sliding and scanning C along the time axis. raw Capture the local shape of the waveform and generate the signal feature sequence H. signal ; Matrix Flow Coding: Using Multilayer Perceptron (MLP) to encode V matrix Perform a dimensionality-up mapping to transform it into a global context vector H that implicitly contains specific matrix noise patterns. context ; Noise locking and suppression based on cross-attention mechanism: Query construction: The system constructs the matrix context vector H context Convert to a query vector (Query, Q); Key-value matching: The system matches the signal feature sequence H signal Convert to a key vector (Key, K) and a value vector (Value, V); calculate the dot product of Q and K to generate the attention score matrix; Inverse inhibition logic: In the scoring matrix, waveform segments that have a very high degree of matching with the visual prior Q are given high weights; Based on this, the model performs weighted fusion and residual connection, specifically suppressing these high-scoring interference terms, preserving and enhancing the features of mismatched real chemical signals, and generating fusion features H. fused ; Signal reconstruction and quantitative inversion: Temporal decoding: Utilizing a one-dimensional deconvolutional decoder to fuse features H fused Restore to the temporal space; this step outputs the reconstructed pure curve C. clean — Its baseline drift was corrected to zero, random spikes were smoothed, and only specific peaks related to pollutant concentration were retained; Numerical Regression: The parallel regression head directly reads the purity features, performs linear regression calculations through a fully connected layer, and outputs the final predicted pollutant concentration value Y. pred This process is subject to a dual loss function and a morphological constraint L. MSE+ Precision Constraint L MAE The supervision ensured the synchronous accuracy of waveform reconstruction and numerical calculation; Summary output: Quantitative results: Precise pollutant concentration values Y pred ; Qualitative evidence: A visual pure reaction curve C clean This is used to compare the curve with the original curve in the next stage, demonstrating the reliability of the detection to the user.
7. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 6, characterized in that: C raw The mathematical form is a one-dimensional tensor of length L; the signal feature sequence H signal Do it in the following way: A one-dimensional convolutional neural network (1D-CNN) is used as a feature extractor; it contains multiple layers of convolutional kernels that slide along the time axis to capture the rising slope, peak curvature, and high-frequency glitches of the waveform; the weight parameters of the convolutional layers are set to W. CNN The bias is b cnn After convolution and the ReLU activation function, the original time-series data is transformed into a feature sequence H. signal : ; Where L' is the length of the downsampled sequence, which represents the time segment of the waveform, and d is the feature embedding dimension; The global context vector H context Do it in the following way: The matrix vector is obtained through linear transformation and nonlinear activation. ; Here, d is consistent with the dimension of the signal coding branch.
8. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 6, characterized in that: The noise locking and suppression based on the cross-attention mechanism is performed in the following manner: Using three learnable linear projection matrices, W Q W K W V The generated encoded features are converted into the query vector Q, key vector K, and value vector V required by the attention mechanism; Generation of query vector Q: using matrix context vector H context Generate Q; ; Q represents a judgment based on the current visual perception; Generation of key vector K and value vector V: using signal feature sequence H signal Generate K and V; ; K is the index label of each local waveform segment in the dynamic curve, and V is the information content actually carried by these segments; Calculate the dot product of Q and K, and scale it to obtain the attention score matrix A; ; Softmax is a normalization exponential function used to transform an input numerical vector into a probability distribution vector. K T This refers to the transpose of K; d k This refers to the feature dimension of K; During the full-domain scanning process, the similarity between the "visual prior" and "each waveform segment" is calculated. If the visual module identifies a sample with extremely high oil content, then the Q vector will contain information about the "oil film noise pattern." At this point, any segment in the curve that matches the "random spike" characteristic, i.e., noise caused by oil, will have a very high matching degree between its corresponding K value and Q, thus obtaining a high score in matrix A. Conversely, the actual pollutant signal waveform will have a low score due to feature mismatch. A high score means that in the attention scoring matrix A, the value at that position is significantly higher than the uniform distribution probability value in the same dimension or significantly higher than the weight of other non-interfering segments; this means that the features of the identified waveform segment are highly similar to Q. A low score refers to a value in the attention scoring matrix A that is close to zero or significantly lower than the weight of the main interference item; this means that the characteristics of the waveform segment do not match the matrix noise pattern and are judged as real chemical signals or non-specific background noise. The value vector V is weighted using the scoring matrix A, and residual connections are introduced to generate the fusion feature H. fused ; In this step, the model employs a reverse suppression strategy: features with high attention scores are weighted or filtered out in subsequent feature reorganization; features with low attention scores are retained and enhanced; thus completing the qualitative change from "noisy features" to "clean features". It also includes the processes of signal reconstruction and concentration inversion: Using a one-dimensional deconvolutional neural network, Transposed Conv1D, as a decoder, the fused features H are... fused Reconstructing the temporal space from the feature space; The decoder gradually recovers the data length through upsampling operations, ultimately outputting a reconstructed, clean response curve C. clean ; Compared to the original input curve, this reconstructed curve has significant features: baseline drift is corrected to zero, random spikes caused by plant fibers are smoothed, background absorbance caused by pigments is subtracted, and only specific peaks directly related to pollutant concentration are retained. To directly output the detection results, a regression head is connected in parallel; it consists of a global average pooling layer and a fully connected layer, directly reading the peak height and peak area characteristics of the purity curve, and outputting the final pollutant concentration value Y through linear regression calculation. pred .
9. The method for detecting food safety in complex matrices using multimodal fusion as described in claim 6, characterized in that: The system's built-in memory contains a database of food safety testing standards, including maximum residue limits (T) for various pesticide residues, heavy metals, and illegal additives. limit When the final pollutant concentration value Y is output... pred Then, numerical comparisons are performed: If Y pred <T limit The test result was determined to be "qualified"; If Y pred ≥T limit The test result was determined to be "unqualified", and the excess multiple was automatically calculated.
10. A detection system for implementing a multimodal fusion-based food safety detection method according to any one of claims 1-9, characterized in that: Includes the following modules: The data acquisition module is used to acquire the global image, microstructure, and diffuse reflectance spectrum of the dish, and analyze it to obtain the matrix feature vector V. matrix ; The selection module is used to select based on the matrix feature vector V. matrix Select a detection strategy; The kinetic detection module is used to perform kinetic detection according to the detection strategy to obtain the original reaction kinetic curve; Analysis module, used for analyzing matrix feature vector V matrix The original reaction kinetic curves were analyzed to obtain the analytical results.