An information processing method and system for improving food detection effect
By combining the multi-head self-attention mechanism and the Transformer model, the bottleneck problem of multimodal data fusion in food inspection is solved, high precision and robustness of food quality assessment are achieved, and the food inspection effect is improved.
Patent Information
- Application Number
- CN202510323785.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing food inspection methods have bottlenecks in multimodal data fusion, dynamic association modeling and global feature expression. In particular, the temporal asynchrony of multi-sensor data makes feature alignment difficult. Traditional models lack explicit modeling of their implicit dependence on the time dimension, making it difficult to capture the continuous dynamic characteristics of food quality degradation.
Using a multi-head self-attention mechanism and Transformer model, the food data is preprocessed, feature extracted, and timestamp aligned to generate a time-enhanced embedding vector. Multi-head attention is used to calculate the interaction relationship, and layer normalization and feedforward neural network fusion are performed. Combined with position encoding and depth analysis, the food quality assessment results are generated.
It improves the accuracy of food detection, enhances the robustness to food quality degradation paths, reduces time series alignment errors, realizes differentiated expression of multimodal data and noise-resistant feature capture, and improves the accuracy and reliability of food quality assessment.
Smart Images

Figure CN120145316B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of food detection, and in particular to an information processing method and system for improving food detection effects. Background Art
[0002] Food quality testing technology is gradually evolving toward multi-source, heterogeneous data fusion and intelligent analysis. Traditional testing methods primarily rely on single physical and chemical indicators or microbiological testing, resulting in long testing cycles and difficulty reflecting dynamic changes in food quality. With the widespread adoption of technologies such as spectral analysis, image recognition, and gas sensing, multimodal data acquisition and joint analysis have become a research hotspot. Machine learning models are widely used for feature extraction and classification, and deep learning techniques have further enhanced the ability to model complex nonlinear relationships.
[0003] However, existing methods still face significant bottlenecks in cross-modal time series data fusion, dynamic correlation modeling, and global feature expression. The temporal asynchrony of multi-sensor data makes feature alignment difficult, while traditional models lack explicit modeling of the implicit dependency on time, making it difficult to capture the continuous dynamic characteristics of food quality degradation. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an information processing method for improving food detection results to solve the problem of limited detection accuracy caused by the lack of temporal correlation of multimodal data and insufficient interaction of high-order features.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an information processing method for improving food detection results, which comprises:
[0008] Collecting food data, including pH, moisture content, appearance image, spectral data, and VOCs concentration data, and preprocessing the food data;
[0009] Extracting features from the food data, aligning and concatenating them according to timestamps to obtain an input matrix, mapping them to a high-dimensional space through linear transformation to generate an embedding vector, and adding position encoding to the embedding vector to obtain a time-enhanced embedding vector;
[0010] The interaction between the temporal enhancement embedding vectors is calculated through a multi-head self-attention mechanism to generate a multi-head attention output. The multi-head attention output is residually connected with the temporal enhancement embedding vector, and then layer normalization is performed. The multi-head attention output is fused through a feedforward neural network and then global average pooling is performed to obtain a comprehensive feature vector.
[0011] Transformer is used to perform in-depth analysis on the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
[0012] As a preferred solution of the information processing method for improving food detection effect of the present invention, the specific steps of generating multi-head attention output are as follows:
[0013] Map the temporal augmented embedding vector to the query matrix, key matrix, and value matrix respectively through linear transformation;
[0014] Split the query matrix, key matrix, and value matrix into 4 attention heads along the feature dimension;
[0015] Calculate the output of each attention head and concatenate the outputs of the four attention heads along the feature dimension to generate the output of the multi-head attention.
[0016] As a preferred solution of the information processing method for improving food detection effect described in the present invention, wherein: adding position coding to the embedding vector to obtain a time-enhanced embedding vector means dividing each dimension of the embedding vector into two categories, odd and even, using a sine function to generate even dimensions, and using a cosine function to generate odd dimensions, and adding the generated position coding to the embedding vector element by element to obtain a time-enhanced embedding vector.
[0017] As a preferred embodiment of the information processing method for improving food inspection results of the present invention, generating food quality assessment results refers to configuring the hyperparameters of a Transformer, mapping the comprehensive feature vector to a high-dimensional space through a linear transformation, and generating a deep analysis embedding vector;
[0018] Add position encoding to the deep analysis embedding vector, calculate the interaction between the deep analysis embedding vectors through the multi-head self-attention mechanism, generate the deep analysis multi-head self-attention output, perform residual connection on the deep analysis multi-head self-attention output and the deep analysis embedding vector, and then perform layer normalization to generate the deep processing comprehensive feature vector;
[0019] The fully connected layer and Softmax function are used to classify the deeply processed comprehensive feature vector to generate the food safety assessment result. The fully connected layer and Sigmoid function are used to regress the deeply processed comprehensive feature vector to generate the nutritional composition and freshness score of the food. The food safety assessment results, nutritional composition and freshness score of the food are integrated into the food quality assessment result.
[0020] As a preferred embodiment of the information processing method for improving food detection results of the present invention, the preprocessing of the food data refers to converting the appearance image into a grayscale image, normalizing the spectral data, and normalizing the spectral data, pH value, moisture content, and VOCs concentration data using a minimum-maximum normalization method;
[0021] A sliding average filter was used to smooth the pH value, moisture content, and VOCs concentration. The 3σ principle was used to detect and eliminate outliers in the food data, and linear interpolation was used to fill in the eliminated outliers.
[0022] As a preferred solution of the information processing method for improving food detection results of the present invention, the food data is subjected to feature extraction, and aligned and spliced according to timestamps to obtain an input matrix. The specific steps are as follows:
[0023] Select MobileNet to extract features from the appearance image and obtain the image feature vector;
[0024] Select partial least squares regression to extract features from spectral data and obtain spectral feature vectors;
[0025] Principal component analysis was used to extract features from VOCs concentration data to obtain odor feature vectors;
[0026] The image feature vector, spectral feature vector and odor feature vector are aligned and concatenated according to the timestamp to form the input matrix.
[0027] As a preferred embodiment of the information processing method for improving food inspection performance described in the present invention, obtaining the comprehensive feature vector refers to adding the output of the multi-head attention to the time-enhanced embedding vector element by element, then calculating the mean and variance of each feature dimension, and normalizing them to obtain a layer-normalized output;
[0028] Map the normalized output of the layer to the intermediate dimension through linear transformation to generate an intermediate feature vector. Use the ReLU activation function to perform nonlinear transformation on the intermediate feature vector, and then map it back to the original dimension through linear transformation to generate a fused feature vector.
[0029] The values of each feature dimension of the fused feature vector at all time steps are averaged to obtain the comprehensive feature vector.
[0030] In a second aspect, the present invention provides an information processing system for improving food detection results, comprising:
[0031] A preprocessing module collects food data, including pH, moisture content, appearance image, spectral data, and VOCs concentration data, and preprocesses the food data;
[0032] A position encoding module extracts features from the food data, aligns and concatenates them according to timestamps to obtain an input matrix, maps it to a high-dimensional space through linear transformation, generates an embedding vector, and adds position encoding to the embedding vector to obtain a time-enhanced embedding vector;
[0033] The pooling module calculates the interaction between the temporal enhancement embedding vectors through a multi-head self-attention mechanism, generates a multi-head attention output, performs a residual connection on the multi-head attention output and the temporal enhancement embedding vector, then performs layer normalization and fuses them through a feedforward neural network, followed by global average pooling to obtain a comprehensive feature vector;
[0034] The classification and regression module uses Transformer to perform in-depth analysis of the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
[0035] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the information processing method for improving food detection effects as described in the first aspect of the present invention is implemented.
[0036] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the information processing method for improving food detection effects as described in the first aspect of the present invention is implemented.
[0037] The beneficial effects of the present invention are as follows: by linearly mapping the time-enhanced embedding vector into a query matrix, a key matrix, and a value matrix, and splitting it into four attention heads along the feature dimension, each attention head independently calculates the interaction relationship and then splices the results, achieving multi-subspace differentiated feature interaction modeling, thereby avoiding the modeling limitations of a single attention mechanism for complex associations, improving the differentiated expression ability of cross-modal dynamic features, and reducing misjudgments caused by single attention weight bias. By using sine and cosine functions to generate position encodings for the odd and even dimensions of the embedding vector, respectively, and adding them element-by-element to the embedding vector, noise-resistant embedding of temporal dynamic features is achieved, reducing temporal alignment errors and enhancing robustness to food quality degradation paths. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 This is a flow chart of the information processing method for improving food detection results in Example 1.
[0040] Figure 2 Schematic diagram of the information processing system for improving food detection results in Example 1.
[0041] Figure 3 Schematic diagram of the calculation process of the multi-head attention mechanism in Example 1.
[0042] Figure 4 Schematic diagram of the Transformer deep analysis process in Example 1. DETAILED DESCRIPTION
[0043] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0045] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0046] Example 1, with reference to Figures 1 to 4 , which is the first embodiment of the present invention, provides an information processing method for improving food detection results, including the following steps:
[0047] S1: Collect food data, including pH, moisture content, appearance image, spectral data and VOCs concentration data, and pre-process the food data.
[0048] The specific steps are as follows:
[0049] The pH sensor model SEN0161 is selected to measure the pH of food. The measurement range is 0-14 and the accuracy is ±0.1.
[0050] The moisture sensor model YL-69 is selected to measure the moisture content of food, with a measuring range of 0-100% and an accuracy of ±1%.
[0051] A camera with a resolution of 1920x1080 is selected to capture the appearance images of the food, with a frame rate of 30fps.
[0052] A near-infrared spectrometer model NIRQuest512 was selected to collect near-infrared spectral data of food, with a wavelength range of 900-1700 nm and a resolution of 3.5 nm.
[0053] The MQ-3 odor sensor is selected to detect the concentration of volatile organic compounds (VOCs) in food. The MQ-3 has high sensitivity to VOCs.
[0054] Start all sensors and begin collecting food data in real time. The collection frequency is set to once per second.
[0055] All sensors are connected to edge computing devices (such as Raspberry Pi 4) via Internet of Things (IoT) technology, ensuring real-time transmission of food data.
[0056] The collected food data is transmitted to the edge computing device via the MQTT protocol.
[0057] The appearance image captured by the high-resolution camera is converted into a grayscale image with a grayscale value range of 0-255. The RGB three channels are converted into a single-channel grayscale value using the weighted average method, with weights of 0.2989 (red channel), 0.5870 (green channel), and 0.1140 (blue channel).
[0058] The spectral data collected by the near-infrared spectrometer were normalized to a uniform wavelength range (900–1700 nm) and scaled to the range of 0–1 using the minimum–maximum normalization method.
[0059] The pH value and water content were scaled to the range of 0–1 using the min–max normalization method.
[0060] The volatile organic compound (VOCs) concentration data detected by the odor sensor were scaled to the range of 0–1 using the minimum–maximum normalization method, with a concentration range of 0–1000 ppm.
[0061] The pH, moisture content, and VOCs concentration data were smoothed using a sliding average filter with a window size of 5 data points.
[0062] The 3σ principle was used to detect outliers in food data. For each data point, the deviation from the mean was calculated. If the deviation exceeded three standard deviations, the data point was marked as an outlier and removed.
[0063] For the outliers removed, linear interpolation is used to fill in the gaps and ensure data continuity.
[0064] Using a SQLite database to manage food data means creating multiple data tables to store pH, moisture content, appearance images, spectral data, and VOCs concentration. Each table contains a timestamp field for data alignment.
[0065] It should also be noted that the sliding average window and 3σ outlier elimination improve the noise suppression effect of multi-source data compared to traditional single-sensor or static threshold detection methods. In addition, through IoT real-time transmission and SQLite timestamp alignment, the timing misalignment problem caused by the asynchrony of multimodal data in existing technologies is solved, providing a high-precision timing benchmark for subsequent feature fusion.
[0066] S2: Extract features from the food data, align them according to timestamps to obtain an input matrix, map them to a high-dimensional space through linear transformation, generate an embedding vector, add position encoding to the embedding vector, and obtain a time-enhanced embedding vector.
[0067] The specific steps are as follows:
[0068] Extract food data from a SQLite database, including pH, moisture content, appearance images, spectral data, and VOCs concentration.
[0069] MobileNet is selected to perform feature extraction on the appearance image to obtain an image feature vector. The dimension of the image feature vector is 1280, which represents the appearance features of the image, such as color, texture, and shape.
[0070] Partial least squares regression (PLSR) is selected to extract features from spectral data and obtain spectral feature vectors. The dimension of the spectral feature vector is 10, which represents the chemical composition characteristics of food such as moisture content and fat content.
[0071] Principal component analysis (PCA) was selected to extract features from the volatile organic compound (VOCs) concentration data to obtain the odor feature vector. The dimension of the odor feature vector is 5, which represents the type and concentration of volatile organic compounds in food.
[0072] The image feature vector, spectral feature vector, and odor feature vector are aligned and concatenated according to the timestamp to form an input matrix (with a dimension of T × 1295, where T is the time step and 1295 is the total feature dimension).
[0073] The input matrix is mapped to a high-dimensional space through a linear transformation to generate an embedding vector. The dimension of the linear transformation matrix is 1295×256, and the input matrix is mapped from T×1295 to T×256. Each row of the embedding vector represents the feature representation of a time step, and 256 is the dimension of the embedding vector.
[0074] Each dimension of the embedding vector is divided into odd and even categories. The even dimensions are generated using the sine function, and the odd dimensions are generated using the cosine function. The generated positional encoding is added to the embedding vector element by element to obtain the time-enhanced embedding vector.
[0075] Through the odd-even dimension divide-and-conquer design, different dimensions of the embedded vector are given differentiated time series modeling capabilities. The even dimension uses the sine function to generate low-frequency oscillation components, focusing on capturing the macroscopic laws of the time span (such as the temperature decay trend within the storage period); the odd dimension uses the cosine function to generate high-frequency oscillation components, focusing on fine-grained time fluctuations (such as the impact of day and night temperature differences on the food metabolism rate). Through progressive wavelength control, the shallow dimension strengthens short-term differences (such as minute-level sampling intervals), and the deep dimension establishes long-term associations (such as the cumulative effect of corruption across periods), forming multi-scale time perception capabilities. At the same time, the orthogonality of the sine and cosine functions is used to ensure information independence between dimensions and avoid coding redundancy.
[0076] Dynamic calibration and interference suppression technology are introduced in the position coding fusion stage to achieve precise adaptation of the time signal and the original features. Through amplitude normalization and phase offset compensation, the numerical distribution of the position coding and the embedded vector is aligned to avoid signal strength imbalance; radial constraint mask and directional filter are used to suppress frequency interference between odd and even dimensions (such as high-frequency noise masks low-frequency trends), respectively, to ensure that the spatial topological structure of the original features is retained after fusion, and the semantic association of the temporal context is enhanced. This mechanism can dynamically adapt to the complex needs of industrial scenarios, such as compensating for equipment sampling anomalies through wavelength parameter reorganization, or adjusting the coding weights in real time based on environmental parameters to improve the robustness of modeling for emergencies (such as cold chain breaks).
[0077] It should also be noted that introducing position encoding and enhancing the time perception ability of the embedding vector through odd-even dimension divide-and-conquer design can better capture the subtle differences in food quality over time, thereby improving the effect of food quality assessment.
[0078] S3: The interaction relationship between the temporal enhancement embedding vectors is calculated through the multi-head self-attention mechanism to generate the multi-head attention output, which is residually connected with the temporal enhancement embedding vector, and then layer normalization is performed. The multi-head attention output is fused through a feedforward neural network, and then global average pooling is performed to obtain a comprehensive feature vector.
[0079] The specific steps are as follows:
[0080] The temporal augmented embedding vectors are mapped to the query matrix, key matrix, and value matrix respectively through linear transformation. The dimension of each matrix is T×64 (the number of attention heads is 4 and the dimension of each head is 64).
[0081] The query matrix, key matrix, and value matrix are evenly divided into 4 independent attention heads along the feature dimension. Each head corresponds to a 64-dimensional feature space, focusing on the interaction patterns of different time spans or feature combinations. For example, the first attention head focuses on the local correlation between adjacent time steps (such as the synchronization of moisture content mutation and appearance change), the second attention head captures long-range temporal dependencies (such as the cumulative effect of initial pH on the final VOCs concentration), the third attention head analyzes cross-modal associations (such as the correspondence between specific spectral bands and odor characteristics), and the fourth attention head learns global feature collaboration (such as the comprehensive weight distribution of multi-sensor data).
[0082] Calculate the dot product of the query matrix and the key matrix and divide it by the square root of the key vector dimension. Then generate the attention weights through the Softmax function. Use the attention weights to perform weighted summation on the value matrix to generate the output of each attention head. The outputs of the four attention heads are spliced along the feature dimension to generate the output of the multi-head attention. The spliced dimension is T×256.
[0083] The output of the multi-head attention is added element-wise to the temporal enhanced embedding vector (TEEV) to obtain the output of the residual connection, preserving the original feature information. The output dimension of the residual connection is T×256.
[0084] The mean and variance of each feature dimension of the residual connection output are calculated and normalized to obtain the layer-normalized output to ensure the stability of the feature distribution. The layer-normalized output dimension is T×256.
[0085] The normalized output of the layer is input into a feedforward neural network, and the normalized output of the layer is mapped to the intermediate dimension through linear transformation to generate an intermediate feature vector. The intermediate feature vector is nonlinearly transformed using the ReLU activation function, and then mapped back to the original dimension through linear transformation to generate a fused feature vector with a dimension of T×256. The feedforward neural network consists of two fully connected layers, with the ReLU activation function used in the middle.
[0086] The fused feature vector is average pooled along the time dimension, that is, the values of each feature dimension (256) at all time steps are averaged to generate a comprehensive feature vector.
[0087] It should also be noted that the multi-head self-attention mechanism performs better than traditional single-layer attention models in capturing different aspects of input data. This approach allows for simultaneous attention to data in multiple modalities (such as images, spectra, and odors), providing richer context for subsequent analysis. In addition, the use of residual connections and layer normalization helps alleviate the vanishing gradient problem during deep network training, and global average pooling further refines key features, making the final integrated feature vector more compact and meaningful.
[0088] S4: Use Transformer (converter model) to conduct in-depth analysis of the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
[0089] The specific steps are as follows:
[0090] The comprehensive feature vector is encrypted using the AES-256 encryption algorithm to ensure the security of data transmission. The encrypted data is transmitted to the central server through the HTTPS protocol. After the central server receives the data, it uses the same AES-256 key to decrypt it and restore the original comprehensive feature vector.
[0091] Transformer is used to perform in-depth analysis of the comprehensive feature vector. Transformer hyperparameters, such as learning rate, batch size, and number of training rounds, are configured to meet the needs of real-time analysis. The initial learning rate is set to 1e-5, increasing by 10% for every 100 samples. The default batch size is 32. When the request queue exceeds 50, it switches to sample-by-sample processing (batch size of 1). At night, when the load is low, the batch size is restored to 64. Online single-round training (1 epoch), incremental fine-tuning for 5 epochs per hour, and full parameter training for 3 epochs per week.
[0092] The comprehensive feature vector is mapped to a high-dimensional space through a linear transformation to generate a deep analysis embedding vector (DAEV). The DAEV has a dimension of 512 to capture deeper semantic information. Positional encoding is added to the DAEV using sine and cosine functions. The interaction between the DAEVs is calculated using a multi-head self-attention mechanism. Each attention head dynamically calculates the interaction between the query matrix, key matrix, and value matrix to generate deep analysis attention weights. The output of the multi-head attention is connected through residual connections and layer normalization to generate a deep processing comprehensive feature vector.
[0093] The fully connected layer and the Softmax function are used to classify the deeply processed comprehensive feature vector to generate the food safety assessment result (such as safe or unsafe). The fully connected layer and the Sigmoid function are used to regress the deeply processed comprehensive feature vector to generate the nutritional components (such as protein content and fat content) and freshness score (range 0-1) of the food. The food safety assessment result, the nutritional components and the freshness score of the food are integrated into the food quality assessment result.
[0094] It should also be noted that using the Transformer architecture for deep analysis offers significant advantages over other machine learning or deep learning methods, particularly when processing sequence data. Transformers not only effectively capture long-range dependencies but also dynamically adjust hyperparameters to adapt to diverse application scenarios. The AES-256 encryption algorithm ensures secure data transmission and protects the privacy of sensitive data during food quality assessment. Combining the outputs of classification and regression tasks provides detailed and reliable assessments of food safety and nutritional content, which is crucial for improving food quality and safety.
[0095] This embodiment also provides an information processing system for improving food detection results, including:
[0096] A preprocessing module collects food data, including pH, moisture content, appearance image, spectral data, and VOCs concentration data, and preprocesses the food data;
[0097] A position encoding module extracts features from the food data, aligns and concatenates them according to timestamps to obtain an input matrix, maps it to a high-dimensional space through linear transformation, generates an embedding vector, and adds position encoding to the embedding vector to obtain a time-enhanced embedding vector;
[0098] The pooling module calculates the interaction between the temporal enhancement embedding vectors through a multi-head self-attention mechanism, generates a multi-head attention output, performs a residual connection on the multi-head attention output and the temporal enhancement embedding vector, then performs layer normalization and fuses them through a feedforward neural network, followed by global average pooling to obtain a comprehensive feature vector;
[0099] The classification and regression module uses Transformer to perform in-depth analysis of the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
[0100] This embodiment also provides a computer device suitable for an information processing method for improving food inspection results, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the information processing method for improving food inspection results proposed in the above embodiment.
[0101] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0102] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the information processing method for improving food inspection results proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0103] In summary, the present invention achieves multi-subspace differentiated feature interaction modeling by linearly mapping the temporally enhanced embedding vector into a query matrix, a key matrix, and a value matrix, respectively, and then segmenting it into four attention heads along the feature dimension. Each attention head independently calculates the interaction relationship and then concatenates the results. This avoids the limitations of a single attention mechanism in modeling complex associations, improves the ability to differentiate cross-modal dynamic features, and reduces misjudgments caused by bias in a single attention weight. By using sine and cosine functions to generate positional encodings for the odd and even dimensions of the embedding vector, respectively, and adding them element-by-element to the embedding vector, noise-resistant embedding of temporal dynamic features is achieved, reducing temporal alignment errors and enhancing robustness to food quality degradation paths.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An information processing method for improving food inspection results, characterized by: include, Collecting food data, including pH, moisture content, appearance image, spectral data, and VOCs concentration data, and preprocessing the food data; Extracting features from the food data, aligning and concatenating them according to timestamps to obtain an input matrix, mapping them to a high-dimensional space through linear transformation to generate an embedding vector, and adding position encoding to the embedding vector to obtain a time-enhanced embedding vector; The interaction between the temporal enhancement embedding vectors is calculated through a multi-head self-attention mechanism to generate a multi-head attention output. The multi-head attention output is residually connected with the temporal enhancement embedding vector, and then layer normalization is performed. The multi-head attention output is fused through a feedforward neural network and then global average pooling is performed to obtain a comprehensive feature vector. Transformer is used to perform in-depth analysis on the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
2. The information processing method for improving food inspection results according to claim 1, characterized in that: The specific steps for generating multi-head attention output are as follows: Map the temporal augmented embedding vector to the query matrix, key matrix, and value matrix respectively through linear transformation; Split the query matrix, key matrix, and value matrix into 4 attention heads along the feature dimension; Calculate the output of each attention head and concatenate the outputs of the four attention heads along the feature dimension to generate the output of the multi-head attention.
3. The information processing method for improving food inspection results according to claim 1, characterized in that: Adding position coding to the embedding vector to obtain the time-enhanced embedding vector means dividing each dimension of the embedding vector into odd and even categories, generating even dimensions using a sine function, generating odd dimensions using a cosine function, and adding the generated position coding to the embedding vector element by element to obtain the time-enhanced embedding vector.
4. The information processing method for improving food inspection results according to claim 1, wherein: Generating the food quality assessment results refers to configuring the hyperparameters of the Transformer, mapping the comprehensive feature vector to a high-dimensional space through a linear transformation, and generating a deep analysis embedding vector; Add position encoding to the deep analysis embedding vector, calculate the interaction between the deep analysis embedding vectors through the multi-head self-attention mechanism, generate the deep analysis multi-head self-attention output, perform residual connection on the deep analysis multi-head self-attention output and the deep analysis embedding vector, and then perform layer normalization to generate the deep processing comprehensive feature vector; The fully connected layer and Softmax function are used to classify the deeply processed comprehensive feature vector to generate the food safety assessment result. The fully connected layer and Sigmoid function are used to regress the deeply processed comprehensive feature vector to generate the nutritional composition and freshness score of the food. The food safety assessment results, nutritional composition and freshness score of the food are integrated into the food quality assessment result.
5. The information processing method for improving food inspection results according to claim 1, wherein: Preprocessing the food data includes converting the appearance image into a grayscale image, normalizing the spectral data, and normalizing the spectral data, pH value, moisture content, and VOCs concentration data using a minimum-maximum normalization method; A sliding average filter was used to smooth the pH value, moisture content, and VOCs concentration. The 3σ principle was used to detect and eliminate outliers in the food data, and linear interpolation was used to fill in the eliminated outliers.
6. The information processing method for improving food inspection results according to claim 1, characterized in that: Feature extraction is performed on the food data, and alignment and splicing are performed according to the timestamp to obtain the input matrix. The specific steps are as follows: Select MobileNet to extract features from the appearance image and obtain the image feature vector; Select partial least squares regression to extract features from spectral data and obtain spectral feature vectors; Principal component analysis was used to extract features from VOCs concentration data to obtain odor feature vectors; The image feature vector, spectral feature vector and odor feature vector are aligned and concatenated according to the timestamp to form the input matrix.
7. The information processing method for improving food inspection results according to claim 2, characterized in that: Obtaining the comprehensive feature vector refers to adding the output of the multi-head attention to the time-enhanced embedding vector element by element, then calculating the mean and variance of each feature dimension and normalizing them to obtain the layer-normalized output; Map the normalized output of the layer to the intermediate dimension through linear transformation to generate an intermediate feature vector. Use the ReLU activation function to perform nonlinear transformation on the intermediate feature vector, and then map it back to the original dimension through linear transformation to generate a fused feature vector. The values of each feature dimension of the fused feature vector at all time steps are averaged to obtain the comprehensive feature vector.
8. An information processing system for improving food inspection results, based on the information processing method for improving food inspection results according to any one of claims 1 to 7, characterized in that: include, A preprocessing module collects food data, including pH, moisture content, appearance image, spectral data, and VOCs concentration data, and preprocesses the food data; A position encoding module extracts features from the food data, aligns and concatenates them according to timestamps to obtain an input matrix, maps it to a high-dimensional space through linear transformation, generates an embedding vector, and adds position encoding to the embedding vector to obtain a time-enhanced embedding vector; The pooling module calculates the interaction between the temporal enhancement embedding vectors through a multi-head self-attention mechanism, generates a multi-head attention output, performs a residual connection on the multi-head attention output and the temporal enhancement embedding vector, then performs layer normalization and fuses them through a feedforward neural network, followed by global average pooling to obtain a comprehensive feature vector; The classification and regression module uses Transformer to perform in-depth analysis of the comprehensive feature vector, classify and regress the analysis results, and generate food quality assessment results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the information processing method for improving food detection effect according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the information processing method for improving food detection effect according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
System and method for monitoring and quality evaluation of perishable food items
US20200250531A1
Method for multimodal emotion classification based on modal space assimilation and contrastive learning
US20240119716A1