A deep learning-based food classification method, system, and readable storage medium

By transforming one-dimensional time series data into two-dimensional data and combining feature extraction methods with multi-branch convolutional neural networks and LSTM layers, the problem of low accuracy in beef classification in existing technologies has been solved, achieving more efficient food supervision.

CN116933156BActive Publication Date: 2025-11-14CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310879868.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-17
Publication Date
2025-11-14
Estimated Expiration
2043-07-17

AI Technical Summary

Technical Problem

Existing technologies for classifying beef using food spectrometers fail to fully utilize the advantages of convolutional networks in processing two-dimensional data, resulting in low classification accuracy.

Method used

One-dimensional time series data is transformed into two-dimensional data, and features are extracted by constructing a multi-branch convolutional neural network, including two-dimensional processing of first-order difference sequences and Fourier sequences. Spatial attention and channel attention mechanisms are combined, LSTM layers are used for temporal feature extraction, and finally classification is performed through SoftMax layers.

Benefits of technology

This improves the accuracy of food classification, enabling better capture of spectral differences between beef samples and enhancing the accuracy and efficiency of food supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116933156B_ABST
    Figure CN116933156B_ABST
Patent Text Reader

Abstract

This application provides a deep learning-based food classification method, system, and readable storage medium. The method includes acquiring a training dataset, which comprises a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type; using the training dataset as the original time series and transforming the original time series to obtain a first-order difference sequence and a Fourier sequence; performing two-dimensional processing on each sequence using a tiling and sliding window approach to obtain corresponding two-dimensional data; constructing an initial classification model and inputting the obtained two-dimensional data and Fourier sequences into the initial classification model for training, obtaining a target classification model upon training termination; and inputting the acquired two-dimensional data to be processed and the Fourier sequences into the target classification model to obtain the food classification result. The implementation of this method can improve the accuracy of food classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning application technology, and more specifically, to a food classification method, system, and readable storage medium based on deep learning. Background Technology

[0002] Food safety and quality assurance have always been a global focus, and food spectrometers can provide significant support in these areas. Currently, food spectrometers are used for food classification and management. For example, in the beef sector, the offal content is a crucial assessment indicator, affecting the flavor, nutritional value, and quality of beef. Specifically, food spectrometers are used to analyze the offal content of beef samples, categorizing them into different types, including pure beef and beef with varying degrees of offal. This further assists regulatory agencies and food production companies in accurately assessing the composition and quality of beef, thereby improving the efficiency and accuracy of food safety supervision.

[0003] Currently, to improve the speed of food quality testing, some researchers have introduced deep learning algorithms for food classification and management. However, most current research focuses on directly inputting the original spectral data of the time series and using the original time series for classification training. This does not fully utilize the advantages of convolutional networks in processing two-dimensional data or the feature extraction advantages of multi-branch convolutional networks, resulting in low classification accuracy. Summary of the Invention

[0004] The purpose of this application is to provide a food classification method, system, and readable storage medium based on deep learning, which can improve the accuracy of food classification.

[0005] This application also provides a food classification method based on deep learning, including the following steps:

[0006] S1. Obtain a training dataset, which includes a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type.

[0007] S2. The training dataset is used as the original time series, and the original time series is transformed to obtain the first-order difference sequence and the Fourier sequence.

[0008] S3. Each sequence is processed into two dimensions using a tiling and sliding window method to obtain the corresponding two-dimensional data;

[0009] S4. Construct an initial classification model, and input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training. When the training terminates, the target classification model is obtained.

[0010] S5. Input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

[0011] Secondly, this application also provides a food classification system based on deep learning, the system comprising a data acquisition module, a sequence conversion module, a two-dimensionalization module, a model training module, and a food classification module, wherein:

[0012] The data acquisition module is used to acquire a training dataset, which includes a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type.

[0013] The sequence conversion module is used to take the training dataset as the original time series and convert the original time series to obtain a first-order difference sequence and a Fourier sequence.

[0014] The two-dimensionalization module is used to perform two-dimensionalization processing on each sequence using a tiling and sliding window mode to obtain the corresponding two-dimensional data;

[0015] The model training module is used to construct an initial classification model, and input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training, and obtain the target classification model when the training terminates;

[0016] The food classification module is used to input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

[0017] Thirdly, this application also provides a readable storage medium including a deep learning-based food classification method program, which, when executed by a processor, implements the steps of a deep learning-based food classification method as described in any of the preceding claims.

[0018] As shown above, this application provides a deep learning-based food classification method, system, and readable storage medium. It transforms the extracted one-dimensional time series into two-dimensional data, inputs this two-dimensional data into a convolutional neural network (RNN) for feature classification, fully leveraging the advantages of RNNs in two-dimensional data feature extraction. This addresses the problem of traditional RNN models performing poorly in classifying entire time series, thus improving the accuracy of food classification. In summary, this application demonstrates superior accuracy compared to most current classic machine learning and deep learning methods, and can advance research into food classification management using existing artificial intelligence, thereby improving the regulatory level of the food industry and ensuring consumers receive safe and high-quality food.

[0019] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0020] To more clearly illustrate the technical solution of this application, the accompanying drawings used in this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating a deep learning-based food classification method provided for this application;

[0022] Figure 2 This is a schematic diagram of a portion of the training dataset;

[0023] Figure 3 This is a schematic diagram illustrating the principle of two-dimensional processing;

[0024] Figure 4 This is a schematic diagram of the classification model.

[0025] Figure 5 A flowchart illustrating the implementation of the beef classification method;

[0026] Figure 6 This is a schematic diagram of the structure of a food classification system based on deep learning, which is provided for this application. Detailed Implementation

[0027] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0028] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0029] Please refer to Figure 1 , Figure 1 This is a flowchart of a deep learning-based food classification method according to some embodiments of this application. The method includes the following steps:

[0030] Step S1: Obtain the training dataset, which includes a first time-series spectral signal corresponding to the pure sample type and a second time-series spectral signal corresponding to the adulterated sample type.

[0031] Step S2: Use the training dataset as the original time series and transform the original time series to obtain a first-order difference sequence and a Fourier sequence.

[0032] Step S3: Use a tiling and sliding window approach to process each sequence into two dimensions to obtain the corresponding two-dimensional data.

[0033] Step S4: Construct an initial classification model, and input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training. When the training terminates, the target classification model is obtained.

[0034] Step S5: Input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

[0035] As can be seen from the above, the food classification method based on deep learning disclosed in this application transforms the extracted one-dimensional time series into two-dimensional data. This two-dimensional data is then input into a convolutional neural network (RNN) for feature classification, fully leveraging the advantages of RNNs in two-dimensional data feature extraction. This addresses the problem of traditional RNN models performing poorly in classifying entire time series, thus improving the accuracy of food classification. In summary, this application demonstrates superior accuracy compared to most current classical machine learning and deep learning methods. It can advance research into food classification management using existing artificial intelligence, thereby improving the regulatory level of the food industry and ensuring consumers receive safe and high-quality food.

[0036] In one embodiment, step S2, which involves converting the original time series to obtain a first-order difference sequence, includes: sequentially obtaining the values ​​of two adjacent items in the original time series, and calculating the difference between the two adjacent items to obtain the corresponding difference value; and sequentially combining the obtained difference values ​​to obtain a first-order difference sequence.

[0037] For specific details, please refer to the data format in the training dataset. Figure 2 In the process of transforming the original time series into a first-order difference series, each adjacent term is traversed, and the difference value of each adjacent term is calculated. Then, according to the original sequence arrangement, the recombination arrangement order between the difference values ​​is determined, and the desired first-order difference series is obtained by sequentially combining them according to this recombination arrangement order.

[0038] For example, for each adjacent two terms in a time series T, through t' i =t i+1 -t i Calculate the corresponding difference values. Then, determine the order of recombination among the obtained difference values, and sort and combine them according to this order to form a new sequence t' = {t'1, t'2, ..., t'}. i ,…,t' n-1 The sequence T' is the first-order difference sequence that is required.

[0039] In one embodiment, step S2, transforming the original time series to obtain a Fourier sequence, includes: performing a Fourier transform on the original time series to obtain a Fourier sequence.

[0040] Specifically, in the current embodiment, the Fast Fourier Transform will be used to convert the original time series into a Fourier sequence.

[0041] In the above embodiments, the original time series is converted into a one-dimensional time series by performing first-order difference sequence and Fourier sequence transformation, so that it can be input into the deep learning model for data processing.

[0042] In one embodiment, step S3, which involves performing two-dimensional processing on each sequence using a tiling and sliding window approach to obtain corresponding two-dimensional data, includes:

[0043] Step S31: Obtain the total length of the corresponding one-dimensional sequence and the length of the sequence retained features.

[0044] Specifically, assume that the total length of the obtained one-dimensional sequence is n, and the length of the sequence retaining features is d.

[0045] Step S32: Determine the window sliding step size based on the total length of the sequence.

[0046] Specifically, in this embodiment, the total length of the sequence is square-rooted to obtain the window sliding step size.

[0047] Step S33: Determine the sliding window value based on the sum of the window sliding step size and the sequence retained feature length.

[0048] Specifically, given the window sliding step size is... When the sequence retains a feature length of d, based on The value of the sliding window is obtained by summing the sums of d and d.

[0049] Step S34: Determine the number of rows of the two-dimensional data based on the window sliding step size, determine the number of columns of the two-dimensional data based on the sliding window value, and convert the corresponding one-dimensional sequence into the corresponding two-dimensional data based on the number of rows and columns of the two-dimensional data.

[0050] Specifically, for a one-dimensional time series of length n, the current embodiment will adjust the window sliding step size. Set as rows of two-dimensional data, and slide window values The data is set as a column of two-dimensional data, and based on this, a one-dimensional time series of length n is transformed into... Two-dimensional data.

[0051] For example, please refer to Figure 3 , Figure 3 Assuming the total length n of the corresponding one-dimensional time series is 16, and the feature length d is 2, then according to the formula for calculating the window sliding step size shown in step S32, the window sliding step size can be calculated to be 4. Furthermore, according to the formula for calculating the sliding window value shown in step S33, the sliding window value can be calculated to be 6. Finally, the one-dimensional time series with a total length of 16 is transformed into (4*6) two-dimensional data. For details, please refer to... Figure 3 To understand.

[0052] In one embodiment, please refer to Figure 3 Given that the number of columns in a two-dimensional data set is greater than the number of rows, for each row of the two-dimensional data set, the numerical content of the next row header will be added to the end of each row. This preserves the continuity between the upper and lower rows and improves the richness of feature extraction.

[0053] In the above embodiments, a tiling and sliding window approach is used to transform one-dimensional time series into two-dimensional data, which facilitates subsequent data processing by the classification network, further leveraging the advantages of the classification network in processing two-dimensional data and improving the accuracy of feature extraction.

[0054] In one embodiment, please refer to Figure 4 In step S4, the initial classification model includes a first convolutional layer, a first cascaded operation layer, a second convolutional layer, and a global average pooling layer connected in sequence, wherein:

[0055] The first convolutional layer is used to extract features from each of the obtained two-dimensional data to obtain multiple initial feature terms.

[0056] For details, please refer to Figure 4 The first convolutional layer consists of multiple convolutional network branches, wherein the number of convolutional network branches is equal to the number of types of two-dimensional data input to the first convolutional layer. For example, if the two-dimensional data input to the first convolutional layer includes two-dimensional processed data corresponding to the original time series, two-dimensional processed data corresponding to the first-order difference sequence, and two-dimensional processed data corresponding to the Fourier sequence (i.e., the number of types is 3), then the number of convolutional network branches in the first convolutional layer is 3, meaning that the current first convolutional layer consists of 3 convolutional network branches.

[0057] In one embodiment, each branch of the convolutional network includes multiple Conv convolutional layers, each followed by a Batch Normalization (BN) layer. Each Conv convolutional layer extracts features from the input. The first convolutional layer may only extract low-level features, while higher-level convolutional layers iteratively extract more complex features from these low-level features to improve feature richness. The BN layer transforms the data distribution of each layer to a state with a mean of zero and a variance of 1. This standardization process ensures that the data distribution of each layer remains roughly the same, making training easier to converge and improving training speed. In this embodiment, by fusing the Conv and BN layers, the training speed, model performance, and stability of the convolutional neural network can be improved, making the network more expressive and more robust to input data.

[0058] In one specific embodiment, the first convolutional layer consists of three first convolutional network branches, and each of the first convolutional network branches consists of three Conv convolutional layers, wherein each Conv convolutional layer is followed by a Batch Normalization (BN) layer. In the current embodiment, the three first convolutional network branches of the first convolutional layer extract corresponding initial feature terms from the input two-dimensional processed data corresponding to the original time series, the two-dimensional processed data corresponding to the first-order difference sequence, and the two-dimensional processed data corresponding to the Fourier sequence, respectively. The extracted initial feature terms are then input into the first cascaded operation layer for cascaded processing.

[0059] The first cascaded operation layer is used to perform cascaded processing on the extracted initial feature terms to obtain cascaded features.

[0060] For details, please refer to Figure 4 The first cascaded operation layer is Figure 4 The connection shown is in the concatenate layer after the three first convolutional network branches.

[0061] It should be noted that the Concatenate layer concatenates the initial feature terms output from the three first convolutional network branches through a concatenation operation to obtain the corresponding concatenated features.

[0062] The second convolutional layer is used to perform feature concatenation by analyzing the correlation between each initial feature item based on the cascaded features.

[0063] Specifically, the second convolutional layer consists of one second convolutional network branch, and the second convolutional network branch consists of multiple Conv convolutional layers, each of which is followed by a Batch Normalization (BN) layer. Since the functions of the Conv convolutional layers and the BN layers have already been explained, they will not be repeated here.

[0064] In a specific embodiment, such as Figure 4 As shown, the second convolutional network branch consists of two Conv convolutional layers, each followed by a Batch Normalization (BN) layer. In this embodiment, the second convolutional network branch analyzes the correlation between the initial feature terms extracted by the three first convolutional network branches from the input cascaded features, and then performs feature concatenation.

[0065] The global average pooling layer is used to reduce the dimensionality of the obtained concatenated features to obtain the target feature term.

[0066] For details, please refer to Figure 4 The global average pooling layer is Figure 4The diagram illustrates the connection to the GloabaAveragePooling operation layer following the second convolutional layer.

[0067] In one embodiment, both the first convolutional layer and the second convolutional layer employ a combination of spatial attention and channel attention mechanisms to extract attention weights for both spatial and channel attention dimensions. These attention weights are then weighted in a weighted manner to improve the accuracy of feature prediction.

[0068] Specifically, this application employs a combination of spatial attention and channel attention mechanisms. It first extracts the attention weights for both spatial and channel attention dimensions separately, and then aggregates the two attention mechanisms using a weighted average. The specific weighting process is illustrated in the following formula:

[0069] γ n =ω1(α) n )+ω2(β n );

[0070] Where, γ n The final weighted attention weights are α. n For spatial attention mechanisms, β n For channel attention mechanism, ω1 and ω2 are preset weight parameters.

[0071] In one embodiment, a residual mechanism is introduced in both the first and second convolutional layers to avoid gradient vanishing and improve model stability.

[0072] It should be noted that residual mechanisms are introduced in both the first and second convolutional layers. Without residual mechanisms, the parameter passing method between convolutional layers would be as follows:

[0073] x n+1 =P(x n W n );

[0074] Where, x n With x n+1 These represent two adjacent convolutional layers, P(x) n W n ) indicates that the convolutional layer is processed in a W manner, which may include BN layers, etc., and can also be regarded as a residual block.

[0075] However, if direct mapping is introduced, the parameter passing method between convolutional layers becomes:

[0076] x n+1 =xn +P(x n W n );

[0077] Among them, the parameter x of the next layer n+1 The determination is made by x n It is obtained by weighting the parameters after the corresponding convolution operation, which can avoid the gradient vanishing problem as much as possible and improve the stability of the model.

[0078] In one embodiment, the initial classification network further includes a second cascaded operation layer connected to the global average pooling layer, an LSTM layer connected to the second cascaded operation layer, a feature fusion layer connected to the second cascaded operation layer, and a classification layer connected to the feature fusion layer, wherein:

[0079] The LSTM layer is used to perform feature extraction processing on the one-dimensional data of the corresponding Fourier sequence input to obtain temporal features.

[0080] It should be noted that the time-series features extracted based on the LSTM layer represent the feature vector of the entire sequence. The reason for using the LSTM layer to extract features from the Fourier sequence is that the model only needs to fuse features from one type of time series. Through practical analysis, by using the LSTM layer to extract features from the one-dimensional data corresponding to the original time series, the one-dimensional data corresponding to the first-order difference sequence, and the one-dimensional data corresponding to the Fourier sequence, it has been verified that the one-dimensional data based on the corresponding Fourier sequence can extract sequence features better, thus solving the problem of poor classification performance of traditional RNN models for the entire time series.

[0081] The second cascaded operation layer is used to cascade the temporal features and the feature terms output by the global average pooling layer to obtain the target cascaded features.

[0082] For details, please refer to Figure 4 The second cascaded operation layer is Figure 4 The connection shown is the second Concatenate layer after the GloabaAveragePooling operation layer. The specific process of the cascading operation has been explained earlier and will not be repeated here.

[0083] The feature fusion layer is used to extract deep features based on the target cascaded features to obtain deep features.

[0084] For details, please refer to Figure 4 The feature fusion layer is Figure 4The connection shown is in the Feature integration layer after the second Concatenate layer.

[0085] The classification layer is used to output corresponding food classification results based on the depth features.

[0086] For details, please refer to Figure 4 The classification layer is Figure 4 The diagram illustrates a connection following the Feature integration layer and then the SoftMax layer. In practice, the classification categories are set based on the actual classification results of the time series data within the network, which determines the number of nodes in the SoftMax layer. Ultimately, the SoftMax layer performs food classification to obtain the corresponding results.

[0087] Taking beef classification based on the above embodiments as an example, the entire operation process can be referred to... Figure 5 To understand. From Figure 5 As is known, the deep learning-based beef classification method disclosed in this application first acquires a beef dataset (specifically, it uses a food spectrometer to obtain spectral data from raw beef samples and beef samples cooked using two different cooking methods, establishing a beef dataset including pure and adulterated (heart, tripe, kidney, and liver) sample types). The beef dataset is then used as the original sequence, which undergoes first-order difference sequence transformation and Fourier sequence transformation to obtain the corresponding difference sequence and Fourier sequence. Next, the original sequence, difference sequence, and Fourier sequence obtained in the previous step are processed into two dimensions. The results of these two-dimensional transformations are then input into a multi-branch convolutional neural network for preliminary extraction of target features. Subsequently, the target features are aggregated and fused with the temporal features obtained through LSTM layer processing to obtain deep features. Finally, the classification result is output based on these deep features.

[0088] In the experimental verification, five true classes were set, and the sequence length was 470. Regarding model parameters, the learning rate of the network model was set to 0.001. Furthermore, since spatial and channel attention mechanisms were introduced in the convolutional layers to improve feature prediction accuracy, the ratio of spatial attention to channel attention was set to 1:1 to avoid result bias. Finally, the number of iterations was set to 1500, the number of units in the LSTM hidden layer and the output embedding layer were set to 100, and the parameters of the fully connected neural network in the feature fusion layer were set to 64.

[0089] During the verification process, this application also selected 10 existing network models, including Encoder, Mcdcnn, MLP, FCN, CNN, and ResNet, to conduct experiments and comparisons with the model in this application. The results are shown in Table 1 below:

[0090] Table 1 Experimental Results and Comparison

[0091]

[0092] As shown in Table 1, the network model constructed in this application has significant advantages in revealing the spectral differences between different beef samples. It can better capture the local temporal features and global correlation features within and between single variables, and is more conducive to mining the feature patterns of beef spectral data.

[0093] Please refer to Figure 6 This application discloses a deep learning-based food classification system 600, which includes a data acquisition module 601, a sequence conversion module 602, a two-dimensionalization module 603, a model training module 604, and a food classification module 605, wherein:

[0094] The data acquisition module 601 is used to acquire a training dataset, which includes a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type.

[0095] The sequence conversion module 602 is used to take the training dataset as the original time series and convert the original time series to obtain a first-order difference sequence and a Fourier sequence.

[0096] The two-dimensionalization module 603 is used to perform two-dimensionalization processing on each sequence using a tiling and sliding window mode to obtain the corresponding two-dimensional data.

[0097] The model training module 604 is used to construct an initial classification model, input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training, and obtain the target classification model when the training terminates.

[0098] The food classification module 605 is used to input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

[0099] As can be seen from the above, the food classification system based on deep learning disclosed in this application transforms the extracted one-dimensional time series into two-dimensional data. This two-dimensional data is then input into a convolutional neural network for feature classification, fully leveraging the advantages of convolutional neural networks in two-dimensional data feature extraction. This addresses the problem of traditional RNN models performing poorly in classifying entire time series, thus improving the accuracy of food classification. In summary, this application demonstrates superior accuracy compared to most current classical machine learning and deep learning methods. It can advance existing research on artificial intelligence for food classification management, thereby improving the regulatory level of the food industry and ensuring consumers receive safe and high-quality food.

[0100] In one embodiment, the sequence conversion module 602 is further configured to sequentially obtain the values ​​of two adjacent items in the original time series, and perform a difference calculation on the values ​​of the two adjacent items to obtain the corresponding difference values; and combine the obtained difference values ​​in order to obtain a first-order difference sequence.

[0101] In one embodiment, the sequence conversion module 602 is further configured to perform a Fourier transform on the original time series to obtain a Fourier sequence.

[0102] In one embodiment, the two-dimensionalization module 603 is further configured to obtain the total sequence length and the sequence-preserving feature length of the corresponding one-dimensional sequence; determine the window sliding step size based on the total sequence length; determine the sliding window value based on the sum of the window sliding step size and the sequence-preserving feature length; determine the number of rows of the two-dimensional data based on the window sliding step size; determine the number of columns of the two-dimensional data based on the sliding window value; and convert the corresponding one-dimensional sequence into corresponding two-dimensional data based on the number of rows and columns of the two-dimensional data.

[0103] In one embodiment, the initial classification model includes a first convolutional layer, a first cascaded operation layer, a second convolutional layer, and a global average pooling layer connected in sequence, wherein: the first convolutional layer is used to extract features from each of the obtained two-dimensional data to obtain multiple initial feature terms; the first cascaded operation layer is used to perform cascaded processing on the extracted initial feature terms to obtain cascaded features; the second convolutional layer is used to analyze the correlation between the initial feature terms based on the cascaded features to perform feature concatenation; and the global average pooling layer is used to perform dimensionality reduction processing on the obtained concatenated features to obtain target feature terms.

[0104] In one embodiment, both the first convolutional layer and the second convolutional layer employ a combination of spatial attention and channel attention mechanisms to extract attention weights for both spatial and channel attention dimensions. These attention weights are then weighted in a weighted manner to improve the accuracy of feature prediction.

[0105] In one embodiment, a residual mechanism is introduced in both the first and second convolutional layers to avoid gradient vanishing and improve model stability.

[0106] In one embodiment, the initial classification network further includes a second cascaded operation layer connected to the global average pooling layer, an LSTM layer connected to the second cascaded operation layer, a feature fusion layer connected to the second cascaded operation layer, and a classification layer connected to the feature fusion layer, wherein: the LSTM layer is used to perform feature extraction processing on the one-dimensional data of the corresponding Fourier sequence of the input to obtain temporal features; the second cascaded operation layer is used to perform cascaded processing on the temporal features and the feature terms output by the global average pooling layer to obtain target cascaded features; the feature fusion layer is used to perform deep feature extraction based on the target cascaded features to obtain deep features; and the classification layer is used to output the corresponding food classification result based on the deep features.

[0107] This application provides a readable storage medium, wherein when the computer program is executed by a processor, it performs the method in any optional implementation of the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0108] The aforementioned readable storage medium transforms the extracted one-dimensional time series into two-dimensional data. This two-dimensional data is then input into a convolutional neural network (RNN) for feature classification, fully leveraging the advantages of RNNs in two-dimensional data feature extraction. This addresses the problem of traditional RNN models performing poorly in classifying entire time series, thus improving the accuracy of food classification. In summary, this application demonstrates superior accuracy compared to most current classical machine learning and deep learning methods, and can advance existing artificial intelligence research in food classification management, thereby improving the regulatory level of the food industry and ensuring consumers receive safe and high-quality food.

[0109] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0110] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0111] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0112] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0113] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A food classification method based on deep learning, characterized in that, Includes the following steps: S1. Obtain a training dataset, which includes a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type. S2. The training dataset is used as the original time series, and the original time series is transformed to obtain the first-order difference sequence and the Fourier sequence. S3. Each sequence is processed into two dimensions using a tiling and sliding window method to obtain the corresponding two-dimensional data; S4. Construct an initial classification model, and input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training. When the training terminates, the target classification model is obtained, wherein: The initial classification model comprises a first convolutional layer, a first cascaded operation layer, a second convolutional layer, and a global average pooling layer connected in sequence, wherein: The first convolutional layer is used to extract features from each of the obtained two-dimensional data to obtain multiple initial feature terms; The first cascaded operation layer is used to perform cascaded processing on the extracted initial feature terms to obtain cascaded features; The second convolutional layer is used to analyze the correlation between each initial feature item based on the cascaded features in order to perform feature concatenation; The global average pooling layer is used to reduce the dimensionality of the obtained concatenated features to obtain the target feature term. Both the first and second convolutional layers employ a combination of spatial attention and channel attention mechanisms, extracting attention weights for both dimensions and then weighting them to improve feature prediction accuracy. Both the first and second convolutional layers incorporate residual mechanisms to prevent gradient vanishing and improve model stability. The initial classification model further includes a second cascaded operation layer connected to the global average pooling layer, an LSTM layer connected to the second cascaded operation layer, a feature fusion layer connected to the second cascaded operation layer, and a classification layer connected to the feature fusion layer, wherein: The LSTM layer is used to perform feature extraction processing on the one-dimensional data of the corresponding Fourier sequence input to obtain temporal features; The second cascaded operation layer is used to cascade the temporal features and the feature terms output by the global average pooling layer to obtain the target cascaded features. The feature fusion layer is used to extract deep features based on the target cascaded features to obtain deep features; The classification layer is used to output corresponding food classification results based on the depth features; S5. Input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

2. The method according to claim 1, characterized in that, In step S2, the transformation of the original time series to obtain a first-order difference sequence includes: sequentially obtaining the values ​​of two adjacent items in the original time series, and calculating the difference between the two adjacent items to obtain the corresponding difference value; and combining the obtained difference values ​​in order to obtain a first-order difference sequence. In step S2, the transformation of the original time series to obtain a Fourier sequence includes: performing a Fourier transform on the original time series to obtain a Fourier sequence.

3. The method according to claim 1, characterized in that, In step S3, the two-dimensional processing of each sequence using a tiling and sliding window method is performed to obtain the corresponding two-dimensional data, including: S31. Obtain the total length of the corresponding one-dimensional sequence and the length of the sequence features retained; S32. Determine the window sliding step size based on the total length of the sequence; S33. Determine the sliding window value based on the sum of the window sliding step size and the sequence retained feature length; S34. Determine the number of rows of the two-dimensional data based on the sliding step size of the window, determine the number of columns of the two-dimensional data based on the sliding window value, and convert the corresponding one-dimensional sequence into the corresponding two-dimensional data based on the number of rows and columns of the two-dimensional data.

4. A food classification system based on deep learning, characterized in that, The system includes a data acquisition module, a sequence conversion module, a two-dimensionalization module, a model training module, and a food classification module, wherein: The data acquisition module is used to acquire a training dataset, which includes a first time-series spectral signal corresponding to a pure sample type and a second time-series spectral signal corresponding to an adulterated sample type. The sequence conversion module is used to take the training dataset as the original time series and convert the original time series to obtain a first-order difference sequence and a Fourier sequence. The two-dimensionalization module is used to perform two-dimensionalization processing on each sequence using a tiling and sliding window mode to obtain the corresponding two-dimensional data; The model training module is used to construct an initial classification model, and input the obtained two-dimensional data and the Fourier sequence into the initial classification model for training. Upon termination of training, the target classification model is obtained, wherein: The initial classification model comprises a first convolutional layer, a first cascaded operation layer, a second convolutional layer, and a global average pooling layer connected in sequence, wherein: The first convolutional layer is used to extract features from each of the obtained two-dimensional data to obtain multiple initial feature terms; The first cascaded operation layer is used to perform cascaded processing on the extracted initial feature terms to obtain cascaded features; The second convolutional layer is used to analyze the correlation between each initial feature item based on the cascaded features in order to perform feature concatenation; The global average pooling layer is used to reduce the dimensionality of the obtained concatenated features to obtain the target feature term. Both the first and second convolutional layers employ a combination of spatial attention and channel attention mechanisms, extracting attention weights for both dimensions and then weighting them to improve feature prediction accuracy. Both the first and second convolutional layers incorporate residual mechanisms to prevent gradient vanishing and improve model stability. The initial classification model further includes a second cascaded operation layer connected to the global average pooling layer, an LSTM layer connected to the second cascaded operation layer, a feature fusion layer connected to the second cascaded operation layer, and a classification layer connected to the feature fusion layer, wherein: The LSTM layer is used to perform feature extraction processing on the one-dimensional data of the corresponding Fourier sequence input to obtain temporal features; The second cascaded operation layer is used to cascade the temporal features and the feature terms output by the global average pooling layer to obtain the target cascaded features. The feature fusion layer is used to extract deep features based on the target cascaded features to obtain deep features; The classification layer is used to output corresponding food classification results based on the depth features; The food classification module is used to input the acquired two-dimensional data to be processed and the Fourier sequence into the target classification model to obtain the food classification result.

5. A readable storage medium, characterized in that, The readable storage medium includes a deep learning-based food classification method program, which, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Near infrared spectrum classification model training method and system and classification method and system

    CN113378971A

  • Wireless channel scenario identification method and system

    US20210399817A1