A lightweight beef quality assessment method based on unsupervised feature learning

By comprehensively applying technologies such as generative adversarial networks, unsupervised feature learning models, MSCNN, lightweight Transformer models and adversarial sample generators, the problems of subjectivity and low computational efficiency of traditional beef quality assessment methods are solved, and high-precision, fast and lossless beef quality assessment is achieved.

CN119418330BActive Publication Date: 2025-05-06YUNNAN AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411448900.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-05-06
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

Traditional beef quality evaluation methods have disadvantages such as strong subjectivity, poor repetition, time-consuming and labor-consuming, destructive, and difficulty in real-time online monitoring. Although hyperspectral imaging technology has the advantages of non-contact, damage-free, fast and efficient, it faces challenges such as high-dimensionality, large data volume, complexity, noise interference, and the existing methods rely on prior knowledge, difficulty in selecting parameters, insufficient generalization capabilities, and low computing efficiency.

Method used

A lightweight beef quality evaluation method for unsupervised feature learning is proposed. By comprehensively applying network models such as generative adversarial network model, unsupervised feature learning model, deep learning model MSCNN, lightweight Transformer model and adversarial sample generator, preprocessing, feature extraction, dimensionality reduction and classification of hyperspectral data is carried out to achieve high-precision evaluation of beef quality.

Benefits of technology

It improves the accuracy, objectivity, speed, losslessness and comprehensiveness of beef quality evaluation, reduces computational complexity and memory consumption, enhances the stability and security of the model, and adapts to the nonlinear and non-Gaussian distribution characteristics of hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418330B_ABST
    Figure CN119418330B_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of quality assessment, and discloses a lightweight beef quality assessment method based on unsupervised feature learning, the method comprising: obtaining original hyperspectral data of beef as real data, and generating high-quality sample data based on a generative adversarial network model; preprocessing the high-quality sample data, mapping, reconstructing and fusing the preprocessed sample data, extracting spectral and spatial features at different levels, and fusing or splicing the eigenvalues ​​of other auxiliary data; mapping and fusing the fused sample features, generating sample data with specific small disturbances similar to the real data, and alternately inputting the sample data with the real data into a Transformer model, and updating the parameters of the Transformer model; and evaluating the beef quality based on the Transformer model with updated parameters. The invention improves the accuracy, objectivity, rapidity, non-destructiveness and comprehensiveness of beef quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of quality assessment, and more particularly to a lightweight beef quality assessment method based on unsupervised feature learning. Background Art

[0002] Beef is one of the important sources of animal protein for humans, and its quality directly affects the health and satisfaction of consumers. However, the quality of beef is affected by many factors, such as the breed, age, feeding method, slaughter conditions, storage and processing of cattle, which leads to differences in the color, tenderness, flavor, juiciness, nutritional value and other aspects of beef. Therefore, how to accurately evaluate the quality of beef and improve its quality and safety is an important issue in the field of beef production and processing.

[0003] Traditional beef quality assessment methods mainly rely on manual sensory evaluation or chemical analysis, which have the disadvantages of strong subjectivity, poor repeatability, time-consuming and labor-intensive, high destructiveness, and difficulty in real-time online monitoring. In recent years, beef quality assessment methods based on hyperspectral imaging technology have gradually attracted people's attention. Hyperspectral imaging technology is a non-contact, non-destructive, fast and efficient spectral analysis technology. It can obtain spatial information and spectral information of beef at the same time, and can reflect the intrinsic properties and external characteristics of beef. Using hyperspectral imaging technology, beef can be comprehensively, meticulously and objectively evaluated for quality, providing an important basis for the classification, grading, traceability and detection of beef. However, hyperspectral imaging technology also faces challenges such as high dimensionality, large data volume, complexity and noise interference of hyperspectral data, which brings difficulties to beef quality assessment. In order to solve the above problems, it is necessary to effectively preprocess, extract features, reduce dimensionality and classify hyperspectral data to extract useful information of beef, eliminate irrelevant information, reduce the complexity of data and improve the accuracy of classification. At present, some researchers have tried to use machine learning, pattern recognition, image processing and other methods to process and analyze hyperspectral data, and have achieved certain results. However, the above methods also have limitations such as reliance on prior knowledge, difficulty in parameter selection, insufficient generalization ability, and low computational efficiency.

[0004] In recent years, deep learning (DL), as a powerful machine learning method, has made breakthrough progress in computer vision, natural language processing, speech recognition and other fields. Deep learning has the advantages of automatically learning data features, expressing data abstraction, and mining data potential laws, providing new ideas and methods for the processing and analysis of hyperspectral data. Deep learning automatically extracts high-level spectral and spatial features from hyperspectral data, which can effectively reduce the dimension and redundancy of data and improve the expression and identification capabilities of data. Using a multi-layer neural network structure to achieve end-to-end data mapping and classification can reduce computational complexity and memory consumption, and improve classification performance and efficiency. Deep learning has achieved remarkable results in hyperspectral image classification, target detection, change detection and other aspects, expanding new areas for the application of hyperspectral images.

[0005] Traditional hyperspectral data processing and analysis methods mainly rely on machine learning, pattern recognition, image processing and other methods. These methods have certain limitations: (1) They rely on prior knowledge and require manual design and selection of appropriate feature extraction and dimensionality reduction methods such as principal component analysis (PCA), linear discriminant analysis (LDA), and independent component analysis (ICA). They are difficult to fully utilize the spectral and spatial information of hyperspectral data and are difficult to adapt to different data scenarios and task requirements. (2) Parameter selection is difficult. They require manual adjustment and optimization of different parameters such as feature dimensions, classifier types, and classifier parameters, which will affect the data processing and analysis results and increase the calculation time and cost. (3) The generalization ability is insufficient. Typical support vector machines (SVM), random forests (RF), and K-nearest neighbors (KNN) require a large amount of labeled data for training and verification. These methods may be difficult to adapt to data changes and noise, and are also difficult to handle data deficiencies and imbalances. (4) The computational efficiency is low. Multiple steps such as feature extraction, dimensionality reduction, and classification are required for data processing and analysis. The redundancy and incompatibility of these steps and modules increase the complexity of the calculation and memory consumption. Summary of the invention

[0006] In view of this, the purpose of the present invention is to propose a lightweight beef quality assessment method based on unsupervised feature learning, which improves the accuracy, objectivity, rapidity, non-destructiveness and comprehensiveness of beef quality assessment by comprehensively using network models such as generative adversarial network model, unsupervised feature learning model, deep learning model MSCNN, lightweight Transformer model and adversarial sample generator.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A lightweight beef quality assessment method based on unsupervised feature learning, including the following methods:

[0009] S1. Obtaining original hyperspectral data of beef as real data, and generating high-quality sample data based on a generative adversarial network model, wherein the generative adversarial network model includes a generator and a discriminator;

[0010] S2. Determine the dynamic range of the high-quality sample data, and determine the corresponding threshold value according to the distribution characteristics of the high-quality sample data within the dynamic range;

[0011] Compressing the dynamic range of high-quality sample data at each wavelength to a corresponding threshold, and performing preprocessing to obtain preprocessed sample data;

[0012] S3. Based on the unsupervised feature learning model, the preprocessed sample data is mapped to a low-dimensional latent space and reconstructed into the original sample data, and the most relevant feature information in the original sample data is fused to obtain the fused original sample data;

[0013] S4, based on the deep learning model MSCNN, extract different levels of spectral and spatial features from the fused original sample data at different scales to obtain the characteristic values ​​of the original sample data;

[0014] Based on a multi-scale and multi-modal feature fusion network, the feature values ​​of the original sample data and the feature values ​​of other auxiliary data are fused or spliced ​​to obtain feature values ​​of different data sources, and the most useful feature values ​​are selected from the feature values ​​of different data sources for fusion to obtain fused sample features;

[0015] S5. Based on a lightweight Transformer model, the fused sample features are mapped to a high-dimensional latent space to obtain the fused sample features in the latent space, and the fused sample features in the latent space are mapped to the category label to obtain the probability of the category label of the fused sample features; the lightweight Transformer model includes an encoder and a classifier;

[0016] S6. Inputting the probability of the category label of the fused sample feature and the real data into an adversarial sample generator, wherein the adversarial sample generator includes: a perturbation generator and a perturbation evaluator;

[0017] Alternately inputting the sample data with slight perturbations generated by the perturbation generator and the real data into the lightweight Transformer model, performing forward propagation and back propagation, and updating the parameters of the lightweight Transformer model;

[0018] Beef quality is evaluated based on the lightweight Transformer model with updated parameters.

[0019] Furthermore, in step S1, the original hyperspectral data of beef is obtained as real data, and high-quality sample data is generated based on a generative adversarial network model, specifically: the generative adversarial network includes a generator and a discriminator, and the loss functions of the generator and the discriminator are calculated respectively; the original gradients of the generator and the discriminator are clipped using an adaptive gradient clipping method; and the above process is repeated until the generator and the discriminator reach a balance, thereby generating high-quality sample data;

[0020] The loss functions of the generator and the discriminator are calculated separately, specifically: the Wasserstein distance is set to define the loss function of the discriminator, and the expression is:

[0021]

[0022] Where: L D represents the loss function of the discriminator, Indicates that in the real data distribution P r expectations, Indicates that in generating data distribution P g The expectation on , x is the real data, is the data generated by the generator, P r is the distribution of real data, P g is the distribution of generated data, and D is the discriminator;

[0023] The generator is a reverse convolutional neural network, the dimension of the input layer of the reverse convolutional neural network is the dimension m of random noise, and the dimension of the output layer of the reverse convolutional neural network is the same as the dimension of the input layer of the discriminator, which is n;

[0024] The loss function expression of the generator is:

[0025]

[0026] Where z is random noise, P z is the distribution of random noise, G is the generator;

[0027] The calculation method of the adaptive gradient clipping is:

[0028]

[0029] In the above formula, is the clipped gradient, g is the original gradient of the generator and discriminator, and c is the clipping threshold, where c = 0.01.

[0030] Furthermore, in step S2, the dynamic range of the high-quality sample data is determined, and a corresponding threshold is determined according to the distribution characteristics of the high-quality sample data within the dynamic range;

[0031] The dynamic range of the high-quality sample data at each wavelength is compressed to a corresponding threshold value, and preprocessed to obtain the preprocessed sample data, specifically: calculating the maximum value and the minimum value of the high-quality sample data to determine the dynamic range of the high-quality sample data; determining the corresponding threshold value according to the distribution characteristics of the high-quality sample data in the dynamic range, compressing the dynamic range of the high-quality sample data at each wavelength to the corresponding threshold value, and normalizing the compressed high-quality sample data to obtain the preprocessed sample data;

[0032] The maximum and minimum values ​​of the high-quality sample data are calculated to determine the dynamic range of the high-quality sample data. The expression is:

[0033] The high-quality sample dataset X∈R n×d , where n is the number of high-quality samples and d is the number of hyperspectral wavelengths;

[0034] The maximum value of the high-quality sample data is:

[0035]

[0036] Among them, X ij is the reflectance of the i-th sample at the j-th wavelength;

[0037] The minimum value of the high-quality sample data is:

[0038]

[0039] The dynamic range of determining high-quality sample data is:

[0040] R=X max -X min ;

[0041] Determine the corresponding threshold value based on the distribution characteristics of high-quality sample data within the dynamic range, specifically: calculate the mean and standard deviation of each sample data at each wavelength:

[0042] The expression of the mean is:

[0043]

[0044] Among them, μ j represents the mean value at the jth wavelength;

[0045] The expression of the standard deviation is:

[0046]

[0047] Among them, σ jrepresents the standard deviation at the jth wavelength;

[0048] The threshold is: [μ j -3σ j , μ j +3σ j ];

[0049] Compress the dynamic range of high-quality sample data at each wavelength to [μ-3σ, μ+3σ] to obtain compressed sample data:

[0050]

[0051] Furthermore, the compressed high-quality sample data are normalized to obtain the preprocessed sample data, specifically:

[0052] Using linear normalization, the expression is:

[0053]

[0054] Among them, X' max is the maximum value of the compressed sample data, X' min is the minimum value of the compressed data;

[0055]

[0056] X” ij is the reflectance of the i-th sample at the j-th wavelength after normalization;

[0057] The preprocessed sample data is X”∈R n×d , where n is the number of samples and d is the number of wavelengths.

[0058] Furthermore, in step S3, the preprocessed sample data is mapped to a low-dimensional latent space and reconstructed into original sample data, and the most relevant feature information in the original sample data is fused to obtain the fused original sample data, specifically:

[0059] The unsupervised feature learning model is an improved diffusion model, and the improved diffusion model includes an encoder and a decoder, specifically: the encoder is used to map the preprocessed sample data to a low-dimensional latent space, and the decoder is used to reconstruct the preprocessed sample data in the latent space into the original sample data;

[0060] Using the feature information of the original sample data at different time steps, a feature library and a dynamic feature fusion module are constructed. The feature library is used to store the feature information of the original sample data at different time steps. The dynamic feature fusion is used to adaptively learn the feature information of the original sample data at different time steps. According to the feature information of the original sample data at the current time step, the feature information of the most relevant original sample data is selected for fusion.

[0061] Among them, the preprocessed hyperspectral data is X″∈R n×d , and input into the encoder, the encoder output is:

[0062] Z=E(X″)∈R n×k

[0063] In the above formula, E is the encoder, k is the dimension of the latent space, Z is the output matrix after encoding, and Z ij is the value of the jth dimension of the i-th sample in the latent space;

[0064] The decoder is used to reconstruct the preprocessed sample data in the latent space into the original sample data, specifically:

[0065] The output of the decoder is:

[0066]

[0067] In the above formula, D is the decoder, is the reconstructed reflectance of the i-th sample at the j-th wavelength;

[0068] Each convolution layer in the encoder adopts a residual connection, and the last layer of the encoder is a self-attention layer; each reverse convolution layer in the decoder adopts a residual connection, and the loss function of the decoder is an adaptive reconstruction loss;

[0069] Each convolutional layer in the encoder uses a residual connection, expressed as:

[0070] H l =F l (H l-1 )+H l-1

[0071] Among them, H l is the output of the lth layer, F l is the convolution operation of the lth layer, H l-1 is the output of the l-1th layer;

[0072] The last layer of the encoder is the self-attention layer, which is expressed as:

[0073] Z = softmax(H L WQ (H L W K ) T )H L W V

[0074] Among them, H L is the output of the last layer, W Q , W K , W V is the parameter of the attention layer, and softmax is the normalization function;

[0075] The loss function of the decoder is the adaptive reconstruction loss, specifically:

[0076]

[0077] Among them, w ij is the reconstruction weight of the i-th sample at the j-th wavelength, which is calculated as:

[0078]

[0079] in, is the variance of the jth wavelength, which is calculated as:

[0080]

[0081] Among them, μ j is the mean value of the jth wavelength, which is calculated as:

[0082]

[0083] The constructing of the feature library is specifically as follows:

[0084] Assume that the current time step is t, then the feature library is a three-dimensional tensor with a shape of T×n×k;

[0085] Where: T is the total number of time steps, n is the number of samples, k is the dimension of the latent space, and the tth slice of the feature library is Z t ∈R n×k , which is the feature of the current time step;

[0086] The input of the dynamic feature fusion module is the feature library, and the output is a two-dimensional matrix with a shape of n×k. The calculation method is:

[0087] Z′ t =softmax(Z t W Q (Z t W K ) T )Z t WV +softmax(Z t W Q (Z 1:t-1 W K ) T )Z 1:t-1 W V

[0088] In the above formula, Z′ t is the fused feature, Z 1:t-1 is the first t-1 slices of the feature library, that is, the features of the historical time step, W Q , W K , W V is the parameter of the dynamic feature fusion module, and softmax is the normalization function;

[0089] The original sample data after fusion is: Z′ t ∈R n×k , where n is the number of samples and k is the dimension of the latent space.

[0090] Furthermore, in step S4, based on the deep learning model MSCNN, spectral and spatial features of different levels are extracted from the fused original sample data at different scales to obtain the characteristic values ​​of the original sample data, specifically:

[0091] The characteristic value of the original sample data is Z′ t ∈R n×k , Z′ t,ij is the eigenvalue of the i-th sample in the j-th latent space dimension;

[0092] Then the characteristic value of the original sample data is:

[0093] F t =M(Z′ t )∈R n×l

[0094] Among them, M is the deep learning model MSCNN, l is the dimension of the features of the original sample data, and F t,ij is the feature value of the i-th sample in the j-th feature dimension;

[0095] The corresponding scale of each convolution layer and pooling layer in the deep learning model MSCNN is expressed as:

[0096] H t,0 =Z′ t

[0097] H t,i =P i (C i (H t,i-1 ))

[0098] F t =H t,s

[0099] Among them, H t,i is the output of the i-th scale, C i is the i-th convolutional layer, P i is the i-th pooling layer, and s is the total number of scales.

[0100] Furthermore, in step S4, based on a multi-scale and multi-modal feature fusion network, the feature values ​​of the original sample data and the feature values ​​of other auxiliary data are fused or concatenated to obtain feature values ​​of different data sources, and the most useful features are selected from the feature values ​​of different data sources to obtain fused sample features, specifically:

[0101] The eigenvalue of the original sample data is F t ∈R n×l ;

[0102] Where n is the number of samples, l is the dimension of the feature, and F t,ij is the feature value of the i-th sample in the j-th feature dimension;

[0103] Other auxiliary data is A t ∈R n×m ;

[0104] Among them, m is the dimension of auxiliary data, A t,ij is the eigenvalue of the i-th sample in the j-th auxiliary data dimension;

[0105] Get the characteristic values ​​of different data sources, the expression is:

[0106] G t =S(F t , A t )∈R n×p

[0107] Among them, S is the deep learning model MMFN, p is the dimension of the final feature, G t,ij is the feature value of the i-th sample in the j-th final feature dimension;

[0108] The deep learning model MSCNN includes multiple convolutional layers and pooling layers;

[0109] The multi-scale and multi-modal feature fusion network includes a feature fusion layer and a feature selection layer, wherein the feature fusion layer is used to fuse feature information from different data sources, and the feature selection layer is used to select useful feature information from the fused feature information;

[0110] The multi-scale and multi-modal feature fusion network includes a feature fusion layer and a feature selection layer, and the expression is:

[0111] B t =F(F t , A t )

[0112] G t =S(B t )

[0113] Among them, B t is the output of the feature fusion layer, F is the feature fusion layer, and S is the feature selection layer;

[0114] The feature fusion layer is used to fuse the features of different data sources, specifically by using a weighted average or concatenation method, and the expression is:

[0115] B t =αF t +(1-α)A t

[0116] or

[0117] B t =[F t ; A t ]

[0118] Among them, α is the fusion weight, [;] is the splicing operation;

[0119] The feature selection layer is used to select useful feature information from the fused feature information, specifically by using sparse coding or principal component analysis, the expression is:

[0120] G t =B t W

[0121] or

[0122] G t = PCA(B t )

[0123] Among them, W is the parameter of sparse coding, and PCA is the function of principal component analysis.

[0124] Furthermore, in step S5, based on the lightweight Transformer model, the fused sample features are mapped to a high-dimensional latent space to obtain the fused sample features in the latent space, and the fused sample features in the latent space are mapped to the category labels to obtain the probability of the category labels of the fused sample features, specifically:

[0125] The lightweight Transformer model includes an encoder and a classifier, wherein the encoder is used to map the fused sample features to a high-dimensional latent space, and the classifier is used to map the fused sample features in the latent space to a category label and output the probability of the category label;

[0126] The encoder is a hybrid attention network, including several hybrid attention layers, each of which includes a global self-attention sublayer and a local convolutional attention sublayer, the global self-attention layer is used to capture the long-distance dependency between feature information, and the local convolutional attention layer is used to capture the short-distance correlation between feature information;

[0127] Among them, the output of the multi-scale and multi-modal feature fusion is G t ∈R n×p ;

[0128] Among them, n is the number of samples, p is the dimension of the final feature, G t,ij is the feature value of the i-th sample in the j-th final feature dimension;

[0129] Then the output of the encoder is:

[0130] H t =E(G t )∈R n×q

[0131] Among them, E is the encoder, q is the dimension of the latent space, and H t,ij is the eigenvalue of the jth dimension of the i-th sample in the latent space;

[0132] The output of the classifier is:

[0133] Y t =C(H t )∈R n×c

[0134] Among them, C is the classifier, c is the number of categories, and Y t,ij is the probability that the i-th sample belongs to the j-th category;

[0135] In step S5, each of the hybrid attention layers includes a global self-attention sublayer and a local convolutional attention sublayer, and the expression is:

[0136] H t,0 =G t

[0137] H t,i =LayerNorm(H t,i-1 +GSA(H t,i-1 ))+LayerNorm(H t,i-1 +LCA(Ht,i-1 ))

[0138] H t =H t,r

[0139] Among them, H t,i is the output of the i-th hybrid attention layer, LayerNorm is the layer normalization function, and r is the total number of hybrid attention layers;

[0140] In step S5, the global self-attention layer is used to capture the long-distance dependencies between features, and the expression is:

[0141] GSA(H t,i-1 )=softmax(H t,i-1 W Q (H t,i-1 W K ) T )H t,i-1 W V

[0142] Among them, W Q , W K , W V is the parameter of the global self-attention sublayer, and softmax is the normalization function;

[0143] The role of the local convolutional attention sublayer is to capture the short-range correlation between features, and its formula is:

[0144] LCA(H t,i-1 )=softmax(H t,i-1 W Q (Conv(H t,i-1 , W K )) T )Conv(H t,i-1 , W V )

[0145] Among them, W Q , W K , W V are the parameters of the local convolutional attention sublayer, softmax is the normalization function, and Conv is the convolution operation;

[0146] The probability of obtaining the category label of the fused sample features is specifically:

[0147] The lightweight Transformer model outputs a probability distribution Y t ∈R n×c , where n is the number of samples and c is the number of beef quality grades.

[0148] Furthermore, in step S6, the probability of the category label of the fused sample feature and the real data are input into the adversarial sample generator, specifically:

[0149] The output of the lightweight Transformer model is: t ∈R n×c ;

[0150] Where n is the number of samples, c is the number of categories, and Y t,ij is the probability that the i-th sample belongs to the j-th category;

[0151] The real data is X∈R n×d , where d is the wavelength, X ij is the reflectance of the i-th sample at the j-th wavelength;

[0152] The perturbation generator and the perturbation evaluator are used to perform adversarial training, so that the perturbation generator generates sample data with specific small perturbations similar to the real data, specifically: the perturbation generator is used to generate sample data with small perturbations similar to the real data; the perturbation evaluator in the adversarial sample generator is used to output a perturbation score of the difference between the real data and the perturbed sample data;

[0153] Then the output of the disturbance generator is:

[0154]

[0155] Where G is the disturbance generator, is the disturbed reflectivity of the i-th sample at the j-th wavelength;

[0156] The output of the perturbation evaluator is:

[0157]

[0158] Where E is the disturbance estimator, S i Score the perturbation of the i-th sample;

[0159] The perturbation evaluator in the adversarial sample generator is used to output a perturbation score of the difference between the real data and the perturbed sample data, expressed as: The loss function of the perturbation estimator is:

[0160]

[0161] Among them, |·| is the norm function.

[0162] Furthermore, in step S6, the sample data with slight perturbations and the real data generated by the perturbation generator are alternately input into the lightweight Transformer model, and forward propagation and backward propagation are performed to update the parameters of the lightweight Transformer model; specifically, for each training batch, half of the sample data with slight perturbations and the real data are randomly selected, and the corresponding perturbation data is generated by the perturbation generator, and the perturbation data and the original data are spliced ​​into a new input, while the category label is kept unchanged, and then the new input and the category label are input into the lightweight Transformer model, and forward propagation and backward propagation are performed to update the parameters of the model;

[0163] The loss function of the perturbation generator and the perturbation evaluator for adversarial training is:

[0164]

[0165] Among them, [;] is the concatenation operation, E is the encoder, and C is the classifier.

[0166] According to the specific embodiments provided by the present invention, the present invention has the following technical effects: the spectral and spatial features of beef are automatically extracted and integrated from the hyperspectral image using a deep learning method, thereby achieving high-precision classification of beef quality. The method of the present invention can adapt to the nonlinear and non-Gaussian distribution characteristics of hyperspectral images, resist potential attacks and interference, and improve the stability and security of classification. Compared with traditional beef quality evaluation methods such as sensory evaluation, physical and chemical evaluation, and near-infrared spectroscopy analysis, the method of the present invention has the advantages of strong objectivity, good repeatability, and easy standardization. It can effectively reduce human errors and deviations and improve the accuracy and credibility of beef quality evaluation.

[0167] The present invention uses an end-to-end method to directly obtain the classification results of beef quality from hyperspectral images, without the need for complex data preprocessing, feature selection, model training and other steps, greatly shortening the time and cost of beef quality assessment. The use of a lightweight Transformer model, combined with a global self-attention mechanism and a local convolutional attention mechanism, can reduce computational complexity and memory consumption, and improve the speed and efficiency of classification. Compared with traditional beef quality assessment methods such as machine vision, computer image processing, and bioelectrical impedance analysis, the method of the present invention has the characteristics of non-destructiveness, rapidity, efficiency, and intelligence, and can meet the needs of large-scale, real-time, and online beef quality assessment, thereby improving the efficiency and quality of beef production, processing, and consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0168] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0169] The following is a further description of the AC and DC side coordinated control system and method of the grid-type energy storage system of the present invention in conjunction with the accompanying drawings;

[0170] Figure 1 It is an overall flow chart of the lightweight beef quality assessment method based on unsupervised feature learning provided by the present invention. DETAILED DESCRIPTION

[0171] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0172] In order to better understand the purpose, structure and function of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.

[0173] like Figure 1 As shown, the present invention proposes an "end-to-end" beef quality assessment method based on deep learning for the hyperspectral image classification task. The main steps are as follows: (1) Use WGAN-AGC to generate more high-quality beef samples to enhance the scale and quality of the data set. (2) Use the dynamic range compression method to preprocess the real data to reduce the noise and outliers of the data. (3) Use the improved diffusion model to automatically learn the spectral and spatial features from the preprocessed sample data, and use the feature library and dynamic feature fusion module to adaptively learn the information of multiple time steps. (4) Use MSCNN to extract spectral and spatial features of different levels from hyperspectral data of different scales, and use MMFN to fuse hyperspectral data and other auxiliary data to improve the expression and identification capabilities of the features. (5) Use a lightweight Transformer model to classify the fused features, combine the global self-attention mechanism and the local convolutional attention mechanism to reduce the computational complexity and memory consumption while maintaining a high classification performance. (6) ASG is used to generate some hyperspectral data with slight disturbances during the training process, so that the model can resist some potential attacks and interferences, and improve the stability and security of the model; and an embodiment of the lightweight beef quality assessment method based on unsupervised feature learning is provided, which specifically includes the following steps:

[0174] S1: Data Augmentation

[0175] S1.1: The original hyperspectral data is taken as real data and input into the discriminator. The goal of the discriminator is to distinguish between real data and generated data.

[0176] The structure of the discriminator is a multi-layer perceptron (MLP), whose input layer has a dimension of n, where n is the number of wavelengths of the hyperspectral data, and whose output layer has a dimension of 1, which represents the probability of the real data.

[0177] S1.2: Random noise is used as input to the generator, whose goal is to generate data similar to real data while deceiving the discriminator.

[0178] The structure of the generator is a Deconvolutional Neural Network (DeCNN), whose input layer has a dimension of m, where m is the dimension of random noise, and whose output layer has a dimension of n, which is the same as the input layer dimension of the discriminator.

[0179] S1.3: Adaptive gradient clipping is used to limit the gradients of the generator and discriminator, thus avoiding gradient explosion or disappearance and improving the efficiency and stability of training.

[0180] The criterion for balancing the generator and the discriminator is that the Wasserstein distance tends to be stable. In this step, the stability threshold is taken as ∈=0.001.

[0181] It should be noted that the Wasserstein distance tends to be stable, which can be defined as: within N consecutive training cycles (epochs), the change amplitude of the Wasserstein distance is less than a given threshold ε, including the steps:

[0182] 1. Set an observation window N, for example N = 10 training cycles.

[0183] 2. At the end of each training cycle, calculate and record the current Wasserstein distance.

[0184] 3. For the Wasserstein distance values ​​of the last N training cycles, calculate the difference between the maximum and minimum values.

[0185] 4. If this difference is less than the preset threshold ε (ε = 0.001), it is considered that the Wasserstein distance tends to be stable; the specific value needs to be adjusted according to the specific data set and training process to achieve the best effect.

[0186] Generate more high-quality beef samples based on WGAN-AGC, increase the scale and quality of the data set, improve the diversity and representativeness of the data, and improve the generalization ability of the model. Use Wasserstein distance as the loss function of the discriminator to avoid the problems of mode collapse and vanishing gradient, and improve the generation quality and diversity of the generator. Adopt the method of adaptive gradient clipping to avoid the problem of exploding gradient and improve the efficiency and stability of training.

[0187] S2: Data preprocessing

[0188] The dynamic range compression method is used to preprocess the hyperspectral data, reduce the noise and outliers of the data, improve the signal-to-noise ratio and contrast of the data, and reduce the redundancy and interference of the data.

[0189] S2.2: Determine an appropriate threshold based on the distribution characteristics of the data, compress the dynamic range of the data to within the threshold, and eliminate data that exceeds the threshold.

[0190] S2.3: The outlier detection method based on triple standard deviation is based on the assumption that the data follows a normal distribution. Then the data in the interval [μ-3σ, μ+3σ] accounts for 99.7% of the total, where μ is the mean and σ is the standard deviation. Therefore, this interval is used as the threshold, and the data outside this interval is regarded as an outlier and removed.

[0191] S2.4: Get the preprocessed hyperspectral data for subsequent feature extraction and classification. The preprocessed hyperspectral data is X”∈R n×d , the value of each element is between [0, 1], and it has a high signal-to-noise ratio and contrast, reducing data redundancy and interference.

[0192] The dynamic range compression method is used to preprocess the hyperspectral data to reduce the noise and outliers of the data, improve the signal-to-noise ratio and contrast of the data, and reduce the redundancy and interference of the data; the outlier detection method based on triple standard deviation is used to effectively remove outliers and noise points in the data and improve the quality and consistency of the data; the linear normalization method is used to map the data range to between [0, 1], which facilitates the subsequent feature extraction and classification, while avoiding the influence of the dimension and scale of the data.

[0193] S3: Unsupervised Feature Learning

[0194] The improved diffusion model is used to automatically learn spectral and spatial features from hyperspectral data. At the same time, the feature library and dynamic feature fusion module are used to adaptively learn information from multiple time steps, thereby improving the effectiveness and richness of features and capturing the complex spectral and spatial relationships of the data. The specific implementation of this step is as follows:

[0195] S3.1: The preprocessed hyperspectral data is taken as input and input into the improved diffusion model. The improved diffusion model consists of an encoder and a decoder. The goal of the encoder is to map the input data to a low-dimensional latent space, and the goal of the decoder is to reconstruct the data in the latent space into the original data.

[0196] S3.2: In order to solve the problems of gradient vanishing and fuzzy reconstruction in the original version of the diffusion model, the encoder and decoder are improved. The encoder adopts residual connection and attention mechanism to enhance the transmission and fusion of information, and the decoder adopts adaptive reconstruction loss to dynamically adjust the reconstruction weight according to the complexity and importance of the data.

[0197] Specifically, the structure of the encoder is a deep convolutional neural network (DCNN), and each convolutional layer is followed by a residual connection (Residual Connection) self-attention layer to enhance information fusion and improve the expressiveness of features.

[0198] The structure of the decoder is a Deconvolutional Neural Network (DeCNN), and each deconvolution layer is followed by a residual connection, and its formula is the same as that of the encoder; the loss function of the decoder is the Adaptive Reconstruction Loss.

[0199] S3.3: Using the features of different time steps, we build a feature library and a dynamic feature fusion module. The feature library is used to store the features of different time steps. The dynamic feature fusion module is used to adaptively learn information from multiple time steps. The attention mechanism is used to select the most relevant features for fusion based on the features of the current time step, thereby improving the richness and effectiveness of the features.

[0200] S3.4: Obtain the output of unsupervised feature learning for subsequent feature fusion and classification.

[0201] The output of unsupervised feature learning is Z' t ∈R n×k , the value of each element represents the eigenvalue of the sample in the latent space dimension, which integrates spectral and spatial features, as well as information from multiple time steps, and has high expressive and distinguishing capabilities.

[0202] S4: Multi-scale and multi-modal feature fusion

[0203] A multi-scale convolutional neural network (MSCNN) is used to extract spectral and spatial features at different levels from hyperspectral data of different scales. A multi-modal fusion network (MMFN) is used to fuse hyperspectral data with other auxiliary data to improve the expression and identification capabilities of features. The information from different data sources is used to improve the diversity and complementarity of features.

[0204] The specific implementation of this step is as follows:

[0205] S4.1: The output of unsupervised feature learning is used as input to MSCNN. MSCNN consists of multiple convolutional layers and pooling layers. Each convolutional layer and pooling layer corresponds to a scale. The goal of the convolutional layer is to extract spectral and spatial features, and the goal of the pooling layer is to reduce the dimension of the features and increase the invariance of the features.

[0206] S4.2: The output of MSCNN and other auxiliary data (such as infrared images, visible light images, etc.) are taken as input into the Multi-Modal Fusion Network (MMFN). The MMFN consists of a feature fusion layer and a feature selection layer. The goal of the feature fusion layer is to fuse the features of different data sources using weighted averaging or concatenation methods. The goal of the feature selection layer is to select the most useful features from the fused features using sparse coding or principal component analysis methods.

[0207] S4.3: Get the output of multi-scale and multi-modal feature fusion for subsequent classification. The output of multi-scale and multi-modal feature fusion is G t ∈R n×p , the value of each element represents the eigenvalue of the sample in the final feature dimension, which integrates the spectral and spatial features of different scales and the information of different data sources, and has high expression and identification capabilities.

[0208] It should be noted that: in this step, MSCNN is used to extract spectral and spatial features at different levels from hyperspectral data of different scales, improve the diversity and complementarity of features, and utilize information from different data sources; the multimodal fusion network MMFN is used to fuse hyperspectral data with other auxiliary data to improve the expression and identification capabilities of features, and the weighted average or splicing method and sparse coding or principal component analysis method are used to achieve effective fusion and selection of features.

[0209] S5: Lightweight Transformer Model Classification

[0210] This step uses a lightweight Transformer model to classify the fused features, combining the global self-attention mechanism and the local convolutional attention mechanism to reduce computational complexity and memory consumption while maintaining high classification performance, and uses the powerful self-attention mechanism and multi-head attention mechanism of the Transformer model to improve classification accuracy and robustness.

[0211] The specific implementation of this step is as follows:

[0212] S5.1: The output of multi-scale and multi-modal feature fusion is taken as input and input into the lightweight Transformer model. The lightweight Transformer model consists of an encoder and a classifier. The goal of the encoder is to map the input features to a high-dimensional latent space, and the goal of the classifier is to map the features of the latent space to category labels.

[0213] S5.2: To reduce computational complexity and memory consumption, the encoder is improved. The encoder adopts a hybrid attention mechanism that combines the global self-attention mechanism and the local convolutional attention mechanism. The goal of the global self-attention mechanism is to capture the long-distance dependencies between features, and the goal of the local convolutional attention mechanism is to capture the short-distance dependencies between features.

[0214] Specifically, the structure of the encoder is a hybrid attention network (HAN), which consists of r hybrid attention layers (HAL), each of which contains a global self-attention sublayer (GSA) and a local convolutional attention sublayer (LCA).

[0215] S5.3: The output of the lightweight Transformer model directly corresponds to the quality assessment results of the beef samples. The model outputs a probability distribution Y t ∈R n×c, where n is the number of samples and c is the number of beef quality grades. The probability distribution corresponding to each sample reflects the probability that the beef sample belongs to each quality grade. Beef quality is divided into four grades: special grade, excellent grade, good grade and ordinary grade. For a specific beef sample, the model may output the following probability distribution: [0.75, 0.20, 0.04, 0.01], indicating that the sample has a 75% probability of being special grade beef, a 20% probability of being excellent grade beef, a 4% probability of being good grade beef, and a 1% probability of being ordinary beef. The probabilistic output gives the most likely quality grade of the beef sample and also provides confidence information for the evaluation result. The output of the lightweight Transformer model comprehensively considers multiple features of the beef sample, including spectral features, texture features, fat distribution, etc. Through the global self-attention mechanism, the model is able to capture the complex interactions between these features, such as the relationship between protein content and meat tenderness, or the association between fat distribution and flavor. At the same time, the local convolutional attention mechanism enables the model to accurately identify local features, such as the distribution pattern of intramuscular fat.

[0216] A lightweight Transformer model is used to classify the fused features to improve the accuracy and robustness of classification, and the powerful self-attention mechanism and multi-head attention mechanism of the Transformer model are utilized; the encoder is improved by adopting a hybrid attention mechanism combined with the global self-attention mechanism and the local convolutional attention mechanism, which reduces the computational complexity and memory consumption while maintaining high classification performance.

[0217] It should be noted that the output of the lightweight Transformer model is a four-dimensional vector, and each dimension corresponds to the probability of a quality grade. Assume that for a specific beef sample, the model may output [0.70, 0.25, 0.04, 0.01], indicating that the sample has a 70% probability of being premium beef, a 25% probability of being excellent beef, a 4% probability of being good beef, and a 1% probability of being ordinary beef. The following are the three main steps to implement the example of the present invention:

[0218] (1) Data Acquisition

[0219] Equipment selection: Hyperspectral imager is a device used to collect hyperspectral image data. The appropriate model should be selected according to different band ranges, resolutions, signal-to-noise ratios and other parameters. The indicators of hyperspectral imagers include the number of bands, wavelength range, spectral resolution, spatial resolution, signal-to-noise ratio, acquisition speed, etc. Typical: a hyperspectral imager with 256 bands, a wavelength range of 400-1000nm, a spectral resolution of 2.34nm, a spatial resolution of 0.5mm, a signal-to-noise ratio of 1000:1, and an acquisition speed of 10ms. Computers are devices used to store, process and analyze hyperspectral image data. They need to have sufficient memory, hard disk, CPU, GPU and other resources. The indicators of computers include memory, hard disk, CPU and GPU, etc. Typical: a computer with 16GB of memory, 1TB of hard disk, Intel Core i7 CPU and NVIDIA GeForce RTX 2080 GPU.

[0220] The above data can be obtained in the following ways:

[0221] Self-collection: Use appropriate instruments and equipment in the laboratory to collect different beef samples to build your own data set; ensure the quality and quantity of the collected data as well as the annotation and storage of the data.

[0222] Public access: Download publicly available hyperspectral datasets and auxiliary datasets from the Internet to supplement the dataset. Ensure the copyright and license of the data as well as the format and content of the data.

[0223] (2) Data processing

[0224] Data preprocessing: The dynamic range compression (DRC) method is applied to the hyperspectral beef sample data to reduce the noise and outliers of the data and improve the signal-to-noise ratio and dynamic range of the data. The exposure.adjust_log function in the scikit-image library of Python is used to perform logarithmic transformation on the hyperspectral data to implement the DRC method. It is also necessary to perform necessary cleaning, normalization, encoding and other operations on the auxiliary data to ensure the validity and consistency of the data.

[0225] Data enhancement: Use the WGAN-AGC (Watermark-GAN with Adaptive Gradient Clipping) method to generate more high-quality beef samples and enhance the scale and quality of the dataset. Use the torch.nn module in Python's PyTorch library to build a generative adversarial network (GAN) and use the torch.optim.Adam function in the torch.optim module to implement the adaptive gradient clipping (AGC) method to optimize the GAN training process. Perform random rotation, translation, scaling, flipping and other operations on the auxiliary data to increase the diversity and robustness of the data.

[0226] Data organization: Rationally organize and divide the hyperspectral data and auxiliary data to facilitate the input and output of the model. Use the DataFrame class in Python's pandas library to integrate the hyperspectral data and auxiliary data into a data frame (DataFrame), and use the label column as the data label (beef grade, freshness, nutritional value, etc.); use the train_test_sp lit function in the sklearn.model_selection module to divide the data frame into training set, validation set, and test set to facilitate model training, validation, and testing.

[0227] (3) Model construction

[0228] In order to implement the model of the present invention, it is necessary to build the following three core model components:

[0229] Diffusion model: used to automatically learn spectral and spatial features from hyperspectral data, while using the feature library and dynamic feature fusion module to adaptively learn information from multiple time steps. Use the torch.nn module in Python's Py Torch library to build a diffusion-based convolutional neural network (DCNN), and use the torch.nn.functional.conv2d function in the torch.nn.functional module to implement the diffusion convolution operation to extract the spectral and spatial features of hyperspectral data. Use the torch.nn.LSTM class in the torch.nn module to build a long short-term memory network (LSTM), and use the torch.nn.Linear class in the torch.nn module to build a fully connected layer, implement the feature library and dynamic feature fusion module, and adaptively learn information from multiple time steps. Refer to the following model core parameter configuration (Table 1):

[0230] Table 1

[0231]

[0232] MSCNN: It is used to extract spectral and spatial features at different levels from hyperspectral data of different scales, and use MMFN to fuse hyperspectral data with other auxiliary data to improve the expressiveness and identification ability of features. Use the torch.nn module in Python's PyTorch library to build a multi-scale convolutional neural network (MSCNN), and use the torch.nn.Conv2d class in the torch.nn module to implement multi-scale convolution operations to extract spectral and spatial features of hyperspectral data of different scales. Use the torch.nn.ModuleList class in the torch.nn module to build a module list for storing convolutional layers of different scales, which is convenient for dynamically adjusting the number and parameters of convolutional layers. Use the torch.nn.MultiheadAttention class in the torch.nn module to build a multi-head attention network (MMFN), and use the torch.nn.Linear class in the torch.nn module to build a fully connected layer to achieve the fusion of hyperspectral data and other auxiliary data, and improve the expressiveness and identification ability of features. Refer to the following model core parameter configuration (see Table 2)

[0233] Table 2

[0234]

[0235] Lightweight Transformer model: used to classify the fused features, combining the global self-attention mechanism and the local convolutional attention mechanism to reduce computational complexity and memory consumption while maintaining high classification performance. Use the torch.nn module in Python's PyTorch library to build a lightweight Transformer model, and use the torch.nn.TransformerEncoder class in the torch.nn module to implement the global self-attention mechanism to capture the long-distance dependencies between features. Use the torch.nn.Conv1d class in the torch.nn module to implement the local convolutional attention mechanism to capture the short-distance dependencies between features. Use the torch.nn.Linear class in the torch.nn module to build a fully connected layer to classify features and output the quality assessment results of beef. Refer to the following model core parameter configuration (see Table 3):

[0236] Table 3

[0237]

[0238]

[0239] S6: Robust Optimization

[0240] In this step, an adversarial sample generator (ASG) is used to generate some hyperspectral data with slight perturbations during the training process, so that the model can resist some potential attacks and interferences and improve the stability and security of the model.

[0241] The specific implementation of this step is as follows:

[0242] S6.1: The output of the lightweight Transformer model and the original hyperspectral data are taken as input into ASG. ASG consists of a perturbation generator and a perturbation evaluator. The goal of the perturbation generator is to generate some hyperspectral data with small perturbations without changing the category of the data. The goal of the perturbation evaluator is to evaluate the size and effectiveness of the perturbation.

[0243] S6.2: Using the adversarial training method, during the training process, the hyperspectral data generated by the perturbation generator and the original hyperspectral data are alternately input into the lightweight Transformer model, so that the model can adapt to different data distributions and improve the robustness of the model.

[0244] Specifically, for each training batch, half of the samples are randomly selected, and the corresponding perturbation data is generated by the perturbation generator. The perturbation data and the original data are concatenated into a new input while keeping the category label unchanged. The new input and category label are then input into the lightweight Transformer model for forward and backward propagation to update the model parameters.

[0245] S6.3: The robustly optimized lightweight Transformer model is used to perform the final beef quality assessment. The model directly outputs the quality assessment results of beef samples, and divides the beef quality into four grades: special grade, excellent grade, good grade, and ordinary grade. Based on multiple key indicators, including meat tenderness, marbling (intermuscular fat distribution), meat color, and meat flavor. The output of the model is a four-dimensional vector, with each dimension corresponding to the probability of a quality grade. For a specific beef sample, the model may output [0.70, 0.25, 0.04, 0.01], indicating that the sample has a 70% probability of being special grade beef, a 25% probability of being excellent grade beef, a 4% probability of being good grade beef, and a 1% probability of being ordinary grade beef. The robustly optimized lightweight Transformer model resists some potential attacks and interferences, improves the stability and security of the model, and maintains high classification accuracy and robustness.

[0246] An adversarial sample generator (ASG) is used to generate some hyperspectral data with slight perturbations during the training process, so that the model can resist some potential attacks and interferences, thereby improving the stability and security of the model. Using the adversarial training method, during the training process, the hyperspectral data generated by the perturbation generator and the original hyperspectral data are alternately input into the lightweight Transformer model, so that the model can adapt to different data distributions and improve the robustness of the model.

[0247] The present invention also has the following technical effects: (1) Based on the generative adversarial network (GAN) and data enhancement, it can use a small amount of real beef hyperspectral images to generate more high-quality beef samples, enhance the scale and quality of the data set, and solve the problems of insufficient data and category imbalance in hyperspectral image classification. The Wasserstein distance is used as the loss function to improve the training stability and convergence of GAN; at the same time, the adaptive gradient clipping (AGC) mechanism is introduced to dynamically adjust the gradient range of the generator and the discriminator, which can prevent the gradient from disappearing or exploding, thereby further improving the generation effect of GAN. (2) Based on the convolutional neural network (CNN) feature extraction and fusion method (MSCNN-MMFN), it is possible to extract spectral and spatial features of different levels from hyperspectral data of different scales, and at the same time use the multimodal fusion network (MMFN) to fuse hyperspectral data and other auxiliary data (such as cattle breed, age, feeding method, etc.), thereby improving the expression and identification capabilities of features. MSCNN-MMFN uses multi-scale convolutional module (MSCM) and multi-scale fusion module (MFM) to construct a multi-scale convolutional neural network (MSCNN), which can effectively extract the spectral and spatial features of hyperspectral data; at the same time, it uses multi-head self-attention mechanism (MHA) and multi-layer perceptron (MLP) to construct MMFN, which can effectively fuse hyperspectral data and other auxiliary data. (3) The sequence-to-sequence model (Transformer) based on the self-attention mechanism can classify the fused features, and combines the global self-attention mechanism and the local convolutional attention mechanism to reduce the computational complexity and memory consumption, while maintaining a high classification performance. Transformer adopts an encoder-decoder structure, uses the self-attention mechanism to capture the long-distance dependency between features, uses the convolutional attention mechanism to capture the local correlation between features, uses the multi-head attention mechanism to increase the expression ability of the model, and uses residual connection and layer normalization to increase the stability of the model.

[0248] The present invention utilizes WGAN-AGC to generate more high-quality beef samples, utilizes MSCNN-MMFN to extract and fuse the spectral and spatial features of beef, and utilizes Transformer to classify the quality of beef, thereby realizing an end-to-end beef quality assessment method based on deep learning. The method has the characteristics of non-destructiveness, rapidity, high efficiency, and intelligence, and can effectively solve the problems of insufficient data, feature extraction, feature fusion, and classification accuracy in hyperspectral image classification, thus providing a new technical means for beef production, processing, and consumption.

[0249] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A lightweight beef quality assessment method based on unsupervised feature learning, characterized in that: Includes the following methods: S1. Obtaining original hyperspectral data of beef as real data, and generating high-quality sample data based on a generative adversarial network model, wherein the generative adversarial network model includes a generator and a discriminator; S2. Determine the dynamic range of the high-quality sample data, and determine the corresponding threshold value according to the distribution characteristics of the high-quality sample data within the dynamic range; Compressing the dynamic range of high-quality sample data at each wavelength to a corresponding threshold value, and performing preprocessing to obtain preprocessed sample data; S3. Based on the unsupervised feature learning model, the preprocessed sample data is mapped to a low-dimensional latent space and reconstructed into the original sample data, and the most relevant feature information in the original sample data is fused to obtain the fused original sample data; S4, based on the deep learning model MSCNN, extract different levels of spectral and spatial features from the fused original sample data at different scales to obtain the characteristic values ​​of the original sample data; Based on a multi-scale and multi-modal feature fusion network, the feature values ​​of the original sample data and the feature values ​​of other auxiliary data are fused or spliced ​​to obtain feature values ​​of different data sources, and the most useful feature values ​​are selected from the feature values ​​of different data sources for fusion to obtain fused sample features; S5. Based on a lightweight Transformer model, the fused sample features are mapped to a high-dimensional latent space to obtain the fused sample features in the latent space, and the fused sample features in the latent space are mapped to the category label to obtain the probability of the category label of the fused sample features; the lightweight Transformer model includes an encoder and a classifier; S6. Inputting the probability of the category label of the fused sample feature and the real data into an adversarial sample generator, wherein the adversarial sample generator includes: a perturbation generator and a perturbation evaluator; Alternately inputting the sample data with slight perturbations generated by the perturbation generator and the real data into the lightweight Transformer model, performing forward propagation and back propagation, and updating the parameters of the lightweight Transformer model; Beef quality is evaluated based on the lightweight Transformer model with updated parameters.

2. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S1, the original hyperspectral data of beef is obtained as real data, and high-quality sample data is generated based on a generative adversarial network model, specifically: the generative adversarial network includes a generator and a discriminator, and the loss functions of the generator and the discriminator are calculated respectively; the original gradients of the generator and the discriminator are clipped using an adaptive gradient clipping method; and the above process is repeated until the generator and the discriminator reach a balance, thereby generating high-quality sample data; The loss functions of the generator and the discriminator are calculated separately, specifically: the Wasserstein distance is set to define the loss function of the discriminator, and the expression is: Where: L D represents the loss function of the discriminator, Indicates that in the real data distribution P r expectations, Indicates that in generating data distribution P g The expectation on , x is the real data, is the data generated by the generator, P r is the distribution of real data, P g is the distribution of generated data, and D is the discriminator; The generator is a reverse convolutional neural network, the dimension of the input layer of the reverse convolutional neural network is the dimension m of random noise, and the dimension of the output layer of the reverse convolutional neural network is the same as the dimension of the input layer of the discriminator, which is n; The loss function expression of the generator is: Where z is random noise, P z is the distribution of random noise, G is the generator; The calculation method of the adaptive gradient clipping is: In the above formula, is the clipped gradient, g is the original gradient of the generator and discriminator, and c is the clipping threshold, where c = 0.

01.

3. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S2, the dynamic range of the high-quality sample data is determined, and the corresponding threshold is determined according to the distribution characteristics of the high-quality sample data within the dynamic range; The dynamic range of the high-quality sample data at each wavelength is compressed to a corresponding threshold value, and preprocessed to obtain the preprocessed sample data, specifically: calculating the maximum value and the minimum value of the high-quality sample data to determine the dynamic range of the high-quality sample data; determining the corresponding threshold value according to the distribution characteristics of the high-quality sample data in the dynamic range, compressing the dynamic range of the high-quality sample data at each wavelength to the corresponding threshold value, and normalizing the compressed high-quality sample data to obtain the preprocessed sample data; The maximum and minimum values ​​of the high-quality sample data are calculated to determine the dynamic range of the high-quality sample data. The expression is: The high-quality sample dataset X∈R n×d , where n is the number of high-quality samples and d is the number of hyperspectral wavelengths; The maximum value of the high-quality sample data is: Among them, X ij is the reflectance of the i-th sample at the j-th wavelength; The minimum value of the high-quality sample data is: The dynamic range of determining high-quality sample data is: R=X max -X min ; Determine the corresponding threshold value based on the distribution characteristics of high-quality sample data within the dynamic range, specifically: calculate the mean and standard deviation of each sample data at each wavelength: The expression of the mean is: Among them, μ j represents the mean value at the jth wavelength; The expression of the standard deviation is: Among them, σ j represents the standard deviation at the jth wavelength; The threshold is: [μ j -3σ j ,μ j +3σ j ]; Compress the dynamic range of high-quality sample data at each wavelength to [μ-3σ,μ+3σ] to obtain compressed sample data:

4. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 3 is characterized in that: The compressed high-quality sample data are normalized to obtain pre-processed sample data, specifically: Using linear normalization, the expression is: Among them, X' max is the maximum value of the compressed sample data, X' min is the minimum value of the compressed data; X” ij is the reflectance of the i-th sample at the j-th wavelength after normalization; The preprocessed sample data is X”∈R n×d , where n is the number of samples and d is the number of wavelengths.

5. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S3, the preprocessed sample data is mapped to a low-dimensional latent space and reconstructed into the original sample data, and the most relevant feature information in the original sample data is fused to obtain the fused original sample data, specifically: The unsupervised feature learning model is an improved diffusion model, and the improved diffusion model includes an encoder and a decoder, specifically: the encoder is used to map the preprocessed sample data to a low-dimensional latent space, and the decoder is used to reconstruct the preprocessed sample data in the latent space into the original sample data; Using the feature information of the original sample data at different time steps, a feature library and a dynamic feature fusion module are constructed. The feature library is used to store the feature information of the original sample data at different time steps. The dynamic feature fusion is used to adaptively learn the feature information of the original sample data at different time steps. According to the feature information of the original sample data at the current time step, the feature information of the most relevant original sample data is selected for fusion. Among them, the preprocessed hyperspectral data is X”∈R n×d , and input into the encoder, the encoder output is: Z=E(X”)∈R m×k In the above formula, E is the encoder, k is the dimension of the latent space, Z is the output matrix after encoding, and Z ij is the value of the jth dimension of the i-th sample in the latent space; The decoder is used to reconstruct the preprocessed sample data in the latent space into the original sample data, specifically: The output of the decoder is: In the above formula, D is the decoder, is the reconstructed reflectance of the i-th sample at the j-th wavelength; Each convolution layer in the encoder adopts a residual connection, and the last layer of the encoder is a self-attention layer; each reverse convolution layer in the decoder adopts a residual connection, and the loss function of the decoder is an adaptive reconstruction loss; Each convolutional layer in the encoder uses a residual connection, expressed as: H l =F l (H l-1 )+H l-1 Among them, H l is the output of the lth layer, F l is the convolution operation of the lth layer, H l-1 is the output of the l-1th layer; The last layer of the encoder is the self-attention layer, which is expressed as: Z=softmax(H L W Q (H L W K ) T )H L W V Among them, H L is the output of the last layer, W Q , W K , W V is the parameter of the attention layer, and softmax is the normalization function; The loss function of the decoder is the adaptive reconstruction loss, specifically: Among them, w ij is the reconstruction weight of the i-th sample at the j-th wavelength, which is calculated as: in, is the variance of the jth wavelength, which is calculated as: Among them, μ j is the mean value of the jth wavelength, which is calculated as: The constructing of the feature library is specifically as follows: Assume that the current time step is t, then the feature library is a three-dimensional tensor with a shape of T×n×k; Where: T is the total number of time steps, n is the number of samples, k is the dimension of the latent space, and the tth slice of the feature library is Z t ∈R n ×k , which is the feature of the current time step; The input of the dynamic feature fusion module is the feature library, and the output is a two-dimensional matrix with a shape of n×k. The calculation method is: WITH' t =softmax(Z t IN Q (WITH t IN K ) T )WITH t IN V +softmax(Z t IN Q (WITH 1:t-1 IN K ) T )WITH 1:t-1 IN V In the above formula, Z' t is the fused feature, Z 1:t-1 is the first t-1 slices of the feature library, that is, the features of the historical time step, W Q , W K , W V is the parameter of the dynamic feature fusion module, and softmax is the normalization function; The original sample data after fusion is: Z' t ∈R n×k , where n is the number of samples and k is the dimension of the latent space.

6. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S4, based on the deep learning model MSCNN, spectral and spatial features of different levels are extracted from the fused original sample data at different scales to obtain the characteristic values ​​of the original sample data, specifically: The characteristic value of the original sample data is Z′ t ∈R n×k , Z' t,ji is the eigenvalue of the i-th sample in the j-th latent space dimension; Then the characteristic value of the original sample data is: F t =M(Z′ t )∈R n×l Among them, M is the deep learning model MSCNN, l is the dimension of the features of the original sample data, and F t,ij is the feature value of the i-th sample in the j-th feature dimension; The corresponding scale of each convolution layer and pooling layer in the deep learning model MSCNN is expressed as: H t,0 =Z' t H t,i =P i (C i (H t,i-1 )) F t =H t,s Among them, H t,i is the output of the i-th scale, C i is the i-th convolutional layer, P i is the i-th pooling layer, and s is the total number of scales.

7. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S4, based on a multi-scale and multi-modal feature fusion network, the feature values ​​of the original sample data and the feature values ​​of other auxiliary data are fused or spliced ​​to obtain feature values ​​of different data sources, and the most useful features are selected from the feature values ​​of different data sources to obtain fused sample features, specifically: The eigenvalue of the original sample data is F t ∈R n×l ; Where n is the number of samples, l is the dimension of the feature, and F t,ij is the feature value of the i-th sample in the j-th feature dimension; Other auxiliary data is A t ∈R n×m ; Among them, m is the dimension of auxiliary data, A t,ij is the eigenvalue of the i-th sample in the j-th auxiliary data dimension; Get the characteristic values ​​of different data sources, the expression is: G t =S(F t ,A t )∈R n×p Among them, S is the deep learning model MMFN, p is the dimension of the final feature, G t,ij is the feature value of the i-th sample in the j-th final feature dimension; The deep learning model MSCNN includes multiple convolutional layers and pooling layers; The multi-scale and multi-modal feature fusion network includes a feature fusion layer and a feature selection layer, wherein the feature fusion layer is used to fuse feature information from different data sources, and the feature selection layer is used to select useful feature information from the fused feature information; The multi-scale and multi-modal feature fusion network includes a feature fusion layer and a feature selection layer, and the expression is: B t =F(F t ,A t ) G t =S(B t ) Among them, B t is the output of the feature fusion layer, F is the feature fusion layer, and S is the feature selection layer; The feature fusion layer is used to fuse the features of different data sources, specifically by using a weighted average or concatenation method, and the expression is: B t =αF t +(1-a)A t or B t =[F t ;A t ] Among them, α is the fusion weight, [;] is the splicing operation; The feature selection layer is used to select useful feature information from the fused feature information, specifically by using sparse coding or principal component analysis, the expression is: G t =B t W or G t =PCA(B t ) Among them, W is the parameter of sparse coding, and PCA is the function of principal component analysis.

8. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S5, based on the lightweight Transformer model, the fused sample features are mapped to a high-dimensional latent space to obtain the fused sample features in the latent space, and the fused sample features in the latent space are mapped to the category label to obtain the probability of the category label of the fused sample features, specifically: The lightweight Transformer model includes an encoder and a classifier, wherein the encoder is used to map the fused sample features to a high-dimensional latent space, and the classifier is used to map the fused sample features in the latent space to a category label and output the probability of the category label; The encoder is a hybrid attention network, including several hybrid attention layers, each of which includes a global self-attention sublayer and a local convolutional attention sublayer, the global self-attention layer is used to capture the long-distance dependency between feature information, and the local convolutional attention layer is used to capture the short-distance correlation between feature information; Among them, the output of the multi-scale and multi-modal feature fusion is G t ∈R n×p ; Among them, n is the number of samples, p is the dimension of the final feature, G t,ij is the feature value of the i-th sample in the j-th final feature dimension; Then the output of the encoder is: H t =E(G t )∈R n×q Among them, E is the encoder, q is the dimension of the latent space, and H t,ij is the eigenvalue of the jth dimension of the i-th sample in the latent space; The output of the classifier is: Y t =C(H t )∈R n×c Among them, C is the classifier, c is the number of categories, and Y t,ij is the probability that the i-th sample belongs to the j-th category; In step S5, each of the hybrid attention layers includes a global self-attention sublayer and a local convolutional attention sublayer, and the expression is: H t,0 =G t H t,i =LayerNorm(H t,i-1 +GSA(H t,i-1 ))+LayerNorm(H t,i-1 +LCA(H t,i-1 )) H t =H t,r Among them, H t,i is the output of the i-th hybrid attention layer, LayerNorm is the layer normalization function, and r is the total number of hybrid attention layers; In step S5, the global self-attention layer is used to capture the long-distance dependencies between features, and the expression is: GSA(H t,i-1 )=softmax(H t,i-1 W Q (H t,i-1 W K ) T )H t,i-1 W V Among them, W Q , W K , W V is the parameter of the global self-attention sublayer, and softmax is the normalization function; The role of the local convolutional attention sublayer is to capture the short-range correlation between features, and its formula is: LCA(H t,i-1 )=softmax(H t,i-1 W Q (Conv(H t,i-1 ,W K )) T )Conv(H t,i-1 ,W V ) Among them, W Q , W K , W V are the parameters of the local convolutional attention sublayer, softmax is the normalization function, and Conv is the convolution operation; The probability of obtaining the category label of the fused sample features is specifically: The lightweight Transformer model outputs a probability distribution Y t ∈R n×c , where n is the number of samples and c is the number of beef quality grades.

9. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S6, the probability of the category label of the fused sample feature and the real data are input into the adversarial sample generator, specifically: The output of the lightweight Transformer model is: t ∈R n×c ; Where n is the number of samples, c is the number of categories, and Y t,ij is the probability that the i-th sample belongs to the j-th category; The real data is X∈R n×d , where d is the wavelength, X ij is the reflectance of the i-th sample at the j-th wavelength; The perturbation generator and the perturbation evaluator are used to perform adversarial training, so that the perturbation generator generates sample data with specific small perturbations similar to the real data, specifically: the perturbation generator is used to generate sample data with small perturbations similar to the real data; the perturbation evaluator in the adversarial sample generator is used to output a perturbation score of the difference between the real data and the perturbed sample data; Then the output of the disturbance generator is: Where G is the disturbance generator, is the disturbed reflectivity of the i-th sample at the j-th wavelength; The output of the perturbation evaluator is: Where E is the disturbance estimator, S i Score the perturbation of the i-th sample; The perturbation evaluator in the adversarial sample generator is used to output a perturbation score of the difference between the real data and the perturbed sample data, expressed as: The loss function of the perturbation estimator is: Among them, |·| is the norm function.

10. The lightweight beef quality assessment method based on unsupervised feature learning according to claim 1, characterized in that: In step S6, the sample data with slight perturbations and the real data generated by the perturbation generator are alternately input into the lightweight Transformer model, and forward propagation and back propagation are performed to update the parameters of the lightweight Transformer model; specifically, for each training batch, half of the sample data with slight perturbations and the real data are randomly selected, and the corresponding perturbation data is generated by the perturbation generator, and the perturbation data and the original data are spliced ​​into a new input, while the category label is kept unchanged, and then the new input and the category label are input into the lightweight Transformer model, and forward propagation and back propagation are performed to update the parameters of the model; The loss function of the perturbation generator and the perturbation evaluator for adversarial training is: Among them, [;] is the concatenation operation, E is the encoder, and C is the classifier.

Citation Information

Patent Citations

  • Panchromatic and multispectral remote sensing image fusion method based on GAN and Transform

    CN117350923A

  • Hyperspectral and laser radar multilayer fusion classification method based on adversarial learning

    CN117934978A