Glucose degree detection model acquisition method based on hyperspectrum and self-supervised learning

Through the glucose degree detection model with hyperspectral and self-supervised learning, the problem of low glucose degree detection efficiency is solved, lossless and accurate grape quality sorting is achieved, and detection accuracy and robustness are improved.

CN120354174APending Publication Date: 2025-07-22INST OF INTELLIGENT MFG GUANGDONG ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511066.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, glucose detection depends on lossy manual measurement methods, resulting in low detection efficiency and high cost, making it difficult to meet the quality sorting needs of grapes, and industrial hyperspectral non-destructive detection methods lack effective means for single grapes.

Method used

Combining hyperspectral and self-supervised learning, a glucose degree detection model based on Transformer is constructed. By obtaining grape hyperspectral data samples, pre-processing and black-and-white frame correction, training sets, verification sets and prediction sets are constructed, and a self-supervised learning training detection network is used to achieve non-destructive detection of single grapes and whole bunches of glucose degrees.

Benefits of technology

It improves the accuracy and robustness of glucose detection, realizes fast and lossless grape quality sorting, reduces manpower, time and financial costs, and realizes accurate glucose sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354174A_ABST
    Figure CN120354174A_ABST
Patent Text Reader

Abstract

The invention provides a glucose degree detection model obtaining method and device based on hyperspectrum and self-supervised learning and terminal equipment, and the method comprises the steps: obtaining a grape hyperspectral data sample, and measuring the glucose degree of each grape and the average sugar degree of a whole string of grapes to obtain a measured sugar degree true value; the method comprises the following steps: constructing a training set, a verification set and a prediction set through hyperspectral data and measured sugar degree true values, and constructing a single sugar degree and average sugar degree detection network based on Transform; constructing a learning task of self-supervised upstream whole string comparison and intra-string comparison for glucose degree prediction and a learning task of downstream single sugar degree and whole string average sugar degree, and training a detection network by using the training set; finely adjusting the detection network according to the verification set and the prediction set to obtain a glucose degree detection model. According to the method, manual detection is replaced by combining hyperspectral and self-supervised learning, so that the grape detection efficiency is improved, and the quality accurate sorting of the glucose degree is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-destructive detection of grapes in the sorting of agricultural fruit quality. Specifically, it relates to a method, device and terminal device for obtaining a glucose detection model based on hyperspectral and self-supervised learning. Background Art

[0002] The glucose content of a single grape is an important indicator for evaluating grape quality. At present, traditional grape quality sorting still relies on manual refractometers to measure the sugar content by damaging the grapes, which is a destructive measurement method. There are problems such as product damage and low detection efficiency, resulting in reduced economic benefits for enterprises. During the grape season, measuring the sugar content of each grape consumes a large amount of manpower, material resources and financial resources, making it difficult to meet the needs of grape quality sorting during picking. In addition, there are already sugar content detection methods for different fruits in the field of industrial hyperspectral non-destructive detection, but there are few non-destructive detection methods for the glucose content of each grape in a freshly picked bunch of grapes, making it difficult to popularize the single-grape glucose content technology based on industrial hyperspectral in the industry. Summary of the Invention

[0003] The main object of the present invention is to provide a method, device and terminal device for obtaining a glucose detection model based on hyperspectral and self-supervised learning. By combining hyperspectral and self-supervised learning to replace manual detection, the grape detection efficiency is improved, and accurate sorting of grape quality based on sugar content is achieved, aiming to solve the technical problem that the existing glucose detection technology is difficult to meet the needs of grape quality sorting during picking.

[0004] In a first aspect, the present invention provides a method for obtaining a glucose detection model based on hyperspectral and self-supervised learning, including:

[0005] Obtain a grape hyperspectral data sample, wherein the grape hyperspectral data sample includes a hyperspectral data block of a bunch of grapes, and measure the glucose content of each grape and the average sugar content of the bunch of grapes in the grape hyperspectral data sample to obtain the true value of the measured sugar content;

[0006] Use a preset data processing method to preprocess and correct the black and white frames of the hyperspectral data in the grape hyperspectral data sample to obtain a corrected hyperspectral data block;

[0007] Construct a training set, a validation set and a prediction set for the glucose content of a single grape based on hyperspectral and self-supervised learning through the hyperspectral data in the corrected hyperspectral data block and the true value of the measured sugar content;

[0008] Construct a single-sugar-content and average-sugar-content detection network based on Transformer according to the training set, the validation set and the prediction set;

[0009] Construct a self-supervised upstream whole-cluster contrast and intra-cluster contrast learning task for glucose degree prediction, and use the training set to train the single-glucose-degree and average-glucose-degree detection network based on Transformer;

[0010] Construct a downstream learning task for single-glucose-degree and whole-cluster average-glucose-degree, and use the training set to further train the single-glucose-degree and average-glucose-degree detection network based on Transformer;

[0011] Fine-tune the trained single-glucose-degree and average-glucose-degree detection network based on Transformer according to the validation set and the prediction set to obtain a glucose-degree detection model that combines single-glucose-degree and whole-cluster glucose-degree.

[0012] In a specific embodiment, the use of a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data sample and perform black-and-white frame correction to obtain a corrected hyperspectral data block includes:

[0013] Use a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data sample and extract a single-grape mask to obtain an enhanced processed hyperspectral data block;

[0014] Perform black-and-white frame correction on the hyperspectral data in the enhanced processed hyperspectral data block to obtain a corrected hyperspectral data block.

[0015] In a specific embodiment, the single-glucose-degree and average-glucose-degree detection network based on Transformer includes:

[0016] A single-grape spectral Transformer encoder for sorting the input grapes from the top to the end in combination with positional encoding, and learning to extract feature of effective glucose-degree information for single-grape spectra at different positions;

[0017] A single-glucose-degree Transformer decoder for learning to predict the glucose-degree of each grape through an attention mechanism;

[0018] A whole-cluster average-glucose-degree Transformer decoder for learning to predict the whole-cluster average-glucose-degree in combination with the whole-cluster average spectrum, average-glucose-degree Token, and all extracted single-grape spectra;

[0019] Among them, each Transformer encoder and decoder is composed of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a positional encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.

[0020] In a specific embodiment, the construction of the self-supervised upstream whole-cluster contrast and intra-cluster contrast learning tasks for glucose degree prediction includes:

[0021] Construct two upstream task loss functions for self-supervised learning of glucose degree prediction, where the two upstream task loss functions are the grape cluster discrimination task loss function and the grape berry discrimination task loss function respectively;

[0022] The calculation formula of the grape cluster discrimination task loss function is:

[0023]

[0024] where e is the exponential function, g(x i,* ; θ h , θ s ) is the potential representation of a certain cluster of grapes, θ h is the weight of the single grape spectral Transformer encoder, θ s is the model pre-training weight of the whole-cluster grape average sugar degree Transformer decoder, S is the batch size of the grape cluster discrimination task, and B is the batch size of the negative samples in this task;

[0025] The calculation formula of the grape berry discrimination task loss function is:

[0026]

[0027] where e is the exponential function, f(x i,n ; θ h , θ b ) is the potential representation of a single grape, θ h is the weight of the single grape spectral Transformer encoder, θ b is the model pre-training weight of the single grape glucose degree Transformer decoder, N is the batch size of the grape berry discrimination task, and T is the batch size of the negative samples in this task;

[0028] The calculation formula of the objective function after combining the two is:

[0029] L SSL = (1 - λ)L 串 + λL 颗

[0030] where the hyperparameter λ controls the adjustment of the learned representation.

[0031] In a specific embodiment, the construction of the downstream single grape sugar degree and whole-cluster average sugar degree learning tasks includes:

[0032] Construct a downstream task loss function for single - grape glucose content and the glucose content of the whole bunch of grapes. Among them, the downstream task loss function adopts the mean - square error loss function, and the calculation formula of the mean - square error loss function is:

[0033]

[0034] where N is the total number of grape bunches, C is the number of grapes per bunch, b i,t is the single - grape sugar content label, f(x i,t ; θ h , θ b ) is the predicted value of the single - grape sugar content of the t - th grape sample in the i - th bunch, is the average sugar content label of all grapes in the i - th bunch, g(x i ; θ h , θ s ) is the predicted value of the sugar content of the grape sample in the i - th bunch.

[0035] Second, the present invention provides a device for obtaining a glucose content detection model based on hyperspectral and self - supervised learning, including:

[0036] A grape hyperspectral data sample and measured sugar content true - value acquisition module, which is used to acquire grape hyperspectral data samples. Among them, the grape hyperspectral data samples contain hyperspectral data blocks of the whole bunch of grapes, and measure the glucose content of each grape and the average glucose content of the whole bunch of grapes in the grape hyperspectral data samples to obtain the measured sugar content true - value;

[0037] A data processing module, which is used to pre - process and correct black - and - white frames for the hyperspectral data in the grape hyperspectral data samples by using a preset data processing method to obtain corrected hyperspectral data blocks;

[0038] A sample data - set construction module, which is used to construct a training set, a validation set, and a prediction set for single - grape glucose content based on hyperspectral and self - supervised learning through the hyperspectral data in the corrected hyperspectral data blocks and the measured sugar content true - value;

[0039] A network construction module, which is used to construct a single - grape sugar content and average sugar content detection network based on Transformer according to the training set, the validation set, and the prediction set;

[0040] An upstream training module, which is used to construct a self - supervised upstream whole - bunch contrast and intra - bunch contrast learning task for glucose content prediction, and use the training set to train the single - grape sugar content and average sugar content detection network based on Transformer;

[0041] A downstream training module for constructing learning tasks of single grape sugar content and average sugar content of the whole bunch, and further training the Transformer-based single grape sugar content and average sugar content detection network by using the training set;

[0042] A glucose content detection model acquisition module for fine-tuning the trained Transformer-based single grape sugar content and average sugar content detection network according to the validation set and the prediction set to obtain a glucose content detection model that combines single grape glucose content and whole bunch glucose content.

[0043] In a specific embodiment, the data processing module is specifically configured to:

[0044] Use a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data sample and extract a single grape mask to obtain an enhanced hyperspectral data block;

[0045] Perform black and white frame correction on the hyperspectral data in the enhanced hyperspectral data block to obtain a corrected hyperspectral data block.

[0046] In a specific embodiment, the Transformer-based single grape sugar content and average sugar content detection network includes:

[0047] A single grape spectrum Transformer encoder for sorting the input grapes from the top to the end in combination with positional encoding, and learning to extract feature of effective sugar content information from the spectra of single grapes at different positions;

[0048] A single grape sugar content Transformer decoder for learning to predict the sugar content of each grape through an attention mechanism;

[0049] An average sugar content Transformer decoder for the whole bunch of grapes for learning to predict the average sugar content of the whole bunch of grapes by combining the average spectrum of the whole bunch of grapes, the average sugar content token, and the spectra of all extracted single grapes;

[0050] Wherein, each Transformer encoder and decoder is composed of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a positional encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.

[0051] In a specific embodiment, the upstream training module is specifically configured to:

[0052] Construct two upstream task loss functions for self-supervised learning of glucose content prediction, wherein the two upstream task loss functions are a grape bunch discrimination task loss function and a grape berry discrimination task loss function respectively;

[0053] The loss function for the grape cluster discrimination task has the following calculation formula:

[0054]

[0055] where \(e\) is the exponential function, \(g(x i,* ; \(\theta h , \(\theta s )\) is the potential representation of a single grape in a cluster, \(\theta h \) is the weight of the single grape spectral Transformer encoder, \(\theta s \) is the pre-trained weight of the average sugar content Transformer decoder for the entire grape cluster, \(S\) is the batch size of the grape cluster discrimination task, and \(B\) is the batch size of the negative samples in this task;

[0056] The loss function for the single grape discrimination task has the following calculation formula:

[0057]

[0058] where \(e\) is the exponential function, \(f(x i,n ; \(\theta h , \(\theta b )\) is the potential representation of a single grape, \(\theta h \) is the weight of the single grape spectral Transformer encoder, \(\theta b \) is the pre-trained weight of the single grape sugar content Transformer decoder, \(N\) is the batch size of the single grape discrimination task, and \(T\) is the batch size of the negative samples in this task;

[0059] The calculation formula for the objective function after combining the two is:

[0060] \(L SSL =(1 - \lambda)L 串 +\lambda L 颗

[0061] where the hyperparameter \(\lambda\) controls the adjustment of the learned representation;

[0062] The downstream training module is specifically used for:

[0063] Construct a downstream task loss function for the single grape sugar content and the entire grape cluster sugar content. Among them, the downstream task loss function uses the mean squared error loss function, and the calculation formula of the mean squared error loss function is:

[0064]

[0065] where \(N\) is the total number of grape clusters, \(C\) is the number of grapes in each cluster, \(b i,t \) is the single grape sugar content label, \(f(x i,t ; \(\thetah , θ b ) is the predicted value of the sugar content of the t-th grape sample in the i-th bunch, is the average sugar content label of all grapes in the i-th bunch, g(x i ; θ h , θ s ) is the predicted value of the sugar content of the grape sample in the i-th bunch.

[0066] In a third aspect, the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a method for obtaining a glucose content detection model based on hyperspectral and self-supervised learning as described in the first aspect.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention combines grape hyperspectral with deep learning technology to model and select model parameters, constructs a Transformer glucose content detection model based on hyperspectral and self-supervised learning, which can improve the accuracy and robustness of the model; in addition, through the Transformer glucose content detection model based on hyperspectral and self-supervised learning, the glucose content detection model can be used to quickly and nondestructively obtain the glucose content of each grape and the average sugar content of the whole bunch of grapes, greatly reducing the labor, time, and financial costs of glucose content measurement during the grape season, and realizing rapid and accurate sorting of grape quality according to sugar content. Description of the Drawings

[0068] Figure 1 is a schematic flowchart of a method for obtaining a glucose content detection model based on hyperspectral and self-supervised learning provided by an embodiment of the present invention;

[0069] Figure 2 is a schematic structural diagram of a single-grape sugar content and average sugar content detection network based on Transformer provided by an embodiment of the present invention;

[0070] Figure 3 is a schematic diagram of the true sugar content and prediction result of a single grape in an embodiment of the present invention;

[0071] Figure 4 is a schematic structural diagram of a device for obtaining a glucose content detection model based on hyperspectral and self-supervised learning provided by an embodiment of the present invention;

[0072] Figure 5 is a schematic structural diagram of a terminal device provided by an embodiment of the present invention.

[0073] Wherein:

[0074] The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0075] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically and clearly defined.

[0077] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection, a direct connection, or an indirect connection through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0078] In the present invention, unless otherwise clearly specified and limited, the first feature being "above" or "below" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features between them. Moreover, the first feature being "above", "over", and "on" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "below", "under", and "beneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely indicates that the horizontal height of the first feature is lower than that of the second feature.

[0079] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for obtaining a glucose detection model based on hyperspectral and self-supervised learning provided by an embodiment of the present invention.

[0080] A method for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to an embodiment of the present invention includes the following steps:

[0081] S100. Obtain a grape hyperspectral data sample, where the grape hyperspectral data sample includes a hyperspectral data block of a whole bunch of grapes, and measure the glucose content of each grape and the average sugar content of the whole bunch of grapes in the grape hyperspectral data sample to obtain the true value of the measured sugar content.

[0082] In this embodiment, a grape hyperspectral data sample set is obtained. The number of samples is as sufficient as possible. During the sample collection process, grape samples can be collected in multiple scenarios such as before branch pruning, after branch pruning, before bad fruit pruning, and after bad fruit pruning, covering multiple links such as after picking and before sorting.

[0083] In one embodiment, the total number of grape bunches should be greater than 800 bunches.

[0084] Exemplarily, the embodiment of the present invention selects 1262 bunches, a total of 21,892 grapes, to form a single grape hyperspectral data sample set.

[0085] S200. Use a preset data processing method to preprocess and perform black and white frame correction on the hyperspectral data in the grape hyperspectral data sample to obtain a corrected hyperspectral data block.

[0086] In step S200, the hyperspectral data in the grape hyperspectral data sample is preprocessed and black and white frame corrected to obtain a corrected hyperspectral data block.

[0087] Further, step S200 includes the following steps S210 to S220:

[0088] S210. Use a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data sample and extract a single grape mask to obtain an enhanced hyperspectral data block.

[0089] In this embodiment, for the acquired data, three-channel bands are selected as the separation map, the image is subjected to Gaussian smoothing and noise reduction processing, and the whole bunch of grapes is segmented according to the deep model YOLO, so as to obtain the masks of the whole bunch and single grapes that filter the background and non-grape objects.

[0090] S220. Perform black and white frame correction on the hyperspectral data in the enhanced hyperspectral data block to obtain a corrected hyperspectral data block.

[0091] In this embodiment, the white frame and black frame data captured by the hyperspectral camera are obtained, and black and white frame correction is performed on the subsequent acquired hyperspectral samples:

[0092]

[0093] Wherein, I is the hyperspectral data before correction, W is the white frame, D is the black frame, I is the original data, and X is the hyperspectral data after correction. Further, for the hyperspectral data after correction, overexposed pixels, abnormal spikes or valleys such as sudden drops in spectral lines, and pixels without energy are filtered to remove interference factors such as abnormal spectra, abnormal pixels, and dead pixels.

[0094] S300. Construct a training set, a validation set, and a prediction set for single grape sugar content based on hyperspectral and self-supervised learning through the hyperspectral data in the corrected hyperspectral data block and the measured true sugar content value.

[0095] In this embodiment, the hyperspectral data, the mask, and the measured true sugar content value in the corrected hyperspectral data block are combined into a data set, and the data set is divided into a training set, a validation set, and a prediction set according to a ratio. The training set is used to train a deep model for glucose detection based on hyperspectral and self-supervised learning, the validation set is used to adjust the model hyperparameters, and the prediction set is used to evaluate the model performance and evaluate the model accuracy.

[0096] Exemplarily, the embodiment of the present invention collects a total of 1262 bunches of grapes, uses 1000 bunches as the training set to train the deep model, 131 bunches as the validation set to adjust the model hyperparameters, and the remaining 131 bunches as the prediction set to evaluate the model performance.

[0097] S400. Construct a single grape sugar content and average sugar content detection network based on Transformer according to the training set, the validation set, and the prediction set.

[0098] The single grape sugar content and average sugar content detection network based on Transformer consists of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.

[0099] In a specific embodiment, as Figure 2 shown, Figure 2 FIG. is a schematic structural diagram of a single grape sugar content and average sugar content detection network based on Transformer provided by an embodiment of the present invention. The single grape sugar content and average sugar content detection network based on Transformer is divided into three modules:

[0100] Module 1 is a single grape spectrum Transformer encoder, which is used to sort the input grapes from the top to the end in combination with position encoding, and learn to extract features of effective sugar content information for the spectra of single grapes at different positions;

[0101] Module 2 serves as a single - glucose - degree Transformer decoder, which is used to learn to predict the glucose degree of each grape through the attention mechanism;

[0102] Module 3 serves as an average - glucose - degree Transformer decoder for the whole bunch of grapes, which is used to learn to predict the average glucose degree of the whole bunch of grapes by combining the average spectrum of the whole bunch of grapes, the average - glucose - degree Token, and all the spectra of single grapes extracted.

[0103] Among them, each Transformer encoder and decoder is composed of a fully - connected layer, a batch normalization layer, a non - linear operation layer, an embedding layer, a position encoder, a self - attention layer, a multi - head attention layer, a normalization layer, a residual connection layer, and an output layer.

[0104] S500. Construct self - supervised upstream whole - bunch contrast and intra - bunch contrast learning tasks for glucose - degree prediction, and use the training set to train the Transformer - based single - glucose - degree and average - glucose - degree detection network.

[0105] In step S500, the construction of the self - supervised upstream whole - bunch contrast and intra - bunch contrast learning tasks for glucose - degree prediction includes the following steps:

[0106] S510. Construct two upstream task loss functions for self - supervised learning of glucose - degree prediction. Among them, the two upstream task loss functions are the grape - bunch discrimination task loss function and the grape - berry discrimination task loss function respectively;

[0107] The grape - bunch discrimination task loss function has the following calculation formula:

[0108]

[0109] where \(e\) is the exponential function, \(g(x\) i,* ;\(\theta\) h ,\(\theta\) s ) is the latent representation of a certain bunch of grapes, \(\theta\) h is the weight of the single - grape - spectrum Transformer encoder, \(\theta\) s is the pre - trained weight of the average - glucose - degree Transformer decoder for the whole bunch of grapes, \(S\) is the batch size of the grape - bunch discrimination task, and \(B\) is the batch size of the negative samples in this task;

[0110] The grape - berry discrimination task loss function has the following calculation formula:

[0111]

[0112] where \(e\) is the exponential function, \(f(x\) i,n ;\(\theta\) h ,\(\theta\) b) is the potential representation of a single grape, θ h are the weights of the spectral Transformer encoder for the single grape, θ b are the pre-trained weights of the model of the single grape sugar content Transformer decoder, N is the batch size of the grape grain discrimination task, and T is the batch size of the negative samples in this task;

[0113] The calculation formula of the objective function after combining the two is:

[0114] L SSL =(1 - λ)L 串 +λL 颗

[0115] Among them, the hyperparameter λ controls the adjustment of the learned representation.

[0116] S600. Construct the learning tasks of the downstream single grape sugar content and the average sugar content of the whole bunch, and use the training set to further train the Transformer-based single grape sugar content and average sugar content detection network.

[0117] In step S600, the construction of the learning tasks of the downstream single grape sugar content and the average sugar content of the whole bunch includes the steps:

[0118] S610. Construct the loss function of the downstream tasks of the single grape sugar content and the whole bunch grape sugar content in the downstream. Among them, the loss function of the downstream tasks uses the mean square error loss function, and the calculation formula of the mean square error loss function is:

[0119]

[0120] Among them, N is the total number of grape bunches, C is the number of grape grains in each bunch, b i,t is the single grape sugar content label, f(x i,t ; θ h ,θ b ) is the predicted value of the single grape sugar content of the t-th grape sample in the i-th bunch, is the average sugar content label of all grapes in the i-th bunch, g(x i ; θ h ,θ s ) is the predicted value of the sugar content of the grape sample in the i-th bunch.

[0121] S700. Fine-tune the trained Transformer-based single grape sugar content and average sugar content detection network according to the validation set and the prediction set to obtain a glucose content detection model that combines the single grape sugar content and the whole bunch grape sugar content.

[0122] In this embodiment, the trained Transformer-based single grape sugar content and average sugar content detection network is a preliminary glucose sugar content detection model. After obtaining the preliminary glucose sugar content detection model, it is necessary to verify the detection model in real time. The model established based on the training samples is used to predict the sugar content of newly collected grape samples, and the indexes between the predicted data and the measured data are calculated to verify the accuracy of the model, so as to evaluate whether the model is applicable to the sugar content detection of grapes.

[0123] Figure 3 It is a schematic diagram of the true value of glucose sugar content and the prediction result in an embodiment of the present invention. From the above Figure 3 results, it can be seen that the method based on hyperspectral and deep learning realizes the simultaneous detection of the sugar content of single grapes in a whole bunch, and the detection accuracy is high. Therefore, using the glucose sugar content detection model constructed by the present invention to detect grape samples with unknown size and weight has high detection accuracy and can meet the requirements of rapid and non-destructive detection of glucose sugar content in large-scale production.

[0124] In summary, the present invention combines grape hyperspectral with deep learning technology to model and select model parameters, constructs a hyperspectral and self-supervised learning Transformer glucose sugar content detection model, which can improve the accuracy and robustness of the model; in addition, through the hyperspectral and self-supervised learning Transformer glucose sugar content detection model, the glucose sugar content detection model can be used to quickly and non-destructively obtain the glucose sugar content of each grape and the average sugar content of the whole bunch of grapes, greatly reducing the labor, time, and financial costs of glucose sugar content measurement during the grape season, and realizing rapid and accurate sorting of grape quality for sugar content.

[0125] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a device for obtaining a glucose sugar content detection model based on hyperspectral and self-supervised learning provided by an embodiment of the present invention.

[0126] An apparatus for obtaining a glucose sugar content detection model based on hyperspectral and self-supervised learning according to an embodiment of the present invention includes:

[0127] A grape hyperspectral data sample and measured sugar content true value acquisition module, configured to acquire grape hyperspectral data samples, where the grape hyperspectral data samples include hyperspectral data blocks of the whole bunch of grapes, and measure the glucose sugar content of each grape and the average sugar content of the whole bunch of grapes in the grape hyperspectral data samples to obtain the measured sugar content true value;

[0128] A data processing module, configured to preprocess and correct black and white frames for the hyperspectral data in the grape hyperspectral data samples by using a preset data processing method to obtain corrected hyperspectral data blocks;

[0129] A sample data set construction module for constructing a training set, a validation set, and a prediction set of single grape sugar content based on hyperspectral and self-supervised learning by using the hyperspectral data in the corrected hyperspectral data block and the true value of the measured sugar content;

[0130] A network construction module for constructing a single grape sugar content and average sugar content detection network based on Transformer according to the training set, the validation set, and the prediction set;

[0131] An upstream training module for constructing self-supervised upstream whole bunch contrast and intra-bunch contrast learning tasks for glucose content prediction, and using the training set to train the single grape sugar content and average sugar content detection network based on Transformer;

[0132] A downstream training module for constructing learning tasks of downstream single grape sugar content and whole bunch average sugar content, and using the training set to further train the single grape sugar content and average sugar content detection network based on Transformer;

[0133] A glucose content detection model acquisition module for fine-tuning the trained single grape sugar content and average sugar content detection network based on Transformer according to the validation set and the prediction set to obtain a glucose content detection model that combines single grape glucose content and whole bunch glucose content.

[0134] The data processing module is specifically used for:

[0135] Using a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data sample and extract a single grape mask to obtain an enhanced processed hyperspectral data block;

[0136] Performing black and white frame correction on the hyperspectral data in the enhanced processed hyperspectral data block to obtain a corrected hyperspectral data block.

[0137] In a specific embodiment, the single grape sugar content and average sugar content detection network based on Transformer includes:

[0138] A single grape spectrum Transformer encoder for sorting the input grapes from the top to the end in combination with position encoding, and learning to extract features of effective sugar content information for single grape spectra at different positions;

[0139] A single grape sugar content Transformer decoder for learning to predict the sugar content of each grape through an attention mechanism;

[0140] The average sugar content Transformer decoder for a whole bunch of grapes is used to learn to combine the average spectrum of the whole bunch of grapes, the average sugar content Token, and all the spectra of individual grapes extracted to predict the average sugar content of the whole bunch of grapes;

[0141] Among them, each Transformer encoder and decoder is composed of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.

[0142] In a specific embodiment, the upstream training module is specifically used for:

[0143] Construct two upstream task loss functions for self-supervised learning of glucose content prediction, where the two upstream task loss functions are the grape bunch discrimination task loss function and the grape berry discrimination task loss function respectively;

[0144] The calculation formula of the grape bunch discrimination task loss function is:

[0145]

[0146] Among them, e is the exponential function, g(x i,* ; θ h , θ s ) is the latent representation of a certain bunch of grapes, θ h is the weight of the single grape spectrum Transformer encoder, θ s is the pre-trained weight of the model of the average sugar content Transformer decoder for the whole bunch of grapes, S is the batch size of the grape bunch discrimination task, and B is the batch size of the negative samples in this task;

[0147] The calculation formula of the grape berry discrimination task loss function is:

[0148]

[0149] Among them, e is the exponential function, f(x i,n ; θ h , θ b ) is the latent representation of a single grape, θ h is the weight of the single grape spectrum Transformer encoder, θ b is the pre-trained weight of the model of the single grape sugar content Transformer decoder, N is the batch size of the grape berry discrimination task, and T is the batch size of the negative samples in this task;

[0150] The calculation formula of the objective function after combining the two is:

[0151] L SSL= (1 - λ)L 串 + λL 颗

[0152] where the hyperparameter λ controls the adjustment of the learned representation;

[0153] The downstream training module is specifically configured to:

[0154] Construct a downstream task loss function for the downstream single grape glucose level and the whole bunch of grape glucose level. Among them, the downstream task loss function adopts the mean square error loss function, and the calculation formula of the mean square error loss function is:

[0155]

[0156] where N is the total number of grape bunches, C is the number of grapes per bunch, b i,t is the single grape glucose level label, f(x i,t ; θ h , θ b ) is the predicted value of the single grape glucose level of the t-th grape sample in the i-th bunch, is the average glucose level label of all grapes in the i-th bunch, g(x i ; θ h , θ s ) is the predicted value of the glucose level of the grape sample in the i-th bunch.

[0157] The glucose level detection model acquisition device based on hyperspectral and self-supervised learning provided by the embodiments of the present invention can execute all steps and functions of the glucose level detection model acquisition method based on hyperspectral and self-supervised learning provided in any of the above embodiments. The specific functions of this device will not be elaborated here.

[0158] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a terminal device provided by an embodiment of the present invention.

[0159] The terminal device includes: a processor, a memory, and a computer program stored in the memory and configured to be run by the processor. When the processor executes the computer program, it implements the steps of the glucose level detection model acquisition method based on hyperspectral and self-supervised learning in each of the above embodiments, such as Figure 1 the steps S100 - S700 shown. Or, when the processor executes the computer program, it implements the functions of each module in each of the above device embodiments.

[0160] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device. For example, the computer program may be divided into several modules, and the specific functions of each module have been described in detail in the method for obtaining a glucose detection model based on hyperspectral and self-supervised learning provided in any of the above embodiments. Therefore, the specific functions of this device will not be elaborated herein.

[0161] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.

[0162] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, and connects various parts of the entire terminal device through various interfaces and circuits.

[0163] The memory can be used to store the computer program and / or module. By running or executing the computer program and / or module stored in the memory, and invoking the data stored in the memory, the processor can implement various functions of the method for obtaining a glucose concentration detection model based on hyperspectral and self-supervised learning. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0164] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for obtaining a glucose concentration detection model based on hyperspectral and self-supervised learning in each of the above embodiments.

[0165] If the module integrated in the terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0166] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for obtaining a glucose concentration detection model based on hyperspectral and self-supervised learning, characterized in that, Including: Obtain grape hyperspectral data samples, where the grape hyperspectral data samples contain hyperspectral data blocks of whole bunches of grapes, and measure the glucose content of each grape and the average sugar content of the whole bunch of grapes in the grape hyperspectral data samples to obtain the true value of the measured sugar content; Use a preset data processing method to preprocess and perform black-and-white frame correction on the hyperspectral data in the grape hyperspectral data samples to obtain corrected hyperspectral data blocks; Construct a training set, a validation set, and a prediction set for the glucose content of single grapes based on hyperspectral and self-supervised learning through the hyperspectral data in the corrected hyperspectral data blocks and the true value of the measured sugar content; Construct a detection network for single-grape sugar content and average sugar content based on Transformer according to the training set, the validation set, and the prediction set; Construct a self-supervised upstream whole-bunch contrast and intra-bunch contrast learning task for glucose content prediction, and use the training set to train the detection network for single-grape sugar content and average sugar content based on Transformer; Construct a downstream learning task for single-grape sugar content and whole-bunch average sugar content, and use the training set to further train the detection network for single-grape sugar content and average sugar content based on Transformer; Fine-tune the trained detection network for single-grape sugar content and average sugar content based on Transformer according to the validation set and the prediction set to obtain a glucose content detection model that combines single-grape glucose content and whole-bunch glucose content.

2. The method for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 1, wherein The using a preset data processing method to preprocess and perform black-and-white frame correction on the hyperspectral data in the grape hyperspectral data samples to obtain corrected hyperspectral data blocks includes: Use a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data samples and extract single-grape masks to obtain enhanced hyperspectral data blocks; Perform black-and-white frame correction on the hyperspectral data in the enhanced hyperspectral data blocks to obtain corrected hyperspectral data blocks.

3. The method for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 1, wherein, The detection network for single-grape sugar content and average sugar content based on Transformer includes: A single-grape spectral Transformer encoder for sorting the input grapes from top to bottom in combination with position encoding and learning to extract features of effective sugar content information for single-grape spectra at different positions; A single-grape sugar content Transformer decoder for learning to predict the glucose content of each grape through an attention mechanism; A whole-bunch grape average sugar content Transformer decoder for learning to predict the average sugar content of the whole bunch of grapes by combining the average spectrum of the whole bunch of grapes, the average sugar content Token, and all the extracted single-grape spectra; Among them, each Transformer encoder and decoder is composed of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.

4. The method for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 3, wherein The constructing a self-supervised upstream whole-bunch contrast and intra-bunch contrast learning task for glucose content prediction includes: Construct two upstream task loss functions for self-supervised learning of glucose degree prediction, where the two upstream task loss functions are respectively the grape cluster discrimination task loss function and the grape berry discrimination task loss function; The calculation formula of the grape cluster discrimination task loss function is: where e is the exponential function, g(x i,* ; θ h , θ s ) is the potential representation of a bunch of grapes, θ h are the weights of the spectral Transformer encoder for a single grape, θ s are the pre-trained weights of the model of the average sugar content Transformer decoder for the whole bunch of grapes, S is the batch size of the grape bunch discrimination task, and B is the batch size of negative samples in this task; The calculation formula of the grape berry discrimination task loss function is: where e is the exponential function, f(x i,n ; θ h , θ b ) is the potential representation of a single grape, θ h are the weights of the spectral Transformer encoder of the single grape, θ b are the model pre-training weights of the glucose Transformer decoder of the single grape, N is the batch size of the grape discrimination task, and T is the batch size of the negative samples in this task; The calculation formula of the objective function after combining the two is: LSSL = (1 - λ)L 串 + λL 颗 Among them, the hyperparameter λ controls the adjustment of the learned representation.

5. The method for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 4, wherein The construction of the learning task for the single-berry sugar degree and the average sugar degree of the whole cluster includes: Construct a downstream task loss function for the single-berry glucose degree and the whole-cluster glucose degree. Among them, the downstream task loss function adopts the mean square error loss function, and the calculation formula of the mean square error loss function is: Among them, N is the total number of grape clusters, C is the number of grapes per cluster, and b i,t is the single grape sugar content label, f(x i,t ; θ h , θ b ) is the predicted single grape sugar content value of the t-th grape sample in the i-th cluster, is the average sugar content label of all grapes in the i-th cluster, g(x i ; θ h , θ s ) is the predicted sugar content value of the grape sample in the i-th cluster.

6. A device for obtaining a glucose concentration detection model based on hyperspectral and self-supervised learning, characterized in that, Including: A grape hyperspectral data sample and measured sugar degree true value acquisition module, which is used to acquire grape hyperspectral data samples. Among them, the grape hyperspectral data samples include hyperspectral data blocks of the whole cluster of grapes, and measure the glucose degree of each grape and the average sugar degree of the whole cluster of grapes in the grape hyperspectral data samples to obtain the measured sugar degree true value; A data processing module, which is used to preprocess and correct black and white frames for the hyperspectral data in the grape hyperspectral data samples by using a preset data processing method to obtain corrected hyperspectral data blocks; A sample data set construction module, which is used to construct a training set, a validation set and a prediction set for the single-berry glucose degree based on hyperspectral and self-supervised learning through the hyperspectral data in the corrected hyperspectral data blocks and the measured sugar degree true value; A network construction module, which is used to construct a single-berry sugar degree and average sugar degree detection network based on Transformer according to the training set, the validation set and the prediction set; An upstream training module, which is used to construct a self-supervised upstream whole-cluster contrast and intra-cluster contrast learning task for glucose degree prediction, and use the training set to train the single-berry sugar degree and average sugar degree detection network based on Transformer; A downstream training module, which is used to construct a learning task for the single-berry sugar degree and the average sugar degree of the whole cluster, and use the training set to further train the single-berry sugar degree and average sugar degree detection network based on Transformer; A glucose degree detection model acquisition module, which is used to fine-tune the trained single-berry sugar degree and average sugar degree detection network based on Transformer according to the validation set and the prediction set to obtain a glucose degree detection model that combines the single-berry glucose degree and the whole-cluster glucose degree.

7. The apparatus for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 6, wherein The data processing module is specifically used for: Using a preset data processing method to preprocess the hyperspectral data in the grape hyperspectral data samples and extract single-berry grape masks to obtain enhanced processed hyperspectral data blocks; Performing black and white frame correction on the hyperspectral data in the enhanced processed hyperspectral data blocks to obtain corrected hyperspectral data blocks.

8. The apparatus for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 6, wherein The single-berry sugar degree and average sugar degree detection network based on Transformer includes: Single grape spectral Transformer encoder, which is used to sort the input grapes from the top to the end in combination with positional encoding, and learn to extract the features of effective sugar content information from the spectra of single grapes at different positions; Single sugar content Transformer decoder, which is used to learn to predict the sugar content of each grape through the attention mechanism; Average sugar content Transformer decoder for the whole bunch of grapes, which is used to learn to combine the average spectrum of the whole bunch of grapes, the average sugar content Token and the spectra of all single grapes extracted, and predict the average sugar content of the whole bunch of grapes; Among them, each Transformer encoder and decoder is composed of a fully connected layer, a batch normalization layer, a non-linear operation layer, an embedding layer, a positional encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer and an output layer.

9. The device for obtaining a glucose detection model based on hyperspectral and self-supervised learning according to claim 8, characterized in that, The upstream training module is specifically used for: Construct two upstream task loss functions for self-supervised learning of grape sugar content prediction, where the two upstream task loss functions are the grape bunch discrimination task loss function and the grape berry discrimination task loss function respectively; The calculation formula of the grape bunch discrimination task loss function is: where, e is the exponential function, g(x i,* ; θ h , θ s ) is the potential representation of a bunch of grapes, θ h are the weights of the spectral Transformer encoder of the single grape, θ s are the pre-trained weights of the model of the average sugar content Transformer decoder of the bunch of grapes, S is the batch size of the grape bunch discrimination task, and B is the batch size of the negative samples in this task; The calculation formula of the grape berry discrimination task loss function is: where e is the exponential function, f(x i,n ; θ h , θ b ) is the potential representation of a single grape, θ h are the weights of the spectral Transformer encoder of the single grape, θ b are the pre-trained weights of the model of the single grape sugar content Transformer decoder, N is the batch size of the grape grain discrimination task, and T is the batch size of the negative samples in this task; The calculation formula of the objective function after combining the two is: L SSL = (1 - λ)L 串 + λL 颗 Among them, the hyperparameter λ controls the adjustment of the learned representation; The downstream training module is specifically used for: Construct a downstream task loss function for the single grape sugar content and the whole bunch of grape sugar content, where the downstream task loss function uses the mean square error loss function, and the calculation formula of the mean square error loss function is: Among them, N is the total number of grape clusters, C is the number of grapes per cluster, and b i,t is the single grape sugar content label, f(x i,t ; θ h , θ b ) is the predicted single grape sugar content value of the t-th grape sample in the i-th cluster, is the average sugar content label of all grapes in the i-th cluster, g(x i ; θ h , θ s ) is the predicted sugar content value of the grape sample in the i-th cluster.

10. A terminal device, characterized in that, Including: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a method for obtaining a glucose detection model based on hyperspectral and self-supervised learning as described in any one of claims 1 to 5.