Agricultural product detection method and device, electronic equipment and computer readable storage medium

CN117392432BActive Publication Date: 2026-09-22SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311187235.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2026-09-22
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

但是,针对近红外光谱和图像的融合无论是特征还是决策域的融合都没有很好地利用两个域和样本之间的关系,只是将两者的信息都利用起来,而现有的跨域特征融合模型都需要两个域的信息具有对应关系,一维近红外光谱无法与数字图像建立很好的对应关系,因此无法有效地对提取得到的特征进行融合处理

Benefits of technology

[0053]根据本发明实施例的农产品检测方法,至少具有如下有益效果:在进行农产品检测的过程中,首先获取农产品图像信息和农产品一维近红外光谱数据;接着基于格拉姆角场将农产品一维近红外光谱数据转换为农产品二维近红外光谱数据;接着基于预训练的卷积自编码器分别对农产品图像信息和农产品二维近红外光谱数据进行特征提取得到图像特征和近红外光谱特征;接着基于对比学习模型对图像特征和近红外光谱特征进行特征融合处理就可以得到农产品检测特征;最后对农产品检测特征进行分类处理就可以得到农产品检测分类结果。通过上述技术方案,能够实现近红外光谱和图像域的有效特征融合,提升农产品的检测速度以及提高分类的准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117392432B_ABST
    Figure CN117392432B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of agricultural product detection method, device, electronic equipment and computer readable storage medium.The method comprises: obtaining agricultural product image information and agricultural product one-dimensional near infrared spectrum data;Gram angle field is based on agricultural product one-dimensional near infrared spectrum data is converted into agricultural product two-dimensional near infrared spectrum data;Based on pre-trained convolutional autoencoder, agricultural product image information and agricultural product two-dimensional near infrared spectrum data are respectively extracted to obtain image feature and near infrared spectrum feature;Image feature and near infrared spectrum feature are based on contrast learning model and are carried out feature fusion processing to obtain agricultural product detection feature;Agricultural product detection feature is classified and processed to obtain agricultural product detection classification result.According to the scheme of the embodiment of the present application, effective feature fusion of near infrared spectrum and image domain can be realized, the detection speed of agricultural product is improved and the accuracy of classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural product classification and testing technology, and in particular to an agricultural product testing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Food hygiene and safety is a major issue affecting people's livelihoods and has received increasing attention and importance in recent years. Classifying agricultural products by quality, origin, and maturity is a crucial step in the commercial processing of harvested agricultural products. Unlike previous biochemical methods such as enzyme inhibition, rapidly developing non-destructive testing technologies can obtain internal and external information without damaging the object being tested, thus enabling the detection and classification of agricultural products. Various non-destructive testing technologies for agricultural products exist, including spectral detection, machine vision inspection, acoustic detection, mechanical detection, X-ray detection, electronic nose, and electronic tongue technologies. Digital image recognition technology uses computers to process images and extract image feature parameters such as color, shape, and texture. By analyzing images of agricultural products, its appearance information can be effectively extracted. Near-infrared spectroscopy offers advantages such as non-destructiveness, convenience, environmental friendliness, and safety. Near-infrared light is an electromagnetic wave between visible and mid-infrared light. Its main source is the absorption of overtones and combination frequencies of the XH vibration of hydrogen-containing groups. Its reflection information contains the composition and molecular structure information of most types of organic compounds. The spectral characteristics of a sample will change with changes in its internal composition or structure. Different samples will also have different absorption levels of near-infrared light at different frequencies. In some wavelength ranges, the absorption is weaker due to absorption, while in other wavelength ranges, the absorption is stronger due to non-absorption. Therefore, based on the different absorption peaks of various hydrocarbons (sugars, acids, moisture, vitamins, etc.) in the near-infrared band within agricultural products, it is possible to detect agricultural products.

[0003] Currently, there are various classification algorithms that use near-infrared spectroscopy and image fusion. There are two main approaches to fusing information between the two domains. The first involves extracting features from both the near-infrared spectrum and the image separately, fusing the extracted features, and then making a decision. The second approach involves performing fusion decision-making within the decision domain, i.e., making separate decisions in each domain first and then fusing the results. However, neither the feature fusion nor the decision domain fusion approach effectively utilizes the relationship between the two domains and the samples; it merely uses information from both. Existing cross-domain feature fusion models require a correspondence between the information from the two domains. One-dimensional near-infrared spectroscopy cannot establish a good correspondence with digital images, thus hindering the effective fusion of extracted features. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0005] Therefore, this invention proposes an agricultural product detection method that can effectively fuse features from near-infrared spectroscopy and the image domain, thereby improving the detection speed and classification accuracy of agricultural products.

[0006] The present invention also proposes an apparatus that applies the above-mentioned method for detecting agricultural products.

[0007] The present invention also proposes an electronic device that applies the above-mentioned method for detecting agricultural products.

[0008] The present invention also proposes a computer-readable storage medium that applies the above-described method for detecting agricultural products.

[0009] According to a first aspect of the present invention, the method for detecting agricultural products includes:

[0010] Acquire image information and one-dimensional near-infrared spectral data of agricultural products;

[0011] Based on the Gram angle field, the one-dimensional near-infrared spectral data of the agricultural products are converted into two-dimensional near-infrared spectral data of the agricultural products.

[0012] Based on a pre-trained convolutional autoencoder, feature extraction is performed on the agricultural product image information and the agricultural product two-dimensional near-infrared spectral data to obtain image features and near-infrared spectral features, respectively.

[0013] Based on a contrastive learning model, the image features and the near-infrared spectral features are fused to obtain agricultural product detection features;

[0014] The detection characteristics of the agricultural products are classified to obtain the agricultural product detection classification results.

[0015] According to some embodiments of the present invention, the conversion of the one-dimensional near-infrared spectral data of the agricultural product into two-dimensional near-infrared spectral data of the agricultural product based on the Gram angle field includes:

[0016] The one-dimensional near-infrared spectral data of the agricultural products are scaled to obtain near-infrared spectral scaled data.

[0017] The near-infrared spectral scaling data is subjected to coordinate transformation to obtain the two-dimensional near-infrared spectral data of the agricultural product.

[0018] According to some embodiments of the present invention, the step of performing feature fusion processing on the image features and the near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features includes:

[0019] The image features and the near-infrared spectral features are jointly processed to obtain the joint features;

[0020] The joint features are divided into multiple subspace features by subspace partitioning.

[0021] Multiple feature mapping results are obtained by mapping multiple subspace features based on a preset weight matrix, wherein each feature mapping result includes a projected feature matrix and an input feature matrix;

[0022] The attention weight matrix is ​​determined based on the projected feature matrix.

[0023] Multiple weight transformation feature matrices are obtained based on each of the attention weight matrices and the corresponding input feature matrices;

[0024] The feature joint matrix is ​​obtained by performing feature joint processing on multiple weight transformation feature matrices;

[0025] A consistency matrix is ​​obtained by performing a linear transformation on the joint feature matrix, and the consistency matrix is ​​used as the detection feature of the agricultural product.

[0026] According to some embodiments of the present invention, the step of classifying the detection features of the agricultural products to obtain the agricultural product detection classification result includes:

[0027] Based on a preset clustering algorithm, the sample features in the consistency matrix are classified to obtain multiple sample clusters;

[0028] The agricultural product detection and classification results are obtained by clustering multiple samples and labeling them with categories.

[0029] According to some embodiments of the present invention, the convolutional autoencoder includes an encoding component and a decoding component, and the training process of the convolutional autoencoder is as follows:

[0030] Acquire training samples, wherein the training samples include image sample data and near-infrared spectral sample data;

[0031] The image sample data and the near-infrared spectral sample data are respectively input into the encoding component to obtain image domain information data corresponding to the image sample data and near-infrared spectral domain information data corresponding to the near-infrared spectral sample data;

[0032] The image domain information data and the near-infrared spectral domain information data are respectively input into the decoding component to obtain the image domain decoding result corresponding to the image domain information data and the near-infrared spectral domain decoding result corresponding to the near-infrared spectral domain information data.

[0033] The image sample data, the image domain decoding result, the near-infrared spectral sample data, and the near-infrared spectral domain decoding result are input into a preset reconstruction loss algorithm to obtain a reconstruction loss value.

[0034] The parameters of the encoding and decoding components are adjusted based on the reconstruction loss value.

[0035] According to some embodiments of the present invention, the training process of the weight matrix is ​​as follows:

[0036] A linear transformation is performed on the image features to obtain an image domain representation matrix, and a linear transformation is performed on the near-infrared spectral features to obtain the near-infrared spectral domain representation matrix;

[0037] The image domain representation matrix and the near-infrared spectral domain representation matrix are determined as specific domain representation matrices;

[0038] The cosine similarity of the samples is obtained based on the consistency matrix and the specific domain representation matrix.

[0039] The sample relationship value is determined based on multiple subspace features;

[0040] The contrast loss value is determined based on the sample cosine similarity and the sample relationship value;

[0041] The weight matrix is ​​trained based on the contrast loss value.

[0042] According to some embodiments of the present invention, the projection feature matrix includes a first feature mapping matrix and a second feature mapping matrix, and the step of determining the attention weight matrix based on the projection feature matrix includes:

[0043] The transpose of the second feature mapping matrix is ​​used to obtain the transpose mapping matrix;

[0044] The attention weight matrix is ​​obtained based on the first feature mapping matrix, the transpose mapping matrix, and the preset activation function.

[0045] According to a second aspect of the present invention, an agricultural product testing apparatus includes:

[0046] The first processing module is used to acquire agricultural product image information and one-dimensional near-infrared spectral data of agricultural products.

[0047] The second processing module is used to convert the one-dimensional near-infrared spectral data of the agricultural products into two-dimensional near-infrared spectral data of the agricultural products based on the Gram angle field.

[0048] The third processing module is used to extract image features and near-infrared spectral features from the agricultural product image information and the agricultural product two-dimensional near-infrared spectral data based on a pre-trained convolutional autoencoder.

[0049] The fourth processing module is used to perform feature fusion processing on the image features and the near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features;

[0050] The fifth processing module is used to classify the detection features of the agricultural products to obtain the agricultural product detection classification results.

[0051] An electronic device according to a third aspect of the present invention includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the agricultural product detection method as described above.

[0052] According to a fourth aspect of the present invention, a computer-readable storage medium stores computer-executable instructions that, when executed by a control processor, implement the agricultural product detection method as described above.

[0053] The agricultural product detection method according to embodiments of the present invention has at least the following beneficial effects: In the process of agricultural product detection, firstly, image information and one-dimensional near-infrared spectral data of the agricultural product are acquired; then, the one-dimensional near-infrared spectral data of the agricultural product is converted into two-dimensional near-infrared spectral data of the agricultural product based on Gram angle field; next, feature extraction is performed on the image information and the two-dimensional near-infrared spectral data of the agricultural product using a pre-trained convolutional autoencoder to obtain image features and near-infrared spectral features, respectively; then, feature fusion processing is performed on the image features and near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features; finally, classification processing is performed on the agricultural product detection features to obtain the agricultural product detection classification result. Through the above technical solution, effective feature fusion of the near-infrared spectral and image domains can be achieved, improving the detection speed of agricultural products and increasing the accuracy of classification.

[0054] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0055] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0056] Figure 1 This is a flowchart of an agricultural product testing method provided in one embodiment of the present invention;

[0057] Figure 2This is a detailed flowchart of S200 provided in one embodiment of the present invention;

[0058] Figure 3 This is a detailed flowchart of S400 provided in one embodiment of the present invention;

[0059] Figure 4 This is a detailed flowchart of S500 provided in one embodiment of the present invention;

[0060] Figure 5 This is a flowchart illustrating the training process of a convolutional autoencoder according to an embodiment of the present invention.

[0061] Figure 6 This is a flowchart illustrating the training process of a weight matrix according to an embodiment of the present invention.

[0062] Figure 7 This is a flowchart illustrating the determination of the attention weight matrix according to an embodiment of the present invention;

[0063] Figure 8 This is a flowchart illustrating a specific method for detecting agricultural products according to an embodiment of the present invention;

[0064] Figure 9 This is a schematic diagram illustrating the working process of a sample feature fusion model provided in one embodiment of the present invention;

[0065] Figure 10 This is a schematic diagram of the structure of an agricultural product testing device provided in one embodiment of the present invention;

[0066] Figure 11 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0068] In the description of this invention, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0069] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0070] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for agricultural product detection. The method includes: firstly, acquiring image information and one-dimensional near-infrared spectral data of the agricultural product during detection; then, converting the one-dimensional near-infrared spectral data into two-dimensional near-infrared spectral data based on Gram angle field; next, extracting features from the image information and the two-dimensional near-infrared spectral data using a pre-trained convolutional autoencoder to obtain image features and near-infrared spectral features, respectively; then, performing feature fusion processing on the image features and near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features; finally, classifying the agricultural product detection features to obtain the agricultural product detection classification result. Through the above technical solution, effective feature fusion of the near-infrared spectral and image domains can be achieved, improving the detection speed and classification accuracy of agricultural products.

[0071] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0072] like Figure 1 As shown, Figure 1 This is a flowchart of an agricultural product testing method provided by an embodiment of the present invention. The method includes, but is not limited to, steps S100, S200, S300, S400, and S500.

[0073] Step S100: Obtain image information and one-dimensional near-infrared spectral data of agricultural products;

[0074] Step S200: Convert the one-dimensional near-infrared spectral data of agricultural products into two-dimensional near-infrared spectral data of agricultural products based on the Gram angle field;

[0075] Step S300: Based on the pre-trained convolutional autoencoder, feature extraction is performed on the agricultural product image information and the agricultural product two-dimensional near-infrared spectral data to obtain image features and near-infrared spectral features respectively;

[0076] Step S400: Based on the contrastive learning model, feature fusion processing is performed on image features and near-infrared spectral features to obtain agricultural product detection features;

[0077] Step S500: Classify the detection characteristics of agricultural products to obtain the agricultural product detection classification results.

[0078] It should be noted that the process of detecting agricultural products involves first acquiring image information and one-dimensional near-infrared spectral data of the agricultural products; then, converting the one-dimensional near-infrared spectral data into two-dimensional near-infrared spectral data based on Gram angle field; next, using a pre-trained convolutional autoencoder, extracting features from both the image information and the two-dimensional near-infrared spectral data to obtain image features and near-infrared spectral features; finally, using a contrastive learning model, fusing the image features and near-infrared spectral features to obtain the detection features of the agricultural products; and finally, classifying the detection features to obtain the classification results. This technical solution achieves effective feature fusion between the near-infrared spectral and image domains, improving the detection speed and classification accuracy of agricultural products.

[0079] It is worth noting that agricultural product image information can be obtained by capturing images of agricultural products using cameras, mobile phones, or other image acquisition devices; one-dimensional near-infrared spectral data of agricultural products can be obtained by detecting agricultural products using a near-infrared spectrometer. However, since the one-dimensional near-infrared spectral data of agricultural products is one-dimensional, it needs to be converted into two-dimensional data to facilitate subsequent feature fusion. In this embodiment, the one-dimensional near-infrared spectral data of agricultural products can be converted using Gram angle field to obtain two-dimensional near-infrared spectral data.

[0080] It is worth noting that by transforming one-dimensional near-infrared spectral data of agricultural products based on Gram angle fields, two-dimensional near-infrared spectral data of agricultural products can be obtained. Gram angle fields can convert time into images. Based on polar coordinates, Gram angle fields generate Gram matrices through polar coordinate encoding. If polar coordinates are not used and Gram matrices are used directly, the images will be noisy and the data will be normally or Gaussian distributed, resulting in insufficient dimensionality of the data. While temporal information is present, deep connections are lacking. Using polar coordinates can increase angular information, making the data more sparse while simultaneously increasing its dimensionality.

[0081] It is worth noting that a convolutional autoencoder (CAE) is an unsupervised deep learning algorithm that learns an encoded representation of the input data and then reconstructs the same input into an output. It consists of two networks: an encoder and a decoder. The encoder compresses the high-dimensional input into a low-dimensional latent code, also known as the latent code or encoding space, to extract the most relevant information from it, while the decoder decompresses the encoded data and recreates the original input. The CAE maximizes information and minimizes reconstruction error during encoding. The reconstruction error, also known as the reconstruction loss, is typically the mean squared error between the reconstructed input and the original input when the input is real-valued. If the input data is categorical data, the loss function used is the cross-entropy loss.

[0082] It is worth noting that feature fusion processing of image features and near-infrared spectral features based on a contrastive learning model can yield agricultural product detection features. During the feature fusion process, cross-domain and cross-sample feature fusion is employed between the near-infrared spectral domain and the digital image domain, utilizing the global structure between samples and the joint information of the two domains to improve the effectiveness of feature fusion. In the embodiments of this application, a combination of near-infrared spectroscopy and digital images is used for agricultural product classification to increase the amount of information and effectively improve the accuracy of agricultural product classification.

[0083] It is worth noting that, in this embodiment, the near-infrared spectrum is converted into a two-dimensional digital image, and a convolutional autoencoder is used to extract features from the image and near-infrared spectrum information, leveraging the advantages of convolutional networks to extract the most comprehensive spatial information features. Secondly, cross-domain and cross-sample feature fusion is adopted between the near-infrared spectral domain and the digital image domain, utilizing the global structure between samples and the joint information of the two domain features to improve the effectiveness of feature fusion. Finally, for training processes that do not require training samples and manually labeled tags, training is completed through self-supervised and unsupervised training, reducing human and material costs.

[0084] Additionally, in one embodiment, such as Figure 2 As shown, step S200 may include, but is not limited to, steps S210 and S220.

[0085] Step S210: Scale the one-dimensional near-infrared spectral data of agricultural products to obtain scaled near-infrared spectral data;

[0086] Step S220: Perform coordinate transformation on the near-infrared spectral scaling data to obtain two-dimensional near-infrared spectral data of agricultural products.

[0087] It should be noted that in the process of converting one-dimensional near-infrared spectral data of agricultural products into two-dimensional near-infrared spectral data, the one-dimensional near-infrared spectrum of agricultural products is first scaled to obtain scaled near-infrared spectral data; then, coordinate transformation is performed on the near-infrared spectral data to obtain the two-dimensional near-infrared spectral data of agricultural products. Converting one-dimensional data into two-dimensional data makes subsequent feature fusion more efficient and accurate.

[0088] It is worth noting that during the scaling process of one-dimensional near-infrared spectral data of agricultural products, the data range can be scaled to [-1,1] or [0,1]. During the coordinate transformation process of the scaled near-infrared spectral data, the scaled sequence data can be converted to a polar coordinate system, that is, the numerical value is regarded as the cosine of the included angle and the timestamp is regarded as the radius.

[0089] Additionally, in one embodiment, such as Figure 3As shown, the above step S400 may include, but is not limited to, steps S410, S420, S430, S440, S450, S460 and S470.

[0090] Step S410: Jointly process the image features and near-infrared spectral features to obtain joint features;

[0091] Step S420: Subspace partitioning is performed on the joint features to obtain multiple subspace features;

[0092] Step S430: Based on the preset weight matrix, the features of multiple subspaces are mapped to obtain multiple feature mapping results, wherein each feature mapping result includes a projected feature matrix and an input feature matrix;

[0093] Step S440: Determine the attention weight matrix based on the projected feature matrix;

[0094] Step S450: Based on each attention weight matrix and the corresponding input feature matrix, obtain multiple weight transformation feature matrices;

[0095] Step S460: Perform feature joint processing on multiple weight transformation feature matrices to obtain a feature joint matrix;

[0096] Step S470: Perform a linear transformation on the joint feature matrix to obtain a consistency matrix, and use the consistency matrix as the detection feature for agricultural products.

[0097] It should be noted that in the process of obtaining agricultural product detection features by fusing image features and near-infrared spectral features based on a contrastive learning model, the image features and near-infrared spectral features are first jointly processed to obtain joint features; then, the joint features are subdivided to obtain multiple subspace features; multiple subspace features are mapped based on a preset weight matrix to obtain multiple feature mapping results, where each feature mapping result includes a projected feature matrix and an input feature matrix; next, an attention weight matrix is ​​determined based on the projected feature matrix; then, multiple weight transformation feature matrices are obtained based on each attention weight matrix and the corresponding input feature matrix; then, multiple weight transformation feature matrices are jointly processed to obtain a feature joint matrix; finally, a linear transformation is performed on the feature joint matrix to obtain a consistency matrix, which is used as the agricultural product detection feature.

[0098] It is worth noting that since samples of the same category should be similar, a sample should be reinforced by other samples in the same category. This requires learning the relationships between samples to help samples in the same category reinforce each other, thereby using global sample information as well as cross-domain information, including near-infrared spectroscopy and digital images, to increase classification accuracy.

[0099] Additionally, in one embodiment, such as Figure 4 As shown, step S500 may include, but is not limited to, steps S510 and S520.

[0100] Step S510: Classify the sample features in the consistency matrix based on a preset clustering algorithm to obtain multiple sample clusters;

[0101] Step S520: Perform category labeling on multiple sample clusters to obtain the agricultural product detection and classification results.

[0102] It should be noted that in the process of classifying the detection features of agricultural products, the sample features in the consistency matrix are first classified based on the preset clustering algorithm, so that multiple sample clusters can be obtained; then, the category labeling process of the obtained multiple sample clusters can be performed to obtain the agricultural product detection classification results.

[0103] It is worth noting that a pre-defined clustering algorithm is used to classify the sample features in the consistency representation matrix. Through continuous iteration, multiple samples are divided into multiple clusters, ensuring that each point belongs to the cluster corresponding to the nearest mean. After clustering, manual sampling can be used to determine the specific actual category corresponding to each category, thus classifying the already clustered samples. This technical solution makes the classification process of agricultural products more accurate and efficient.

[0104] Additionally, in one embodiment, such as Figure 5 As shown, a convolutional autoencoder includes an encoding component and a decoding component, and the training process of a convolutional autoencoder may include, but is not limited to, the following steps.

[0105] Step S310: Obtain training samples, wherein the training samples include image sample data and near-infrared spectral sample data;

[0106] Step S320: Input the image sample data and near-infrared spectral sample data into the encoding component respectively to obtain image domain information data corresponding to the image sample data and near-infrared spectral domain information data corresponding to the near-infrared spectral sample data.

[0107] Step S330: Input the image domain information data and the near-infrared spectral domain information data into the decoding component respectively to obtain the image domain decoding result corresponding to the image domain information data and the near-infrared spectral domain decoding result corresponding to the near-infrared spectral domain information data.

[0108] Step S340: Input the image sample data, image domain decoding result, near-infrared spectral sample data, and near-infrared spectral domain decoding result into the preset reconstruction loss algorithm to obtain the reconstruction loss value;

[0109] Step S350: Adjust the parameters of the encoding and decoding components based on the reconstruction loss value.

[0110] It should be noted that during the training process of the convolutional autoencoder, training samples are first acquired, which may include image sample data and near-infrared spectral sample data. Next, the image sample data and near-infrared spectral sample data are input into the encoding component to obtain image domain information data corresponding to the image sample data and near-infrared spectral domain information data corresponding to the near-infrared spectral sample data, respectively. Then, the image domain information data and near-infrared spectral domain information data are input into the decoding component to obtain image domain decoding results and near-infrared spectral domain decoding results, respectively. Finally, the image sample data, image domain decoding results, near-infrared spectral sample data, and near-infrared spectral domain decoding results are input into a preset reconstruction loss algorithm to obtain the reconstruction loss value. Finally, based on the reconstruction loss value, the parameters of the encoding and decoding components are adjusted to ensure the convolutional autoencoder can be trained, preparing for subsequent feature extraction and enabling faster and more accurate feature extraction.

[0111] Additionally, in one embodiment, such as Figure 6 As shown, the training process of the weight matrix may include, but is not limited to, the following steps.

[0112] Step S431: Perform a linear transformation on the image features to obtain the image domain representation matrix, and perform a linear transformation on the near-infrared spectral features to obtain the near-infrared spectral domain representation matrix;

[0113] Step S432: Determine the image domain representation matrix and the near-infrared spectral domain representation matrix as specific domain representation matrices;

[0114] Step S433: Obtain the sample cosine similarity based on the consistency matrix and the specific domain representation matrix;

[0115] Step S434: Determine the sample relationship value based on multiple subspace features;

[0116] Step S435: Determine the contrast loss value based on the sample cosine similarity and sample relationship value;

[0117] Step S436: Train the weight matrix based on the contrast loss value.

[0118] It should be noted that the training process of the weight matrix can be as follows: First, a linear transformation is performed on the image features to obtain the image domain representation matrix, and a linear transformation is performed on the near-infrared spectral features to obtain the near-infrared spectral domain representation matrix. Next, the image domain representation matrix and the near-infrared spectral domain representation matrix are determined as specific domain representation matrices. Then, the sample cosine similarity is obtained based on the consistency matrix and the specific domain representation matrix. Next, the sample relationship value is determined based on multiple subspace features. Then, the contrast loss value is determined based on the sample cosine similarity and the sample relationship value. Finally, the weight matrix is ​​trained based on the contrast loss value. This training and adjustment of the weight matrix improves the effectiveness of subsequent feature fusion.

[0119] Additionally, in one embodiment, such as Figure 7 As shown, the projection feature matrix includes a first feature mapping matrix and a second feature mapping matrix, and the above step S440 may include, but is not limited to, steps S441 and S442.

[0120] Step S441: Transpose the second feature mapping matrix to obtain the transposed mapping matrix;

[0121] Step S442: Obtain the attention weight matrix based on the first feature mapping matrix, the transpose mapping matrix, and the preset activation function.

[0122] It should be noted that in determining the attention weight matrix based on the projected feature matrix, the transpose of the second feature mapping matrix is ​​first performed to obtain the transpose mapping matrix; then, the attention weight matrix is ​​obtained based on the first feature mapping matrix, the transpose mapping matrix, and the preset activation function. The attention weight matrix reflects the relationships between samples in the matrix.

[0123] To more clearly illustrate the specific process of the agricultural product testing method provided in the embodiments of the present invention, specific examples are given below.

[0124] like Figure 8 As shown, GAF represents Gram angle field, CAE represents convolutional autoencoder, Decoder represents decoder, Cat represents concatenating feature matrices, and Linear represents linear transformation.

[0125] The technical solution adopted in the embodiments of this application includes the following steps:

[0126] Step 1: Obtain information from different samples, including image information and near-infrared spectral information:

[0127] First, a camera is used to capture images of agricultural products, and then the image information from different samples is concatenated to obtain the input X of the image domain. 1 Simultaneously, near-infrared spectra of the same batch of agricultural products were collected using a near-infrared spectrometer. Since near-infrared spectra are one-dimensional data, and building deep learning models for one-dimensional data is difficult, while convolutional networks have achieved great success in processing two-dimensional image classification tasks, Gram's angle field can be used to convert the one-dimensional near-infrared spectrum into a two-dimensional image to obtain the input X in the near-infrared spectral domain. 2 This allows for more effective use of computer vision processing techniques to extract feature information.

[0128] Step 2: Feature extraction using a convolutional autoencoder.

[0129] The initial input data typically contains redundant information and random noise, therefore a convolutional autoencoder is needed for feature extraction and redundant information removal. Image features and near-infrared spectral features are generated by their respective convolutional encoders.

[0130]

[0131]

[0132] As shown in the formula above, where and f represents the features extracted from the i-th sample after convolution. 1 and f 2 Encoders representing the image domain and near-infrared spectral domain. After concatenation, image features Z are obtained respectively. 1 and near-infrared spectral characteristics Z 2 And the total feature Z. X1i should refer to the input in the image domain, and X2i should refer to the input in the near-infrared spectral domain.

[0133] Step 3: Utilize contrastive learning to achieve cross-domain and cross-sample feature fusion.

[0134] Since samples of the same category should be similar, a sample should be reinforced by other samples in the same category. This requires learning the relationships between samples to help samples in the same category reinforce each other, thereby using global sample information as well as cross-domain information, including near-infrared spectroscopy and digital images, to increase the accuracy of classification.

[0135] Reference Figure 9 As shown, Figure 9 This is a cross-domain and cross-sample feature fusion model. W 11 W 12 , W 21 W22 , W 31 W 32 , These are weight matrices, and features from different domains are fused and projected into different spaces using these different weight matrices. Q 11 Q 12 Q 21 Q 22 Q 31 Q 32 The new feature matrices are used to calculate similarity and obtain attention weight matrices S1, S2, and S3. R1, R2, and R3 are of the same length as the input sequence and store the feature information of the input matrix. Z1, Z2, and Z3 are the feature matrices after weight transformation. The features are then combined to obtain... for The consistency matrix obtained after linear transformation.

[0136] The image above shows a model for cross-domain and cross-sample feature fusion based on a multi-head attention mechanism. The input is the total feature Z, Z = [Z1, Z2]. The attention mechanism in the transformer is then used to fuse features from different domains and samples. The model is divided into three heads, forming three subspaces, allowing each subspace to focus on different aspects of information. Taking the first subspace as an example, for each feature mapping subspace, different weights are first used to map the fused features from different domains to a new feature space, resulting in three different feature mapping results: Q... 11 Q 12 R1. The feature map can be represented as:

[0137]

[0138] Therefore, the result of the feature mapping can be obtained as follows: Q 11 =ZW 11 W 12 =ZW 12 , Q 21 =ZW 21 Q 22 =ZW 22 , Q 31 =ZW 31 Q 32 =ZW 32 .

[0139] Both the Q matrix and the R matrix are results of fusing cross-domain features and mapping them to a new feature space. 11 Q 12 Q21 Q 22 Q 31 Q 32 It is used to further calculate the global structural relationships between samples.

[0140]

[0141]

[0142]

[0143] According to the above formula, S1, S2, and S3 are the global structure relation matrices corresponding to the three heads. Dividing by d is to adjust the scale and prevent the input value from being too large, causing the gradient of the softmax function to tend to 0. In the S matrix, S... ij S represents the relationship between sample i and sample j. ij The larger the value, the greater the probability that sample i and sample j belong to the same class.

[0144]

[0145] As shown in the formula above, R1, R2, and R3 are representation matrices. Using S1, S2, and S3 as attention matrices, we recombine the representation matrices based on the correlation between samples to obtain Z1 = S1R1, Z2 = S2R2, and Z3 = S3R3. Finally, we concatenate Z1, Z2, and Z3 together to obtain... The final features are obtained through cross-domain and cross-sample feature fusion. To reduce redundant information generated during feature fusion, a linear transformation is needed to remove some redundant information to obtain the final consistent representation matrix. Where n represents the total number of samples, d f This represents the length of the feature vector of each sample obtained after the linear transformation.

[0146] Step 4, Classification based on clustering algorithm

[0147] The K-means algorithm is used to classify the sample features in the consistency representation matrix. Through iterative processing, n samples are divided into C clusters, ensuring that each point belongs to the cluster corresponding to the nearest mean. C represents the number of categories to be classified. After clustering, manual sampling can be used to determine the specific actual category corresponding to each category, thus classifying the already clustered samples.

[0148] Furthermore, the training model can be performed as follows:

[0149] A convolutional autoencoder is an unsupervised model that can effectively extract features from an image and remove redundant information. The features obtained during encoding and decoding, and the decoded result, can be represented as follows:

[0150]

[0151]

[0152]

[0153]

[0154] According to the above formula, where Information data representing the d-th field of the i-th sample. This represents the result of recovering the d-th domain of the i-th sample after decoding. f and g represent the encoder and decoder, respectively. The superscripts 1 and 2 represent the domain of the digital image and the domain of the near-infrared spectrum, respectively. The reconstruction loss is defined by the following formula, which is used to train the network to better reconstruct the input image.

[0155]

[0156] For the final consistency representation matrix And the domain-specific (image domain and near-infrared spectral domain) representation matrix H 1 and H 2 The features mapped to samples belonging to the same category should be consistent. In addition, to meet the classification requirements, the features mapped to samples belonging to the same category should also be consistent. This can be achieved through a contrastive loss guided by structural relationships, defined as follows (13):

[0157]

[0158] in, Cosine similarity between features in the consistent representation matrix of the i-th sample in the d-th domain and features in the specific representation matrix of the i-th sample in the d-th domain. This relates to the relationship between the i-th sample and the j-th sample in the k-th head. (The rest of the text is incomplete and cannot be translated.) The introduction of this can force the features in the representation matrices of samples belonging to the same category to be more similar, when When the value is large, it indicates a strong relationship between the i-th and j-th samples. As can be seen from the contrastive learning formula above, this can preserve... It will not decrease, thus ensuring consistency among samples in the same category, while when When the value is small, it indicates a weak relationship between the i-th and j-th samples. In this case, the contrastive loss can force... The similarity is smaller, thus reducing the similarity between samples in different categories. Therefore, the contrastive loss guided by global structure increases the consistency of representations from the same sample and category while ensuring that the similarity is reduced only between consistent representations from different categories and representations in a specific domain, thereby achieving classification.

[0159] Total loss is defined as:

[0160]

[0161] α and β represent hyperparameters, where The encoder and decoder for two domains are trained using the difference between the original data and the reconstructed data. The linear transformation matrix for feature mapping is trained by reducing the similarity between consistent representation matrices from different categories and domain-specific representation matrices.

[0162] The above technical solution employs a convolutional autoencoder to extract features from images and two-dimensional near-infrared spectral images converted using Gram angle field transformation. The extracted digital image and near-infrared spectral features are then fused across domains and samples based on a global structure. First, this application converts the near-infrared spectrum into a two-dimensional digital image and uses a convolutional autoencoder to extract features from the image and near-infrared spectral information, leveraging the advantages of convolutional networks to extract the most comprehensive spatial information features. Second, it employs cross-domain and cross-sample feature fusion between the near-infrared spectral domain and the digital image domain, utilizing the global structure between samples and the joint information of the two domains to improve the effectiveness of feature fusion. Finally, for training processes that do not require training samples or manually labeled data, training is completed through self-supervised and unsupervised training, reducing human and material costs.

[0163] In some embodiments of the present invention, such as Figure 10 As shown, one embodiment of the present invention also provides an agricultural product testing device 10, which includes:

[0164] The first processing module 100 is used to acquire agricultural product image information and one-dimensional near-infrared spectral data of agricultural products.

[0165] The second processing module 200 is used to convert one-dimensional near-infrared spectral data of agricultural products into two-dimensional near-infrared spectral data of agricultural products based on Gram angle field.

[0166] The third processing module 300 is used to extract features from agricultural product image information and agricultural product two-dimensional near-infrared spectral data based on a pre-trained convolutional autoencoder to obtain image features and near-infrared spectral features respectively.

[0167] The fourth processing module 400 is used to perform feature fusion processing on image features and near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features;

[0168] The fifth processing module 500 is used to classify the detection characteristics of agricultural products to obtain the detection classification results of agricultural products.

[0169] The specific implementation of the agricultural product testing device 10 is basically the same as the specific implementation of the above-mentioned agricultural product testing method, and will not be described again here.

[0170] In some embodiments of the present invention, such as Figure 11 As shown, one embodiment of the present invention also provides an electronic device 700, including: a memory 720, a processor 710, and a computer program stored in the memory 720 and executable on the processor 710. When the processor 710 executes the computer program, it implements the agricultural product detection method described in the above embodiment, for example, performing the above-described... Figure 1 Method steps S100 to S500 Figure 2 Method steps S210 to S220, Figure 3 Method steps S410 to S470 Figure 4 Method steps S510 to S520 Figure 5 Method steps S310 to S350 Figure 6 Method steps S431 to S436 and Figure 7 The method steps S441 to S442.

[0171] In some embodiments of the present invention, one embodiment further provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor or controller, for example, by a processor in the above-described device embodiments, causing the processor to perform the agricultural product detection method described above, for example, performing the above-described... Figure 1 Method steps S100 to S500 Figure 2 Method steps S210 to S220, Figure 3 Method steps S410 to S470 Figure 4 Method steps S510 to S520 Figure 5 Method steps S310 to S350 Figure 6 Method steps S431 to S436 and Figure 7 The method steps S441 to S442.

[0172] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0173] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A method for detecting agricultural products, characterized in that, The method includes: Acquire image information and one-dimensional near-infrared spectral data of agricultural products; Based on the Gram angle field, the one-dimensional near-infrared spectral data of the agricultural products are converted into two-dimensional near-infrared spectral data of the agricultural products. Based on a pre-trained convolutional autoencoder, feature extraction is performed on the agricultural product image information and the agricultural product two-dimensional near-infrared spectral data to obtain image features and near-infrared spectral features, respectively. Based on a contrastive learning model, the image features and the near-infrared spectral features are fused to obtain agricultural product detection features; The agricultural product detection characteristics are classified to obtain agricultural product detection classification results; The step of obtaining agricultural product detection features by fusing the image features and the near-infrared spectral features based on a contrastive learning model includes: The image features and the near-infrared spectral features are jointly processed to obtain the joint features; The joint features are divided into multiple subspace features by subspace partitioning. Multiple feature mapping results are obtained by mapping multiple subspace features based on a preset weight matrix, wherein each feature mapping result includes a projected feature matrix and an input feature matrix; The attention weight matrix is ​​determined based on the projected feature matrix. Multiple weight transformation feature matrices are obtained based on each of the attention weight matrices and the corresponding input feature matrices; The feature joint matrix is ​​obtained by performing feature joint processing on multiple weight transformation feature matrices; A consistency matrix is ​​obtained by performing a linear transformation on the joint feature matrix, and the consistency matrix is ​​used as the detection feature of the agricultural product.

2. The method for detecting agricultural products according to claim 1, characterized in that, The conversion of the one-dimensional near-infrared spectral data of agricultural products into two-dimensional near-infrared spectral data based on Gram angle field includes: The one-dimensional near-infrared spectral data of the agricultural products are scaled to obtain near-infrared spectral scaled data. The near-infrared spectral scaling data is subjected to coordinate transformation to obtain the two-dimensional near-infrared spectral data of the agricultural product.

3. The method for detecting agricultural products according to claim 1, characterized in that, The process of classifying the detection features of agricultural products to obtain agricultural product classification results includes: Based on a preset clustering algorithm, the sample features in the consistency matrix are classified to obtain multiple sample clusters; The agricultural product detection and classification results are obtained by clustering multiple samples and labeling them with categories.

4. The method for detecting agricultural products according to claim 1, characterized in that, The convolutional autoencoder includes an encoding component and a decoding component, and the training process of the convolutional autoencoder is as follows: Acquire training samples, wherein the training samples include image sample data and near-infrared spectral sample data; The image sample data and the near-infrared spectral sample data are respectively input into the encoding component to obtain image domain information data corresponding to the image sample data and near-infrared spectral domain information data corresponding to the near-infrared spectral sample data; The image domain information data and the near-infrared spectral domain information data are respectively input into the decoding component to obtain the image domain decoding result corresponding to the image domain information data and the near-infrared spectral domain decoding result corresponding to the near-infrared spectral domain information data. The image sample data, the image domain decoding result, the near-infrared spectral sample data, and the near-infrared spectral domain decoding result are input into a preset reconstruction loss algorithm to obtain a reconstruction loss value. The parameters of the encoding and decoding components are adjusted based on the reconstruction loss value.

5. The method for detecting agricultural products according to claim 1, characterized in that, The training process for the weight matrix is ​​as follows: A linear transformation is performed on the image features to obtain an image domain representation matrix, and a linear transformation is performed on the near-infrared spectral features to obtain the near-infrared spectral domain representation matrix; The image domain representation matrix and the near-infrared spectral domain representation matrix are determined as specific domain representation matrices; The cosine similarity of the samples is obtained based on the consistency matrix and the specific domain representation matrix. The sample relationship value is determined based on multiple subspace features; The contrast loss value is determined based on the sample cosine similarity and the sample relationship value; The weight matrix is ​​trained based on the contrast loss value.

6. The method for detecting agricultural products according to claim 1, characterized in that, The projection feature matrix includes a first feature mapping matrix and a second feature mapping matrix. The step of determining the attention weight matrix based on the projection feature matrix includes: The transpose of the second feature mapping matrix is ​​used to obtain the transpose mapping matrix; The attention weight matrix is ​​obtained based on the first feature mapping matrix, the transpose mapping matrix, and the preset activation function.

7. An agricultural product testing device, characterized in that, The device includes: The first processing module is used to acquire agricultural product image information and one-dimensional near-infrared spectral data of agricultural products. The second processing module is used to convert the one-dimensional near-infrared spectral data of the agricultural products into two-dimensional near-infrared spectral data of the agricultural products based on the Gram angle field. The third processing module is used to extract image features and near-infrared spectral features from the agricultural product image information and the agricultural product two-dimensional near-infrared spectral data based on a pre-trained convolutional autoencoder. The fourth processing module is used to perform feature fusion processing on the image features and the near-infrared spectral features based on a contrastive learning model to obtain agricultural product detection features; The fifth processing module is used to classify the detection features of the agricultural products to obtain the agricultural product detection classification results; The step of obtaining agricultural product detection features by fusing the image features and the near-infrared spectral features based on a contrastive learning model includes: The image features and the near-infrared spectral features are jointly processed to obtain the joint features; The joint features are divided into multiple subspace features by subspace partitioning. Multiple feature mapping results are obtained by mapping multiple subspace features based on a preset weight matrix, wherein each feature mapping result includes a projected feature matrix and an input feature matrix; The attention weight matrix is ​​determined based on the projected feature matrix. Multiple weight transformation feature matrices are obtained based on each of the attention weight matrices and the corresponding input feature matrices; The feature joint matrix is ​​obtained by performing feature joint processing on multiple weight transformation feature matrices; A consistency matrix is ​​obtained by performing a linear transformation on the joint feature matrix, and the consistency matrix is ​​used as the detection feature of the agricultural product.

8. An electronic device, characterized in that, include: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the agricultural product detection method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the control processor, they implement the agricultural product detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Fusion method for optimization results of different near infrared spectrum variables and application

    CN107179292A

  • Drug classification method based on self-encoding and extreme learning machine

    CN113627554A