A method and system for detecting tomato maturity

By combining the illumination calibration and feature fusion technology of digital cameras and hyperspectral imaging systems, the accuracy and efficiency issues of tomato maturity detection were solved, and efficient and accurate detection was achieved under different lighting environments.

CN120219858BActive Publication Date: 2025-09-16FUJIAN PROV AGRI MACHANIZATION INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510694847.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-16
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing tomato maturity detection technology is sensitive to lighting conditions, has poor accuracy and low efficiency. Machine vision technology relies on the color and texture characteristics of tomatoes, and hyperspectral imaging technology requires manual analysis, which is time-consuming and labor-intensive.

Method used

Combining a digital camera and a hyperspectral imaging system, RGB image features are extracted through illumination calibration and spectral response function mapping methods. An image encoder is used for feature extraction, and a dynamic gated fusion mechanism and a spectrally sensitive attention mechanism are used to achieve cross-modal feature fusion. Finally, a maturity decoder generates detection results.

Benefits of technology

It achieves accurate and rapid tomato maturity detection under different lighting environments, improves detection accuracy and efficiency, reduces dependence on lighting environment, and is suitable for a variety of scenarios and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219858B_ABST
    Figure CN120219858B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for detecting tomato maturity. The method performs illumination calibration on a first real-time tomato image captured by a digital camera. A spectral response function mapping method is used to extract N real-time RGB images corresponding to N preset indicators from a second real-time tomato image captured by a spectral imaging system. Image feature extraction is performed on the first real-time tomato image and the N real-time RGB images after illumination calibration to obtain a first feature and N second features. A first fusion weight for the first feature and a second fusion weight for the N second features are dynamically adjusted based on the illumination calibration first real-time tomato image, the N real-time RGB images, and a dynamic gated fusion mechanism. A spectral sensitivity matrix and an attention mechanism are introduced into the fusion of the first and second features to achieve cross-modal feature fusion. The resulting fused multimodal features are decoded to obtain a maturity result. Consequently, the present invention improves the accuracy of maturity detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and system for detecting tomato maturity. Background Art

[0002] Tomato maturity detection is a critical step in agricultural production, harvesting, and quality control. Picking unripe or overripe tomatoes can affect consumer purchasing intentions and eating experience. Therefore, accurate tomato maturity detection can help improve tomato quality. Existing methods for tomato maturity detection primarily rely on machine vision or hyperspectral imaging. However, machine vision relies on characteristics such as tomato color and texture, and is highly sensitive to environmental factors such as light, which affects the accuracy of tomato maturity detection. Hyperspectral imaging, on the other hand, requires extensive and complex manual analysis, which is time-consuming and labor-intensive, impacting the efficiency of tomato maturity detection. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: the present invention provides a method and system for detecting tomato maturity, which improves the accuracy of tomato maturity detection while improving the detection efficiency.

[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0005] In a first aspect, the present invention provides a method for detecting tomato maturity, comprising:

[0006] Acquire a first real-time tomato image captured by a digital camera and a second real-time tomato image captured by a hyperspectral imaging system, perform illumination calibration on the first real-time tomato image to obtain a first real-time tomato image after illumination calibration, and simultaneously extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image using a spectral response function mapping method;

[0007] Inputting the first real-time tomato image after illumination calibration into a first image encoder for image feature extraction to obtain a first feature, inputting the N real-time RGB images into a second image encoder for image feature extraction to obtain N second features, dynamically adjusting a first fusion weight of the first feature and a second fusion weight of the N second features based on the first real-time tomato image after illumination calibration, the N real-time RGB images, and a dynamic gated fusion mechanism, fusing the first and second features based on the first and second fusion weights, and introducing a spectrally sensitive and product fusion attention mechanism during the fusion process to achieve cross-modal feature fusion, thereby obtaining a fused multimodal feature;

[0008] The multimodal features are input into a maturity decoder for decoding to obtain a maturity result.

[0009] The beneficial effects of the present invention are as follows: when detecting the maturity of tomatoes, a first real-time tomato image captured by a digital camera and a second real-time tomato image captured by a hyperspectral imaging system are fused, that is, both the external features and the internal features of the tomato image are taken into account, thereby improving the accuracy of tomato maturity detection. Before fusion, the first real-time tomato image is first subjected to illumination calibration to improve the quality of the first real-time tomato image, and a spectral response function mapping method is used to extract N real-time RGB images corresponding to N preset indicators from the second tomato image, thereby improving the analysis efficiency of the second real-time tomato image captured by the spectral imaging system and ensuring the accuracy of the N real-time RGB images. The first feature obtained from the first real-time tomato image and the second feature obtained from N real-time RGB images are then fused. During the fusion process, not only the first fusion weight of the first feature and the second fusion weight of the second feature are dynamically adjusted through a dynamic gating mechanism, but also the spectral sensitivity matrix and attention mechanism are introduced to enhance the ability to capture complex features while allocating attention more stably, improving the accuracy and completeness of the fused multimodal features, thereby improving the accuracy of the obtained maturity results. This makes it possible to detect tomato maturity at any time, place, and environment without considering external lighting environment factors, thereby improving the detection efficiency of tomato maturity.

[0010] Optionally, the extracting N real-time RGB images corresponding to N preset indicators from the second real-time tomato image by using a spectral response function mapping method includes:

[0011] Obtain a second historical tomato image, randomly extract N bands from the second historical tomato image using a spectral response function mapping method, and convert the N bands into corresponding N historical band RGB images;

[0012] Inputting the N historical band RGB images into a third image encoder to extract corresponding band features to obtain N band features, and simultaneously inputting N preset indicators into the third image encoder to extract corresponding indicator features to obtain N indicator features;

[0013] Matching the N band features with the N indicator features to obtain a matching result, and training the spectral response function mapping method based on the matching result. If the matching result is successful, a trained spectral response function mapping method is obtained. If the matching result is unsuccessful, new N bands are extracted from the second historical tomato image and band features are extracted from the new N bands until the new N band features are successfully matched with the N indicator features, thereby obtaining a trained spectral response function mapping method.

[0014] The trained spectral response function mapping method is used to extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image.

[0015] As can be seen from the above description, the spectral response function mapping method is trained by using the second historical tomato image and the third image encoder, so that the trained spectral response function mapping method can extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image, eliminating the need for complex manual analysis, thereby improving the efficiency and accuracy of the analysis of the second real-time tomato image.

[0016] Optionally, inputting the first real-time tomato image after illumination calibration into a first image encoder for image feature extraction includes:

[0017] Obtaining a first historical tomato image after illumination calibration, inputting the first historical tomato image after illumination calibration into a first image encoder so that the first image encoder randomly masks the first historical tomato image after illumination calibration according to a first masking strategy to obtain the randomly masked first historical tomato image, and performing image restoration training on the first image encoder using the randomly masked first historical tomato image until an image that achieves a preset effect can be restored, thereby completing the training of the first image encoder and obtaining a trained first image encoder;

[0018] Inputting the first real-time tomato image after illumination calibration into the trained first image encoder for image feature extraction;

[0019] Inputting the N real-time RGB images into a second image encoder for image feature extraction includes:

[0020] Acquire a historical RGB image, input the historical RGB image into a second image encoder so that the second image encoder randomly masks the bands of the historical RGB image according to a second masking strategy to obtain a randomly masked historical RGB image, and perform image restoration training on the second image encoder using the randomly masked historical RGB image and an image restored by the trained first encoder to achieve a preset effect, until an image that achieves the preset effect can be restored to complete the training of the second image encoder, thereby obtaining a trained second image encoder;

[0021] The N real-time RGB images are input into the trained second image encoder to extract image features.

[0022] As can be seen from the above description, the first image encoder is trained for image restoration using the first historical tomato image after illumination calibration and the first masking strategy, so that the trained first image encoder is capable of restoring images that achieve the desired effect. This improves the accuracy of the first features obtained when the trained first image encoder is input with the first real-time tomato image after illumination calibration for image feature extraction. When training the second image encoder for image restoration, not only the historical RGB image and the second masking strategy are utilized, but also the image restored by the trained first image encoder to achieve the desired effect is combined. This enhances the second image encoder's understanding of cross-modal image correlations, improves the feature capture capability of the trained second image encoder, and further improves the accuracy of the second features obtained when the trained second image encoder is input with the real-time RGB image for image feature extraction.

[0023] Optionally, the training of the first image encoder is completed until an image that achieves a preset effect can be restored, and obtaining the trained first image encoder includes:

[0024] Calculating a first loss function value between the image restored by the first image encoder and the first historical tomato image after illumination calibration using a first composite loss function, and penalizing the first image encoder using the first loss function value to continuously perform image restoration training until the first loss function value reaches a preset convergence value, whereby the first image encoder is able to restore an image that achieves a preset effect, thereby completing the training of the first image encoder and obtaining a trained first image encoder;

[0025] The first composite loss function is:

[0026] ;

[0027] in, represents the first loss function value, represents the first constant coefficient, represents the second constant coefficient, n represents the total number of pixels in the first historical tomato image after illumination calibration, represents the first historical tomato image after illumination calibration pixel values, The first image encoder restores the image pixel values, represents the pixel mean of the first historical tomato image after illumination calibration, represents the pixel mean of the image restored by the first image encoder, represents the pixel covariance of the image restored by the first image encoder, represents the pixel covariance of the first historical tomato image after illumination calibration, represents the pixel variance of the image restored by the first image encoder, is the first constant, is the second constant;

[0028] The training of the second image encoder is completed until an image that achieves a preset effect can be restored, and the trained second image encoder is obtained, comprising:

[0029] Calculating a second loss function value of the image restored by the second image encoder and the historical RGB image using a second composite loss function, and penalizing the second image encoder by the second loss function value to continuously perform image restoration training until the second loss function value reaches a preset convergence value, so that the second image encoder can restore an image that achieves a preset effect, thereby completing the training of the second image encoder, and obtaining a trained second image encoder;

[0030] The second composite loss function is:

[0031] ;

[0032] in, represents the second loss function value, represents the third constant coefficient, represents the fourth constant coefficient, Represents the total number of pixels in the historical RGB image, Represents the first historical RGB image pixel values, The image restored by the second image encoder is represented by pixel values, represents the pixel mean of the historical RGB image, represents the pixel mean of the image restored by the second image encoder, represents the pixel covariance of the image restored by the second image encoder, represents the pixel covariance of the historical RGB image, represents the pixel variance of the image restored by the second image encoder, is the third constant, is the fourth constant.

[0033] According to the above description, the first composite loss function and the second composite loss function are respectively used as the judgment criteria for whether the first image encoder and the second image encoder can restore the image that achieves the preset effect, thereby ensuring the accuracy of the trained first image encoder and the trained second image encoder.

[0034] Optionally, inputting the first real-time tomato image after illumination calibration into a first image encoder for image feature extraction to obtain a first feature, and inputting the N real-time RGB images into a second image encoder for image feature extraction to obtain N second features includes:

[0035] The first image encoder and the second image encoder perform L×L image segmentation on the input first real-time tomato image after illumination calibration and the N real-time RGB images, respectively, to obtain corresponding S first image blocks and J second image blocks;

[0036] Straightening the S first image blocks and the J second image blocks respectively to obtain S straightened first image blocks and J straightened second image blocks, and performing one-dimensional vector conversion on the S straightened first image blocks and the J straightened second image blocks respectively to obtain S first one-dimensional vectors and J second one-dimensional vectors;

[0037] Performing a linear transformation on the S first one-dimensional vectors to obtain S first eigenvectors, performing a linear transformation on the J second one-dimensional vectors to obtain J second eigenvectors,

[0038] Using a cosine self-attention mechanism, respectively, calculate the similarities between the S first eigenvectors and the similarities between the J second eigenvectors to obtain corresponding first similarities and second similarities; using a multi-layer perception mechanism to adjust the dimensions of the S first eigenvectors based on the first similarities to obtain S dimensionally adjusted eigenvectors; and using a multi-layer perception mechanism to adjust the dimensions of the J second eigenvectors based on the second similarities to obtain J dimensionally adjusted eigenvectors;

[0039] The downsampling operation is used to extract features from the S first feature vectors and the J second feature vectors after dimension adjustment to obtain first features and N second features.

[0040] As can be seen from the above description, whether image feature extraction is performed on the first real-time tomato image after lighting standardization or on the real-time RGB image, the process involves segmenting the image to obtain image blocks, straightening the image blocks and converting them into one-dimensional vectors, linearly transforming the one-dimensional vectors obtained from the one-dimensional vector conversion, calculating the similarity between feature vectors using the cosine self-attention mechanism, adjusting the dimensions based on the similarity using the multi-layer perception mechanism, and finally performing feature extraction through downsampling to ensure the accuracy and completeness of the obtained first and second features.

[0041] Optionally, the first fusion weight of the first feature and the second fusion weights of the N second features are dynamically adjusted based on the first real-time tomato image after illumination calibration, the N real-time RGB images, and the dynamic gated fusion mechanism. The first feature and the second feature are fused based on the first fusion weight and the second fusion weight. During the fusion process, a spectrally sensitive and product fusion attention mechanism is introduced to achieve cross-modal feature fusion. The fused multimodal features include:

[0042] The first real-time tomato image after illumination calibration and the N real-time RGB images are input into the dynamic gating formula for calculation to obtain the dynamic gating value. The dynamic gating formula is:

[0043] ;

[0044] in, represents the dynamic gate value, represents the sigmoid function, W g represents the weight matrix, represents the i-th real-time RGB image, represents the first real-time tomato image after light calibration, represents the bias term;

[0045] The dynamic gating value is input into the first fusion weight formula for calculation to obtain the first fusion weight, and the dynamic gating value is used as the second fusion weight. The first fusion weight formula is:

[0046] ;

[0047] in, ' represents the first fusion weight;

[0048] The first fusion weight and the first real-time tomato image after illumination calibration are input into the first vector formula for calculation to obtain a first vector. Each real-time RGB image and the second fusion weight are sequentially input into the second vector formula for calculation to obtain N second vectors. The first vector and the N second vectors are fused. During the fusion process, a spectral sensitivity matrix and an attention mechanism are introduced to achieve cross-modal feature fusion to obtain a fused multimodal feature. The first vector formula is:

[0049]

[0050] Wherein, f1 represents the first vector;

[0051] The second vector formula is:

[0052]

[0053] in, represents the second vector corresponding to the i-th real-time RGB image, represents the i-th real-time RGB image.

[0054] As can be seen from the above description, the first fusion weight and the second fusion weight are calculated based on the input first real-time tomato image after illumination calibration and each real-time RGB image to adapt to the importance differences under different scene modalities and ensure the accuracy of the fused multimodal features.

[0055] Optionally, N is 2, and the attention mechanism includes a spectral sensitivity matrix, a sum attention mechanism, and a product attention mechanism. The first vector and the N second vectors are fused, and during the fusion process, a spectral sensitivity sum product fusion attention mechanism is introduced to achieve cross-modal feature fusion. The fused multimodal features include:

[0056] Generate a first matrix and a first projection matrix according to the first vector, generate corresponding second matrices and second projection matrices according to the two second vectors, input the first matrix, the first projection matrix, all second matrices, and all second projection matrices into the sum attention formula in the sum attention mechanism for calculation to obtain the first attention, and the sum attention formula is:

[0057] ;

[0058] Among them, AddA (Q, K, V) represents the first attention, tanh represents the hyperbolic tangent function, W q represents the first projection matrix, Q represents the first matrix, ⊙ represents the element-by-element multiplication of the matrix, W k represents the first second projection matrix, K represents the first second matrix, V represents the second second matrix, W v Represents the second projection matrix, V T represents the transposed matrix of the second matrix, and α represents a constant;

[0059] The spectral sensitivity matrix, the first matrix, and all the second matrices are input into the product attention formula in the product attention mechanism containing M attention heads to calculate the second attention. The product attention formula is:

[0060] ;

[0061] in, Indicates the second attention, Q m represents the semantic relevance of the features of the first matrix captured by the m-th attention head, V represents the feature semantic relevance of the transposed matrix of the first and second matrices captured by the m-th attention head, m represents the semantic relevance of the features of the second matrix captured by the mth attention head, W spec represents the spectral sensitivity matrix, SoftMax represents the normalized exponential function;

[0062] The first attention and the second attention are superimposed and fused to obtain a fused multimodal feature, and the multimodal feature is linearly transformed to obtain a final multimodal feature.

[0063] According to the above description, the matrix and projection matrix generated by the first vector and the second vector are input into the sum attention formula in the sum attention mechanism for calculation to obtain the first attention. All matrices are combined with the spectral sensitivity matrix and input into the product attention formula in the product attention mechanism containing M attention heads to obtain the second attention. The first attention and the second attention are superimposed, fused, and linearly transformed to obtain the final multimodal feature. The sum attention mechanism and the product attention mechanism perfectly capture and fuse the first vector and the second vector, that is, the first real-time tomato image and the real-time RGB image are perfectly fused. While utilizing the appearance features of the tomato image, its internal quality information is fully utilized to ensure the accuracy of the final multimodal feature.

[0064] Optionally, performing a linear transformation on the multimodal features to obtain final multimodal features includes:

[0065] The multimodal features are input into a feature enhancement formula to perform feature enhancement processing to obtain the multimodal features after feature enhancement processing. The multimodal features after feature enhancement processing are linearly transformed to obtain the final multimodal features. The feature enhancement formula is:

[0066] ;

[0067] Among them, y represents the multimodal features after feature enhancement processing, represents the scaling parameter, h represents the multimodal feature, u represents the mean of the multimodal feature, represents the variance of multimodal features, represents a very small constant, Represents the translation parameter.

[0068] According to the above description, before performing linear transformation on the multimodal features, the model features are first enhanced, that is, the multimodal features are regularized layer by layer to alleviate the drastic changes in the input distribution caused by the parameter update of the previous layer, making the gradient more stable and improving the accuracy of the final multimodal features.

[0069] Optionally, inputting the multimodal features into a maturity decoder for decoding to obtain a maturity result includes:

[0070] A maturity boundary level table and a target maturity level are obtained, and the multimodal features, the maturity boundary level table, and the target maturity level are input into a maturity decoder, so that the maturity decoder can hierarchically decode the multimodal features according to the maturity boundary level table to obtain a hierarchical decoding result, and calculate a target cosine similarity between the hierarchical decoding result and the target maturity level. A target maturity probability distribution is generated from the target cosine similarity using a Softmax function, and a maturity result is obtained based on the target maturity probability distribution.

[0071] According to the above description, when decoding multimodal features, the maturity results can be obtained according to the real-time maturity boundary level table and the target maturity level, ensuring the flexibility of decoding to be applicable to various scenarios and optimize the user experience.

[0072] In a second aspect, the present invention provides a system for detecting tomato maturity, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for detecting tomato maturity described in the first aspect is implemented.

[0073] The technical effects corresponding to the system for detecting tomato maturity provided in the second aspect refer to the relevant description of the method for detecting tomato maturity provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 This is a flow chart of a method for detecting tomato maturity provided in this embodiment;

[0075] Figure 2 This is a schematic diagram of the overall process of a method for detecting tomato maturity provided in this embodiment;

[0076] Figure 3 This is a schematic diagram of a process for extracting N real-time RGB images corresponding to N preset indicators from a second real-time tomato image using a spectral response function mapping method according to this embodiment;

[0077] Figure 4 A schematic diagram of a process for performing image feature extraction by the first image encoder and the second image encoder involved in this embodiment;

[0078] Figure 5 This is a schematic diagram of the structure of a tomato maturity detection system provided in this embodiment.

[0079] Description of Reference Numerals

[0080] 1. A system for detecting the maturity of tomatoes;

[0081] 2. Processor;

[0082] 3. Memory. DETAILED DESCRIPTION

[0083] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0084] Example 1

[0085] Please refer to Figures 1 to 2 The present invention provides a method for detecting tomato maturity, comprising the steps of:

[0086] S1. Obtain a first real-time tomato image captured by a digital camera and a second real-time tomato image captured by a hyperspectral imaging system, perform illumination calibration on the first real-time tomato image to obtain a first real-time tomato image after illumination calibration, and simultaneously extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image using a spectral response function mapping method;

[0087] In this embodiment, if Figure 2 As shown, a first real-time tomato image captured by a digital camera and a second real-time tomato image captured by a hyperspectral imaging system are obtained. A gamma function is used to perform illumination calibration on the first real-time tomato image to address the uneven illumination problem in the first real-time tomato image, thereby obtaining the first real-time tomato image after illumination calibration. Simultaneously, a spectral response function mapping method is used to extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image. In this case, N is 2, i.e., two preset indicators and two real-time RGB images corresponding to the preset indicators, such as preset indicator 1: sugar-acidity and preset indicator 2: lycopene content. The specific settings of N and the preset indicators can be adjusted according to actual conditions.

[0088] At this time, the step S1 of extracting N real-time RGB images corresponding to N preset indicators from the second real-time tomato image by using the spectral response function mapping method includes:

[0089] S101, obtaining a second historical tomato image, randomly extracting N bands from the second historical tomato image using a spectral response function mapping method, and converting the N bands into corresponding N historical band RGB images;

[0090] S102: Inputting the N historical band RGB images into a third image encoder to extract corresponding band features to obtain N band features, and simultaneously inputting N preset indicators into the third image encoder to extract corresponding indicator features to obtain N indicator features;

[0091] S103, matching the N band features with the N indicator features to obtain a matching result, and training the spectral response function mapping method based on the matching result. If the matching result is successful, a trained spectral response function mapping method is obtained. If the matching result is unsuccessful, new N bands are extracted from the second historical tomato image and band features are extracted from the new N bands until the new N band features are successfully matched with the N indicator features, thereby obtaining a trained spectral response function mapping method.

[0092] S104: Extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image using the trained spectral response function mapping method.

[0093] In this embodiment, if Figure 3 As shown, a second historical tomato image is obtained, and N bands are randomly extracted from the second historical tomato image using the spectral response function mapping method. At this time, N is 2, that is, 2 bands are randomly extracted, and the 2 bands are converted into corresponding 2 historical band RGB images. The 2 historical band RGB images are input into the third image encoder for corresponding band feature extraction to obtain 2 band features. At the same time, 2 preset indicators, that is, preset indicator 1: sugar acidity and preset indicator 2: lycopene content, are input into the third image encoder for indicator feature extraction to obtain 2 indicator features. The 2 band features are combined with the 2 indicators. The features are matched to obtain a matching result. If the matching result fails, the spectral response function mapping method re-extracts two new bands from the second historical tomato image and performs new band feature extraction on the new two bands until the new two band features are successfully matched with the two indicator features. The training of the spectral response function mapping method is completed, and the trained spectral response function mapping method is obtained. The trained spectral response function mapping method is used to extract two real-time RGB images corresponding to the two preset indicators from the second real-time tomato image. The specific steps for converting the extracted bands into RGB images are as follows:

[0094] 1. Let H( , λ) represents the position ( ) reflectivity at wavelength λ;

[0095] 2. Obtain the camera's red, green, and blue channel response functions R(λ), G(λ), and B(λ);

[0096] 3. For each wavelength position ( ), calculate the RGB three-channel intensity, and then perform discrete integration to calculate the RGB channel value, and finally use the calculated value as the position ( ), specifically the RGB channel value calculation formula is as follows, where Δλ is the spectral resolution, that is, the interval between adjacent wavelengths:

[0097] ;

[0098] ;

[0099] ;

[0100] 4. Finally, the discrete integration result is normalized to the range of [0, 1], and then gamma correction is applied to convert it into an RGB color image of [0, 255].

[0101] S2. Input the first real-time tomato image after illumination calibration into a first image encoder for image feature extraction to obtain a first feature. Input the N real-time RGB images into a second image encoder for image feature extraction to obtain N second features. Dynamically adjust a first fusion weight of the first feature and a second fusion weight of the N second features based on the first real-time tomato image after illumination calibration, the N real-time RGB images, and a dynamic gated fusion mechanism. Fuse the first and second features based on the first and second fusion weights. In the fusion process, introduce a spectral sensitivity matrix and an attention mechanism to achieve cross-modal feature fusion to obtain a fused multimodal feature.

[0102] In this embodiment, the spectrally sensitive and product fusion attention mechanism refers to: combining the spectrally sensitive matrix, the sum attention mechanism and the product attention mechanism, and the specific combination is described in steps S2121-S2122.

[0103] At this time, the step S2 of inputting the first real-time tomato image after illumination calibration into the first image encoder for image feature extraction includes:

[0104] S201: Obtain a first historical tomato image after illumination calibration, input the first historical tomato image after illumination calibration into a first image encoder, causing the first image encoder to randomly mask the first historical tomato image after illumination calibration according to a first masking strategy, thereby obtaining a randomly masked first historical tomato image. The first image encoder is then trained on the randomly masked first historical tomato image until an image that achieves a preset effect can be restored, thereby completing the training of the first image encoder and obtaining a trained first image encoder.

[0105] At this time, in step S201, the training of the first image encoder is completed until an image that achieves a preset effect can be restored, and the trained first image encoder is obtained, including:

[0106] S2011. Calculating a first loss function value between an image restored by the first image encoder and the first historical tomato image after illumination calibration using a first composite loss function, and penalizing the first image encoder using the first loss function value to continuously perform image restoration training until the first loss function value reaches a preset convergence value, whereupon the first image encoder is able to restore an image that achieves a preset effect, thereby completing the training of the first image encoder and obtaining a trained first image encoder.

[0107] The first composite loss function is:

[0108] ;

[0109] in, represents the first loss function value, represents the first constant coefficient, represents the second constant coefficient, n represents the total number of pixels in the first historical tomato image after illumination calibration, represents the first historical tomato image after illumination calibration pixel values, The first image encoder restores the image pixel values, represents the pixel mean of the first historical tomato image after illumination calibration, represents the pixel mean of the image restored by the first image encoder, represents the pixel covariance of the image restored by the first image encoder, represents the pixel covariance of the first historical tomato image after illumination calibration, represents the pixel variance of the image restored by the first image encoder, is the first constant, is the second constant;

[0110] S202: Input the first real-time tomato image after illumination calibration into the trained first image encoder to extract image features;

[0111] In this embodiment, if Figure 4As shown, the first real-time tomato image after illumination calibration is input into the trained first image encoder for image feature extraction. When training the first image encoder, the first historical tomato image after illumination calibration is input into the first image encoder so that the first image encoder can randomly mask the first historical tomato image after illumination calibration according to a first masking strategy. The first masking strategy is: using a 4×4 block-level masking strategy, first randomly masking the first historical tomato image after illumination calibration by 75% to obtain the first randomly masked historical tomato image. The first image encoder is then trained for image restoration using the first randomly masked historical tomato image until an image that achieves the preset effect can be restored to complete the training of the first image encoder. The training of the first image encoder is carried out, and whether the image that achieves the preset effect can be restored is judged by the first composite loss function. The first composite loss function is used to calculate the first loss function value of the image restored by the first image encoder and the first historical tomato image after illumination calibration. The first image encoder is punished by the first loss function value to continuously perform image restoration training until the first loss function value reaches the preset convergence value. It is considered that the first image encoder can restore the image that achieves the preset effect to complete the training of the first image encoder, and the trained first image encoder is obtained. The first real-time tomato image after illumination calibration is input into the trained first image encoder for image feature extraction to obtain the first feature. In the first composite loss function, the first constant coefficient and the second constant coefficient The value range of is (0,1), the first constant and the second constant To prevent the denominator from being 0, adaptive adjustment settings can be made.

[0112] Inputting the N real-time RGB images into a second image encoder for image feature extraction includes:

[0113] S203. Obtain a historical RGB image, input the historical RGB image into a second image encoder so that the second image encoder randomly masks the bands of the historical RGB image according to a second masking strategy to obtain a randomly masked historical RGB image, and perform image restoration training on the second image encoder using the randomly masked historical RGB image and an image restored by the trained first encoder to achieve a preset effect, until an image that achieves the preset effect can be restored to complete the training of the second image encoder, thereby obtaining a trained second image encoder.

[0114] At this time, the training of the second image encoder is completed in step S203 until an image that achieves a preset effect can be restored, and the trained second image encoder is obtained, including:

[0115] S2031. Calculating a second loss function value of the image restored by the second image encoder and the historical RGB image using a second composite loss function, and penalizing the second image encoder using the second loss function value to continuously perform image restoration training until the second loss function value reaches a preset convergence value, so that the second image encoder can restore an image that achieves a preset effect, thereby completing the training of the second image encoder, and obtaining a trained second image encoder.

[0116] The second composite loss function is:

[0117] ;

[0118] in, represents the second loss function value, represents the third constant coefficient, represents the fourth constant coefficient, Represents the total number of pixels in the historical RGB image, Represents the first historical RGB image pixel values, The image restored by the second image encoder is represented by pixel values, represents the pixel mean of the historical RGB image, represents the pixel mean of the image restored by the second image encoder, represents the pixel covariance of the image restored by the second image encoder, represents the pixel covariance of the historical RGB image, represents the pixel variance of the image restored by the second image encoder, is the third constant, is the fourth constant.

[0119] S204: Input the N real-time RGB images into the trained second image encoder to extract image features.

[0120] In this embodiment, if Figure 4As shown, the second image encoder after training inputs the real-time RGB image to perform image feature extraction. When training the second image encoder, the historical RGB image is input into the second image encoder so that the second image encoder can randomly mask the bands of the historical RGB image according to the second masking strategy. The second masking strategy is to perform a 50% random mask on the historical RGB image, forcing the second image encoder to be able to perform image restoration training based on the tomato internal quality information requirements input by the user, forcibly combining the image restored by the trained first encoder to achieve the preset effect, until the image that achieves the preset effect can be restored. As for whether it can When an image that achieves a preset effect is restored, it is judged by a second composite loss function. The second composite loss function is used to calculate the second loss function value of the image restored by the second image encoder and the historical RGB image. The second image encoder is punished by the second loss function value to continuously perform image restoration training until the second loss function value reaches a preset convergence value. It is then considered that the second image encoder can restore the image that achieves the preset effect to complete the training of the second image encoder, and the trained second image encoder is obtained. The real-time RGB image is input into the trained second image encoder for image feature extraction to obtain the second feature. In the second composite loss function, the third constant coefficient and the fourth constant coefficient The value range of is (0,1), the third constant and the fourth constant To prevent the denominator from being 0, adaptive adjustment settings can be made.

[0121] At this time, in step S2, the first real-time tomato image after illumination calibration is input into the first image encoder for image feature extraction to obtain the first feature, and the N real-time RGB images are input into the second image encoder for image feature extraction to obtain N second features including:

[0122] S205: The first image encoder and the second image encoder perform L×L image segmentation on the input first real-time tomato image after illumination calibration and the N real-time RGB images, respectively, to obtain corresponding S first image blocks and J second image blocks;

[0123] S206: Straighten the S first image blocks and the J second image blocks to obtain S straightened first image blocks and J straightened second image blocks, and perform one-dimensional vector conversion on the S straightened first image blocks and the J straightened second image blocks to obtain S first one-dimensional vectors and J second one-dimensional vectors.

[0124] S207, performing a linear transformation on the S first one-dimensional vectors to obtain S first eigenvectors, performing a linear transformation on the J second one-dimensional vectors to obtain J second eigenvectors,

[0125] S208. Using a cosine self-attention mechanism, respectively, calculate the similarities between the S first eigenvectors and the similarities between the J second eigenvectors to obtain corresponding first similarities and second similarities. Based on the first similarities, use a multi-layer perception mechanism to perform dimension adjustment on the S first eigenvectors to obtain S dimension-adjusted eigenvectors. Based on the second similarities, use a multi-layer perception mechanism to perform dimension adjustment on the J second eigenvectors to obtain J dimension-adjusted eigenvectors.

[0126] S209 , performing feature extraction on the S first feature vectors and the J second feature vectors after dimension adjustment through a downsampling operation to obtain first features and N second features.

[0127] In this embodiment, if Figure 2 As shown, the first image encoder and the second image encoder perform L×L image segmentation on the input first real-time tomato image after illumination calibration and N real-time RGB images, respectively, to obtain corresponding S first image blocks and J second image blocks, where L is 4, that is, 4×4 image segmentation is performed, and the S first image blocks and the J second image blocks are straightened to obtain S straightened first image blocks and J straightened second image blocks, and the S straightened first image blocks and the J straightened second image blocks are converted into one-dimensional vectors to obtain S first one-dimensional vectors and J second one-dimensional vectors, and the S first one-dimensional vectors and the J second one-dimensional vectors are linearly transformed to obtain S first eigenvectors and J second eigenvectors, where the linear transformation formula is:

[0128] Eigenvector = one-dimensional vector × preset first constant matrix + preset second constant matrix;

[0129] The cosine self-attention mechanism is used to calculate the similarities between the S first eigenvectors and the similarities between the J second eigenvectors, respectively, to obtain the corresponding high first similarities and second similarities. The multi-layer perception mechanism is used to adjust the dimensions of the S first eigenvectors according to the first similarity to obtain S eigenvectors after dimension adjustment. The multi-layer perception mechanism is used to adjust the dimensions of the J second eigenvectors according to the second similarity to obtain J eigenvectors after dimension adjustment. Then, feature extraction is performed on the S first eigenvectors after dimension adjustment and the J second eigenvectors after dimension adjustment through downsampling operations to obtain the initial first feature and the initial N second features. Steps S208-S209 are repeated until the loop number threshold is reached to obtain the final first feature and the final N second features, where the loop number threshold is 4 and can be adjusted according to actual conditions.

[0130] At this time, in step S2, the first fusion weight of the first feature and the second fusion weights of the N second features are dynamically adjusted based on the first real-time tomato image after illumination calibration, the N real-time RGB images, and the dynamic gated fusion mechanism. The first feature and the second feature are fused based on the first fusion weight and the second fusion weight. During the fusion process, a spectrally sensitive and product fusion attention mechanism is introduced to achieve cross-modal feature fusion. The fused multimodal features include:

[0131] S210: Input the first real-time tomato image after illumination calibration and the N real-time RGB images into a dynamic gating formula for calculation to obtain a dynamic gating value. The dynamic gating formula is:

[0132] ;

[0133] in, represents the dynamic gate value, represents the sigmoid function, W g represents the weight matrix, represents the i-th real-time RGB image, represents the first real-time tomato image after light calibration, represents the bias term;

[0134] S211: Input the dynamic gating value into a first fusion weight formula for calculation to obtain a first fusion weight, and use the dynamic gating value as a second fusion weight. The first fusion weight formula is:

[0135] ;

[0136] in, ' represents the first fusion weight;

[0137] S212: Input the first fusion weight and the first real-time tomato image after illumination calibration into a first vector formula for calculation to obtain a first vector. Input each real-time RGB image and the second fusion weight into a second vector formula for calculation in sequence to obtain N second vectors. Fuse the first vector with the N second vectors. During the fusion process, introduce a spectral sensitivity matrix and an attention mechanism to achieve cross-modal feature fusion to obtain a fused multimodal feature. The first vector formula is:

[0138]

[0139] Wherein, f1 represents the first vector;

[0140] The second vector formula is:

[0141]

[0142] in, represents the second vector corresponding to the i-th real-time RGB image, represents the i-th real-time RGB image.

[0143] In this embodiment, if Figure 2 As shown, the first real-time tomato image after illumination calibration and N real-time RGB images are input into the dynamic gating formula for calculation to obtain a dynamic gating value. The dynamic gating value is then input into the first fusion weight formula for calculation to obtain a first fusion weight. The dynamic gating value is also used as the second fusion weight. That is, the dynamic gating formula can adaptively adjust the importance differences of modalities in different scenarios based on the input first real-time tomato image after illumination calibration and N real-time RGB images. In this embodiment, it can be seen that the dynamic gating value is the second fusion weight. Therefore, if the N real-time RGB images are more reliable than the first real-time tomato image after illumination calibration, the calculated will be better than For example, if the N real-time RGB images obtained in a strong light environment are more reliable than the first real-time tomato image after illumination calibration, the second fusion weight will be set higher than the first fusion weight. The first fusion weight and the first real-time tomato image after illumination calibration are input into the first vector formula for calculation to obtain the first vector. Each real-time RGB image and the second fusion weight are input into the second vector formula in sequence for calculation to obtain N second vectors. All first vectors are fused with the N second vectors. During the fusion process, the spectral sensitivity matrix and attention mechanism are introduced to achieve cross-modal feature fusion to obtain the fused multimodal features.

[0144] At this time, the N in step S212 is 2, and the spectrally sensitive and product fusion attention mechanism includes a spectrally sensitive matrix, a sum attention mechanism, and a product attention mechanism. The first vector and the N second vectors are fused, and in the fusion process, the spectrally sensitive and product fusion attention mechanism is introduced to achieve cross-modal feature fusion. The fused multimodal features include:

[0145] S2121. Generate a first matrix and a first projection matrix based on the first vector, generate corresponding second matrices and second projection matrices based on the two second vectors, input the first matrix, the first projection matrix, all second matrices, and all second projection matrices into the sum attention formula in the sum attention mechanism for calculation to obtain a first attention. The sum attention formula is:

[0146] ;

[0147] Among them, AddA (Q, K, V) represents the first attention, tanh represents the hyperbolic tangent function, W q represents the first projection matrix, Q represents the first matrix, ⊙ represents the element-by-element multiplication of the matrix, W k represents the first second projection matrix, K represents the first second matrix, V represents the second second matrix, W v Represents the second projection matrix, V T represents the transposed matrix of the second matrix, and α represents a constant;

[0148] S2122: Input the spectral sensitivity matrix, the first matrix, and all the second matrices into the product attention formula in the product attention mechanism including M attention heads to calculate and obtain the second attention. The product attention formula is:

[0149] ;

[0150] in, Indicates the second attention, Q m represents the semantic relevance of the features of the first matrix captured by the m-th attention head, V represents the feature semantic relevance of the transposed matrix of the first and second matrices captured by the m-th attention head, m represents the semantic relevance of the features of the second matrix captured by the mth attention head, W spec represents the spectral sensitivity matrix, SoftMax represents the normalized exponential function;

[0151] S2123. Superimpose and fuse the first attention and the second attention to obtain a fused multimodal feature, and perform a linear transformation on the multimodal feature to obtain a final multimodal feature.

[0152] In this embodiment, when performing linear transformation on multimodal features, feature enhancement processing is first performed. The specific steps are as follows:

[0153] The multimodal features are input into a feature enhancement formula to perform feature enhancement processing to obtain the multimodal features after feature enhancement processing. The multimodal features after feature enhancement processing are linearly transformed to obtain the final multimodal features. The feature enhancement formula is:

[0154] ;

[0155] Among them, y represents the multimodal features after feature enhancement processing, represents the scaling parameter, h represents the multimodal feature, u represents the mean of the multimodal feature, represents the variance of multimodal features, represents a very small constant, Represents the translation parameter.

[0156] In this embodiment, if Figure 2 As shown, the first matrix Q and the first projection matrix W are generated according to the first vector q , according to the two second vectors, generate the corresponding first second matrix K, second second matrix V, and first second projection matrix W k And the second projection matrix W v , all matrices and projection matrices are input into the sum attention formula in the sum attention mechanism for calculation to obtain the first attention, and the spectral sensitivity matrix, the first matrix and all the second matrices are input into the product attention formula in the product attention mechanism containing M attention heads for calculation to obtain the second attention, where M is 8, that is, the product attention mechanism containing 8 attention heads. The attention head is used to capture the semantic relevance of the matrix features. The obtained first attention and second attention are superimposed and fused to obtain the fused multimodal features, and the multimodal features are enhanced. Feature enhancement refers to layer regularization of the multimodal features to alleviate the drastic changes in the input distribution caused by the parameter update of the previous layer, making the gradient more stable, thereby obtaining the multimodal features after feature enhancement processing, and performing linear transformation on the multimodal features after feature enhancement processing to obtain the final multimodal features.

[0157] S3. Input the multimodal features into a maturity decoder for decoding to obtain a maturity result.

[0158] At this time, step S3 includes:

[0159] S301. Obtain a maturity boundary level table and a target maturity level, input the multimodal features, the maturity boundary level table, and the target maturity level into a maturity decoder, so that the maturity decoder can hierarchically decode the multimodal features according to the maturity boundary level table to obtain a hierarchical decoding result, calculate a target cosine similarity between the hierarchical decoding result and the target maturity level, generate a target maturity probability distribution for the target cosine similarity using a Softmax function, and obtain a maturity result based on the target maturity probability distribution.

[0160] In this embodiment, if Figure 2 As shown, the maturity boundary level table and the target maturity level are obtained as the maturity level specified by the user, and the multimodal features, the maturity boundary level table and the target maturity level are input into the maturity decoder for decoding, so that the maturity decoder can perform hierarchical decoding of the multimodal features according to the maturity boundary level table. When performing hierarchical decoding according to the maturity boundary level table, the maturity decoder will fuse the details of the multimodal features by calculating cosine self-attention, and then adjust the feature dimension vector corresponding to the multimodal features through the multi-layer perception mechanism. Finally, the hierarchical decoding results are visualized through upsampling, and the target cosine similarity between the hierarchical decoding results and the target maturity level is calculated. The target maturity probability distribution is generated by the target cosine similarity through the Softmax function, and finally the maturity result is obtained by drawing a bounding box. The maturity boundary level table and the target maturity level can support real-time adjustment and modification by the user.

[0161] Example 2

[0162] Please refer to Figure 5 The present invention provides a system 1 for detecting tomato maturity, comprising a memory 3, a processor 2, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, the steps in the first embodiment are implemented.

[0163] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structures and variations of these systems / devices based on the methods described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.

[0164] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0165] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process flow and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0166] It should be noted that, in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims enumerating several means, several of these means may be embodied by one and the same hardware. The use of the words first, second, third etc. is for convenience only and does not indicate any order. These words may be understood as part of the component name.

[0167] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0168] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0169] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.

Claims

1. A method for detecting tomato maturity, characterized in that: include: Acquire a first real-time tomato image captured by a digital camera and a second real-time tomato image captured by a hyperspectral imaging system, perform illumination calibration on the first real-time tomato image to obtain a first real-time tomato image after illumination calibration, and simultaneously extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image using a spectral response function mapping method; Inputting the first real-time tomato image after illumination calibration into a first image encoder for image feature extraction to obtain a first feature, and inputting the N real-time RGB images into a second image encoder for image feature extraction to obtain N second features; The first real-time tomato image after illumination calibration and N real-time RGB images are input into the dynamic gating formula for calculation to obtain the dynamic gating value And as the second fusion weight, then (1- ) is the first fusion weight, and the dynamic gating formula is: ; in, represents the dynamic gate value, represents the sigmoid function, W g represents the weight matrix, represents the i-th real-time RGB image, represents the first real-time tomato image after light calibration, represents the bias term; The first fusion weight and the first feature are input into the first vector formula for calculation to obtain a first vector, each second feature and the second fusion weight are sequentially input into the second vector formula for calculation to obtain N second vectors, the first vector and the N second vectors are fused, and in the fusion process, a spectrally sensitive sum product fusion attention mechanism is introduced to achieve cross-modal feature fusion. The spectrally sensitive sum product fusion attention mechanism includes a spectrally sensitive matrix, a sum attention mechanism, and a product attention mechanism to obtain fused multimodal features, specifically including: The first vector formula is: Wherein, f1 represents the first vector; The second vector formula is: in, Represents the second vector corresponding to the i-th real-time RGB image; The number N is 2. A first matrix and a first projection matrix are generated according to the first vector. The corresponding second matrices and second projection matrices are generated according to the two second vectors. The first matrix, the first projection matrix, all the second matrices, and all the second projection matrices are input into the sum attention formula in the sum attention mechanism for calculation to obtain the first attention. The sum attention formula is: ; Among them, AddAttn(Q,K,V) represents the first attention, tanh represents the hyperbolic tangent function, W q represents the first projection matrix, Q represents the first matrix, ⊙ represents the element-by-element multiplication of the matrix, W k represents the first second projection matrix, K represents the first second matrix, V represents the second second matrix, W v Represents the second projection matrix, V T represents the transposed matrix of the second matrix, and α represents a constant; The spectral sensitivity matrix, the first matrix, and all the second matrices are input into the product attention formula in the product attention mechanism containing M attention heads to calculate the second attention. The product attention formula is: ; in, Indicates the second attention, Q m represents the semantic relevance of the features of the first matrix captured by the m-th attention head, V represents the feature semantic relevance of the transposed matrix of the first and second matrices captured by the m-th attention head, m represents the semantic relevance of the features of the second matrix captured by the mth attention head, W spec represents the spectral sensitivity matrix, SoftMax represents the normalized exponential function; The first attention and the second attention are superimposed and fused to obtain a fused multimodal feature, and the multimodal feature is linearly transformed to obtain a final multimodal feature; The multimodal features are input into a maturity decoder for decoding to obtain a maturity result.

2. The method for detecting tomato maturity according to claim 1, wherein: The extracting N real-time RGB images corresponding to N preset indicators from the second real-time tomato image by using a spectral response function mapping method comprises: Obtain a second historical tomato image, randomly extract N bands from the second historical tomato image using a spectral response function mapping method, and convert the N bands into corresponding N historical band RGB images; Inputting the N historical band RGB images into a third image encoder to extract corresponding band features to obtain N band features, and simultaneously inputting N preset indicators into the third image encoder to extract corresponding indicator features to obtain N indicator features; Matching the N band features with the N indicator features to obtain a matching result, and training the spectral response function mapping method based on the matching result. If the matching result is successful, a trained spectral response function mapping method is obtained. If the matching result is unsuccessful, new N bands are extracted from the second historical tomato image and band features are extracted from the new N bands until the new N band features are successfully matched with the N indicator features, thereby obtaining a trained spectral response function mapping method. The trained spectral response function mapping method is used to extract N real-time RGB images corresponding to N preset indicators from the second real-time tomato image.

3. The method for detecting tomato maturity according to claim 1, wherein: Inputting the first real-time tomato image after illumination calibration into the first image encoder to extract image features includes: Obtaining a first historical tomato image after illumination calibration, inputting the first historical tomato image after illumination calibration into a first image encoder so that the first image encoder randomly masks the first historical tomato image after illumination calibration according to a first masking strategy to obtain the randomly masked first historical tomato image, and performing image restoration training on the first image encoder using the randomly masked first historical tomato image until an image that achieves a preset effect can be restored, thereby completing the training of the first image encoder and obtaining a trained first image encoder; Inputting the first real-time tomato image after illumination calibration into the trained first image encoder for image feature extraction; Inputting the N real-time RGB images into a second image encoder for image feature extraction includes: Acquire a historical RGB image, input the historical RGB image into a second image encoder so that the second image encoder randomly masks the bands of the historical RGB image according to a second masking strategy to obtain a randomly masked historical RGB image, and perform image restoration training on the second image encoder using the randomly masked historical RGB image and an image restored by the trained first encoder to achieve a preset effect, until an image that achieves the preset effect can be restored to complete the training of the second image encoder, thereby obtaining a trained second image encoder; The N real-time RGB images are input into the trained second image encoder to extract image features.

4. A method for detecting tomato maturity as claimed in claim 3, characterized in that: The training of the first image encoder is completed until an image that achieves a preset effect can be restored, and the trained first image encoder is obtained, comprising: Calculating a first loss function value between the image restored by the first image encoder and the first historical tomato image after illumination calibration using a first composite loss function, and penalizing the first image encoder using the first loss function value to continuously perform image restoration training until the first loss function value reaches a preset convergence value, whereby the first image encoder is able to restore an image that achieves a preset effect, thereby completing the training of the first image encoder and obtaining a trained first image encoder; The first composite loss function is: ; in, represents the first loss function value, represents the first constant coefficient, represents the second constant coefficient, n represents the total number of pixels in the first historical tomato image after illumination calibration, represents the first historical tomato image after illumination calibration pixel values, The first image encoder restores the image pixel values, represents the pixel mean of the first historical tomato image after illumination calibration, represents the pixel mean of the image restored by the first image encoder, represents the pixel covariance of the image restored by the first image encoder, represents the pixel covariance of the first historical tomato image after illumination calibration, represents the pixel variance of the image restored by the first image encoder, is the first constant, is the second constant; The training of the second image encoder is completed until an image that achieves a preset effect can be restored, and the trained second image encoder is obtained, comprising: Calculating a second loss function value of the image restored by the second image encoder and the historical RGB image using a second composite loss function, and penalizing the second image encoder by the second loss function value to continuously perform image restoration training until the second loss function value reaches a preset convergence value, so that the second image encoder can restore an image that achieves a preset effect, thereby completing the training of the second image encoder, and obtaining a trained second image encoder; The second composite loss function is: ; in, represents the second loss function value, represents the third constant coefficient, represents the fourth constant coefficient, Represents the total number of pixels in the historical RGB image, Represents the first historical RGB image pixel values, The image restored by the second image encoder is represented by pixel values, represents the pixel mean of the historical RGB image, represents the pixel mean of the image restored by the second image encoder, represents the pixel covariance of the image restored by the second image encoder, represents the pixel covariance of the historical RGB image, represents the pixel variance of the image restored by the second image encoder, is the third constant, is the fourth constant.

5. The method for detecting tomato maturity according to claim 1, wherein: Inputting the first real-time tomato image after illumination calibration into the first image encoder for image feature extraction to obtain the first feature, and inputting the N real-time RGB images into the second image encoder for image feature extraction to obtain N second features include: The first image encoder and the second image encoder perform L×L image segmentation on the input first real-time tomato image after illumination calibration and the N real-time RGB images, respectively, to obtain corresponding S first image blocks and J second image blocks; Straightening the S first image blocks and the J second image blocks respectively to obtain S straightened first image blocks and J straightened second image blocks, and performing one-dimensional vector conversion on the S straightened first image blocks and the J straightened second image blocks respectively to obtain S first one-dimensional vectors and J second one-dimensional vectors; Performing a linear transformation on the S first one-dimensional vectors to obtain S first eigenvectors, performing a linear transformation on the J second one-dimensional vectors to obtain J second eigenvectors, Using a cosine self-attention mechanism, respectively, calculate the similarities between the S first eigenvectors and the similarities between the J second eigenvectors to obtain corresponding first similarities and second similarities; using a multi-layer perception mechanism to adjust the dimensions of the S first eigenvectors based on the first similarities to obtain S dimensionally adjusted eigenvectors; and using a multi-layer perception mechanism to adjust the dimensions of the J second eigenvectors based on the second similarities to obtain J dimensionally adjusted eigenvectors; The downsampling operation is used to extract features from the S first feature vectors and the J second feature vectors after dimension adjustment to obtain first features and N second features.

6. The method for detecting tomato maturity according to claim 1, wherein: The linear transformation of the multimodal features to obtain the final multimodal features includes: The multimodal features are input into a feature enhancement formula to perform feature enhancement processing to obtain the multimodal features after feature enhancement processing. The multimodal features after feature enhancement processing are linearly transformed to obtain the final multimodal features. The feature enhancement formula is: ; Among them, y represents the multimodal features after feature enhancement processing, represents the scaling parameter, h represents the multimodal feature, u represents the mean of the multimodal feature, represents the variance of multimodal features, represents a very small constant, Represents the translation parameter.

7. The method for detecting tomato maturity according to claim 1, wherein: Inputting the multimodal features into a maturity decoder for decoding to obtain a maturity result includes: A maturity boundary level table and a target maturity level are obtained, and the multimodal features, the maturity boundary level table, and the target maturity level are input into a maturity decoder, so that the maturity decoder can hierarchically decode the multimodal features according to the maturity boundary level table to obtain a hierarchical decoding result, and calculate a target cosine similarity between the hierarchical decoding result and the target maturity level. A target maturity probability distribution is generated from the target cosine similarity using a Softmax function, and a maturity result is obtained based on the target maturity probability distribution.

8. A system for detecting tomato maturity, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Transform-based cross-modal fusion hyperspectral image super-resolution reconstruction method

    CN118628357A

  • Deep learning-based tomato maturity detection system and method

    CN119540944A