Fruit and vegetable identification and identification model and optical element combined training method and device

Through the joint training method of recognition model and optical components, the problem of poor image quality in traditional fruit and vegetable recognition technology in low temperature environments is solved, and high-accurate fruit and vegetable category and freshness recognition is achieved.

CN119942529APending Publication Date: 2025-05-06SHPHOTONICS LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202412000336.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional fruit and vegetable recognition technology is affected by factors such as light changes, reflections and occlusion in low-temperature environments, resulting in poor image quality, which in turn reduces the accuracy of recognition results, and it is difficult to achieve category recognition and freshness evaluation of fruits and vegetables at the same time.

Method used

The joint training method of recognition model and optical elements is adopted to modulate the optical signal through optical elements to generate high-quality target images, and the category and freshness recognition results of fruits and vegetables are generated using deep learning technology.

Benefits of technology

It improves image quality and accuracy of recognition results, realizes category recognition and freshness evaluation of fruits and vegetables, and improves the comprehensiveness of recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942529A_ABST
    Figure CN119942529A_ABST
Patent Text Reader

Abstract

The invention provides a fruit and vegetable identification and identification model and optical element combined training method and device, and relates to the artificial intelligence fields of deep learning, computer vision, image processing, micro-nano optical design and the like. The fruit and vegetable recognition method comprises the steps that a to-be-processed target image is obtained, the target image is a to-be-recognized fruit and vegetable image generated by a photosensitive element according to an obtained second optical signal, and the second optical signal is an optical signal obtained after an optical element modulates an obtained first optical signal, the first optical signal is an optical signal that light emitted by the light source irradiates to-be-identified fruits and vegetables and then enters the optical element through reflection; and generating a recognition result corresponding to the target image by using the recognition model, wherein the recognition result comprises a category recognition result and a freshness recognition result of the to-be-recognized fruits and vegetables.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, computer vision, image processing and micro-nano optical design, and specifically to a method and device for joint training of fruit and vegetable recognition and recognition models and optical elements. Background Art

[0002] Traditional fruit and vegetable recognition mainly involves identifying the categories of fruits and vegetables, such as collecting images of fruits and vegetables in refrigeration equipment such as refrigerators and freezers, and determining the categories of fruits and vegetables in the images by analyzing the collected images. Summary of the invention

[0003] The present disclosure provides a method and device for fruit and vegetable recognition and joint training of a recognition model and an optical element.

[0004] A method for identifying fruits and vegetables, comprising:

[0005] Acquire a target image to be processed, wherein the target image is an image of fruits and vegetables to be identified generated by a photosensitive element according to an acquired second light signal, wherein the second light signal is an optical signal obtained by modulating the acquired first light signal by an optical element, and the first light signal is an optical signal that enters the optical element after the light emitted by a light source irradiates the fruits and vegetables to be identified and is reflected;

[0006] A recognition model is used to generate a recognition result corresponding to the target image, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

[0007] A joint training method of a recognition model and an optical element, comprising:

[0008] Acquire a sample image, and construct a training sample according to the sample image;

[0009] The recognition model and the optical element are jointly trained according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to a target image to be processed, the target image is an image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, the first light signal is a light signal that is emitted by a light source and then reflected by the light source and enters the optical element to be identified, and the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be identified.

[0010] A fruit and vegetable identification device, comprising: an image acquisition module and an image recognition module;

[0011] The image acquisition module is used to acquire a target image to be processed, wherein the target image is an image of fruits and vegetables to be identified generated by a photosensitive element according to an acquired second light signal, wherein the second light signal is an optical signal obtained by modulating an acquired first light signal by an optical element, and the first light signal is an optical signal that enters the optical element after light emitted by a light source irradiates the fruits and vegetables to be identified and is reflected;

[0012] The image recognition module is used to generate a recognition result corresponding to the target image using a recognition model, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

[0013] A joint training device for a recognition model and an optical element, comprising: a sample construction module and a model training module;

[0014] The sample construction module is used to obtain a sample image and construct a training sample according to the sample image;

[0015] The model training module is used to jointly train the recognition model and the optical element according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to a target image to be processed, the target image is an image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, the first light signal is a light signal emitted by a light source and then reflected by the light source and entering the optical element to be identified, and the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be identified.

[0016] An electronic device, comprising:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0020] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0021] A computer program product comprises a computer program / instruction, wherein the computer program / instruction implements the method described above when executed by a processor.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0024] Figure 1 This is a flow chart of an embodiment of the fruit and vegetable identification method disclosed in the present invention;

[0025] Figure 2 A schematic diagram of a method for acquiring a target image according to the present disclosure;

[0026] Figure 3 A flowchart of an embodiment of a joint training method of a recognition model and an optical element disclosed in the present invention;

[0027] Figure 4 It is a schematic diagram of the overall implementation process of the joint training of the recognition model and the optical element and the fruit and vegetable recognition method disclosed in the present invention;

[0028] Figure 5 It is a schematic diagram of the composition structure of the fruit and vegetable identification device embodiment 500 of the present disclosure;

[0029] Figure 6 It is a schematic diagram of the composition structure of an embodiment 600 of a joint training device for a recognition model and an optical element according to the present disclosure;

[0030] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0031] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] In addition, it should be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0033] Figure 1FIG. 1 is a flow chart of an embodiment of the fruit and vegetable identification method disclosed in the present invention. Figure 1 As shown, the following specific implementation methods are included.

[0034] In step 101, a target image to be processed is obtained. The target image is an image of fruits and vegetables to be identified generated by a photosensitive element according to an acquired second light signal. The second light signal is a light signal obtained by modulating an acquired first light signal by an optical element. The first light signal is a light signal emitted by a light source and then irradiated onto the fruits and vegetables to be identified and then reflected into the optical element.

[0035] In step 102, a recognition model is used to generate a recognition result corresponding to the target image, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

[0036] In traditional fruit and vegetable identification methods, the quality of collected images is usually poor due to factors such as light changes, reflection and occlusion in low-temperature environments, which in turn reduces the accuracy of the identification results.

[0037] By adopting the scheme described in the above method embodiment, the optical element can be used to modulate the light signal, and then the target image can be generated based on the modulated light signal, thereby improving the image quality and correspondingly improving the accuracy of the subsequent recognition results. Moreover, the recognition model can be used to generate the recognition result, thereby utilizing the powerful reasoning ability of the recognition model, thereby further improving the accuracy of the recognition result. In addition, not only the category of fruits and vegetables can be identified, but also the freshness of fruits and vegetables can be identified, that is, fruit and vegetable classification and freshness evaluation can be achieved simultaneously, thereby improving the comprehensiveness of the recognition results.

[0038] Figure 2 Schematic diagram of the method for acquiring the target image described in the present disclosure. Figure 2 As shown, after the light emitted by the light source irradiates the fruits and vegetables to be identified, it is reflected and enters the optical element. The optical element modulates the incoming light signal and outputs it to the photosensitive element, which then generates a target image through photoelectric conversion and other processing. The fruits and vegetables to be identified can refer to one or more fruits and vegetables of a single category, such as one or more kiwis, or multiple fruits and vegetables of multiple categories. Figure 2 As shown, both the optical element and the photosensitive element can be located in the imaging structure, that is, the optical element and the photosensitive element can be arranged in one piece, or they can be arranged separately, depending on actual needs. Figure 2 As shown, the recognition model, light source and imaging structure can together form a fruit and vegetable recognition system.

[0039] In some embodiments of the present disclosure, the light source may include: a wide-band light source, and the target image may include: a target image including both red, green, and blue (RGB) image information and near infrared (NIR) image information.

[0040] The wide wavelength band may refer to a wavelength band from 450 nanometers (nm) to 1400 nm.

[0041] Accordingly, the scheme disclosed in the present invention can combine a target image including both RGB image information and NIR image information and deep learning technology to identify fruits and vegetables to be identified, thereby improving the accuracy of the identification results. For example, fruits and vegetables with similar appearance but different categories, such as Golden Delicious apples and Crown Pears, can be more accurately distinguished. In addition, fruits and vegetables that are difficult for the human eye to distinguish, such as internal rot and other unfresh conditions, can also be accurately identified.

[0042] For the recognition model, multimodal input can be constructed through specific channel encoding methods, such as adding NIR image information as an additional channel to RGB image information. The recognition model may include feature extraction modules, multimodal fusion modules, and classification heads. The feature extraction module can be used to extract multimodal features from the target image. The multimodal fusion module can splice the extracted features in the channel dimension and filter out key features through 1*1 convolution or attention mechanism. Feature interaction can be performed in the middle and high layers of the network to improve the model's ability to express the collaborative characteristics of RGB and NIR. The fully connected layer can be used in the classification head to output the recognition results, such as using the normalized exponential (Softmax) function to determine the category recognition results and freshness recognition results.

[0043] The recognition model may be a convolutional neural network (CNN) model, such as a residual network (ResNet) model or a mobile convolutional network (MobileNet) model.

[0044] In some embodiments of the present disclosure, the optical element may include: a metasurface element. As an emerging micro-nano optical design method, metasurface technology can accurately control the light propagation characteristics of the RGB band and the NIR band, maximize the light transmittance and optimize the response to light of different bands, etc., thereby improving image quality, such as enhancing key information in the image, including the color, shape and texture of fruits and vegetables, especially having significant advantages under low temperature and complex lighting conditions.

[0045] It should be noted that, in addition to being a metasurface element, the optical element in the solution described in the present disclosure may also be a refractive optical element, a diffractive optical element or a scattering medium element.

[0046] The recognition model and the optical element may be obtained in advance through joint training. The above mainly describes the model reasoning process. The following will describe the joint training process of the recognition model and the optical element in conjunction with the embodiments.

[0047] Figure 3 Flow chart of an embodiment of the joint training method of the recognition model and the optical element disclosed in the present invention. Figure 3 As shown, the following specific implementation methods are included.

[0048] In step 301, a sample image is obtained, and a training sample is constructed according to the sample image.

[0049] In step 302, the recognition model and the optical element are jointly trained according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to the target image to be processed, the target image is a fruit and vegetable image to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by the optical element after modulating the acquired first light signal, the first light signal is a light signal emitted by the light source and then irradiated on the fruit and vegetable to be identified and then reflected into the optical element, and the recognition result includes the category recognition result and freshness recognition result of the fruit and vegetable to be identified.

[0050] In some embodiments of the present disclosure, multiple sample image pairs can be obtained, and each sample image pair can include: for the same sample fruits and vegetables, when the band of the light source is set to the RGB band, the RGB image acquired by the first image collector, and when the band of the light source is set to the NIR band, the NIR image acquired by the second image collector. Then, corresponding training samples can be constructed according to each sample image pair.

[0051] For example, for a certain sample of fruits and vegetables, the wavelength band of the light source can be first set to the RGB band of 380-750nm, and the RGB image can be collected using the first image collector. Then, the wavelength band of the light source can be set to the NIR band of 850nm or 940nm, and the NIR image can be collected using the second image collector. When collecting images, the imaging effects of the sample fruits and vegetables in the RGB image and the NIR image are required to be as clear as possible.

[0052] The first image collector can be an RGB camera, and the second image collector can be a hyperspectral imaging (HSI) camera. The RGB camera receives light from the RGB band and generates an RGB image, which is mainly used for fruit and vegetable category recognition. The HSI camera receives light from the NIR band and generates an NIR image, which can be used to obtain internal information of fruits and vegetables, such as moisture, sugar, and chemical composition, and is mainly used for freshness recognition of fruits and vegetables.

[0053] According to each sample image pair, corresponding training samples can be constructed respectively. In some embodiments of the present disclosure, for each sample image pair, the following processing can be performed respectively: a first label corresponding to the RGB image in the sample image pair is obtained, the first label includes the category information of the sample fruits and vegetables, a second label corresponding to the NIR image in the sample image pair is obtained, the second label includes the freshness level information of the sample fruits and vegetables, and at least the RGB image, the first label, the NIR image and the second label are used to form a training sample.

[0054] There is no restriction on how to obtain the first label and the second label. For example, for RGB images, they can be processed directly, that is, the categories of sample fruits and vegetables in the image can be manually labeled according to clear visual features. For NIR images, since the sample fruits and vegetables may have internal rot and other conditions that are difficult for the human eye to distinguish, a clustering algorithm can be used to analyze the spectral data in the NIR image to achieve accurate classification of freshness. For example, five freshness levels can be included, represented by 0, 1, 2, 3 and 4, respectively, 0 means fresh, 1 means relatively fresh, 2 means generally fresh, 3 means poor freshness, and 4 means not fresh.

[0055] After respectively obtaining the first label corresponding to the RGB image and the second label corresponding to the NIR image, the RGB image, the first label, the NIR image and the second label may be used to form a training sample, which may be summarized into one file.

[0056] In practical applications, each training sample may also include: the position information of each category of fruits and vegetables in the sample image. Assuming that the image in the sample image pair includes cabbage (for ease of description, referred to as target fruits and vegetables), the corresponding first label, second label and position information may be saved in the following format: [<category (class_id)>, <freshness level (freshness_id)>, <center point horizontal coordinate (x_center)>, <center point vertical coordinate (y_center)>, <width (width)>, <height (heig ht)>], specifically, it can be: [0, 2, 0.2604, 0.3704, 0.1042, 0.1852], where "0" indicates that the category of the target fruit and vegetable is "cabbage", "2" indicates that the freshness level of the target fruit and vegetable is "generally fresh", <0.2604, 0.3704, 0.1042, 0.1852> indicates the center point coordinates of the target fruit and vegetable in the image and the width and height information. Through <0.2604, 0.3704, 0.1042, 0.1852>, a rectangular frame including the target fruit and vegetable can be determined. Accordingly, if the training sample includes the position information, the recognition result generated by the recognition model can also include the position information to further enrich the content of the recognition result and facilitate user viewing.

[0057] After constructing a sufficient number of training samples, the true value data set is obtained. In practical applications, the true value data set can be divided in a random order according to the ratio of training set: validation set: test set = 6:2:2. The training set and validation set can be used for the iterative training process of the model, and the test set can be used to test the effect of the model training.

[0058] Furthermore, the recognition model and the optical element may be jointly trained according to the training samples. In some embodiments of the present disclosure, after each round of iterative training, the joint loss and classification accuracy may be determined according to the training samples used in the current round of iterative training and the output of the recognition model, respectively. In response to determining that the end condition is met according to the joint loss and the classification accuracy, the training process may be terminated. In response to determining that the end condition is not met according to the joint loss and the classification accuracy, the model parameters of the recognition model and / or the optical parameters of the optical element may be updated, and the next round of iterative training process may be performed.

[0059] In some embodiments of the present disclosure, the method of determining the joint loss may include: obtaining the classification task loss, and obtaining the product of the classification task loss and the corresponding first weight coefficient, obtaining the freshness task loss, and obtaining the product of the freshness task loss and the corresponding second weight coefficient, obtaining the parameter regularization loss, and obtaining the product of the parameter regularization loss and the corresponding third weight coefficient, and determining the sum of the three products as the joint loss.

[0060] That is, the model objective mainly includes two tasks, namely the classification task and the freshness task. Each task corresponds to its own loss. In addition, parameter regularization loss can be obtained to avoid model overfitting.

[0061] Accordingly, we have:

[0062] L=λ1L class (φ,θ)+λ2L freshness (φ,θ)+λ3L reg (φ,θ); (1)

[0063] Among them, L represents the joint loss, L class (φ,θ) represents the classification task loss, L freshness (φ,θ ) Freshness task loss, L reg (φ,θ) represents parameter regularization loss, λ1 represents the first weight coefficient, λ2 represents the second weight coefficient, and λ3 represents the third weight coefficient. The specific value of each weight coefficient can be determined according to actual needs.

[0064] According to the above processing method, the obtained joint loss can include loss information corresponding to different tasks at the same time, and overfitting of the model can be avoided, thereby improving the accuracy of the obtained joint loss.

[0065] In some embodiments of the present disclosure, the classification task loss can be determined by a cross-entropy loss function based on the number of target samples, the number of fruit and vegetable categories, the first label in each target sample, and the category recognition results generated by the recognition model for each target sample to optimize the classification accuracy. The target samples are the training samples used in this round of iterative training.

[0066] If available:

[0067]

[0068] Where N represents the number of training samples used in this round of iterative training, that is, the number of target samples, and C represents the number of fruit and vegetable categories. For example, if 50 different fruit and vegetable categories are defined in advance, then the value of C is 50. i,cIndicates the true label of target sample i belonging to category c, that is, if target sample i belongs to category c, then y i,c The value is 1, otherwise it is 0. It represents the probability that the target sample i identified by the recognition model belongs to category c, and the probability is greater than or equal to 0 and less than or equal to 1.

[0069] In addition, in some embodiments of the present disclosure, the freshness task loss can be determined by a mean square error (MSE) function based on the number of target samples, the second label in each target sample, and the freshness recognition results generated by the recognition model for each target sample. The target samples are the training samples used in this round of iterative training.

[0070] If available:

[0071]

[0072] Among them, f i represents the true label of the freshness of the target sample i, that is, the probability value corresponding to the freshness level of the target sample i, It represents the freshness probability value of the target sample i identified by the recognition model. Different freshness levels can correspond to their own freshness probability value ranges.

[0073] In the scheme disclosed in the present invention, φ is used to represent the optical parameters of the optical element, and θ is used to represent the model parameters of the recognition model. In order to achieve the joint optimization of fruit and vegetable classification and freshness assessment, a joint training framework is proposed to incorporate the optical parameters φ and the model parameters θ into the joint optimization process.

[0074] The joint optimization strategy can be as follows: the optical parameter φ is optimized using a gradient-based algorithm, such as the Adaptive Moment Estimation (ADAM) algorithm, and the model parameter θ is optimized using back-propagation to optimize the network weights.

[0075] That is, the joint update rule can be:

[0076]

[0077] Among them, η φ and η θ Respectively represent the learning rate, and the specific value can be determined according to actual needs. They represent the gradients of optical parameters and model parameters respectively, and can be determined based on the joint loss, etc.

[0078] In the joint training process, strategies such as cosine annealing or learning rate hot restart can be used to make the training process more stable. In addition, after each round of iterative training process, it can be determined whether the end condition is met based on the joint loss and classification accuracy. If so, the training process can be ended, otherwise, the next round of iterative training process can be executed.

[0079] Among them, meeting the termination condition may mean that the joint loss converges to a first threshold and the classification accuracy reaches a second threshold. The specific values ​​of the first threshold and the second threshold may be determined according to actual needs.

[0080] How to determine the classification accuracy can also be determined according to actual needs. For example, assuming that the validation set used in this round of iterative training includes a total of 100 kinds of fruits and vegetables, after this round of iterative training, the recognition model correctly identifies the categories of 85 kinds of fruits and vegetables in the validation set, then the classification accuracy can be considered to be 85%.

[0081] The updated model parameters may include convolution kernel weights and fully connected layer weights, etc. In addition, in some embodiments of the present disclosure, the optical element may include: a metasurface element, the metasurface element includes a plurality of nanostructure units, the nanostructure units include titanium dioxide square columns and / or silicon nitride square columns, titanium dioxide and silicon nitride have the characteristics of low absorption, broadband response and high refractive index in the visible and near-infrared regions, and have high chemical and thermal stability and corrosion resistance, which help to extend the life of the metasurface element, and the updated optical parameters may include: the rotation angle and aspect ratio of the nanostructure units in the metasurface element.

[0082] The metasurface element in the scheme described in the present disclosure needs to have excellent optical performance in both the RGB band and the NIR band to meet the multi-band requirements of fruit and vegetable identification. In the design process, it is necessary to ensure that when light passes through the column structure, the light in the RGB band and the NIR band maintains a high transmittance within the same column parameter variation range. To achieve this goal, the nanostructure unit in the scheme described in the present disclosure may include titanium dioxide square columns and / or silicon nitride square columns.

[0083] In addition, analysis shows that the design requirements of high transmittance and consistent phase shift in both bands can be achieved by combining the following two key means: 1) Adjustment of the rotation angle: The rotation angle can effectively adjust the polarization state and phase response of the light, especially balancing the transmission performance of the light between the two bands; 2) Adjustment of the aspect ratio: By adjusting the aspect ratio, the transmittance and phase shift response can be finely adjusted, so as to find a compatible optical parameter region in the RGB band and the NIR band.

[0084] Accordingly, in the solution disclosed in the present invention, the rotation angle and aspect ratio of the nanostructure unit can be set as learnable parameters. These parameters can be automatically adjusted in a data-driven manner during the joint optimization process with the recognition model, thereby finding the optimal solution with excellent transmittance and phase shift characteristics in both the RGB band and the NIR band.

[0085] During the training phase, the output of the metasurface element can be determined by the optical response or optical model of the metasurface determined by existing optical simulation methods, that is, the metasurface element can be a simulated optical response relationship, and its output can be determined by calculation. In addition, the output of the metasurface element can also be determined by constructing a physical metasurface for experiments.

[0086] By jointly optimizing the metasurface elements and recognition models and performing end-to-end training, the fruit and vegetable recognition system can work stably in various environments such as low temperature and low light and provide efficient and accurate fruit and vegetable recognition capabilities, thereby improving the accuracy of fruit and vegetable classification and freshness assessment, avoiding local optimality, simplifying the design process and shortening the R&D cycle. It is an innovative breakthrough in the field of combining multi-band optics and deep learning, and has broad application potential.

[0087] After the training is completed, the metasurface element can be prepared based on the optical parameters obtained by training, that is, the metasurface element can be designed based on the learned optical parameters, and can be realized through technologies such as lithography and packaging. Some process errors will inevitably occur during the preparation process. Therefore, when the metasurface element is prepared, the recognition model can also be tested, such as using the recognition model to perform category recognition on the target image generated by the prepared metasurface element, and it can be determined whether the classification accuracy of the recognition model meets the requirements based on the category recognition results. If so, the metasurface element and the recognition model can be used for actual fruit and vegetable recognition. Otherwise, the recognition model can be fine-tuned to make the classification accuracy of the recognition model meet the requirements. If necessary, the metasurface element can also be fine-tuned.

[0088] Combined with the above introduction, Figure 4 Schematic diagram of the overall implementation process of the joint training of the recognition model and the optical element and the fruit and vegetable recognition method disclosed in the present invention. Taking the optical element as a metasurface element as an example, Figure 4 As shown, sample image pairs can be obtained first, and then labels corresponding to each sample image pair, i.e., annotation results, can be obtained respectively, and then training samples can be constructed according to the annotation results. Figure 4As shown, the recognition model and the metasurface element can be jointly trained according to the training samples, that is, multiple rounds of iterative training can be performed, wherein after each round of iterative training, the joint loss and classification accuracy can be determined according to the training samples used in the current round of iterative training and the output of the recognition model, respectively. If the end condition is determined to be met according to the joint loss and the classification accuracy, the training process can be terminated, otherwise, the model parameters of the recognition model and / or the optical parameters of the metasurface element can be updated, and the next round of iterative training process can be performed. Figure 4 As shown, after the training process is completed, a metasurface element can be prepared according to the learned optical parameters such as the rotation angle and aspect ratio, and the metasurface element can be used to obtain the target image to be processed, and then the recognition model can be used to generate the recognition result corresponding to the target image, including the category recognition result and freshness recognition result of the fruits and vegetables to be identified.

[0089] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the described order of actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0090] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0091] Figure 5 FIG. 5 is a schematic diagram of the composition structure of the fruit and vegetable identification device embodiment 500 of the present disclosure. Figure 5 As shown, it includes: an image acquisition module 501 and an image recognition module 502.

[0092] The image acquisition module 501 is used to acquire a target image to be processed. The target image is an image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal. The second light signal is a light signal obtained by modulating the acquired first light signal by the optical element. The first light signal is a light signal emitted by the light source and then irradiated on the fruits and vegetables to be identified and then reflected into the optical element.

[0093] The image recognition module 502 is used to generate a recognition result corresponding to the target image using a recognition model, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

[0094] In some embodiments of the present disclosure, the light source may include: a wide-band light source, and the target image may include: a target image including both RGB image information and NIR image information.

[0095] In addition, in some embodiments of the present disclosure, the optical element may include: a metasurface element.

[0096] Figure 6 FIG. 6 is a schematic diagram of the structure of an embodiment 600 of a combined training device for a recognition model and an optical element according to the present disclosure. Figure 6 As shown, it includes: a sample construction module 601 and a model training module 602.

[0097] The sample construction module 601 is used to obtain sample images and construct training samples according to the sample images.

[0098] The model training module 602 is used to jointly train the recognition model and the optical element according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to the target image to be processed, the target image is the image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by the optical element after modulating the acquired first light signal, the first light signal is a light signal emitted by the light source and then irradiated by the light to be identified and then enters the optical element through reflection, and the recognition result includes the category recognition result and freshness recognition result of the fruits and vegetables to be identified.

[0099] In some embodiments of the present disclosure, the sample construction module 601 can obtain multiple sample image pairs, each of which can include: for the same sample fruits and vegetables, an RGB image collected by the first image collector when the band of the light source is set to the RGB band, and a NIR image collected by the second image collector when the band of the light source is set to the NIR band. Thereafter, corresponding training samples can be constructed according to each sample image pair.

[0100] According to each sample image pair, the sample construction module 601 can respectively construct corresponding training samples. In some embodiments of the present disclosure, the sample construction module 601 can respectively perform the following processing for each sample image pair: obtain a first label corresponding to the RGB image in the sample image pair, the first label includes the category information of the sample fruits and vegetables, obtain a second label corresponding to the NIR image in the sample image pair, the second label includes the freshness level information of the sample fruits and vegetables, and at least use the RGB image, the first label, the NIR image, and the second label to form a training sample.

[0101] Further, the model training module 602 may perform joint training of the recognition model and the optical element according to the training samples. In some embodiments of the present disclosure, after each round of iterative training, the model training module 602 may respectively determine the joint loss and the classification accuracy rate according to the training samples used in the current round of iterative training and the output of the recognition model, and in response to determining that the end condition is met according to the joint loss and the classification accuracy rate, the training process may be terminated, and in response to determining that the end condition is not met according to the joint loss and the classification accuracy rate, the model parameters of the recognition model and / or the optical parameters of the optical element may be updated, and the next round of iterative training process may be executed.

[0102] In some embodiments of the present disclosure, the model training module 602 may determine the joint loss in a manner that includes: obtaining the classification task loss, and obtaining the product of the classification task loss and the corresponding first weight coefficient, obtaining the freshness task loss, and obtaining the product of the freshness task loss and the corresponding second weight coefficient, obtaining the parameter regularization loss, and obtaining the product of the parameter regularization loss and the corresponding third weight coefficient, and determining the sum of the three products as the joint loss.

[0103] In some embodiments of the present disclosure, the model training module 602 can determine the classification task loss through a cross entropy loss function based on the number of target samples, the number of fruit and vegetable categories, the first label in each target sample, and the category recognition results generated by the recognition model for each target sample. The target samples are the training samples used in this round of iterative training.

[0104] In addition, in some embodiments of the present disclosure, the model training module 602 can determine the freshness task loss through the mean square error function according to the number of target samples, the second label in each target sample, and the freshness recognition results generated by the recognition model for each target sample, and the target sample is the training sample used in this round of iterative training.

[0105] Furthermore, in some embodiments of the present disclosure, the optical element may include: a metasurface element, the metasurface element includes a plurality of nanostructure units, the nanostructure units include titanium dioxide square pillars and / or silicon nitride square pillars, and the optical parameters may include: a rotation angle and an aspect ratio of the nanostructure units in the metasurface element.

[0106] The specific working processes of the above-mentioned device embodiments can refer to the relevant descriptions in the above-mentioned method embodiments and will not be repeated here.

[0107] In summary, by adopting the scheme described in the present disclosure, optical elements can be used to generate target images in real time, and based on the target images, the categories and freshness of fruits and vegetables can be accurately identified in combination with deep learning technology, thereby achieving efficient fruit and vegetable management and quality monitoring, and providing reliable technical support for the fields of fruit and vegetable preservation, intelligent inventory management and expiration reminders.

[0108] The scheme disclosed in the present invention can be applied to the field of artificial intelligence, especially to the fields of deep learning, computer vision, image processing and micro-nano optical design. Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0109] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0110] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0111] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 to a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0112] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0113] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described in the present disclosure may be executed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).

[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0116] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Electronically Programmable Read-Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0118] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0119] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0120] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0121] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for identifying fruits and vegetables, comprising: Acquire a target image to be processed, wherein the target image is an image of fruits and vegetables to be identified generated by a photosensitive element according to an acquired second light signal, wherein the second light signal is an optical signal obtained by modulating the acquired first light signal by an optical element, and the first light signal is an optical signal that enters the optical element after the light emitted by a light source irradiates the fruits and vegetables to be identified and is reflected; A recognition model is used to generate a recognition result corresponding to the target image, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

2. The method according to claim 1, wherein: The light source comprises: a broadband light source; The target image includes: a target image including red, green, and blue image information and near-infrared image information.

3. The method according to claim 1 or 2, wherein: The optical element includes: a metasurface element.

4. A joint training method of a recognition model and an optical element, comprising: Acquire a sample image, and construct a training sample according to the sample image; The recognition model and the optical element are jointly trained according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to a target image to be processed, the target image is an image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, the first light signal is a light signal that is emitted by a light source and then reflected by the light source and enters the optical element to be identified, and the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be identified.

5. The method according to claim 4, wherein: The acquiring of sample images and constructing training samples according to the sample images comprises: Acquire a pair of sample images, each of which includes: for the same sample fruits and vegetables, a red, green, and blue image acquired by a first image collector when the wavelength band of the light source is set to the red, green, and blue wavelength band, and a near-infrared image acquired by a second image collector when the wavelength band of the light source is set to the near-infrared wavelength band; According to each sample image pair, corresponding training samples are constructed respectively.

6. The method according to claim 5, wherein: The step of constructing corresponding training samples according to each sample image pair comprises: For each sample image pair, the following processing is performed: Obtaining a first label corresponding to the red, green and blue image in the sample image pair, wherein the first label includes category information of the sample fruits and vegetables; Acquire a second label corresponding to the near infrared image in the sample image pair, wherein the second label includes freshness level information of the sample fruits and vegetables; The training sample is composed of at least the red, green and blue image, the first label, the near infrared image and the second label.

7. The method according to claim 6, wherein: The joint training of the recognition model and the optical element according to the training sample comprises: After each round of iterative training, the joint loss and the classification accuracy are determined according to the training samples used in the current round of iterative training and the output of the recognition model; In response to determining that an end condition is met according to the joint loss and the classification accuracy, ending the training process; In response to determining that the termination condition is not met according to the joint loss and the classification accuracy, the model parameters of the recognition model and / or the optical parameters of the optical element are updated, and the next round of iterative training process is performed.

8. The method according to claim 7, wherein: Determining the joint loss includes: Obtaining a classification task loss, and obtaining a product of the classification task loss and a corresponding first weight coefficient; Obtaining a freshness task loss, and obtaining a product of the freshness task loss and a corresponding second weight coefficient; Obtaining a parameter regularization loss, and obtaining a product of the parameter regularization loss and a corresponding third weight coefficient; The sum of the three products is determined as the joint loss.

9. The method according to claim 8, wherein: The acquisition of classification task loss includes: According to the number of target samples, the number of fruit and vegetable categories, the first label in each target sample, and the category recognition results generated by the recognition model for each target sample, the classification task loss is determined by a cross entropy loss function, and the target samples are the training samples used in this round of iterative training.

10. The method according to claim 8, wherein: The freshness acquisition task loss includes: The freshness task loss is determined by a mean square error function according to the number of target samples, the second label in each target sample, and the freshness recognition results generated by the recognition model for each target sample. The target sample is a training sample used in this round of iterative training.

11. The method according to any one of claims 7 to 10, wherein: The optical element comprises: a metasurface element, wherein the metasurface element comprises a plurality of nanostructure units, wherein the nanostructure units comprise titanium dioxide square pillars and / or silicon nitride square pillars; The optical parameters include: a rotation angle and an aspect ratio of the nanostructure unit in the metasurface element.

12. A fruit and vegetable identification device, comprising: Image acquisition module and image recognition module; The image acquisition module is used to acquire a target image to be processed, wherein the target image is an image of fruits and vegetables to be identified generated by a photosensitive element according to an acquired second light signal, wherein the second light signal is an optical signal obtained by modulating an acquired first light signal by an optical element, and the first light signal is an optical signal that enters the optical element after light emitted by a light source irradiates the fruits and vegetables to be identified and is reflected; The image recognition module is used to generate a recognition result corresponding to the target image using a recognition model, wherein the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be recognized.

13. The device according to claim 12, wherein: The light source comprises: a broadband light source; The target image includes: a target image including red, green, and blue image information and near-infrared image information.

14. The device according to claim 12 or 13, wherein: The optical element includes: a metasurface element.

15. A joint training device for a recognition model and an optical element, comprising: Sample construction module and model training module; The sample construction module is used to obtain a sample image and construct a training sample according to the sample image; The model training module is used to jointly train the recognition model and the optical element according to the training samples, wherein the recognition model is used to generate a recognition result corresponding to a target image to be processed, the target image is an image of fruits and vegetables to be identified generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, the first light signal is a light signal emitted by a light source and then reflected by the light source and entering the optical element to be identified, and the recognition result includes a category recognition result and a freshness recognition result of the fruits and vegetables to be identified.

16. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 11.

18. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Cited By

  • Collaborative optimization and task processing method and device and optical encryption system

    CN121000828A

  • Joint training and image processing method, device and system

    CN121330429A