Food maturity identification method and device, electronic equipment and storage medium

By extracting and processing multi-view images and combining color feature information, inputting into the food maturity recognition network, the problem of inaccurate judgment of food maturity in the prior art is solved, and higher recognition accuracy is achieved.

CN120388366APending Publication Date: 2025-07-29NINGBO FOTILE KITCHEN WARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261199.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the judgment of food maturity depends on experience, and it is difficult to accurately determine the maturity of diverse foods, and the food safety is low.

Method used

By obtaining multiple perspective images of the food to be identified, feature extraction, homography transformation and aggregation are performed, and combined with color feature information, input the food maturity recognition network for maturity recognition.

Benefits of technology

It improves the accuracy of food maturity recognition and enhances the effectiveness of food characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388366A_ABST
    Figure CN120388366A_ABST
Patent Text Reader

Abstract

The invention discloses a food maturity identification method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction processing of each to-be-identified image of to-be-identified food, obtaining the initial depth feature information and color feature information corresponding to each to-be-identified image, and enabling each to-be-identified image to correspond to a visual angle; homography transformation and aggregation processing are carried out on the initial depth feature information corresponding to each to-be-recognized image, and aggregation depth feature information corresponding to each to-be-recognized image is obtained; splicing the aggregated depth feature information corresponding to each to-be-recognized image and the corresponding color feature information to obtain target feature information corresponding to each to-be-recognized image; and inputting the target feature information corresponding to each to-be-identified image into a food maturity identification network for maturity identification to obtain target maturity information of the to-be-identified food. According to the embodiment of the invention, the accuracy of food maturity identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly relates to a method, device, electronic device and storage medium for identifying the maturity of food. Background Art

[0002] With the development of the catering industry, people pay more and more attention to food safety, and the safe cooking process of food has received increasing attention. For cooking safety, the judgment of food maturity during the cooking process is of utmost importance. However, currently, the maturity of food usually depends on empirical judgment. For a variety of foods, it is difficult to accurately determine the maturity of food, and the food safety is relatively low. Summary of the Invention

[0003] In view of the above problems in the prior art, the present invention discloses a method, device, electronic device and storage medium for identifying the maturity of food, which can enhance the effectiveness of food features in the process of identifying food maturity, thereby improving the accuracy of identifying food maturity. The technical solutions disclosed by the present invention are as follows:

[0004] According to one aspect of the disclosed embodiments of the present invention, a method for identifying the maturity of food is provided, including:

[0005] Obtain at least two images to be recognized of the food to be recognized, and each image to be recognized corresponds to a viewing angle;

[0006] Perform feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized;

[0007] Perform homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized;

[0008] Perform splicing processing on the aggregated depth feature information corresponding to each image to be recognized and the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized;

[0009] Input the target feature information corresponding to each image to be recognized into a food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized.

[0010] Optionally, the performing homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized includes:

[0011] Take the initial depth feature information corresponding to any image to be recognized as the main feature information, and map the initial depth feature information corresponding to other images to be recognized onto the main feature information to obtain the updated depth feature information corresponding to each image to be recognized;

[0012] Perform aggregation processing on the updated depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized.

[0013] Optionally, the feature extraction process for each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized includes:

[0014] Based on a preset multi-layer convolutional network, perform feature extraction processing on each image to be recognized to obtain multiple types of feature information corresponding to each image to be recognized, and one type of feature information corresponding to each image to be recognized corresponds to a channel number;

[0015] Perform splicing processing on the multiple types of feature information corresponding to each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized.

[0016] Optionally, the performing feature extraction processing on each image to be recognized based on a preset multi-layer convolutional network to obtain multiple types of feature information corresponding to each image to be recognized includes:

[0017] Based on a preset multi-layer deformable convolutional network, perform feature extraction processing on each image to be recognized to obtain multiple types of feature information corresponding to each image to be recognized.

[0018] Optionally, the food maturity recognition network is trained through the following steps:

[0019] Obtain at least two sample images of the sample food and the preset maturity information of the sample food, and each sample image corresponds to a viewing angle;

[0020] Perform feature extraction processing on each sample image to obtain the initial sample depth feature information corresponding to each sample image and the sample color feature information corresponding to each sample image;

[0021] Perform homography transformation and aggregation processing on the initial sample depth feature information corresponding to each sample image to obtain the sample aggregated depth feature information corresponding to each sample image;

[0022] Perform splicing processing on the sample aggregated depth feature information corresponding to each sample image and the corresponding sample color feature information to obtain the target sample feature information corresponding to each sample image;

[0023] Input the target sample feature information corresponding to each of the sample images into the food maturity recognition network to be trained for maturity recognition, so as to obtain the predicted maturity information of the sample food;

[0024] Based on the predicted maturity information of the sample food and the preset maturity information, train the food maturity recognition network to be trained to obtain the trained food maturity recognition network.

[0025] Optionally, the training of the food maturity recognition network to be trained based on the predicted maturity information of the sample food and the preset maturity information to obtain the trained food maturity recognition network includes:

[0026] Determine loss information based on the predicted maturity information of the sample food and the preset maturity information;

[0027] Based on the loss information, train the food maturity recognition network to be trained to obtain the trained food maturity recognition network.

[0028] According to another aspect of the disclosed embodiments of the present invention, there is provided a food maturity recognition device, including:

[0029] A first acquisition module, configured to acquire at least two images to be recognized of the food to be recognized, and each image to be recognized corresponds to a perspective;

[0030] A first feature extraction module, configured to perform feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized;

[0031] A first aggregation module, configured to perform homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized;

[0032] A first splicing module, configured to splice the aggregated depth feature information corresponding to each image to be recognized with the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized;

[0033] A first recognition module, configured to input the target feature information corresponding to each image to be recognized into the food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized.

[0034] According to another aspect of the disclosed embodiments of the present invention, there is provided an electronic device for food maturity recognition, including a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the food maturity recognition method described in any one of the above.

[0035] According to another aspect of the disclosed embodiments of the present invention, there is provided a computer-readable storage medium. At least one instruction is stored in the computer storage medium, and the at least one instruction is loaded and executed by a processor to implement the food maturity recognition method described in any one of the above.

[0036] According to another aspect of the disclosed embodiments of the present invention, there is provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the food maturity recognition method described in any one of the above disclosed embodiments of the present invention.

[0037] The food maturity recognition method provided by the present invention has the following technical effects:

[0038] The present invention obtains at least two images to be recognized of the food to be recognized, where each image to be recognized corresponds to a perspective. Then, feature extraction processing is performed on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized. Furthermore, homography transformation and aggregation processing are performed on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized, and the aggregated depth feature information corresponding to each image to be recognized is spliced with the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized, which can enhance the effectiveness of food features in the food maturity recognition process. Furthermore, the target feature information corresponding to each image to be recognized is input into a food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized, thereby improving the accuracy of food maturity recognition.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is a schematic diagram of the application environment of a food maturity recognition method shown according to an exemplary embodiment;

[0042] Figure 2 is a schematic flowchart of a method for identifying food maturity shown according to an exemplary embodiment;

[0043] Figure 3 is a schematic flowchart of a method for determining initial depth feature information shown according to an exemplary embodiment;

[0044] Figure 4 is a schematic diagram of a method for determining initial depth feature information shown according to an exemplary embodiment;

[0045] Figure 5 is a schematic flowchart of a method for determining aggregated depth feature information shown according to an exemplary embodiment;

[0046] Figure 6 is a schematic structural diagram of an aggregation network shown according to an exemplary embodiment;

[0047] Figure 7 is a schematic structural diagram of a regression depth information network shown according to an exemplary embodiment;

[0048] Figure 8 is a block diagram of a device for identifying food maturity shown according to an exemplary embodiment;

[0049] Figure 9 is a block diagram of a terminal electronic device for identifying food maturity shown according to an exemplary embodiment;

[0050] Figure 10 is a block diagram of a server electronic device for identifying food maturity shown according to an exemplary embodiment. Detailed implementation manners

[0051] To enable those of ordinary skill in the art to better understand the technical solutions disclosed in the present invention, the technical solutions in the disclosed embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings disclosed in the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention disclosed herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0053] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0054] Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0055] The solution provided in the embodiments of this application involves technologies such as deep learning in artificial intelligence. Specifically, it may involve processing such as food maturity recognition based on deep learning, which will be specifically described through the following embodiments:

[0056] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the application environment of a food maturity recognition method shown according to an exemplary embodiment. The application environment may at least include a server 100 and a terminal 200.

[0057] In an optional embodiment, the server 100 can be used to perform food maturity recognition processing. For example, it can perform food maturity recognition processing based on a food maturity recognition network. The server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0058] In an optional embodiment, the terminal 200 can be used to provide services such as food maturity information to users. Specifically, the terminal 200 can include, but is not limited to, electronic devices such as smartphones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, vehicle-mounted terminals, smart TVs, etc.; it can also be software running on the above-mentioned electronic devices, such as application programs, applets, etc.; the terminal 200 can also be a smart cooking device. The operating systems running on the electronic devices in the embodiments of the present application can include, but are not limited to, Android system, IOS system, linux, windows, etc.

[0059] In addition, it should be noted that Figure 1 The shown is only an application environment of a food maturity recognition method, and the embodiments of this specification are not limited thereto.

[0060] In the embodiments of this specification, the above-mentioned server 100 and terminal 200 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0061] The following introduces a food maturity recognition method of the present application. Figure 2 It is a flowchart showing a food maturity recognition method according to an exemplary embodiment. This specification provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, there can be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the above method can include:

[0062] S201: Obtain at least two images to be recognized of the food to be recognized.

[0063] In a specific embodiment, the food to be recognized can be any food that needs to have its maturity recognized. Each image to be recognized can be collected by a camera, and each image to be recognized can correspond to a shooting angle.

[0064] S203: Perform feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized.

[0065] In a specific embodiment, each image to be recognized can be subjected to feature extraction processing through a convolutional network to obtain the initial depth feature information and the corresponding color feature information corresponding to each image to be recognized. Specifically, the initial depth feature information can be two-dimensional depth feature information, and the depth feature information can be used to characterize the distance information between the camera and the food to be recognized. In practical applications, during the cooking process of food, due to the degree of ripeness of the food, phenomena such as expansion or shrinkage will occur, and the above distance will change. Therefore, depth feature information that can characterize the distance information can be extracted from images of different perspectives of the food to be recognized, which helps to recognize the ripeness of the food subsequently.

[0066] In an alternative embodiment, as Figure 3 shown, the above-mentioned feature extraction processing of each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized may include:

[0067] S301: Based on a preset multi-layer convolutional network, perform feature extraction processing on each image to be recognized to obtain various feature information corresponding to each image to be recognized.

[0068] In a specific embodiment, one kind of feature information corresponding to each image to be recognized may correspond to a channel number. The network structure of the preset multi-layer convolutional network can be set according to actual application requirements. Specifically, the preset multi-layer convolutional network can be a three-layer convolutional network. Inputting each image to be recognized into the three-layer convolutional network for feature extraction, feature maps of three different image dimensions can be obtained. The image dimensions may include dimensions such as the height H of the image, the width W of the image, and the channel number. For example, a feature map with a resolution of (H, W) and a channel number of 8; a feature map with a resolution of (H / 2, W / 2) and a channel number of 16; a feature map with a resolution of (H / 4, W / 4) and a channel number of 32, etc.

[0069] Optionally, the above-mentioned feature extraction processing of each image to be recognized based on a preset multi-layer convolutional network to obtain various feature information corresponding to each image to be recognized may include:

[0070] Based on a preset multi-layer deformable convolutional network, perform feature extraction processing on each image to be recognized to obtain various feature information corresponding to each image to be recognized.

[0071] Specifically, the deformable convolutional network can adaptively adjust the width of the convolutional kernel to adaptively adjust the receptive field. In an environment with weak texture or local lack of texture, the width of the convolutional kernel becomes larger, thereby increasing the receptive field of the convolution, while in a region with rich texture, the width of the convolutional kernel becomes smaller, thereby reducing the receptive field, so as to extract more features in different situations.

[0072] S303: Concatenate the multiple feature information corresponding to each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized.

[0073] In a specific embodiment, multiple feature maps with different resolutions (i.e., feature information corresponding to different numbers of channels) can be concatenated to obtain a concatenated feature map (i.e., the initial depth feature), so as to aggregate features, which helps to improve the accuracy of subsequent food maturity recognition. The initial depth feature information can be a two-dimensional depth feature map corresponding to each image to be recognized.

[0074] In the above embodiment, as Figure 4 shown, concatenate the multiple feature information corresponding to multiple different image dimensions, and finally obtain the initial depth feature information corresponding to one image dimension. Each initial image to be recognized corresponds to 3 channels, namely the R channel (red channel), the G channel (green channel), and the B channel (blue channel), and the resolution is (H, W). Input each initial image to be recognized into the convolutional layer, the batch normalization layer, and the non-linear activation layer (Relu layer) in sequence, and feature maps with a resolution of (H, W) and 8 channels, a feature map with a resolution of (H / 2, W / 2) and 16 channels, and a feature map with a resolution of (H / 4, W / 4) and 32 channels can be obtained respectively. The feature map can also be input into the deformable convolutional layer to extract more features. After that, the feature maps corresponding to the 32 channels and the 16 channels can be upsampled to the same resolution (H, W) as the initial image to be recognized and concatenated to obtain a 32-channel feature map.

[0075] S205: Perform homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized.

[0076] In a specific embodiment, the aggregated depth feature information can be an aggregated depth feature map (i.e., the cost body) corresponding to the image to be recognized.

[0077] In an alternative embodiment, as Figure 5 shown, the above-mentioned homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized may include:

[0078] S501: Use the initial depth feature information corresponding to any image to be recognized as the main feature information, and map the initial depth feature information corresponding to other images to be recognized onto the main feature information to obtain the updated depth feature information corresponding to each image to be recognized.

[0079] In a specific embodiment, other images to be recognized may be images other than any one of the at least two images to be recognized. Updating the depth feature information may be the three-dimensional depth feature map (i.e., the feature volume) corresponding to each image to be recognized.

[0080] S503: Aggregate the updated depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized.

[0081] In the embodiments of this specification, for the two-dimensional depth feature map corresponding to the image to be recognized from any perspective, taking it as the main feature map, based on the principle of homography transformation and the corresponding preset homography transformation matrix, a corresponding mapping matrix can be generated to map the two-dimensional depth feature maps corresponding to the images to be recognized from other perspectives onto the main feature map, obtaining the three-dimensional depth feature maps after mapping from other perspectives corresponding to this main feature map. At the same time, this main feature map also correspondingly obtains the three-dimensional depth feature map after mapping from this perspective, that is, obtaining the three-dimensional depth feature maps from all perspectives corresponding to the main feature map from this perspective, and aggregating these three-dimensional depth feature maps to obtain an aggregated depth feature map corresponding to the image to be recognized from this perspective. Repeating the above process can obtain the aggregated depth feature maps corresponding to the images to be recognized from each perspective.

[0082] In practical applications, as Figure 6 shown, assuming there are N images to be recognized, based on the principle of homography transformation, the feature maps corresponding to the other N - 1 perspectives can be mapped onto the feature map corresponding to one perspective, thereby obtaining the feature volume corresponding to this perspective. The dimension of the feature volume is (D, C, H, W), where D is the pre-set number of depth channels, C is the number of feature channels, and H and W are the height and width of the feature volume respectively; at the same time, this perspective can also obtain a feature volume in the same way, so that N feature volumes corresponding to one perspective can be obtained. Repeating the above process multiple times can obtain N feature volumes corresponding to each perspective. Furthermore, the N feature volumes corresponding to each perspective can be input into an aggregation network for aggregation processing to obtain the cost volume corresponding to each perspective. The dimension of the cost volume is (D, C, H, W). After that, the cost volumes corresponding to each perspective obtained by aggregation can be input into a preset network to regress the depth map corresponding to each perspective. Specifically, the above preset network can be a network for regressing depth information; the specific structure of the network for regressing depth information can be determined according to actual application requirements. For example, the structure of the network for regressing depth information can be as Figure 7As shown in the figure, the specific steps to obtain the depth map corresponding to each perspective are as follows: First, the obtained cost volume (D, C, H, W) is input into the network for feature aggregation processing, where D is the number of depth channels for regression, C is the number of feature channels of the cost volume, 32 channels are taken in the figure, and H and W are the height and width of each layer of features of the cost volume; then, the network aggregates the cost volume with 32 channels into a cost volume with 1 channel, also known as the probability volume, and the dimension of the probability volume is (D, H, W); finally, the expectation of each pixel point of the probability volume is calculated along the depth direction, which is the aggregated depth feature information of the pixel point, and the dimension of the depth feature information map is (H, W).

[0083] S207: Concatenate the aggregated depth feature information corresponding to each image to be recognized with the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized.

[0084] In a specific embodiment, the aggregated depth feature information corresponding to each image to be recognized can be concatenated with the corresponding color feature information end to end to obtain the above-mentioned target feature information. It is also possible to concatenate the depth map of each perspective obtained by the above regression with the initial image to be recognized to obtain the concatenated image.

[0085] S209: Input the target feature information corresponding to each image to be recognized into the food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized.

[0086] In an optional embodiment, the above food maturity recognition network can be trained through the following steps:

[0087] Obtain at least two sample images of the sample food and the preset maturity information of the sample food;

[0088] Perform feature extraction processing on each sample image to obtain the initial sample depth feature information corresponding to each sample image and the sample color feature information corresponding to each sample image;

[0089] Perform homography transformation and aggregation processing on the initial sample depth feature information corresponding to each sample image to obtain the sample aggregated depth feature information corresponding to each sample image;

[0090] Concatenate the sample aggregated depth feature information corresponding to each sample image with the corresponding sample color feature information to obtain the target sample feature information corresponding to each sample image;

[0091] Input the target sample feature information corresponding to each sample image into the food maturity recognition network to be trained for maturity recognition to obtain the predicted maturity information of the sample food;

[0092] Based on the predicted maturity information and the preset maturity information of the sample food, the food maturity recognition network to be trained is trained to obtain a trained food maturity recognition network.

[0093] In a specific embodiment, each sample image can correspond to a perspective. Each sample food corresponds to a preset maturity information.

[0094] In a specific embodiment, the food maturity recognition network to be trained can be a deep learning network for performing maturity recognition processing by combining the images to be recognized of multiple different perspectives of the sample food; specifically, the network structure of the food maturity recognition network to be trained can be set according to the actual application requirements, and the network structure of the food maturity recognition network to be trained can be the same as the network structure of the above-mentioned food maturity recognition network.

[0095] In an alternative embodiment, the above-mentioned training of the food maturity recognition network to be trained based on the predicted maturity information and the preset maturity information of the sample food to obtain a trained food maturity recognition network may include:

[0096] Determine the loss information based on the predicted maturity information and the preset maturity information of the sample food;

[0097] Train the food maturity recognition network to be trained based on the loss information to obtain a trained food maturity recognition network.

[0098] In a specific embodiment, the loss information can be calculated in combination with a preset loss function; optionally, the preset loss function can be set according to the actual application requirements, such as an exponential loss function, a cross-entropy loss function, etc. The above loss information can characterize the accuracy of the current food maturity recognition network to be trained in performing food maturity recognition based on the color and depth feature information of the images to be recognized of multiple different perspectives of the sample food.

[0099] In a specific embodiment, the above-mentioned training of the food maturity recognition network to be trained based on the loss information to obtain a trained food maturity recognition network may include: updating the network parameters of the food maturity recognition network to be trained based on the loss information, and repeating the above training iteration steps based on the updated food maturity recognition network to be trained until the preset training convergence condition is met. The above-mentioned meeting the preset training convergence condition can be that the loss information is less than or equal to a preset loss threshold, or the number of training iteration steps reaches a preset number, etc. Specifically, the preset loss threshold and the preset number can be set according to the network accuracy and training speed requirements in actual applications.

[0100] As can be seen from the technical solutions provided in the embodiments of this specification, at least two to-be-recognized images of the to-be-recognized food are obtained in this specification, where each to-be-recognized image corresponds to a perspective. Then, feature extraction processing is performed on each to-be-recognized image to obtain the initial depth feature information corresponding to each to-be-recognized image and the color feature information corresponding to each to-be-recognized image. Furthermore, homography transformation and aggregation processing are performed on the initial depth feature information corresponding to each to-be-recognized image to obtain the aggregated depth feature information corresponding to each to-be-recognized image, and the aggregated depth feature information corresponding to each to-be-recognized image is concatenated with the corresponding color feature information to obtain the target feature information corresponding to each to-be-recognized image, which can enhance the effectiveness of food features in the food maturity recognition process. Furthermore, the target feature information corresponding to each to-be-recognized image is input into the food maturity recognition network for maturity recognition to obtain the target maturity information of the to-be-recognized food, thereby improving the accuracy of food maturity recognition.

[0101] An embodiment of the present invention further provides a food maturity recognition device, as Figure 8 shown. The device includes:

[0102] A first acquisition module 810, configured to acquire at least two to-be-recognized images of the to-be-recognized food, and each to-be-recognized image corresponds to a perspective;

[0103] A first feature extraction module 820, configured to perform feature extraction processing on each to-be-recognized image to obtain the initial depth feature information corresponding to each to-be-recognized image and the color feature information corresponding to each to-be-recognized image;

[0104] A first aggregation module 830, configured to perform homography transformation and aggregation processing on the initial depth feature information corresponding to each to-be-recognized image to obtain the aggregated depth feature information corresponding to each to-be-recognized image;

[0105] A first splicing module 840, configured to splice the aggregated depth feature information corresponding to each to-be-recognized image with the corresponding color feature information to obtain the target feature information corresponding to each to-be-recognized image;

[0106] A first recognition module 850, configured to input the target feature information corresponding to each to-be-recognized image into a food maturity recognition network for maturity recognition to obtain the target maturity information of the to-be-recognized food.

[0107] Optionally, the first aggregation module 830 includes:

[0108] A mapping unit, configured to use the initial depth feature information corresponding to any image to be recognized as the main feature information, and map the initial depth feature information corresponding to other images to be recognized onto the main feature information, so as to obtain the updated depth feature information corresponding to each image to be recognized;

[0109] An aggregation unit, configured to perform an aggregation process on the updated depth feature information corresponding to each image to be recognized, so as to obtain the aggregated depth feature information corresponding to each image to be recognized.

[0110] Optionally, the first feature extraction module 820 includes:

[0111] A first feature extraction unit, configured to perform feature extraction processing on each image to be recognized based on a preset multi-layer convolutional network, so as to obtain multiple types of feature information corresponding to each image to be recognized, and one type of feature information corresponding to each image to be recognized corresponds to a channel number;

[0112] A splicing unit, configured to perform a splicing process on the multiple types of feature information corresponding to each image to be recognized, so as to obtain the initial depth feature information corresponding to each image to be recognized.

[0113] Optionally, the first feature extraction unit includes:

[0114] A second feature extraction unit, configured to perform feature extraction processing on each image to be recognized based on a preset multi-layer deformable convolutional network, so as to obtain multiple types of feature information corresponding to each image to be recognized.

[0115] Optionally, the food maturity recognition network is trained through the following modules:

[0116] A second acquisition module, configured to acquire at least two sample images of a sample food and the preset maturity information of the sample food, and each sample image corresponds to a viewing angle;

[0117] A second feature extraction process, configured to perform feature extraction processing on each sample image, so as to obtain the initial sample depth feature information corresponding to each sample image and the sample color feature information corresponding to each sample image;

[0118] A second aggregation module, configured to perform a homography transformation and an aggregation process on the initial sample depth feature information corresponding to each sample image, so as to obtain the sample aggregated depth feature information corresponding to each sample image;

[0119] A second splicing module, configured to splice the sample aggregated depth feature information corresponding to each sample image with the corresponding sample color feature information, so as to obtain the target sample feature information corresponding to each sample image;

[0120] A second recognition module, configured to input the target sample feature information corresponding to each sample image into a food maturity recognition network to be trained for maturity recognition, so as to obtain the predicted maturity information of the sample food;

[0121] A training module, configured to train the food maturity recognition network to be trained based on the predicted maturity information of the sample food and the preset maturity information, so as to obtain the trained food maturity recognition network.

[0122] Optionally, the training module includes:

[0123] A loss information determination unit, configured to determine loss information based on the predicted maturity information of the sample food and the preset maturity information;

[0124] A training unit, configured to train the food maturity recognition network to be trained based on the loss information, so as to obtain the trained food maturity recognition network.

[0125] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0126] Figure 9 is a block diagram of a terminal electronic device for food maturity recognition shown according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a food maturity recognition method. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0127] Figure 10 is a block diagram of a server electronic device for food maturity recognition shown according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as Figure 10As shown. The electronic device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a method for identifying the maturity of food.

[0128] Those skilled in the art can understand that Figure 9 or Figure 10 the structure shown in is only a block diagram of some structures related to the disclosed solution of the present invention, and does not constitute a limitation on the electronic device to which the disclosed solution of the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0129] In an exemplary embodiment, there is also provided an electronic device for identifying the maturity of food, including a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for identifying the maturity of food in the disclosed embodiment of the present invention.

[0130] In an exemplary embodiment, there is also provided a computer-readable storage medium. At least one instruction is stored in the computer storage medium, and the at least one instruction is loaded and executed by a processor to implement the method for identifying the maturity of food in the disclosed embodiment of the present invention.

[0131] In an exemplary embodiment, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the method for identifying the maturity of food in the disclosed embodiment of the present invention.

[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate

[0133] SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0134] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the invention disclosed. The present invention is intended to cover any variations, uses, or adaptations of the invention disclosed, which follow the general principles of the invention disclosed and include known common general knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the invention disclosed are pointed out by the following claims.

[0135] It should be understood that the present invention is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for identifying the maturity of food, characterized in that The method includes: Obtaining at least two images to be recognized of the food to be recognized, with each image to be recognized corresponding to a perspective; Performing feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized; Performing homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized; Performing splicing processing on the aggregated depth feature information corresponding to each image to be recognized and the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized; Inputting the target feature information corresponding to each image to be recognized into a food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized.

2. The method according to claim 1, wherein The performing homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized includes: Using the initial depth feature information corresponding to any one image to be recognized as the main feature information, and mapping the initial depth feature information corresponding to other images to be recognized onto the main feature information to obtain the updated depth feature information corresponding to each image to be recognized; Performing aggregation processing on the updated depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized.

3. The method according to claim 1, wherein The performing feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized includes: Based on a preset multi-layer convolutional network, performing feature extraction processing on each image to be recognized to obtain various feature information corresponding to each image to be recognized, and one kind of feature information corresponding to each image to be recognized corresponds to a channel number; Performing splicing processing on the various feature information corresponding to each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized.

4. The method according to claim 3, wherein The based on a preset multi-layer convolutional network, performing feature extraction processing on each image to be recognized to obtain various feature information corresponding to each image to be recognized includes: Based on a preset multi-layer deformable convolutional network, performing feature extraction processing on each image to be recognized to obtain various feature information corresponding to each image to be recognized.

5. The method according to claim 1, wherein The food maturity recognition network is trained through the following steps: Obtaining at least two sample images of the sample food and the preset maturity information of the sample food, with each sample image corresponding to a perspective; Performing feature extraction processing on each sample image to obtain the initial sample depth feature information corresponding to each sample image and the sample color feature information corresponding to each sample image; Performing homography transformation and aggregation processing on the initial sample depth feature information corresponding to each sample image to obtain the sample aggregated depth feature information corresponding to each sample image; Performing splicing processing on the sample aggregated depth feature information corresponding to each sample image and the corresponding sample color feature information to obtain the target sample feature information corresponding to each sample image; Input the target sample feature information corresponding to each of the sample images into the food maturity recognition network to be trained for maturity recognition, and obtain the predicted maturity information of the sample food; Based on the predicted maturity information of the sample food and the preset maturity information, train the food maturity recognition network to be trained to obtain the trained food maturity recognition network.

6. The method according to claim 5, characterized in that, The training of the food maturity recognition network to be trained based on the predicted maturity information of the sample food and the preset maturity information to obtain the trained food maturity recognition network includes: Determine loss information based on the predicted maturity information of the sample food and the preset maturity information; Based on the loss information, train the food maturity recognition network to be trained to obtain the trained food maturity recognition network.

7. A food maturity recognition device, characterized in that, The device includes: A first acquisition module, configured to acquire at least two images to be recognized of the food to be recognized, and each image to be recognized corresponds to a perspective; A first feature extraction module, configured to perform feature extraction processing on each image to be recognized to obtain the initial depth feature information corresponding to each image to be recognized and the color feature information corresponding to each image to be recognized; A first aggregation module, configured to perform a homography transformation and aggregation processing on the initial depth feature information corresponding to each image to be recognized to obtain the aggregated depth feature information corresponding to each image to be recognized; A first splicing module, configured to splice the aggregated depth feature information corresponding to each image to be recognized with the corresponding color feature information to obtain the target feature information corresponding to each image to be recognized; A first recognition module, configured to input the target feature information corresponding to each image to be recognized into the food maturity recognition network for maturity recognition to obtain the target maturity information of the food to be recognized.

8. An electronic device for food ripeness recognition, characterized in that, The electronic device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the food maturity recognition method according to any one of claims 1 to 6.

9. A computer storage medium, characterized in that, At least one instruction is stored in the computer storage medium, and the at least one instruction is loaded and executed by a processor to implement the food maturity recognition method according to any one of claims 1 to 6.