Food material position determination method and device based on image recognition, and storage medium

By using image recognition technology to acquire and process the planar and cross-sectional perspective features of the ingredients in the oven, and utilizing edge detection and multiple image recognition models, the problem of inaccurate position recognition of ingredients in the oven is solved, achieving precise positioning of the ingredients in the oven and improving the baking effect.

CN120708214APending Publication Date: 2025-09-26QINGDAO HAIER TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510803050.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The oven cannot accurately identify the location of the ingredients, resulting in poor baking results.

Method used

Image recognition technology is used to obtain the planar and cross-sectional perspective features of the food, and perspective segmentation processing is used to generate the first image and the second image. Combined with edge detection and multiple image recognition models, the three-dimensional spatial position of the food in the oven is determined.

Benefits of technology

It achieves precise positioning of ingredients in the oven, improving the accuracy and consistency of baking results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708214A_ABST
    Figure CN120708214A_ABST
Patent Text Reader

Abstract

The invention discloses a food material position determination method and device based on image recognition and a storage medium, and relates to the technical field of image data processing, and the method comprises the steps: obtaining an image of a food material placed on a shelf of an oven, and carrying out the visual angle segmentation processing of the image, and obtaining a first image and a second image; and performing edge detection on the first image to obtain a boundary position of a plane where the food material is located, and determining first position information of the food material in the plane based on a relative position between the boundary position and the food material. The second image is input into the first image recognition model and the second image recognition model, and second position information of the baking tray where the food materials are placed is determined according to the relative relation between the first recognition result and the second recognition result output by the first image recognition model and the second image recognition model. And determining the spatial position of the food material in the oven according to the first position information and the second position information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image data processing, and more specifically, to a method, device, and storage medium for determining the location of food based on image recognition. Background Art

[0002] With the improvement of living standards, baking has become an increasingly popular way to heat food. In order to meet people's demand for baking, ovens have become one of the increasingly popular household appliances in modern family kitchens.

[0003] In the related art, in order to improve the effect of baking food in an oven, the oven usually guides the user to place the food in a designated position for baking.

[0004] However, users bake the ingredients according to their own habits, and the ovens in the related art cannot accurately identify the positions of the ingredients in the oven, resulting in poor baking effects. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for determining the location of ingredients based on image recognition, a storage medium, and an electronic device, to at least solve the problem in the related art of inaccurate identification of the location of ingredients in an oven, resulting in poor baking results.

[0006] According to one embodiment of the present invention, a method for determining the position of ingredients based on image recognition is provided, comprising: acquiring an image of ingredients placed on a shelf of an oven, and performing perspective segmentation processing on the image based on planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the ingredient from a planar perspective, and the second image includes an image of the oven shelf from a cross-sectional perspective; performing edge detection on the first image to obtain a boundary position of the plane where the ingredients are located, and determining first position information of the ingredients in the plane based on the relative position between the boundary position and the ingredients; inputting the second image into a first image recognition model and a second image recognition model, respectively, and determining second position information of a baking tray on which the ingredients are placed based on a relative relationship between first and second recognition results output by the first and second image recognition models, respectively, wherein the first image recognition model is configured to recognize baking tray features in the second image and enhance the baking tray features, and the second image recognition model is pre-trained using a baking tray training set, the baking tray training set including images of baking trays occluded by ingredients, so that the second image recognition model can predict the outline of the baking tray when the baking tray in the image is occluded; and determining the spatial position of the ingredients within the oven based on the first and second position information.

[0007] According to another embodiment of the present invention, there is provided a device for determining the position of ingredients based on image recognition, comprising: an image acquisition module for acquiring an image of ingredients placed on a shelf of an oven, and performing perspective segmentation processing on the image based on the plane perspective feature and the cross-sectional perspective feature of the image to obtain a first image and a second image, wherein the first image includes an image of the plane perspective of the ingredients, and the second image includes an image of the cross-sectional perspective of the shelf of the oven; a first position determination module for performing edge detection on the first image to obtain the boundary position of the plane where the ingredients are located, and determining the first position information of the ingredients in the plane based on the relative position between the boundary position and the ingredients; a second position determination module for The two images are respectively input into the first image recognition model and the second image recognition model, and the second position information of the baking tray where the ingredients are placed is determined based on the relative relationship between the first recognition result and the second recognition result output by the first image recognition model and the second image recognition model respectively, wherein the first image recognition model is used to identify the baking tray features in the second image and enhance the baking tray features, and the second image recognition model is pre-trained using a baking tray training set, and the baking tray training set includes baking tray images obscured by ingredients, so that the second image recognition model can predict the baking tray outline when the baking tray in the image is obscured; the ingredient positioning module is used to determine the spatial position of the ingredients in the oven through the first position information and the second position information.

[0008] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for determining the location of ingredients based on image recognition when running.

[0009] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for determining the location of food based on image recognition through the computer program.

[0010] The above-mentioned method for determining the position of ingredients based on image recognition first captures an image of the ingredients placed on an oven shelf. The image is captured from a perspective that includes both the planar perspective of the ingredients and the cross-sectional perspective of the oven shelf, ensuring that the captured image can be used for subsequent recognition. The image is then segmented based on the planar and cross-sectional perspective features of the image to produce a first image and a second image. The first image is then used to analyze the position of the ingredients on the plane, while the second image is used to analyze the shelf on which the baking tray is located. Edge detection is then performed on the first image, and the ingredients are accurately located using the boundary position as a reference, determining the first position information of the ingredients in the plane. The second image is then input into the first image recognition model and the second image recognition model respectively. The first image recognition model focuses on enhancing the baking tray features in the image, thereby outputting a first recognition result of baking tray recognition, while the second image recognition model focuses on the situation where the baking tray is blocked by food. Even when the food blocks the baking tray, the outline of the baking tray can be accurately predicted, thereby outputting a second recognition result of baking tray recognition. The second position information of the baking tray is obtained by combining the first recognition result and the second recognition result, thereby realizing precise positioning of the baking tray. The position of the baking tray is the vertical position of the food in the oven. Based on the first position information and the second position information, the three-dimensional spatial position of the food inside the oven can be accurately determined, thereby realizing precise recognition of the position of the food. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 1 is a schematic diagram of a hardware environment for a method for determining food location based on image recognition according to an embodiment of the present application;

[0014] Figure 2 is one of the flow charts of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0015] Figure 3 This is a second flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0016] Figure 4 is a third flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0017] Figure 5 is a fourth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0018] Figure 6 is a fifth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0019] Figure 7 is a sixth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0020] Figure 8 is a seventh flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0021] Figure 9 is a seventh flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0022] Figure 10 is an eighth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0023] Figure 11 is a ninth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0024] Figure 12 is a tenth flowchart of a method for determining food location based on image recognition according to an embodiment of the present invention;

[0025] Figure 13 4 is a structural block diagram of a device for determining food location based on image recognition according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] According to one aspect of the embodiment of the present application, a method for determining the location of food based on image recognition is provided. The method for determining the location of food based on image recognition is widely used in smart home, smart home, smart home device ecology, smart home ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned interaction method of smart home devices can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0029] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may include, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projector, a smart TV, a smart clothes drying rack, smart curtains, a smart audio / video system, a smart socket, a smart speaker, a smart fresh air system, smart kitchen and bathroom equipment, smart bathroom equipment, a smart robot vacuum, a smart window cleaning robot, a smart robot mop, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen appliance, a smart purifier, a smart water dispenser, a smart door lock, and the like.

[0030] In this embodiment, a method for determining the location of food ingredients based on image recognition is provided, which can be applied to cloud servers, ovens, edge servers, or the highest computing power device in the home. For example, when the execution subject of this application is an edge server, multiple edge servers can be deployed in a cell or adjacent cells as redundant nodes. When the main server executing the method for determining the location of food ingredients based on image recognition fails, the backup server in the cell or the backup server in other cells can seamlessly take over the responsibilities of the main server to ensure uninterrupted service. Figure 2 : is a flowchart of a method for determining food location based on image recognition according to an embodiment of the present application, the process comprising the following steps:

[0031] Step S200 , obtaining an image of food placed on a shelf of an oven, and performing perspective segmentation processing on the image based on the plane perspective feature and the cross-sectional perspective feature of the image to obtain a first image and a second image.

[0032] The first image includes an image of the plane perspective of the food, and the second image includes an image of the cross-sectional perspective of the oven shelf.

[0033] Specifically, the image of the ingredients placed on the oven shelf is captured by a camera set in the oven. The image needs to include multiple perspectives, such as a bird's-eye view and a perspective inclined to the oven wall. The plane perspective can be the perspective of the ingredients seen from above, and the cross-sectional perspective is the perspective of the ingredients and the baking tray seen from the side or perpendicular to the oven wall. Based on the specific perspective features in the image, the captured comprehensive perspective image is divided into two images, namely the first image and the second image, which are plane perspective features and cross-sectional perspective features. The first image is mainly used to show the perspective of the plane where the ingredients are located, while the second image focuses on the cross-sectional perspective of the oven shelf, so as to determine the position of the ingredients on the plane based on the first image, and determine the shelf in the oven where the baking tray carrying the ingredients is placed based on the second image.

[0034] For example, image processing techniques, such as a deep learning-based perspective separation algorithm, can be used to identify and segment image components from different perspectives. This can be achieved using a pre-trained multi-perspective recognition model that can distinguish and isolate the planar perspective and the cross-sectional perspective of the food image to generate the first image and the second image.

[0035] Step S210 , performing edge detection on the first image to obtain the boundary position of the plane where the food is located, and determining first position information of the food in the plane based on the relative position between the boundary position and the food.

[0036] Specifically, edge detection is performed on the first image to determine the outline and boundary position of the ingredient, facilitating the determination of the distribution of the ingredient on the plane of the baking pan. Based on the relative relationship between the boundary position of the ingredient and the ingredient itself, the specific position of the ingredient on the baking pan is calculated. For example, using the boundary position of the ingredient as a reference point, the distance between the ingredient and the reference point is determined.

[0037] In step S220, the second image is input into the first image recognition model and the second image recognition model respectively, and the second position information of the baking tray where the ingredients are placed is determined based on the relative relationship between the first recognition result and the second recognition result output by the first image recognition model and the second image recognition model respectively.

[0038] Among them, the first image recognition model is used to identify the baking pan features in the second image and enhance the baking pan features. The second image recognition model is pre-trained using a baking pan training set, which includes baking pan images occluded by ingredients, so that the second image recognition model can predict the baking pan outline when the baking pan in the image is occluded.

[0039] Specifically, the second image is input into two different image recognition models: a first image recognition model and a second image recognition model. These two models recognize the characteristics of the ingredients and the baking pan, respectively. The first image recognition model focuses on enhancing the characteristics of the baking pan, improving the accuracy of baking pan recognition. The second image recognition model focuses on predicting the baking pan outline even when the ingredients are obstructing it. The recognition results of the two models are compared and verified to improve the accuracy of the second position information. The specific shelf position of the ingredients is determined based on the relative relationship between the first and second recognition results.

[0040] Step S230: determining the spatial position of the food in the oven using the first position information and the second position information.

[0041] Specifically, the first position information (i.e., the position of the ingredients on the baking tray) and the second position information (i.e., the position of the shelf holding the ingredients) can be used to comprehensively determine the three-dimensional spatial position of the ingredients within the oven. For example, a three-dimensional coordinate system can be established within the oven, where the x- and y-axes represent the position of the baking tray, and the z-axis represents the height of the shelf. The first and second position information are mapped into this coordinate system, and mathematical operations (such as coordinate transformations) are used to determine the three-dimensional position coordinates of the ingredients, i.e., their specific spatial location within the oven.

[0042] In this embodiment, an image of the food placed on the oven shelf is first acquired. The image is captured from a perspective that includes both the plane view of the food and the cross-sectional view of the oven shelf, ensuring that the captured image can be used for subsequent recognition. Based on the image's plane view and cross-sectional view features, the image is segmented to produce a first image and a second image. The first image is used to analyze the food's position on the plane, while the second image is used to analyze the shelf on which the baking tray is located. Edge detection is then performed on the first image, and the food is accurately located using the boundary position as a reference, determining its first position in the plane. The second image is then input into the first image recognition model and the second image recognition model respectively. The first image recognition model focuses on enhancing the baking tray features in the image, thereby outputting a first recognition result of baking tray recognition, while the second image recognition model focuses on the situation where the baking tray is blocked by food. Even when the food blocks the baking tray, the outline of the baking tray can be accurately predicted, thereby outputting a second recognition result of baking tray recognition. The second position information of the baking tray is obtained by combining the first recognition result and the second recognition result, thereby realizing precise positioning of the baking tray. The position of the baking tray is the vertical position of the food in the oven. Based on the first position information and the second position information, the three-dimensional spatial position of the food inside the oven can be accurately determined, thereby realizing precise recognition of the position of the food.

[0043] In one embodiment, Figure 3 As shown, step S210, edge detection is performed on the first image to obtain the boundary position of the plane where the food is located, and based on the relative position between the boundary position and the food, the first position information of the food in the plane is determined. It includes: steps S300-S330, wherein:

[0044] Step S300: performing edge detection on the first image to determine at least three reference points at the boundary positions of the plane where the food is located.

[0045] Among them, at least three benchmark points are non-collinearly distributed.

[0046] Specifically, by performing edge detection on the first image of the plane perspective in which the food is located, at least three non-collinear reference points of the plane in which the food is located are determined, so that the plane in which the food is located can be located by at least three reference points. Edge detection is an image processing technology used to identify the boundaries or outlines of objects in an image, so that the specific shape and position of the object can be determined. Reference points refer to the basic reference points used to construct a coordinate system. Non-collinear distribution means that these points are not on the same straight line. This is to allow a plane to be determined by at least three reference points.

[0047] Step S310 : Taking a first reference point among the at least three reference points as an origin, determining a coordinate axis according to a first vector pointing from the first reference point to the second reference point and a second vector pointing from the first reference point to the third reference point.

[0048] Step S320: establishing a first coordinate system of the plane where the food is located according to the origin and the coordinate axes.

[0049] Specifically, a first coordinate system of the plane where the food is located is established through the selected reference points, and then the position of the food in the plane is quantified, that is, the first position information of the food.

[0050] For example, a reference point is chosen as the origin, typically the lower left or lower right corner of the plane. The coordinate axes are then determined by the first and second vectors pointing to the other two reference points in the plane. If the lower right corner is chosen as the origin, the vector pointing to the lower left corner of the plane can be used as the x-axis vector, and the vector pointing to the top of the plane can be used as the y-axis vector.

[0051] Step S330 , determining the center of mass position of the food according to the outline image of the food, and indicating first position information of the food by the coordinates of the center of mass position in the first coordinate system.

[0052] Specifically, by quantifying the center of mass position of the food in the first coordinate system, specific position information of the food on the plane of the baking tray, that is, first position information, is obtained.

[0053] For example, the center of mass coordinates of the food are converted from pixel coordinates to real-world coordinates in the first coordinate system. This can be done by performing coordinate mapping based on the dimensions of the baking pan or oven, ensuring that the units of the coordinate system are consistent with real-world spatial units, such as centimeters or millimeters. The resulting center of mass coordinates are the first position information of the food on the baking pan plane.

[0054] For example, the location of the food in the first coordinate system can be determined using a method for calculating the center of mass. The center of mass is determined by the food outline image, and the center of mass is found by calculating the average pixel position of the outline. The position coordinates of this center of mass are then quantized in the established coordinate system.

[0055] The outline image of the food is determined based on the first image.

[0056] In this embodiment, edge detection technology is first used to accurately identify the boundaries of the plane containing the ingredients in the first image. Based on these boundaries, at least three non-collinear reference points are then determined, providing a stable reference for the subsequent establishment of a coordinate system. This coordinate system quantifies the position of the ingredients. By calculating the center of mass of the ingredient's outline and converting this location into actual coordinates in the coordinate system, the precise position of the ingredients on the baking sheet, or first position information, is obtained.

[0057] In one embodiment, Figure 4 As shown, step S300 is to perform edge detection on the first image to determine at least three reference points at the boundary position of the plane where the food is located. It includes steps S400-S430, wherein:

[0058] Step S400: performing corner detection on the first image to determine first corner point coordinates of at least three corner points of a baking tray carrying food.

[0059] Among them, at least three corner points serve as at least three reference points.

[0060] Specifically, a corner detection algorithm can be used to identify at least three corners of the baking sheet where the ingredients are placed. These corners serve as the basis for subsequent positioning and coordinate transformation. Corner detection is a computer vision technique used to identify points in an image with large variations in pixel intensity. These points are typically located at the corners of an object and are a key step in image feature extraction.

[0061] For example, the Harris corner detection algorithm or the Shi-Tomasi corner detection method can be used to identify the corners of the baking tray. These algorithms can detect the edges and corners of the baking tray and determine the corner locations by finding areas with high gradient changes on the baking tray contour.

[0062] Step S410 , determining a perspective transformation matrix according to a first viewing angle of a plane on which the food is located and a second viewing angle perpendicular to the plane on which the baking tray is located, used to obtain the first image.

[0063] The perspective transformation matrix is ​​used to indicate the conversion relationship between the coordinates of the target point in the first perspective and the coordinates of the target point in the second perspective.

[0064] Specifically, calibration is performed based on images captured from a first perspective and images captured from a second perspective. A perspective transformation matrix is ​​calculated. This matrix describes the conversion relationship between coordinate points between the two perspectives. This facilitates converting the camera's oblique perspective to a vertical perspective, resolving the issue of near objects appearing larger than far objects within the field of view. Perspective transformation is an image processing technique used to simulate the visual effects of objects in three-dimensional space from different perspectives. Matrix operations can be used to transform an image from one perspective into an image from another.

[0065] For example, the detected corner coordinates can be used in combination with a pre-calibrated perspective transformation matrix to perform coordinate transformation, converting the corner coordinates under the first perspective into the corner coordinates under the second perspective, thereby eliminating the influence of different perspectives on the coordinates.

[0066] Step S420: Using a perspective transformation matrix, convert the first corner point coordinates of at least three corner points into second corner point coordinates.

[0067] The coordinates of the second corner point are the coordinates at the second viewing angle.

[0068] Step S430: Using the coordinates of the second corner points of the at least three corner points as the coordinates of the at least three reference points to determine the at least three reference points.

[0069] In this embodiment, at least three non-collinear corner points of the baking pan are determined in the first image as reference points through corner detection technology, which provides a basis for the subsequent establishment of a coordinate system. Then, based on the conversion correspondence between the first perspective and the second perspective, a perspective transformation matrix is ​​calculated, which describes the geometric transformation relationship between the two perspectives. Then, using the perspective transformation matrix, the corner point coordinates under the first perspective are converted to the corner point coordinates under the second perspective, realizing the coordinate conversion from the oblique perspective to the normal perspective. Finally, the converted coordinates are used as the reference point coordinates for the subsequent establishment of the coordinate system of the plane where the food is located, thereby accurately quantifying the position information of the food. Through corner detection and perspective transformation technology, the problem of positioning the food under different perspectives is effectively solved, ensuring the accuracy and reliability of the position information of the food on the baking pan.

[0070] In one embodiment, Figure 5 As shown, step S330 determines the center of mass of the food according to the outline image of the food. It includes steps S500-S520, wherein:

[0071] Step S500: performing boundary smoothing processing on the outline image of the food.

[0072] Specifically, we reduce noise on the edges of the food outline image to make the edges smoother and facilitate subsequent centroid calculation. Boundary smoothing is a common preprocessing step in image processing, primarily used to remove irregularities or noise points on image boundaries, making them clearer and more coherent.

[0073] For example, the boundary smoothing can be achieved by using techniques such as mean filtering or Gaussian filtering. Specifically, by setting an appropriate filter window size, the food outline image can be filtered to remove small noise points on the edge and make the edge smoother.

[0074] Step S510 : determining the zero-order moment and the first-order moment of the contour image according to the positions and grayscale values ​​of the pixels included in the contour image after the boundary smoothing process.

[0075] The zero-order moment is used to indicate the area of ​​the contour image, and the first-order moment is used to indicate the distribution of the pixels of the contour image.

[0076] Specifically, the area and pixel distribution of the food outline are statistically analyzed and quantified using zero-order and first-order moment calculations. The zero-order moment reflects the total area of ​​the outline image, equivalent to the number of pixels within the outline. The first-order moment reflects the offset of the pixel distribution within the outline image and is used to calculate the coordinates of the center of mass.

[0077] Step S520 : determining the center of mass position of the food according to the zero-order moment and the first-order moment of the contour image.

[0078] Specifically, based on the zero-order moment and the first-order moment of the contour image, the center of mass position of the food is calculated, that is, a point that can be used to represent the overall position of the food.

[0079] For example, the center of mass coordinates are calculated based on the first-order and zero-order moments of the contour image. The coordinates of the center of mass in the x-axis direction can be expressed as Cx = M10 / M00, where Cx is the coordinate of the center of mass in the x-axis direction, M00 is the zero-order moment, and M10 is the first-order moment in the x-axis direction. The coordinates of the center of mass on the y-axis are Cy = M01 / M00, where Cy is the coordinate of the center of mass in the y-axis direction, M00 is the zero-order moment, and M01 is the first-order moment in the y-axis direction. This can be used to obtain the coordinates of the center of mass of the food contour (Cx, Cy).

[0080] In this example, we first smooth the edges of the ingredient contour image to remove edge noise, ensuring the continuity and clarity of the contour boundary and providing high-quality input for subsequent moment calculations. Next, we calculate the zero-order and first-order moments of the contour image to quantify the area and pixel distribution of the ingredient, a step crucial for calculating the center of mass. Finally, based on the calculated first-order and zero-order moments, we accurately calculate the coordinates of the ingredient's center of mass. These coordinates reflect the relative position of the ingredient on the baking sheet, enabling precise positioning of the ingredient.

[0081] In one embodiment, Figure 6 As shown, step S200 is to obtain an image of food placed on a shelf of an oven. The process includes steps S600-S650, wherein:

[0082] Step S600: Acquire an image of a calibration plate placed on a preset calibration shelf of an oven.

[0083] Among them, a plurality of calibration points are set on the calibration plate.

[0084] Specifically, before beginning food recognition, the camera is calibrated to ensure the accuracy and resolvability of the captured food images. A calibration plate is a flat surface with a specific pattern (e.g., a checkerboard) used for camera calibration. Multiple calibration points with known world coordinates are set on it, which are used to calibrate camera parameters and calculate the coordinate mapping between images and the real world.

[0085] For example, a camera installed inside the oven captures a photo of a calibration plate placed on a pre-set calibration shelf within the oven, generating an image containing the calibration points. This image is then preprocessed, including brightness and contrast adjustments, and filtering and denoising, to improve the accuracy of subsequent feature detection.

[0086] Step S610: determining first world coordinates of a plurality of calibration points according to the image of the calibration plate.

[0087] Specifically, image processing techniques are used to identify the positions of the calibration points on the calibration plate and convert them into coordinate values ​​within the oven's three-dimensional space. The world coordinate system is used to describe the positions of objects in the real world. The first world coordinates of the calibration points refer to their three-dimensional coordinate values ​​within the oven's internal coordinate system.

[0088] Exemplarily, based on the position of the preset calibration plate in the oven coordinate system, the position of the detected feature point in the image coordinate system is converted into the first world coordinate, that is, the actual three-dimensional space coordinate.

[0089] Step S620 : determining coordinate errors based on the determined first world coordinates of the plurality of calibration points and the actual second world coordinates of the plurality of calibration points.

[0090] Specifically, by comparing the actual second-world coordinates of the calibration points with the calculated first-world coordinates, the accuracy of the mapping relationship between camera parameters and coordinates is evaluated. Coordinate error is the deviation between the actual measured coordinates and the theoretical coordinates, and is used to evaluate the accuracy of the calibration results.

[0091] For example, the Euclidean distance between the first world coordinate and the second world coordinate can be calculated to obtain the coordinate error of each calibration point. If the coordinates of the points on the calibration plate are known, the error can be calculated directly. If they are not completely known, the least squares method or other optimization algorithm can be used to estimate the coordinates of these points and calculate the error.

[0092] Step S630: Determine whether the coordinate error is greater than the error threshold. If the coordinate error is greater than the error threshold, execute step S640. If the coordinate error is less than or equal to the error threshold, execute step S650.

[0093] In step S640, the parameters of the camera used to capture the image are adjusted based on the coordinate error, as well as the mapping relationship between the first world coordinate and the second world coordinate. The process then returns to step S610 to determine the first world coordinates of the plurality of calibration points based on the image of the calibration plate until the coordinate error is less than or equal to the error threshold.

[0094] Specifically, when the coordinate error exceeds a predetermined threshold, the camera's intrinsic parameters and mapping relationships are adjusted until the calibration result is sufficiently accurate. Camera parameter adjustment primarily involves adjusting intrinsic parameters such as the camera's focal length and optical center to reduce image distortion. Coordinate mapping relationship adjustment involves correcting the conversion relationship from image coordinates to world coordinates.

[0095] Step S650: Acquire an image of the food placed on the oven shelf.

[0096] Specifically, after the camera calibration is complete and the error is within the threshold range, images of the food placed on the oven shelves are captured for subsequent food identification and location determination. Once the camera parameters are adjusted and the calibration is verified, the camera can be used to capture images of food placed on the oven shelves.

[0097] In this embodiment, the oven system first obtains the camera's intrinsic parameters and the first-world coordinates of the calibration points in the oven coordinate system through calibration plate image and feature point detection. This forms the basis of the calibration process. Subsequently, by comparing the first-world coordinates with the theoretical second-world coordinates of the calibration points, the coordinate error is calculated to assess the accuracy of the calibration. If the coordinate error exceeds the threshold range, the system automatically adjusts the camera parameters and coordinate mapping relationship until the calibration error meets the requirements. This process ensures that the camera can accurately capture food images and reduces recognition errors caused by improper device parameter settings or inaccurate coordinate conversion. Finally, only after the calibration error meets the standard will the camera capture the food image for subsequent food identification and spatial position determination.

[0098] In one embodiment, Figure 7 As shown, step S220 inputs the second image into the first image recognition model and the second image recognition model respectively, and determines the second position information of the baking tray where the ingredients are placed based on the relative relationship between the first recognition result and the second recognition result output by the first image recognition model and the second image recognition model respectively. It includes steps S700-S750, wherein:

[0099] Step S700: input the second image into the first image recognition model to obtain a first recognition result output by the first image recognition model.

[0100] Specifically, the collected image of the interior of the oven (i.e., the second image) is input into the first image recognition model, the baking tray is identified by the model, and the shelf information where the baking tray is placed and the model's confidence in the recognition result are output.

[0101] Step S720: When the first recognition result and the second recognition result are inconsistent, determine a first weighted confidence level according to the first confidence level carried by the first recognition result and a preset first weight corresponding to the first image recognition model.

[0102] Wherein, when the first recognition result and the second recognition result are consistent, the second position information of the baking tray is determined according to the first recognition result or the second recognition result.

[0103] Specifically, the second image is also input into the second image recognition model, which is designed to identify the ingredients inside the oven, but the output result also includes the shelf information and confidence level of the baking tray where the ingredients are located.

[0104] Step S730: Determine a second weighted confidence level based on the second confidence level carried by the second recognition result and a preset second weight corresponding to the second image recognition model.

[0105] Specifically, when the shelf information given by the first recognition result and the second recognition result are inconsistent, the weighted confidence is calculated based on the first confidence of the first recognition result and the second confidence of the second recognition result, combined with the weights preset by the model (the first weight and the second weight).

[0106] For example, based on a pre-assessment of model performance, weights can be set for the first and second image recognition models (e.g., the first weight is 0.3 and the second weight is 0.7). Weighted confidence is calculated using these weights and the confidence level of the model output. The stronger the model's predicted performance, the higher the corresponding weight can be set. For example, the second image recognition model is able to identify an obscured baking tray, so it may have stronger recognition capabilities and can be designed with a higher weight. The weighted confidence is obtained by multiplying the confidence level by the corresponding weight.

[0107] Step S740: Determine the target result according to the relative relationship between the first weighted confidence and the second weighted confidence.

[0108] The target result is one of the first recognition result and the second recognition result.

[0109] Specifically, the first weighted confidence and the second weighted confidence are compared, and the recognition result with the higher weighted confidence is selected as the final target result.

[0110] Step S750: determining second position information according to the target shelf indicated by the target result.

[0111] Specifically, the second location information, namely the specific location and shelf of the baking tray within the oven, is determined based on the shelf information specified in the target result. Since the positions of the multiple shelves within the oven (e.g., six shelves) are fixed and known in advance, once the target shelf for the baking tray is identified, the second location information can be obtained by querying the location of this target shelf.

[0112] In this embodiment, the baking tray is identified by two different image recognition models. By using a weighted confidence method, based on the preset weights of the models and the confidence of the recognition results, the more reliable recognition result is intelligently selected as the target result, which can improve the recognition accuracy of the second position information. This method enhances the robustness and accuracy of the recognition system by combining model confidence with preset weights, especially when the internal environment of the oven is complex, the lighting conditions are changeable, or the baking tray is blocked by food, and can provide more stable and accurate baking tray position information. For example, in one scenario, the first image recognition model may have a low recognition confidence for the baking tray position due to food blocking, while the second image recognition model has good recognition accuracy even when blocked by food, and may provide a higher confidence. Therefore, the recognition results of the two models are combined to determine the baking tray position, which is more accurate and reliable.

[0113] In one embodiment, Figure 8 As shown, step S700, inputting the second image into the first image recognition model to obtain the first recognition result output by the first image recognition model. It includes: steps S800-S850, wherein:

[0114] Step S800: Input the second image into the first image recognition model, and extract baking pan features of the second image through multiple convolutional layers of the first image recognition model.

[0115] The multiple convolutional layers of the first image recognition model are jump-connected, and the baking tray features include baking tray edge features and baking tray seam features.

[0116] Specifically, the captured second image is fed into a pre-trained first image recognition model. The model's multiple convolutional layers extract features of the baking pan, including its edges and seams. The convolutional layers in the model use skip connections to enhance feature extraction. Skip connections are a technique used to mitigate vanishing or exploding gradients as network depth increases. They allow the network to propagate gradients more efficiently during training, thereby helping to learn deeper feature representations.

[0117] For example, a deep learning model that includes skip connections, such as a Residual Network (ResNet) or a Convolutional Network for Biomedical Image Segmentation (U-Net), can be used as the first image recognition model. Ensure that the model includes skip connections to preserve the hierarchical structure and details of the baking sheet features.

[0118] Step S810: Determine the channel attention weights corresponding to the output channels of the multiple convolutional layers according to the baking tray features output by the output channels of the multiple convolutional layers.

[0119] Among them, the channel attention weight corresponding to the output channel of each convolutional layer is determined by learning the dependency relationship between the output channels of multiple convolutional layers. The channel attention weight is related to the degree of influence of the output channel on the first recognition result.

[0120] Specifically, based on the outputs of multiple convolutional layers in the first image recognition model, the channel attention weights are calculated for each convolutional layer output channel. These weights reflect the contribution of each channel to the first recognition result (the baking pan position and its confidence). Channel attention weights measure the importance of each model channel's output to the recognition result. Learned by analyzing the dependencies between channels, they significantly enhance the model's ability to capture key features.

[0121] For example, a squeeze-and-excitation (SE) network structure can be used to perform a channel-wise attention mechanism on the model's convolutional layer outputs to obtain weights for each output channel. Adding an SE module to the model performs channel-wise weighting on the output of each convolutional layer to enhance the representation of baking pan features.

[0122] Step S820: Determine the key area in the second image based on the baking tray features output by the output channels of the multiple convolutional layers.

[0123] Among them, the key areas include the edge area of ​​the baking pan and the seam area of ​​the baking pan.

[0124] Specifically, based on the baking tray features extracted by the convolutional layer, key areas in the image are determined. These areas include the edges and seams of the baking tray. The features of these areas are crucial for identifying the location of the baking tray.

[0125] For example, the pan features are mapped back into the second image, and image processing techniques such as edge detection or the Canny edge extraction algorithm are used to locate the pan edges and seam areas. Feature enhancement is then performed on the located key areas, for example by increasing the pixel values ​​in these areas or applying a Gaussian filter, to increase the model's sensitivity to key features.

[0126] Step S830: enhancing the spatial attention weight corresponding to the key area in the second image.

[0127] Specifically, through iterative learning, the model increases the spatial attention weights corresponding to key areas of the baking pan. This helps the model focus more on these areas and improves the accuracy of baking pan recognition. Spatial attention weights measure the contribution of different image regions to the model's recognition results. By increasing the weights of key areas, the model is encouraged to pay more attention to key areas such as the edges and seams of the baking pan.

[0128] For example, a spatial attention module (such as the Spatial Attention Module) is used to introduce an attention mechanism into the model, enhancing the spatial attention weights for key areas of the baking pan. Spatial attention weights are calculated based on the feature importance of key areas. This can be achieved through global average pooling after the convolutional layer and a fully connected layer combined with a softmax function.

[0129] Step S840: Adjust the weights of corresponding model parameters in the first image recognition model according to the channel attention weight and the spatial attention weight.

[0130] Specifically, based on the calculated channel attention weights and spatial attention weights, the corresponding model parameter weights in the first image recognition model are adjusted to optimize the model's performance for baking pan recognition. Model parameter weights are the weight coefficients of the convolutional layers in the model, which determine how the model processes input features.

[0131] For example, during model training, backpropagation is used to update the weights of the convolutional layers based on the channel-wise and spatial-wise attention weights to optimize the model's ability to recognize baking pans. Optimizers such as Adaptive Moment Estimation (Adam) and Stochastic Gradient Descent (SGD) can be used, combined with channel-wise and spatial-wise attention mechanisms, to adjust model parameters and improve the model's performance on the baking pan recognition task.

[0132] Step S850: Using the first image recognition model after the weights of the model parameters are adjusted, the second image is recognized to obtain a first recognition result.

[0133] The first recognition result includes the shelf where the baking tray is placed and the corresponding first confidence level.

[0134] Specifically, the second image is recognized using the first image recognition model after the channel attention and spatial attention weights are adjusted, and the shelf information where the baking tray is placed and the corresponding confidence level are output as the first recognition result.

[0135] For example, the second image is fed into a weighted first image recognition model. The model outputs a first recognition result, including the shelf information and confidence level of the baking tray. The model output is post-processed, such as using a non-maximum suppression algorithm, to optimize the recognition result, remove redundant predictions, and ensure the most accurate baking tray location recognition.

[0136] In this example, skip connections are first used to enhance convolutional layer feature extraction, ensuring that the model captures the edges and seams of the baking pan. Next, a channel-wise attention mechanism is introduced to learn and determine the contribution of each channel to the recognition result, thereby optimizing the model's capture of baking pan features. Then, a spatial attention mechanism is used to identify key regions in the image (i.e., the edges and seams of the baking pan) and enhance the feature representation of these regions. Subsequently, based on the channel-wise and spatial-wise attention weights, the model parameters are adjusted to further optimize the model's recognition performance, focusing more on the features of key regions. Finally, the adjusted model is used to perform a prediction, outputting the first recognition result, which includes the baking pan's location and its confidence level.

[0137] In one embodiment, Figure 9 As shown, step S710, inputting the second image into the second image recognition model to obtain the second recognition result output by the second image recognition model. It includes: steps S900-S930, wherein:

[0138] Step S900: Obtain a baking tray training set.

[0139] The baking tray training set includes an unobstructed first baking tray image placed in an oven and an unobstructed second baking tray image placed in an oven and obstructed by food.

[0140] Specifically, obtaining a baking tray training set involves collecting a series of images containing baking trays placed inside an oven. These images are divided into two categories: one is unobstructed baking tray images, meaning there is no food obstructing the baking tray, clearly showing the complete outline of the baking tray; the other is food obstructed baking tray images, meaning the food is placed on the baking tray, partially or completely blocking the view of the baking tray.

[0141] For example, we collected images using the oven's built-in camera, capturing images of both unobstructed and obscured baking pans on different shelves, ensuring that the images covered all possible positions and angles of baking pans inside the oven. These images were organized and labeled to create a baking pan training set for subsequent model training.

[0142] Step S910 : performing a delineation and annotation process on the second baking pan image in the baking pan training set to outline the outline of the baking pan that is blocked in the second baking pan image.

[0143] Specifically, the depiction and annotation processing is to manually or automatically annotate the image of the baking pan that is blocked by the food, outlining the outline of the baking pan in the picture. Even if it is partially blocked, the complete outline of the baking pan must be restored through technical means such as edge completion or interpolation.

[0144] For example, annotation software can be used to manually outline the complete outline of the obscured baking pan in the image. For the obscured baking pan portion, image processing techniques such as edge detection and contour connection can be used to assist the annotator in completing the outline of the baking pan.

[0145] Step S920 , using the first baking tray image and the second baking tray image after depiction and annotation processing to train the initial image recognition model to obtain a second image recognition model.

[0146] Specifically, the unobstructed baking tray image and the processed obstructed baking tray image are used to train the initial image recognition model so that it can accurately identify the specific position and shelf of the baking tray inside the oven.

[0147] For example, a pre-trained image recognition model, such as an existing deep learning model for image recognition, can be selected as the starting point for model training. Labeled baking pan images and unobstructed baking pan images are combined to form the final training dataset, which is then fed into the model for training. The model parameters are adjusted until the model can accurately identify the baking pan's position within the oven, including the shelf and specific coordinates.

[0148] Step S930: Use a second image recognition model to recognize the second image to obtain a second recognition result.

[0149] The second recognition result includes the shelf where the baking tray is placed and the corresponding second confidence level.

[0150] In this example, a training set covering all possible configurations of baking pans inside an oven is first constructed by collecting images of both unobstructed and obscured baking pans, providing a comprehensive data foundation for model training. Next, images of obscured baking pans are meticulously delineated and annotated, ensuring that the model learns to recognize the outline of baking pans even under partial occlusion. This annotated data is then used to train an initial image recognition model. Through iterative training, the model gradually becomes capable of accurately identifying the position of baking pans on the oven shelves. This is achieved through intelligent analysis, even when ingredients obscure the edges or seams of the baking pans. Finally, the trained model is applied to a real-world scenario, identifying baking pans obscured by ingredients inside the oven. A second recognition result, including the shelf location and its confidence score, is output. This allows for precise identification of baking pans inside the oven even under complex occlusion conditions.

[0151] In one embodiment, Figure 10 As shown, the method further includes: steps S1000-S1070, wherein:

[0152] Step S1000: Obtain food training set.

[0153] The food training set includes pictures of various food items inside ovens.

[0154] Specifically, we collected pictures of a variety of different ingredients taken inside an oven environment to provide rich and diverse data for subsequent model training, ensuring that the model can learn to recognize various types of ingredients.

[0155] For example, we collected images by placing various ingredients on different shelves in an oven and capturing them using an RGB camera inside the oven. We ensured that the images covered different types of ingredients, different positions, and different lighting conditions, making the dataset comprehensive and realistic.

[0156] Step S1010: Process the food training set using a pre-trained image recognition network to extract the border features of the food in each picture.

[0157] Specifically, an image recognition network that has been trained on large-scale datasets (such as real-time object detection models (You Only Look Once, YOLO, Faster Regions with Convolutional Neural Networks, Faster R-CNN)) is used to process each image in the food training set to identify and extract the bounding box features of the food, including the location, size, and boundary information of the food.

[0158] Step S1020: Clustering the border features of the food in each picture of the food training set using a clustering algorithm to obtain multiple prior anchor frames corresponding to the preset baking tray size.

[0159] Specifically, the extracted food border features are analyzed using a clustering algorithm (such as the K-means clustering algorithm) to find common patterns of food borders that match the size of the baking tray inside the oven, thereby obtaining a set of prior anchor frames for initializing model detection.

[0160] For example, a K-means algorithm is used to cluster the width and height of the food borders to obtain multiple cluster centers, i.e., the size parameters of multiple prior anchor boxes. This ensures that the number and shape of the prior anchor boxes can cover most of the food in the oven to improve the accuracy of food recognition.

[0161] Step S1030: Generate an anchor frame for food identification based on multiple prior anchor frames.

[0162] Specifically, based on the multiple prior anchor frames obtained, a set of anchor frames specifically for food identification is generated as the initial detection frames for model training, helping the model to initialize the detection process more accurately.

[0163] For example, the size and position information of these prior anchor boxes are set in the model training configuration to ensure that the model can refer to these preset anchor boxes when learning the food bounding boxes, thereby improving the initialization accuracy and efficiency of the detection.

[0164] Step S1040 , adjusting the shape of the food recognition anchor frame according to the image of the food placed on the shelf of the oven so that the shape of the food recognition anchor frame matches the shape of the food image.

[0165] Specifically, based on the actual food pictures taken, the shape and size of the food identification anchor frame are dynamically adjusted to ensure that the anchor frame can more closely match the actual shape of the food, thereby improving the accuracy of the detection frame.

[0166] For example, during model training, the loss function and backpropagation algorithm are used to dynamically adjust the size of the anchor boxes so that they more closely match the actual boundaries of the ingredients. This can be achieved by modifying the YOLO model's loss function, ensuring that the model can optimize the shape of the anchor boxes to accommodate the diversity of ingredients.

[0167] Step S1050: Use the shape-adjusted food recognition anchor frame to recognize the image of the food on the shelf to obtain at least one detection frame.

[0168] Each detection frame includes one food ingredient.

[0169] Specifically, the adjusted food identification anchor frame is used to identify the food on the oven shelf to obtain one or more detection frames containing different food ingredients, each detection frame corresponding to one food ingredient.

[0170] For example, a picture of food taken inside an oven is input into the trained model, and the model outputs a series of detection frames and their corresponding food categories.

[0171] Step S1060: Determine at least one target detection frame based on the overlap of at least one detection frame and the position of the center point of at least one detection frame.

[0172] Specifically, based on the multiple detection frames output by the model, by analyzing the overlap and center point position of the detection frames, the most accurate detection frame is selected as the target detection frame to ensure that the food information contained in each detection frame is the most accurate.

[0173] For example, based on the intersection over union (IoU) and center point position of the detection frames, some overlapping detection frames are removed, some overlapping targets are retained, and the detection frame with an accurate center point position and including the retained overlapping targets is used as the target detection frame.

[0174] Step S1070: outputting food identification results based on the detection results of at least one target detection frame.

[0175] Among them, the food identification results include the label of the food category, the food border, and the confidence level.

[0176] Specifically, based on the object detection frame information, the final ingredient recognition results are generated, including the ingredient category label, the ingredient bounding box, and the model's confidence in the ingredient classification. Based on the object detection frame output by the model, the ingredient category label and confidence information are extracted, and the coordinates of the detection frame are recorded.

[0177] In this embodiment, first, by collecting and organizing the food training set, the data foundation required for model training is ensured, and the model can learn the skills of identifying various ingredients in the internal environment of the oven. Then, using the pre-trained model and clustering algorithm, a set of prior anchor frames that match the size of the oven baking tray are obtained. These anchor frames provide the model with initial shape assumptions for food detection, simplifying the detection difficulty and improving detection efficiency and accuracy. Then, by dynamically adjusting the shape of the anchor frame, the model can more accurately detect based on the true shape of the food, no longer restricted by the fixed shape anchor frame. Finally, the most accurate food detection frame is screened out through the non-maximum suppression algorithm, and the recognition results including food category, border and confidence are output. This provides accurate data support for the oven to dynamically adjust the cooking parameters, realizing intelligent identification and precise positioning of food.

[0178] In one embodiment, Figure 11 As shown, the method further includes: steps S1100-S1110, wherein:

[0179] Step S1100 , according to the type information of the food and the baking requirements input by the target object, the baking parameters of the food are searched in a preset knowledge base.

[0180] Among them, baking parameters include baking temperature field and baking time.

[0181] Specifically, based on the identified food type information and the baking requirements input by the user, the smart oven searches for baking parameters that match the food type and requirements from a preset knowledge base, including the temperature field distribution and baking time required for baking.

[0182] For example, a knowledge base containing various common ingredients and their optimal baking parameters (including temperature distribution and baking time) is established. This knowledge base can be stored in the oven's memory or in the cloud in the form of a database or file. The ingredient type and user requirements are used as query parameters to retrieve the corresponding baking parameters from the knowledge base.

[0183] Step S1110: adjusting the temperature field at the spatial position according to the spatial position of the food in the oven so that the temperature field at the spatial position is consistent with the baking temperature field.

[0184] Specifically, based on the spatial position information of the ingredients inside the oven, the heat source distribution inside the oven is intelligently adjusted, that is, the power of the heating elements at specific positions is adjusted to ensure that the temperature field at the spatial position of the ingredients is consistent with the baking temperature field obtained by the query.

[0185] For example, the power of the oven's heating elements can be adjusted based on the food's location. A proportional-integral-derivative (PID) control algorithm can be used to dynamically adjust the heating element power to achieve the target temperature. Temperature sensors can be installed at multiple locations within the oven to monitor the temperature distribution in real time. Combining sensor data with heating element power adjustments creates a closed-loop control loop, ensuring the target temperature is achieved at the food's location.

[0186] In this embodiment, the smart oven first retrieves precise baking parameters, including specific temperature distribution and baking time, from a pre-set knowledge base based on the type of ingredients and user baking requirements. This knowledge base query mechanism, powered by big data and user needs, ensures optimal cooking conditions for the ingredients during baking, improving both cooking results and the taste of the ingredients. Next, the oven intelligently adjusts the power of the heating elements based on the specific spatial location of the ingredients to create a target temperature field, ensuring even heating and ideal cooking conditions for the ingredients regardless of their location. The core of this technical solution lies in dynamic temperature field control and intelligent power allocation of the heating elements. This system combines ingredient recognition results with pre-set baking parameters to automatically adjust the heat source distribution within the oven to adapt to the spatial location of the ingredients, achieving targeted heating. For example, if an ingredient is detected on the top layer of the oven, the oven will increase the power of the top layer's heating elements while reducing the power of the lower layers' heating elements to ensure rapid and even heating of the top layer's ingredients without overheating the lower layers. This targeted heating method not only reduces uneven baking caused by varying location, but also shortens cooking time, saves energy, and enhances the personalized and intelligent level of cooking, providing a more personalized and intelligent user experience. In addition, through a closed-loop control mechanism and real-time monitoring data from temperature sensors, the oven can precisely control the power of each heating element, ensuring that the temperature field where the ingredients are located is consistent with the preset baking temperature field. Even if external conditions change during the cooking process, such as frequent opening and closing of the oven door, the heat source distribution can be quickly adjusted to maintain a stable temperature field, providing users with stable and high-quality baking results.

[0187] In one embodiment, Figure 12 As shown, step S1110, according to the spatial position of the food in the oven, adjusts the temperature field at the spatial position so that the temperature field at the spatial position is consistent with the baking temperature field. It includes: steps S1200-S1230, wherein:

[0188] Step S1200: continuously monitor the actual temperature field at the spatial position.

[0189] Specifically, during the entire baking process in the oven, the intelligent system needs to continuously monitor the actual temperature distribution (temperature field) of the spatial location of the ingredients to ensure dynamic adjustment of heating parameters and maintain a stable cooking environment.

[0190] Step S1210: determining the target temperature field at the spatial position at the current moment according to the baking parameters.

[0191] Specifically, based on the baking parameters obtained from the query, the intelligent system needs to determine the ideal temperature field distribution at the spatial location of the ingredients at the current time point, that is, the target temperature field, as a reference standard for temperature field adjustment.

[0192] For example, the intelligent control unit analyzes baking parameters obtained from a knowledge base and calculates the target temperature distribution for the ingredients at the current baking stage based on the ingredient type and baking time. A predictive model can be used to dynamically calculate the target temperature distribution based on the ingredient's baking progress and the current baking stage to accommodate different cooking requirements.

[0193] Step S1220: determining a temperature field deviation based on the actual temperature field and the target temperature field at the spatial position at the current moment.

[0194] Specifically, the actual temperature field is compared with the target temperature field, and the gap (deviation) between the two is calculated to provide a basis for adjusting the heating parameters.

[0195] Step S1230: Adjust the heating parameters of the oven according to the temperature field deviation so that the actual temperature field reaches the target temperature field.

[0196] Specifically, based on the temperature field deviation, the smart oven adjusts the power or distribution of the heating elements to optimize the temperature field at the spatial location of the food so that it is close to the target temperature field.

[0197] For example, a proportional-integral-derivative (PID) controller can be used to adjust the power of the heating element in real time based on temperature field deviations, quickly responding to temperature changes and maintaining the target temperature field. By calling the control interface of the heating element, the intelligent control unit adjusts the operating state of the heating element (such as power, on or off) in real time to achieve the target temperature field.

[0198] In this embodiment, temperature sensors inside the oven continuously monitor the actual temperature field in the food space, ensuring real-time data feedback and providing a foundation for precise temperature field control. Next, the intelligent system calculates the target temperature field for the current baking stage based on the type of food and baking parameters. This serves as the goal and direction for temperature field adjustment. Subsequently, by comparing the actual temperature field with the target temperature field, the temperature field deviation is calculated, providing a quantitative basis for adjusting the heating parameters and ensuring accuracy. Finally, based on the temperature field deviation, the intelligent oven adjusts the power and distribution of the heating elements in real time to optimize the temperature field in the food space, aligning it with the target temperature field. The key to this technical solution lies in its closed-loop control mechanism. This mechanism, through continuous monitoring, deviation calculation, and dynamic adjustment of heating parameters, forms a closed-loop feedback control system that ensures uniform heating of food in all oven levels, achieving the ideal cooking state. Dynamically adjusting the heating parameters based on the spatial position of the food inside the oven ensures that the temperature field in the food area is consistent with the preset baking temperature field, significantly improving cooking uniformity and efficiency.

[0199] In one embodiment, in step S230, the spatial position of the food in the oven is determined using the first position information and the second position information. Thereafter, the method further includes:

[0200] According to the spatial position of the food in the oven, the surface image of the food at the spatial position and the thermal imaging data of the food are obtained in real time.

[0201] Specifically, based on the specific spatial position of the food in the oven, the surface image and thermal imaging data of the food can be captured accurately and in real time to obtain the appearance characteristics and temperature distribution information of the food.

[0202] For example, inside the oven, high-resolution cameras and infrared thermal imagers are installed to ensure that food images and temperature information can be fully captured from all angles.

[0203] Analyze the surface image to determine the appearance information of the food, and analyze the thermal imaging data to determine the surface temperature distribution of the food.

[0204] Specifically, the smart oven system uses multimodal perception and deep learning algorithms to analyze surface images of ingredients, identifying their type and appearance. Simultaneously, it analyzes thermal imaging data to determine the temperature distribution on the ingredient surface, enabling a comprehensive assessment of the cooking progress.

[0205] For example, a food identification model can be built, using a convolutional neural network to classify images to determine food type and appearance information, such as color and texture. Separately, a thermal distribution recognition model can be created, using a convolutional neural network model with a similar architecture based on thermal imaging data to determine the temperature distribution on the food surface.

[0206] During model training, a large amount of food surface images and thermal imaging data, including food at different stages of cooking, is used to improve the model's generalization capabilities. The model should be able to adapt to varying lighting conditions and oven environment variations to ensure accurate recognition in real-world cooking scenarios. The food surface image output is combined with the results of the thermal distribution recognition model for analysis to determine the food's real-time cooking status, such as whether it is golden and crispy, and whether it is cooked or overcooked.

[0207] Determine the real-time baking status of the ingredients based on their appearance information and surface temperature distribution.

[0208] Specifically, by combining the appearance and temperature distribution of the ingredients, the system can infer the real-time baking status of the ingredients. This can be done through color analysis: observing the color changes on the surface of the ingredients. For example, a golden yellow surface on bread usually indicates that baking is nearly complete. It can also be done through expansion analysis: monitoring the expansion of the ingredients. For example, when a cake is ripe, it will expand to a certain extent and then fall back slightly. It can also be done through surface texture analysis: checking for cracks or other abnormal textures on the surface of the ingredients, which may indicate overbaking.

[0209] According to the real-time baking status of the ingredients, the baking parameters of the oven are dynamically adjusted to ensure that the ingredients achieve the best baking effect.

[0210] Specifically, based on the real-time baking status of the ingredients, the smart oven system can adjust baking parameters such as temperature, time, and hot air circulation speed in real time to ensure the best baking results. It can also monitor the status of ingredients in real time, detect abnormalities (such as burning or overcooking), and issue timely warning signals.

[0211] Specifically, the method may also include monitoring the real-time weight changes of the ingredients in the oven through a weight sensor provided on the baking tray, determining the moisture loss of the ingredients based on the real-time weight changes, and further judging the baking state of the ingredients based on the moisture loss of the ingredients.

[0212] Specifically, the weight sensor inside the baking tray monitors the real-time weight changes of the ingredients during the cooking process, thereby judging the degree of moisture loss of the ingredients and assisting in evaluating the baking status of the ingredients.

[0213] In this embodiment, an intelligent oven system based on multimodal perception is provided, which can significantly improve the accuracy of food identification and the intelligence level of the baking process. First, through the combination of images and thermal imaging, the system can obtain the appearance and temperature distribution information of the food, and accurately judge the type and baking status of the food even in complex scenarios such as similar shapes of food or changes in lighting. Secondly, the addition of weight sensors provides real-time monitoring of weight changes of food during baking, which makes up for the limitations of visual perception and enhances the multi-dimensional perception capabilities of the system. Thus, the system can integrate various sensory information, accurately judge the degree of baking of food, and realize intelligent food status recognition. Finally, the ability to dynamically adjust baking parameters enables the oven to dynamically optimize the heating strategy according to the real-time status of the food, ensuring that the food is evenly heated to achieve the ideal cooking effect.

[0214] Through the description of the above embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is the more preferred embodiment. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0215] Figure 13 FIG is a structural block diagram of a device for determining the location of food based on image recognition according to an embodiment of the present invention. Figure 13 As shown, including:

[0216] The image acquisition module 1301 is used to acquire an image of the food placed on the shelf of the oven, and perform perspective segmentation processing on the image based on the planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the planar perspective of the food, and the second image includes an image of the cross-sectional perspective of the shelf of the oven.

[0217] The first position determination module 1302 is configured to perform edge detection on the first image to obtain the boundary position of the plane where the food is located, and determine first position information of the food in the plane based on the relative position between the boundary position and the food.

[0218] The second position determination module 1303 is used to input the second image into the first image recognition model and the second image recognition model respectively, and determine the second position information of the baking tray where the ingredients are placed according to the relative relationship between the first recognition result and the second recognition result output by the first image recognition model and the second image recognition model respectively, wherein the first image recognition model is used to identify the baking tray features in the second image and enhance the baking tray features, and the second image recognition model is pre-trained using a baking tray training set, and the baking tray training set includes baking tray images obscured by ingredients, so that the second image recognition model can predict the baking tray outline when the baking tray in the image is obscured.

[0219] The food positioning module 1304 is configured to determine the spatial position of the food in the oven using the first position information and the second position information.

[0220] In an exemplary embodiment, the first position determination module 1302 is further configured to perform edge detection on the first image to determine at least three reference points at the boundary position of the plane where the food is located, wherein the at least three reference points are non-collinearly distributed. The first reference point of the at least three reference points is used as the origin, and the coordinate axis is determined based on the first vector pointing from the first reference point to the second reference point and the second vector pointing from the first reference point to the third reference point. A first coordinate system for the plane where the food is located is established based on the origin and the coordinate axis. The center of mass position of the food is determined based on the outline image of the food, and the first position information of the food is indicated by the coordinates of the center of mass position in the first coordinate system, wherein the outline image of the food is determined based on the first image.

[0221] In an exemplary embodiment, the first position determination module 1302 is further configured to perform corner detection on the first image to determine the coordinates of first corner points of at least three corner points of the baking tray carrying the ingredients, wherein the at least three corner points serve as at least three reference points. A perspective transformation matrix is ​​determined based on a first perspective of the plane in which the ingredients are located, used to acquire the first image, and a second perspective perpendicular to the plane in which the baking tray is located, wherein the perspective transformation matrix is ​​used to indicate the conversion relationship between the coordinates of the target point under the first perspective and the coordinates of the target point under the second perspective. The perspective transformation matrix is ​​used to convert the first corner point coordinates of the at least three corner points into second corner point coordinates, wherein the second corner point coordinates are the coordinates under the second perspective. The second corner point coordinates of the at least three corner points are used as the coordinates of the at least three reference points to determine the at least three reference points.

[0222] In one exemplary embodiment, first position determination module 1302 is further configured to perform boundary smoothing on the contour image of the food. Based on the positions and grayscale values ​​of the pixels in the contour image after boundary smoothing, the zero-order moment and first-order moment of the contour image are determined. The zero-order moment indicates the area of ​​the contour image, and the first-order moment indicates the distribution of the pixels in the contour image. The center of mass of the food is determined based on the zero-order moment and first-order moment of the contour image.

[0223] In an exemplary embodiment, the image acquisition module 1301 is further used to acquire an image of a calibration plate placed on a preset calibration shelf of the oven, wherein a plurality of calibration points are provided on the calibration plate. The first world coordinates of the plurality of calibration points are determined based on the image of the calibration plate. The coordinate error is determined based on the determined first world coordinates of the plurality of calibration points and the actual second world coordinates of the plurality of calibration points. When the coordinate error is greater than an error threshold, the parameters of the camera used to capture the image are adjusted according to the coordinate error, and the mapping relationship between the first world coordinate and the second world coordinate is adjusted, and the step of determining the first world coordinates of the plurality of calibration points based on the image of the calibration plate is returned to execution until the coordinate error is less than or equal to the error threshold. When the coordinate error is less than or equal to the error threshold, an image of the food placed on the shelf of the oven is acquired.

[0224] In an exemplary embodiment, the second position determination module 1303 is also used to input the second image into the first image recognition model to obtain a first recognition result output by the first image recognition model. And, input the second image into the second image recognition model to obtain a second recognition result output by the second image recognition model. In the case that the first recognition result and the second recognition result are inconsistent, a first weighted confidence is determined based on the first confidence carried by the first recognition result and a preset first weight corresponding to the first image recognition model. A second weighted confidence is determined based on the second confidence carried by the second recognition result and a preset second weight corresponding to the second image recognition model. Based on the relative relationship between the first weighted confidence and the second weighted confidence, a target result is determined, wherein the target result is one of the first recognition result and the second recognition result. The second position information is determined based on the target shelf indicated by the target result.

[0225] In one exemplary embodiment, the second position determination module 1303 is further configured to input the second image into the first image recognition model and extract pan features from the second image using multiple convolutional layers of the first image recognition model, wherein the multiple convolutional layers of the first image recognition model are skip-connected, and the pan features include pan edge features and pan seam features. Based on the pan features output by the output channels of the multiple convolutional layers, channel attention weights corresponding to the output channels of the multiple convolutional layers are determined. The channel attention weights corresponding to the output channels of each convolutional layer are determined by learning the dependencies between the output channels of the multiple convolutional layers, and the channel attention weights are related to the influence of the output channel on the first recognition result. Based on the pan features output by the output channels of the multiple convolutional layers, key regions in the second image are determined, wherein the key regions include pan edge regions and pan seam regions. Spatial attention weights corresponding to the key regions in the second image are enhanced. The weights of corresponding model parameters in the first image recognition model are adjusted based on the channel attention weights and the spatial attention weights. The second image is recognized using the first image recognition model after the weights of the model parameters are adjusted to obtain a first recognition result, wherein the first recognition result includes the shelf on which the baking tray is placed and a corresponding first confidence level.

[0226] In an exemplary embodiment, second position determination module 1303 is further configured to obtain a baking tray training set, wherein the baking tray training set includes an image of a first baking tray placed in an oven without obstructions, and an image of a second baking tray placed in the oven with obstructions caused by food. The second baking tray image in the baking tray training set is delineated and annotated to outline the obstructed baking tray in the second baking tray image. An initial image recognition model is trained using the first baking tray image and the delineated and annotated second baking tray image to obtain a second image recognition model. The second image is recognized using the second image recognition model to obtain a second recognition result, wherein the second recognition result includes the shelf on which the baking tray is placed and a corresponding second confidence level.

[0227] In an exemplary embodiment, the apparatus further includes: a training set acquisition module, a feature extraction module, a clustering module, an anchor frame generation module, an anchor frame adjustment module, an identification module, a detection frame determination module, and an ingredient identification module, wherein:

[0228] The training set acquisition module is used to acquire an ingredient training set, wherein the ingredient training set includes pictures of a variety of different ingredients located inside the oven.

[0229] The feature extraction module is used to process the food training set using a pre-trained image recognition network and extract the border features of the food in each image.

[0230] The clustering module is used to cluster the border features of the ingredients in each picture of the ingredient training set using a clustering algorithm to obtain multiple prior anchor frames corresponding to the preset baking tray size.

[0231] The anchor frame generation module is used to generate an anchor frame for food identification based on multiple prior anchor frames.

[0232] The anchor frame adjustment module is used to adjust the shape of the food identification anchor frame according to the image of the food placed on the shelf of the oven so that the food identification anchor frame matches the shape of the food image.

[0233] The recognition module is used to use the shape-adjusted food recognition anchor frame to recognize the image of the food on the shelf to obtain at least one detection frame, wherein each detection frame includes one food.

[0234] The detection frame determination module is used to determine at least one target detection frame based on the overlap of at least one detection frame and the position of the center point of at least one detection frame.

[0235] The food identification module is used to output food identification results based on the detection results of at least one target detection frame, wherein the food identification results include the label of the category to which the food belongs, the food border, and the confidence level.

[0236] In an exemplary embodiment, the device further includes: a parameter determination module and a temperature field adjustment module, wherein:

[0237] The parameter determination module is used to query and obtain the baking parameters of the ingredients in a preset knowledge base according to the type information of the ingredients and the baking requirements input by the target object, wherein the baking parameters include baking temperature field and baking time.

[0238] The temperature field adjustment module is used to adjust the temperature field at the spatial position according to the spatial position of the food in the oven so that the temperature field at the spatial position is consistent with the baking temperature field.

[0239] In one exemplary embodiment, the temperature field adjustment module is further configured to continuously monitor the actual temperature field at a spatial location. Based on the baking parameters, the module determines a target temperature field at the current spatial location. Based on the actual temperature field at the current spatial location and the target temperature field, the module determines a temperature field deviation. Based on the temperature field deviation, the module adjusts the oven's heating parameters to bring the actual temperature field to the target temperature field.

[0240] An embodiment of the present invention further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.

[0241] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0242] S1, obtain an image of food placed on a shelf of an oven, and perform perspective segmentation processing on the image based on the planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the planar perspective of the food, and the second image includes an image of the cross-sectional perspective of the shelf of the oven.

[0243] S2, performing edge detection on the first image to obtain a boundary position of the plane where the food is located, and determining first position information of the food in the plane based on the relative position between the boundary position and the food.

[0244] S3, inputting the second image into the first image recognition model and the second image recognition model respectively, and determining the second position information of the baking tray where the ingredients are placed according to the relative relationship between the first recognition result and the second recognition result respectively output by the first image recognition model and the second image recognition model, wherein the first image recognition model is used to recognize the baking tray features in the second image and enhance the baking tray features, and the second image recognition model is pre-trained using a baking tray training set, and the baking tray training set includes baking tray images obscured by ingredients, so that the second image recognition model can predict the baking tray contour when the baking tray in the image is obscured.

[0245] S4: Determine the spatial position of the food in the oven using the first position information and the second position information.

[0246] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0247] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0248] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0249] S1, obtain an image of food placed on a shelf of an oven, and perform perspective segmentation processing on the image based on the planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the planar perspective of the food, and the second image includes an image of the cross-sectional perspective of the shelf of the oven.

[0250] S2, performing edge detection on the first image to obtain a boundary position of the plane where the food is located, and determining first position information of the food in the plane based on the relative position between the boundary position and the food.

[0251] S3, inputting the second image into the first image recognition model and the second image recognition model respectively, and determining the second position information of the baking tray where the ingredients are placed according to the relative relationship between the first recognition result and the second recognition result respectively output by the first image recognition model and the second image recognition model, wherein the first image recognition model is used to recognize the baking tray features in the second image and enhance the baking tray features, and the second image recognition model is pre-trained using a baking tray training set, and the baking tray training set includes baking tray images obscured by ingredients, so that the second image recognition model can predict the baking tray contour when the baking tray in the image is obscured.

[0252] S4: Determine the spatial position of the food in the oven using the first position information and the second position information.

[0253] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0254] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0255] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0256] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for determining food location based on image recognition, characterized in that: include: Acquire an image of food placed on a shelf of an oven, and perform perspective segmentation processing on the image based on planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the food at a planar perspective, and the second image includes an image of the shelf of the oven at a cross-sectional perspective; Performing edge detection on the first image to obtain a boundary position of the plane where the food is located, and determining first position information of the food in the plane based on a relative position between the boundary position and the food; inputting the second image into a first image recognition model and a second image recognition model, respectively, and determining second position information of the baking pan on which the ingredients are placed based on a relative relationship between first and second recognition results output by the first and second image recognition models, wherein the first image recognition model is used to identify features of the baking pan in the second image and perform enhancement processing on the baking pan features, and the second image recognition model is pre-trained using a baking pan training set, wherein the baking pan training set includes images of baking pans occluded by ingredients, so that the second image recognition model can predict the outline of the baking pan when the baking pan in the image is occluded; The spatial position of the food in the oven is determined by the first position information and the second position information.

2. The method for determining food location based on image recognition according to claim 1, characterized in that: The step of performing edge detection on the first image to obtain a boundary position of the plane where the food is located, and determining first position information of the food based on a relative position between the boundary position and the food, includes: Performing edge detection on the first image to determine at least three reference points at boundary positions of a plane where the food is located, wherein the at least three reference points are non-collinearly distributed; Taking a first reference point of the at least three reference points as an origin, and determining a coordinate axis according to a first vector pointing from the first reference point to a second reference point and a second vector pointing from the first reference point to a third reference point; Establishing a first coordinate system on the plane where the food is located according to the origin and the coordinate axes; The center of mass position of the food is determined according to the outline image of the food, and the first position information of the food is indicated by the coordinates of the center of mass position in the first coordinate system, wherein the outline image of the food is determined according to the first image.

3. The method for determining food location based on image recognition according to claim 2, characterized in that: The performing edge detection on the first image to determine at least three reference points at boundary positions of the plane where the food is located includes: Performing corner point detection on the first image to determine first corner point coordinates of at least three corner points of the baking tray carrying the food, wherein the at least three corner points serve as the at least three reference points; determining a perspective transformation matrix based on a first perspective of a plane on which the food is located and a second perspective perpendicular to the plane on which the baking tray is located, wherein the perspective transformation matrix indicates a conversion relationship between coordinates of a target point at the first perspective and coordinates of the target point at the second perspective; Using the perspective transformation matrix, converting first corner point coordinates of the at least three corner points into second corner point coordinates, wherein the second corner point coordinates are coordinates at the second viewing angle; The coordinates of the second corner points of the at least three corner points are used as the coordinates of the at least three reference points to determine the at least three reference points.

4. The method for determining food location based on image recognition according to claim 2, characterized in that: Determining the center of mass position of the food according to the contour image of the food includes: Performing boundary smoothing processing on the contour image of the food; Determining a zero-order moment and a first-order moment of the contour image according to positions and grayscale values ​​of pixels included in the contour image after boundary smoothing, wherein the zero-order moment is used to indicate the area of ​​the contour image, and the first-order moment is used to indicate the distribution of pixels in the contour image; The center of mass position of the food is determined according to the zero-order moment and the first-order moment of the contour image.

5. The method for determining food location based on image recognition according to any one of claims 1 to 4, characterized in that: The acquiring of an image of food placed on a shelf of an oven includes: Acquire an image of a calibration plate placed on a preset calibration shelf of the oven, wherein a plurality of calibration points are set on the calibration plate; Determining first world coordinates of the plurality of calibration points according to the image of the calibration plate; Determining coordinate errors based on the determined first world coordinates of the plurality of calibration points and the actual second world coordinates of the plurality of calibration points; When the coordinate error is greater than the error threshold, adjusting parameters of a camera used to acquire the image and a mapping relationship between the first world coordinate and the second world coordinate according to the coordinate error, and returning to the step of determining the first world coordinates of the plurality of calibration points according to the image of the calibration plate until the coordinate error is less than or equal to the error threshold; When the coordinate error is less than or equal to the error threshold, an image of the food placed on the shelf of the oven is acquired.

6. The method for determining food location based on image recognition according to any one of claims 1 to 4, characterized in that: Inputting the second image into the first image recognition model and the second image recognition model respectively, and determining the second position information of the baking tray where the ingredients are placed according to the relative relationship between the first recognition result and the second recognition result respectively output by the first image recognition model and the second image recognition model, includes: Inputting the second image into the first image recognition model to obtain a first recognition result output by the first image recognition model; and inputting the second image into the second image recognition model to obtain a second recognition result output by the second image recognition model; In a case where the first recognition result and the second recognition result are inconsistent, determining a first weighted confidence level according to a first confidence level carried by the first recognition result and a first weight preset corresponding to the first image recognition model; and determining a second weighted confidence level according to the second confidence level carried by the second recognition result and a preset second weight corresponding to the second image recognition model; Determining a target result according to a relative relationship between the first weighted confidence and the second weighted confidence, wherein the target result is one of the first recognition result and the second recognition result; The second position information is determined according to the target shelf indicated by the target result.

7. The method for determining food location based on image recognition according to claim 6, characterized in that: Inputting the second image into the first image recognition model to obtain a first recognition result output by the first image recognition model includes: Inputting the second image into the first image recognition model, and extracting baking pan features of the second image through multiple convolutional layers of the first image recognition model, wherein the multiple convolutional layers of the first image recognition model are skip-connected, and the baking pan features include baking pan edge features and baking pan seam features; Determining, based on the baking pan features respectively output by the output channels of the multiple convolutional layers, channel attention weights corresponding to the output channels of the multiple convolutional layers, respectively, wherein the channel attention weight corresponding to the output channel of each convolutional layer is determined by learning a dependency relationship between the output channels of the multiple convolutional layers, and the channel attention weight is related to the degree of influence of the output channel on the first recognition result; Determining a key area in the second image based on the baking pan features respectively output by the output channels of the multiple convolutional layers, wherein the key area includes an edge area of ​​the baking pan and a seam area of ​​the baking pan; Enhancing the spatial attention weight corresponding to the key area in the second image; Adjusting the weights of corresponding model parameters in the first image recognition model according to the channel attention weight and the spatial attention weight; The second image is recognized by using the first image recognition model after the weight adjustment of the model parameters to obtain the first recognition result, wherein the first recognition result includes the shelf on which the baking tray is placed and the corresponding first confidence level.

8. The method for determining food location based on image recognition according to claim 6, characterized in that: Inputting the second image into the second image recognition model to obtain a second recognition result output by the second image recognition model includes: Obtaining a baking pan training set, wherein the baking pan training set includes an unobstructed first baking pan image placed in the oven and an unobstructed second baking pan image placed in the oven and obstructed by food; performing a delineation and annotation process on the second baking pan image in the baking pan training set to outline the outline of the baking pan that is blocked in the second baking pan image; Using the first baking pan image and the second baking pan image after depiction and annotation processing to train an initial image recognition model, to obtain a second image recognition model; The second image is recognized using the second image recognition model to obtain a second recognition result, wherein the second recognition result includes the shelf on which the baking tray is placed and a corresponding second confidence level.

9. The method for determining food location based on image recognition according to any one of claims 1 to 4, characterized in that: The method further comprises: Obtaining an ingredient training set, wherein the ingredient training set includes multiple images of different ingredients located inside an oven; Using a pre-trained image recognition network to process the food training set, extracting the border features of the food in each picture; Using a clustering algorithm to cluster the bounding box features of the food in each picture of the food training set, and obtain multiple prior anchor frames corresponding to the preset baking pan size; generating an ingredient recognition anchor frame based on the multiple prior anchor frames; Adjusting the shape of the food identification anchor frame according to the image of the food placed on the shelf of the oven so that the food identification anchor frame matches the shape of the food image; Using the shape-adjusted food identification anchor frame to identify the image of the food on the shelf, to obtain at least one detection frame, wherein each detection frame includes one food; Determining at least one target detection frame according to an overlap condition of the at least one detection frame and a position of a center point of the at least one detection frame; Outputting an ingredient recognition result according to the detection result of the at least one target detection frame, wherein the ingredient recognition result includes a label of the category to which the ingredient belongs, an ingredient border, and a confidence level.

10. The method for determining food location based on image recognition according to claim 9, characterized in that: The method further comprises: According to the type information of the food and the baking requirements input by the target object, the baking parameters of the food are queried in a preset knowledge base, wherein the baking parameters include baking temperature and baking time; According to the spatial position of the food in the oven, the temperature field at the spatial position is adjusted so that the temperature field at the spatial position is consistent with the baking temperature field.

11. The method for determining food location based on image recognition according to claim 10, characterized in that: The step of adjusting the temperature field at the spatial position according to the spatial position of the food in the oven so that the temperature field at the spatial position is consistent with the baking temperature field includes: Continuously monitoring the actual temperature field at the spatial position; determining a target temperature field at the spatial position at a current moment according to the baking parameters; determining a temperature field deviation according to the actual temperature field at the spatial position at a current moment and the target temperature field; The heating parameters of the oven are adjusted according to the temperature field deviation so that the actual temperature field reaches the target temperature field.

12. A device for determining the location of food based on image recognition, characterized in that: include: an image acquisition module, configured to acquire an image of food placed on a shelf of an oven, and perform perspective segmentation processing on the image based on the planar perspective features and cross-sectional perspective features of the image to obtain a first image and a second image, wherein the first image includes an image of the food from a planar perspective, and the second image includes an image of the shelf from a cross-sectional perspective of the oven; a first position determination module, configured to perform edge detection on the first image to obtain a boundary position of the plane where the food is located, and determine first position information of the food in the plane based on a relative position between the boundary position and the food; a second position determination module, configured to input the second image into a first image recognition model and a second image recognition model, respectively, and determine second position information of the baking pan on which the ingredients are placed based on a relative relationship between first and second recognition results output by the first and second image recognition models, respectively, wherein the first image recognition model is configured to identify and enhance baking pan features in the second image, and the second image recognition model is pre-trained using a baking pan training set that includes images of baking pans occluded by ingredients, so that the second image recognition model can predict the outline of the baking pan when the baking pan in the image is occluded; The food positioning module is used to determine the spatial position of the food in the oven according to the first position information and the second position information.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 11 when executed.

14. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 11 through the computer program.

Citation Information

Cited By

  • Food material baking method based on image recognition, storage medium and electronic device

    CN121459343A

  • Food material data generation method and device based on image processing, storage medium and electronic device

    CN121482526A

  • Oven-based mode matching method and device, storage medium and electronic device

    CN121482771A