Object recognition method and device based on intelligent glasses, terminal, medium and product

Through user gesture selection combined with preset reference object calibration, smart glasses can accurately identify the target object, solve the problems of convenience and recognition accuracy, and provide intuitive health management data.

CN120673175APending Publication Date: 2025-09-19BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510874457.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, smart glasses lack convenience and accuracy in object recognition, making it difficult to accurately select target objects. They are easily affected by background interference and changes in shooting distance lead to inaccurate size estimation.

Method used

By using intuitive user gestures to select the target object and combining it with preset reference objects for size calibration, the smart glasses automatically identify the composition information of the target object, avoiding background interference and improving size recognition accuracy.

Benefits of technology

It improves the accuracy and convenience of object recognition, ensures the accuracy of ingredient information recognition, and provides intuitive health management data reference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673175A_ABST
    Figure CN120673175A_ABST
Patent Text Reader

Abstract

The invention relates to an object recognition method and device based on intelligent glasses, a terminal, a medium and a product. A target image can be acquired through an image acquisition unit of the intelligent glasses, wherein the target image comprises a gesture and at least one to-be-recognized object; determining a target object selected by the gesture from the at least one object to be recognized; and identifying component information of the target object. According to the embodiment of the invention, the target object of which the component information needs to be recognized is selected through the visual gesture of the user, and the intelligent glasses can automatically determine the target object from the at least one to-be-recognized object based on the gesture, thereby avoiding the interference of other objects, improving the component recognition accuracy, and improving the interaction convenience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a method, device, terminal, medium, and product for object recognition based on smart glasses. Background Art

[0002] With the development of artificial intelligence, computer vision, and the Internet of Things (IoT), smart wearable devices are becoming increasingly popular and play an important role in people’s daily lives. For example, computer vision technology can be used to identify the composition of target objects.

[0003] The object recognition methods provided in the related art need to be improved in terms of convenience and recognition accuracy. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides an object recognition method, device, terminal, medium and product based on smart glasses.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for object recognition based on smart glasses is provided, comprising: Acquire a target image through an image acquisition unit of the smart glasses, wherein the target image includes a gesture and at least one object to be recognized; determining a target object selected by the gesture from the at least one object to be identified; Identify composition information of the target object.

[0006] Using this method, the user can use intuitive gestures to select the target object whose component information needs to be identified. The smart glasses can automatically determine the target object from at least one object to be identified based on the gesture, avoiding interference from other objects, improving the accuracy of component identification, and also improving the user's interaction convenience.

[0007] In some possible implementations, determining the target object selected by the gesture from the at least one object to be identified includes: performing gesture detection on the target image to determine the pixel position of the key point corresponding to the gesture; and determining the target object based on the pixel position.

[0008] By adopting this implementation mode, gesture detection can be performed on the user's gestures to help the user quickly and accurately specify the target object for their intention analysis, avoiding interference from non-target objects, improving the accuracy of component recognition, and improving the convenience of interaction.

[0009] In some possible implementations, determining the target object according to the pixel position includes: determining a target image area selected by the gesture in the target image according to the pixel position; The object to be identified in the target image area is used as the target object.

[0010] By adopting this embodiment, the target image area where the target object is located can be located based on the gesture detection result, and the target object that meets the user's intention can be accurately identified, thereby improving the accuracy and efficiency of component recognition of the target object.

[0011] In some possible implementations, the target object includes at least one object; and identifying the component information of the target object includes: determining, for each object, the object type and object size of the object; and determining the component information of the object based on the object type and the object size.

[0012] With this implementation, the smart glasses can automatically determine at least one object that meets the user's intention based on the user's gesture, and perform component identification on each object separately, thereby improving the accuracy of component identification.

[0013] In some possible implementations, determining the object type and object size of each object includes: performing target detection on the target image area selected by the gesture in the target image to determine the object type corresponding to each object; and determining the object size of each object according to the object type of the object.

[0014] In some possible implementations, determining the object type and object size of each object includes: performing image recognition on the target image area selected by the gesture in the target image to obtain an image recognition result; performing target detection on the target image to determine whether the target image also includes a preset reference object; in response to the target image also including the preset reference object, determining the object type and the object size corresponding to each object according to the preset reference object and the image recognition result.

[0015] In some possible implementations, determining the object type and the object size corresponding to each object respectively based on the preset reference object and the image recognition result includes: determining the object type of each object and the image area of ​​each object in the target image based on the image recognition result; identifying the first pixel size of the preset reference object in the target image and the second pixel size of the image area corresponding to each object respectively; and determining the object size corresponding to each object respectively based on the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object respectively.

[0016] In some possible implementations, determining the object size corresponding to each object based on the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object includes: determining a proportional coefficient based on the actual size and the first pixel size, the proportional coefficient representing a ratio of the actual size of the object to the pixel size; and determining, for each object, the object size of the object based on the proportional coefficient and the second pixel size.

[0017] By adopting the above-mentioned embodiment, by introducing a preset reference object with a known actual size, the size of the target object is calibrated, so that the system can more accurately convert the actual size of the target object in the image, thereby improving the accuracy of component information recognition and solving the problem of inaccurate object size estimation of the target object caused by changes in shooting distance.

[0018] In some possible embodiments, the ingredient information includes the weight of each ingredient in at least one ingredient corresponding to the object; determining the ingredient information of the object based on the object type and the object size includes: determining the content of each ingredient in at least one ingredient corresponding to the object based on the object type; determining the portion of the object based on the object size; and determining the weight of each ingredient corresponding to the object based on the content of each ingredient in the at least one ingredient and the portion.

[0019] With this implementation, the user intuitively selects the target object whose composition information needs to be identified through a gesture. Based on this gesture, the smart glasses automatically identify the target object from at least one other object to be identified, eliminating interference from other objects and improving the accuracy of identifying the target object's type and size. By introducing a preset reference object of known actual size, the target object's size is calibrated, enabling more accurate identification of the target object's size. This accurate identification of object type and size further improves the accuracy of composition information identification.

[0020] In some possible implementations, the object to be identified includes food, and the ingredient information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

[0021] In this implementation, the object to be identified includes food. Users can intuitively lock onto the target image area of ​​the food they wish to analyze using gestures. The system then automatically crops the image, effectively eliminating interference from irrelevant food or other objects in the background and preventing misidentification. Furthermore, by introducing pre-set reference objects for size calibration, the system effectively overcomes the problem of inaccurate portion size estimation caused by varying shooting distances and angles, resulting in nutritional analysis results that are closer to actual values.

[0022] In some possible implementations, the method further includes: determining the calories of the target food based on the weight of each nutrient in the at least one nutrient.

[0023] By adopting this implementation, users can be provided with intuitive calorie data, providing an intuitive data reference for their health management.

[0024] In some possible implementations, the method further includes: obtaining user data, wherein the user data represents the nutritional needs of the user; and outputting nutritional recommendations based on the user data and the weight of each nutrient component corresponding to the target food, wherein the nutritional recommendations include a diet plan that meets the nutritional needs.

[0025] By adopting this implementation method, nutritional plans can be recommended to users, thereby providing services for users' health management and improving user experience.

[0026] In some possible implementations, the method further includes: outputting interaction information through the smart glasses; wherein the interaction information includes the ingredient information and / or the calories of the target food.

[0027] With this implementation, the interactive information is output to the user via the smart glasses, thereby improving the convenience for the user to view the interactive information.

[0028] According to a second aspect of an embodiment of the present disclosure, there is provided an object recognition device based on smart glasses, comprising: an acquisition module, configured to acquire a target image through an image acquisition unit of the smart glasses, wherein the target image includes a gesture and at least one object to be recognized; a determination module, configured to determine a target object selected by the gesture from the at least one object to be identified; The recognition module is configured to recognize the component information of the target object.

[0029] In some possible implementations, the determination module is configured to perform gesture detection on the target image to determine pixel positions of key points corresponding to the gesture; and determine the target object based on the pixel positions.

[0030] In some possible implementations, the target object includes at least one object; the recognition module is configured to determine the object type and object size of each object; and determine the component information of the object according to the object type and the object size.

[0031] In some possible embodiments, the recognition module is configured to perform image recognition on the target image area selected by the gesture in the target image to obtain an image recognition result; perform target detection on the target image to determine whether the target image also includes a preset reference object; in response to the target image also including the preset reference object, determine the object type and the object size corresponding to each object according to the preset reference object and the image recognition result.

[0032] In some possible embodiments, the recognition module is configured to determine the object type of each object and the image area of ​​each object in the target image based on the image recognition result; identify the first pixel size of the preset reference object in the target image and the second pixel size of the image area corresponding to each object; and determine the object size corresponding to each object based on the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object.

[0033] In some possible embodiments, the ingredient information includes the weight of each ingredient in at least one ingredient corresponding to the object; the identification module is configured to determine the content of each ingredient in at least one ingredient corresponding to the object based on the object type; determine the portion of the object based on the object size; and determine the weight of each ingredient corresponding to the object based on the weight of each ingredient in the at least one ingredient and the portion.

[0034] In some possible implementations, the object to be identified includes food, and the ingredient information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

[0035] In some possible implementations, the identification module is further configured to determine the calories of the target food based on the weight of each nutrient in the at least one nutrient.

[0036] According to a third aspect of an embodiment of the present disclosure, a terminal is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to execute the steps of the object recognition method based on smart glasses provided in the first aspect of the present disclosure.

[0037] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the object recognition method based on smart glasses provided in the first aspect of the present disclosure are implemented.

[0038] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the object recognition method based on smart glasses provided in the first aspect of the present disclosure.

[0039] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: through the user's intuitive gesture to select the target object whose component information needs to be identified, the smart glasses can automatically determine the target object from at least one object to be identified based on the gesture, without the user having to manually crop the image, to achieve analysis of only the target object that meets the user's intention, avoiding interference from other objects, improving the accuracy of component recognition, and also improving the convenience of user interaction.

[0040] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0042] Figure 1 The figure is a flowchart showing an object recognition method based on smart glasses according to an exemplary embodiment.

[0043] Figure 2 is based on Figure 1 The illustrated embodiment shows a flow chart of an object recognition method based on smart glasses.

[0044] Figure 3 is based on Figure 1 The illustrated embodiment shows a flow chart of an object recognition method based on smart glasses.

[0045] Figure 4 is based on Figure 3 The illustrated embodiment shows a flow chart of an object recognition method based on smart glasses.

[0046] Figure 5 is based on Figure 3 The illustrated embodiment shows a flow chart of an object recognition method based on smart glasses.

[0047] Figure 6 The figure is a flow chart showing a method for identifying food ingredients based on smart glasses according to an exemplary embodiment.

[0048] Figure 7 The figure is a block diagram showing an object recognition device based on smart glasses according to an exemplary embodiment.

[0049] Figure 8 is based on Figure 7 The illustrated embodiment shows a block diagram of an object recognition device based on smart glasses.

[0050] Figure 9 The figure is a block diagram of a terminal for object recognition according to an exemplary embodiment.

[0051] Figure 10 It is a block diagram showing a device for object recognition according to an exemplary embodiment. DETAILED DESCRIPTION

[0052] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0053] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0054] This disclosure is primarily applicable to scenarios where computer vision technology is used to analyze the composition of a target object. For example, the target object could be food, and computer vision technology can be used to identify the food and estimate its nutritional content (e.g., calories, carbohydrates, fat, protein, etc.). Another example is a material composed of multiple substances, and computer vision technology can be used to identify the material composition of the material.

[0055] Taking food nutritional analysis as an example, users expect to be able to record their daily diet conveniently and accurately to assist in health management and weight control. The solutions for food nutritional analysis provided in related technologies are as follows: Solution 1: A food calorie estimation application based on general image recognition. Users use the target app on their smartphone to take a photo of food and upload it to a cloud server. The AI ​​model on the server identifies the food in the image (e.g., "rice" or "chicken breast") and attempts to estimate the portion size based on the area in the image, ultimately returning nutritional information such as calories.

[0056] Solution 2: Some smart devices have built-in food recognition capabilities. Some high-end smart refrigerators or smart kitchen assistants may be equipped with cameras that can identify ingredients placed inside. However, these solutions primarily focus on food management, rather than real-time dining records, and are limited in scope. For portable devices (such as smart glasses), food recognition capabilities are primarily limited to identifying specific types and lack precise quantification capabilities.

[0057] The technical solutions provided in the related art for identifying food and estimating nutritional content by taking photos have the following problems: It's difficult to accurately select the target food, and it's easily affected by background interference. In complex dining scenarios, the camera may capture multiple foods, tableware, or other irrelevant objects simultaneously. This may include foods that the user didn't intend to record, or require the user to perform tedious manual selection on the captured image, affecting user interaction convenience and the accuracy of ingredient recognition.

[0058] Furthermore, most nutrient content recognition solutions offered in related technologies directly analyze the pixel size of food in an image to estimate portion size. However, this is sensitive to the shooting distance, leading to inaccurate volume estimation of the food being identified. For example, when a user captures food from different distances, the pixel area occupied by the same food in the image varies significantly, resulting in significant fluctuations in nutrient content estimation and low accuracy.

[0059] To solve the above problems, the present disclosure provides an object recognition method, device, terminal, medium and product based on smart glasses. The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0060] Figure 1 FIG. 1 is a flow chart showing an object recognition method based on smart glasses according to an exemplary embodiment. The method can be applied to smart glasses and can also be applied to a server communicating with the smart glasses. Figure 1 As shown, the method includes the following steps.

[0061] In step S11, a target image is acquired by an image acquisition unit of the smart glasses, where the target image includes a gesture and at least one object to be identified.

[0062] The object to be identified may include food, composite materials, and other objects that require component analysis.

[0063] In response to the user starting the component information recognition function, the smart glasses can activate the image acquisition unit on the smart glasses to capture the target image in the user's current field of view in real time.

[0064] To avoid including objects not intended by the user in object recognition and reduce background interference during object recognition, the present disclosure allows the user to intuitively indicate the target image area where the target object is located through gestures. Therefore, the target image captured in this step may include the user's gesture and at least one object to be recognized.

[0065] The gesture is a user gesture used to select a target object that meets the user's intent. This gesture can be performed with one hand (e.g., a semicircle formed by the thumb and index finger of one hand, or a finger pointing at an object), with two hands (e.g., a rectangle formed by the thumb and index finger of both hands, which is used to select the target object that meets the user's intent), or with both hands and the arm.

[0066] In some implementations, after determining that the image acquisition unit of the smart glasses has been activated, a prompt (in the form of voice and / or text) may be output to the user, prompting the user to use gestures to select the target object whose composition information is to be analyzed. For example, the prompt may be "Please use gestures to select the food you wish to analyze." In this way, the user can gesture to select the food according to the voice prompt, and the smart glasses will initiate a photo capture action after recognizing the user's gesture within the current field of view, thereby capturing the target image.

[0067] In another implementation, the smart glasses may initiate a photo-taking action after a preset time period after determining that the image acquisition unit is activated, thereby acquiring the target image. The preset time period may be greater than or equal to 0 seconds, for example.

[0068] For example, when a user wears smart glasses and wants to record information such as the calories of the food they consume while dining in a restaurant or at home, the user can activate the food recording function of the smart glasses. In response to the user activating the food recording function of the smart glasses, the smart glasses activate the camera module of the smart glasses and preview or capture the target image in the user's current field of view in real time. At the same time, the system outputs a voice prompt message "Please use gestures to select the food you want to analyze." In this way, the user can follow the voice prompt or directly perform gestures to select the food you want to analyze. The smart glasses capture the target image including the gesture and at least one food to be identified. This example is for illustrative purposes only and is not limited to this disclosure.

[0069] In step S12, a target object selected by the gesture is determined from the at least one object to be identified.

[0070] In this step, the target image region selected by the user's gesture can be determined based on the user's gesture, and the object to be identified within the target image region is then used as the target object. The target object is the object to be identified for component information recognition that meets the user's intention.

[0071] In step S13, component information of the target object is identified.

[0072] The component information may include the weight of at least one component corresponding to the target object. The weight of the at least one component corresponding to the target object may also vary depending on the size (including volume or weight) of the target object.

[0073] For example, if the target object is food, the ingredient information may include the nutritional composition of the food (such as calories, carbohydrates, fat, and protein), along with the weight of each nutrient. For example, if the target object is a 100-gram serving of white rice, the ingredient information corresponding to this 100-gram serving of white rice would be: protein: approximately 7.0 grams, fat: approximately 0.6 grams, carbohydrates: approximately 80.0 grams, sugar: approximately 0.1 grams, and fiber: approximately 1.0 grams. This is merely an example and is not intended to be limiting.

[0074] If the target object is a material composed of multiple substances, the material may be a metal material, for example, and the composition information may be the weight of each metal element constituting the metal material.

[0075] Using the above method, the user can use intuitive gestures to select the target object whose component information needs to be identified. The smart glasses can automatically determine the target object from at least one object to be identified based on the gesture. Without the user manually cropping the image, only the target object that meets the user's intention can be analyzed, avoiding interference from other objects. While improving the accuracy of component recognition, it also improves the convenience of user interaction.

[0076] Figure 2 is based on Figure 1 The embodiment shown is a flow chart of an object recognition method based on smart glasses, as shown in FIG. Figure 2 As shown, step S12 includes the following sub-steps: In step S121 , gesture detection is performed on the target image to determine the pixel positions of key points corresponding to the gesture.

[0077] The key points corresponding to the gesture may include, for example, one or more of the following positions: fingertip position, finger joint position, the connection position between the index finger and thumb, wrist position, etc. The pixel position refers to the position of the pixel corresponding to each key point in the target image.

[0078] During this step, the target image may first be detected for a valid gesture. This valid gesture may be a user-preset "frame selection" gesture (for example, the user's index fingers and thumbs forming an L-shape, forming a rectangular frame, or the user's hands and arms forming a closed frame). If the valid gesture is detected in the target image, gesture keypoints may be extracted to obtain the pixel locations of the keypoints corresponding to the gesture.

[0079] In one embodiment, a traditional image processing method may be used to detect whether the valid gesture exists in the target image. The traditional image processing method includes, for example, edge detection, contour extraction, and the like.

[0080] In another embodiment, a pre-trained gesture detection model can be used to perform feature extraction and target detection to identify whether there is a valid gesture in the target image, and further use a key point detection algorithm (such as OpenPose or other human key point detection algorithms) to accurately identify the key points of the gesture.

[0081] In step S122, the target object is determined according to the pixel position.

[0082] In this step, the target image area selected by the gesture in the target image can be determined according to the pixel position, and the object to be identified in the target image area is used as the target object.

[0083] In one implementation, the gesture includes forming an L-shape with the index fingers and thumbs of both hands, forming a rectangular frame relative to each other. Accordingly, the key points of the gesture include the fingertips of the index fingers and thumbs of both hands, and the connection between the index fingers and thumbs. Thus, a rectangular area with these four key points of both hands as vertices can be used as the target image area.

[0084] In another implementation, the gesture includes forming an L-shape with the index fingers and thumbs of both hands, forming a rectangular frame. Accordingly, the key points of the gesture can include the connection points between the index fingers and thumbs of both hands. Thus, the two key points of the hands can be used as the two vertices of a rectangular diagonal, and the rectangular area corresponding to the diagonal is used as the target image area.

[0085] In another implementation, the gesture includes a finger pointing at a target object. Accordingly, the key point of the gesture includes the fingertip position of the finger. In this case, the pixel at the fingertip position is typically located on the target object. Therefore, the object where the pixel at the fingertip position is located can be used as the target object. After edge detection is performed on the target object, the target image area can be obtained.

[0086] In some implementations, the target object may include at least one object.

[0087] For example, if the target object is food, the target object may include one or more foods, and each food may further include at least one.

[0088] Figure 3 is based on Figure 1 The embodiment shown is a flow chart of an object recognition method based on smart glasses, as shown in FIG. Figure 3 As shown, step S13 includes the following sub-steps: In step S131 , for each object, the object type and object size of the object are determined.

[0089] In step S132 , component information of the object is determined according to the object type and the object size.

[0090] In this way, the corresponding component information of each object selected by the user with a gesture can be identified and obtained, thereby improving user satisfaction.

[0091] During the execution of step S131, one implementation method may be to perform target detection on the target image area selected by the gesture in the target image to determine the object type corresponding to each object; for each object, determine the object size of the object according to the object type of the object.

[0092] As described above, after performing gesture detection on the target image and obtaining gesture key points, the target image region can be determined based on these gesture key points. The target image can then be cropped, retaining only the image data within this target image region. Subsequent reference object recognition and object type identification will then be performed on this cropped target image region, preventing interference from unintended objects and image backgrounds. This cropping operation can be performed by the smart glasses' processor, or the target image with the gesture key point pixel locations can be sent to the cloud for cropping.

[0093] In one embodiment of the present disclosure, a pre-trained target detection model may be used to perform target detection on the target image area to determine the object type of each object in the target image area and the image area corresponding to each object (i.e., a mask for each object).

[0094] After determining the object type of each object, the object size of each object can be determined based on the object type based on the first preset correspondence. This method is an estimation method, and the first preset correspondence includes a correspondence between object type and object size. Among them, the object sizes corresponding to different object types in the first preset correspondence can be set based on the average size of objects of each type. The object size of each object can also be estimated based on the object type of each object and the relative size of the objects in the image. It should be noted that this estimation method may have errors. Therefore, the present disclosure can output a prompt information for the object size of each object determined in this way. The prompt information is used to prompt the user that there may be errors in the recognition result or allow the user to manually set the object size of each object.

[0095] Figure 4 is based on Figure 3 The embodiment shown is a flow chart of an object recognition method based on smart glasses, as shown in FIG. Figure 4 As shown, step S131 includes the following sub-steps: In step S1311 , image recognition is performed on the target image area selected by the gesture in the target image to obtain an image recognition result.

[0096] After performing gesture detection on the target image and obtaining the gesture key points, the target image area can be determined based on the gesture key points. The target image can then be cropped to retain only the image data within the target image area. In this way, subsequent reference object recognition and object type recognition will be performed on this cropped target image area, avoiding interference from objects and image background not intended by the user on the current recognition task.

[0097] The image recognition result may include the object type of each object and the image area of ​​each object in the target image.

[0098] In this step, performing image recognition on the target image area selected by the gesture in the target image may include performing target detection on the target image area using a pre-trained target detection model to determine the object type of each object in the target image area, and the image area corresponding to each object (i.e., a mask for each object).

[0099] In step S1312 , target detection is performed on the target image to determine whether the target image further includes a preset reference object.

[0100] The present disclosure utilizes computer vision technology to determine the composition information of a user-selected target object based on image recognition results of a target image. This process requires identifying the target object's object size, and further identifying the composition information based on the object size. Therefore, the accuracy of object size recognition affects the accuracy of the composition information recognition results. Most related technologies directly analyze the pixel size of the target object in the image to estimate its weight. However, when users photograph the target object at different distances, the pixel area occupied by the same target object in the image varies significantly, resulting in significant fluctuations in the object size recognition results for the target object, which in turn affects the accuracy of the composition recognition results.

[0101] In order to solve the problem of inaccurate object size estimation of the target object caused by changes in shooting distance, the present disclosure introduces a preset reference object with a known actual size to calibrate the size of the target object, so that the system can more accurately convert the actual size of the target object in the image, thereby improving the accuracy of component information recognition.

[0102] The preset reference object may include, for example, a user's hand or a preset object of known size (such as a card or a mobile phone). The user may customize the preset reference object or use the system default reference object as the preset reference object.

[0103] In actual application scenarios, in response to the image acquisition unit on the smart glasses being activated, a prompt message can be output, and the prompt message can also be used to prompt the user to place a preset reference object next to the target object to be identified.

[0104] During this step, target detection may be performed on the original target image to determine whether the target image includes the preset reference object. Alternatively, target detection may be performed on a cropped target image region (i.e., the target image region where the target object is located) to determine whether the target image region includes the preset reference object. If it is determined that the target image region includes the preset reference object, the target image is determined to include the preset reference object.

[0105] Among them, a pre-trained preset reference object detection model can be used to detect whether the target image includes the preset reference object. The preset reference object detection model can be a deep learning model (such as CNN, Transformer and other models).

[0106] In step S1313 , in response to the target image further including a preset reference object, the object type and object size corresponding to each object are determined according to the preset reference object and the image recognition result.

[0107] In this step, the object type of each object and the image area of ​​each object in the target image can be determined based on the image recognition result; the first pixel size of the preset reference object in the target image and the second pixel size of the image area corresponding to each object can be identified; and the object size corresponding to each object can be determined based on the actual size of the preset reference object, the first pixel size and the second pixel size corresponding to each object.

[0108] In one implementation, a proportionality coefficient may be determined based on the actual size of a preset reference object and a first pixel size, where the proportionality coefficient represents the ratio of the actual size of the object to the pixel size; and for each object, the object size of the object is determined based on the proportionality coefficient and the second pixel size.

[0109] For example, feature extraction and matching can be performed on the preset reference object to identify the first pixel size of the preset reference object in the target image. For example, if the preset reference object is a user's hand, the first pixel size may include the pixel size of the hand's width in the target image (which can be understood as the image size of the hand's width in the image). If the preset reference object is a mobile phone, the first pixel size may include the pixel size of the phone's length in the target image and / or the pixel size of the phone's width in the target image. Then, based on the known actual size of the preset reference object and its first pixel size in the target image, the ratio coefficient between the pixel size and the actual size of the target object in the target image under the current shooting conditions is calculated.

[0110] For example, the scale factor K=the actual size of the preset reference object (mm) / the first pixel size (pixel).

[0111] After the scale factor is calculated, the object size of each object can be determined according to the scale factor and the second pixel size. The object size is the actual size of the object.

[0112] For example, for each object, the object size L of the object is calculated as follows: L=K×P, where K represents the scale factor and P represents the second pixel size of the object. The above examples are merely illustrative and are not limited in this disclosure.

[0113] Figure 5 is based on Figure 3 The embodiment shown is a flow chart of a method for object recognition, as shown in FIG. Figure 5 As shown, step S132 includes the following sub-steps: In step S1321 , the content of each component in at least one component corresponding to the object is determined according to the object type.

[0114] For each object, after determining the object type of the object, the content of each component corresponding to the object of this type can be obtained by looking up a table based on a preset database.

[0115] For example, if the object type is white rice, the content of each component corresponding to white rice is obtained by looking up the table in the nutritional composition database: per 100 grams of white rice contains: approximately 7.0 grams of protein, approximately 0.6 grams of fat, approximately 80.0 grams of carbohydrates, approximately 0.1 grams of sugar, and approximately 1.0 grams of fiber. This is merely an example and is not intended to be limiting.

[0116] In step S1322, the amount of the object is determined according to the size of the object.

[0117] The weight of the object may include the volume or weight of the object.

[0118] After determining the size of each object, the volume of the object can be calculated based on the size of the object, and then the weight of the object can be determined based on the object type. For example, the weight of food can be estimated using a food density database (which converts volume to weight) or directly using a vision-based weight estimation model.

[0119] In step S1323, the weight of each component corresponding to the object is determined according to the content of each component in the at least one component and the portion.

[0120] In one embodiment, the object to be identified includes food, the target object includes target food, and the component information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

[0121] For each food, after determining the portion size and the weight of each nutrient component of the food, the weight of each component of the food can be calculated based on the portion size and nutrient content of the food. For example, assuming that the food is a 200-gram portion of white rice, the content information per 100 grams of white rice is: approximately 7.0 grams of protein, approximately 0.6 grams of fat, approximately 80.0 grams of carbohydrates, approximately 0.1 grams of sugar, and approximately 1.0 grams of fiber. Then, the component information of the 200 grams of white rice can be further calculated as: approximately 14.0 grams of protein, approximately 1.2 grams of fat, approximately 160 grams of carbohydrates, approximately 0.2 grams of sugar, and approximately 2.0 grams of fiber. This is merely an example and is not limited in this disclosure.

[0122] If the target food includes multiple types, the total nutritional components of the target food consumed by the user can be further calculated.

[0123] In another embodiment of the present disclosure, the calorie content of the target food can be determined based on the weight of each nutrient in the at least one nutrient corresponding to the target food, where the calorie content is measured in calories. This provides the user with intuitive calorie data and a straightforward data reference for health management.

[0124] For example, if the target food is 200 grams of white rice, based on the above ingredient information corresponding to this 200 grams of white rice (i.e., containing approximately 14.0 grams of protein, approximately 1.2 grams of fat, approximately 160 grams of carbohydrates, approximately 0.2 grams of sugar, and approximately 2.0 grams of fiber), the total calories corresponding to this 200 grams of white rice can be converted.

[0125] In another embodiment of the present disclosure, nutritional plans can also be recommended to users, thereby providing services for users' health management and improving user experience.

[0126] In one implementation, user data may be obtained, the user data representing the user's nutritional needs; nutritional recommendations may be output based on the user data and the weight of each nutrient component corresponding to the target food, the nutritional recommendations including a diet plan that meets the nutritional needs.

[0127] For example, the user data may include the user's age, gender, eating habits, and health goals (such as weight control, weight loss, weight gain, muscle gain, etc.). Based on the above user data, the nutritional needs of the corresponding user can be determined, and then the weight of each nutrient component of the target food currently identified by the user is compared with the nutritional needs, and nutritional recommendations are output based on the comparison results. For example, if the current weight of the nutrient components of the target food cannot meet the user's nutritional needs, other foods can be recommended to the user for nutritional supplementation. If the current weight of the nutrient components of the target food has exceeded the user's nutritional needs, a prompt message can be output to the user, and the prompt message is used to prompt the user to reduce the intake of the target food. This is only an example, and the present disclosure is not limited to this.

[0128] In another embodiment of the present disclosure, interactive information may also be output through smart glasses; wherein the interactive information includes the identified ingredient information of the target object and / or the calorie content of the target food.

[0129] After identifying and obtaining the component information of the target object, the present disclosure may also output the component information to the user, wherein the output form of the component information may be in the form of voice and / or text display.

[0130] For example, if the target objects are multiple foods, the analyzed nutritional information (such as total calories and weight of each nutrient) can be fed back to the user through the smart glasses' display (e.g., micro-projection onto the lenses) or voice broadcast module. The user can confirm the record, and the data can be saved to the user's health record, providing data reference for health management.

[0131] Figure 6 FIG. 1 is a flow chart showing a method for identifying food ingredients based on smart glasses according to an exemplary embodiment. Figure 6 As shown, the method includes the following steps: In step S601, in response to the food recording function of the smart glasses being activated, a target image within the current field of view is captured.

[0132] In step S602 , gesture detection is performed on the target image.

[0133] In step S603, in response to detecting a valid gesture in the target image, the target image area where the target food is located is identified and cropped according to the gesture, and a reference object detection is performed on the target image area to determine whether the target image includes a preset reference object.

[0134] In step S604 , in response to not detecting a valid gesture in the target image, reference object detection is performed on the target image to determine whether the target image includes a preset reference object.

[0135] In step S605, in response to the target image including a preset reference object, a proportional coefficient is calculated based on the actual size of the preset reference object and the first pixel size, and the size of the target food is calibrated according to the proportional coefficient to obtain the food size of each food in the target food.

[0136] In step S606 , in response to the target image not including a preset reference object, the food size of each food in the target food is estimated.

[0137] In step S607 , for each food in the target food, the portion size of the food is determined based on the food size, and the nutritional components of the food are determined based on the portion size and the food type of the food.

[0138] In step S608 , the ingredient information and / or calorie content of the target food is output to the user.

[0139] The implementation of each step in the above steps S601-S608 can refer to the relevant description in the above embodiment and will not be repeated here.

[0140] Using this method, users can lock the target image area of ​​the target food they want to analyze through intuitive gestures. The system automatically crops the image, effectively eliminating interference from irrelevant food or other objects in the background, avoiding misidentification and misrecording, while simplifying user operations and improving the accuracy and convenience of target food selection.

[0141] Furthermore, by introducing preset reference objects for size calibration, the problem of inaccurate portion estimation caused by varying shooting distances and angles is effectively overcome, making nutritional analysis results closer to actual values. In this disclosure, using a single camera with preset reference object calibration can, to a certain extent, replace or reduce the need for expensive and power-hungry depth sensors, thereby reducing hardware costs and power consumption.

[0142] Using the user's hands or everyday objects (mobile phones) as preset reference objects, and the interactive method of using both hands to gesture and select frames, conforms to people's natural behavioral habits and has a low learning cost.

[0143] The smart glasses device that executes the food ingredient identification method can be, for example, a portable device such as smart glasses. By executing the above-mentioned food ingredient identification method, the application of smart glasses in dietary health management becomes more accurate, reliable and easy to use, expanding its function as a personal health assistant.

[0144] Figure 7 FIG. 1 is a block diagram of an object recognition device based on smart glasses according to an exemplary embodiment. Figure 7 As shown, the device includes: An acquisition module 701 is configured to acquire a target image through an image acquisition unit of the smart glasses, where the target image includes a gesture and at least one object to be recognized; A determination module 702 is configured to determine a target object selected by the gesture from the at least one object to be identified; The identification module 703 is configured to identify the component information of the target object.

[0145] Optionally, the determination module 702 is configured to perform gesture detection on the target image to determine pixel positions of key points corresponding to the gesture; and determine the target object according to the pixel positions.

[0146] Optionally, the determination module 702 is configured to determine a target image area selected by the gesture in the target image according to the pixel position; and use the object to be identified in the target image area as the target object.

[0147] Optionally, the target object includes at least one object; the recognition module 703 is configured to determine the object type and object size of each object; and determine the component information of the object according to the object type and the object size.

[0148] Optionally, the recognition module 703 is configured to perform target detection on the target image area selected by the gesture in the target image to determine the object type corresponding to each object; for each of the objects, determine the object size of the object according to the object type of the object.

[0149] Optionally, the recognition module 703 is configured to perform image recognition on the target image area selected by the gesture in the target image to obtain an image recognition result; perform target detection on the target image to determine whether the target image also includes a preset reference object; in response to the target image also including the preset reference object, determine the object type and the object size corresponding to each object according to the preset reference object and the image recognition result.

[0150] Optionally, the recognition module 703 is configured to determine the object type of each object and the image area of ​​each object in the target image based on the image recognition result; identify the first pixel size of the preset reference object in the target image and the second pixel size of the image area corresponding to each object; and determine the object size corresponding to each object based on the actual size of the preset reference object, the first pixel size and the second pixel size corresponding to each object.

[0151] Optionally, the recognition module 703 is configured to determine a proportional coefficient based on the actual size and the first pixel size, where the proportional coefficient represents the ratio of the actual size of the object to the pixel size; and for each object, determine the object size of the object based on the proportional coefficient and the second pixel size.

[0152] Optionally, the ingredient information includes the weight of each ingredient in at least one ingredient corresponding to the object; the identification module 703 is configured to determine the content of each ingredient in at least one ingredient corresponding to the object according to the object type; determine the portion of the object according to the object size; and determine the weight of each ingredient corresponding to the object based on the content of each ingredient in the at least one ingredient and the portion.

[0153] Optionally, the object to be identified includes food, and the ingredient information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

[0154] Optionally, the identification module 703 is further configured to determine the calories of the target food according to the weight of each nutrient in the at least one nutrient.

[0155] Optionally, Figure 8 is based on Figure 7 The embodiment shown is a block diagram of an object recognition device based on smart glasses, such as Figure 8 As shown, the device also includes: The interaction module 704 is configured to obtain user data, which represents the user's nutritional needs; and output nutritional recommendations based on the user data and the weight of each nutrient component corresponding to the target food, wherein the nutritional recommendations include a diet plan that meets the nutritional needs.

[0156] Optionally, the interaction module 704 is further configured to output interaction information through the smart glasses; wherein the interaction information includes the ingredient information and / or the calories of the target food.

[0157] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0158] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the object recognition method provided by the present disclosure are implemented.

[0159] Figure 9 This is a block diagram of a terminal for object recognition according to an exemplary embodiment. For example, terminal 800 may be a wearable device (such as glasses or a watch), a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0160] Reference Figure 9 The terminal 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0161] The processing component 802 generally controls the overall operation of the terminal 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the object recognition method described above. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0162] The memory 804 is configured to store various types of data to support operations on the terminal 800. Examples of such data include instructions for any application or method operating on the terminal 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0163] Power supply component 806 provides power to various components of terminal 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to terminal 800.

[0164] The multimedia component 808 includes a screen that provides an output interface between the terminal 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the terminal 800 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.

[0165] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the terminal 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0166] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0167] The sensor assembly 814 includes one or more sensors for providing various aspects of the terminal 800's status assessment. For example, the sensor assembly 814 can detect the open / closed state of the terminal 800, the relative positioning of components, such as the display and keypad of the terminal 800. The sensor assembly 814 can also detect changes in the position of the terminal 800 or a component of the terminal 800, the presence or absence of user contact with the terminal 800, the orientation or acceleration / deceleration of the terminal 800, and temperature changes of the terminal 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0168] The communication component 816 is configured to facilitate wired or wireless communication between the terminal 800 and other devices. The terminal 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0169] In an exemplary embodiment, the terminal 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned object recognition method.

[0170] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions. The instructions may be executed by the processor 820 of the terminal 800 to perform the object recognition method described above. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0171] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for performing the above-mentioned object recognition method when executed by the programmable device.

[0172] Figure 10 9 is a block diagram of an apparatus for object recognition according to an exemplary embodiment. For example, apparatus 900 may be provided as a server. Figure 10 The apparatus 900 includes a processing component 922, which further includes one or more processors, and a memory resource represented by a memory 932 for storing instructions, such as an application, that can be executed by the processing component 922. The application stored in the memory 932 can include one or more modules, each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute the instructions to perform the object recognition method described above.

[0173] The device 900 may also include a power supply component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or similar.

[0174] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented through electronic hardware, computer software, or a combination of both. Whether such functions are implemented through hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0175] In the foregoing detailed description, reference is made to the accompanying drawings, which illustrate, by way of illustration, specific aspects of the present disclosure in which it may be practiced. In this regard, terms indicating directions or expressing positional relationships, such as "center," "longitudinal," "lateral," "length," "width," "thickness," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," "clockwise," "counterclockwise," "axial," "radial," "circumferential," etc., may be used with reference to the orientation of the figures being described. Since the components of the described devices may be positioned in a plurality of different orientations, the directional terms may be used for illustrative purposes rather than restrictive. It should be understood that other aspects may be utilized and structural or logical changes may be made without departing from the concepts of the present disclosure. Therefore, the following detailed description should not be taken in a limiting sense.

[0176] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.

[0177] It should be understood that, unless otherwise expressly specified or limited, the terms "join," "attach," "install," "connect," "connect," "fix," etc. used in the embodiments of the present disclosure should be understood in a broad sense. For example, they can be fixedly connected, detachably connected, or integrated; they can be mechanically connected, electrically connected, or communicable with each other; they can be directly connected, or indirectly connected through an intermediate medium, and they can be internally connected between two elements or an interactive relationship between two elements, unless otherwise expressly limited. For those skilled in the art, the specific meanings of the above terms in this article can be understood according to specific circumstances.

[0178] Additionally, the term "over" as used in reference to a component, element, or material layer being formed "over" or located "over" a surface may be used herein to mean that the component, element, or material layer is "indirectly" positioned (e.g., placed, formed, deposited, etc.) on the surface such that one or more additional components, elements, or layers are disposed between the surface and the component, element, or material layer. However, the term "over" as used in reference to a component, element, or material layer being formed "over" or located "over" a surface may alternatively have a specific meaning: the component, element, or material layer is "directly" positioned (e.g., placed, formed, deposited, etc.) on the surface, e.g., in direct contact with the surface.

[0179] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one such feature. In the description herein, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined.

[0180] It should be understood that spatially relative terms, such as "above," "upper," "below," and "lower," are used herein to describe the relationship of one element to another element shown in the figures. Such spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, an element described as being "above" or "upper" relative to another element would then be "below" or "lower" relative to the other element. Thus, the term "above" encompasses both above and below orientations, depending on the spatial orientation of the device. The device may be oriented in other ways (e.g., rotated 90 degrees or in other orientations), and the spatially relative terms used herein should be interpreted accordingly.

[0181] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.

[0182] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. With particular regard to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. In addition, although particular features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include," "have," "have," "have," or variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0183] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

[0184] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An object recognition method based on smart glasses, characterized in that: The method comprises: Acquire a target image through an image acquisition unit of the smart glasses, wherein the target image includes a gesture and at least one object to be recognized; determining a target object selected by the gesture from the at least one object to be identified; Identify composition information of the target object.

2. The method according to claim 1, characterized in that The determining the target object selected by the gesture from the at least one object to be identified comprises: Performing gesture detection on the target image to determine pixel positions of key points corresponding to the gesture; The target object is determined according to the pixel position.

3. The method according to claim 2, characterized in that Determining the target object according to the pixel position includes: determining a target image area selected by the gesture in the target image according to the pixel position; The object to be identified in the target image area is used as the target object.

4. The method according to claim 1, wherein The target object includes at least one object; the component information for identifying the target object includes: For each object, determining an object type and an object size of the object; The composition information of the object is determined according to the object type and the object size.

5. The method according to claim 4, characterized in that The determining, for each object, the object type and object size of the object includes: performing object detection on the target image area selected by the gesture in the target image to determine the object type corresponding to each object; For each of the objects, the object size of the object is determined according to the object type of the object.

6. The method according to claim 4, characterized in that The determining, for each object, the object type and object size of the object includes: Performing image recognition on the target image area selected by the gesture in the target image to obtain an image recognition result; Performing target detection on the target image to determine whether the target image also includes a preset reference object; In response to the target image further including the preset reference object, the object type and the object size corresponding to each object are determined according to the preset reference object and the image recognition result.

7. The method according to claim 6, characterized in that The determining, based on the preset reference object and the image recognition result, the object type and the object size corresponding to each object includes: determining an object type of each object and an image area of ​​each object in the target image according to the image recognition result; Identifying a first pixel size of the preset reference object in the target image and a second pixel size of the image region corresponding to each object; The object size corresponding to each object is determined according to the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object.

8. The method according to claim 7, characterized in that The determining the object size corresponding to each object according to the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object includes: determining a scale factor according to the actual size and the first pixel size, wherein the scale factor represents a ratio of the actual size of the object to the pixel size; For each object, the object size of the object is determined according to the scale factor and the second pixel size.

9. The method according to claim 4, characterized in that The component information includes the weight of each component of the at least one component corresponding to the object; and determining the component information of the object according to the object type and the object size includes: determining, according to the object type, the content of each component in at least one component corresponding to the object; determining the portion of the object based on the size of the object; The weight of each component corresponding to the object is determined according to the content of each component in the at least one component and the portion size.

10. The method according to any one of claims 1 to 9, characterized in that The object to be identified includes food, and the ingredient information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

11. The method according to claim 10, characterized in that The method further comprises: The calories of the target food are determined according to the weight of each nutrient component in the at least one nutrient component.

12. The method according to claim 10, characterized in that The method further comprises: obtaining user data, wherein the user data represents nutritional needs of the user; Nutritional advice is output based on the user data and the weight of each nutrient component corresponding to the target food, where the nutritional advice includes a diet plan that meets the nutritional needs.

13. The method according to claim 10, characterized in that The method further comprises: outputting interactive information through the smart glasses; The interactive information includes the ingredient information and / or the calories of the target food.

14. An object recognition device based on smart glasses, characterized in that: include: an acquisition module, configured to acquire a target image through an image acquisition unit of the smart glasses, wherein the target image includes a gesture and at least one object to be recognized; a determination module, configured to determine a target object selected by the gesture from the at least one object to be identified; The recognition module is configured to recognize the component information of the target object.

15. The device according to claim 14, characterized in that The determination module is configured to perform gesture detection on the target image to determine pixel positions of key points corresponding to the gesture; and determine the target object according to the pixel positions.

16. The device according to claim 14, characterized in that The target object includes at least one object; the recognition module is configured to determine the object type and object size of each object; and determine the component information of the object according to the object type and the object size.

17. The device according to claim 16, characterized in that The recognition module is configured to perform image recognition on the target image area selected by the gesture in the target image to obtain an image recognition result; perform target detection on the target image to determine whether the target image also includes a preset reference object; in response to the target image also including the preset reference object, determine the object type and the object size corresponding to each object according to the preset reference object and the image recognition result.

18. The device according to claim 17, characterized in that The recognition module is configured to determine the object type of each object and the image area of ​​each object in the target image according to the image recognition result; Identifying a first pixel size of the preset reference object in the target image and a second pixel size of the image region corresponding to each object; The object size corresponding to each object is determined according to the actual size of the preset reference object, the first pixel size, and the second pixel size corresponding to each object.

19. The device according to claim 16, characterized in that The ingredient information includes the weight of each ingredient in at least one ingredient corresponding to the object; the identification module is configured to determine the content of each ingredient in at least one ingredient corresponding to the object according to the object type; determine the portion of the object according to the object size; and determine the weight of each ingredient corresponding to the object based on the content of each ingredient in the at least one ingredient and the portion.

20. The device according to any one of claims 14 to 19, characterized in that The object to be identified includes food, and the ingredient information includes the weight of each nutrient component in at least one nutrient component corresponding to the target food.

21. The device according to claim 20, characterized in that The identification module is further configured to determine the calories of the target food according to the weight of each nutrient component in the at least one nutrient component.

22. A terminal, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to: execute the steps of the method according to any one of claims 1 to 13.

23. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

24. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 13 when executed by a processor.