Object recognition method and apparatus, robot, and readable storage medium

By receiving user instructions through intelligent robots, and using recognition and mobile devices to autonomously identify and move to the target object within the target scene, the problem of lack of autonomy and accuracy in identifying and finding target objects in existing technologies is solved, and the effect of autonomous identification and finding people is achieved.

CN115909463BActive Publication Date: 2026-05-29MIDEA GRP (SHANGHAI) CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MIDEA GRP (SHANGHAI) CO LTD
Filing Date
2022-12-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, home service robots lack autonomy and accuracy in identifying and locating target objects, and cannot effectively meet the diverse needs of users.

Method used

The intelligent robot receives user commands and uses recognition and mobile devices to autonomously identify and move to the target object in the target scene. By combining facial feature matching, target recognition model and tracking algorithm, it can achieve the function of autonomously finding the target object.

Benefits of technology

This technology enables intelligent robots to autonomously identify and move to the location of target objects within a target scene, improving the accuracy and efficiency of identification and search, enriching the application functions of robots, and meeting user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909463B_ABST
    Figure CN115909463B_ABST
Patent Text Reader

Abstract

The application provides an object recognition method and device, a robot and a readable storage medium. The object recognition method comprises the following steps: determining a target object to be recognized according to a target instruction; performing object recognition in a target scene, determining a target position of the target object in the target scene according to a recognition result; and moving to the target position based on the target position. The technical scheme provided by the application can realize active recognition and searching of a target object in a target scene by a robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and more specifically, to an object recognition method and apparatus, a robot, and a readable storage medium. Background Technology

[0002] Currently, with the development of robotics technology, the application of home service robots is becoming increasingly widespread, and users' functional requirements for these robots are gradually increasing. Therefore, functional research on robots by those skilled in the art has become particularly important.

[0003] Application content

[0004] This application aims to address at least one of the technical problems existing in the prior art or related technologies.

[0005] Therefore, the first aspect of this application is to propose an object recognition method.

[0006] The second aspect of this application is to propose an object recognition device.

[0007] The third aspect of this application is to propose a robot.

[0008] The fourth aspect of this application is to propose a readable storage medium.

[0009] In view of this, according to one aspect of this application, an object recognition method is proposed, the method comprising: determining a target object to be recognized according to a target instruction; performing object recognition within a target scene, and determining the target position of the target object in the target scene based on the recognition result; and moving towards the target position based on the target position.

[0010] The object recognition method provided in this application can be implemented by an intelligent robot, or by an object recognition device within an intelligent robot, or it can be determined according to actual usage requirements, without specific limitations here. To more clearly describe the object recognition method provided in this application, the following explanation will use an intelligent robot as the implementing entity for the object recognition method.

[0011] The object recognition method provided in this application is used to control an intelligent robot to perform tasks, enabling the robot to actively identify and locate target objects within a target scene. The intelligent robot is equipped with a recognition device and a mobility device. In practical applications, the intelligent robot can specifically be a robotic vacuum cleaner, a lobby robot, a child companion robot, an airport intelligent robot, an intelligent question-and-answer robot, or other intelligent robots with autonomous mobility capabilities; no specific limitations are imposed here.

[0012] Specifically, in the object recognition method provided in this application, when a user uses the aforementioned intelligent robot, the user can issue a target command to the intelligent robot through button input, voice input, or other means. This target command instructs the intelligent robot to recognize a target object. Based on this, after receiving the user's target command, the intelligent robot analyzes the semantic information of the user's target command and determines the target object associated with that command, i.e., determines the target object to be recognized by the intelligent robot. Further, the aforementioned intelligent robot stores a target scene map. The intelligent robot retrieves the pre-stored target scene map from its storage device and, based on the target scene map, the target recognition model, and the user's target command, autonomously identifies the target object indicated by the target command within the target scene using its onboard recognition device. Based on the recognition result of its onboard recognition device, the intelligent robot determines the specific location information of the target object within the target scene, i.e., the target position. Further, after the intelligent robot determines the specific location information of the target object within the target scene, the intelligent robot uses its onboard movement device to move itself towards the target object until the intelligent robot reaches the location of the target object, at which point the intelligent robot stops moving. In this way, without the need for manual control by the user, the intelligent robot can autonomously identify and find the target object indicated by the user's target command within the target scene, enriching the application functions of the intelligent robot and meeting the user's needs.

[0013] The object recognition method described above in this application may also have the following additional technical features:

[0014] In the above technical solution, determining the target location of the target object in the target scene based on the recognition result includes: determining at least one facial feature based on the recognition result; in response to the first facial feature among the at least one facial features successfully matching the target facial feature, determining the first recognition object corresponding to the first facial feature as the target object, and determining the target location of the target object; wherein, the target facial feature corresponds to the target object.

[0015] In this technical solution, during the process of the intelligent robot determining the specific location information of the target object within the target scene (i.e., the target location) based on the recognition results of its onboard recognition device, the intelligent robot extracts facial features of at least one object within the target scene. The intelligent robot then matches the extracted facial features of each object with the target facial features of the target object pre-stored in its storage device. If any facial feature of any object (i.e., the aforementioned first facial feature) matches the target facial feature, the intelligent robot identifies that first object as the target object and determines its specific location. In this way, the intelligent robot determines the target object and its specific location based on the matching results of the pre-stored target facial features and the facial features of the identified objects, ensuring the accuracy of the intelligent robot's autonomous identification and search capabilities.

[0016] In any of the above technical solutions, the object recognition method further includes: in response to the failure of matching the first facial feature with the target facial feature and the successful matching of the first facial feature with the facial feature in the cache interval, moving at a first speed to continue target recognition; in response to the failure of matching the first facial feature with the target facial feature and the facial feature in the cache interval, storing the first facial feature in the cache interval, and moving at a first speed to continue target recognition.

[0017] In this technical solution, during the process of the intelligent robot matching the facial features of each extracted object with the target facial features of the target object pre-stored in its storage device, if any facial feature of the object being identified (i.e., the aforementioned first facial feature) does not match the target facial feature (i.e., the match fails), the intelligent robot will continue to compare and match the first object feature with the facial features temporarily cached in its cache interval. Based on this, if there is a facial feature in the cache interval that matches the first facial feature (i.e., the match is successful), it indicates that the first facial feature is not the facial feature of the target object to be identified, and that the first facial feature has already been identified and analyzed. At this point, there is no need to analyze and process the first facial feature again, and the intelligent robot moves at a first speed to continue target identification within the target scene. Conversely, if every facial feature in the cache interval does not match the first facial feature (i.e., the match fails), the intelligent robot will temporarily cache the first facial feature and move at the first speed to continue target identification within the target scene. In this way, the intelligent robot controls its identification speed based on the object identification results, thus controlling the robot's movement speed during the identification and search process, improving the efficiency of the intelligent robot's identification and search.

[0018] In any of the above technical solutions, in response to the failure to match the first facial feature with the target facial feature and the facial features in the cache interval, the object recognition method further includes: obtaining the target identifier in the recognition interface; when the target identifier indicates that the facial feature information of the recognition interface has changed, adjusting the recognition angle according to the recognition box information of the target recognition model, and moving according to the second speed; wherein, the first speed is greater than the second speed.

[0019] In this technical solution, if the first facial feature does not match the target facial feature (i.e., matching fails), and every facial feature in the cache interval does not match the first facial feature (i.e., matching fails), the intelligent robot will also acquire a target identifier on its recognition interface. This target identifier indicates changes in facial feature information on the recognition interface. Based on this, if the intelligent robot determines that the facial feature information on its recognition interface has not changed, it moves at a first speed without adjusting its recognition angle. However, if the intelligent robot determines that the facial feature information on its recognition interface has changed, such as the appearance of a new first facial feature, it moves at a second speed and adjusts its recognition angle according to the recognition box information of the target recognition model to track and recognize the newly appearing first facial feature on the recognition interface. The second speed is less than the first speed, and the target recognition model is the recognition model used by the intelligent robot when performing object recognition. Thus, the intelligent robot adjusts its recognition speed and recognition angle based on the object recognition result and changes in facial feature information on its recognition interface, i.e., adjusting the intelligent robot's movement speed and head posture during the identification and search process. This improves the image quality captured by the intelligent robot during its search for people and reduces the impact of the robot's head posture on the recognition results. In other words, while maintaining efficiency in finding people, it also improves the image quality captured during object recognition, thereby increasing the success rate of the intelligent robot in identifying and finding people while in motion.

[0020] In any of the above technical solutions, adjusting the recognition angle based on the recognition box information of the target recognition model includes: obtaining a first facial recognition box corresponding to the first facial feature; determining a corresponding first head recognition box based on the first facial recognition box; determining a first distance based on the first head recognition box; and adjusting the recognition angle based on the first distance and the position information of the first facial recognition box in the recognition interface.

[0021] In this technical solution, during the process of the intelligent robot adjusting its recognition angle based on the recognition box information of the aforementioned target recognition model to track and recognize the newly emerging first facial feature in the recognition interface, the intelligent robot acquires the first facial recognition box output by the target recognition model, which corresponds to the newly emerging first facial feature. Based on this, the intelligent robot determines the first head recognition box with which it has a belonging relationship, i.e., determines the head recognition box of the first recognition object corresponding to the first facial feature, based on the recognition box information of the first head recognition box. Further, the intelligent robot determines the first distance between itself and the first recognition object based on the recognition box information of the first head recognition box, and then adjusts its recognition angle based on the position information of the first facial recognition box in the recognition interface and the determined first distance to track and recognize the first facial feature. In this way, the intelligent robot adjusts its recognition angle based on the recognition box information of the target recognition model it uses to achieve tracking and recognition, ensuring the accuracy of the intelligent robot's recognition of the first facial feature, thereby ensuring the accuracy of the intelligent robot's autonomous identification and search for people.

[0022] In any of the above technical solutions, determining the corresponding first head recognition box based on the first face recognition box includes: determining a target value between the first face recognition box and the head recognition box based on the size correspondence and area correspondence between the first face recognition box and each head recognition box output by the target recognition model; performing target matching on the target value; and determining the first head recognition box based on the matching result; wherein, the area correspondence includes the intersection area and the union area of ​​the first face recognition box and each head recognition box, and the target value is used to indicate the degree of correlation between the first face recognition box and the head recognition box.

[0023] In this technical solution, during the process of the intelligent robot determining the first head recognition frame with which it has a belonging relationship based on the recognition frame information of the first face recognition frame, specifically, the intelligent robot acquires each head recognition frame output by its target recognition model, and determines the target value between the first face recognition frame and each head recognition frame based on the area correspondence between the first face recognition frame and each head recognition frame, such as the intersection area and the union area, as well as the size correspondence between the first face recognition frame and each head recognition frame, to obtain at least one target value. This target value is used to indicate the degree of association between the first face recognition frame and the head recognition frame. Based on this, the intelligent robot then matches the at least one target value calculated above according to a target matching algorithm, and filters the head recognition frames based on the comparison results to determine the first head recognition frame that best matches the first face recognition frame. Thus, by analyzing the intersection and union relationships and size relationships between the first face recognition frame and the head recognition frames, the first head recognition frame with a belonging relationship to the first face recognition frame is determined, ensuring the accuracy of the determination of the first head recognition frame, and thus ensuring the accuracy of object tracking and recognition.

[0024] In any of the above technical solutions, the target value between the first face recognition box and the head recognition box is determined based on the size correspondence and area correspondence between the first face recognition box and each head recognition box output by the target recognition model. This includes: determining a first height ratio between the first face recognition box and each head recognition box, a first area of ​​the intersection region, and a second area of ​​the union region; when the first height ratio is greater than or equal to a first threshold, the first area is greater than or equal to a second threshold, and the second area is equal to the area of ​​the corresponding head recognition box, determining a first ratio between the first area and the second area, and determining a second ratio between the first pixel distance from the first edge of the head recognition box to the second edge of the first face recognition box and the height of the head recognition box; and determining the target value based on the first ratio and the second ratio.

[0025] In this technical solution, in the process of determining the target values ​​between the first face recognition frame and each head recognition frame, specifically, the intelligent robot determines the ratio between the height of the first face recognition frame and the height of each head recognition frame, respectively obtaining the corresponding first height ratio. Furthermore, the intelligent robot determines the first area of ​​the intersection region between the first face recognition frame and each head recognition frame, and simultaneously determines the second area of ​​the union region between the first face recognition frame and each head recognition frame.

[0026] Furthermore, based on the specific values ​​of the first area, second area, and first height ratio determined above, the intelligent robot performs a first screening of each head recognition box output by the target recognition model. Specifically, the intelligent robot filters out head recognition boxes with a first height ratio less than a first threshold, head recognition boxes with a first area less than a second threshold, and head recognition boxes with a second area greater than the corresponding head recognition box area. On this basis, for the remaining head recognition boxes, the intelligent robot calculates a first ratio between the first area and the second area corresponding to each head recognition box, and calculates a first pixel distance between the second edge of the first face recognition box and the first edge of each head recognition box, and then calculates a second ratio between the first pixel distance and the height of the corresponding head recognition box. Further, the intelligent robot calculates at least one target value based on each calculated second ratio and the corresponding first ratio. In this way, after filtering the head recognition boxes based on the intersection and union relationship and size relationship between the first face recognition box and the head recognition box, the degree of correlation between the first face recognition box and each of the filtered head recognition boxes is determined by analyzing the intersection and union relationship and size relationship between the first face recognition box and the head recognition box. This ensures the accuracy of the correlation between the first face recognition box and the head recognition box, while reducing the amount of computation and improving the efficiency of object recognition.

[0027] In any of the above technical solutions, adjusting the recognition angle based on the first distance and the position information of the first facial recognition box in the recognition interface includes: adjusting the recognition angle to the target angle when the first distance is greater than a distance threshold; and adjusting the recognition angle based on the position information when the first distance is less than or equal to the distance threshold, so that the first facial recognition box is located in the target area of ​​the recognition interface.

[0028] In this technical solution, during the process of adjusting the intelligent robot's recognition angle based on the position information of the first facial recognition frame in the recognition interface and the first distance between the intelligent robot and the first object to be recognized, the intelligent robot compares the determined first distance with a set distance threshold. If the set distance threshold is less than the first distance, it indicates that the distance between the intelligent robot and the first object to be recognized is large. In this case, the intelligent robot's recognition angle does not need to be adjusted in the pitch direction, and the first object to be recognized is within the intelligent robot's field of vision. The intelligent robot directly adjusts its pitch recognition angle to the target angle, such as 0 degrees, for object recognition. Conversely, if the set distance threshold is greater than or equal to the first distance, it indicates that the distance between the intelligent robot and the first object to be recognized is small. In this case, the intelligent robot adjusts its pitch angle for object recognition based on the position information of the first facial recognition frame in the recognition interface. Simultaneously, the intelligent robot also adjusts its horizontal rotation angle for object recognition based on changes in the position information of the first facial recognition frame in the recognition interface. In this way, the tilt angle of the intelligent robot when performing object recognition can be flexibly adjusted based on the distance between the intelligent robot and the first object to be recognized, and the horizontal rotation angle of the intelligent robot when performing object recognition can be flexibly adjusted based on the position information of the first facial recognition box in the recognition interface. This ensures the flexibility of adjusting the recognition angle of the intelligent robot, thereby ensuring the flexibility and accuracy of the intelligent robot in object tracking and recognition, and ensuring the accuracy of the intelligent robot in autonomously identifying and finding people.

[0029] In any of the above technical solutions, moving towards the target location based on the target location includes: obtaining a target face recognition box corresponding to the target object based on the recognition result; determining a corresponding target head recognition box based on the target face recognition box; determining a corresponding target body recognition box based on the target head recognition box; determining the target distance based on the target head recognition box; and moving the target body recognition box according to the target tracking algorithm and the target distance until the target location is reached.

[0030] In this technical solution, as the intelligent robot moves towards the target object based on the target location using its onboard mobile device, the robot obtains a target facial recognition box output by the target recognition model based on the recognition results of the recognition device. This target facial recognition box corresponds to the target object. Based on this, the intelligent robot determines a target head recognition box with a corresponding relationship to the target facial recognition box, and then determines a target torso recognition box with a corresponding relationship to the target head recognition box. Further, the intelligent robot determines the target distance between itself and the target object based on the target head recognition box information. Based on the determined target distance and the target tracking algorithm, it autonomously plans the necessary route to reach the target object's location. Following this route, the robot moves towards the target object using its onboard mobile device, tracking the determined target torso recognition box, until it reaches the target object's location, at which point the intelligent robot stops moving.

[0031] In any of the above technical solutions, object recognition within the target scene includes: determining the target sub-scene in the target scene according to the target instruction when the target instruction contains scene semantic information; moving to the target sub-scene and performing object recognition in the target sub-scene.

[0032] In this technical solution, during the process of the intelligent robot autonomously identifying the target object indicated by the user's target command within the target scene using its onboard recognition device, specifically, the intelligent robot analyzes the semantic information of the target command and determines the target sub-scene indicated by the command. Based on this, the intelligent robot determines the necessary movement route to the target sub-scene using the target scene map and obstacle information within the target scene, and controls its onboard movement device to move the entire robot towards the target sub-scene according to this route. Further, after the intelligent robot moves to the target sub-scene, it uses its onboard recognition device to perform target recognition within the target sub-scene, autonomously identifying the target object indicated by the user's command. In this way, without manual user control, the intelligent robot can autonomously move to the corresponding target sub-scene based on the user's target command and autonomously identify the target object indicated by the command within the target sub-scene, enriching the application functions of the intelligent robot and meeting the user's needs.

[0033] In any of the above technical solutions, object recognition within the target scene includes: determining a first movement route based on the target scene map when the target instruction does not contain scene semantic information; and performing object recognition within the target scene based on the first movement route.

[0034] In this technical solution, during the process of the intelligent robot autonomously identifying the target object indicated by the user's target command within the target scene using its onboard recognition device, specifically, the intelligent robot analyzes the semantic information of the target command to determine whether it contains scene semantic information. If the intelligent robot determines that the user's target command does not contain scene semantic information, it activates its whole-house recognition and person-finding function and autonomously plans a person-finding route, namely the first movement route. Based on this, the intelligent robot controls its onboard movement device to move within the target scene according to the first movement route. During its movement, it uses its onboard recognition device to identify objects in the target scene environment it traverses, thereby identifying the target object indicated by the user command from multiple objects within the target scene. In this way, without manual user control, the intelligent robot can autonomously plan a person-finding route based on the user's target command and autonomously identify the target object indicated by the target command within the target scene according to that route, enriching the application functions of the intelligent robot and meeting the user's needs.

[0035] In any of the above technical solutions, determining the first movement route based on the target scene map includes: traversing the target scene map and determining the first movement route based on the traversal results; or determining at least one detection point based on the target scene map and determining the first movement route based on at least one detection point.

[0036] In this technical solution, during the process of the intelligent robot autonomously planning its person-finding route (i.e., the first movement route) based on the target scene map, specifically, the intelligent robot traverses and analyzes the target scene map pre-stored in its storage device to analyze and process scene information such as obstacle information and road information in the target scene. Based on this, the intelligent robot plans a person-finding route (i.e., the first movement route) according to the traversal results of the target scene map, so that when the intelligent robot moves according to the planned first movement route, it can traverse the aforementioned target scene. In this way, the intelligent robot can identify the target object by traversing the entire target scene, achieving blind-spot-free detection of the target scene, ensuring the comprehensiveness and accuracy of the identification results, and improving the intelligent robot's person-finding performance.

[0037] In this technical solution, further, during the process of the intelligent robot autonomously planning its search route (i.e., the first movement route) based on the target scene map, specifically, the intelligent robot can also analyze and process the target scene map pre-stored in its storage device to autonomously plan at least one detection point within the target scene. Based on this, after planning at least one detection point, the intelligent robot then autonomously plans a search route (i.e., the first movement route) based on the target scene map, the specific coordinates of each detection point, and obstacle information in the target scene. This first movement route passes through each of the planned detection points. Therefore, by performing rotational recognition and detection at each detection point within the target scene, the intelligent robot can achieve comprehensive, blind-spot-free recognition and detection of the target scene, thus ensuring the completeness of the intelligent robot's autonomous search for people.

[0038] In any of the above technical solutions, the object recognition method is applied to a robot, and the method further includes: receiving a target instruction from a user, the target instruction being used to instruct the robot to perform a target task on a target object; and moving to the target position and performing the target task on the target object based on the target position.

[0039] In this technical solution, the aforementioned target command not only instructs the robot to identify and locate the target object within the target scene, but also instructs the robot to perform a target task on the identified target object, such as a relaying task or a delivery task. Based on this, after the robot identifies the target object within the target scene according to the user-issued target command, thereby determining the target object's location, the robot moves to that location. Upon reaching the target object, it interacts with the target object based on the target task indicated by the target command to complete the task. In this way, without manual user control, the intelligent robot can autonomously identify and locate the target object indicated by the user-issued target command within the target scene and perform the corresponding target task, enriching the application functions of the intelligent robot and meeting user needs.

[0040] According to a second aspect of this application, an object recognition device is proposed, comprising: a memory storing a program or instructions; and a processor, which, when executing the program or instructions, implements the steps of the object recognition method as described in any of the above-described technical solutions. Therefore, the object recognition device proposed in the second aspect of this application possesses all the beneficial effects of the object recognition method in any of the technical solutions of the first aspect, and will not be elaborated further here.

[0041] According to a third aspect of this application, a robot is proposed, comprising: a chassis, on which a moving device is disposed; a main body disposed on the chassis, on which a recognition device is disposed; and the object recognition device of the second aspect of the technical solution described above; wherein the object recognition device controls the recognition device to perform object recognition, and the object recognition device controls the moving device to move or rotate.

[0042] The robot proposed in the third aspect of this application includes the object recognition device in the technical solution of the second aspect described above. Therefore, the robot proposed in the third aspect of this application possesses all the beneficial effects of the object recognition device in the technical solution of the second aspect described above, which will not be repeated here.

[0043] Furthermore, the robot also includes a chassis and a main body. The chassis is equipped with a moving device that can move or rotate the chassis and main body together. The main body is equipped with a recognition device. During the robot's operation, the object recognition device can control the recognition mechanism to identify objects and, based on the recognition results, control the moving device to move or rotate, adjusting the recognition angle of the recognition mechanism to achieve the robot's tracking and recognition function.

[0044] Furthermore, in practical applications, the aforementioned robots include, but are not limited to: robot vacuum cleaners, lobby robots, child companion robots, airport intelligent robots, intelligent question-and-answer robots, and home service robots, as well as other robots with object recognition capabilities.

[0045] The robot described above according to this application may also have the following additional technical features:

[0046] In the above technical solution, the main body includes: a first main body, which is mounted on a chassis; a second main body, which is mounted on the first main body and rotatably connected to the first main body; an identification device is mounted on the second main body; the second main body can swing along the axis of the robot and rotate around the axis of the robot.

[0047] In this technical solution, the main body of the robot includes a first body and a second body. The second body can be the robot head. The second body is equipped with a recognition device and can rotate or pitch relative to the first body. Specifically, the second body can swing along the robot's axis and rotate around the robot's axis.

[0048] Based on this, during the process of adjusting the pitch angle when the intelligent robot recognizes an object, the intelligent robot can specifically drive the recognition device to perform pitch movement through its second main body, such as the robot head, in order to adjust the pitch recognition angle of the recognition device, thereby adjusting the pitch recognition angle of the intelligent robot.

[0049] Furthermore, during the process of the intelligent robot adjusting its horizontal rotation angle for object recognition based on changes in the position of the first facial recognition frame within the recognition interface, the intelligent robot can rotate independently via a movable device on its chassis; alternatively, it can rotate the recognition device independently via its second main body, such as the robot head; or it can rotate in conjunction with its second main body and the movable device on its chassis. Specifically, when the position of the first facial recognition frame within the recognition interface changes significantly, the intelligent robot adjusts its horizontal rotation angle by rotating its entire body via the movable device on its chassis. Conversely, when the position change of the first facial recognition frame within the recognition interface is minor, the intelligent robot adjusts its horizontal rotation angle by rotating the recognition device via its second main body. In practical applications, those skilled in the art can set the specific method for adjusting the horizontal rotation angle of the intelligent robot according to the actual situation, and no specific limitations are imposed here.

[0050] According to a fourth aspect of this application, a readable storage medium is proposed, on which a program or instructions are stored, which, when executed by a processor, implement the object recognition method as described in any of the above-described technical solutions. Therefore, the readable storage medium proposed in the fourth aspect of this application possesses all the beneficial effects of the object recognition method in any of the technical solutions of the first aspect, and will not be elaborated further here.

[0051] Additional aspects and advantages of this application will become apparent in the following description or may be learned by practice of this application. Attached Figure Description

[0052] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0053] Figure 1 One of the flowcharts of the object recognition method according to an embodiment of this application is shown;

[0054] Figure 2 A second schematic flowchart of the object recognition method according to an embodiment of this application is shown;

[0055] Figure 3 The third schematic flowchart of the object recognition method according to an embodiment of this application is shown;

[0056] Figure 4 The fourth schematic flowchart of the object recognition method according to an embodiment of this application is shown;

[0057] Figure 5 The fifth schematic flowchart of the object recognition method according to an embodiment of this application is shown;

[0058] Figure 6 The sixth schematic flowchart of the object recognition method according to an embodiment of this application is shown;

[0059] Figure 7 This paper illustrates a diagram showing the correspondence between the recognition boxes of the target recognition model according to an embodiment of this application.

[0060] Figure 8 A structural block diagram of an object recognition device according to an embodiment of this application is shown;

[0061] Figure 9 A structural block diagram of the robot according to an embodiment of this application is shown;

[0062] Figure 10 A flowchart illustrating the workflow framework of a robot according to an embodiment of this application is shown. Detailed Implementation

[0063] To better understand the above-mentioned objectives, features, and advantages of this application, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0064] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.

[0065] The following is combined with Figures 1 to 10 The object recognition method, apparatus, robot, and readable storage medium provided in this application will be described in detail through specific embodiments and application scenarios.

[0066] In one embodiment of this application, such as Figure 1 As shown, the object recognition method may specifically include the following steps 102 to 108:

[0067] Step 102: Determine the target object to be identified based on the target instruction;

[0068] Step 104: Perform object recognition within the target scene;

[0069] Step 106: Determine the target location of the target object in the target scene based on the recognition results;

[0070] Step 108: Based on the target location, move towards the target location. The object recognition method provided in this application is used to control an intelligent robot to perform tasks, enabling the robot to actively identify and locate target objects within a target scene. The intelligent robot is equipped with a recognition device and a movement device. In practical applications, the intelligent robot can specifically be a sweeping robot, a lobby robot, a child companion robot, an airport intelligent robot, an intelligent question-and-answer robot, a home service robot, or other intelligent robots with autonomous movement capabilities; no specific limitations are imposed here.

[0071] Specifically, in the object recognition method provided in this application, when a user uses the aforementioned intelligent robot, the user can issue a target instruction to the intelligent robot through button input, voice input, or other means. This target instruction is used to instruct the intelligent robot to recognize the target object.

[0072] After receiving the user's target instruction, the intelligent robot analyzes the semantic information of the user's target instruction and determines the target object associated with the target instruction, thereby identifying the target object to be recognized by the intelligent robot.

[0073] The intelligent robot stores a target scene map of the target scene. The intelligent robot retrieves the target scene map pre-stored in its storage device. Based on the target scene map, the target recognition model and the user's target command, the intelligent robot autonomously identifies the target object indicated by the target command in the target scene through its recognition device. Based on the recognition result of its recognition device, the intelligent robot determines the specific location information of the target object in the target scene, i.e., the target location.

[0074] After determining the specific location of the target object within the target scene, the intelligent robot, based on the target tracking algorithm, the recognition results of the recognition device, and the bounding box information of the aforementioned target recognition model, moves itself towards the target object using its onboard mobility device until it reaches the target object's location, at which point it stops moving. In this way, without manual user control, the intelligent robot can autonomously identify the target object indicated by the user's target command within the target scene, enriching the robot's application functions and meeting user needs.

[0075] The intelligent robot, through its onboard recognition device, autonomously identifies the target object indicated by the aforementioned target command within the target scene. Specifically, based on the user's target command and the aforementioned target scene map, the intelligent robot autonomously plans its search route within the target scene. Based on the planned search route, the intelligent robot controls its onboard movement device to move its entire body within the target scene, and based on the aforementioned target recognition model, uses its onboard recognition device to perform object recognition within the target scene, thereby autonomously identifying the target object.

[0076] In the process of determining the specific location information of a target object within a target scene, i.e., the target location, based on the recognition results of its onboard recognition device, the intelligent robot extracts object features of at least one object within the target scene. The intelligent robot then matches the extracted object features of each object with the target object features pre-stored in its storage device. If a match is found between the object features of an object and the target object features, the intelligent robot identifies that object as the target object and determines its specific location information. In practical applications, these target object features include, but are not limited to, biometric features such as facial features, voice features, and body features; no specific limitations are imposed here.

[0077] As the intelligent robot moves toward the target object based on the target tracking algorithm, the recognition results of the recognition device, and the recognition box information of the aforementioned target recognition model, it determines the distance between the target object and the intelligent robot based on the recognition results of the recognition device and the recognition box information of the aforementioned target recognition model. Then, based on this distance information and the target tracking algorithm, it autonomously plans the movement route required to move to the location of the target object, and moves toward the target object by means of its moving device according to the movement route, until the intelligent robot moves to the location of the target object and stops moving.

[0078] In practical applications, the aforementioned target recognition models can specifically be detection models such as the YOLO series and the Faster R-CNN target detection model, which can simultaneously detect and output face recognition information, head recognition information, and body recognition information. The aforementioned target recognition models can also be recognition models composed of single models such as face detection models, head detection models, and body detection models. Those skilled in the art can choose the specific type of the aforementioned target recognition model according to the actual situation; no specific selection is made here.

[0079] The bounding box information of the aforementioned target recognition model may specifically include face bounding box information, head bounding box information, and body bounding box information. Specifically, the bounding box information may include the position information and size information of each bounding box, etc., without specific limitations.

[0080] In practical applications, the target tracking algorithm mentioned above may be a REID-based target tracking algorithm, etc. Those skilled in the art can choose the specific type of the target tracking algorithm according to the actual situation, without making specific restrictions here.

[0081] In practical applications, after the intelligent robot moves to the location of the target object, it can actively interact with the target object based on the target instructions issued by the user in order to complete the target task indicated by the user's instructions.

[0082] In summary, the object recognition method provided in this application enables the intelligent robot to automatically identify the target object indicated by the user's target command within the target scene after receiving the command, thereby determining the specific location information of the target object and moving towards it based on that location. This achieves the intelligent robot's autonomous person-finding function, eliminating the need for manual user control. The intelligent robot can automatically plan its route and perform autonomous person-finding tasks based on the user's target command, enriching the application functions of the intelligent robot and meeting user needs.

[0083] In the embodiments of this application, further, as Figure 2 As shown, step 104 above may specifically include steps 104a and 104b as follows:

[0084] Step 104a: If the target instruction contains scene semantic information, determine the target sub-scene in the target scene based on the scene semantic information;

[0085] Step 104b: Move to the target sub-scene and perform object recognition in the target sub-scene.

[0086] In the above embodiments, during the process of the intelligent robot autonomously identifying the target object indicated by the target command within the target scene based on the target scene map, target recognition model, and user target command, the intelligent robot specifically analyzes the semantic information of the target command and determines the target sub-scene indicated by the target command. Based on this, the intelligent robot determines the movement route required to move to the target sub-scene based on the target scene map and obstacle information within the target scene, and controls its onboard movement device to move the entire robot to the target sub-scene according to the movement route. After the intelligent robot moves to the target sub-scene, it performs target recognition within the target sub-scene using its onboard recognition device to autonomously identify the target object indicated by the user command. In this way, without manual user control, the intelligent robot can autonomously move to the corresponding target sub-scene according to the target command issued by the user, and autonomously identify the target object indicated by the target command within the target sub-scene, enriching the application functions of the intelligent robot and meeting the user's needs.

[0087] For example, in the case of the aforementioned intelligent robot being a home service robot, when a user issues a voice command to the robot, "Go to the bedroom and call Dad for dinner," the robot receives the user's voice command and analyzes its semantic information to determine that the target object to be identified by the user is the person identified as "Dad" in the target biometric database. Simultaneously, the robot determines that the target sub-scene indicated by the user's command is the "bedroom." Based on this, the robot autonomously plans its movement route to the "bedroom" using a target scene map and controls its mobile device to move the entire robot towards the "bedroom" according to this route. After entering the "bedroom," the robot activates its identification and search function, performing object recognition within the "bedroom" based on a target recognition model to determine the specific location of "Dad." After determining the specific location of "Dad," the robot autonomously moves towards "Dad" based on the target tracking algorithm, the recognition results of the recognition device, and the recognition frame information of the aforementioned target recognition model. After recognizing the target person "Dad," the robot sends a voice message "Dad, it's time to eat!" to notify "Dad" to come to the dining room for dinner.

[0088] In the embodiments of this application, further, as Figure 3 As shown, step 104 above may specifically include steps 104c and 104d as follows:

[0089] Step 104c: If the target instruction does not contain scene semantic information, determine the first movement route based on the target scene map;

[0090] Step 104d: Perform object recognition within the target scene based on the first movement route.

[0091] In the above embodiments, the intelligent robot, based on a target scene map, a target recognition model, and the user's target command, autonomously identifies the target object indicated by the target command within the target scene using its onboard recognition device. Specifically, the intelligent robot analyzes the semantic information of the target command to determine whether it contains scene semantic information. If the intelligent robot determines that the user's target command does not contain scene semantic information, it activates the whole-house recognition and person-finding function and autonomously plans a person-finding route, namely the first movement route. Based on this, the intelligent robot controls its onboard movement device to move its entire body within the target scene according to the first movement route. During its movement, it uses its onboard recognition device to identify objects in the target scene environment it passes through, thereby identifying the target object indicated by the user command from multiple objects within the target scene. In this way, without manual user control, the intelligent robot can autonomously plan a person-finding route based on the user's target command and autonomously identify the target object indicated by the target command within the target scene according to the target command, enriching the application functions of the intelligent robot and meeting the user's needs.

[0092] For example, when the aforementioned intelligent robot is a home service robot, when a user issues a voice command to the robot, "Call Grandpa for dinner," the robot receives the user's voice command and analyzes its semantic information to determine that the target object to be identified is the person identified as "Grandpa" in the target biometric database. Simultaneously, the robot determines that the user's command does not contain scene semantic information. At this point, the robot activates its whole-house identification and search function and autonomously plans a search route based on the target scene map. Further, the robot moves within the target scene according to the planned search route, using the biometric features identified as "Grandpa" in its pre-stored biometric database as a matching standard, and performs object recognition within the target scene through a target recognition model to determine the specific location of "Grandpa." After determining the specific location of "Grandpa," the robot autonomously moves towards "Grandpa" based on the target tracking algorithm, the recognition results of the recognition device, and the recognition box information of the aforementioned target recognition model. Further, after the robot recognizes the target person "Grandpa," it sends a voice message "Dinner's ready!" to the target person "Grandpa" to notify him to come to the dining room for dinner.

[0093] In the embodiments of this application, further, as Figure 4 As shown, step 104c above may specifically include the following step 104c1:

[0094] Step 104c1: If the target instruction does not contain scene semantic information, traverse the target scene map to determine the first movement route, or determine at least one identification point based on the target scene map and determine the first movement route based on the at least one identification point.

[0095] In the above embodiments, during the process of the intelligent robot autonomously planning its person-finding route, i.e., the first movement route, based on the target scene map, the intelligent robot specifically performs a traversal analysis of the target scene map pre-stored in its storage device to analyze and process scene information such as obstacle information and road information in the target scene. Based on this, the intelligent robot plans a person-finding route, i.e., the first movement route, according to the traversal results of the target scene map, so that when the intelligent robot moves according to the planned first movement route, it can traverse the aforementioned target scene. In this way, the intelligent robot can identify the target object by traversing the entire target scene, achieving comprehensive and accurate detection of the target scene, ensuring the comprehensiveness and accuracy of the identification results, and improving the intelligent robot's person-finding performance.

[0096] In this technical solution, further, during the process of the intelligent robot autonomously planning its search route (i.e., the first movement route) based on the target scene map, specifically, the intelligent robot can also analyze and process the target scene map pre-stored in its storage device to autonomously plan at least one detection point within the target scene. Based on this, after planning at least one detection point, the intelligent robot then autonomously plans a search route (i.e., the first movement route) based on the target scene map, the specific location coordinates of each detection point, and the obstacle information of the target scene. This first movement route passes through each of the planned detection points.

[0097] It should be noted that the aforementioned detection points are located in relatively open areas within the target scene. In practical applications, these detection points can be located at the relative center of each room. That is, when the intelligent robot performs 360° rotational detection at each detection point using its onboard recognition device, there are no obstacles obstructing the recognition device's path, meaning the device can perform a 360° all-around scan without blind spots at each detection point. Based on this, by performing rotational detection at each detection point within the target scene, the intelligent robot can achieve comprehensive detection of the target scene without blind spots, thus ensuring the completeness of the intelligent robot's autonomous identification and search capabilities.

[0098] In this embodiment of the application, the above-described object recognition method is further applied to a robot, and the method may further include the following steps 126 and 128:

[0099] Step 126: Receive the user's target instruction;

[0100] Step 128: Based on the target location, move to the target location and perform the target task on the target object;

[0101] Among them, the target instruction is used to instruct the robot to perform the target task on the target object.

[0102] In the above embodiments, the target instruction not only instructs the robot to identify and locate the target object within the target scene, but also instructs the robot to perform a target task on the identified target object, such as a relaying task or a delivery task. Based on this, after the robot identifies the target object within the target scene according to the user-issued target instruction, thereby determining the target location of the target object, the robot moves to that location. Upon reaching the target object, it interacts with the target object based on the target task indicated by the target instruction to complete the target task. In this way, without manual user control, the intelligent robot can autonomously identify and locate the target object indicated by the user-issued target instruction within the target scene and perform the corresponding target task on the target object, enriching the application functions of the intelligent robot and meeting the user's needs.

[0103] For example, if the user's target instruction to the intelligent robot is a message relay command, the intelligent robot, after recognizing the target object, automatically plays the voice message the user wants to convey to that object. Alternatively, the aforementioned intelligent robot can be equipped with an item storage and retrieval device. Building upon this, if the user's target instruction to the intelligent robot is an object transfer command, after receiving the user's target instruction, the intelligent robot automatically rotates to bring its item storage and retrieval device closer to the user, allowing the user to place the desired item within the device. Furthermore, after recognizing the target object, the intelligent robot automatically rotates to bring its item storage and retrieval device closer to the target object, allowing the target object to retrieve the item placed within the device. Thus, in situations where the user is too busy to personally deliver items or messages, the user can simply rely on the intelligent robot to deliver the desired item or voice message to the target object, fulfilling the user's needs.

[0104] In the embodiments of this application, further, as Figure 5 As shown, step 106 above may specifically include steps 106a and 106b as follows:

[0105] Step 106a: Determine at least one facial feature based on the recognition result;

[0106] Step 106b: If the first facial feature and the target facial feature are successfully matched, the first identification object corresponding to the first facial feature is identified as the target object, and the target location of the target object is determined.

[0107] Among them, the target facial features correspond to the target object.

[0108] In the above embodiments, during the process of the intelligent robot determining the specific location information of the target object within the target scene, i.e., the target location, based on the recognition results of its onboard recognition device, the intelligent robot extracts facial features of at least one object within the target scene. The intelligent robot then matches the extracted facial features of each object with the target facial features of the target object pre-stored in its storage device. If any facial feature of any object, i.e., the aforementioned first facial feature, matches the aforementioned target facial feature, a successful match is achieved. The intelligent robot then identifies this first object as the target object and determines its specific location. In this way, the intelligent robot determines the target object based on the pre-stored target facial features and the matching results of the facial features of the identified objects, thereby determining the target object's specific location and ensuring the accuracy of the intelligent robot's autonomous identification and search for people.

[0109] Furthermore, in practical applications, intelligent robots can also identify and analyze the biometric features such as voice and body characteristics of various objects within the target scene, and determine the specific location of the target object based on the matching results between the identified biometric features and the preset target biometric features. Those skilled in the art can set the specific criteria for the intelligent robot to determine the aforementioned target object according to the actual situation, and no specific restrictions are imposed here.

[0110] In this embodiment of the application, the object identification method may further include the following steps 110 and 112:

[0111] Step 110: In response to the failure of the target facial feature to match the first facial feature, and the successful match of the facial feature in the buffer interval with the first facial feature, move at the first speed to continue target recognition;

[0112] Step 112: In response to the failure to match the first facial feature with the target facial feature and the facial features in the cached area, the first facial feature is stored in the cached area, and the target recognition continues to be performed by moving at the first speed.

[0113] In the above embodiments, during the process of the intelligent robot matching the facial features of each extracted object with the target facial features of the target object pre-stored in its storage device, if any facial feature of an object (i.e., the first facial feature) does not match the target facial feature (i.e., the matching fails), the intelligent robot will continue to compare and match the first object feature with the facial features temporarily cached in its cache interval. If there is a facial feature in the cache interval that matches the first facial feature (i.e., the matching is successful), it indicates that the first facial feature is not the facial feature of the target object to be identified, and the first facial feature has already been identified and analyzed. At this time, there is no need to analyze and process the first facial feature again, and the intelligent robot moves at the first speed to continue target identification in the target scene. However, if every facial feature in the cache interval does not match the first facial feature (i.e., the matching fails), it indicates that the first facial feature is not the facial feature of the target object to be identified, and the first facial feature is a newly appearing facial feature in the intelligent robot's field of view. At this time, the intelligent robot will temporarily cache the first facial feature and move at the first speed to continue target identification in the target scene. In this way, the intelligent robot controls its recognition speed based on the object recognition results, that is, controls the movement speed of the intelligent robot in the process of identifying and finding people, thereby improving the efficiency of the intelligent robot in identifying and finding people.

[0114] In this embodiment of the application, further, after the step of failing to match the first facial feature with the target facial feature and the facial features in the cache interval, the above object recognition method may specifically include the following steps 114 and 116:

[0115] Step 114: Obtain the target identifier in the recognition interface;

[0116] Step 116: If the facial feature information of the target identifier recognition interface changes, adjust the recognition angle according to the recognition box information of the target recognition model, and move according to the second speed.

[0117] In the above embodiments, if the first facial feature does not match the target facial feature (i.e., matching fails), and every facial feature in the cache interval does not match the first facial feature (i.e., matching fails), the intelligent robot will also acquire a target identifier from its recognition interface. This target identifier is used to indicate changes in facial feature information within the recognition interface. Specifically, the target identifier can be a trackID. If the trackID changes, it indicates that the facial feature information within the intelligent robot's recognition interface has changed; for example, a new facial feature appears in the recognition interface, or a facial feature disappears and then reappears. If the trackID does not change, it indicates that the facial feature information within the intelligent robot's recognition interface has not changed. Based on this, if the intelligent robot determines that the target identifier has not changed, that is, if the intelligent robot determines that the facial feature information within its recognition interface has not changed, the intelligent robot moves at a first speed without adjusting its recognition angle. When the intelligent robot determines that the target identifier has changed, that is, when the intelligent robot determines that the facial feature information in its recognition interface has changed, it is necessary to re-extract and analyze the facial features that appear in the recognition interface, namely the first facial feature. At this time, the intelligent robot moves at the second speed and adjusts its recognition angle according to the recognition box information of the target recognition model to track and recognize the newly appearing first facial feature.

[0118] The second speed is less than the first speed. That is, during the process of the intelligent robot recognizing objects in a target scene using its onboard recognition device, if the intelligent robot fails to recognize the target object's facial features and detects a change in the facial feature information on its recognition interface, the intelligent robot needs to re-extract and re-analyze the facial features appearing on the recognition interface. In this case, the intelligent robot will move slowly. Conversely, if the intelligent robot fails to recognize the target object's facial features and detects no change in the facial feature information on its recognition interface, the intelligent robot will move quickly. In this way, by adjusting the intelligent robot's recognition speed based on changes in the facial feature information on the recognition interface, the phenomenon of blurred recognition images is avoided, thereby reducing the false recognition rate of the recognition device, ensuring the accuracy and precision of the recognition device in identifying objects, and improving the efficiency of the intelligent robot's autonomous identification and search for people.

[0119] In addition, in practical applications, when the intelligent robot tracks and identifies the first newly appearing facial feature within its recognition interface, the intelligent robot can specifically use tracking algorithms based on IOU or Kalman filtering, without making specific restrictions here.

[0120] In this way, the intelligent robot adjusts its recognition speed and angle based on the object recognition results and changes in facial features on its recognition interface. This means adjusting the robot's movement speed and head posture during the person-finding process. This improves the image quality captured by the robot during its search and reduces the impact of head posture on the recognition results. In other words, while maintaining efficiency in finding people, it also improves the image quality captured during object recognition, thereby increasing the success rate of the intelligent robot in identifying and finding people while in motion.

[0121] In this embodiment of the application, the step of adjusting the recognition angle based on the recognition box information of the target recognition model may specifically include the following steps 118 to 124:

[0122] Step 118: Obtain the first facial recognition box corresponding to the first facial feature;

[0123] Step 120: Determine the corresponding first head recognition frame based on the first face recognition frame;

[0124] Step 122: Determine the first distance based on the first head recognition frame;

[0125] Step 124: Adjust the recognition angle based on the position information of the first facial recognition frame in the recognition interface and the first distance.

[0126] In the above embodiments, during the process of the intelligent robot adjusting its recognition angle based on the recognition box information of the target recognition model to track and recognize the newly emerging first facial feature, the intelligent robot acquires the first facial recognition box output by the target recognition model, which corresponds to the first facial feature. Based on this, the intelligent robot determines the first head recognition box with which it has a belonging relationship, i.e., determines the head recognition box of the first recognition object corresponding to the first facial feature, based on the recognition box information of the first head recognition box. Further, the intelligent robot determines the first distance between itself and the first recognition object based on the recognition box information of the first head recognition box, and then adjusts its recognition angle based on the position information of the first facial recognition box in the recognition interface and the determined first distance to track and recognize the first facial feature. In this way, the intelligent robot adjusts its recognition angle based on the recognition box information of the target recognition model it uses to achieve tracking and recognition, ensuring the accuracy of the intelligent robot's recognition of the first facial feature, thereby ensuring the accuracy of the intelligent robot's autonomous identification and person-finding.

[0127] In the process of determining the first head recognition frame with which it belongs based on the recognition frame information of the first face recognition frame, the intelligent robot can specifically perform a fast matching between the first face recognition frame and each head recognition frame in the recognition interface based on the intersection and union relationship and size relationship between the first face recognition frame and each head recognition frame in the recognition interface, so as to select the first head recognition frame that best matches the first face recognition frame from each head recognition frame in the recognition interface.

[0128] Furthermore, in practical applications, the intelligent robot can also directly determine the first distance between itself and the first object to be identified based on a regression algorithm, or it can directly determine the first distance between itself and the first object to be identified using a TOF camera. Those skilled in the art can choose the specific method for determining the aforementioned first distance according to the actual situation, and no specific restrictions are imposed here.

[0129] In this embodiment of the application, step 120 may further include steps 120a and 120b as follows:

[0130] Step 120a: Determine the target value between the first face recognition box and the head recognition box based on the size correspondence and area correspondence between the first face recognition box and each head recognition box output by the target recognition model;

[0131] Step 120b: Perform target matching on the target value and determine the first head recognition box based on the matching result;

[0132] The area correspondence includes the intersection area and the union area of ​​the first face recognition box and each head recognition box. The target value is used to indicate the degree of correlation between the first face recognition box and the head recognition box.

[0133] In the above embodiments, during the process of the intelligent robot determining the first head recognition frame with which it belongs based on the recognition frame information of the first face recognition frame, specifically, the intelligent robot acquires each head recognition frame output by the target recognition model it uses, and determines the target value between the first face recognition frame and each head recognition frame based on the area correspondence between the first face recognition frame and each head recognition frame, such as the intersection area and the union area, as well as the size correspondence between the first face recognition frame and each head recognition frame, to obtain at least one target value. This target value is used to indicate the degree of association between the first face recognition frame and the head recognition frame. On this basis, the intelligent robot then matches the at least one target value calculated above according to the target matching algorithm, and filters the head recognition frames according to the comparison results to determine the first head recognition frame that best matches the first face recognition frame. In this way, by analyzing the intersection and union relationships and size relationships between the first face recognition frame and the head recognition frame, the first head recognition frame with which it belongs is determined, ensuring the accuracy of the determination of the first head recognition frame, and thus ensuring the accuracy of object tracking and recognition.

[0134] The target matching algorithm described above can be a greedy matching algorithm. In practical applications, those skilled in the art can choose the specific type of target matching algorithm according to the actual situation, and no specific restrictions are imposed here.

[0135] In this embodiment of the application, step 120a may further include steps 120a1 to 120a3 as follows:

[0136] Step 120a1: Determine the first height ratio of the first face recognition box to each head recognition box output by the target recognition model, the second area of ​​the union region, and the first area of ​​the intersection region;

[0137] Step 120a2: If the first threshold is less than or equal to the first height ratio, the second threshold is less than or equal to the first area, and the area of ​​the head recognition box is greater than the corresponding second area, determine the first ratio between the first area and the second area, determine the first pixel distance between the second edge of the first face recognition box and the first edge of the head recognition box, and determine the second ratio between the first pixel distance and the height of the head recognition box.

[0138] Step 120a3: Determine the target value based on the second ratio and the first ratio.

[0139] In the above embodiments, in the process of determining the target value between the first face recognition frame and each head recognition frame, specifically, the intelligent robot determines the ratio between the height of the first face recognition frame and the height of each head recognition frame, respectively obtaining the corresponding first height ratio. Furthermore, the intelligent robot determines the first area of ​​the intersection region between the first face recognition frame and each head recognition frame, and simultaneously determines the second area of ​​the union region between the first face recognition frame and each head recognition frame.

[0140] Furthermore, based on the specific values ​​of the first area, second area, and first height ratio determined above, the intelligent robot performs a first screening of the head recognition boxes output by the target recognition model. Specifically, the intelligent robot filters out head recognition boxes with a first height ratio less than a first threshold, head recognition boxes with a first area less than a second threshold, and head recognition boxes with a second area greater than the area of ​​the corresponding head recognition box. For example, Figure 7 As shown, for the same object to be identified, the head recognition bounding box output by the target recognition model includes the face recognition bounding box. When the first height ratio is less than a first threshold, it indicates that the ratio between the height of the first face recognition bounding box and the height of the head recognition bounding box is small, meaning the first face recognition bounding box and the corresponding head recognition bounding box do not belong to the same object. When the first area is less than a second threshold, it indicates that the intersection area between the first face recognition bounding box and the head recognition bounding box is small, meaning the first face recognition bounding box and the corresponding head recognition bounding box do not belong to the same object. When the second area is greater than the area of ​​the corresponding head recognition bounding box, it indicates that the union area between the first face recognition bounding box and the head recognition bounding box exceeds the selection range of the head recognition bounding box, meaning the first face recognition bounding box and the corresponding head recognition bounding box do not belong to the same object.

[0141] Furthermore, in practical applications, those skilled in the art can set the specific values ​​of the first threshold and the second threshold according to the actual situation, such as setting the first threshold to 0.1 and the second threshold to 10, without making specific restrictions here.

[0142] Based on this, for the remaining head recognition boxes, the intelligent robot calculates a first ratio between the first area and the second area corresponding to each head recognition box, and calculates the first pixel distance between the second edge of the first face recognition box and the first edge of each head recognition box, and then calculates a second ratio between the first pixel distance and the height of the corresponding head recognition box. Further, the intelligent robot substitutes each calculated second ratio and its corresponding first ratio into the first formula for calculation, thereby obtaining at least one target value. In this way, after filtering the head recognition boxes based on the intersection and union relationship and size relationship between the first face recognition box and the head recognition boxes, and then analyzing the intersection and union relationship and size relationship between the first face recognition box and the filtered head recognition boxes, the degree of correlation between the first face recognition box and each of the head recognition boxes is determined. This ensures the accuracy of the correlation between the first face recognition box and the head recognition boxes while reducing the computational load and improving object recognition efficiency.

[0143] Specifically, the first edge can be the upper edge of each head recognition frame, and the second edge can be the lower edge of the first face recognition frame.

[0144] In practical applications, the first formula mentioned above can be specifically defined as follows:

[0145] score1=λ×ration_box1+β×ratio_h1

[0146] Where score1 represents the target value, ration_box1 represents the first ratio, ratio_h1 represents the second ratio, λ and β are constant coefficients, and λ+β=1.

[0147] In this embodiment of the application, step 124 may further include steps 124a and 124b as follows:

[0148] Step 124a: If the distance threshold is less than the first distance, adjust the recognition angle to the target angle;

[0149] Step 124b: If the distance threshold is greater than or equal to the first distance, adjust the recognition angle according to the location information.

[0150] In the above embodiments, during the process of adjusting the intelligent robot's recognition angle based on the position information of the first facial recognition frame in the recognition interface and the first distance between the intelligent robot and the first recognition object, the intelligent robot compares the determined first distance with a set distance threshold. If the set distance threshold is less than the first distance, it indicates that the distance between the intelligent robot and the first recognition object is large. In this case, the intelligent robot's recognition angle does not need to be adjusted in the pitch direction, and the first recognition object is within the intelligent robot's field of vision for object recognition. The intelligent robot directly adjusts its pitch recognition angle to the target angle, such as 0 degrees, for object recognition. Conversely, if the set distance threshold is greater than or equal to the first distance, it indicates that the distance between the intelligent robot and the first recognition object is small. In this case, the intelligent robot adjusts its pitch angle for object recognition based on the position information of the first facial recognition frame in the recognition interface. Simultaneously, the intelligent robot also adjusts its horizontal rotation angle for object recognition based on changes in the position information of the first facial recognition frame in the recognition interface. In this way, the tilt angle of the intelligent robot when performing object recognition can be flexibly adjusted based on the distance between the intelligent robot and the first object to be recognized, and the horizontal rotation angle of the intelligent robot when performing object recognition can be flexibly adjusted based on the position information of the first facial recognition box in the recognition interface. This ensures the flexibility of adjusting the recognition angle of the intelligent robot, thereby ensuring the flexibility and accuracy of the intelligent robot in object tracking and recognition, and ensuring the accuracy of the intelligent robot in autonomously identifying and finding people.

[0151] The specific value of the aforementioned distance threshold can be set by those skilled in the art according to the actual situation, for example, the distance threshold can be set to 3 meters, and no specific restrictions are made here.

[0152] In practical applications, the aforementioned intelligent robot may specifically include a chassis and a main body. The chassis is equipped with a moving device that can move or rotate the chassis and main body together. The main body includes a first main body and a second main body. The second main body may specifically be the robot's head. The second main body is equipped with a recognition device and can rotate or pitch independently relative to the first main body. Furthermore, during the process of adjusting the pitch angle when the intelligent robot is recognizing an object, the intelligent robot can adjust the pitch angle of the recognition device by independently driving the recognition device through its second main body (such as the robot's head), thereby adjusting the overall pitch angle of the intelligent robot.

[0153] Furthermore, during the process of the intelligent robot adjusting its horizontal rotation angle for object recognition based on changes in the position of the first facial recognition frame within the recognition interface, the intelligent robot can rotate independently via a movable device on its chassis; alternatively, it can rotate the recognition device independently via its second main body, such as the robot head; or it can rotate in conjunction with its second main body and the movable device on its chassis. Specifically, when the position of the first facial recognition frame within the recognition interface changes significantly, the intelligent robot adjusts its horizontal rotation angle by rotating its entire body via the movable device on its chassis. Conversely, when the position change of the first facial recognition frame within the recognition interface is minor, the intelligent robot adjusts its horizontal rotation angle by rotating the recognition device via its second main body. In practical applications, those skilled in the art can set the specific method for adjusting the horizontal rotation angle of the intelligent robot according to the actual situation, and no specific limitations are imposed here.

[0154] In the embodiments of this application, further, as Figure 6 As shown, step 108 above may specifically include steps 108a to 108d as follows:

[0155] Step 108a: Obtain the target face recognition box corresponding to the target object based on the recognition result;

[0156] Step 108b: Determine the corresponding target head recognition box based on the target face recognition box, and determine the corresponding target body recognition box based on the target head recognition box;

[0157] Step 108c: Determine the target distance based on the target head recognition box;

[0158] Step 108d: Based on the target distance and the target tracking algorithm, the target body recognition box is moved.

[0159] In the above embodiments, during the process of the intelligent robot moving towards the target object based on the target tracking algorithm, the recognition results of the recognition device, and the recognition box information of the target recognition model, the intelligent robot determines the target facial features of the target object according to the recognition results of the recognition device and obtains the target facial recognition box output by the target recognition model. This target facial recognition box corresponds to the aforementioned target facial features. Based on this, the intelligent robot determines the target head recognition box with which it belongs, and determines the target body recognition box with which it belongs, based on the recognition box information of the determined target head recognition box. Further, the intelligent robot determines the target distance between itself and the target object based on the recognition box information of the target head recognition box, and autonomously plans the movement route required to reach the target object's location based on the determined target distance and the aforementioned target tracking algorithm. Following this movement route, the intelligent robot moves towards the target object by tracking the determined target body recognition box, until it reaches the target object's location, at which point the intelligent robot stops moving.

[0160] In the process of the intelligent robot determining the target body recognition box that has a belonging relationship with the target head recognition box based on the recognition box information of the determined target head recognition box, the intelligent robot can specifically perform rapid matching between the target head recognition box and each body recognition box in the recognition interface based on the intersection and union relationship and size relationship between the target head recognition box and each body recognition box in the recognition interface, so as to filter out the target body recognition box that best matches the target head recognition box from each body recognition box in the recognition interface.

[0161] Specifically, the intelligent robot acquires the body recognition bounding boxes output by its target recognition model. Based on this, the intelligent robot determines the ratio between the height of the target head recognition box and the height of each body recognition box, obtaining the corresponding second height ratio. Furthermore, the intelligent robot determines the third area of ​​the intersection region between the target head recognition box and each body recognition box, and simultaneously determines the fourth area of ​​the union region between the target head recognition box and each body recognition box.

[0162] Furthermore, based on the specific values ​​of the third area, fourth area, and second height ratio determined above, the intelligent robot performs a first screening of the body recognition boxes output by the target recognition model. Specifically, the intelligent robot filters out body recognition boxes whose second height ratio is less than the third threshold, whose third area is less than the fourth threshold, and whose fourth area is greater than the corresponding body recognition box area. For example, Figure 7As shown, for the same object to be identified, the body recognition bounding box output by the target recognition model includes the face recognition bounding box. When the second height ratio is less than the third threshold, it indicates that the ratio between the height of the target head recognition bounding box and the height of the body recognition bounding box is small, meaning the target head recognition bounding box and the corresponding body recognition bounding box do not belong to the same object. When the third area is less than the fourth threshold, it indicates that the intersection area between the target head recognition bounding box and the body recognition bounding box is small, meaning the target head recognition bounding box and the corresponding body recognition bounding box do not belong to the same object. When the fourth area is greater than the area of ​​the corresponding body recognition bounding box, it indicates that the union area between the target head recognition bounding box and the body recognition bounding box exceeds the selection range of the body recognition bounding box, meaning the target head recognition bounding box and the corresponding body recognition bounding box do not belong to the same object.

[0163] In addition, in practical applications, those skilled in the art can set the specific values ​​of the third and fourth thresholds according to the actual situation, such as setting the third threshold to 0.1 and the fourth threshold to 10, without making specific restrictions here.

[0164] Based on this, for the remaining body recognition boxes, the intelligent robot calculates the third ratio between the third and fourth areas corresponding to each body recognition box, and calculates the second pixel distance between the fourth edge of the target head recognition box and the third edge of each body recognition box. Then, it calculates the fourth ratio between the second pixel distance and the height of the corresponding body recognition box. Further, the intelligent robot substitutes each calculated fourth ratio and its corresponding third ratio into the second formula to calculate at least one target score. Based on this, the intelligent robot then matches the calculated at least one target score according to the target matching algorithm, and performs a second screening of the body recognition boxes based on the comparison results to determine the target body recognition box that best matches the target head recognition box.

[0165] Specifically, the third edge can be the upper edge of each body recognition box, and the fourth edge can be the lower edge of the target head recognition box.

[0166] In practical applications, the second formula mentioned above can be specifically defined as follows:

[0167] score2=a×ration_box2+b×ratio_h2

[0168] Where score2 represents the target score, ration_box2 represents the third ratio, ratio_h2 represents the fourth ratio, and a and b are constant coefficients, with a+b=1.

[0169] In one embodiment of this application, an object recognition device is also proposed. For example... Figure 8 As shown, Figure 8 A structural block diagram of an object recognition device 800 provided in an embodiment of this application is shown. The object recognition device 800 includes:

[0170] Memory 802, on which programs or instructions are stored;

[0171] The processor 804 executes the above-described program or instructions to implement the steps of the object recognition method as described in any of the above embodiments.

[0172] The object recognition device 800 provided in this embodiment includes a memory 802 and a processor 804. When the program or instructions in the memory 802 are executed by the processor 804, they implement the steps of the object recognition method as described in any of the above embodiments. Therefore, the object recognition device 800 has all the beneficial effects of the object recognition method in any of the above embodiments, which will not be repeated here.

[0173] Specifically, the memory 802 and the processor 804 can be connected via a bus or other means. The processor 804 may include one or more processing units, and the processor 804 may be a chip such as a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).

[0174] In one embodiment of this application, a robot is also proposed. For example... Figure 9 As shown, Figure 9 A structural block diagram of a robot 900 provided in an embodiment of this application is shown. The robot 900 includes an object recognition device 800, a chassis 902, and a main body 904, as described in the above embodiment.

[0175] The robot 900 proposed in the third aspect of this application includes the object recognition device 800 in the second aspect embodiment described above. Therefore, the robot 900 proposed in the third aspect of this application possesses all the beneficial effects of the object recognition device 800 in the fourth aspect embodiment described above, which will not be repeated here.

[0176] Furthermore, the robot 900 also includes a chassis 902 and a main body 904. The chassis 902 is equipped with a moving device 906, which can move or rotate the chassis 902 and the main body 904 together. The main body 904 is equipped with a recognition device 908. During the operation of the robot 900, the object recognition device 800 can control the recognition device 908 to perform object recognition, and control the moving device 906 to move or rotate based on the recognition result of the recognition device 908, thereby adjusting the recognition angle of the recognition device 908 and realizing the tracking and recognition function of the robot 900.

[0177] Furthermore, in practical applications, the aforementioned robots 900 include, but are not limited to: sweeping robots, lobby robots, child companion robots, airport intelligent robots, intelligent question-and-answer robots, home service robots, and other robots with object recognition capabilities.

[0178] In the embodiments of this application, further, as Figure 9 As shown, the main body 904 of the robot 900 includes a first main body 910 and a second main body 912. The second main body 912 can specifically be the head of the robot 900. The second main body 912 is provided with a recognition device 908, and the second main body 912 can rotate or pitch independently relative to the first main body 910. Specifically, the second main body 912 can swing along the axis of the robot 900, and the second main body 912 can rotate around the axis of the robot 900.

[0179] Based on this, during the process of adjusting the pitch angle of the robot 900 when it recognizes objects, the robot 900 can specifically drive the recognition device 908 to perform pitch movement through its second main body 912, such as the head of the robot 900, so as to adjust the pitch recognition angle of the recognition device 908, thereby adjusting the pitch recognition angle of the robot 900.

[0180] Furthermore, during the process of adjusting the horizontal rotation angle of the robot 900 when recognizing an object, the robot 900 can rotate independently via the moving device 906 on its chassis 902; the robot 900 can also rotate independently via its second body 912 (such as the robot 900 head); or the robot 900 can rotate in conjunction with its second body 912 and the moving device 906 on its chassis 902. Specifically, when the position of the first facial recognition frame changes significantly in the recognition interface, the robot 900 adjusts its horizontal rotation angle by rotating its entire body via the moving device 906 on its chassis 902; conversely, when the position of the first facial recognition frame changes only slightly, the robot 900 adjusts its horizontal rotation angle by rotating the recognition device 908 via its second body 912. In practical applications, those skilled in the art can set the specific method for adjusting the horizontal rotation angle of the robot 900 according to the actual situation, and no specific limitations are imposed here.

[0181] In practical applications, the working system of Robot 900 can be specifically divided into a vision end and a control end. The vision end is responsible for controlling the movement speed and head posture of Robot 900 when it is searching for people. The control end uses semantic map information stored in Robot 900 to autonomously navigate to each family room to search for people.

[0182] In the working process of robot 900, specifically, such as Figure 10 As shown, after the user wakes up the robot 900 with their voice and issues a search command, the robot 900 initiates its autonomous movement and search function. Specifically, the camera on the robot 900's head is turned on to obtain video stream data. The visual end calls a single model to perform face-head-body detection and determine the association between the face, head, and body. Based on this, during the automatic navigation process of the robot 900 to find the person using semantic map information, the camera on the robot 900's head will initiate face recognition to extract and match facial features. The robot 900 will also control its movement speed, head posture, and proximity following based on the visual recognition results.

[0183] Specifically, Robot 900 first performs target tracking based on all facial targets in the video stream. Then, it extracts facial features from the video stream and matches these extracted features with those in the registered feature library and the temporary facial feature cache. If both matches fail, Robot 900 updates the temporary facial feature cache with the extracted features and checks if the trackID has changed. If the trackID remains unchanged, Robot 900 disables motion control and head pose control; if the trackID changes, it enables motion control and head pose control. Furthermore, if the extracted facial features fail to match only those in the registered feature library but successfully match those in the temporary facial feature cache, Robot 900 also directly disables control over motion speed and head pose. Finally, if the extracted facial features successfully match those in the registered feature library, Robot 900 initiates proximity and following control. Specifically, the robot 900 first finds the corresponding human target based on the matched face, then performs REID-based target tracking on the human target, and sends the distance estimated based on the head corresponding to the matched face to the robot control terminal to control the robot 900 to walk in front of the user, thereby realizing interactions such as handing over objects or relaying messages.

[0184] An embodiment of the fourth aspect of this application provides a readable storage medium. A program or instructions are stored thereon, which, when executed by a processor, implement the steps of the object recognition method as described in any of the above embodiments.

[0185] The readable storage medium provided in this application provides a program or instruction that, when executed by a processor, can implement the steps of the object recognition method as described in any of the above embodiments. Therefore, this readable storage medium possesses all the beneficial effects of the object recognition method in any of the above embodiments, which will not be elaborated further here.

[0186] Specifically, the aforementioned readable storage medium can include any medium capable of storing or transmitting information. Examples of readable storage media include electronic circuits, semiconductor memory devices, read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), flash memory, erasable ROM (EROM), magnetic tape, floppy disk, optical disk, hard disk, fiber optic media, radio frequency (RF) links, optical data storage devices, etc. Code segments can be downloaded via computer networks such as the Internet and intranets.

[0187] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance, unless otherwise expressly specified and limited. The terms "connection," "installation," and "fixing," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0188] In the description of this specification, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0189] Furthermore, the technical solutions of the various embodiments of this application can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this application.

[0190] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An object recognition method, characterized in that, The method includes: The target object to be identified is determined based on the target instructions; Perform object recognition within the target scene, and determine the target location of the target object in the target scene based on the recognition results; Based on the target location, move towards the target location; Determining the target location of the target object in the target scene based on the recognition result includes: At least one facial feature is determined based on the recognition result, wherein the target facial feature corresponds to the target object; The object recognition method further includes: In response to the failure of the first facial feature to match the target facial feature, and the successful match of the first facial feature with the facial feature in the cache interval, the target recognition continues to be performed by moving at the first speed; In response to the failure of the first facial feature to match the target facial feature and the facial features in the cache interval, the first facial feature is stored in the cache interval, and the target recognition is continued by moving at the first speed. After the first facial feature fails to match the target facial feature and the facial features in the cached interval, the object recognition method further includes: Obtain the target identifier from the recognition interface; When the target identifier indicates that the facial feature information of the recognition interface has changed, the recognition angle is adjusted according to the recognition box information of the target recognition model, and the movement is performed at the second speed. Wherein, the first speed is greater than the second speed.

2. The object recognition method according to claim 1, characterized in that, The step of determining the target location of the target object in the target scene based on the recognition result further includes: In response to a successful match between a first facial feature and a target facial feature among the at least one facial features, a first identification object corresponding to the first facial feature is identified as the target object, and the target location of the target object is determined.

3. The object recognition method according to claim 2, characterized in that, The step of adjusting the recognition angle based on the recognition bounding box information of the target recognition model includes: Obtain the first facial recognition box corresponding to the first facial feature; Determine the corresponding first head recognition frame based on the first face recognition frame; Determine the first distance based on the first head recognition frame; The recognition angle is adjusted based on the first distance and the position information of the first facial recognition frame in the recognition interface.

4. The object recognition method according to claim 3, characterized in that, The step of determining the corresponding first head recognition frame based on the first facial recognition frame includes: Based on the size and area correspondence between the first face recognition box and each head recognition box output by the target recognition model, the target value between the first face recognition box and the head recognition box is determined. The target value is matched, and the first head recognition box is determined based on the matching result; The area correspondence includes the intersection area and the union area of ​​the first face recognition box and each head recognition box, and the target value is used to indicate the degree of correlation between the first face recognition box and the head recognition box.

5. The object recognition method according to claim 4, characterized in that, The step of determining the target value between the first face recognition box and the head recognition box based on the size correspondence and area correspondence between the first face recognition box and each head recognition box output by the target recognition model includes: Determine the first height ratio of the first face recognition frame to each head recognition frame, the first area of ​​the intersection region, and the second area of ​​the union region; When the first height ratio is greater than or equal to the first threshold, the first area is greater than or equal to the second threshold, and the second area is equal to the area of ​​the corresponding head recognition frame, a first ratio between the first area and the second area is determined, and a second ratio between the first pixel distance from the first edge of the head recognition frame to the second edge of the first face recognition frame and the height of the head recognition frame is determined. The target value is determined based on the first ratio and the second ratio.

6. The object recognition method according to claim 3, characterized in that, The step of adjusting the recognition angle based on the first distance and the position information of the first facial recognition frame in the recognition interface includes: If the first distance is greater than the distance threshold, the recognition angle is adjusted to the target angle; If the first distance is less than or equal to the distance threshold, the recognition angle is adjusted according to the location information so that the first face recognition box is located in the target area of ​​the recognition interface.

7. The object recognition method according to any one of claims 1 to 6, characterized in that, The movement towards the target location based on the target location includes: Based on the recognition result, obtain the target facial recognition box corresponding to the target object; The target head recognition box is determined based on the target face recognition box, and the target body recognition box is determined based on the target head recognition box; Determine the target distance based on the target head recognition frame; Based on the target tracking algorithm and the target distance, the target body recognition box is tracked and moved until it reaches the target position.

8. The object recognition method according to claim 1, characterized in that, The object recognition within the target scene includes: If the target instruction contains scene semantic information, a target sub-scene in the target scene is determined based on the target instruction; Move to the target sub-scene and perform object recognition in the target sub-scene.

9. The object recognition method according to claim 1, characterized in that, The object recognition within the target scene includes: If the target instruction does not contain scene semantic information, a first movement route is determined based on the target scene map; Object recognition is performed within the target scene based on the first movement route.

10. The object recognition method according to claim 9, characterized in that, Determining the first movement route based on the target scene map includes: Traverse the target scene map and determine the first movement route based on the traversal results; or At least one detection point is determined based on the target scene map, and the first movement route is determined based on the at least one detection point.

11. The object recognition method according to claim 1, characterized in that, When applied to robots, the method further includes: Receive a target instruction from a user, the target instruction being used to instruct the robot to perform a target task on the target object; Based on the target location, move to the target location and perform the target task on the target object.

12. An object recognition device, characterized in that, include: Memory, which stores programs or instructions; A processor that, when executing the program or instructions, implements the steps of the object recognition method as described in any one of claims 1 to 11.

13. A robot, characterized in that, include: A chassis on which a moving device is mounted; The main body is mounted on the chassis, and an identification device is provided on the main body; The object recognition device as described in claim 12; The object recognition device controls the recognition device to perform object recognition, and the object recognition device controls the moving device to move or rotate.

14. The robot according to claim 13, characterized in that, The subject includes: The first main body is mounted on the chassis; The second body is disposed on the first body and rotatably connected to the first body. The identification device is disposed on the second body. The second body can swing along the axis of the robot and rotate about the axis of the robot.

15. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps of the object recognition method as described in any one of claims 1 to 11.