Object recognition method in dynamic scenes and its applicable intelligent robot

By establishing an object memory bank in an intelligent robot and dynamically update it, the problem of inaccurate object recognition in complex dynamic scenarios is solved, and the recognition accuracy and adaptability are improved.

CN119380252BActive Publication Date: 2025-05-13BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411961475.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In complex dynamic scenarios, intelligent robots are not accurate enough to recognize objects, resulting in task failure.

Method used

By obtaining target video and embodied sensing data, an object memory database is established, the perception data of each object is stored, and the memory database is dynamically updated when the scene changes, so as to achieve long-term tracking and identification of objects.

Benefits of technology

It improves the accuracy and adaptability of object recognition of intelligent robots in dynamic scenarios, ensuring efficient and accurate task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380252B_ABST
    Figure CN119380252B_ABST
Patent Text Reader

Abstract

The present application discloses an object recognition method in a dynamic scene and an intelligent robot applicable thereto, belonging to the field of artificial intelligence. The method comprises: obtaining a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame; based on at least one video frame and embodied sensor data, respectively determining the perception data of the intelligent robot for at least one object in the target scene; establishing an object memory library corresponding to the target scene, and storing the perception data of at least one object in the object memory library in separate entries; when the target scene changes, determining the target object related to the change in the target scene, and updating the object memory library based on the target object. The present application realizes persistent memory of objects in the scene, and enhances the intelligent robot's understanding and adaptability to dynamic scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to an object recognition method in a dynamic scene and an intelligent robot applicable thereto. Background Art

[0002] With the development of artificial intelligence technology and intelligent robots, the perception and understanding capabilities of intelligent robots for complex dynamic scenes have been significantly improved. However, there are still many challenges in efficiently managing and utilizing perception data in changing scenes, especially information related to scene changes.

[0003] In the prior art, intelligent robots detect objects in the scene by performing target detection on images or video frames, and then perform various tasks based on these objects, such as path planning, obstacle avoidance, or providing services to users. However, when objects in the scene move frequently or the scene changes dynamically, this will cause the intelligent robot to identify the objects inaccurately or even make mistakes, which will lead to task failure. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems existing in the related art. To this end, the present application proposes an object recognition method in a dynamic scene and an intelligent robot applicable thereto to improve the accuracy of object recognition.

[0005] In a first aspect, the present application provides an object recognition method in a dynamic scene, which is applied to an intelligent robot; the method comprises:

[0006] Acquire a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame;

[0007] Based on the at least one video frame and the embodied sensor data, respectively determining perception data of the intelligent robot for at least one object in the target scene;

[0008] Establishing an object memory library corresponding to the target scene, and storing the perception data of the at least one object in the object memory library;

[0009] When the target scene changes, a target object related to the target scene change is determined, and the object memory library is updated based on the target object.

[0010] In the above technical scheme, by acquiring the target video and the collected embodied sensor data, and performing multimodal data processing based on the target video and the embodied sensor data, each object in the target scene is identified, and the corresponding perception data is extracted, and then stored in an object memory library. By constructing an object memory library, persistent memory of objects in the scene is achieved, so that multimodal perception data can be stored and managed in a structured manner, and long-term tracking of objects can be achieved through dynamic maintenance, which is convenient for efficient execution of subsequent analysis and decision-making; when the target scene changes, the object memory library can be updated in real time based on the target objects related to the change, so that the intelligent robot can quickly identify and respond to changes in the target scene, thereby enhancing the adaptability of the intelligent robot in complex scenes.

[0011] According to an embodiment of the present application, when the target scene changes, determining a target object related to the target scene change, and updating the object memory based on the target object includes:

[0012] When an unknown object is identified in the target scene, re-identifying the unknown object to obtain a re-identification result;

[0013] If the re-identification result indicates that the unknown object is the same as any identified object, updating the perception data corresponding to the identified object in the object memory;

[0014] If the re-identification result indicates that the unknown object is not any identified object, determining the perception data corresponding to the unknown object;

[0015] The perception data of the unknown object is stored in the object memory bank.

[0016] In a second aspect, the present application provides an object recognition device in a dynamic scene, which is applied to an intelligent robot; the device comprises:

[0017] An acquisition module, used to acquire a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame;

[0018] A determination module, configured to determine, based on the at least one video frame and the embodied sensor data, the perception data of the intelligent robot for at least one object in the target scene;

[0019] A memory module, used to establish an object memory library corresponding to the target scene, and store the perception data of the at least one object in the object memory library;

[0020] The updating module is used to determine the target object related to the change of the target scene when the target scene changes, and update the object memory library based on the target object.

[0021] In the above technical scheme, by acquiring the target video and the collected embodied sensor data, and performing multimodal data processing based on the target video and the embodied sensor data, each object in the target scene is identified, and the corresponding perception data is extracted, and then stored in an object memory library. By constructing an object memory library, persistent memory of objects in the scene is achieved, so that multimodal perception data can be stored and managed in a structured manner, and long-term tracking of objects can be achieved through dynamic maintenance, which is convenient for efficient execution of subsequent analysis and decision-making; when the target scene changes, the object memory library can be updated in real time based on the target objects related to the change, so that the intelligent robot can quickly identify and respond to changes in the target scene, thereby enhancing the adaptability of the intelligent robot in complex scenes.

[0022] In a third aspect, the present application provides an intelligent robot, comprising an object recognition device in a dynamic scene as described in the second aspect.

[0023] In a fourth aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for object recognition in a dynamic scene as described in the first aspect above is implemented.

[0024] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the object recognition method in a dynamic scene as described in the first aspect above is implemented.

[0025] In a sixth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the object recognition method in a dynamic scene as described in the first aspect above.

[0026] In a seventh aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the object recognition method in a dynamic scene as described in the first aspect above.

[0027] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0029] Figure 1 is a schematic diagram of an application scenario of an object recognition method in a dynamic scene provided in some embodiments of the present application;

[0030] Figure 2 is a flowchart of an object recognition method in a dynamic scene provided in some embodiments of the present application;

[0031] Figure 3 It is a schematic diagram of the question-answering principle under the association of actions and object states provided in some embodiments of the present application;

[0032] Figure 4 is a schematic diagram of the principle of human-computer interaction provided by the present application in other embodiments;

[0033] Figure 5 is a schematic diagram of the overall principle of the object recognition method in a dynamic scene provided in some embodiments of the present application;

[0034] Figure 6 is a schematic diagram of the structure of an object recognition device in a dynamic scene provided in some embodiments of the present application;

[0035] Figure 7 It is a schematic diagram of the structure of a computer device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0037] Unless otherwise defined, all technical and scientific terms used in this application have the same meanings as those commonly understood by technicians in the technical field of this application; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned drawings and any variations thereof are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order or a primary and secondary relationship.

[0038] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0039] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", and "attached" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0040] The term "and / or" in this application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this application generally indicates that the associated objects before and after are in an "or" relationship.

[0041] The term "multiple" as used in the present application refers to more than two (including two). Similarly, the term "multiple groups" refers to more than two groups (including two groups), and the term "multiple sheets" refers to more than two sheets (including two sheets).

[0042] With the development of artificial intelligence technology, the perception, understanding and task execution capabilities of embodied agents in complex scenarios have gradually become a research hotspot. In dynamic scenes, object recognition and manipulation through multimodal data such as vision, touch, and language are important foundations for achieving intelligence.

[0043] Currently, many embodied agents can collect video data through cameras and use deep learning algorithms to identify objects in the scene. However, the complexity of dynamic scenes (such as object movement, state changes, or new objects) places higher demands on the perception and operation of the agent. In addition, natural language interaction, as an important way of human-computer interaction, requires embodied agents to accurately understand user text descriptions and associate them with object information in the actual scene, so as to efficiently execute user instructions.

[0044] In view of this, the embodiment of the present application provides an object recognition method in a dynamic scene and an intelligent robot applicable thereto, which constructs an object memory corresponding to the target scene through multimodal perception technology. The memory records the perception data of all objects in the scene, including object categories, state descriptions, three-dimensional boundary data, positional relationships, visual features, etc. For newly added or changed objects in dynamic scenes, the integrity and accuracy of the memory are dynamically maintained through real-time re-identification and state update mechanisms, so that the embodied intelligent body can maintain high task execution efficiency and accuracy.

[0045] Furthermore, based on the constructed object memory library, the embodiment of the present application can provide an object recognition and operation method based on natural language interaction. The embodied intelligent agent can respond to the user's text description, retrieve the target object and its perception data matching the description in the object memory library, and execute the user's instructions in combination with the operation requirements. For example, if the user specifies "the blue water cup on the table" in the command, the intelligent agent can quickly locate the target object and complete the indicated task.

[0046] Among them, the intelligent robot involved in the embodiments of the present application is, for example, an embodied intelligent agent. An embodied intelligent agent refers to an intelligent agent equipped with artificial intelligence and having the ability to perceive and interact with the outside world, which can interact with the environment in real time through perception and interaction, and can perform various tasks in a virtual environment or a real physical environment. An intelligent robot can be a robot with a physical entity, etc.

[0047] The embodiment of the present application also proposes an embodied video agent (Embodied VideoAgent), which can process videos to explore scenes, make decisions and plans, and perform actions.

[0048] The object recognition method, device, computer equipment and storage medium in dynamic scenes provided in the embodiments of the present application can be applied in various fields, such as smart home, security monitoring, smart assistant, automatic driving and games.

[0049] The object recognition method in a dynamic scene provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0050] The object recognition method in a dynamic scene provided by the embodiment of the present application can be applied to Figure 1In the application scenario shown. Among them, the embodied intelligent body 10 is placed in a scene, which can be a real physical scene or a virtual scene, such as a simulator scene, a game scene, etc. The embodied intelligent body 10 can be provided with a camera and a depth sensor, etc., for detecting the external scene. The embodied intelligent body 10 can explore the scene by itself to identify and remember the objects in the scene. In some embodiments, the embodied intelligent body 10 can also analyze based on a given video recording the scene to identify and remember the objects in the scene. For example, the embodied intelligent body 10 senses the objects in the target scene by acquiring the target video and the collected embodied sensor data, and obtains the perception data of at least one object. Then, an object memory library corresponding to the target scene is established, and the perception data of at least one object is stored in the object memory library by entries. In this way, the initial recognition and memory of each object in the scene are achieved. Of course, the embodied intelligent body 10 can also explore the environment by moving, and continuously take images / record videos in the process, and then identify and memorize objects based on each image / video frame. When the target scene changes, for example, the embodied intelligent agent 10 moves, causing the video frame captured by the camera to change, or the object in the target scene is moved, the embodied intelligent agent 10 determines the target object related to the target scene change and updates the object memory based on the target object.

[0051] The object recognition method in a dynamic scene provided by the embodiment of the present application can be executed by an intelligent robot or a functional module or functional entity in the intelligent robot that can implement the object recognition method in a dynamic scene. The object recognition method in a dynamic scene provided by the embodiment of the present application is described below by taking an intelligent robot as an example of the execution subject.

[0052] like Figure 2 As shown, the object recognition method in a dynamic scene includes: steps 210 to 240.

[0053] Step 210: Acquire a target video and collected embodied sensor data; the target video is used to present a target scene, and includes at least one video frame.

[0054] The target video is the video data used to present the target scene, which is usually composed of multiple consecutive frames of images. A video frame is the basic unit of a video, and each frame is a static image, representing a picture captured by the video at a certain moment.

[0055] In the embodiment of the present application, the target video provides visual information of the target scene. By processing multiple video frames, it is possible to extract information such as object features, spatial layout, dynamic changes, etc. in the scene, thereby assisting the intelligent robot in environmental understanding and decision-making.

[0056] Exemplarily, the target video may be a first-person video. The intelligent robot may be an embodied video agent.

[0057] The target scene refers to the spatial range recorded by the target video and perceived by the intelligent robot, including objects, structures and their relationships within the range. The target scene is the core background for the intelligent robot to analyze and make decisions. The intelligent robot builds an object memory based on the perception data in the target scene, thereby realizing dynamic monitoring of the scene and object management. The target scene can vary depending on the specific application, for example, it can be a scene in a room, a section of road, or a storage area.

[0058] Embodied sensor data refers to the sensory data collected by various sensors on intelligent robots, including but not limited to physical, chemical or spatial attribute information, such as distance, temperature, vibration, sound or smell, etc. Embodied sensor data can be combined with video data to provide multimodal data support, thereby improving the accuracy of object recognition and feature analysis in the target scene, and helping intelligent robots interact with the environment and perceive dynamic changes, such as detecting the movement of objects.

[0059] Exemplarily, the embodied sensing data includes depth data and camera 6D pose data. Depth data refers to information reflecting the distance between an object or surface in a scene and a sensor, usually expressing the depth value of each point in the form of pixels. Depth data can be obtained by a depth camera (such as LiDAR, ToF camera, or structured light camera), or calculated from multi-view images by a stereo vision algorithm. The camera 6D pose data describes the comprehensive information of the position (3 degrees of freedom) and direction (3 degrees of freedom) of the camera in three-dimensional space, including the three-dimensional position of the camera in a reference coordinate system, usually represented by a three-dimensional vector; and the posture of the camera in the reference coordinate system, that is, its shooting direction and rotation state, usually represented by a rotation matrix, Euler angles, or quaternions. The reference coordinate system is, for example, the world coordinate system.

[0060] The intelligent robot can obtain the target video by real-time shooting / recording with its own visual acquisition device, or by receiving a given video. The visual acquisition device includes but is not limited to a camera.

[0061] Exemplarily, the intelligent robot extracts a historically recorded video from a local storage space as a target video, downloads a video from the Internet as a target video, or receives a video transmitted by a user or other device as a target video, etc.

[0062] Step 220: Based on at least one video frame and the embodied sensor data, respectively determine the perception data of the intelligent robot for at least one object in the target scene.

[0063] The intelligent robot recognizes objects in the target scene based on the video frames in the target video and the embodied sensory data corresponding to each video frame, and extracts the sensory data of the objects in the target scene. The objects may be biological objects or non-biological objects, such as tables, refrigerators, cats, humans, or other robots.

[0064] The embodied sensor data corresponding to each video frame refers to, for example, the embodied sensor data collected at the time corresponding to the corresponding video frame.

[0065] For example, in an indoor environment, an intelligent robot can identify objects such as tables, chairs, and cups through video frames, and combine video frames and embodied sensor data to further identify the positions and sizes of these objects, as well as the positional relationships between these objects, for example, the chair is next to the table and the cup is placed on the table.

[0066] Step 230: Establish an object memory library corresponding to the target scene, and store the perception data of at least one object in the object memory library.

[0067] The intelligent robot associates this data with the corresponding objects and stores them in the object memory library in separate entries to achieve persistent memory of each object in the scene.

[0068] Among them, the object memory is a data structure dedicated to storing and managing information related to objects in the target scene. Its storage content includes the object's perception data (such as location, shape, size and other attributes) and the object's historical records (such as appearance time, interaction records, dynamic changes, etc.). The memory is usually organized in the form of entries, each entry corresponds to an object or a class of objects, and a unique identifier is assigned to each object.

[0069] Exemplarily, each entry includes but is not limited to the following fields: object identification, object category, object state description, positional relationship between objects, three-dimensional boundary data of the object, and visual data of the object.

[0070] Object identification is usually automatically assigned by intelligent robots or generated by the object's own attributes (for example, its color, shape, label and other features can be combined) to ensure the distinction of different objects in the scene and the management of corresponding memory, and to provide a reference for subsequent dynamic tracking and retrieval interaction of objects.

[0071] Object categories are classifications of object properties, usually based on the function, form or purpose of the object. In some embodiments, object categories can be automatically identified by a machine learning algorithm carried by an intelligent robot.

[0072] The state description of an object refers to a comprehensive expression of the state of the object at a certain point in time, which may include but is not limited to position changes, physical forms, etc. For example, the state description of a refrigerator may include "normal", "refrigerator door open" or "refrigerator door closed", and the state description of an object may include "on the table", "in hand" or "not in hand", etc. Exemplarily, the state description of an object may be set to "normal" by default at the initial time (for example, when it is first recognized), and may be updated according to real-time changes in the scene.

[0073] The positional relationship between objects refers to the spatial description of the relative positions of different objects in the target scene, including distance, orientation, contact relationship, etc. For example, object A is placed on the upper surface of object B, object C is located inside object D, or object E is located on the right side of object F and the distance is 1 meter, etc. The positional relationship between the objects can be recorded and stored in the form of a list, for example.

[0074] The three-dimensional boundary data of an object is an accurate description of the shape, size, and spatial occupancy range of the object in three-dimensional space, and is usually represented by a three-dimensional coordinate point cloud or voxel data.

[0075] The visual data of an object refers to the feature information of the object extracted from the image / video frame, including but not limited to the image / video frame or a part thereof (such as an image area where the object is located) and feature vectors.

[0076] Of course, it is not limited to the above data. Intelligent robots can store different types of perception data according to the needs of actual scenarios.

[0077] Step 240: When the target scene changes, a target object related to the target scene change is determined, and the object memory library is updated based on the target object.

[0078] Among them, the change of the target scene may be that the objects in the target scene have changed, such as an object has been moved, or a new object has been added; it may also be a change caused by the change of visual acquisition, such as the intelligent robot has moved, or the position of the visual acquisition device carried by it has changed.

[0079] When the target scene changes, the intelligent robot can collect new images / video frames and identify the objects in the new images / video frames, determine whether they are memorized objects or newly added objects, or update the existence status of memorized objects in the scene.

[0080] Specifically, the intelligent robot determines the target object associated with the change by comparing the newly collected data with the data in the object memory, and updates the object memory. For example, when a chair is detected to move from a corner of the room to the side of the table, the target object associated with the change is the chair, and the intelligent robot will update the position information of the chair to reflect the latest status.

[0081] For example, the intelligent robot can compare the continuous video frames in the target video, detect the difference between frames, identify the area in the scene that has changed, and thus determine the target object related to the change of the target scene. Alternatively, the intelligent robot can also analyze the changes in the relative position, posture or distance between objects through other sensor information such as depth data and camera posture data, so as to determine the target object related to the change of the target scene. For another example, the intelligent robot can determine which objects have changed their state due to the interaction by monitoring the interaction between objects (such as collision, contact, occlusion, etc.).

[0082] In some embodiments, the intelligent robot can understand and analyze the scene through machine learning or deep learning models, so as to identify dynamic objects in the scene and determine whether the perception data of these objects in the object memory has changed, such as whether its state description has changed.

[0083] The object recognition method in a dynamic scene provided by the embodiment of the present application obtains a target video and collected embodied sensor data, and performs multimodal data processing based on the target video and the embodied sensor data, identifies each object in the target scene, extracts corresponding perception data, and then stores the data through an object memory library. By constructing an object memory library, persistent memory of objects in the scene is achieved, so that multimodal perception data can be stored and managed in a structured manner, and long-term tracking of objects can be achieved through dynamic maintenance, which is convenient for efficient execution of subsequent analysis and decision-making; when the target scene changes, the object memory library can be updated in real time based on the target objects related to the change, so that the intelligent robot can quickly identify and respond to changes in the target scene, thereby enhancing the adaptability of the intelligent robot in complex scenes.

[0084] Among them, based on the fusion of multimodal data of video frames and embodied sensor data, the intelligent robot can effectively perceive and process the object information in the target scene, and accurately identify the object and construct its three-dimensional information, thereby achieving comprehensive perception and understanding of the object. To this end, in some embodiments, based on at least one video frame and embodied sensor data, the perception data of the intelligent robot for at least one object in the target scene is determined separately, including steps 310 to 340:

[0085] Step 310: for any video frame, identify at least one object in the video frame, and obtain feature data, attribute data and two-dimensional boundary data of the at least one object;

[0086] Step 320: Perform dimensionality upscaling on the two-dimensional boundary data using the embodied sensor data corresponding to the targeted video frame to obtain three-dimensional boundary data corresponding to at least one object;

[0087] Step 330: determining a positional relationship between at least one object based on the three-dimensional boundary data corresponding to at least one object;

[0088] Step 340: Based on the feature data, attribute data, three-dimensional boundary data and position relationship, the intelligent robot obtains perception data of at least one object in the target scene.

[0089] Among them, feature data includes visual features. Attribute data includes but is not limited to object identification, object category, object state description, etc. Two-dimensional boundary data is, for example, coordinate data of a two-dimensional detection box of an object. Specifically, for each video frame, the intelligent robot extracts features from it, detects the object included in the video frame, and extracts its feature data, attribute data, and two-dimensional boundary data through image processing algorithms, etc. Exemplarily, the intelligent robot can extract image features in the video frame through its onboard target detection algorithm, etc., and classify based on the image features, thereby obtaining the object identification and object category of the object in the video frame.

[0090] In order to more accurately identify the existence status of objects in the target scene, the intelligent robot also obtains the embodied sensor data corresponding to the video frame, including but not limited to depth maps and camera 6D pose data.

[0091] Furthermore, based on the embodied sensor data, the intelligent robot can map the position of the object in the two-dimensional image to the three-dimensional physical space, thereby determining the existence state of the object in the three-dimensional scene. Specifically, the intelligent robot uses the embodied sensor data to perform dimensional processing on the two-dimensional boundary data to obtain the three-dimensional boundary data of the object. The three-dimensional boundary data is, for example, the coordinate data of the three-dimensional detection box of the object.

[0092] In some embodiments, the intelligent robot can extract the two-dimensional detection frame of the object from the video frame through a machine learning model or a deep learning model. On this basis, combined with the depth map and the camera 6D pose data, the two-dimensional boundary data is combined with the depth information, and the pixel points in each two-dimensional boundary box are deep matched to obtain the three-dimensional space coordinates of each pixel. For example, the intelligent robot can map the two-dimensional coordinates to the three-dimensional space using the perspective projection formula through the intrinsic and extrinsic parameters of the camera and the depth value of each pixel. Furthermore, the intelligent robot can generate the three-dimensional boundary data of the object by calculating the position of each pixel point in the three-dimensional space.

[0093] Therefore, for the video frame, the intelligent robot can judge the spatial relationship between the object and other objects based on the three-dimensional boundary data of the object and the three-dimensional boundary data of other objects, such as the distance between object A and object B, the up and down position relationship between object C and object D, etc.

[0094] Finally, the intelligent robot can generate the perception data of the object based on the above data for memory and storage. For example, the intelligent robot can directly memorize the above data as perception data, or the intelligent robot can further process the above data, such as deduplication and noise reduction, and then memorize the processed data as perception data. For example, the three-dimensional boundary data obtained after dimensionality increase is often affected by noise or errors. The intelligent robot can process the three-dimensional boundary data through geometric optimization algorithms (such as point cloud fitting, boundary smoothing, etc.) to make it more consistent with the actual shape of the object, and so on.

[0095] In the above embodiments, by combining video frames and embodied sensor data and performing processing based on multimodal data, a more accurate and comprehensive perception of objects is achieved, thereby improving the intelligent robot's perception ability of space and objects in the space, enabling it to better understand and interact with its environment, thereby performing more intelligent tasks.

[0096] In the process of feature extraction from video frames, the intelligent robot can not only identify the position of each object in the image, but also fully understand the nature and environment of the target object through the combination of object features and context features, thereby perceiving the object more accurately.

[0097] To this end, in some embodiments, for any video frame, identifying at least one object in the targeted video frame and obtaining feature data of the at least one object includes steps 410 to 430:

[0098] Step 410: for any video frame, identifying the image region where at least one object in the targeted video frame is located;

[0099] Step 420: extract features from the image regions where at least one object is located, and obtain object features corresponding to the at least one object.

[0100] Step 430: extract features from the targeted video frame to obtain context features corresponding to at least one object;

[0101] Step 440: Obtain feature data of at least one object in the targeted video frame based on the object features and the context features.

[0102] Specifically, for any video frame, the intelligent robot uses an object detection model (such as YOLO or Faster R-CNN based on convolutional neural networks) to detect objects in the video frame and obtain at least one object and its image region in the video frame, such as a part of the video frame. In the image region where the object is located, the intelligent robot uses a visual feature extraction algorithm (such as ResNet, CNN, etc.) to extract the visual features of the object, which are called object features.

[0103] In addition to the features of the object itself, the intelligent robot will also extract features from the entire video frame to obtain contextual features related to the object. Contextual features can reflect the position of the object in the scene, etc. As a result, the intelligent robot can generate complete feature data of the target object based on object features and contextual features. These data combine the information of the object itself and its relationship information in the scene, which is convenient for subsequent storage, analysis and decision-making. In some embodiments, the intelligent robot uses the object features and contextual features as feature data of the object and stores them in the object memory.

[0104] In the above embodiment, by performing multi-level feature extraction and fusion of the visual information and scene context information of the object in the target video frame, the target object can be accurately identified and its rich feature data can be obtained, thereby improving the accuracy of object recognition; and, by combining object features with context features, the intelligent robot can more comprehensively perceive the object's ontological properties and its relative relationship in the scene, thereby improving the accuracy and robustness of object recognition.

[0105] It should be noted that context features can reflect the position and state of an object in the global context, and can provide reliable data support when the intelligent robot performs interactive tasks. For example, the intelligent robot can answer the user's questions such as "Where is object A" based on the object features and context features.

[0106] When the target scene changes, the intelligent robot can understand the changed scene based on the object memory library that has been built. When the target scene changes, the intelligent robot can not only identify unknown objects in time and classify them, but also update the existing data in the object memory library to ensure the accuracy and completeness of the object perception data. At the same time, relying on the object memory library, the intelligent robot can also accurately identify objects that have moved.

[0107] To this end, in some embodiments, when the target scene changes, a target object related to the target scene change is determined, and the object memory is updated based on the target object, including steps 510 to 540:

[0108] Step 510: when an unknown object is identified in the target scene, the unknown object is re-identified to obtain a re-identification result;

[0109] Step 520: if the re-identification result indicates that the unknown object is the same as any identified object, the perception data corresponding to the identified object in the object memory is updated;

[0110] Step 530: If the re-identification result indicates that the unknown object is not any identified object, determine the perception data corresponding to the unknown object;

[0111] Step 540: Store the perception data of the unknown object into the object memory.

[0112] When the target scene changes, if the intelligent robot detects an unknown object, it first uses re-identification technology to determine whether it matches the identified object. If the match is successful, it means that the unknown object is an identified object, and its position, shape, etc. may have changed. The intelligent robot then updates the perception data of the identified object in the object memory library; if the match fails, it means that the unknown object may be a newly added object, and the perception data of the object is added as a new entry to the object memory library. Therefore, through the dynamic update mechanism, the intelligent robot can adapt to changes in the target scene and continue to maintain a comprehensive perception of the scene.

[0113] Specifically, the intelligent robot can identify the set of objects in the target scene through the target detection algorithm. If an unknown object (i.e., an object that is not in the object memory) is detected, the re-identification process is triggered. For example, there was originally a red cup on the table in the target scene. When someone puts a green cup on the table, the intelligent robot identifies the red cup and the green cup, and determines that the red cup is a recognized object, while the green cup is an unknown object.

[0114] For unknown objects, the intelligent robot can re-identify the unknown object by judging the similarity between the visual features and spatial position of the unknown object and the identified objects, thereby determining whether the unknown object is an identified object or a new object.

[0115] If the re-identification result shows that the unknown object is the same as an identified object, the perception data of the object in the object memory is updated, such as the object's position, state or other attribute information. For example, if the intelligent robot recognizes that the green cup matches a memorized green cup, the position and state description of the green cup will be updated, such as its position is moved from the refrigerator to the table.

[0116] If the re-identification result shows that the unknown object does not match the object in the memory bank, it is added as a new entry to the object memory bank, and its perception data is stored at the same time, including feature data, attribute data, three-dimensional boundary data, and the positional relationship between the object and other objects.

[0117] In the above embodiment, through real-time recognition and updating, the intelligent robot can quickly respond to the addition, removal or state change of objects in the target scene, ensuring the accuracy of the perceived data; and, by introducing a re-recognition algorithm, the intelligent robot can distinguish between known objects and unknown objects, avoid repeated recording or omission of important information in the scene, and achieve persistent object memory and accurate tracking of the object state.

[0118] In the process of re-identifying objects, the embodiment of the present application also improves the accuracy of unknown object re-identification by constructing a re-identification mechanism of stereo similarity comparison. In addition, for unknown objects in the target scene, the identity of the unknown object and the identified object can be efficiently determined through differentiated processing strategies under static and dynamic scene conditions. To this end, in some embodiments, the unknown object is re-identified to obtain a re-identification result, including steps 610 to 650:

[0119] Step 610, obtaining first three-dimensional boundary data of an unknown object, and extracting second three-dimensional boundary data corresponding to each identified object from an object memory library;

[0120] Step 620: Based on the first three-dimensional boundary data and each second three-dimensional boundary data, a stereo similarity comparison is performed to obtain a stereo similarity result;

[0121] Step 630: determine whether the unknown object is a static object or a dynamic object;

[0122] Step 640: When the unknown object is a static object, if the stereo similarity result satisfies the first similarity condition, a first recognition result is obtained;

[0123] Step 650: When the unknown object is a dynamic object, if the stereo similarity result satisfies a second similarity condition, a second recognition result is obtained; wherein the first recognition result and the second recognition result indicate that the unknown object is the same object as a recognized object.

[0124] That is, the intelligent robot obtains the first three-dimensional boundary data of the unknown object and compares it with the second three-dimensional boundary data in the object memory library to match the object. Specifically, the matching degree of the unknown object is judged by the stereo similarity of the three-dimensional shape. In addition, the intelligent robot also uses different similarity conditions for re-identification according to the dynamic properties of the object (static object or dynamic object). Therefore, in the process of re-identifying the object, not only the shape characteristics of the object are considered, but also its dynamic properties are considered, which can enhance the adaptability of the specific intelligent agent to the changing objects in the dynamic scene.

[0125] Specifically, the intelligent robot obtains the three-dimensional boundary data of the unknown object. The specific steps can refer to the above embodiment. For the convenience of distinction, the three-dimensional boundary data of the unknown object is called the first three-dimensional boundary data, and the three-dimensional boundary data of the identified object in the object memory is called the second three-dimensional boundary data. Then the intelligent robot compares the first three-dimensional boundary data with each second three-dimensional boundary data in three-dimensional and four-dimensional order to obtain a similarity result.

[0126] In addition, the intelligent robot also determines whether the unknown object is a static object or a dynamic object. In some embodiments, the intelligent robot can observe the historical motion trajectory of the position object based on the historical multi-frame video frame, and determine whether it is a static object or a dynamic object. For example, when the intelligent robot detects that the red cup has not changed position in the past 10 seconds, it determines that it is a static object.

[0127] For unknown objects with different dynamic attributes, the intelligent robot makes judgments based on different similarity conditions. That is, if the unknown object is a static object, the intelligent robot determines whether the stereo similarity meets the first similarity condition. If so, a first recognition result is obtained, which indicates that the unknown object is the same object as a recognized object.

[0128] If the unknown object is a dynamic object, the intelligent robot determines whether the stereo similarity satisfies a second similarity condition. If so, a second recognition result is obtained, which indicates that the unknown object is the same object as a recognized object.

[0129] It is easy to understand that if the unknown object is a static object and the stereo similarity does not meet the first similarity condition, the intelligent robot obtains a third recognition result, which indicates that the unknown object is not the same object as a recognized object. If the unknown object is a dynamic object and the stereo similarity does not meet the second similarity condition, the intelligent robot obtains a fourth recognition result, which indicates that the unknown object is not the same object as a recognized object.

[0130] Therefore, the intelligent robot updates the perception data in the object memory based on the re-identification results, including the position, state and other dynamic attributes. If any similarity conditions cannot be met, the unknown object is treated as a new object and its perception data is added to the object memory.

[0131] In the above embodiment, by comparing the stereo similarity based on the three-dimensional boundary data, it is possible to effectively distinguish unknown objects from identified objects to avoid misidentification or missed identification; and by introducing the classification processing of static and dynamic objects, the intelligent robot can have a deeper understanding of the changes and interaction relationships of objects in the scene, thereby improving the accuracy and robustness of object recognition.

[0132] In some embodiments, for static objects, the first similarity condition includes: the degree of overlap exceeds a first threshold, or the maximum inclusion ratio exceeds a second threshold and the object categories are the same.

[0133] The Intersection over Union (IoU) represents the overlap between the 3D bounding boxes of an unknown object and an identified object. The greater the overlap, the more likely it is that the unknown object and the identified object are the same object. For example, the overlap can be expressed by the following formula (1):

[0134] (1)

[0135] in, is the intersection of their three-dimensional bounding boxes, is the union of their 3D bounding boxes.

[0136] The Maximum Ratio of Intersection over Subsets (MaxIoS) represents the inclusion relationship between the 3D bounding box of the unknown object and the 3D bounding box of the identified object. When the two bounding boxes show a strong inclusion relationship, MaxIoS will be close to its maximum value of 1, which means that the unknown object is likely to be the same object as the identified object. For example, the maximum inclusion ratio can be expressed by the following formula (2):

[0137] (2)

[0138] in, represents the volume of the unknown object, Indicates that an object has been recognized.

[0139] Assumptions and are all detected as "table", where The volume is one tenth of The 3D bounding box of If the object is within the 3D bounding box, MaxIoS will reach 1, while IoU will be only 0.1. By introducing the maximum inclusion ratio and combining it with the object category for discrimination, it is possible to re-identify some observable objects in the case of occlusion. For example, may be Some observations (such as is part of the table), since they have overlapping bounding boxes and belong to the same object category.

[0140] Since objects are changing dynamically, their 3D detection frames may change significantly in space, and may be blocked by other objects at certain moments, so static overlap and inclusion are not applicable. Therefore, in some embodiments, for dynamic objects, the second similarity condition includes: the volume similarity exceeds the third threshold, or the visual features match.

[0141] Volume similarity (Bounding Box Volume Similarity, Vol_Sim) characterizes the volume similarity between the unknown object and the identified object. When two 3D bounding boxes have similar volumes, the value of Vol_Sim will be larger. The intelligent robot determines the volume of the unknown object and the volume of each identified object based on the first 3D boundary data and each second 3D boundary data, and determines the volume similarity. The volume similarity can be expressed by the following formula (3):

[0142] (3)

[0143] When the volume similarity between the unknown object and the volume of a recognized object exceeds a third threshold, it indicates that the unknown object and the recognized object are the same object.

[0144] If the visual features of the unknown object match the visual features of an identified object, the intelligent robot can also determine that the unknown object and the identified object are the same object.

[0145] Exemplarily, whether the visual features of the unknown object match the visual features of the identified object can be determined based on the respective object features of the two, such as comparing the similarity between the visual features.

[0146] In the above embodiment, by comparing the stereo similarity based on the three-dimensional boundary data, unknown objects and identified objects can be effectively distinguished, thereby improving the accuracy of object recognition and further enhancing the intelligent robot's understanding of object changes and interaction relationships in dynamic scenes; in addition, according to the scene requirements, the similarity conditions of static objects and dynamic objects can be dynamically adjusted to adapt to scenes of different complexities.

[0147] Among them, in some embodiments, determining whether an unknown object is a static object or a dynamic object includes: obtaining two-dimensional boundary data of the unknown object, and determining a first image area in the targeted video frame based on the two-dimensional boundary data; performing feature extraction on the first image area to obtain a first object feature corresponding to the unknown object; obtaining a previous video frame, and determining a second image area in the previous video frame that has the same position as the first image area; performing feature extraction on the second image area to obtain a second object feature; if the difference between the first object feature and the second object feature is less than a preset threshold, determining that the unknown object is a static object; if the difference between the first object feature and the second object feature is not less than a preset threshold, determining that the unknown object is a dynamic object.

[0148] That is, the intelligent robot locates the image area of ​​the unknown object in the current video frame and the previous video frame based on the two-dimensional boundary data, extracts the visual features of the image area, and compares the feature differences between the two frames. Then, the intelligent robot can determine whether the unknown object is static or dynamic based on the size of the difference and the preset threshold. As a result, the intelligent robot can quickly identify the dynamic properties of objects in the target scene, providing a reliable basis for the subsequent identification of unknown objects.

[0149] Specifically, the intelligent robot extracts the two-dimensional boundary data of the unknown object from the current video frame, such as a two-dimensional bounding box, and determines the first image area accordingly. The first image area is the pixel range where the unknown object is located in the current video frame. For example, in the current video frame, the xy coordinates of the upper left corner of the two-dimensional bounding box of the water bottle and the width and height of the two-dimensional bounding box are (200, 300, 50, 150) respectively, and the corresponding first image area is the pixel data within the two-dimensional bounding box.

[0150] The intelligent robot may, for example, extract the first object feature of the unknown object from the first image region through a feature extraction algorithm (such as a convolutional neural network).

[0151] In addition, the intelligent robot determines the second image area with the same position in the previous video frame according to the position of the first image area, and extracts the second object feature of the image area. For example, in the previous video frame, the intelligent robot also extracts features of the pixel data in the image area defined by the two-dimensional bounding box of (200, 300, 50, 150) to obtain the second object feature. Therefore, the intelligent robot determines whether the difference between the visual feature of the first object and the visual feature of the second object (such as Euclidean distance or cosine similarity, etc.) is less than a preset threshold. If the difference between the first object feature and the second object feature is less than the preset threshold, the intelligent robot determines that the unknown object is a static object; otherwise, it means that the posture of the object in two adjacent frames has changed significantly, and the intelligent robot determines that the unknown object is a dynamic object.

[0152] In the above embodiment, by comparing the visual features of the object image regions in the consecutive video frames, the dynamic properties of the object can be accurately determined, and then the object can be accurately re-identified, which has strong robustness.

[0153] When an unknown object is re-identified as an identified object, in some embodiments, the intelligent robot can update the three-dimensional boundary data, object features, and context features corresponding to the identified object in the object memory by means of moving average update, and re-determine the positional relationship between the identified object and other objects to update the positional relationship corresponding to the identified object in the object memory. The moving average update method can effectively smooth errors and eliminate noise in a single detection, ensuring the stability and reliability of object data.

[0154] In addition to the changes in the target scene in the above embodiments, the target scene may also change according to the actions of animals, humans or other robots in the scene, thereby causing the state of the object to change. For example, in a first-person video, when a user picks up an object or moves an object, the perceived data of the object will change due to the interactive action, so it is necessary to update the corresponding perceived data in the object memory. However, updating the state changes of an object caused by an interactive action is a key challenge, especially in the case of visual occlusion. Therefore, in some embodiments, the object recognition method in a dynamic scene provided by an embodiment of the present application also includes: when an action is detected in any video frame, the intelligent robot determines the object associated with the action and retrieves the perceived data corresponding to the object in the object memory. If the object associated with the action includes multiple objects belonging to the same object category, the intelligent robot determines the perceived data corresponding to each object respectively.

[0155] For each object, the intelligent robot renders the corresponding 3D boundary data in the object memory to the current frame, and uses the Vision Language Model (VLM) to determine whether the object defined by the 3D boundary data (e.g., the object within the 3D boundary box) is the target of the action. If so, the intelligent robot updates the perception data of the object in the object memory that is the target of the action. For example, if the object was originally placed on the table, when the action indicates that the object is picked up, the intelligent robot updates the state description of the object to "in hand", and so on.

[0156] Furthermore, when faced with a user's question-and-answer task regarding an action, the intelligent robot can answer based on the perception data corresponding to the object associated with the action. Figure 3 As shown in the figure, there are jars A and B placed in the target scene respectively. When the user asks "Which jar is currently picked up", the intelligent robot can determine the perception data corresponding to jars A and B associated with the action. For example, based on the three-dimensional boundary data corresponding to jars A and B, the intelligent robot can determine whether jars A and B are the targets of the action through the visual language model. Then, the intelligent robot outputs the corresponding answer, such as "The jar currently picked up is A". Therefore, through this visual prompt method, the action can be associated with the object and the object memory, the scene can be more accurately identified, and then a more accurate answer to the question can be output.

[0157] In order to further enhance the intelligent robot's understanding of dynamic scenes, in an embodiment of the present application, in addition to maintaining an object memory library (as a persistent object memory), the intelligent robot also sets up a temporary memory buffer for storing temporary memory.

[0158] In some embodiments, the temporary memory buffer includes an action buffer and an object buffer. The action buffer is used to record the action detected by the intelligent robot and the timestamp of the action, the name of the action, the object as the target of the action (for example, the object identifier can be recorded), and the visual data of the current frame when the action is detected (including object features and context features). The object buffer is used to record each recognized object and its recognition timestamp, object identifier, and three-dimensional boundary data. As a result, temporary memory can be quickly referenced in subsequent tasks, assisting scene understanding and task execution, and improving task execution efficiency.

[0159] In actual application scenarios, intelligent robots can interact with users and answer questions raised by users. For example, when a user asks, "Where is the milk?", the intelligent robot retrieves the object memory bank and answers, "The milk is in the refrigerator."

[0160] For another example, the intelligent robot can also execute instructions given by the user. For example, if the user gives the instruction to "put the milk on the table", the intelligent robot determines that the milk is inside the refrigerator by retrieving the object memory bank. It can then go to the location of the refrigerator and perform the actions of "opening the refrigerator door", "picking up the milk", and "closing the refrigerator door" in sequence, and then go to the location of the table and perform the action of "putting the milk on the table" to complete the task execution of the instruction.

[0161] Exemplarily, the pre-configured executable actions of the intelligent robot include, but are not limited to: one or more of: search, go to, open, close, pick up, or place.

[0162] To this end, in some embodiments, the object recognition method in a dynamic scene provided by the embodiments of the present application also includes: in response to a description text sent by a user for a target scene, determining a target object associated with the description text, and a target action corresponding to the target object; searching a pre-established object memory library corresponding to the target scene to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene; and based on the target perception data, performing a target action on the target object.

[0163] In dynamic scenes, users can interact with intelligent robots through natural language to efficiently specify task content. Users can input description text through voice or other means, and the description text is used to describe an interactive task, such as "pass me the blue water cup on the table" or "where is the blanket placed", etc. The former is mainly used to instruct the intelligent robot to perform specific actions to complete the task, while the latter is mainly used to instruct the intelligent robot to answer questions.

[0164] For example, Figure 4 As shown, the user can send instructions by voice or manually input description text. The intelligent robot analyzes the description text and determines the target object associated with the description text. For example, for the user's description text "What is on the microwave oven", the intelligent robot can determine that the target object is "microwave oven"; for another example, for the user's description text "Can you put the bowl by the sink", the intelligent robot can determine that the target objects are "bowl" and "sink". Then, the intelligent robot searches the object memory library and obtains the perception data corresponding to one or more target objects, which are called target perception data for distinction.

[0165] As a result, the intelligent robot can perform the target action on the target object based on the retrieved target perception data. For example, when the intelligent robot performs the interactive task of "putting the bowl next to the pool", it determines the location of each according to the perception data of the "bowl" and "pool" in the object memory, and performs path planning and navigation from the current location to the location of the "bowl". After that, the intelligent robot performs the "pick" operation on the object "bowl", moves it to the side of the "pool", and returns the corresponding task completion feedback to the user.

[0166] Furthermore, if the intelligent robot does not find the "bowl", the intelligent robot can also explore the surrounding environment to find the location of the "bowl" and update the perception data of the "bowl". Alternatively, the intelligent robot can also return the task execution result of "no bowl found" to the user.

[0167] In the above embodiment, the target object and target operation can be efficiently specified through natural language description, which reduces the complexity of user operations and enhances the user experience; and, by combining the user text description and the object memory library data, the intelligent robot can quickly locate the target object, and even if the scene changes dynamically or the scene is highly complex (such as a large number of objects), efficient and accurate recognition can be achieved.

[0168] By combining natural language text with multimodal data from the object memory library, the intelligent robot has stronger reasoning and response capabilities. Even if the text description is ambiguous, such as the user only indicates "the cup on the table", the intelligent robot can further resolve the ambiguity through the object's perception data. For example, if the object memory library shows that cup A is placed in the refrigerator, cup B is by the sink, and cup C is on the table, the intelligent robot can determine that the cup the user is referring to is cup C based on the object memory library.

[0169] In a specific example, Figure 5As shown, taking the intelligent robot as an embodied video agent as an example, the embodied video agent is equipped with a camera and a depth detection device, etc. The embodied video agent can combine the depth data, the camera 6D posture and the video data for multimodal processing, identify each object in the target scene, and extract the perception data of each object. For example, the embodied video agent performs dimensionality-upgrading processing on the two-dimensional bounding box obtained by object detection from the video frame by combining the depth data and the camera 6D posture data to obtain the three-dimensional bounding box of the object. The specific processing flow can refer to the above embodiment, which will not be repeated here. Thus, the embodied video agent stores the perception data of each object in the object memory. Exemplarily, the object memory is distinguished by the object identification (ID), and the state description (STATE) of the object is "normal" at the beginning. The state of the object can be updated synchronously based on the real-time detection of dynamic changes in the scene. For example, the state description of the object "microwave oven" is updated from "normal" to "closed" to indicate that the door of the microwave oven is in a closed state. In addition, the object memory also stores the positional relationship (RO) between objects, such as the roll of paper O 0Place in microwave O 1, the corresponding RO is, for example, (on, O 1); and microwave oven O 1Support roll paper O 0, the corresponding RO is, for example, (uphold, O 0). When an action is detected in the scene, the embodied video agent also updates the object memory based on the objects related to the action, thereby achieving accurate tracking of dynamic scenes or dynamic changes of objects.

[0170] The object recognition method in a dynamic scene provided in the embodiment of the present application can be executed by an object recognition device in a dynamic scene. In the embodiment of the present application, an object recognition device in a dynamic scene performs the object recognition method in a dynamic scene as an example to illustrate the object recognition device in a dynamic scene provided in the embodiment of the present application.

[0171] The present application also provides an object recognition device in a dynamic scene, which is applied to an intelligent robot. Figure 6 As shown, the object recognition device in the dynamic scene includes an acquisition module 601, a determination module 602, a memory module 603 and an update module 604. Among them:

[0172] The acquisition module 601 is used to acquire a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame.

[0173] The determination module 602 is used to determine the perception data of the intelligent robot for at least one object in the target scene based on at least one video frame and the embodied sensor data.

[0174] The memory module 603 is used to establish an object memory library corresponding to the target scene, and store the perception data of at least one object in the object memory library.

[0175] The updating module 604 is used to determine the target object related to the target scene change when the target scene changes, and update the object memory based on the target object.

[0176] According to the object recognition device in a dynamic scene provided by the embodiment of the present application, by acquiring the target video and the collected embodied sensor data, and performing multimodal data processing based on the target video and the embodied sensor data, each object in the target scene is identified, and the corresponding perception data is extracted, and then stored in an object memory library by entries. Thus, by constructing the object memory library, persistent memory of objects in the scene is achieved, so that multimodal perception data can be stored and managed in a structured manner, and long-term tracking of objects can be achieved through dynamic maintenance, which is convenient for efficient execution of subsequent analysis and decision-making; when the target scene changes, the object memory library can be updated in real time based on the target objects related to the change, so that the intelligent robot can quickly identify and respond to changes in the target scene, thereby enhancing the adaptability of the intelligent robot in complex scenes.

[0177] In some embodiments, the determination module is also used to identify at least one object in any video frame, and obtain feature data, attribute data and two-dimensional boundary data of at least one object; perform dimensionality upscaling on the two-dimensional boundary data through the embodied sensor data corresponding to the video frame, and obtain three-dimensional boundary data corresponding to at least one object; determine the positional relationship between at least one object based on the three-dimensional boundary data corresponding to at least one object; obtain the perception data of the intelligent robot for at least one object in the target scene based on the feature data, attribute data, three-dimensional boundary data and positional relationship.

[0178] In some embodiments, the determination module is also used to identify, for any video frame, the image areas where at least one object in the targeted video frame is located; perform feature extraction on the image areas where at least one object is located to obtain object features corresponding to the at least one object; perform feature extraction on the targeted video frame to obtain context features corresponding to the at least one object; and obtain feature data of at least one object in the targeted video frame based on the object features and the context features.

[0179] In some embodiments, the update module is also used to re-identify the unknown object when an unknown object is identified in the target scene to obtain a re-identification result; if the re-identification result indicates that the unknown object is the same object as any identified object, the perception data corresponding to the identified object in the object memory is updated; if the re-identification result indicates that the unknown object is not any identified object, the perception data corresponding to the unknown object is determined; and the perception data of the unknown object is stored in the object memory.

[0180] In some embodiments, the update module is also used to obtain first three-dimensional boundary data of the unknown object, and extract second three-dimensional boundary data corresponding to each identified object from the object memory; based on the first three-dimensional boundary data and each second three-dimensional boundary data, respectively perform stereo similarity comparison to obtain a stereo similarity result; determine whether the unknown object is a static object or a dynamic object; in the case where the unknown object is a static object, if the stereo similarity result satisfies a first similarity condition, obtain a first recognition result; in the case where the unknown object is a dynamic object, if the stereo similarity result satisfies a second similarity condition, obtain a second recognition result; wherein the first recognition result and the second recognition result characterize that the unknown object and an identified object are the same object.

[0181] In some embodiments, the update module is also used to obtain two-dimensional boundary data of the unknown object, and determine a first image area in the targeted video frame based on the two-dimensional boundary data; perform feature extraction on the first image area to obtain a first object feature corresponding to the unknown object; obtain a previous video frame, and determine a second image area in the previous video frame that has the same position as the first image area, and perform feature extraction on the second image area to obtain a second object feature; if the difference between the first object feature and the second object feature is less than a preset threshold, determine that the unknown object is a static object; if the difference between the first object feature and the second object feature is not less than a preset threshold, determine that the unknown object is a dynamic object.

[0182] The object recognition device in the dynamic scene in the embodiment of the present application can be a computer device, or a component in the computer device, such as an integrated circuit or a chip. The computer device can be an intelligent robot. Exemplarily, the computer device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted computer device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0183] The present application also provides an intelligent robot, which includes: Figure 6 An object recognition device in a dynamic scene is shown.

[0184] The object recognition device in a dynamic scene in the embodiment of the present application may be a device having an operating system. The operating system may be a Microsoft (Windows) operating system, an Android (Android) operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0185] The object recognition device in a dynamic scene provided by the embodiment of the present application can achieve Figure 2 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0186] In some embodiments, Figure 7 As shown, an embodiment of the present application also provides a computer device 700, including a processor 701, a memory 702, and a computer program stored in the memory 702 and executable on the processor 701. When the program is executed by the processor 701, each process of the above-mentioned method embodiments is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0187] It should be noted that the computer device in the embodiment of the present application includes the mobile computer device and the non-mobile computer device mentioned above.

[0188] The embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned object recognition method embodiment in the dynamic scene are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0189] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0190] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned object recognition method in a dynamic scene when executed by a processor.

[0191] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0192] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned object recognition method embodiment in the dynamic scene, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0193] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0194] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0195] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0196] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0197] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0198] If not otherwise specified, all embodiments and optional embodiments of the present application can be combined with each other to form a new technical solution.

[0199] Unless otherwise specified, all technical features and optional technical features of this application can be combined with each other to form a new technical solution.

[0200] If there is no special explanation, all steps of the present application can be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), which means that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, it is mentioned that the method may also include step (c), which means that step (c) can be added to the method in any order. For example, the method may include steps (a), (b) and (c), or may include steps (a), (c) and (b), or may include steps (c), (a) and (b), etc.

[0201] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for object recognition in a dynamic scene, characterized in that: Applied to intelligent robots; the method comprises: Acquire a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame; Based on the at least one video frame and the embodied sensor data, respectively determining the perception data of the intelligent robot for at least one object in the target scene; the perception data includes object identification, object category, state description of the object, positional relationship between objects, three-dimensional boundary data of the object, and visual data of the object; Establishing an object memory library corresponding to the target scene, and storing the perception data of the at least one object in the object memory library; When an unknown object is identified in the target scene, determining whether the unknown object is a static object or a dynamic object; Acquire first three-dimensional boundary data of the unknown object, and extract second three-dimensional boundary data corresponding to all identified objects from the object memory; When the unknown object is a static object, if the overlap degree of the first three-dimensional boundary data and any second three-dimensional boundary data exceeds a first threshold, or the maximum inclusion ratio exceeds a second threshold, and the object categories are the same, it is determined that the unknown object and an identified object are the same object; When the unknown object is a dynamic object, if the volume similarity between the first three-dimensional boundary data and any second three-dimensional boundary data exceeds a third threshold or the visual features match, it is determined that the unknown object is the same object as an identified object; The perception data corresponding to the recognized objects in the object memory that are the same object as the unknown object are updated.

2. The method according to claim 1, characterized in that The determining, based on the at least one video frame and the embodied sensor data, respectively, the perception data of the intelligent robot for at least one object in the target scene comprises: For any video frame, identifying at least one object in the targeted video frame, and obtaining feature data, attribute data and two-dimensional boundary data of the at least one object; Performing dimensionality upscaling processing on the two-dimensional boundary data using the embodied sensor data corresponding to the targeted video frame to obtain three-dimensional boundary data corresponding to the at least one object; Determining a positional relationship between the at least one object based on the three-dimensional boundary data respectively corresponding to the at least one object; Based on the feature data, the attribute data, the three-dimensional boundary data and the positional relationship, perception data of the intelligent robot for at least one object in the target scene is obtained.

3. The method according to claim 2, characterized in that The step of identifying at least one object in any video frame and obtaining feature data of the at least one object includes: For any video frame, identifying the image region where at least one object in the targeted video frame is located; Extracting features from the image regions where the at least one object is located to obtain object features corresponding to the at least one object; Performing feature extraction on the targeted video frame to obtain context features corresponding to the at least one object; Based on the object features and the context features, feature data of at least one object in the targeted video frame is obtained.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In the case where the unknown object is a static object, if the overlap degree of the first three-dimensional boundary data and the second three-dimensional boundary data does not exceed a first threshold, and does not satisfy the condition that the maximum inclusion ratio exceeds a second threshold and the object categories are the same, it is determined that the unknown object is not any identified object; In the case where the unknown object is a dynamic object, if the volume similarity between the first three-dimensional boundary data and the second three-dimensional boundary data does not exceed a third threshold, and the visual features do not match, it is determined that the unknown object is not any recognized object; determining sensory data corresponding to the unknown object; The perception data of the unknown object is stored in the object memory bank.

5. The method according to claim 1, characterized in that The determining whether the unknown object is a static object or a dynamic object includes: Acquire two-dimensional boundary data of the unknown object, and determine a first image area in the targeted video frame based on the two-dimensional boundary data; Performing feature extraction on the first image region to obtain a first object feature corresponding to the unknown object; Acquire a previous video frame, determine a second image region in the previous video frame that is located at the same position as the first image region, and perform feature extraction on the second image region to obtain a second object feature; If the difference between the first object feature and the second object feature is less than a preset threshold, determining that the unknown object is a static object; If the difference between the first object feature and the second object feature is not less than the preset threshold, the unknown object is determined to be a dynamic object.

6. An object recognition device in a dynamic scene, characterized in that: Applied to intelligent robots; the device comprises: An acquisition module, used to acquire a target video and collected embodied sensor data; the target video is used to present a target scene, including at least one video frame; A determination module, configured to determine, based on the at least one video frame and the embodied sensor data, the perception data of the intelligent robot for at least one object in the target scene; the perception data includes an object identifier, an object category, a state description of the object, a positional relationship between objects, three-dimensional boundary data of the object, and visual data of the object; A memory module, used to establish an object memory library corresponding to the target scene, and store the perception data of the at least one object in the object memory library; The updating module is used to determine whether the unknown object is a static object or a dynamic object when an unknown object is recognized in a target scene; obtain the first three-dimensional boundary data of the unknown object, and extract the second three-dimensional boundary data corresponding to all recognized objects from the object memory library; when the unknown object is a static object, if the overlap degree of the first three-dimensional boundary data and any second three-dimensional boundary data exceeds a first threshold, or the maximum inclusion ratio exceeds a second threshold and the object categories are the same, determine that the unknown object is the same object as a recognized object; when the unknown object is a dynamic object, if the volume similarity of the first three-dimensional boundary data and any second three-dimensional boundary data exceeds a third threshold or the visual features match, determine that the unknown object is the same object as a recognized object; and update the perception data corresponding to the recognized object in the object memory library that is the same object as the unknown object.

7. An intelligent robot, characterized in that: It comprises the object recognition device in a dynamic scene as described in claim 6.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the object recognition method in a dynamic scene as described in any one of claims 1-5 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the object recognition method in a dynamic scene as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Object tracking method, tracking model training method, device, equipment and medium

    CN115909173A

  • Large-model-driven intelligent agent scene exploration and memory management method and system

    CN117854059A

  • Video object detection method and device, electronic equipment and storage medium

    CN118644811A