Embodied robot and its environment perception method, device, storage medium, and computer program product

By obtaining pre-stored information through wireless communication modules and basic information from information acquisition sensors, combined with interaction order priority, the problems of high computing resource consumption and perception delay of embodied intelligent robots are solved, and efficient environmental perception is achieved.

CN120326637BActive Publication Date: 2025-09-05GREE ELECTRIC APPLIANCE INC OF ZHUHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510812679.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-05
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

When traditional embodied intelligent robots obtain real-time environmental information through sensors, computing resource consumption is high and there is perception delay, which affects the real-time performance and response speed of the system.

Method used

The embodied robot obtains the pre-stored information of the object through the wireless communication module, and obtains basic information in combination with the information acquisition sensor. It determines the priority according to the interaction order, performs real-time perception information fusion, and reduces repeated scanning and computing resource consumption.

Benefits of technology

It improves the response speed and task execution smoothness, reduces the amount of calculation and energy consumption, and achieves efficient environmental perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120326637B_ABST
    Figure CN120326637B_ABST
Patent Text Reader

Abstract

The present invention discloses an embodied robot environmental perception method, device, embodied robot, storage medium, and computer program product. Objects in a perception scene include a first category of objects and a second category of objects; the first category of objects are equipped with a positioning chip that stores pre-stored information about the objects. The method comprises: after the embodied robot enters the perception scene, it obtains pre-stored information via its wireless communication module; it obtains basic information about the second category of objects via its information acquisition sensor; after the embodied robot begins interacting with the objects, it determines interaction priorities based on the order of interaction with the objects; based on the interaction priorities, it sequentially obtains real-time perception information about the objects via the information acquisition sensor; and it fuses the real-time perception information with the pre-stored information and the basic information to obtain accurate perception results of the objects. This solution can efficiently utilize resources, improve response speed and task execution fluency, and reduce computational complexity and energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent perception technology, and specifically relates to an environment perception method and device for an embodied robot, an embodied robot, a storage medium, and a computer program product. Background Art

[0002] Compared to traditional robots, embodied intelligent robots can achieve autonomous learning and evolution through dynamic interaction with their environment. This relies not only on information processing capabilities but also on the perception and action capabilities of the intelligent body, enabling them to solve problems by perceiving the environment and taking appropriate actions. However, traditional perception technologies primarily rely on real-time data collection and processing, acquiring and analyzing environmental information through sensors. This real-time perception approach presents the following problems:

[0003] High computing resource consumption: Real-time perception requires processing a large amount of sensor data, which significantly increases computing resource consumption, especially in complex environments.

[0004] Information redundancy and delay: Real-time perception may generate a large amount of redundant information. At the same time, in some application scenarios, real-time processing may cause perception delays, affecting the real-time performance and response speed of the system.

[0005] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention

[0006] The purpose of the present invention is to provide an environmental perception method and device for an embodied robot, an embodied robot, a storage medium and a computer program product, so as to solve the problem in related schemes that the embodied intelligent robot obtains real-time environmental information through sensors and analyzes it, resulting in perception delay and high consumption of computing resources, so as to achieve the effect of quickly obtaining the pre-stored information of the first-class object positioning chip through the wireless communication module, avoiding repeated scanning, and at the same time, real-time perception information collection driven by basic information collection and interaction priority, efficiently utilizing resources, improving response speed and task execution fluency, and reducing computing power and energy consumption.

[0007] The present invention provides an environmental perception method for an embodied robot, wherein the embodied robot can obtain information about objects in a perception scene through its own information acquisition sensors and wireless communication modules; the objects in the perception scene include a first category of objects and a second category of objects; the first category of objects is equipped with a positioning chip, and the positioning chip stores pre-stored information of the object; the method comprises: after the embodied robot enters the perception scene, obtaining the pre-stored information in the positioning chip through the wireless communication module; obtaining basic information of the second category of objects through the information acquisition sensor; after the embodied robot starts to interact with the objects in the perception scene, determining the interaction priority of the objects in the perception scene according to the interaction order of the embodied robot and the objects in the perception scene; according to the interaction priority of the objects in the perception scene, obtaining real-time perception information of the objects in the perception scene through the information acquisition sensors in sequence; and fusing the real-time perception information with the pre-stored information and the basic information respectively to obtain accurate perception results of the objects in the perception scene.

[0008] In some embodiments, the interaction priority of objects in the perception scene is determined based on the interaction order between the embodied robot and the objects in the perception scene, including: obtaining the expected contact time between the embodied robot and the objects in the perception scene; determining the interaction priority of the objects based on the order of the expected contact times; wherein, the interaction priority of the object with the earlier expected contact time is higher than that of the object with the later expected contact time.

[0009] In some embodiments, the pre-stored information of the first category of objects includes: position information, posture information, and material property information; the method further includes: based on the pre-stored information and combined with preset logical reasoning rules, inferring the material property information and use of the second category of objects.

[0010] In some embodiments, the information acquisition sensor includes a camera and a depth camera; obtaining basic information of the second category of objects through the information acquisition sensor includes: obtaining visual features of the second category of objects through the camera; obtaining three-dimensional geometric features of the second category of objects through the depth camera; based on the visual features and the three-dimensional geometric features, identifying the basic information of the second category of objects through a preset deep learning model, including category labels and posture information.

[0011] In some embodiments, the method further includes: when the position or state of the first type of object changes, acquiring change information through the information acquisition sensor; and writing the change information into the positioning chip of the corresponding object through the wireless communication module.

[0012] In some embodiments, the method further includes: during the process of the embodied robot actively interacting with the first type of object, obtaining detailed information of the object being interacted with; comparing and analyzing the detailed information with information in the positioning chip of the object being interacted with, and correcting missing or inaccurate information in the positioning chip.

[0013] Matching the above method, another aspect of the present invention provides an environmental perception device for an embodied robot, wherein the embodied robot is capable of acquiring information about objects in a perception scene through its own information acquisition sensors and / or wireless communication module; the objects in the perception scene include a first category of objects and a second category of objects; the first category of objects is equipped with a positioning chip, which stores pre-stored information about the objects; the environmental perception device comprises: an information interaction unit configured to acquire pre-stored information in the positioning chip through the wireless communication module after the embodied robot enters the perception scene; the information interaction unit is further configured to acquire basic information about the second category of objects through the information acquisition sensors; a decision unit configured to determine, after the embodied robot begins interacting with the objects in the perception scene, an interaction priority of the objects in the perception scene based on the order of interaction between the embodied robot and the objects in the perception scene; a perception unit configured to acquire real-time perception information of the objects in the perception scene through the information acquisition sensors in sequence according to the interaction priority of the objects in the perception scene; and a fusion unit configured to fuse the real-time perception information with the pre-stored information and the basic information, respectively, to obtain accurate perception results of the objects in the perception scene.

[0014] In some embodiments, the decision unit determines the interaction priority of the objects in the perception scene based on the interaction order between the embodied robot and the objects in the perception scene, including: obtaining the expected contact time between the embodied robot and the objects in the perception scene; determining the interaction priority of the objects based on the order of the expected contact times; wherein, the object with an earlier expected contact time has a higher interaction priority than the object with a later expected contact time.

[0015] In some embodiments, the pre-stored information of the first category of objects includes: position information, posture information, and material property information; the information interaction unit is further configured to: infer the material property information and use of the second category of objects based on the pre-stored information and combined with preset logical reasoning rules.

[0016] In some embodiments, the information acquisition sensor includes a camera and a depth camera; the information interaction unit obtains basic information of the second category of objects through the information acquisition sensor, including: obtaining visual features of the second category of objects through the camera; obtaining three-dimensional geometric features of the second category of objects through the depth camera; based on the visual features and the three-dimensional geometric features, identifying the basic information of the second category of objects through a preset deep learning model, including category labels and posture information.

[0017] In some embodiments, the information interaction unit is further configured to: when the position or state of the first type of object changes, obtain change information through the information acquisition sensor; and write the change information into the positioning chip of the corresponding object through the wireless communication module.

[0018] In some embodiments, the information interaction unit is further configured to: obtain detailed information of the object being interacted with during the process of the embodied robot actively interacting with the first type of object; compare and analyze the detailed information with the information in the positioning chip of the object being interacted with, and correct missing or inaccurate information in the positioning chip.

[0019] Matching the above-mentioned device, the present invention further provides an embodied robot, comprising: the above-mentioned environmental perception device of the embodied robot.

[0020] In accordance with the above method, the present invention provides a storage medium on another aspect, wherein the storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned environmental perception method of the embodied robot.

[0021] In accordance with the above method, the present invention provides a computer program product on another aspect, which includes a computer program, and when the computer program product is processed and executed, it implements the steps of the above-mentioned environmental perception method of the embodied robot.

[0022] The solution of the present invention comprises an embodied robot having an information acquisition sensor and a wireless communication module; objects in a perception scene include a first category of objects and a second category of objects; the first category of objects are equipped with a positioning chip that stores pre-stored information about the objects; after the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module; and obtains basic information about the second category of objects via the information acquisition sensor; after the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene; and based on the interaction priority of the objects, it sequentially obtains real-time perception information of the objects via the information acquisition sensor; and the real-time perception information is integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. Thus, the wireless communication module can quickly obtain the pre-stored information of the positioning chip of the first category of objects, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by the basic information acquisition and interaction priority can efficiently utilize resources, improve response speed and task execution fluency, and reduce computational complexity and energy consumption.

[0023] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of the present invention.

[0024] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic flow chart of an embodiment of an environment perception method for an embodied robot according to the present invention;

[0026] Figure 2 A schematic structural diagram of an embodiment of an environment perception device for an embodied robot according to the present invention;

[0027] Figure 3 A schematic diagram of the process of inputting pre-stored information into the positioning chip;

[0028] Figure 4 FIG. 4 is a flow chart of another embodiment of the environment perception method of the embodied robot of the present invention.

[0029] In conjunction with the accompanying drawings, the reference numerals in the embodiments of the present invention are as follows:

[0030] 101 - information interaction unit; 102 - decision-making unit; 103 - perception unit; 104 - fusion unit. DETAILED DESCRIPTION

[0031] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0032] According to an embodiment of the present invention, a method for environmental perception of an embodied robot is provided, wherein the embodied robot can obtain information about objects in a perception scene through its own information acquisition sensors and wireless communication module; the objects in the perception scene include a first category of objects and a second category of objects; the first category of objects is configured with a positioning chip, and the positioning chip stores pre-stored information about the object.

[0033] Information collection sensors include visual sensors, environmental sensors, and tactile sensors. Visual sensors, including cameras and depth cameras, are used to perceive RGB images and point cloud data, enabling object recognition and scene reconstruction. Environmental sensors, including lidar and infrared sensors, are used to perceive distance data and detect obstacles, enabling embodied robots to perform functions such as navigation and obstacle avoidance, and dynamic object tracking. Tactile sensors, including six-dimensional force sensors and tactile arrays, are used to sense contact force, torque, and the surface texture of objects, enabling embodied robots to perform operations such as fine grasping and compliant interaction.

[0034] The wireless communication module is used to realize data interaction between the robot and external systems. It has the characteristics of low latency, high bandwidth, low power consumption and flexible networking.

[0035] In the perception scenario, the first category of objects are those with fixed shapes and stable structures (such as factory machine tools and warehouse shelves). Positioning chips can be securely installed (either internally or externally) and are not easily deformed by external forces. The second category of objects are dynamically changeable or irregularly shaped (such as flexible packages and scattered parts), or temporarily placed unknown objects, which cannot be pre-installed with positioning chips.

[0036] The positioning chip for the first type of object stores the object's 3D model, location information, orientation information, and material properties. The 3D model is provided by the manufacturer upon shipment. For objects for which a model is not provided, a 3D laser scanner and depth camera are used to create a high-precision model of the target object and obtain its 3D geometric information. The object's fixed position, orientation, and material properties are determined and stored in the positioning chip.

[0037] Figure 3 A flow chart showing the process of inputting pre-stored information into the positioning chip, such as Figure 3 As shown, the method includes:

[0038] Step 1: Model the fixed and stable objects in the perception scene.

[0039] Step 2: Determine whether the manufacturer has provided a 3D model of the object. If so, proceed to Step 3. If not, use a 3D laser scanner and depth camera to perform high-precision modeling of the target object, obtain its 3D geometric information, and proceed to Step 3.

[0040] Step 3: Determine the object's fixed position, orientation information, posture information, and material property information, and store this information and the object's three-dimensional model in the corresponding positioning chip.

[0041] The embodied robot can directly read the information of the first type of objects from the positioning chip, thereby eliminating the need to collect data on the first type of objects through information collection sensors, greatly reducing the amount of data the robot has to process and improving the robot's perception rate.

[0042] like Figure 1 FIG2 is a flow chart of an embodiment of the method of the present invention. The environment perception method of the embodied robot may include steps S110 to S150.

[0043] In step S110 , after the embodied robot enters the perception scene, it obtains pre-stored information in the positioning chip through the wireless communication module.

[0044] In step S120, basic information of the second category of objects is acquired through the information acquisition sensor.

[0045] After the embodied robot enters the scene, it synchronously or asynchronously acquires information about objects in the perceived scene. For the first category of objects, the wireless communication module directly reads the pre-stored information on the positioning chip, including location, material, and posture. For the second category of objects, sensors such as cameras and depth cameras collect visual features and three-dimensional geometric features, including fixed position, color, texture, size, and curvature. For example, after entering the warehouse, the robot reads the pre-stored information on shelf labels (first category objects) within a 20-meter range within 1 second. Simultaneously, the depth camera quickly scans the channel and identifies temporary obstacles (second category objects), reducing the overall initialization time from 5 seconds for traditional solutions to 1.5 seconds. This pre-stored information on the first category of objects provides a basic anchor point for the environment, filling the gaps in the sensor's repeated scanning of known objects. The real-time perception of the second category of objects covers unknown elements, ensuring comprehensive scene information.

[0046] Compared with information collection sensors, the time consumption of obtaining information through wireless communication modules is shorter, which significantly improves the overall perception initialization speed, reduces the invalid data collection of known objects by sensors, and reduces energy consumption.

[0047] In some embodiments, the information acquisition sensor includes a camera and a depth camera. The camera is used to collect two-dimensional visual features of the object, including color, texture, shape and outline. The depth camera is used to obtain the three-dimensional geometric features of the object, that is, depth information, so as to construct spatial coordinates and a point cloud model. In step S120, the specific process of obtaining the basic information of the second category of objects through the information acquisition sensor includes: obtaining the visual features of the second category of objects through the camera; obtaining the three-dimensional geometric features of the second category of objects through the depth camera; and based on the visual features and the three-dimensional geometric features, identifying the basic information of the second category of objects through a preset deep learning model, including category labels and posture information.

[0048] Specifically, the visual features are first filtered and de-noised, and the three-dimensional geometric features are filtered. The two-dimensional visual features and three-dimensional geometric features are then concatenated at the feature level to form a consistent feature vector that includes both appearance and spatial information. The fused feature vector is then fed into a pre-set deep learning model to obtain the object's category label and pose information. Category labels include "obstacle," "electronic device," and so on, accompanied by a confidence level. Pose information includes translation and rotation. For example, if the camera captures the visual features of "silver-gray, irregular shape, and threaded surface," and the depth camera captures the geometric features of "volume 100 cm³, thread depth 0.5 mm, and tilt angle 20°," the model outputs the category label "bolt" (with a 92% confidence level) and the pose information "threaded hole facing right, deviating 15° from the standard angle."

[0049] Through the collaboration of cameras and depth cameras and the application of deep learning models, key technical support is provided for embodied robots to achieve precise object interaction in unknown environments.

[0050] In some embodiments, the pre-stored information of the first category of objects includes: position information, posture information, and material property information. Position information is the absolute coordinates of the object in three-dimensional space or the position data relative to the robot, which can be used for robot path planning and as a spatial reference for logical reasoning. Posture information includes the orientation and state of the object, which is used to assist the robot in understanding the functional state of the object and infer the potential use of the second category of objects based on the position information. Material property information is the physical properties and chemical characteristics of the object, such as material type, physical parameters, safety properties, etc., which is used to guide the robot to perform interactive actions and is also a key basis for logical reasoning.

[0051] The method also includes a process of inferring information about a second type of object based on pre-stored information about the first type of object. The process specifically includes: inferring material property information and usage of the second type of object based on the pre-stored information and in combination with preset logical reasoning rules.

[0052] The inference process includes data collection and feature extraction, rule matching, and confidence calculation. Specifically, the data collection and feature extraction step uses a visual camera to extract the color, shape, and texture of the second-category object, a tactile sensor to detect its hardness and weight, and a radar to measure its distance and relative position to the first-category object. The rule matching step includes spatial association rules, material consistency rules, functional complementarity rules, and causal relationship rules.

[0053] In spatial association rules, if the second-category object is located close to the first-category object and has related functions, it inherits some of its properties. For example, if the first-category object is a microwave donkey and the second-category object is an unknown box placed next to a microwave, it is likely a plastic heating container.

[0054] In the functional complementarity rule, if the first type of object requires a tool to be used, then the nearby second type of object is likely to be the corresponding tool. For example, if the first type of object is a dining table and the second type of object is a flat object on the table, it is likely to be a metal or ceramic cutlery.

[0055] In the material consistency rule, objects in the same area usually have similar materials. For example, if the first type of object is a sofa, and the second type of object is a soft object on the ground, it is inferred that it is made of cloth.

[0056] In causal relationship rules, if a change in the state of a first type of object leads to the appearance of a second type of object, the second type of object is likely a product or related object of the first type. For example, if the first type of object is a printer and the second type of object is an unknown piece of paper that comes out of the printer, then it is inferred that it is a printed document.

[0057] The confidence calculation step calculates the confidence based on the rule matching weight and the degree of agreement with the sensor data. For example, if the spatial association rule weight is 0.4, the sensor data shows a distance of 0.2 meters from the oven, and the degree of agreement is 0.9, the contribution confidence is 0.36.

[0058] By inferring the information of the second category of objects, the perception dependence and computing cost are reduced, the information gap of unknown objects is filled, and the rationality and safety of the robot's interactive decision-making are enhanced.

[0059] At step S130 , after the embodied robot starts interacting with the objects in the perception scene, the interaction priority of the objects in the perception scene is determined according to the interaction order between the embodied robot and the objects in the perception scene.

[0060] After the robot begins interacting, it prioritizes objects based on their expected contact time. This time is calculated for path planning and sorted by time, with objects with earlier contact times receiving higher priority. Traditional methods rely on fixed priorities (e.g., person > obstacle > target) and are unable to cope with dynamic scenarios. This approach, which prioritizes interactions based on order, enables the robot to respond to environmental changes in real time. For example, when planning a path, the robot detects the target cargo (contact time 8 seconds) and a passageway obstacle (contact time 3 seconds). Prioritizing the obstacle, the robot only adds 1.2 seconds to the obstacle avoidance process, reducing latency by 3 seconds compared to fixed-priority methods.

[0061] In some embodiments, in step S130, determining the interaction priority of the objects in the perception scene according to the interaction order between the embodied robot and the objects in the perception scene includes: step S210 and step S220.

[0062] Step S210: obtaining the expected contact time between the embodied robot and the object in the perception scene.

[0063] The expected contact time is the time interval between the robot's current state and the start of its interaction with an object. It can be calculated based on the robot's current position, movement speed, acceleration, and other operating parameters, combined with the object's motion state. For example, for a static object, the expected contact time = the distance between the robot and the object's interaction point / the robot's movement speed; for a dynamic object, the expected contact time = the distance between the robot and the object's interaction point / the robot's speed relative to the object.

[0064] Step S220 , determining the interaction priority of the objects according to the order of the expected contact times; wherein, the interaction priority of the object with the earlier expected contact time is higher than that of the object with the later expected contact time.

[0065] Interaction priorities are sorted by time, with shorter expected contact times giving higher priority. For example, an obstacle 2 meters in front of the robot takes priority over a target 5 meters away. A person approaching the robot dynamically has the highest priority.

[0066] By reasonably setting the interaction priority of different objects according to the expected contact time, the robot's dynamic response capability is improved, solving the problems of decision-making lag and resource waste caused by ambiguous interaction priorities, and avoiding repeated movements or ineffective waiting due to unreasonable path planning.

[0067] In step S140 , according to the interaction priorities of the objects in the perception scene, real-time perception information of the objects in the perception scene is acquired in sequence through the information collection sensors.

[0068] Real-time perception information is the information about the environment and objects collected by embodied robots through various sensors during real-time interaction. It is dynamic and multimodal. This information includes: 2D image information such as object color, texture, and shape captured by cameras; distance information between objects and the robot acquired by depth sensors, used to construct 3D spatial coordinates; force feedback data (contact force) and ultrasonic data (obstacle distance).

[0069] Traditional methods perform full-resolution scans on all objects, resulting in wasted computing resources. This solution uses information collection sensors to obtain real-time depth information of objects in order of priority, avoiding the use of low-value data in computing power. This optimizes computing power, reduces overall data processing volume, and reduces the load on edge computing nodes.

[0070] In step S150, the real-time perception information is fused with the pre-stored information and the basic information respectively to obtain an accurate perception result of the object in the perception scene.

[0071] Compared to pre-stored and basic information, real-time perception information reflects the current state and is highly timely. Furthermore, because it is acquired by multiple sensors, the accuracy of this information is higher.

[0072] When performing information fusion, data alignment and calibration are first performed, including spatiotemporal alignment and outlier filtering. Spatiotemporal alignment unifies the timestamps and spatial coordinate systems of multi-source data, ensuring that the visual image and depth data correspond to the same moment and the same object. Sensor calibration can eliminate perspective deviations and ensure a one-to-one correspondence between 2D pixels and 3D point clouds. Outlier filtering is used to remove noise points in the depth data or blurred areas in the visual data, ensuring reliable input data. For example, red areas in the visual image are matched to corresponding areas in the depth point cloud to eliminate background interference.

[0073] Feature fusion is then performed. Visual data provides semantic information, identifying object category, color, and texture; depth data provides spatial information, determining the object's three-dimensional position, size, and posture. Visual and depth features are combined at the feature level to form composite features that encompass both semantics and spatial information. For example, the visually recognized label "red cylinder" is combined with the depth-measured label "diameter 8 cm, height 10 cm, position (1, 0.5, 0.8)" to form a complete description of a water cup.

[0074] Finally, a data fusion algorithm processes the composite features and outputs accurate object perception results. Data fusion algorithms include Kalman filtering and neural network fusion models. The output includes the object's precise 3D coordinates and pose, as well as semantic labels and attributes. For example, the fusion results can be directly used by the robot to plan grasping actions, such as calculating the robot arm's movement path based on depth data and adjusting the grasping angle based on visual data.

[0075] Real-time perception information is integrated with pre-stored information or sensor data to generate precise perception results. For the first category of objects, real-time perception information can correct pre-stored position errors and supplement surface details. For the second category of objects, real-time perception information is integrated with visual features to infer material and purpose. For example, high-priority obstacles are scanned with high precision point clouds to generate semantically labeled 3D models, guiding the robot to navigate around them with high accuracy. Low-priority shelves are scanned with low precision to update only position deviations, consuming less computing power. This solves the problem of one-sided data collected by a single sensor, making object perception more accurate and complete.

[0076] This solution systematically addresses the balance between efficiency and accuracy in the perception of embodied robots in complex environments through rapid initialization of pre-stored information, precise scheduling based on dynamic priorities, and deep fusion of multi-source data. It leverages positioning chips to reduce the cost of perceiving known objects while expanding adaptability to unknown scenarios through sensors. Computing power and sensor resources are intelligently scheduled based on interaction priorities to achieve efficient perception.

[0077] In some embodiments, a process of updating the information in the positioning chip is also included, which specifically includes: when the position or state of the first type of object changes, obtaining the change information through the information acquisition sensor; and writing the change information into the positioning chip of the corresponding object through the wireless communication module.

[0078] As the robot interacts with objects in the scene or manually handles objects, the pre-stored information in the positioning chip lags behind changes in the actual environment and needs to be updated in a timely manner. When the object moves a distance exceeding a preset threshold or the posture changes by more than a certain angle, the object's position is considered to have changed. The depth camera can compare multiple frames of point cloud data to calculate the object's coordinate offset; the visual camera detects the pixel displacement of the object and converts it into the actual movement distance to determine whether the object's position has changed. When the functional state of the object changes, the object's state is considered to have changed. For example, the object switches from open to closed, or from loading to unloading. The camera can detect changes in the object's appearance, and the force sensor can detect changes in the object's weight to determine whether the object's state has changed.

[0079] After determining that the position or status of an object has changed, the robot collects the changed information and updates it to the positioning chip of the corresponding object through wireless communication, ensuring that the pre-stored information matches the current status, solving the position drift problem of traditional static tags and improving collaborative efficiency and the accuracy of robot perception.

[0080] In some embodiments, the process of correcting the information in the positioning chip is also included, and the process specifically includes: during the process of the embodied robot actively interacting with the first type of object, obtaining detailed information of the object being interacted with; comparing and analyzing the detailed information with the information in the positioning chip of the object being interacted with, and correcting the missing or inaccurate information in the positioning chip.

[0081] Since the information in the object's positioning chip may be missing or erroneous, it is necessary to correct the information in real time. Specifically, when the robot obtains more detailed object information during active interaction, it compares and analyzes it with the information in the object's positioning chip to determine whether the information in the positioning chip is accurate and complete. For example, the compared information includes position coordinates, object weight, material hardness, surface roughness, etc. When the error between a certain information in the positioning chip and the obtained detailed information exceeds the set threshold, it is considered necessary to correct the information in the positioning chip. When the information is missing, the measured value is written directly; when the information is inaccurate, the new information is used to overwrite the old information. Continuous calibration offsets the errors caused by environmental changes to ensure that the robot remains reliable during long-term operations.

[0082] Figure 4 FIG. 1 is a flow chart of another embodiment of the environment perception method of the embodied robot of the present invention. Figure 4 As shown, the method includes:

[0083] In step 11, after entering the sensing scene, the robot activates its wireless communication module and scans the surrounding area for positioning chips. If a positioning chip is present, the robot uses the wireless communication module to retrieve the object's 3D model, position information, material properties, and other information stored in the positioning chip, and then proceeds to step 13. If no positioning chip is present, the robot proceeds to step 12.

[0084] In step 12, sensors such as cameras and depth cameras are used to perform preliminary perception of the object without a positioning chip, obtaining rough information such as the object's location, shape, and size. Simultaneously, information about known objects is combined to perform inference and infer the properties and purpose of the object without a positioning chip. Then, step 13 is executed.

[0085] In step 13, objects are dynamically layered based on the order in which they interact. Information about objects that require immediate robot interaction is prioritized. Based on the layering results, high-priority object information is fused with real-time perception data to generate high-precision perception results.

[0086] In step 14, the robot continuously optimizes its perception model and scene understanding capabilities through active interaction with the environment. The robot interacts with objects in the environment through actions such as touch and movement, acquiring more detailed object information (such as material properties and surface texture). This newly acquired information is compared and supplemented with the information stored in the chip to optimize the perception model. If new objects appear in the scene or existing objects are moved, the robot records these changes and writes the updated information to the relevant chip via wireless communication. The robot stores this updated scene information in a local database to provide reference for subsequent perception tasks.

[0087] By pre-modeling objects in the surrounding scene, the perception system only needs to perceive objects that are not expected to appear in the perception scene. This enables low-latency, real-time perception in complex environments, significantly improving the efficiency of the perception system. By layering objects in the perception scene, it prioritizes objects that require immediate interaction, reducing the system's computing resource consumption and making it suitable for resource-constrained environments.

[0088] According to the technical solution of this embodiment, an embodied robot comprises an information acquisition sensor and a wireless communication module. Objects in a perception scene include a first category of objects and a second category of objects. The first category of objects is equipped with a positioning chip that stores pre-stored information about the objects. After the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module. Basic information about the second category of objects is obtained via the information acquisition sensor. After the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene. Based on the interaction priority of the objects, the robot sequentially obtains real-time perception information of the objects via the information acquisition sensor. The real-time perception information is then integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. Thus, the wireless communication module rapidly obtains the pre-stored information of the positioning chip of the first category of objects, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by basic information acquisition and interaction priority utilizes resources efficiently, improves response speed and task execution fluency, and reduces computational complexity and energy consumption.

[0089] According to an embodiment of the present invention, an embodied robot environment perception device corresponding to the embodied robot environment perception method is also provided. The embodied robot is capable of acquiring information about objects in a perception scene through its own information acquisition sensors and wireless communication module. The objects in the perception scene include first and second categories of objects. The first category of objects is equipped with a positioning chip that stores pre-stored information about the objects.

[0090] Information collection sensors include visual sensors, environmental sensors, and tactile sensors. Visual sensors, including cameras and depth cameras, are used to perceive RGB images and point cloud data, enabling object recognition and scene reconstruction. Environmental sensors, including lidar and infrared sensors, are used to perceive distance data and detect obstacles, enabling embodied robots to perform functions such as navigation and obstacle avoidance, and dynamic object tracking. Tactile sensors, including six-dimensional force sensors and tactile arrays, are used to sense contact force, torque, and the surface texture of objects, enabling embodied robots to perform operations such as fine grasping and compliant interaction.

[0091] The wireless communication module is used to realize data interaction between the robot and external systems. It has the characteristics of low latency, high bandwidth, low power consumption and flexible networking.

[0092] In the perception scenario, the first category of objects are those with fixed shapes and stable structures (such as factory machine tools and warehouse shelves). Positioning chips can be securely installed (either internally or externally) and are not easily deformed by external forces. The second category of objects are dynamically changeable or irregularly shaped (such as flexible packages and scattered parts), or temporarily placed unknown objects, which cannot be pre-installed with positioning chips.

[0093] The positioning chip for the first type of object stores the object's 3D model, location information, orientation information, and material properties. The 3D model is provided by the manufacturer upon shipment. For objects for which a model is not provided, a 3D laser scanner and depth camera are used to create a high-precision model of the target object and obtain its 3D geometric information. The object's fixed position, orientation, and material properties are determined and stored in the positioning chip.

[0094] Figure 3 A flow chart showing the process of inputting pre-stored information into the positioning chip, such as Figure 3 As shown, the method includes:

[0095] Step 1: Model the fixed and stable objects in the perception scene.

[0096] Step 2: Determine whether the manufacturer has provided a 3D model of the object. If so, proceed to Step 3. If not, use a 3D laser scanner and depth camera to perform high-precision modeling of the target object, obtain its 3D geometric information, and proceed to Step 3.

[0097] Step 3: Determine the object's fixed position, orientation information, posture information, and material property information, and store this information and the object's three-dimensional model in the corresponding positioning chip.

[0098] The embodied robot can directly read the information of the first type of objects from the positioning chip, thereby eliminating the need to collect data on the first type of objects through information collection sensors, greatly reducing the amount of data the robot has to process and improving the robot's perception rate.

[0099] See also Figure 2 FIG2 is a schematic structural diagram of an embodiment of the device of the present invention. The environment perception device of the embodied robot may include: an information interaction unit 101 , a decision unit 102 , a perception unit 103 , and a fusion unit 104 .

[0100] The information interaction unit 101 is configured to obtain the pre-stored information in the positioning chip through the wireless communication module after the embodied robot enters the perception scene.

[0101] The information interaction unit 101 is further configured to obtain basic information of the second category of objects through the information collection sensor.

[0102] After the embodied robot enters the scene, it synchronously or asynchronously acquires information about objects in the perceived scene. For the first category of objects, the wireless communication module directly reads the pre-stored information on the positioning chip, including location, material, and posture. For the second category of objects, sensors such as cameras and depth cameras collect visual features and three-dimensional geometric features, including fixed position, color, texture, size, and curvature. For example, after entering the warehouse, the robot reads the pre-stored information on shelf labels (first category objects) within a 20-meter range within 1 second. Simultaneously, the depth camera quickly scans the channel and identifies temporary obstacles (second category objects), reducing the overall initialization time from 5 seconds for traditional solutions to 1.5 seconds. This pre-stored information on the first category of objects provides a basic anchor point for the environment, filling the gaps in the sensor's repeated scanning of known objects. The real-time perception of the second category of objects covers unknown elements, ensuring comprehensive scene information.

[0103] Compared with information collection sensors, the time consumption of obtaining information through wireless communication modules is shorter, which significantly improves the overall perception initialization speed, reduces the invalid data collection of known objects by sensors, and reduces energy consumption.

[0104] In some embodiments, the information acquisition sensor includes a camera and a depth camera. The camera is used to collect two-dimensional visual features of the object, including color, texture, shape outline, etc. The depth camera is used to obtain the three-dimensional geometric features of the object, that is, depth information, so as to construct spatial coordinates and point cloud models. The information interaction unit 101 obtains the basic information of the second category of objects through the information acquisition sensor, including: obtaining the visual features of the second category of objects through the camera; obtaining the three-dimensional geometric features of the second category of objects through the depth camera; based on the visual features and the three-dimensional geometric features, identifying the basic information of the second category of objects through a preset deep learning model, including category labels and posture information.

[0105] Specifically, the visual features are first filtered and de-noised, and the three-dimensional geometric features are filtered. The two-dimensional visual features and three-dimensional geometric features are then concatenated at the feature level to form a consistent feature vector that includes both appearance and spatial information. The fused feature vector is then fed into a pre-set deep learning model to obtain the object's category label and pose information. Category labels include "obstacle," "electronic device," and so on, accompanied by a confidence level. Pose information includes translation and rotation. For example, if the camera captures the visual features of "silver-gray, irregular shape, and threaded surface," and the depth camera captures the geometric features of "volume 100 cm³, thread depth 0.5 mm, and tilt angle 20°," the model outputs the category label "bolt" (with a 92% confidence level) and the pose information "threaded hole facing right, deviating 15° from the standard angle."

[0106] Through the collaboration of cameras and depth cameras and the application of deep learning models, key technical support is provided for embodied robots to achieve precise object interaction in unknown environments.

[0107] In some embodiments, the pre-stored information of the first category of objects includes: position information, posture information, and material property information. Position information is the absolute coordinates of the object in three-dimensional space or the position data relative to the robot, which can be used for robot path planning and as a spatial reference for logical reasoning. Posture information includes the orientation and state of the object, which is used to assist the robot in understanding the functional state of the object and infer the potential use of the second category of objects based on the position information. Material property information is the physical properties and chemical characteristics of the object, such as material type, physical parameters, safety properties, etc., which is used to guide the robot to perform interactive actions and is also a key basis for logical reasoning.

[0108] The information interaction unit 101 is further configured to: infer the material property information and usage of the second type of object based on the pre-stored information and in combination with preset logical reasoning rules.

[0109] The inference process includes data collection and feature extraction, rule matching, and confidence calculation. Specifically, the data collection and feature extraction step uses a visual camera to extract the color, shape, and texture of the second-category object, a tactile sensor to detect its hardness and weight, and a radar to measure its distance and relative position to the first-category object. The rule matching step includes spatial association rules, material consistency rules, functional complementarity rules, and causal relationship rules.

[0110] In spatial association rules, if the second-category object is located close to the first-category object and has related functions, it inherits some of its properties. For example, if the first-category object is a microwave donkey and the second-category object is an unknown box placed next to a microwave, it is likely a plastic heating container.

[0111] In the functional complementarity rule, if the first type of object requires a tool to be used, then the nearby second type of object is likely to be the corresponding tool. For example, if the first type of object is a dining table and the second type of object is a flat object on the table, it is likely to be a metal or ceramic cutlery.

[0112] In the material consistency rule, objects in the same area usually have similar materials. For example, if the first type of object is a sofa and the second type of object is a soft object on the ground, it is inferred that it is made of cloth.

[0113] In causal relationship rules, if a change in the state of a first type of object leads to the appearance of a second type of object, the second type of object is likely a product or related object of the first type. For example, if the first type of object is a printer and the second type of object is an unknown piece of paper that comes out of the printer, then it is inferred that it is a printed document.

[0114] The confidence calculation step calculates the confidence based on the rule matching weight and the degree of agreement with the sensor data. For example, if the spatial association rule weight is 0.4, the sensor data shows a distance of 0.2 meters from the oven, and the degree of agreement is 0.9, the contribution confidence is 0.36.

[0115] By inferring the information of the second type of objects, the perception dependence and computing cost are reduced, the information gap of unknown objects is filled, and the rationality and safety of the robot's interactive decision-making are enhanced.

[0116] The decision unit 102 is configured to determine the interaction priority of the objects in the perception scene according to the interaction order between the embodied robot and the objects in the perception scene after the embodied robot starts to interact with the objects in the perception scene.

[0117] After the robot begins interacting, it prioritizes objects based on their expected contact time. This time is calculated for path planning and sorted by time, with objects with earlier contact times receiving higher priority. Traditional methods rely on fixed priorities (e.g., person > obstacle > target) and are unable to cope with dynamic scenarios. This approach, which prioritizes interactions based on order, enables the robot to respond to environmental changes in real time. For example, when planning a path, the robot detects the target cargo (contact time 8 seconds) and a passageway obstacle (contact time 3 seconds). Prioritizing the obstacle, the robot only adds 1.2 seconds to the obstacle avoidance process, reducing latency by 3 seconds compared to fixed-priority methods.

[0118] In some embodiments, the decision unit 102 determines the interaction priority of the objects in the perceived scene according to the interaction order between the embodied robot and the objects in the perceived scene, including:

[0119] The decision unit 102 is further configured to obtain an expected contact time between the embodied robot and the object in the perception scene.

[0120] The expected contact time is the time interval between the robot's current state and the start of its interaction with an object. It can be calculated based on the robot's current position, movement speed, acceleration, and other operating parameters, combined with the object's motion state. For example, for a static object, the expected contact time = the distance between the robot and the object's interaction point / the robot's movement speed; for a dynamic object, the expected contact time = the distance between the robot and the object's interaction point / the robot's speed relative to the object.

[0121] The decision unit 102 is further configured to determine the interaction priority of the objects according to the order of the expected contact times; wherein the object with the earlier expected contact time has a higher interaction priority than the object with the later expected contact time.

[0122] Interaction priorities are sorted by time, with shorter expected contact times giving higher priority. For example, an obstacle 2 meters in front of the robot takes priority over a target 5 meters away. A person approaching the robot dynamically has the highest priority.

[0123] By reasonably setting the interaction priority of different objects according to the expected contact time, the robot's dynamic response capability is improved, solving the problems of decision-making lag and resource waste caused by ambiguous interaction priorities, and avoiding repeated movements or ineffective waiting due to unreasonable path planning.

[0124] The perception unit 103 is configured to acquire real-time perception information of the objects in the perception scene through the information collection sensors in sequence according to the interaction priorities of the objects in the perception scene.

[0125] Real-time perception information is the information about the environment and objects collected by embodied robots through various sensors during real-time interaction. It is dynamic and multimodal. This information includes: 2D image information such as object color, texture, and shape captured by cameras; distance information between objects and the robot acquired by depth sensors, used to construct 3D spatial coordinates; force feedback data (contact force) and ultrasonic data (obstacle distance).

[0126] Traditional methods perform full-resolution scans on all objects, resulting in wasted computing resources. This solution uses information collection sensors to obtain real-time depth information of objects in order of priority, avoiding the use of low-value data in computing power. This optimizes computing power, reduces overall data processing volume, and reduces the load on edge computing nodes.

[0127] The fusion unit 104 is configured to fuse the real-time perception information with the pre-stored information and the basic information respectively to obtain an accurate perception result of the object in the perception scene.

[0128] Compared to pre-stored and basic information, real-time perception information reflects the current state and is highly timely. Furthermore, because it is acquired by multiple sensors, the accuracy of this information is higher.

[0129] When performing information fusion, data alignment and calibration are first performed, including spatiotemporal alignment and outlier filtering. Spatiotemporal alignment unifies the timestamps and spatial coordinate systems of multi-source data, ensuring that the visual image and depth data correspond to the same moment and the same object. Sensor calibration can eliminate perspective deviations and ensure a one-to-one correspondence between 2D pixels and 3D point clouds. Outlier filtering is used to remove noise points in the depth data or blurred areas in the visual data, ensuring reliable input data. For example, red areas in the visual image are matched to corresponding areas in the depth point cloud to eliminate background interference.

[0130] Feature fusion is then performed. Visual data provides semantic information, identifying object category, color, and texture; depth data provides spatial information, determining the object's three-dimensional position, size, and posture. Visual and depth features are combined at the feature level to form composite features that encompass both semantics and spatial information. For example, the visually recognized label "red cylinder" is combined with the depth-measured label "diameter 8 cm, height 10 cm, position (1, 0.5, 0.8)" to form a complete description of a water cup.

[0131] Finally, a data fusion algorithm processes the composite features and outputs accurate object perception results. Data fusion algorithms include Kalman filtering and neural network fusion models. The output includes the object's precise 3D coordinates and pose, as well as semantic labels and attributes. For example, the fusion results can be directly used by the robot to plan grasping actions, such as calculating the robot arm's movement path based on depth data and adjusting the grasping angle based on visual data.

[0132] Real-time perception information is integrated with pre-stored information or sensor data to generate precise perception results. For the first category of objects, real-time perception information can correct pre-stored position errors and supplement surface details. For the second category of objects, real-time perception information is integrated with visual features to infer material and purpose. For example, high-priority obstacles are scanned with high precision point clouds to generate semantically labeled 3D models, guiding the robot to navigate around them with high accuracy. Low-priority shelves are scanned with low precision to update only position deviations, consuming less computing power. This solves the problem of one-sided data collected by a single sensor, making object perception more accurate and complete.

[0133] This solution systematically addresses the balance between efficiency and accuracy in the perception of embodied robots in complex environments through rapid initialization of pre-stored information, precise scheduling based on dynamic priorities, and deep fusion of multi-source data. It leverages positioning chips to reduce the cost of perceiving known objects while expanding adaptability to unknown scenarios through sensors. Computing power and sensor resources are intelligently scheduled based on interaction priorities to achieve efficient perception.

[0134] In some embodiments, the information interaction unit 101 is further configured to: when the position or state of the first type of object changes, obtain change information through the information acquisition sensor; and write the change information into the positioning chip of the corresponding object through the wireless communication module.

[0135] As the robot interacts with objects in the scene or manually handles objects, the pre-stored information in the positioning chip lags behind changes in the actual environment and needs to be updated in a timely manner. When the object moves a distance exceeding a preset threshold or the posture changes by more than a certain angle, the object's position is considered to have changed. The depth camera can compare multiple frames of point cloud data to calculate the object's coordinate offset; the visual camera detects the pixel displacement of the object and converts it into the actual movement distance to determine whether the object's position has changed. When the functional state of the object changes, the object's state is considered to have changed. For example, the object switches from open to closed, or from loading to unloading. The camera can detect changes in the object's appearance, and the force sensor can detect changes in the object's weight to determine whether the object's state has changed.

[0136] After determining that the position or status of an object has changed, the robot collects the changed information and updates it to the positioning chip of the corresponding object through wireless communication, ensuring that the pre-stored information matches the current status, solving the position drift problem of traditional static tags and improving collaborative efficiency and the accuracy of robot perception.

[0137] In some embodiments, the information interaction unit 101 is further configured to: obtain detailed information of the object being interacted with during the process of the embodied robot actively interacting with the first type of object; compare and analyze the detailed information with the information in the positioning chip of the object being interacted with, and correct missing or inaccurate information in the positioning chip.

[0138] Since the information in the object's positioning chip may be missing or erroneous, it is necessary to correct the information in real time. Specifically, when the robot obtains more detailed object information during active interaction, it compares and analyzes it with the information in the object's positioning chip to determine whether the information in the positioning chip is accurate and complete. For example, the compared information includes position coordinates, object weight, material hardness, surface roughness, etc. When the error between a certain information in the positioning chip and the obtained detailed information exceeds the set threshold, it is considered necessary to correct the information in the positioning chip. When the information is missing, the measured value is written directly; when the information is inaccurate, the new information is used to overwrite the old information. Continuous calibration offsets the errors caused by environmental changes to ensure that the robot remains reliable during long-term operations.

[0139] Figure 4 FIG. 1 is a flow chart of another embodiment of the environment perception method of the embodied robot of the present invention. Figure 4 As shown, the method includes:

[0140] In step 11, after entering the sensing scene, the robot activates its wireless communication module and scans the surrounding area for positioning chips. If a positioning chip is present, the robot uses the wireless communication module to retrieve the object's 3D model, position information, material properties, and other information stored in the positioning chip, and then proceeds to step 13. If no positioning chip is present, the robot proceeds to step 12.

[0141] In step 12, sensors such as cameras and depth cameras are used to perform preliminary perception of the object without a positioning chip, obtaining rough information such as the object's location, shape, and size. Simultaneously, information about known objects is combined to perform inference and infer the properties and purpose of the object without a positioning chip. Then, step 13 is executed.

[0142] In step 13, objects are dynamically layered based on the order in which they interact. Information about objects that require immediate robot interaction is prioritized. Based on the layering results, high-priority object information is fused with real-time perception data to generate high-precision perception results.

[0143] In step 14, the robot continuously optimizes its perception model and scene understanding capabilities through active interaction with the environment. The robot interacts with objects in the environment through actions such as touch and movement, acquiring more detailed object information (such as material properties and surface texture). This newly acquired information is compared and supplemented with the information stored in the chip to optimize the perception model. If new objects appear in the scene or existing objects are moved, the robot records these changes and writes the updated information to the relevant chip via wireless communication. The robot stores this updated scene information in a local database to provide reference for subsequent perception tasks.

[0144] By pre-modeling objects in the surrounding scene, the perception system only needs to perceive objects that are not expected to appear in the perception scene. This enables low-latency, real-time perception in complex environments, significantly improving the efficiency of the perception system. By layering objects in the perception scene, it prioritizes objects that require immediate interaction, reducing the system's computing resource consumption and making it suitable for resource-constrained environments.

[0145] Since the processing and functions implemented by the device of this embodiment basically correspond to the embodiments, principles and examples of the aforementioned method, for any details not fully described in this embodiment, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.

[0146] The present invention employs an embodied robot comprising an information acquisition sensor and a wireless communication module. Objects in a perception scene include a first category and a second category. The first category of objects is equipped with a positioning chip containing pre-stored information about the objects. After the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module. Basic information about the second category of objects is obtained via the information acquisition sensor. After the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene. Based on the interaction priority of the objects, the robot sequentially obtains real-time perception information of the objects via the information acquisition sensor. The real-time perception information is then integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. This allows the robot to quickly obtain the pre-stored information of the positioning chip for the first category of objects via the wireless communication module, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by basic information acquisition and interaction priority utilizes resources efficiently, improves response speed and task execution fluency, and reduces computational complexity and energy consumption.

[0147] According to an embodiment of the present invention, an embodied robot corresponding to an environment perception device of the embodied robot is further provided. The embodied robot may include: the environment perception device of the embodied robot described above.

[0148] Since the processing and functions implemented by the embodied robot of this embodiment basically correspond to the embodiments, principles and examples of the aforementioned devices, for any details not fully described in this embodiment, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.

[0149] The present invention employs an embodied robot comprising an information acquisition sensor and a wireless communication module. Objects in a perception scene include a first category and a second category. The first category of objects is equipped with a positioning chip containing pre-stored information about the objects. After the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module. Basic information about the second category of objects is obtained via the information acquisition sensor. After the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene. Based on the interaction priority of the objects, the robot sequentially obtains real-time perception information of the objects via the information acquisition sensor. The real-time perception information is then integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. This allows the robot to quickly obtain the pre-stored information of the positioning chip for the first category of objects via the wireless communication module, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by basic information acquisition and interaction priority utilizes resources efficiently, improves response speed and task execution fluency, and reduces computational complexity and energy consumption.

[0150] According to an embodiment of the present invention, a storage medium corresponding to the environmental perception method of an embodied robot is also provided, wherein the storage medium includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the environmental perception method of the embodied robot described above.

[0151] Since the processing and functions implemented by the storage medium of this embodiment basically correspond to the embodiments, principles and examples of the aforementioned method, for any details not fully described in this embodiment, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.

[0152] The present invention employs an embodied robot comprising an information acquisition sensor and a wireless communication module. Objects in a perception scene include a first category and a second category. The first category of objects is equipped with a positioning chip containing pre-stored information about the objects. After the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module. Basic information about the second category of objects is obtained via the information acquisition sensor. After the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene. Based on the interaction priority of the objects, the robot sequentially obtains real-time perception information of the objects via the information acquisition sensor. The real-time perception information is then integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. This allows the robot to quickly obtain the pre-stored information of the positioning chip for the first category of objects via the wireless communication module, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by basic information acquisition and interaction priority utilizes resources efficiently, improves response speed and task execution fluency, and reduces computational complexity and energy consumption.

[0153] According to an embodiment of the present invention, a computer program product corresponding to the environment perception method of an embodied robot is also provided. The computer program product includes a computer program, and when the computer program product is processed and executed, the steps of the environment perception method of the embodied robot are implemented.

[0154] Since the processing and functions implemented by the computer program product of this embodiment basically correspond to the embodiments, principles and examples of the aforementioned method, for any details not fully described in this embodiment, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.

[0155] The present invention employs an embodied robot comprising an information acquisition sensor and a wireless communication module. Objects in a perception scene include a first category and a second category. The first category of objects is equipped with a positioning chip containing pre-stored information about the objects. After the embodied robot enters the perception scene, it obtains the pre-stored information in the positioning chip via the wireless communication module. Basic information about the second category of objects is obtained via the information acquisition sensor. After the embodied robot begins interacting with the objects in the perception scene, it determines the interaction priority of the objects based on the order of interaction with the objects in the perception scene. Based on the interaction priority of the objects, the robot sequentially obtains real-time perception information of the objects via the information acquisition sensor. The real-time perception information is then integrated with the pre-stored information and the basic information to obtain accurate perception results of the objects in the perception scene. This allows the robot to quickly obtain the pre-stored information of the positioning chip for the first category of objects via the wireless communication module, avoiding repeated scanning. Simultaneously, the real-time perception information acquisition driven by basic information acquisition and interaction priority utilizes resources efficiently, improves response speed and task execution fluency, and reduces computational complexity and energy consumption.

[0156] In summary, it is easy for those skilled in the art to understand that, under the premise of no conflict, the above-mentioned advantageous methods can be freely combined and superimposed.

[0157] The foregoing description is merely an embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of the claims.

Claims

1. A method for environmental perception of an embodied robot, characterized in that: The embodied robot is capable of acquiring information about objects in a perception scene through its own information acquisition sensors and wireless communication module; the objects in the perception scene include first-category objects and second-category objects; the first-category objects are equipped with a positioning chip, and the positioning chip stores pre-stored information about the objects; The method comprises: After the embodied robot enters the perception scene, obtaining pre-stored information in the positioning chip through the wireless communication module; acquiring basic information of the second category of objects through the information acquisition sensor; After the embodied robot begins to interact with the objects in the perception scene, determining the interaction priority of the objects in the perception scene according to the interaction order between the embodied robot and the objects in the perception scene; acquiring, in sequence, real-time perception information of the objects in the perception scene through the information collection sensors according to the interaction priorities of the objects in the perception scene; fusing the real-time perception information with the pre-stored information and the basic information respectively to obtain accurate perception results of the first category of objects and the second category of objects; The pre-stored information of the first type of objects includes: position information, posture information, and material property information; the basic information of the second type of objects includes category labels and posture information.

2. The environment perception method of an embodied robot according to claim 1, characterized in that: Determining the interaction priority of the objects in the perception scene according to the interaction sequence between the embodied robot and the objects in the perception scene includes: obtaining an expected contact time between the embodied robot and an object in the perception scene; The interaction priority of the objects is determined according to the order of the expected contact times; wherein, the interaction priority of the object with the earlier expected contact time is higher than that of the object with the later expected contact time.

3. The environment perception method of an embodied robot according to claim 1, characterized in that: The method further comprises: Based on the pre-stored information and in combination with preset logical reasoning rules, the material property information and purpose of the second type of object are inferred.

4. The method for environmental perception of an embodied robot according to claim 1, wherein: The information acquisition sensor includes a camera and a depth camera; Acquiring basic information of the second category of objects through the information acquisition sensor includes: acquiring two-dimensional visual features of the second type of object through the camera; Acquire three-dimensional geometric features of the second type of objects through the depth camera; Based on the two-dimensional visual features and the three-dimensional geometric features, basic information of the second category of objects is identified through a preset deep learning model.

5. The environment perception method of an embodied robot according to any one of claims 1 to 4, characterized in that: Also includes: When the position or state of the first type of object changes, acquiring change information through the information collection sensor; The change information is written into the positioning chip of the corresponding object through the wireless communication module.

6. The environment perception method of an embodied robot according to any one of claims 1 to 4, characterized in that: Also includes: During the process of the embodied robot actively interacting with the first type of object, obtaining detailed information of the object being interacted with; The detailed information is compared and analyzed with the information in the positioning chip of the object being interacted with, and missing or inaccurate information in the positioning chip is corrected.

7. An environmental perception device for an embodied robot, characterized in that: The embodied robot can obtain information about objects in a perception scene through its own information collection sensor and wireless communication module; the objects in the perception scene include first-category objects and second-category objects; The first type of object is equipped with a positioning chip, wherein the positioning chip stores pre-stored information of the object; The environment sensing device includes: an information interaction unit, configured to obtain pre-stored information in the positioning chip through the wireless communication module after the embodied robot enters the perception scene; The information interaction unit is further configured to obtain basic information of the second category of objects through the information collection sensor; a decision unit configured to, after the embodied robot begins to interact with the objects in the perception scene, determine an interaction priority of the objects in the perception scene according to an interaction order between the embodied robot and the objects in the perception scene; a perception unit configured to acquire real-time perception information of objects in the perception scene through the information collection sensors in sequence according to the interaction priorities of the objects in the perception scene; a fusion unit configured to fuse the real-time perception information with the pre-stored information and the basic information, respectively, to obtain accurate perception results of the first category of objects and the second category of objects; The pre-stored information of the first type of objects includes: position information, posture information, and material property information; the basic information of the second type of objects includes category labels and posture information.

8. An embodied robot, characterized in that: include: The environmental perception device of the embodied robot as claimed in claim 7.

9. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the environment perception method of the embodied robot according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Visual interaction method and system for robot with body, terminal and medium

    CN119036461A

  • Human-computer interaction method and device based on body-equipped intelligent agent and body-equipped intelligent agent

    CN119376549A