A data acquisition method, device and medium for a body-equipped intelligent robot
By performing uncertainty analysis and scene understanding in embodied intelligent robots, dynamically adjusting the sampling frequency, and generating multimodal acquisition and recording packages, the problem of insufficient coordination between AI vision-driven and embodied motion control in existing technologies is solved, thereby improving data acquisition efficiency and perception quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 台州安先机器人技术有限公司
- Filing Date
- 2026-04-01
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies have failed to effectively integrate AI vision-driven uncertainty analysis with embodied motion control in data acquisition for embodied intelligent robots. This results in redundant data in areas of low information gain and loss of key details in areas of high uncertainty, making it difficult to achieve proactive perception optimization oriented towards task objectives.
By acquiring continuous image frames of the current environment, uncertainty analysis and scene understanding are performed to generate an initial perception data packet. Information gain is evaluated to select the optimal observation action, the sampling frequency is dynamically adjusted, and an uncertainty-weighted visual stream is output. The embodied robot state information is then fused to form a multimodal acquisition and recording packet.
It improves the data acquisition efficiency and perception quality of embodied intelligent robots in complex environments, provides data support with high consistency and high signal-to-noise ratio, and enhances environmental adaptability and long-term learning ability.
Smart Images

Figure CN122253191A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot perception and control technology, and in particular to a data acquisition method, device and medium for an embodied intelligent robot. Background Technology
[0002] In recent years, with the deep integration of artificial intelligence and perception technology, AI vision has become a core support for achieving environmental understanding and autonomous decision-making. Current mainstream data acquisition methods generally rely on fixed-frequency image sensor sampling mechanisms, combined with preset paths or pose adjustment strategies based on simple feedback to acquire environmental information. Some studies have added high-level vision tasks such as semantic segmentation and object detection to enhance scene understanding capabilities, while other works have attempted to couple SLAM (Simultaneous Localization and Mapping) with visual perception to improve the adaptability of embodied robots to dynamic or complex scenes.
[0003] Existing technologies have significant limitations in dealing with environmental uncertainties and dynamic changes. During the data acquisition phase, they typically employ a constant sampling frequency, failing to consider the differences in confidence levels of visual information across different regions or time periods. This leads to redundant data in low-information-gain areas and the potential loss of crucial details in high-uncertainty areas due to insufficient sampling. The generation of observation actions largely relies on predefined rules or offline-trained models, lacking online joint evaluation of the current perception state and potential information gains, making it difficult to achieve proactive perception optimization oriented towards task objectives. Particularly in embodied intelligence scenarios, embodied robots need to actively explore the environment through their own movements to obtain the optimal observation perspective. However, existing methods fail to effectively integrate the collaborative mechanism between AI vision-driven uncertainty analysis and embodied motion control, limiting the improvement of perception efficiency and the overall intelligence level of the system. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a data acquisition method for embodied intelligent robots to solve the problem of collaborative optimization of observation actions and embodied motion control under uncertain environments.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a data acquisition method for an embodied intelligent robot, comprising: acquiring continuous image frames of the current environment and performing uncertainty analysis and scene understanding on the real-time image to generate an initial perception data packet; evaluating the information gain of the initial perception data packet to generate a candidate observation action set, and selecting the optimal observation action from the candidate observation action set to generate observation action decision data; generating embodied robot control instructions based on the observation action decision data, and controlling the embodied robot to move to the target observation pose through the embodied robot control instructions, recording real-time pose state information during the movement, and outputting a stable observation state identifier and pose parameters; performing uncertainty analysis on the real-time image based on the stable observation state identifier and pose parameters, and dynamically adjusting the sampling frequency to obtain an uncertainty-weighted visual stream; performing unified time reference synchronization processing on the uncertainty-weighted visual stream, and fusing the embodied robot state information to output a multimodal acquisition recording packet.
[0008] As a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the steps of acquiring continuous image frames of the current environment, performing uncertainty analysis and scene understanding on the real-time images, and generating an initial perception data packet are as follows.
[0009] Collect continuous image frames of the current environment to form image frame data. Perform noise reduction, brightness normalization and distortion correction on the image frame data to form standardized visual data.
[0010] Based on edge detection, region segmentation and morphological processing, target regions are extracted and spatial regions are divided from standardized visual data, and scene semantic understanding information is output.
[0011] The recognition confidence is calculated based on the feature matching consistency, region boundary stability, and image grayscale change degree of each recognition region in the scene semantic understanding information, and the initial uncertainty distribution information is obtained by performing probability statistics based on the recognition confidence.
[0012] The system collects shooting pose information and encapsulates it together with image frame data, scene semantic understanding information, and initial uncertainty distribution information to form an initial perception data package.
[0013] In a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the steps of evaluating the information gain of the initial perception data packet, generating a candidate observation action set, and selecting the optimal observation action from the candidate observation action set to generate observation action decision data are as follows.
[0014] The scene semantic understanding information, initial uncertainty distribution information, and shooting pose information in the initial perception data packet are read, and the observation position information is obtained by mapping image coordinates to spatial coordinates based on the scene semantic understanding information.
[0015] Based on the shooting pose information, multiple candidate shooting pose information are generated around the area to be observed, and each candidate shooting pose information is associated with the pose movement parameters to form a set of candidate observation actions.
[0016] Based on the candidate observation action set, a virtual observation image is generated by transforming the viewpoint through image frame data and corresponding candidate shooting pose information, and scene understanding processing is performed on the virtual observation image to obtain the prediction uncertainty distribution.
[0017] The information gain value is calculated based on the predicted uncertainty distribution and the initial uncertainty distribution information, and the optimal observation action is selected from the candidate observation actions based on the information gain value;
[0018] The shooting pose information and pose movement parameters corresponding to the optimal observation action are uniformly encapsulated to generate observation action decision data.
[0019] In a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the specific steps for generating control commands for the embodied robot based on observed action decision data are as follows:
[0020] Collect the current pose information of the embodied robot, convert the target shooting pose information and the current pose information in the observation action decision data to the same embodied robot base coordinate system, and output the current pose difference;
[0021] Calculate the pose control quantity based on the current pose difference and pose movement parameters, perform motion trajectory planning based on the pose control quantity, and generate a pose control sequence.
[0022] During the execution of the pose control sequence, real-time pose state information is periodically collected, the residual deviation between the real-time pose and the target pose is calculated, and the pose control sequence is corrected based on the residual deviation, and the embodied robot control commands are output.
[0023] As a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the steps of controlling the embodied robot to move to the target observation pose via embodied robot control commands, recording real-time pose state information during the movement, and outputting a stable observation state identifier and pose parameters are as follows.
[0024] The embodied robot is driven to perform pose adjustment based on the control commands of the embodied robot, and the real-time pose status data of the embodied robot is collected in real time during the pose adjustment process;
[0025] The pose change is calculated based on the real-time pose data of adjacent sampling times. When the pose change is lower than the change threshold within a continuous sampling period, a stable observation state identifier is generated, and the current pose information is extracted as the pose parameter.
[0026] As a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the change threshold is obtained by continuously acquiring multiple sets of pose data when the embodied intelligent robot is in a stationary state, calculating the corresponding pose change based on the pose data at adjacent sampling times, and selecting the change corresponding to the stable fluctuation range after statistical analysis of the pose change.
[0027] As a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the steps of performing uncertainty analysis on the real-time image based on stable observation state identifiers and pose parameters, and dynamically adjusting the sampling frequency to obtain an uncertainty-weighted visual stream, are as follows:
[0028] During the period when the stable observation state indicator is valid, real-time images are continuously acquired to form initial image frame data. Uncertainty analysis is performed on the initial image frame data to generate real-time uncertainty distribution information.
[0029] The overall uncertainty level of the current image is calculated based on the real-time uncertainty distribution information, and the sampling frequency is dynamically adjusted according to the overall uncertainty level to form the adjusted sampling frequency.
[0030] Image frame data is acquired according to the adjusted sampling frequency, and pose parameters and corresponding timestamps are added to form an uncertainty-weighted visual stream.
[0031] As a preferred embodiment of the data acquisition method for the embodied intelligent robot described in this invention, the steps of performing unified time-base synchronization processing on the uncertainty-weighted visual stream and fusing the embodied robot's state information to output a multimodal acquisition record package are as follows:
[0032] Based on the timestamps of each updated image frame in the uncertainty-weighted visual stream, the embodied robot state information corresponding to the timestamps is collected synchronously.
[0033] The updated image frame data is time-aligned with the embodied robot state information to obtain aligned data. The aligned data is then associated and encapsulated with real-time uncertainty distribution information. The captured pose information and pose movement parameters are written into the encapsulation result as identification information of the acquisition process, forming a multimodal acquisition record package.
[0034] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the data acquisition method for an embodied intelligent robot as described in the first aspect of the present invention.
[0035] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the data acquisition method for an embodied intelligent robot as described in the first aspect of the present invention.
[0036] The beneficial effects of this invention are as follows: by adding an AI vision-driven active perception and adaptive acquisition mechanism, the data acquisition efficiency and perception quality of the embodied intelligent robot in complex environments are improved. By synchronizing visual data and robot state information with a unified time base, a structured multimodal acquisition and recording package is formed, which provides highly consistent and high signal-to-noise ratio data support for subsequent AI vision training, environmental modeling and autonomous decision-making. This realizes the leap from "passive recording" to "cognitive-guided acquisition", and enhances the environmental adaptability and long-term learning ability of the embodied intelligent system. Attached Figure Description
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 A flowchart of a data acquisition method for an embodied intelligent robot.
[0039] Figure 2 A flowchart for generating the initial sensing data packet.
[0040] Figure 3 A flowchart for generating observation action decision data.
[0041] Figure 4 This is a flowchart for dynamically adjusting the sampling frequency.
[0042] Figure 5 Generate a curve showing the change in the number of data collection records for different data collection modes of the embodied intelligent robot.
[0043] Figure 6 A comparison chart of the average number of multimodal data acquisition packets for different data acquisition modes of an embodied intelligent robot. Detailed Implementation
[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0046] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0047] Reference Figures 1-6 As one embodiment of the present invention, this embodiment provides a data acquisition method for an embodied intelligent robot, comprising the following steps:
[0048] S1. Acquire continuous image frames of the current environment and perform uncertainty analysis and scene understanding on the real-time image to generate initial perception data packets.
[0049] S1.1. Collect continuous image frames of the current environment to form image frame data. Perform noise reduction, brightness normalization and distortion correction on the image frame data to form standardized visual data.
[0050] Specifically, when acquiring continuous image frames of the current environment, the embodied intelligent robot camera device outputs image frame data arranged in chronological order using continuous imaging, and adds a corresponding timestamp to each image frame data to maintain the consistency of the acquisition order; when performing noise reduction processing on the image frame data, spatial filtering or temporal filtering is performed on the random noise and compressed noise in the image frame data to reduce the impact of noise on subsequent recognition; when performing brightness normalization processing on the image frame data, the brightness is stretched or equalized based on the grayscale distribution of the image frame data to reduce the difference in brightness caused by changes in illumination; when performing distortion correction processing on the image frame data, geometric inverse transformation is performed on radial distortion and tangential distortion according to the pre-stored calibration parameters of the camera device to correct edge stretching and bending deformation, and obtain standardized visual data that corresponds one-to-one with the image frame data.
[0051] S1.2. Based on edge detection, region segmentation and morphological processing, target regions are extracted and spatial regions are divided from standardized visual data, and scene semantic understanding information is output.
[0052] Specifically, edge detection is performed on standardized visual data to obtain object boundary contours and generate edge maps. Then, based on the edge maps, the standardized visual data is segmented into regions to form multiple connected regions and output region label maps. Subsequently, morphological processing is performed on the region label maps to fill holes, break connections, and remove noise spots, resulting in target regions and spatial regions with continuous boundaries and complete areas. After obtaining the target regions and spatial regions, the target region extraction results and spatial region segmentation results are labeled by combining the geometric shape features and relative positional relationships of the target regions and spatial regions. The target region extraction results and spatial region segmentation results are then uniformly encapsulated to output scene semantic understanding information. Scene semantic understanding information refers to the structured environmental semantic description data obtained after parsing the target, region, and spatial relationship of the collected standardized visual data. It is used to characterize the object category, region attributes, spatial relationships, and recognition confidence of each spatial location in the image.
[0053] S1.3. Calculate the recognition confidence based on the feature matching consistency, region boundary stability and image grayscale change degree of each recognition region in the scene semantic understanding information, and obtain the initial uncertainty distribution information by performing probability statistics based on the recognition confidence.
[0054] Specifically, a consistency index is obtained by statistically analyzing the feature matching degree of the same recognition region in continuous image frame data. Simultaneously, a region boundary stability index is obtained by differential calculation of the boundary position changes of the recognition region in continuous image frame data, and a grayscale change index is obtained by statistical analysis of the image grayscale changes within the recognition region. Subsequently, the feature matching consistency index, region boundary stability index, and grayscale change index are normalized and weighted. The weights of the feature matching consistency index, region boundary stability index, and grayscale change index can be determined through offline calibration. On the collected sample data, multiple candidate weight combinations are subjected to grid search or fitted optimization based on minimizing historical recognition errors. The weight combination that minimizes the recognition error or achieves optimal recognition stability is selected as the target weight. The recognition confidence level corresponding to each recognition region is obtained. After obtaining the recognition confidence level, probability statistics are performed based on the recognition confidence level, mapping the recognition confidence level to the corresponding spatial location uncertainty value and arranging them according to image coordinates to obtain initial uncertainty distribution information.
[0055] S1.4. Collect shooting pose information and encapsulate it together with image frame data, scene semantic understanding information and initial uncertainty distribution information to form an initial perception data packet.
[0056] Specifically, when collecting shooting pose information, the current pose information of the embodied intelligent robot is read at the same sampling time when each image frame data is generated as the shooting pose information. A timestamp consistent with the image frame data is added to the shooting pose information to establish a one-to-one correspondence. The shooting pose information includes a set of parameters of the position and attitude state of the camera or robot imaging device in space, which is used to describe the observation perspective and spatial coordinate relationship. After obtaining the shooting pose information, the image frame data, scene semantic understanding information, initial uncertainty distribution information and shooting pose information are written at the field level. The image frame data is written to the image frame field, the scene semantic understanding information is written to the semantic field, the initial uncertainty distribution information is written to the uncertainty field, and the shooting pose information is written to the pose field. All fields are indexed and associated using the same timestamp, and unified encapsulation is completed to form the initial perception data packet.
[0057] S2. Evaluate the information gain of the initial sensing data packet, generate a set of candidate observation actions, select the optimal observation action from the set of candidate observation actions, and generate observation action decision data.
[0058] S2.1. Read the scene semantic understanding information, initial uncertainty distribution information and shooting pose information from the initial perception data packet, and obtain the observation position information based on the scene semantic understanding information by mapping image coordinates to spatial coordinates.
[0059] Specifically, when reading the initial perception data packet, the scene semantic understanding information, initial uncertainty distribution information, and shooting pose information corresponding one-to-one with the image frame data are extracted from the initial perception data packet, and the correspondence between the scene semantic understanding information and the initial uncertainty distribution information in the image coordinates is established. Based on the target object information, spatial region information, and obstacle information in the scene semantic understanding information, the region where the target object to be further observed or the boundary of the spatial region is located in the image frame data is selected as the candidate range. At the same time, the candidate range is narrowed or expanded by referring to the high uncertainty position of the initial uncertainty distribution information within the candidate range, thereby obtaining the region to be observed. The position of the region to be observed in the image coordinates is associated with the shooting pose information, and the position information to be observed is output. The spatial position information is used to generate candidate shooting pose information around the region to be observed in the future.
[0060] S2.2. Based on the shooting pose information, multiple candidate shooting pose information are generated around the area to be observed, and each candidate shooting pose information is associated with the pose movement parameters to form a set of candidate observation actions.
[0061] Specifically, based on the shooting pose information, the current camera extrinsic pose and camera orientation are determined. A candidate observation pose sampling range is constructed in the robot's base coordinate system, centered on the position information to be observed. Within this range, multiple candidate shooting pose information are generated according to different observation azimuth angles, observation distances, and pitch angles. Each candidate shooting pose information includes candidate position coordinates and candidate attitude angle parameters. Reachability checks and collision elimination processing under obstacle position contour constraints are performed on each candidate shooting pose information, retaining those that satisfy kinematic constraints. For each retained candidate shooting pose information, the pose difference from the shooting pose information to the candidate shooting pose information is calculated, and this difference is decomposed into translation and rotation to form pose movement parameters. Each candidate shooting pose information is associated and encapsulated with its corresponding pose movement parameters to form a candidate observation action set.
[0062] It should be noted that obstacle position contour constraint is obtained by performing obstacle detection and spatial boundary extraction on environmental perception data to obtain the spatial range information occupied by the obstacle, and transforming the obstacle spatial range into the robot base coordinate system to form a spatial restriction area that can be used for collision detection. During the candidate shooting pose generation process, the candidate pose and its motion path are judged to see if they enter the spatial restriction area in order to exclude candidate shooting poses that may collide.
[0063] Kinematic constraints are pose-executable limitations determined based on the range of motion of each actuator of the embodied robot, the reachable workspace range, and the attitude adjustment range. After candidate shooting poses are generated, the reachability of the joint range of motion, pose change amplitude, and trajectory execution range corresponding to the candidate poses is judged, and candidate shooting poses that meet the execution capability range are retained.
[0064] S2.3. Based on the candidate observation action set, the virtual observation image is generated by transforming the viewpoint through image frame data and corresponding candidate shooting pose information, and the scene understanding processing is performed on the virtual observation image to obtain the prediction uncertainty distribution.
[0065] Specifically, based on the candidate observation action set, for each candidate observation action, the corresponding candidate shooting pose information is read, and the image frame data and shooting pose information are subjected to geometric viewpoint transformation calculation. Through coordinate transformation relationship, the pixel position of the area to be observed in the image frame data is mapped to the imaging position under the corresponding candidate shooting pose information, thereby generating a virtual observation image that matches the candidate shooting pose information. The virtual observation image is divided into regions, and the spatial change of the recognition confidence of each region in the scene semantic understanding information is statistically analyzed. The prediction uncertainty distribution corresponding to the spatial position of the virtual observation image is obtained by calculating the probability distribution of the region recognition confidence.
[0066] S2.4. Calculate the information gain value based on the predicted uncertainty distribution and the initial uncertainty distribution information, and select the optimal observation action from the candidate observation actions based on the information gain value.
[0067] Specifically, based on the candidate observation action set, the prediction uncertainty distribution and initial uncertainty distribution information corresponding to each candidate observation action are aligned region by region in the same spatial location, and the uncertainty difference between each region is calculated. The uncertainty differences between each region are then weighted and summed according to spatial weights to obtain the information gain value of the corresponding candidate observation action. The spatial weights are set according to the importance of the observed region in the scene semantic understanding information. Subsequently, the information gain values corresponding to all candidate observation actions in the candidate observation action set are compared, and the candidate observation action with the largest information gain value is determined as the optimal observation action. The information gain value reflects the potential contribution of the observation action to improving the perception effectiveness.
[0068] It should be noted that the expression for calculating the information gain value is:
[0069] ;
[0070] in, Indicates the first Information gain of each candidate observation action Indicates the first Initial uncertainty distribution information corresponding to each spatial location Indicates the first The spatial location is executing the first... Distribution of prediction uncertainty after each candidate observation action Indicates the first Spatial weight of each spatial location Indicates the candidate observation action number, Indicates spatial location or pixel region number. This indicates the total number of spatial locations involved in the calculation.
[0071] S2.5. Unify and encapsulate the shooting pose information and pose movement parameters corresponding to the optimal observation action to generate observation action decision data.
[0072] Specifically, the system reads the shooting pose information and pose movement parameters corresponding to the optimal observation action, and sorts the shooting pose information and pose movement parameters at the field level. The spatial position parameters and attitude parameters in the shooting pose information are written into the pose description field, and the movement direction parameters and movement amplitude parameters in the pose movement parameters are written into the motion description field. After the fields are written, the system adds action identification information that matches the corresponding action number in the candidate observation action set to the shooting pose information and pose movement parameters, so that the shooting pose information, pose movement parameters and action identification information form a one-to-one correspondence in the same record structure, generating observation action decision data.
[0073] S3. Based on the observed action decision data, generate embodied robot control commands, and control the embodied robot to move to the target observation pose through the embodied robot control commands. During the movement, record the real-time pose status information and output the stable observation status identifier and pose parameters.
[0074] S3.1. Collect the current pose information of the embodied robot, convert the target shooting pose information and the current pose information in the observation action decision data to the same embodied robot base coordinate system, and output the current pose difference.
[0075] Specifically, when collecting the current pose information of the embodied robot, the spatial position parameters and attitude parameters of the embodied robot at the current moment are read to form the current pose information, and the target shooting pose information in the observation action decision data is also read. The target shooting pose information and the current pose information are respectively transformed into coordinates so that both the target shooting pose information and the current pose information are represented in the robot base coordinate system. The coordinate transformation converts the spatial position parameters and attitude parameters in the sensor coordinates or end effector coordinates into the spatial position parameters and attitude parameters in the robot base coordinate system through known coordinate transformation relationships. After completing the coordinate transformation, the target shooting pose information in the robot base coordinate system and the current pose information in the robot base coordinate system are differentially calculated to obtain the current pose difference, which includes the position difference and the attitude difference.
[0076] S3.2. Calculate the pose control quantity based on the current pose difference and pose movement parameters, perform motion trajectory planning based on the pose control quantity, and generate a pose control sequence.
[0077] Specifically, based on the current pose difference and pose movement parameters, the position difference and attitude difference in the current pose difference are combined with the movement direction parameter and movement amplitude parameter in the pose movement parameters to obtain the pose control quantity used to describe the pose adjustment range of the embodied robot in each control cycle. The pose control quantity is divided into multiple continuous control steps in time sequence. Then, based on the pose control quantity, trajectory interpolation calculation is performed on the pose change process from the current pose information to the target shooting pose information to form a pose control sequence composed of multiple continuous pose control quantities arranged in time sequence.
[0078] S3.3. During the execution of the pose control sequence, real-time pose state information is periodically collected, the residual deviation between the real-time pose and the target pose is calculated, and the pose control sequence is corrected according to the residual deviation, and the embodied robot control command is output.
[0079] Specifically, during the execution of the pose control sequence, the current pose information of the embodied robot is continuously collected to form real-time pose state information. After the real-time pose state information is converted to the robot's base coordinate system, the difference between the real-time pose state information and the target pose is calculated to obtain the residual deviation, which includes spatial position deviation and attitude deviation. Based on the residual deviation, the pose control quantities that have not yet been executed in the pose control sequence are corrected. The corrected pose control quantities are rearranged in chronological order to form an updated pose control sequence. The embodied robot control instructions are generated based on the updated pose control sequence.
[0080] S3.4. Drive the embodied robot to perform pose adjustment based on the embodied robot control commands, and collect the real-time pose status data of the embodied robot in real time during the pose adjustment process.
[0081] Specifically, based on the embodied robot control commands, the embodied robot is driven to gradually complete the spatial position and posture adjustment according to the execution order of each posture control variable in the posture control sequence. During the posture adjustment process, the current posture information of the embodied robot is continuously read, which includes spatial position parameters and posture parameters. Each time the real-time posture state information of the embodied robot is collected, a timestamp of the collection time is added to the corresponding current posture information, so that the current posture information of the embodied robot and the collection time form a one-to-one correspondence. The current posture information of the embodied robot and the timestamp are combined and recorded in chronological order, thereby forming real-time posture data containing continuous time identifiers. Real-time posture data refers to the set of real-time parameters periodically collected during the robot's posture adjustment movement, which describes the robot's current spatial motion state and posture state, and reflects the actual motion execution of the robot at each sampling moment.
[0082] S3.5. Calculate the pose change based on the real-time pose data of adjacent sampling times. When the pose change is lower than the change threshold within a continuous sampling period, generate a stable observation state identifier and extract the current pose information as pose parameters.
[0083] Specifically, two real-time pose data points at adjacent sampling times are read according to the time sequence of the real-time pose data, and the current pose information and corresponding timestamps are extracted from the two real-time pose data points respectively; the pose change is obtained by differential calculation of the current pose information at adjacent sampling times, which includes the change in spatial position parameters and the change in attitude parameters, and the pose change is associated with the timestamp to form a pose change sequence for continuous sampling periods; the continuity of the pose change sequence is judged, and a stable observation state identifier is generated when the pose change corresponding to each sampling period in the continuous sampling period is lower than the change threshold. At the same time, the current pose information is extracted as the pose parameter from the last real-time pose data that meets the continuous sampling period condition.
[0084] It should be noted that the change threshold is obtained by continuously collecting multiple sets of pose data when the embodied intelligent robot is in a stationary state, calculating the corresponding pose change based on the pose data at adjacent sampling times, and then selecting the change corresponding to the stable fluctuation range after statistical analysis of the pose change. The example value is 0.5mm to 5mm.
[0085] S4. Based on the stable observation state identifier and pose parameters, perform uncertainty analysis on the real-time image and dynamically adjust the sampling frequency to obtain the uncertainty-weighted visual stream.
[0086] S4.1. During the period when the stable observation state indicator is valid, continuously acquire real-time images to form initial image frame data, perform uncertainty analysis on the initial image frame data, and generate real-time uncertainty distribution information.
[0087] Specifically, during the period when the stable observation state indicator is valid, the camera device is activated to continuously acquire real-time images and generate initial image frame data in chronological order. At the same time, a corresponding timestamp is added to each initial image frame data to form a temporally continuous initial image frame sequence. Based on the recognition confidence of each pixel region in the initial image frame data in the scene semantic understanding information, the confidence of each spatial position in the initial image frame data is statistically calculated. The spatial distribution of recognition confidence is converted into the uncertainty value of the corresponding pixel region and arranged according to the image coordinate position to form real-time uncertainty distribution information. This real-time uncertainty distribution information can reflect the reliability of the recognition results of different spatial regions in the current image frame data.
[0088] S4.2. Calculate the overall uncertainty level of the current image based on the real-time uncertainty distribution information, and dynamically adjust the sampling frequency according to the overall uncertainty level to form the adjusted sampling frequency.
[0089] Specifically, based on the real-time uncertainty distribution information, the uncertainty values corresponding to each spatial location in the real-time uncertainty distribution information are statistically weighted by region. The uncertainty values of each spatial location are accumulated according to the spatial weight to obtain the overall uncertainty level of the current image. The spatial weight is allocated according to the area range of the observed region in the scene semantic understanding information. Then, the overall uncertainty level is compared with the uncertainty level interval. The corresponding sampling frequency is selected according to the uncertainty level interval in which the overall uncertainty level is located. The current sampling frequency is increased or decreased to obtain an adjusted sampling frequency that matches the current overall uncertainty level.
[0090] It should be noted that the expression for calculating the overall uncertainty level of the current image is:
[0091] ;
[0092] in, This indicates the overall uncertainty level of the current image. Indicates the first Uncertainty values in the real-time uncertainty distribution information corresponding to each spatial location.
[0093] It should be noted that the uncertainty level interval is obtained by statistical analysis of the real-time uncertainty distribution information obtained during the historical acquisition process. The overall uncertainty level calculated in multiple acquisition cycles is sorted according to the numerical value and divided into several stable distribution intervals. The upper and lower boundaries of each stable distribution interval are used as the interval thresholds of the uncertainty level interval, thus forming the uncertainty level interval for dynamic adjustment of the sampling frequency.
[0094] S4.3. Collect updated image frame data according to the adjusted sampling frequency, and attach pose parameters and corresponding timestamps to form an uncertainty-weighted visual stream.
[0095] Specifically, the camera device is controlled to acquire images in real time according to the adjusted sampling frequency. In each sampling period, the corresponding updated image frame data is output. At the same time as acquiring the updated image frame data, the current pose parameters are read. The current pose parameters and the timestamp of the corresponding sampling time are appended to the corresponding updated image frame data, so that each updated image frame data has a one-to-one correspondence with the pose parameters and timestamp. Then, the updated image frame data with the added pose parameters and timestamp are arranged continuously in chronological order. During the arrangement process, the spatial correspondence between the updated image frame data and the real-time uncertainty distribution information is maintained, thereby forming an uncertainty-weighted visual stream containing the characteristics of sampling frequency changes.
[0096] S5. Perform unified time base synchronization processing on the uncertainty-weighted visual stream, and integrate the embodied robot state information to output a multimodal acquisition and recording package.
[0097] S5.1. Based on the timestamps of each updated image frame data in the uncertainty-weighted visual stream, the embodied robot state information corresponding to the timestamps is collected synchronously.
[0098] Specifically, based on the timestamps corresponding to each updated image frame in the uncertainty-weighted visual stream, and using the timestamps as time indexes, the embodied intelligent robot state information at the same sampling moment is retrieved from the embodied intelligent robot operation record. The embodied intelligent robot state information includes pose parameters and real-time pose state information corresponding to the timestamps. Subsequently, a one-to-one correspondence is established between the timestamps and the corresponding embodied intelligent robot state information, and the embodied intelligent robot state information is associated with the image frame data in the uncertainty-weighted visual stream in chronological order, thereby completing the synchronous acquisition of the embodied robot state information corresponding to the timestamps.
[0099] S5.2. Time-align the updated image frame data with the embodied robot state information to obtain aligned data. Associate and encapsulate the aligned data with the real-time uncertainty distribution information, and write the shooting pose information and pose movement parameters as acquisition process identifier information into the encapsulation result to form a multimodal acquisition record package.
[0100] Specifically, the timestamps corresponding to each image frame in the uncertainty-weighted visual stream are read, and the timestamps are matched with the corresponding timestamps recorded in the embodied robot's state information. The image frame data with consistent times are time-aligned with the embodied robot's state information to form aligned data containing image frame data, pose parameters, and real-time pose state information. Subsequently, the aligned data is encapsulated with the real-time uncertainty distribution information under the corresponding timestamp at the field level, so that the image frame data, embodied robot's state information, and real-time uncertainty distribution information establish a spatial and temporal correspondence in the same record structure. The shooting pose information and pose movement parameters recorded in the observation action decision data are written into the encapsulation result as the acquisition process identifier information to form a multimodal acquisition record package.
[0101] It should be noted that, as Figure 5 As shown in the figure, the number of multimodal acquisition record packets generated per unit time in the fixed acquisition mode and the active sensing acquisition mode changes within the acquisition cycle. The upper part is the overall acquisition process curve, and the lower part is a local magnified comparison curve of a selected time interval. As can be seen from the figure, the active sensing acquisition mode generates a higher number of multimodal acquisition record packets in each acquisition cycle than the fixed acquisition mode, and the difference remains stable within local time intervals. This indicates that the observation-action decision-making data generation mechanism can guide the robot to prioritize acquiring high information gain areas, thereby improving data acquisition efficiency.
[0102] Figure 6 The results show the statistical results of the average number of multimodal acquisition record packets generated by the fixed acquisition mode and the active sensing acquisition mode within the same experimental time range. The bar value corresponding to the active sensing acquisition mode is significantly higher than that of the fixed acquisition mode, indicating that more effective acquisition records can be obtained in the same time period through information gain-driven observation action generation and adaptive sampling mechanism, thereby improving the overall data acquisition efficiency.
[0103] This embodiment also provides a computer device applicable to the data acquisition method of an embodied intelligent robot, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the data acquisition method of the embodied intelligent robot as proposed in the above embodiment.
[0104] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0105] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the data acquisition method for implementing an embodied intelligent robot as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0106] In summary, this invention improves the data acquisition efficiency and perception quality of embodied intelligent robots in complex environments by adding an AI vision-driven active perception and adaptive acquisition mechanism. By synchronizing visual data and robot state information with a unified time base, a structured multimodal acquisition and recording package is formed, providing highly consistent and high signal-to-noise ratio data support for subsequent AI vision training, environmental modeling, and autonomous decision-making. This achieves a leap from "passive recording" to "cognitive-guided acquisition," enhancing the environmental adaptability and long-term learning ability of the embodied intelligent system.
[0107] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A data acquisition method for an embodied intelligent robot, characterized in that: include, Collect continuous image frames of the current environment and perform uncertainty analysis and scene understanding on the real-time image to generate an initial perception data packet; The information gain of the initial sensing data packet is evaluated to generate a set of candidate observation actions. The optimal observation action is selected from the set of candidate observation actions to generate observation action decision data. Based on the observed action decision data, control commands for the embodied robot are generated, and the embodied robot is controlled to move to the target observation pose through the control commands. During the movement, real-time pose status information is recorded, and stable observation status identifiers and pose parameters are output. Uncertainty analysis of real-time images is performed based on stable observation state indicators and pose parameters, and the sampling frequency is dynamically adjusted to obtain uncertainty-weighted visual streams. Uncertainty-weighted visual streams are synchronized using a unified time reference and integrated with embodied robot state information to output a multimodal acquisition and recording package.
2. The data acquisition method for the embodied intelligent robot as described in claim 1, characterized in that: The process of acquiring continuous image frames of the current environment, performing uncertainty analysis and scene understanding on the real-time image, and generating an initial perception data packet involves the following steps: Collect continuous image frames of the current environment to form image frame data. Perform noise reduction, brightness normalization and distortion correction on the image frame data to form standardized visual data. Based on edge detection, region segmentation and morphological processing, target regions are extracted and spatial regions are divided from standardized visual data, and scene semantic understanding information is output. The recognition confidence is calculated based on the feature matching consistency, region boundary stability, and image grayscale change degree of each recognition region in the scene semantic understanding information, and the initial uncertainty distribution information is obtained by performing probability statistics based on the recognition confidence. The system collects shooting pose information and encapsulates it together with image frame data, scene semantic understanding information, and initial uncertainty distribution information to form an initial perception data package.
3. The data acquisition method for the embodied intelligent robot as described in claim 2, characterized in that: The steps for evaluating the information gain of the initial sensing data packet, generating a candidate observation action set, and selecting the optimal observation action from the candidate observation action set to generate observation action decision data are as follows. The scene semantic understanding information, initial uncertainty distribution information, and shooting pose information in the initial perception data packet are read, and the observation position information is obtained by mapping image coordinates to spatial coordinates based on the scene semantic understanding information. Based on the shooting pose information, multiple candidate shooting pose information are generated around the area to be observed, and each candidate shooting pose information is associated with the pose movement parameters to form a set of candidate observation actions. Based on the candidate observation action set, a virtual observation image is generated by transforming the viewpoint through image frame data and corresponding candidate shooting pose information, and scene understanding processing is performed on the virtual observation image to obtain the prediction uncertainty distribution. The information gain value is calculated based on the predicted uncertainty distribution and the initial uncertainty distribution information, and the optimal observation action is selected from the candidate observation actions based on the information gain value; The shooting pose information and pose movement parameters corresponding to the optimal observation action are uniformly encapsulated to generate observation action decision data.
4. The data acquisition method for the embodied intelligent robot as described in claim 1, characterized in that: The specific steps for generating control commands for the embodied robot based on observed action decision data are as follows. Collect the current pose information of the embodied robot, convert the target shooting pose information and the current pose information in the observation action decision data to the same embodied robot base coordinate system, and output the current pose difference; Calculate the pose control quantity based on the current pose difference and pose movement parameters, perform motion trajectory planning based on the pose control quantity, and generate a pose control sequence. During the execution of the pose control sequence, real-time pose state information is periodically collected, the residual deviation between the real-time pose and the target pose is calculated, and the pose control sequence is corrected based on the residual deviation, and the embodied robot control commands are output.
5. The data acquisition method for the embodied intelligent robot as described in claim 1, characterized in that: The process involves controlling the android to move to the target observation pose using android control commands, recording real-time pose state information during the movement, and outputting a stable observation state identifier and pose parameters. The specific steps are as follows: The embodied robot is driven to perform pose adjustment based on the control commands of the embodied robot, and the real-time pose status data of the embodied robot is collected in real time during the pose adjustment process; The pose change is calculated based on the real-time pose data of adjacent sampling times. When the pose change is lower than the change threshold within a continuous sampling period, a stable observation state identifier is generated, and the current pose information is extracted as the pose parameter.
6. The data acquisition method for the embodied intelligent robot as described in claim 5, characterized in that: The change threshold is obtained by continuously collecting multiple sets of pose data when the embodied intelligent robot is in a stationary state, calculating the corresponding pose change based on the pose data at adjacent sampling times, and selecting the change corresponding to the stable fluctuation range after statistical analysis of the pose change.
7. The data acquisition method for the embodied intelligent robot as described in claim 5, characterized in that: The uncertainty analysis of the real-time image based on stable observation state indicators and pose parameters, and the dynamic adjustment of the sampling frequency, to obtain an uncertainty-weighted visual flow, are described in the following steps. During the period when the stable observation state indicator is valid, real-time images are continuously acquired to form initial image frame data. Uncertainty analysis is performed on the initial image frame data to generate real-time uncertainty distribution information. The overall uncertainty level of the current image is calculated based on the real-time uncertainty distribution information, and the sampling frequency is dynamically adjusted according to the overall uncertainty level to form the adjusted sampling frequency. The updated image frame data is collected according to the adjusted sampling frequency, and the pose parameters and corresponding timestamps are added to form an uncertainty-weighted visual stream.
8. The data acquisition method for the embodied intelligent robot as described in claim 7, characterized in that: The process of performing unified time-base synchronization processing on the uncertainty-weighted visual stream, fusing embodied robot state information, and outputting a multimodal acquisition and recording package involves the following specific steps. Based on the timestamps of each updated image frame in the uncertainty-weighted visual stream, the embodied robot state information corresponding to the timestamps is collected synchronously. The updated image frame data is time-aligned with the embodied robot state information to obtain aligned data. The aligned data is then associated and encapsulated with real-time uncertainty distribution information. The captured pose information and pose movement parameters are written into the encapsulation result as identification information of the acquisition process, forming a multimodal acquisition record package.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the data acquisition method for the embodied intelligent robot according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the data acquisition method for the embodied intelligent robot according to any one of claims 1 to 8.