Action control method and device based on physical reference, equipment and medium

By using a joint calibration mechanism that generates physical reference baselines and model-estimated baselines, the scale-normalized point cloud output by the visual geometry basic model is converted into a physical space point cloud. This solves the problem that existing VLA models cannot obtain the true physical scale and improves the robot's spatial understanding and operational accuracy in complex scenes.

CN121374618APending Publication Date: 2026-01-23SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202511808372.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing VLA models cannot obtain three-dimensional spatial information with real physical scale significance, making it difficult for robots to achieve precise physical spatial control based on visual reconstruction results.

Method used

By acquiring command information, multi-view images, and pose information of active components, target segmentation information and scale-normalized point clouds are generated. Combined with the calibration information of the view acquisition unit and the pose information of active components, a physical reference baseline is determined, a scale calibration factor is generated, the scale-normalized point cloud is converted into a physical space point cloud, the three-dimensional relative position of the target object with respect to the active components is extracted, and action commands are generated.

Benefits of technology

It realizes the conversion from scale-normalized point clouds output by visual geometric basic models to physical spatial point clouds with real physical scale meaning, thereby improving the robot's spatial understanding and operational accuracy in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374618A_ABST
    Figure CN121374618A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot visual perception and motion control, and discloses a motion control method and device based on physical reference, equipment and a medium, and the method comprises the steps: obtaining instruction information, a multi-view image and movable assembly pose information; processing the multi-view image according to the instruction information to generate target segmentation information; generating a scale normalization point cloud and a model estimation baseline; determining a physical reference baseline and generating a scale calibration factor; converting the scale normalization point cloud into a physical space point cloud by using a scale calibration factor; extracting a three-dimensional relative position of the target object relative to the movable component in combination with the target segmentation information; an action instruction is generated based on the multi-modal input. According to the method, physical scale alignment of the point cloud is realized through physical reference baseline calibration, so that a visual reconstruction result has real space significance, an accurate action instruction is generated, and the robot space understanding and operation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot visual perception and motion control, and in particular to a motion control method and device based on physical reference, equipment and medium. BACKGROUND

[0002] Although the existing visual-language-motion (VLA) model has made some progress in the research of robot operation tasks, there are still significant technical bottlenecks. The traditional VLA model relies on two-dimensional visual input and language instructions for learning and reasoning, and its perception ability is mainly limited to the image plane, making it difficult to truly reflect the depth and scale information of objects in three-dimensional space. This two-dimensional perception method is prone to target recognition bias and spatial reasoning errors when performing operation tasks that require spatial positioning and geometric constraints, resulting in unstable performance of robots in fine control tasks such as grasping, obstacle avoidance, or assembly.

[0003] To compensate for the shortcomings of two-dimensional vision, some research has introduced three-dimensional perception methods, such as obtaining depth data through RGB-D cameras or laser radars. However, such solutions rely on high-cost hardware devices, and the depth sensor is extremely sensitive to environmental lighting, reflective surfaces, and other conditions, making it difficult to ensure data stability and accuracy. At the same time, traditional three-dimensional reconstruction algorithms based on geometric optimization are usually computationally intensive and slow, making it difficult to meet the real-time and high-frequency response requirements of robot control systems.

[0004] In recent years, the visual geometry foundation model (VGGT) has provided a new direction for robot three-dimensional reconstruction. The VGGT model can quickly reconstruct the three-dimensional point cloud of a scene from multiple-view images, significantly improving processing efficiency. However, the output point cloud is a scale-normalized result that lacks physical scale significance. That is, the model can recover the geometric structure of the scene, but it cannot determine the real physical distance between points. For VLA robots that rely on accurate physical spatial coordinates for motion control, this scale ambiguity will directly lead to operation errors, making it impossible to achieve precise grasping or positioning tasks. SUMMARY

[0005] The main purpose of the present application is to provide a motion control method, device, equipment and storage medium based on physical reference, aiming to solve the technical problem that the existing VLA model cannot obtain three-dimensional spatial information with real physical scale significance, making it difficult for robots to achieve precise physical space control based on visual reconstruction results.

[0006] To achieve the above purpose, the present application provides a motion control method based on physical reference, comprising: obtaining instruction information, multi-view images obtained by at least two view acquisition units, and activity component pose information of an activity component; According to the instruction information, the multi-view image is processed to generate target segmentation information of a target object; The multi-view image is processed to generate a scale normalized point cloud and a model estimation baseline of a scene; According to the calibration information of the at least two view acquisition units and the active component pose information, a physical reference baseline between the at least two view acquisition units is determined; Based on the model estimation baseline and the physical reference baseline, a scale calibration factor is generated; The scale normalized point cloud is converted into a physical space point cloud by applying the scale calibration factor; The three-dimensional relative position of the target object relative to the active component is extracted in combination with the target segmentation information and the physical space point cloud; Based on the three-dimensional relative position, the multi-view image, the instruction information, and the active component pose information, an action instruction is generated.

[0007] Further, to achieve the above object, the application provides a physical reference based action control device, comprising: A multi-modal data acquisition module is configured to acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of an active component; A semantic segmentation module is configured to process the multi-view images according to the instruction information to generate target segmentation information of a target object; A three-dimensional reconstruction module is configured to process the multi-view images to generate a scale normalized point cloud and a model estimation baseline of a scene; A physical baseline determination module is configured to determine a physical reference baseline between the at least two view acquisition units according to the calibration information of the at least two view acquisition units and the active component pose information; A scale calibration calculation module is configured to generate a scale calibration factor based on the model estimation baseline and the physical reference baseline; A point cloud scale conversion module is configured to convert the scale normalized point cloud into a physical space point cloud by applying the scale calibration factor; A relative position extraction module is configured to extract a three-dimensional relative position of the target object relative to the active component in combination with the target segmentation information and the physical space point cloud; An action generation module is configured to generate an action instruction based on the three-dimensional relative position, the multi-view image, the instruction information, and the active component pose information.

[0008] Further, to achieve the above object, the present application also provides a computer device, comprising a memory, a processor, and a physical reference based action control program stored in the memory and executable on the processor, wherein the physical reference based action control program, when executed by the processor, implements the steps of the physical reference based action control method.

[0009] Further, to achieve the above object, the present application also provides a computer readable storage medium, wherein the storage medium stores a physical reference based action control program, and the physical reference based action control program, when executed by a processor, implements the steps of the physical reference based action control method.

[0010] Beneficial effects: The present application relates to the technical field of robot visual perception and action control, and discloses a physical reference based action control method, device, equipment and medium, which comprises the following steps: obtaining instruction information, multi-view images and active component pose information; generating target segmentation information by processing the multi-view images according to the instruction information; generating a scale normalized point cloud and a model estimated baseline based on the multi-view images; determining a physical reference baseline according to the calibration information and the active component pose information; generating a scale calibration factor based on the model estimated baseline and the physical reference baseline; converting the scale normalized point cloud into a physical space point cloud by applying the scale calibration factor; extracting a three-dimensional relative position of a target object relative to the active component by combining the target segmentation information and the physical space point cloud; and generating an action instruction based on the three-dimensional relative position, the multi-view images, the instruction information and the active component pose information. The present application converts the scale normalized point cloud output by a visual geometry basic model into a physical space point cloud with real physical scale significance by introducing a joint calibration mechanism of a physical reference baseline and a model estimated baseline, realizes physical alignment of three-dimensional geometric features, and enables the VLA model to generate an accurate and executable action instruction based on multi-modal input, thereby significantly improving the spatial understanding ability and operation precision of a robot in a complex scene. BRIEF DESCRIPTION OF DRAWINGS

[0011] The present application will be further described below in combination with the drawings and embodiments, wherein: Figure 1 An application environment schematic diagram of the physical reference based action control method in an embodiment of the present application; Figure 2 A flow schematic diagram of the physical reference based action control method in an embodiment of the present application; Figure 3 A functional module schematic diagram of the physical reference based action control device in a preferred embodiment of the present application; Figure 4 A structure schematic diagram of the computer device in an embodiment of the present application; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The motion control method based on physical reference provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain command information, multi-view images, and active component pose information from the client; process the multi-view images based on the command information to generate target segmentation information; generate scale-normalized point clouds and model estimation baselines based on the multi-view images; determine physical reference baselines based on calibration information and active component pose information; generate scale calibration factors based on the model estimation baselines and physical reference baselines; apply the scale calibration factors to convert the scale-normalized point clouds into physical space point clouds; extract the three-dimensional relative position of the target object relative to the active components by combining the target segmentation information and physical space point clouds; and generate action commands based on the three-dimensional relative position, multi-view images, command information, and active component pose information. This invention introduces a joint calibration mechanism of physical reference baselines and model estimation baselines to convert the scale-normalized point clouds output by the visual geometric model into physical space point clouds with real physical scale meaning, achieving physical alignment of three-dimensional geometric features. This enables the VLA model to generate accurate and executable action commands based on multimodal inputs, thereby significantly improving the robot's spatial understanding and operational accuracy in complex scenes. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the motion control method based on physical reference provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the motion control method based on physical reference proposed in this invention includes the following steps: S10: Acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of the active component; In this embodiment, the acquisition of instruction information is the starting point of the entire operation process, and its content includes three parts of task target, object description and operation semantics. The task target is used to clarify the operation intention of the robot, such as "grabbing the target object" and "moving to the specified position"; the object description provides semantic labels related to the target, which is used for subsequent image segmentation and feature matching; and the operation semantics describes the execution mode, such as "grabbing from the left side" and "rotating 90 degrees and placing". The instruction information can be input by a natural language input interface, a remote operation terminal or a task scheduling system. When the natural language is input, the key semantic units are first extracted through the semantic analysis module, and then converted into structured task parameters through the embedding model. When the remote terminal is input, the structured parameter set is directly transmitted for system analysis.

[0016] The acquisition of multi-view images is completed by at least two view acquisition units. Each acquisition unit includes a camera module, an imaging control circuit and a time synchronization module. The configuration of two or more acquisition units should cover different viewing angles of the scene, so that the three-dimensional structure of the target object has sufficient parallax information when reconstructed. The acquisition units achieve frame-level time alignment through synchronous trigger signals, thereby avoiding image mismatch caused by motion or delay. After synchronous acquisition, the image data is stored in a cache module and then transmitted to an image processing module for preprocessing. Preprocessing includes color equalization, distortion correction and noise suppression to improve the geometric accuracy of subsequent reconstruction.

[0017] The acquisition of the pose information of the active component is based on the robot body sensing system, including angle sensors, accelerometers, gyroscopes and joint encoders. The pose information covers the position coordinates and attitude matrix of the active component, which is used to describe the current spatial state of the robot. The pose information is transmitted to the control unit through a real-time communication interface to ensure synchronization with image acquisition time. The acquisition process adopts a combination of periodic sampling and time sequence marking, and through a unified timestamp system, the pose information is corresponded with the image acquisition frame to realize the spatio-temporal alignment of multi-source data.

[0018] Image data, instruction information and pose information are converged in the data fusion module. The system is associated through a unique task identifier to ensure that the multi-modal data of the same task corresponds in time and semantic dimensions. When data fusion, the instruction information provides operation semantic guidance to determine the target object category and the region of interest; the multi-view image provides environmental visual data; and the pose information provides spatial reference to establish the coordinate mapping relationship between the robot and the visual scene. The data structure formed in this way provides an input basis for subsequent three-dimensional point cloud reconstruction and scale calibration.

[0019] Multi-view images can be acquired using a binocular camera system, with each camera unit mounted on the robot's end effector or a fixed support. The baseline distance is dynamically adjusted according to the operational scenario. For scenarios requiring wide-area observation, a tri- or quad-camera configuration can be used to enhance depth estimation accuracy. The image acquisition system achieves microsecond-level time alignment through a high-precision clock synchronization chip, ensuring that each frame is captured at the same time.

[0020] Command information can be input via a human-computer interaction terminal, and the natural language parsing module converts the input statement into task parameters. For example, if the input is "grab the red square on the desktop," the system parses it as follows: operation type "grab," target category "square," color feature "red," and location region "desktop." After parsing, a structured command vector is generated and transmitted to the control module.

[0021] The pose information of the active components is read in real time through the controller interface. Each joint encoder outputs an angle signal, which is used to calculate the end-effector pose matrix using kinematics. If the active components are equipped with an IMU module, attitude drift can be further corrected to achieve high-precision pose estimation. The time synchronization between pose data and image acquisition frames is achieved through hardware interrupts, ensuring that the pose state is recorded at the same time.

[0022] In different implementation environments, the data acquisition system can be adapted to different sensing devices. For example, in complex lighting environments, HDR cameras can be used to extend the dynamic range; in space-constrained scenarios, the field of view can be expanded through mirror reflection. For mobile platforms, pose information can be fused with visual odometry to ensure synchronization and accuracy.

[0023] This embodiment acquires instruction information, multi-view images, and pose information of active components simultaneously from multiple sources, enabling the system to form complete perceptual input before task execution. This approach achieves unified spatiotemporal alignment of task semantics, visual information, and spatial pose without increasing hardware complexity, providing a highly consistent data foundation for subsequent 3D reconstruction, scale calibration, and motion generation. This significantly improves the spatial perception accuracy and task response speed of robot operations.

[0024] S20, according to the instruction information, process the multi-view image to generate target segmentation information of the target object; In this embodiment, the instruction information plays a semantic guiding role in the visual processing stage. Its content includes the category features, appearance description, and task semantic constraints of the operation target. The system first performs natural language parsing on the input instruction information, structuring the statement into keyword vectors, such as object category (e.g., "cup," "nut"), attribute features (e.g., "red," "metal"), and operation type (e.g., "grab," "move"). After parsing, a task semantic embedding vector is generated to guide the search for the target region in the image.

[0025] The processing of multi-view images is premised on obtaining clear scene features. After normalization, each view image is sent to a visual feature extraction module to extract color histograms, edge information, and depth-related features. To enable semantic features and visual features to be matched in the same representation space, semantic vectors and image features are aligned in a fusion module, and the correlation weight between each pixel and a semantic keyword is calculated through an attention mechanism to determine the preliminary regional distribution of the target object.

[0026] After completing semantic alignment, the system enters the target detection and segmentation phase. This phase uses a pre-trained image segmentation model to perform pixel-level segmentation operations. The input includes multi-view image features guided by semantics and corresponding semantic vectors, and the output is a pixel-level segmentation mask of the target object. After the mask is generated, morphological filtering is performed to remove isolated noise points, and boundary smoothing is performed to improve contour continuity. For multi-view input, the segmentation module performs mask projection and weighted overlapping regions in each view result to generate a set of target masks with three-dimensional consistency.

[0027] Finally, the fusion module calculates the pixel intersection over union based on the target mask set of all views and outputs the target segmentation information of the target object. This segmentation information includes the pixel region coordinates, region confidence, and cross-view correspondence of the object in each view, providing spatial constraint input for subsequent point cloud mapping and three-dimensional position extraction.

[0028] The multi-modal Transformer structure can be used to realize semantic-guided image segmentation. Instruction information is generated into semantic embeddings by a language encoder, and images are extracted into multi-layer features by convolution or visual Transformer. The two features are fused in an interactive attention layer, and the attention distribution is used to guide pixel selection. This structure can dynamically adjust the attention range under different task semantics to achieve adaptive target segmentation.

[0029] A dual-channel network structure can also be used to achieve fast processing. One channel inputs multi-view images to extract low-level spatial features, and the other channel inputs semantic vectors to generate guidance signals. In the fusion layer, the outputs of different channels are adjusted through a gating mechanism. This approach has an advantage in robot systems with high real-time requirements.

[0030] This embodiment combines semantic instruction information and multi-view images for segmentation processing. The system can accurately locate the task target in a complex scene and avoid target misidentification caused by relying solely on visual features. Semantic constraints enable the model to quickly focus on task-related areas in a multi-object environment, and multi-view consistency ensures the geometric stability of the segmentation results in space, providing accurate target area references for subsequent three-dimensional point cloud mapping and position extraction.

[0031] S30, processing the multi-view images to generate a scale-normalized point cloud and a model estimation baseline of the scene; In this embodiment, the multi-view images are further used to generate a scale-normalized point cloud and a model estimation baseline of the scene after semantic processing. The generation of the scale-normalized point cloud relies on the spatial geometric constraints between different views, aiming to restore the three-dimensional geometric structure of the scene without relying on physical scales. The image of each view is first input into a feature extraction network to extract multi-layer spatial features, including key point positions, texture gradients, and depth clues. After feature extraction, the system uses a disparity estimation module to calculate the matching relationship between corresponding points between adjacent views, constructs a disparity volume through a matching cost volume, and then aggregates and optimizes the cost to obtain a dense depth map. The depth maps of each view are fused in a unified virtual coordinate system, and a scale-normalized point cloud is formed through triangular reconstruction. This point cloud contains relative spatial shape and structure, but its distance information is in a dimensionless state, only maintaining geometric proportionality.

[0032] The generation of the model estimation baseline is an estimated scale reference automatically derived from the reconstruction process. The system calculates the virtual baseline length between views according to the geometric consistency of the point cloud, and estimates the relative scale proportion within the model through the average of the key point triangulation error and the re-projection error. This model estimation baseline is not a real physical length, but a relative reference that reflects the strength of geometric constraints between cameras in the point cloud generation process. By comparing with the subsequent physical reference baseline, it can be used to calculate the scale calibration factor to realize the mapping from virtual scale to real physical scale.

[0033] During the generation process, the system records the confidence weight of each point, which is used to describe the stability of triangulation. Low-confidence points can be filtered out in the post-processing stage to reduce false point clouds caused by matching errors. For scenes with dynamic objects or light changes, the system uses a temporal consistency detection module to identify and remove motion artifact points, thereby improving the stability and repeatability of the reconstruction.

[0034] A Transformer-based visual geometry foundation model can be used to generate a scale-normalized point cloud. This model inputs multi-view images into a unified attention encoder, and the cross-view features are fused through a self-attention mechanism to explicitly capture the spatial correspondence between multiple views. The output point cloud coordinates are automatically normalized within the model, so that the reconstruction results of any scene are mapped to a fixed scale space, thereby eliminating the influence of scale inconsistency between different tasks.

[0035] A geometry reconstruction method based on structured light beam adjustment can also be used. By matching key points in different views, a sparse three-dimensional point set is constructed, and then local optimization is iteratively solved to minimize the re-projection error, generating a three-dimensional model with relative scales. This method is suitable for mechanical assembly or fine operation scenes that require high-precision geometric recovery.

[0036] The embodiment generates a scale normalized point cloud and a model estimated baseline from multi-view images. The system can recover the scene three-dimensional geometry without a depth sensor, providing a reliable reference for subsequent physical scale calibration. This approach reduces hardware costs, simplifies system structure, and ensures the integrity and accuracy of the reconstructed geometry, enabling the robot to have a basic ability of real space geometry reasoning.

[0037] S40, according to the calibration information of the at least two view acquisition units and the active component pose information, determining the physical reference baseline between the at least two view acquisition units; In the embodiment, the determination process of the physical reference baseline is based on the joint calculation of the calibration information of the at least two view acquisition units and the active component pose information. The calibration information includes two parts: intrinsic parameters and extrinsic parameters. The intrinsic parameters are used to describe the imaging characteristics of each view acquisition unit, including focal length, principal point coordinates and radial distortion coefficients; the extrinsic parameters are used to define the spatial position and orientation relationship between the view acquisition units in the unified world coordinate system, including rotation matrix and translation vector. The active component pose information provides the spatial state of the system as a whole at the time of acquisition, including the position vector and attitude quaternion of the active component in the reference coordinate system.

[0038] The system first reads the calibration file of each view acquisition unit, and uses the intrinsic parameter matrix for the camera projection model and the extrinsic parameter matrix to establish the relative pose relationship between the view acquisition units. Then, through the pose information of the active component, the spatial position change of each view acquisition unit at the actual operation time is determined, so that the coordinate system of the acquisition unit can be updated in real time to match the motion state of the current mechanical structure.

[0039] Next, the position difference vector of the two view acquisition units in the global coordinate system is calculated, and this is used as the geometric definition of the physical reference baseline. The length of this difference vector is the physical baseline length between the two acquisition units, and its direction vector represents the baseline direction between the views. In a multi-view scene, the system will calculate the baseline matrix between all view acquisition units, and extract the baseline with the highest geometric stability as the final physical reference baseline.

[0040] To improve accuracy, the system will use the pose change of the active component for dynamic correction. When the active component performs an action, the acquisition unit may have a slight displacement or angular change, at which time the spatial transformation compensation is performed according to the real-time pose data to ensure that the length of the physical reference baseline remains consistent. If there is a baseline drift, the system estimates the stable physical baseline value by regressing the historical pose sequence using the least squares fitting method.

[0041] The calibration information of the view acquisition unit can be obtained through stereo vision calibration. Specifically, during the calibration phase, the calibration board is placed in different poses, and multiple sets of images are simultaneously acquired by two view acquisition units. Zhang Zhengyou's calibration algorithm is used to solve for the camera's intrinsic and extrinsic parameters. The obtained calibration parameters are stored in a calibration matrix file, which is read during runtime to restore the spatial geometric relationship of the view acquisition unit.

[0042] The pose information of the moving components is provided by inertial measurement units and joint angle sensors mounted on the robotic arm or motion platform. Before each image acquisition, the control system acquires real-time pose data and transforms the camera coordinate system to the global coordinate system through a homogeneous transformation matrix to ensure that the calculated baseline position is consistent with the robot's motion.

[0043] This embodiment determines the physical reference baseline by combining the calibration information of the view acquisition unit with the pose information of the moving components. The system can automatically establish the real spatial relationship between the view acquisition units without the need for external ranging equipment, achieving a scale-consistent physical mapping. This method significantly improves the accuracy of subsequent 3D point cloud and spatial positioning calculations, enabling the robot to perform spatial operations at a physical scale and avoiding the scale drift problem caused by relying solely on model estimation.

[0044] S50, Based on the model, estimate the baseline and the physical reference baseline, and generate a scale calibration factor; In this embodiment, the generation of the scale calibration factor depends on the numerical relationship between the model estimation baseline and the physical reference baseline. The model estimation baseline originates from the relative scale generated from multi-view images during virtual space reconstruction, while the physical reference baseline originates from actual hardware calibration and pose measurement. Together, they constitute the scale mapping constraint between virtual and physical spaces. The system first reads the values ​​of the model estimation baseline and the physical reference baseline, and performs format normalization within the same data structure to ensure consistency in units and precision.

[0045] Subsequently, a data validity check is performed to examine the two sets of baseline values ​​for missing data, abrupt changes, or noise contamination. When abnormal data deviating from the expected range is detected, median filtering and a moving average algorithm are used for smoothing correction to eliminate interference caused by short-term sensor errors or visual reconstruction drift. This processing yields valid baseline values, which the system then uses to calculate the ratio of the physical reference baseline to the model-estimated baseline. This ratio is the initial estimate of the scale calibration factor, reflecting the scaling relationship between the virtual and physical spaces.

[0046] To prevent scale factor fluctuations caused by single-calculation errors, the system introduces a dynamic smoothing mechanism based on the initial estimate. This mechanism combines historical scale factor sequences and calculates new calibration factors using weighted moving averages or exponentially decaying weights. The weight allocation can be adaptively adjusted based on the stability of the time series. The smoothed scale calibration factor can maintain stable output under multi-frame input conditions, avoiding scale jitter caused by single-frame errors.

[0047] The system further introduces a numerical range verification step to check the reasonableness of the generated scale calibration factor. If the value exceeds the preset physical range (e.g., 0.5 to 2 times), a re-estimation mechanism will be triggered to recalculate the factor by tracing back the baseline data of the previous few frames, ensuring that the scale mapping result is always within a physically interpretable range.

[0048] Finally, the system stores the finalized scale calibration factor in the calibration parameter table and directly calls it in the subsequent point cloud coordinate mapping and 3D position calculation process, thereby achieving scale consistency transformation from virtual scene data to physical space coordinate system.

[0049] The scale calibration factor can be generated using a linear scaling method, which involves dividing the physical reference baseline length by the model-estimated baseline length to obtain the scaling factor. This method is simple to implement and suitable for stable dual-camera systems or fixed camera architectures.

[0050] Alternatively, a nonlinear optimization method based on robust estimation can be employed. This method uses baseline data obtained from multiple frames of sampling as the observation input, constructs an objective function to minimize the reprojection error, and iteratively solves for the optimal scale calibration factor. This approach is more stable in noisy scenarios and can effectively resist the influence of anomalous data.

[0051] This embodiment achieves scale unification from virtual reconstruction space to physical space by combining the model estimation baseline with the physical reference baseline to calculate the scale calibration factor, thus solving the scale ambiguity problem in the visual geometric model. This method does not rely on additional sensors; it obtains high-precision scale mapping based solely on multi-view images and existing calibration data, significantly improving the physical accuracy of robot spatial manipulation and path planning, and providing a stable geometric scale reference for point cloud accuracy correction and 3D motion generation.

[0052] S60, apply the scale calibration factor to convert the scale-normalized point cloud into a physical space point cloud; In this embodiment, the conversion process from scale-normalized point cloud to physical space point cloud is a crucial step in mapping a virtual point cloud without physical scale to a coordinate system with actual spatial distance meaning. This process uses a scale calibration factor as a scaling parameter to perform a linear scaling transformation on each 3D coordinate point in the point cloud. The system first reads the current value of the scale calibration factor and performs multiplication operations on all coordinate components in the scale-normalized point cloud to obtain the scale-transformed point cloud coordinates. This operation is performed in a unified coordinate system to ensure that the spatial geometry remains consistent, only changing the distribution range of the point cloud within a physical unit of length.

[0053] After coordinate scaling, the system remaps the transformed point cloud to the physical coordinate system used by the robot control. This coordinate system is typically aligned with the world coordinate system of the robot arm base or end effector. To ensure the continuity of coordinate transformation, the system introduces a homogeneous coordinate transformation matrix to transform the scaled point cloud from the virtual reconstructed coordinate system to the robot's global coordinate system, achieving coordinate consistency through rotation and translation parameters. If the point cloud contains data frames from different sources, an attitude alignment algorithm is used to correct registration errors between multiple point cloud frames, ensuring the geometric consistency of the point cloud in physical space.

[0054] The system then performs boundary consistency and data integrity checks. Boundary consistency checks detect whether the physical extent of the point cloud matches the camera's view frustum; if point cloud drift or abnormal expansion is detected, a recalibration process is triggered. Data integrity checks calculate the point cloud density distribution and neighborhood voxel fill rate, removing isolated and noisy points. Point clouds that pass the verification are marked as valid data and proceed to the optimization phase.

[0055] During the optimization phase, the system utilizes spatial smoothing and neighborhood fitting algorithms to correct the continuity of the point cloud surface. Weighted least squares is used to perform local plane fitting on points within the neighborhood. By adjusting the plane normal direction and point weights, the point cloud surface becomes smoother while preserving structural features. This process effectively reduces small numerical drifts after scaling and restores the coherence of the spatial topology.

[0056] A direct scaling transformation can be used to apply the scale calibration factor to the 3D coordinates of the point cloud. For example, for each coordinate point P(x, y, z) in the point cloud, the transformation yields a new coordinate P′(kx, ky, kz), where k is the scale calibration factor. This method is simple to implement, computationally efficient, and suitable for real-time reconstruction and control scenarios.

[0057] Rigid body transformations can also be performed simultaneously with coordinate transformations. By superimposing rotation and translation matrices after the scaling matrix, point cloud coordinates are synchronously mapped from the reconstructed coordinate system to the robot coordinate system. This approach is particularly suitable for multi-view point cloud fusion scenarios, effectively reducing geometric deviations caused by different camera coordinate systems.

[0058] This embodiment applies a scale calibration factor to the scale-normalized point cloud, enabling the system to achieve an accurate mapping from virtual reconstructed space to physical space, thus giving the reconstructed data practical distance significance. This process ensures that the robot can perform precise control based on real dimensions when performing grasping, ranging, or spatial planning tasks. This method eliminates the limitation of scale uncertainty in the output of VGGT-type models, achieving high-precision, low-cost spatial geometric calibration without increasing sensor complexity.

[0059] S70, combining the target segmentation information and the physical space point cloud, extract the three-dimensional relative position of the target object with respect to the active component; In this embodiment, the extraction of three-dimensional relative position is a key step in achieving precise robot control. Its core lies in locating the target object's position in the physical space point cloud using target segmentation information and calculating the spatial difference vector between the two by combining the spatial pose information of the active components. The system first registers and maps the target segmentation information with the physical space point cloud. The target segmentation information consists of pixel-level masks output by the image segmentation model. Through the relationship between the camera's intrinsic and extrinsic parameters and depth mapping, the three-dimensional coordinates corresponding to each pixel are mapped to the physical space point cloud, generating a target object point cloud cluster.

[0060] After generating the point cloud cluster of the target object, the system performs clustering and noise removal on the cluster. Density clustering (DBSCAN) or voxel downsampling methods are used to remove isolated and redundant points to ensure the continuous and stable spatial distribution of the point cloud. Subsequently, the geometric centroid coordinates of the point cloud cluster are calculated, serving as the representative position of the target object in physical space. To enhance positioning accuracy, the system can simultaneously calculate the center point of the point cloud bounding box or the structural center determined based on principal component analysis (PCA), thus achieving robust position estimation even when there is non-uniform sampling on the target surface.

[0061] The spatial pose information of the active components is provided in real time by the robot control system, including position vectors and translation matrices. The system spatially aligns the coordinate system of the end effector of the active components with the global coordinate system of the point cloud, and transforms the pose coordinates of the active components into coordinates in a spatial reference system consistent with the point cloud through a coordinate transformation matrix. Using the two sets of aligned coordinates, the system calculates the three-dimensional difference vector between the centroid coordinates of the target object and the current position of the active component, and represents the three-dimensional relative position of the target object with respect to the active component in the form of this vector.

[0062] This three-dimensional relative position comprises three components: x, y, and z, representing the spatial offset of the target object in the front-back, left-right, and up-down directions, respectively. To avoid the propagation of measurement errors, the system performs spatial filtering and outlier correction before output, using Kalman filtering or low-pass filtering to smooth the position difference sequence and obtain a stable relative position signal. This relative position data will then be used in the motion command generation stage to provide accurate geometric input for robot control.

[0063] The mapping between target segmentation information and physical point cloud can be achieved through direct geometric projection. The system calculates the 3D coordinates of each pixel using a pinhole imaging model based on the depth information of each pixel in the camera coordinate system, projects the corresponding pixels of the target mask onto the point cloud data, and thus extracts the target point cloud clusters.

[0064] Semantic point cloud fusion can also be used to transform the segmentation results into a semantic label map and directly introduce semantic labels during the point cloud generation stage, so that each point cloud point carries target category information. This allows for direct filtering of target object point cloud clusters at the point cloud level, reducing post-processing steps.

[0065] When extracting the centroid of a target object, a weighted centroid calculation strategy can be used. This involves assigning weights to different points based on point cloud density or reflection intensity to prevent centroid shift caused by sparse point cloud regions. In robotic arm or mobile platform scenarios, an IMU-assisted coordinate synchronization mechanism can be used to ensure that the pose of active components is consistent with the point cloud acquisition timestamp, avoiding spatial errors caused by time asynchrony.

[0066] When there are multiple target objects, the system can use the nearest neighbor matching strategy to filter out the most relevant target object based on the current movement direction or instruction information of the active component, so as to achieve accurate relative positioning in a multi-target environment.

[0067] This embodiment extracts the 3D position by combining target segmentation information with physical space point clouds. The system can accurately obtain the positional relationship of the target object in physical space without relying on additional depth sensors. This approach fuses image semantic information with geometric spatial information, preserving the target accuracy of visual recognition while adding maneuverability in 3D space. Therefore, when performing grasping, assembly, or manipulation tasks, the robot can generate more precise motion paths and control decisions based on actual spatial geometric relationships.

[0068] S80, based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component, generate an action instruction.

[0069] In this embodiment, the generation of motion commands is a crucial process for achieving closed-loop control of the vision-language-motion system. Its core lies in the multimodal fusion of 3D spatial information, visual perception results, language commands, and the motion states of active components to generate motion control outputs that conform to physical constraints and task semantics. The system first integrates 3D relative positions, multi-view images, command information, and the pose information of active components to construct a structured multimodal input vector. This input vector includes spatial coordinate features, visual features, semantic features, and dynamic features, which originate from the physical space calculation, visual encoding module, language parsing module, and robot sensing unit in the preceding steps, respectively.

[0070] After the input data is integrated, the system uses a feature encoding network to perform unified vectorization representation of different modalities. Visual features are extracted by convolutional neural networks or Transformer structures to capture local morphology and global semantic information in the image; linguistic features are obtained by contextual embedding representation of instruction text to parse action type, direction, target object, and constraints; the spatial feature encoding module transforms the three-dimensional relative position and the pose of active components into a geometric relationship matrix to describe the spatial relationship between the target object and the robot's current state.

[0071] The encoded multimodal features are input into the fusion layer of the VLA model. The fusion layer calculates the association weights between different modalities through an attention mechanism, achieving multi-layered interaction of semantic, visual, and spatial features. The fused contextual representation contains the semantic intent and spatial constraint information of the task objective. Based on this, the system generates preliminary action instructions through a decoder. The action instructions are represented in parameterized form and typically include translation, rotation, pose adjustment, and end-effector commands.

[0072] The generated preliminary motion commands then enter the verification module for a feasibility assessment. This module verifies whether the motion exceeds the range of motion of the active components through physical simulation or kinematic equations, and whether it will cause collisions, joint overloads, or path conflicts. For commands that do not meet the physical constraints, the system executes a correction strategy, recalculating the pose path parameters through a constraint optimization algorithm to ensure that the final motion commands meet the executable conditions.

[0073] The validated motion commands structurally include the target point, trajectory path, and velocity parameters, and can be directly transmitted to the robot controller for execution. Simultaneously, to ensure the stability of the commands, the system performs dynamic smoothing and delay prediction compensation on the output motion, maintaining its continuity and accuracy under high-speed execution conditions.

[0074] A VLA model based on the Transformer architecture can be used for motion command generation. The model's input consists of four branches: a visual encoding branch processes multi-view images, a language branch parses command information, a spatial branch encodes 3D relative positions, and a pose branch takes the pose matrix of the active components as input. The model fuses features from different modalities through multi-head attention, ultimately outputting a structured sequence of motion parameters.

[0075] Alternatively, a graph neural network can be used to model the target object, active components, and their spatial relationships as nodes and edges, and generate action vectors through a message passing mechanism. This approach is suitable for complex operation scenarios that require high spatial accuracy and structural constraints.

[0076] For high-frequency control scenarios, a hierarchical control mechanism can be adopted, breaking down action commands into policy-level and execution-level components. The policy-level model determines the action type and target position based on linguistic semantics and visual information, while the execution-level model calculates precise trajectory parameters based on the three-dimensional relative position and the state of active components, achieving continuous control with millisecond-level response.

[0077] In multi-object scenarios, target priority encoding can be introduced, which combines the semantic strength of language instructions with the confidence of visual features, and generates action instructions for the corresponding target through attention weighting, thereby enabling the robot to perform multi-object operations in the same task.

[0078] This embodiment generates motion commands by multimodal fusion of 3D relative position, multi-view images, command information, and pose information of active components, achieving a precise combination of visual semantics and physical control. This approach not only understands the operational intent of verbal commands but also outputs precise actions that conform to physical constraints by combining spatial geometric relationships, solving the problem that traditional vision-language models cannot execute actual physical operations. This enables the robot to adaptively manipulate target objects in complex environments, improving operational accuracy, response speed, and motion stability.

[0079] This invention relates to the field of robot visual perception and motion control technology, and discloses a motion control method, device, equipment, and medium based on physical reference. The method includes: acquiring instruction information, multi-view images, and pose information of active components; processing multi-view images based on instruction information to generate target segmentation information; generating scale-normalized point clouds and model estimation baselines based on multi-view images; determining physical reference baselines based on calibration information and pose information of active components; generating scale calibration factors based on model estimation baselines and physical reference baselines; applying scale calibration factors to convert scale-normalized point clouds into physical space point clouds; extracting the three-dimensional relative position of the target object relative to the active components by combining target segmentation information and physical space point clouds; and generating motion commands based on the three-dimensional relative position, multi-view images, instruction information, and pose information of active components. This invention introduces a joint calibration mechanism between the physical reference baseline and the model estimation baseline, converting the scale-normalized point cloud output by the visual geometric model into a physical space point cloud with real physical scale meaning. This achieves physical alignment of three-dimensional geometric features, enabling the VLA model to generate accurate and executable motion commands based on multimodal inputs, thereby significantly improving the robot's spatial understanding and operational accuracy in complex scenes.

[0080] In one embodiment, step S20 above includes: S201, parse the instruction information to extract keywords describing the target object; S202, based on the target object description keywords, the multi-view image is preprocessed to enhance image quality; S203, input the preprocessed multi-view image and the target object description keywords into the pre-trained target detection and segmentation model; S204, Generate a pixel-level segmentation mask for the target object using the target detection and segmentation model; S205, perform morphological optimization on the pixel-level segmentation mask to remove noise; S206, Based on the optimized pixel-level segmentation mask, generate target segmentation information for the target object.

[0081] In this embodiment, the instruction information is used to define the target object and operational intent. First, text parsing is performed to obtain keywords describing the target object. The parsing process starts with word segmentation and part-of-speech tagging, combined with named entity recognition and dependency analysis, to extract a set of phrases representing entity names, appearance attributes, functional attributes, spatial constraints, and operational constraints. To avoid omissions due to differences in expression, synonym expansion and domain thesaurus mapping are introduced to unify "blue wrench," "blue wrench," and "wrench class in blue tools" into the same entity category, and weights are recorded to reflect their relevance to the task intent. The parsing output is in the form of keyword-attribute-weight triples, serving as conditional input for subsequent visual processing.

[0082] The preprocessing of multi-view images uses target object description keywords as conditions for adaptive enhancement. First, geometric correction is performed, utilizing intrinsic parameters and distortion coefficients to achieve distortion removal and principal point alignment. Next, photometric alignment is performed, unifying color distribution under different cameras and exposures through brightness curve matching and white balance estimation. Then, enhancement operators are dynamically selected based on color or texture attribute keywords (such as "blue," "metallic luster," and "cylindrical handle"). For color attributes, color channel reweighting and color constancy are performed; for metallic and texture attributes, local contrast enhancement and multi-scale detail enhancement are performed. To reduce noise, spatiotemporal consistency filtering is introduced, performing joint bilateral filtering between frames of the same viewpoint or between adjacent viewpoints to preserve edge structure. For objects with insufficient size, super-resolution reconstruction is triggered to improve the segmentation input resolution. The preprocessed output is a multi-view image sequence that has undergone geometric and photometric normalization and includes semantically guided enhancement.

[0083] The pre-trained object detection and segmentation model receives joint input of preprocessed multi-view images and keywords describing the target object. The image branch extracts hierarchical features through convolution or a visual Transformer, while the text branch encodes keyword triples into semantic vectors. The multimodal fusion layer employs cross-attention, using the semantic vectors as queries to focus on image feature regions consistent with the target object's attributes. The detection head outputs candidate boxes and class confidence scores, while the segmentation head generates pixel-level segmentation masks within the candidate regions and outputs pixel confidence maps. To adapt to multiple views, weight-sharing branch-based per-view inference is used, and inter-view consistency constraints suppress isolated false detections. When the same object exhibits scale and pose differences across different views, the fusion layer maintains mask stability through scale-invariant positional encoding and a multi-scale feature pyramid.

[0084] After pixel-level segmentation masks are generated, they enter the morphological optimization process. First, confidence-guided binarization is performed on low-confidence boundaries, with the adaptive threshold jointly determined by mask confidence statistics and category priors. Then, opening operations are performed to remove scattered noise and small artifacts, and closing operations are performed to fill boundary gaps. Connected component analysis is used to eliminate regions whose area and aspect ratio do not conform to the object's priors, while retaining connected components with high overlap with the detection box. Hole regions are filled to avoid non-physical cavities during subsequent 3D reconstruction. At the boundaries, thinning and edge alignment are superimposed, and guided filtering or conditional random fields are used to refine strong gradient edges, ensuring the mask boundaries conform to the real object contours. Epipolar correspondences are established between multiple viewpoints using fundamental or essential matrices, and consistency voting is performed on the masks for matched viewpoint pairs to eliminate occasional errors from a single viewpoint. The above optimization outputs a stable, continuous, and cross-viewpoint consistent pixel-level segmentation mask set.

[0085] The target segmentation information is organized in the data structure as a mask set indexed by viewpoint, category labels, mask confidence statistics, and geometric auxiliary information. The geometric auxiliary information includes the mask bounding box, minimum bounding rectangle, principal axis direction, and pixel area, which are used for subsequent integration with the physical point cloud. To ensure seamless integration with subsequent processing chains, the target segmentation information is explicitly labeled with timestamps and viewpoint identifiers, ensuring correct alignment with 3D reconstruction and pose data in both time and viewpoint dimensions. This output is directly used in subsequent stages to filter target object point cloud clusters from the physical point cloud, forming a continuous mapping chain from linguistic semantics to visual masks and then to 3D geometry.

[0086] This embodiment establishes a language-to-visual constraint channel through parsing and synonym expansion based on instruction information, enabling the segmentation of the region of interest to align with the task intent. Figure 1 To achieve the following: improve the segmentability of multi-view inputs under color, texture and scale variations by geometric and photometric normalization and superimposed semantic guidance enhancement; enhance the robustness of pixel-level segmentation masks under occlusion and viewpoint difference conditions by text-image cross-attention and viewpoint consistency constraints; and output a set of masks with well-fitting boundaries and controlled noise by confidence-guided morphological optimization and boundary refinement.

[0087] In one embodiment, step S30 above includes: S301, Input the multi-view image into the pre-trained VGGT model; S302, Extract depth features from multi-view images using the VGGT model; S303, Based on the depth features, reconstruct the three-dimensional point cloud data of the scene; S304, The three-dimensional point cloud data is subjected to scale normalization processing to generate a scale-normalized point cloud; S305, the relative distance between the view acquisition units corresponding to the multi-view images is estimated as the model estimation baseline.

[0088] In this embodiment, multi-view images are acquired by at least two view acquisition units within the same time window. Image frames are labeled with viewpoint identifiers and timestamps for cross-viewpoint registration. During the input phase, geometric and photometric normalization is first performed. Distortion correction and principal point correction are performed using the intrinsic parameters of each view acquisition unit, and the resolution is resampled to the input size supported by the VGGT model. Color channels undergo luminance and white balance alignment to ensure radiometric consistency across viewpoints. Subsequently, the multi-view image sequence and its corresponding viewpoint identifier are encoded into tensor sequences and fed into the pre-trained VGGT model. Pixel positions are embedded using two-dimensional position encoding, and viewpoint identities are distinguished using viewpoint embedding vectors, enabling cross-viewpoint self-attention to establish long-range geometric associations between different observations of the same scene.

[0089] The VGGT model internally uses a visual Transformer as its framework. It first extracts local appearances through patch embedding layers, then aggregates information through alternating stacked self-attention layers (cross-view and single-view) to generate depth features containing pixel-to-pixel and view-to-view geometric constraints. Each pixel location in the depth features carries both appearance descriptions and geometric cues, including disparity cues, occlusion indicators, and matching confidence. The model's decoder further regresses the relative depth or disparity for each view, as well as the relative pose parameters between views, forming the minimum geometric set required for reconstruction. To suppress ambiguity caused by weak and repetitive textures, the decoder outputs a confidence map for subsequent weighting.

[0090] 3D point cloud data is obtained through triangulation reconstruction. The reconstruction process begins with backprojection of pixel coordinates from each viewpoint. An intrinsic parameter matrix is ​​used to convert the pixel coordinates into normalized ray directions, and then intersection is performed to solve for the spatial point coordinates, combining the relative poses between viewpoints. For the same spatial point observed by multiple viewpoints, robust estimation aggregation based on reprojection error is used to eliminate outliers with reprojection errors exceeding the threshold. Weighted least squares calculation is then performed using the confidence map as weights to obtain geometrically consistent point coordinates and observation counts. To improve density and completeness, cross-viewpoint propagation is triggered for low-confidence regions from a single viewpoint, using epipolar lines to guide the search for consistent pixels in other viewpoints to complete the corresponding spatial points. Foreground-background consistency checks are performed on occluded boundary regions to prevent erroneous triangulation into the point cloud. The resulting point set, carrying spatial coordinates, normal approximations, and visibility statistics, constitutes the 3D point cloud data.

[0091] Scale-normalized point clouds are obtained by uniformly scaling the 3D point coordinates and relative poses to eliminate the inherent scale ambiguity in reconstruction. The translation vectors output by VGGT only determine direction and lack absolute scale. The normalization strategy fixes the selected geometric scale to a unit value to establish a uniform scale. Alternatively, the distance between the camera centers of any adjacent viewpoint pair can be normalized to one. This is done by solving for the two camera centers C1 and C2 from the relative pose and calculating the distance norm, then scaling all 3D point coordinates and translation vectors according to the inverse of this norm. Another option is to normalize the median depth of the point cloud to the main camera. This is done by statistically analyzing the depth distribution of 3D points projected onto the main camera coordinate system and calculating the scaling factor accordingly. During normalization, rotation remains unchanged, and only translation and point coordinates are linearly scaled. The type of normalization scale and scaling factor used are recorded to ensure traceability for subsequent docking with the physical scale. After normalization, each 3D point in the output coordinate system has a consistent relative scale, forming a scale-normalized point cloud.

[0092] The model estimation baseline is defined as the relative distance between view acquisition units corresponding to multi-view images at the model scale, used to characterize the length scale of camera geometry in a normalized scale. The calculation process first selects viewpoint pairs for triangulation, recovers the camera center positions in a unified reference coordinate system using the relative pose output by VGGT, and then calculates the Euclidean distance between the two centers to obtain the baseline value. To reduce bias caused by viewpoint selection, distances can be independently calculated on multiple pairs of adjacent viewpoints and then weighted averaged. The weights are determined by the number of in-registration points and the average reprojection error of each pair. When there are faults or strong occlusions in the scene, viewpoint pairs with abnormal reprojection errors are removed to improve robustness. The model estimation baseline shares the same normalized scale as the scale-normalized point cloud and is the direct input for subsequent comparison with the physical reference baseline and generation of scale calibration factors.

[0093] This embodiment reconstructs dense and consistent 3D point cloud data without adding additional depth hardware by utilizing depth features of cross-view attention encoding and confidence-weighted triangulation; it eliminates scale uncertainty by uniform scaling to form a scale-normalized point cloud that can be docked with any physical scale; and it obtains the model estimation baseline by extracting the camera center distance from relative pose stability, providing a quantitative basis for subsequent comparison with the physical reference baseline.

[0094] In one embodiment, step S40 above includes: S401, Load the pre-calibrated calibration information of the at least two view acquisition units; S402, Real-time acquisition of the active component pose information of the active component; S403, Based on the calibration information and the pose information of the active component, calculate the Euclidean distance between the at least two view acquisition units; S404, verify the rationality of the calculated Euclidean distance; S405, the Euclidean distance after the rationality verification is passed is used as the physical reference baseline between the at least two view acquisition units.

[0095] In this embodiment, when loading the calibration information of at least two pre-calibrated view acquisition units, the intrinsic parameter matrix, radial and tangential distortion coefficients, initial installation pose, and calibration timestamp of each view acquisition unit are read. The intrinsic parameters are used for back-projection from pixels to camera coordinates, the distortion coefficients are used for imaging geometric correction, and the initial installation pose provides the static installation relationship of the view acquisition unit relative to the moving component. If the calibration information includes the mapping relationship of external reference coordinates, the homogeneous transformation of the reference coordinates is also read for subsequent coordinate unification. For multiple batches of calibration or online correction records, a set of calibration information consistent with the current hardware state is selected using version number and timestamp, and the residual ranges of intrinsic and extrinsic parameters are verified to avoid misusing incorrect calibration results.

[0096] When acquiring the pose information of active components in real time, the 3D position and orientation quaternions or Euler angles of the active components in a unified reference coordinate system are read from the kinematic chain or localization module, along with a timestamp and validity identifier. The pose information is used to advance the view acquisition unit mounted on the active component from its initial installation pose to its current true pose. To reduce the impact of communication latency, a time synchronization strategy is introduced to align the pose information of the active components with the time axis of the multi-view images, and, if necessary, interpolate the pose to a point consistent with the current image frame. If there are flexible supports or minor deflections on the active component, fine-tuning terms can be introduced into the pose information to compensate for installation deviations.

[0097] When calculating the Euclidean distance between at least two view acquisition units based on calibration information and active component pose information, the camera center of each view acquisition unit is first determined by the initial installation pose value in the active component coordinate system. Then, the camera center is unified to the reference coordinate system using the active component pose information. The camera center position is represented by three-dimensional coordinates in the same metric space, and the norm of the difference vector between the two camera center coordinates gives the Euclidean distance. If there are multiple pairs of view acquisition units, the distance is calculated one-to-one with the viewpoint used for the current reconstruction, and the accompanying information for each calculation is retained, including the calibration version number used, pose timestamp, interpolation method, and coordinate unification link. To mitigate single-measurement jitter, a sliding window statistical method is applied to the Euclidean distances at consecutive time points, outputting the median as an instantaneous estimate, and recording the proportion of outliers within the window for subsequent rationality verification.

[0098] To verify the reasonableness of the calculated Euclidean distance, several independent tests are introduced. The first is a geometric constraint test, comparing the Euclidean distance with the mechanical structure design value or manufacturing tolerance range; if the deviation exceeds the allowable range, it is considered a mismatch. The second is a reprojection consistency test, using the Euclidean distance and the camera's relative attitude to derive the theoretical parallax range, and comparing its consistency with the matching parallax distribution of multi-view images; when the theoretical range deviates significantly from the observed distribution, it is marked as an anomaly. The third is a temporal continuity test, analyzing the first and second differences of the Euclidean distance within the sliding window; if abrupt changes occur that do not conform to the constraints of the maximum speed and maximum acceleration of the active component, it is considered unreasonable. A reasonableness indicator is formed through joint scoring of multiple tests; when a single test fails while the others pass, a conservative strategy can be adopted to reduce the confidence weight but not immediately remove the component, in order to reduce false alarms.

[0099] When using the Euclidean distance, after passing the rationality verification, as the physical reference baseline between at least two view acquisition units, the distance value, reference coordinate identifier, statistical information involved in the verification, and confidence score are retained. This physical reference baseline is aligned with the model-estimated baseline to the same timestamp set, forming a one-to-one data pair, providing stable input for subsequent scale calibration factor calculations. To reduce the impact of high-frequency jitter on downstream processes, the physical reference baseline can be low-pass filtered at the output, while retaining the unfiltered original records for traceability. If the system contains more than two view acquisition units, a graph optimization strategy can be selected to jointly estimate multiple physical reference baselines on the graph structure. The baseline corresponding to the viewpoint pair involved in the reconstruction is then used as the physical reference baseline required for the current stage, thereby improving overall consistency. At extreme moments of viewpoint geometric degradation, such as when two cameras are almost co-located or the lines of sight are almost collinear, a degradation marker is triggered, and the physical reference baseline from the previous moment is reused, while the confidence score is reduced, waiting for the geometric conditions to recover before updating.

[0100] This embodiment precisely couples calibration information with the pose information of active components under the same reference coordinates, enabling the direct calculation of Euclidean distance from the current real camera center position and the formation of a physical reference baseline. By leveraging the joint verification of geometric constraints, reprojection consistency, and temporal continuity, deviations caused by installation errors, sensing delays, and transient noise are significantly reduced. Under multi-view and multi-time conditions, sliding window statistics and low-pass output are introduced to ensure that the physical reference baseline possesses both real-time performance and stability.

[0101] In one embodiment, step S50 above includes: S501, Read the values ​​of the model estimation baseline and the physical reference baseline to obtain the baseline value; S502, check the data validity of the baseline value and output the valid baseline value; S503, Based on the effective baseline values, calculate the ratio of the physical reference baseline to the model-estimated baseline, and generate an initial value for the scale calibration factor; S504, verify the rationality of the numerical range of the initial value of the scale calibration factor, and output the verified initial value of the scale calibration factor; S505, Perform smoothing filtering on the verified initial value of the scale calibration factor to generate a smoothed scale calibration factor; S506, Based on historical scale calibration factor data, the smoothed scale calibration factor is dynamically adjusted to generate the adjusted scale calibration factor. S507, store and output the adjusted scale calibration factor as the final scale calibration factor.

[0102] In this embodiment, when reading the values ​​of the model estimation baseline and the physical reference baseline, the model estimation baseline and the physical reference baseline with timestamps are first obtained from the reconstruction module and the calibration-pose link, and unified to the same length unit and the same reference time axis. To avoid cross-module accuracy loss, double-precision floating-point storage can be agreed upon, and the original measurement accuracy and source identifier are retained, forming a baseline value record containing the value, unit, timestamp, and source. If there are multiple redundant sources, a candidate set is constructed with time proximity and source confidence as weights to prepare diverse inputs for subsequent validity checks.

[0103] When checking the validity of baseline values, the model-estimated baseline and the physical reference baseline are subjected to interval constraints, missing value checks, and outlier detection, respectively. Interval constraints provide upper and lower limits based on structural design and assembly tolerances; missing value checks cover null values, non-numeric values, and missing units; outlier detection uses robust statistics with a sliding window (e.g., median absolute deviation) to identify instantaneous jumps. A scale consistency check is also added to the model-estimated baseline, comparing the current value with the equivalent baseline derived from the disparity distribution of the nearest time points; a geometric continuity check is added to the physical reference baseline, limiting the first and second differences to not exceeding preset velocity and acceleration limits. Valid baseline values ​​are output through a comprehensive multi-check scoring system, along with accompanying confidence levels and rejection reason reports, ensuring that only qualified data is included in the ratio calculation.

[0104] When calculating the ratio of the physical reference baseline to the model-estimated baseline based on effective baseline numerical values, a simultaneous pairing strategy and a time interpolation strategy are executed in parallel. If the two timestamps are inconsistent, a time-aligned model-estimated baseline is first obtained through linear or spline interpolation, and then the ratio is calculated to generate the initial value of the scale calibration factor. To avoid numerical instability caused by zero or extremely small denominators, a minimum denominator threshold and a lower bound are introduced. When multiple pairs of valid samples exist, a robust estimate of the ratio is obtained by weighting by confidence level to suppress the impact of single-point anomalies on the initial value. The ratio calculation simultaneously retains the reference window size, the number of participating samples, and the residuals as supporting information for subsequent validation.

[0105] When verifying the rationality of the initial value range of the scale calibration factor, three types of constraints are superimposed. The numerical boundary constraint derives the scale range based on the imaging baseline, focal length, and typical target distance; values ​​exceeding the boundary are directly judged as failing. The historical distribution constraint refers to the quantile interval of historical scale calibration factor data, limiting the deviation of the initial value relative to the historical median. The closed-loop consistency constraint utilizes the recent 3D reprojection error, substituting the initial value into the reconstruction-projection link to evaluate pixel and 3D errors; if the error exceeds a threshold, it is marked as failing. For those that pass, the verified initial value of the scale calibration factor is output; for those that fail, an automatic rollback to the previous stable value is triggered, and an alarm is recorded.

[0106] When smoothing the validated initial values ​​of the scale calibration factor, an exponential moving average or a one-dimensional Kalman filter is selected based on the application's time delay budget and dynamic requirements. The exponential moving average strikes a balance between suppressing high-frequency jitter and maintaining response speed by setting an attenuation factor; the Kalman filter adaptively estimates process noise and observation noise in a model with the scale calibration factor as the state and the observations as the initial values, improving stability under weak textures or sudden illumination changes. The smoothed scale calibration factor is used as a real-time estimate, simultaneously outputting the filter gain, estimation variance, and update step size for transparent traceability.

[0107] When dynamically adjusting the smoothed scale calibration factor based on historical scale calibration factor data, slow drift compensation and adaptive weighting are introduced. Slow drift compensation monitors and estimates the offset trend within a long-term window, fine-tuning the current value in small steps to offset systematic deviations caused by temperature drift, mechanical loosening, and lens zoom. Adaptive weighting adjusts the adjustment coefficient based on image texture intensity, feature matching confidence, and camera pose baseline geometry. The adjustment magnitude is reduced when texture is weak or geometry is degraded, while tracking performance is improved when geometry is good and confidence is high. This process yields the adjusted scale calibration factor, along with a current condition label and adjustment gain, ensuring background awareness during subsequent use.

[0108] When storing and outputting the adjusted scale calibration factor as the final scale calibration factor, the numerical value, timestamp, confidence level, filtering and adjustment parameters, and historical window statistics are written to an immutable log and online shared memory. This allows for immediate consumption by downstream modules, while also providing version numbers and rollback mechanisms, enabling the control chain to quickly revert to a previous stable version when an anomaly is detected.

[0109] This embodiment establishes a multi-level assurance chain from reading baseline values ​​to the final scale calibration factor output. First, it uses validity checks to eliminate unqualified inputs, then it uses ratio robust estimation and range verification to constrain the physical rationality, then it uses smoothing filtering to suppress high-frequency fluctuations, uses dynamic adjustment to compensate for slow drift, and finally uses versioned storage and atomic output to ensure consistency.

[0110] In one embodiment, step S60 above includes: S601, apply the scale calibration factor to the coordinates of each point in the scale-normalized point cloud, perform multiplication operations, and generate the scale-transformed point cloud coordinates; S602, Map the scale-transformed point cloud coordinates to the robot coordinate system to generate an initial physical space point cloud; S603, verify the boundary consistency and data integrity of the initial physical space point cloud, and output the verified initial physical space point cloud; S604, optimize the verified initial physical space point cloud to generate an optimized physical space point cloud.

[0111] In this embodiment, when applying the scale calibration factor to the coordinates of each point in the scale-normalized point cloud, the scale calibration factor generated during the calibration process is first read and ensured to be consistent with the current point cloud reconstruction timestamp. Each point in the scale-normalized point cloud is stored in three-dimensional coordinates, and its value originates from the unified scale coordinate system output by VGGT reconstruction. By performing element-wise multiplication on the coordinate vector of each point, the scale calibration factor is applied to each component of the three-dimensional coordinates, forming a set of scale-transformed point cloud coordinates. The operation is performed in floating-point precision, and the original point attributes (such as color, normal, and confidence) are preserved in the result after multiplication. To avoid the influence of numerical overflow or rounding errors, double-precision floating-point operations are used throughout the process, and a batch vectorization acceleration mechanism is introduced to support real-time processing. The role of the scale calibration factor is to restore the physical size of the point cloud, making it reflect the true spatial scale relationship, thereby eliminating the virtual size deviation generated by the VGGT model under scale uncertainty.

[0112] When mapping the scale-transformed point cloud coordinates to the robot coordinate system, an extrinsic parameter matrix is ​​used to describe the rigid transformation relationship from the camera coordinate system to the robot coordinate system. This extrinsic parameter matrix, obtained during the calibration process, includes a rotation matrix and a translation vector. The mapping operation is implemented through matrix multiplication and vector addition, specifically performing the calculation R·P + t for each point, where R represents the rotation matrix, P represents the scale-transformed point coordinates, and t represents the translation vector. If the system uses multi-view point clouds, a unified coordinate mapping is required for point cloud data from different viewpoints, followed by fusion in the robot coordinate system. During fusion, redundant points are merged based on a distance threshold for overlapping areas, and the attributes of the fused points are calculated using nearest-neighbor interpolation to obtain a complete initial physical space point cloud. To ensure computational consistency, a single reference frame locking mechanism is used for the mapping matrix to prevent coordinate drift introduced by changes in robot posture.

[0113] To verify the boundary consistency and data integrity of the initial physical space point cloud, a feasible region is first defined based on the geometric boundary model of the robot's workspace, and points in the point cloud that exceed this region are marked as outliers. Next, statistical analysis is performed on the point cloud density, inter-point distance distribution, and confidence level, and outlier detection algorithms are used to identify isolated points or regions with abrupt density changes. If missing regions are found in the point cloud, neighborhood interpolation or multi-frame fusion is used for completion. Data integrity verification also includes checking the index integrity of the point cloud file, the continuity of the coordinate sequence, and the validity of attribute fields. The validated initial physical space point cloud is marked as a valid dataset and output for subsequent geometric feature extraction and relative position calculation.

[0114] When optimizing the validated initial physical point cloud, a multi-stage filtering and geometric reconstruction mechanism is introduced. First, statistical filtering is applied to remove noise points, and outliers are identified by calculating the mean distance to the neighborhood of each point. Second, voxel mesh downsampling is used to reduce data redundancy and improve subsequent computational efficiency. Third, smooth reconstruction is performed on discontinuous areas of the surface, restoring continuous structures through surface fitting or local curvature compensation. During the optimization phase, the point cloud normal vectors are simultaneously re-estimated, and local principal component analysis (PCA) is used to calculate the normal direction of each point, followed by normal consistency adjustment. The result generated after optimization is the optimized physical point cloud, possessing a smooth, complete, and physically scale-accurate spatial representation.

[0115] This embodiment achieves a precise conversion from virtual to physical scale by applying a scale calibration factor to the point cloud and mapping it to the robot coordinate system, enabling the reconstructed point cloud to possess realistic spatial size information. The verification and optimization process ensures the consistency of the point cloud in terms of boundary range, density distribution, and structural integrity, significantly reducing noise and redundant points, and improving data stability and reconstruction quality.

[0116] In one embodiment, step S70 above includes: S701, Map the target segmentation information onto the physical space point cloud, and extract the point cloud clusters of the target object; S702, Calculate the centroid coordinates of the point cloud cluster of the target object; S703, Obtain the current pose coordinates of the active component; S704, calculate the difference between the centroid coordinates and the current pose coordinates to generate a three-dimensional relative position.

[0117] In this embodiment, when mapping target segmentation information onto the physical space point cloud, the two types of inputs are first synchronized in time and aligned to the same coordinate system. The target segmentation information originates from pixel-level segmentation masks of multi-view images, carrying a binary or probability label for each pixel indicating whether it belongs to the target object or the background. The physical space point cloud originates from scale-calibrated 3D reconstruction, with coordinates already in the robot's working coordinate system. To establish the correspondence between pixels and 3D points, the intrinsic and extrinsic parameters of the view acquisition unit used to generate the point cloud are read, and each 3D point in the point cloud is mapped to the imaging plane of the corresponding view via a projection model. When multiple views exist, the projected pixels in all views are calculated for each 3D point, and the label or probability value of the target segmentation information is retrieved. A weighted fusion of view confidence and visibility score is performed to obtain the target attribution score for that 3D point. To suppress the effects of occlusion and edge mismatch, visibility determination and occlusion detection rules are introduced to remove points with excessively large angles to the viewing direction or those that are occluded. 3D points with scores below a threshold are subject to background determination and are excluded. The three-dimensional points retained after filtering by the scoring threshold form a point cloud cluster. This point cloud cluster provides a three-dimensional geometric support set for the target object in the robot coordinate system, and retains the color, normal, timestamp and score of each point, which is convenient for subsequent robust statistics.

[0118] When calculating the centroid coordinates of the target object point cloud cluster, a robust estimation process sensitive to both geometry and confidence is employed. Based on the spatial distribution of the point cloud cluster, local density estimation is used to eliminate isolated points and sparse tails, avoiding the influence of a few outliers on the center position. The remaining points are jointly weighted according to the target attribution score and the point cloud reconstruction confidence, and the weighted center is calculated as the centroid coordinates. If the point cloud cluster contains obvious planar or elongated structures, principal component analysis can be used to infer the principal axis direction, and the center offset can be constrained along the principal axis direction to improve stability. When multiple frames of point cloud clusters are available within a short time window, a time-weighted moving average is used to smooth the centroid coordinates. The window length and weights are adaptively adjusted with the motion speed, thereby reducing instantaneous noise while ensuring responsiveness. The output coordinate system of the centroid coordinates is kept consistent with the physical space point cloud, facilitating direct coupling with downstream pose and trajectory calculations.

[0119] When acquiring the current pose coordinates of the active component, the pose stream provided by the robot controller or external tracking system is read, and the pose is represented jointly by position and orientation. To ensure coordinate consistency in the 3D relative position calculation, the pose coordinates are transformed to a working coordinate system consistent with the point cloud, using a calibrated rigid body transformation. The timestamps of the pose and the point cloud are aligned; when there is a difference in sampling frequency, interpolation or extrapolation is used for time registration of the pose. Simultaneously, the impact of joint states and end-effector compensation parameters on the pose is evaluated. In cases with flexible end effectors or tool coordinate system offsets, tool center point compensation is applied to ensure that the pose coordinates accurately reflect the actual spatial position of the active component. For high-speed motion or vibration scenarios, low-latency filtering and prediction models are introduced to reduce pose shifts caused by communication and sensing delays.

[0120] When calculating the difference between the centroid coordinates and the current pose coordinates to generate the 3D relative position, it is first confirmed that they are in the same coordinate system and the same time reference. If there is a coordinate or time inconsistency, the aforementioned alignment and transformation are performed first. The vector difference between the current position vector and the centroid position vector is used to obtain the 3D displacement from the active component to the target object, and the attributes associated with this displacement, such as timestamp, confidence level, and the most recently seen view set, are retained. To improve the reliability of the 3D relative position under noise and occlusion conditions, an uncertainty ellipsoid can be calculated based on the spatial distribution of point cloud clusters, and a covariance estimate can be added to the 3D relative position for safety determination in subsequent motion planning. For scenarios with non-static targets, the local velocity of the target can be estimated by combining the changes in the two most recent centroid coordinates, and a short-time prediction correction can be added to the 3D relative position to reduce tracking errors caused by control loop delay. The final output 3D relative position satisfies geometric consistency, temporal consistency, and coordinate consistency, and can be directly used as a structured input element in the motion generation stage.

[0121] This embodiment establishes a bidirectional mapping from pixels to 3D points between physical space point clouds and target segmentation information. The point cloud clusters can characterize the true 3D range of the target object with high confidence, avoiding the amplification of spatial calculations caused by misjudgments of 2D masks under complex lighting and occlusion. By implementing weighted robust center estimation and time smoothing on the point cloud clusters, the insensitivity of centroid coordinates to noise and local missing values ​​is improved, reducing the disturbance of position jumps to the control loop. By performing time registration, coordinate unification, and tool center point compensation on the pose coordinates of active components, the consistency of 3D relative positions in space and time is ensured.

[0122] In one embodiment, a physical reference-based motion control device is provided, which corresponds one-to-one with the physical reference-based motion control methods described in the above embodiments. Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the motion control device based on physical reference of the present invention. The modules include a multimodal data acquisition module 10, a semantic segmentation module 20, a 3D reconstruction module 30, a physical baseline determination module 40, a scale calibration calculation module 50, a point cloud scale conversion module 60, a relative position extraction module 70, and a motion generation module 80. Detailed descriptions of each functional module are as follows: The multimodal data acquisition module 10 is used to acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of active components; The semantic segmentation module 20 is used to process the multi-view image according to the instruction information and generate target segmentation information of the target object; The 3D reconstruction module 30 is used to process the multi-view images and generate a scale-normalized point cloud of the scene and a model estimation baseline. The physical baseline determination module 40 is used to determine the physical reference baseline between the at least two view acquisition units based on the calibration information of the at least two view acquisition units and the pose information of the active component. The scale calibration calculation module 50 is used to estimate the baseline and the physical reference baseline based on the model and generate a scale calibration factor. The point cloud scale conversion module 60 is used to apply the scale calibration factor to convert the scale-normalized point cloud into a physical space point cloud. The relative position extraction module 70 is used to extract the three-dimensional relative position of the target object relative to the active component by combining the target segmentation information and the physical space point cloud; The motion generation module 80 is used to generate motion commands based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component.

[0123] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a physical reference-based motion control method on the server side.

[0124] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a physical reference-based motion control method, defining client-side functions or steps.

[0125] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of the active component; Based on the instruction information, process the multi-view image to generate target segmentation information of the target object; The multi-view images are processed to generate a scale-normalized point cloud of the scene and a model estimation baseline; Based on the calibration information of the at least two view acquisition units and the pose information of the active component, a physical reference baseline is determined between the at least two view acquisition units; Based on the model-estimated baseline and the physical reference baseline, a scale calibration factor is generated; The scale-normalized point cloud is converted into a physical space point cloud by applying the scale calibration factor. By combining the target segmentation information and the physical space point cloud, the three-dimensional relative position of the target object with respect to the active component is extracted; Based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component, an action instruction is generated.

[0126] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of the active component; Based on the instruction information, process the multi-view image to generate target segmentation information of the target object; The multi-view images are processed to generate a scale-normalized point cloud of the scene and a model estimation baseline; Based on the calibration information of the at least two view acquisition units and the pose information of the active component, a physical reference baseline is determined between the at least two view acquisition units; Based on the model-estimated baseline and the physical reference baseline, a scale calibration factor is generated; The scale-normalized point cloud is converted into a physical space point cloud by applying the scale calibration factor. By combining the target segmentation information and the physical space point cloud, the three-dimensional relative position of the target object with respect to the active component is extracted; Based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component, an action instruction is generated.

Claims

1. A motion control method based on physical references, characterized in that, Includes the following steps: Acquire instruction information, multi-view images acquired by at least two view acquisition units, and active component pose information of the active component; Based on the instruction information, process the multi-view image to generate target segmentation information of the target object; The multi-view images are processed to generate a scale-normalized point cloud of the scene and a model estimation baseline; Based on the calibration information of the at least two view acquisition units and the pose information of the active component, a physical reference baseline is determined between the at least two view acquisition units; Based on the model-estimated baseline and the physical reference baseline, a scale calibration factor is generated; The scale-normalized point cloud is converted into a physical space point cloud by applying the scale calibration factor. By combining the target segmentation information and the physical space point cloud, the three-dimensional relative position of the target object with respect to the active component is extracted; Based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component, an action instruction is generated.

2. The motion control method based on physical reference as described in claim 1, characterized in that, Based on the instruction information, the multi-view image is processed to generate target segmentation information of the target object, including: Parse the instruction information to extract keywords describing the target object; Based on the keywords describing the target object, the multi-view images are preprocessed to enhance image quality; The preprocessed multi-view images and the target object description keywords are input into the pre-trained target detection and segmentation model; The target detection and segmentation model generates a pixel-level segmentation mask for the target object. Morphological optimization is performed on the pixel-level segmentation mask to remove noise; Based on the optimized pixel-level segmentation mask, target segmentation information of the target object is generated.

3. The motion control method based on physical reference as described in claim 1, characterized in that, Processing the multi-view images to generate a scale-normalized point cloud of the scene and a model estimation baseline includes: Input the multi-view images into the pre-trained VGGT model; The VGGT model is used to extract depth features from multi-view images; Based on the aforementioned depth features, the three-dimensional point cloud data of the scene is reconstructed; The three-dimensional point cloud data is subjected to scale normalization processing to generate a scale-normalized point cloud; The relative distance between the view acquisition units corresponding to the multi-view images is estimated as the model estimation baseline.

4. The motion control method based on physical reference as described in claim 1, characterized in that, Based on the calibration information of the at least two view acquisition units and the pose information of the active component, a physical reference baseline is determined between the at least two view acquisition units, including: Load the pre-calibrated calibration information of the at least two view acquisition units; Real-time acquisition of the active component's pose information; Based on the calibration information and the pose information of the active component, calculate the Euclidean distance between the at least two view acquisition units; Verify the reasonableness of the calculated Euclidean distance; The Euclidean distance, after passing the rationality verification, is used as the physical reference baseline between the at least two view acquisition units.

5. The motion control method based on physical reference as described in claim 1, characterized in that, Based on the model-estimated baseline and the physical reference baseline, a scale calibration factor is generated, including: Read the values ​​of the model-estimated baseline and the physical reference baseline to obtain the baseline values; Check the data validity of the baseline values ​​and output valid baseline values; Based on the effective baseline values, the ratio of the physical reference baseline to the model-estimated baseline is calculated to generate the initial value of the scale calibration factor; Verify the rationality of the numerical range of the initial value of the scale calibration factor, and output the verified initial value of the scale calibration factor; The initial value of the verified scale calibration factor is subjected to a smoothing filter to generate a smoothed scale calibration factor. Based on historical scale calibration factor data, the smoothed scale calibration factor is dynamically adjusted to generate the adjusted scale calibration factor. The adjusted scale calibration factor is stored and output as the final scale calibration factor.

6. The motion control method based on physical reference as described in claim 1, characterized in that, Applying the scale calibration factor, the scale-normalized point cloud is converted into a physical space point cloud, including: The scale calibration factor is applied to the coordinates of each point in the scale-normalized point cloud, and a multiplication operation is performed to generate the scale-transformed point cloud coordinates. The scale-transformed point cloud coordinates are mapped to the robot coordinate system to generate the initial physical space point cloud; Verify the boundary consistency and data integrity of the initial physical space point cloud, and output the verified initial physical space point cloud; The initial physical space point cloud that has passed the verification is optimized to generate an optimized physical space point cloud.

7. The motion control method based on physical reference as described in claim 1, characterized in that, Combining the target segmentation information and the physical space point cloud, the three-dimensional relative position of the target object with respect to the active component is extracted, including: The target segmentation information is mapped onto the physical space point cloud to extract the point cloud clusters of the target object; Calculate the centroid coordinates of the point cloud cluster of the target object; Obtain the current pose coordinates of the active component; Calculate the difference between the centroid coordinates and the current pose coordinates to generate a three-dimensional relative position.

8. A motion control device based on physical reference, characterized in that, The physical reference-based motion control device includes: The multimodal data acquisition module is used to acquire command information, multi-view images acquired by at least two view acquisition units, and active component pose information of active components; The semantic segmentation module is used to process the multi-view image according to the instruction information and generate target segmentation information of the target object; The 3D reconstruction module is used to process the multi-view images and generate a scale-normalized point cloud of the scene and a model estimation baseline. The physical baseline determination module is used to determine the physical reference baseline between the at least two view acquisition units based on the calibration information of the at least two view acquisition units and the pose information of the active component. The scale calibration calculation module is used to estimate the baseline and the physical reference baseline based on the model and generate a scale calibration factor. The point cloud scale conversion module is used to apply the scale calibration factor to convert the scale-normalized point cloud into a physical space point cloud. The relative position extraction module is used to extract the three-dimensional relative position of the target object relative to the active component by combining the target segmentation information and the physical space point cloud; The motion generation module is used to generate motion commands based on the three-dimensional relative position, the multi-view image, the instruction information, and the pose information of the active component.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a physical reference-based motion control program stored in the memory and executable on the processor. When executed by the processor, the physical reference-based motion control program implements the steps of the physical reference-based motion control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a motion control program based on physical references, which, when executed by a processor, implements the steps of the motion control method based on physical references as described in any one of claims 1-7.

Citation Information

Cited By

  • Case material visual guidance grabbing method based on improved calibration model, medium and system

    CN121696983A

  • Off-line data-oriented three-dimensional space mode completion method for database with body

    CN121810955A

  • Embodied database three-dimensional space modal completion method for offline data

    CN121810955B

  • Industrial equipment action parameter generation method and device, equipment and storage medium

    CN122132819A

  • Industrial equipment action parameter generation method and device, equipment and storage medium

    CN122132819B