Mechanical arm control method and system based on image point cloud cross-modal fusion and action block Transformer
Through the method of cross-modal fusion of image point cloud and motion block Transformer, the problems of precise operation and generalization ability of the robotic arm control system in complex environments are solved, efficient cross-modal information complementation and motion prediction are achieved, and the task success rate of the robotic arm in new scenarios is improved.
Patent Information
- Application Number
- CN202510851897.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-12
AI Technical Summary
Existing robotic arm manipulation systems perform poorly in precise object manipulation, are difficult to generalize to unseen environmental configurations, have low success rates in performing tasks in cross-modal fusion and complex environments, and lack robustness and generalization capabilities.
A method based on cross-modal fusion of image point cloud and action block Transformer is adopted. By acquiring multi-view RGB images and depth maps to construct a scene point cloud pyramid, a virtual view of multi-view fused image-point cloud information is generated. The Transformer encoder-decoder architecture is used for feature extraction and action prediction, and continuous action sequences are processed in blocks to improve the consistency and accuracy of execution.
It improves the robot arm's global understanding and decision-making capabilities in complex tasks, reduces its dependence on real multi-camera systems, and improves its adaptability and mission success rate in new scenes and new object combinations.
Smart Images

Figure CN120620192A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of embodied intelligence, and specifically relates to a robotic arm control method and system based on cross-modal fusion of image point clouds and motion block Transformer. Background Art
[0002] Robotic manipulation aims to teach robots to complete three-dimensional manipulation tasks, such as automatically grasping a cup, opening a drawer, and classifying shapes, which is crucial for the development of automated robotics and industrial manufacturing. However, most existing robotic manipulation systems perform poorly in precise object manipulation and are difficult to generalize to unseen environmental configurations. Their motion prediction mainly relies on a single modality (RGB image or 3D point cloud), and cross-modal fusion has not been fully explored in robotic manipulation. In addition, although a large number of imitation learning algorithms can be trained with only a small number of demonstration samples, they have limited adaptability when dealing with complex environments and lack robustness when learning complex skills.
[0003] Although the robot control method based on multimodal fusion has made significant progress with the advancement of deep learning technology and the popularization of low-cost depth cameras, there are still some challenging scenarios that are difficult to solve, mainly including the following aspects: (1) Most of the early robot control methods are based on ordinary RGB images (patent CN108544531B). The lack of depth data will make these methods lack spatial perception ability, thereby reducing the success rate of the method in performing tasks; (2) Among the robot control methods, some methods (patent CN119551336A, patent CN119347785A) use depth cameras to obtain 3D information in the scene, but when they obtain information, they are easily limited by the actual position of the camera, which is not conducive to comprehensive information collection. (3) The current methods that use multimodal information will significantly reduce the success rate of performing tasks when the scene environment changes, and lack generalization. In summary, there is an urgent need for a new robot control method that can overcome the above problems. This method can achieve cross-modal information complementarity, quickly adapt to new scenes and new object combinations with only a few demonstration samples, and have a high success rate in performing tasks. Summary of the Invention
[0004] To solve the problems existing in the above-mentioned prior art, the present invention proposes a robotic arm control method and system based on cross-modal fusion of image point clouds and motion block Transformer, the method comprising: obtaining a multi-perspective RGB image and a corresponding depth map of a scene; constructing a scene point cloud pyramid based on the multi-perspective RGB image and the corresponding depth map; using the scene point cloud pyramid to generate a virtual view of the multi-perspective fused image-point cloud information around the robotic arm workspace; inputting the virtual view of the multi-perspective fused image-point cloud information into a feature extraction network to obtain a feature map; obtaining the states of each joint and the historical motion sequence in the robotic arm, and projecting the joint states; inputting the feature map, the projected joint states and the historical motion sequence into the Transformer encoder to obtain encoded features; inputting the encoded features into the Transformer decoder to predict the robotic arm motion sequence; and controlling the robotic arm according to the robotic arm motion sequence.
[0005] A robotic arm control system based on cross-modal integration of image point clouds and action block transformer, the system includes: a virtual viewpoint re-rendering unit, a feature encoding and decoding unit, and an action block prediction unit;
[0006] The virtual viewpoint re-rendering unit is used to convert the monocular RGB-D image into a multi-view orthogonal projection point cloud, thereby reconstructing a globally consistent 3D scene representation under the virtual viewpoint;
[0007] The feature encoding and decoding unit adopts the standard Transformer encoder-decoder architecture to encode the extracted 3D features, joint states and historical action sequences to obtain cross-modal global perception information;
[0008] The action block prediction unit balances computational efficiency and action continuity by generating block actions, thereby ensuring the real-time and coherence of the predicted action sequence.
[0009] Beneficial effects of the present invention:
[0010] The robotic arm manipulation method based on image-point cloud cross-modal fusion and motion block Transformer provided by the present invention generates multi-view point clouds by re-rendering, which is equivalent to data enhancement and reduces dependence on real multi-camera systems. The present invention uses perception data of different modalities (such as vision, body state, and motion history) to integrate into a unified feature representation, thereby improving the robot's global understanding and decision-making capabilities for complex tasks. The present invention divides continuous action sequences into sub-blocks of fixed length, and predicts and executes them in blocks, effectively alleviating the problem that small action errors gradually accumulate with the number of execution steps. The present invention integrates point cloud image multimodal perception data into a unified feature representation, and combines the motion block Transformer to realize intelligent motion planning of the robotic arm, thereby improving the robot's global understanding and decision-making capabilities for complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is the overall flow chart of the present invention;
[0012] Figure 2 Schematic diagram of the robot arm control method of the present invention;
[0013] Figure 3 The specific process of extracting input information features and integrating features of the present invention;
[0014] Figure 4 Schematic diagram of the action blocks of the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0016] A robotic arm control method and system based on image point cloud cross-modal fusion and action block transformer, such as Figure 1As shown, the method includes: obtaining multi-view RGB images of a scene and corresponding depth maps; constructing a scene point cloud pyramid based on the multi-view RGB images and the corresponding depth maps; generating virtual views of multi-view fusion image-point cloud information around the robotic arm's workspace using the scene point cloud pyramid; inputting the virtual views of multi-view fusion image-point cloud information into a feature extraction network to obtain feature maps; obtaining the states of each joint and the historical action sequence in the robotic arm, and projecting the joint states; inputting the feature maps, the projected joint states, and the historical action sequence into a Transformer encoder to obtain encoded features; inputting the encoded features into a Transformer decoder to predict the robotic arm action sequence; and controlling the robotic arm according to the robotic arm action sequence.
[0017] In this embodiment, generating virtual views of multi-view fusion image-point cloud information includes: based on the constructed three-dimensional scene point cloud pyramid, establishing three orthogonal virtual viewpoints in the robotic arm base coordinate system, and for the point cloud at each resolution, generating corresponding multi-channel virtual images at the virtual viewpoints using different projection parameters to form a multi-scale input.
[0018] Furthermore, the constructed three-dimensional scene point cloud pyramid includes: respectively selecting different radii r1, r2, r3 (r1 < r2 < r3) for screen space sputtering and point cloud reconstruction, and constructing a multi-scale point cloud pyramid after coarse, medium, and fine three-layer point cloud sampling.
[0019] In this embodiment, using a feature extraction network to extract features from the virtual views of multi-view fusion image-point cloud information includes: for the input virtual views, which contain two-dimensional image information and three-dimensional coordinates, separately processing the two-dimensional and three-dimensional information; two-dimensional branch input: combining RGB×3 and depth×1 into a 4-channel virtual view, and using a slightly modified Resnet-30 to extract two-dimensional features; three-dimensional branch input: combining XYZ×3 into a 3-channel virtual view, and using a feature extraction network similar to PointNet++ to extract three-dimensional features; splitting the extracted features into N spatial tokens, adding position encoding to the tokens of the three-dimensional features, linearly mapping them to d dimensions, and finally concatenating the tokens of the two-dimensional and three-dimensional features and inputting them into the Transformer decoder.
[0020] Processing the joint states and historical action sequences involves mapping the joint states and historical action sequences into 512-dimensional features through a linear layer, where the historical action sequences also include positional encodings. The action sequence features, joint state features, and special tags [CLS] are then combined into a (k+1)×512-dimensional feature vector, which is then fed into the Transformer decoder. Four layers of self-attention and reparameterized sampling are used to generate the final style variable z, which is then fed into the Transformer decoder along with the virtual view features.
[0021] The Transformer decoder decodes the encoded features, including: the Transformer decoder generates an action sequence through a cross-attention mechanism, with the encoder output as "key" and "value" and a fixed sine embedding as "query"; the continuous action sequence is divided into action blocks of fixed length, each block contains k time steps of action; prediction and execution are performed in blocks; time series integration is added when predicting the action of the block, that is, at each time t, a length k action sequence is requested from the strategy, and for the same time t, k "predicted actions for t" are collected from the past k frames, and the exponential decay weight w is used. i =exp(-mi) Take the weighted average of k actions and get the final action to be executed.
[0022] In this embodiment, Figure 1 The following is a flow chart of the robotic arm control method based on image-point cloud cross-modal fusion and action segmentation. Figure 1 , explaining each step in detail:
[0023] S101, obtain the RGB image of the scene and its corresponding depth map, such as Figure 2 shown.
[0024] In this example, both simulation and real-world environments are used. The simulation environment used in this example is based on the CoppeliaSim platform, and robotic arm control is implemented through the PyRep interface. All simulation experiments utilize a Franka Panda robotic arm equipped with parallel grippers. The observation camera layout follows the definition of the RLBench benchmark task, deployed at the front, left shoulder, right shoulder, and wrist positions, capturing RGB-D images at a 128×128 resolution.
[0025] In a real-world environment, an Intel RealSense D435i depth camera was mounted in front of a robotic arm with a fixed third-person perspective to simultaneously capture RGB (color) and depth information. The experiment involved a fixed-mounted RM-65 robotic arm, and all operations were performed within a desktop workspace.
[0026] S102: Project the obtained RGB image and depth map, perform depth sorting, and perform screen space spraying noise reduction to construct a scene point cloud. Use the scene point cloud to generate a virtual image around the robotic arm workspace.
[0027] In this step, the center of the robot base is used as the origin, and three orthogonal viewpoints, front view, top view, and right view, are set to eliminate the perspective distortion of the real camera. This rendering process converts the RGB-D image with a shape of (4,128,128,3) into a virtual view with a size of (N,H,W,C). The specific steps are as follows:
[0028] Step 1: Projection: Use the camera's internal and external parameters to project each 3D point (X n ,Y n ,Z n ) is transformed to the camera coordinate system
[0029]
[0030] Where [R|t] is the camera external parameter. Then project it onto the two-dimensional image plane through the internal parameter matrix K and normalize it to get the pixel coordinate (x n ,y n ), depth d n =Z n ,Finally, if the image width is w, the linear index is calculated as i n =x n *w+y n .
[0031]
[0032]
[0033]
[0034] Step 2, depth sorting: sort the 32-bit depth value d n The point index n is packed as a 64-bit integer to support parallel comparisons.
[0035] Step 3, screen space splashing (noise suppression): Discrete point rendering will introduce noise (such as "sparse light points") when the point cloud resolution is insufficient. To alleviate this problem, each 3D point is represented as a disk with a radius of r facing the camera (rather than an infinitesimal point). For each pixel j and its neighboring pixel k, search for a pixel in its neighborhood that satisfies d k <d j And the following conditions:
[0036]
[0037] Where f is the focal length of the camera, and the feature of pixel j is replaced by the feature of pixel k.
[0038] The re-rendering results in a 7-channel image (RGB, world coordinate system XYZ, depth). This provides multi-view observation by covering the entire workspace and ensures that each 3D point maps to the same pixel coordinates in all views. This method perfectly solves the problems of missing pixel-level correspondences and blind spots, while also eliminating the perspective distortion of real cameras, laying a solid foundation for subsequent global reasoning and 3D data enhancement.
[0039] S103: Input the obtained virtual image into the feature extraction network to extract features. The extracted features, along with the projected joint states and historical action sequences, are then input into the Transformer encoder. The Transformer can process long sequences to achieve global information fusion and generate new sequences, effectively processing multi-source high-dimensional input and learning the distribution of action styles.
[0040] This paper adopts a Transformer-based encoder-decoder framework for visual understanding, such as Figure 3 As shown, its structure contains the following parts:
[0041] Multimodal Encoder: The model first projects the current robot joint state and historical motion sequence into vectors through linear projection. It then extracts 3D and 2D features from the re-rendered virtual view. The extracted features, along with the projected joint states and historical motion sequence, are fed into the Transformer encoder. Through a multi-head self-attention mechanism, the model captures long-range dependencies between different viewpoints, joint information, and motion timing, and aggregates this information into a hidden representation of the [CLS] token. Finally, this hidden representation is projected into a latent representation z through linear projection.
[0042] Action Block Decoder: Based on the current observation and the latent representation z, the model fuses point cloud features, joint observations, and historical action sequences, and regresses the output action sequence through a cross-attention mechanism. This Transformer-based architecture not only efficiently integrates cross-modal visual information, but also learns compact and controllable action styles in the latent space, supporting fine-grained continuous operation of the robotic arm. The results of the block encoder are shown in Figure 2. Figure 4 shown.
[0043] S104: Input the encoded result into the Transformer decoder to obtain the final predicted action sequence. As the robot arm gradually performs the action, the local errors caused by the imperfect strategy will accumulate over time, eventually causing the overall trajectory to deviate from the expert demonstration. This phenomenon is called compound error in imitation learning. To correct this problem, this embodiment adopts an action block strategy, which divides the continuous action sequence into "blocks" of fixed length. For each frame t, the strategy is asked for an action block of length k, and the results are as follows:
[0044] C t ={a t|t ,a t+1|t ,......,a t+k-1|t}
[0045] At the same time, the action predictions covering time t have been generated by t-1, t-2, ..., t-(k-1) before. For time t, the action predictions from C t ,C t-1 ,......C t-(k-1) The corresponding "step 0, step 1, ..., step (k-1)" predictions are made, so there are a total of k groups of "actions at time t" candidates. Then use the weights
[0046] w i =exp(-mi)i=0,1,2,...,k-1
[0047] Take a weighted average of these k sets of predictions (where i = 0 corresponds to the latest prediction and i = k-1 corresponds to the oldest prediction) to get the final action to be executed at time t:
[0048]
[0049] This approach does not switch suddenly at the boundary of each action block, but smoothly fuses predictions at different times to avoid jittery robot movements.
[0050] This example also provides a performance comparison of the image-point cloud cross-modal fusion and motion segmentation robotic arm control methods, as follows:
[0051] The simulation environment includes: Each task is derived from RLBench benchmark tasks. Each task contains multiple variants (for example, in the cup stacking task, the positions of the cups are randomized and the colors are modified). All task variants are randomly sampled during data collection to ensure diversity but remain fixed during evaluation to ensure consistency in comparison. The entire evaluation consists of multiple independent benchmark tasks, each of which is evaluated in two distinct phases. In phase one, the training set contains only a small number of demonstration trials (20), and the model is evaluated over 25 test rounds. In phase two, the training set contains a large number of demonstration trials (100), and performance is again evaluated. During testing, the agent continuously executes actions until the task is completed (confirmed by a predefined oracle verification module) or the maximum number of allowed steps is reached. The average success rate (ASR) of the actions is used as the primary evaluation metric. The results are shown in Table 1. As shown in Table 1, four state-of-the-art general-purpose visual imitation learning algorithms (ACT, DP, DP3, and iDP3) are selected as baselines for the experiment. The effectiveness of our proposed method is verified through comprehensive comparisons. Table 1 reports the success rates of the multi-task agents trained on all nine tasks. It can be seen that our proposed method outperforms the other four methods in all baseline tasks.
[0052] Diffusion policy variants (DP, DP3, and iDP3) exhibit lower success rates when insufficient demonstration data is available. This is because the limited demonstration trajectories typically cover only common or simple scenarios, resulting in the diffusion model observing only a subset of state-action pairs during training. Data scarcity not only weakens the model's ability to generalize to out-of-distribution states but also amplifies sampling noise due to the high sensitivity of the inverse diffusion process, ultimately degrading performance. The original ACT, which uses only RGB images lacking 3D features, performs poorly on challenging tasks (such as cup placement and pin insertion) and exhibits weak generalization.
[0053] The present invention significantly improves performance by re-rendering the actual camera point cloud into a virtual viewpoint, applying 3D data augmentation of the point cloud, and adding channels (such as point correspondences) that are not available in the original sensor image.
[0054] The real environment includes: a total of 3 tasks are set up in the real environment: ① "Pick up blocks": pick up blocks and place them at the specified target location; ② "Push Metal Cubes": push metal cubes into the predetermined target area; ③ "Stack Cups": stack cups together. This embodiment obtains training data by remotely operating the right half of the Joy-Con controller. For each real task and scene, the demonstrator uses the controller to drive the robotic arm to complete the required action and records the joint state of the robotic arm at a frequency of 10Hz. After the complete target sequence is recorded, the robotic arm is reset to the starting posture and the process is repeated. In each movement, the RGB-D video stream of the camera is synchronously collected to obtain a set of paired data sets of RGB-D frames and corresponding target joint annotations. This embodiment collects 50 test demonstrations for each task.
[0055] A robotic arm control system based on image point cloud cross-modal fusion and action block Transformer, the system includes: a virtual viewpoint re-rendering unit, a feature encoding and decoding unit, and an action block prediction unit;
[0056] The virtual viewpoint re-rendering unit is used to convert the monocular RGB-D image into a multi-view orthogonal projection point cloud, thereby reconstructing a globally consistent 3D scene representation under the virtual viewpoint;
[0057] The feature encoding and decoding unit adopts the standard Transformer encoder-decoder architecture to encode the extracted 3D features, joint states and historical action sequences to obtain cross-modal global perception information;
[0058] The action block prediction unit balances computational efficiency and action continuity by generating block actions, thereby ensuring the real-time and coherence of the predicted action sequence.
[0059] The specific implementation of the system of the present application is the same as the specific implementation of the method.
[0060] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A robotic arm control method based on image point cloud cross-modal fusion and motion block transformer, characterized by: Comprising: Obtain multi-view RGB images of the scene and corresponding depth maps; Construct a scene point cloud based on the multi-view RGB images and corresponding depth maps; Generate virtual views of multi-view fusion image-point cloud information around the workspace of the robotic arm using the scene point cloud; Input the virtual views of multi-view fusion image-point cloud information into a feature extraction network to obtain feature maps; Obtain the states of each joint and the historical action sequence in the robotic arm, and project the joint states; Input the feature maps, the projected joint states, and the historical action sequence into a Transformer encoder to obtain encoded features; Input the encoded features into a Transformer decoder, and predict the robotic arm action sequence based on an action chunking strategy; Control the robotic arm according to the robotic arm action sequence.
2. The robotic arm control method based on image point cloud cross-modal fusion and motion block transformer according to claim 1 is characterized in that: Obtain multi-view RGB images of the scene and corresponding depth maps, and then project the real point cloud onto a multi-view two-dimensional plane through virtual viewpoint re-rendering technology, and synchronously extract the pixel-level texture features of the RGB image and the three-dimensional geometric features of the point cloud.
3. The robotic arm control method based on image point cloud cross-modal fusion and motion block transformer according to claim 1 is characterized in that: Constructing a scene point cloud package maps each three-dimensional point to pixel coordinates through the internal and external parameters of the camera and calculates the linear pixel index; Select the point with the minimum depth from all the projected points at each pixel position, write its color and depth to the corresponding pixel, expand each point into a disk with a radius of r on the screen, and use neighborhood depth comparison for local replacement at the pixel level to obtain the scene point cloud.
4. The method for controlling a robotic arm based on image point cloud cross-modal fusion and motion block transformer according to claim 1, characterized in that: Generating virtual views of multi-view fusion image-point cloud information includes: Based on the constructed three-dimensional scene point cloud pyramid, establish three orthogonal virtual viewpoints in the base coordinate system of the robotic arm, and for the point cloud at each resolution, use different projection parameters to generate corresponding multi-channel virtual images at the virtual viewpoints to form a multi-scale input.
5. The method for controlling a robotic arm based on image point cloud cross-modal fusion and motion block transformer according to claim 4 is characterized in that: The constructed three-dimensional scene point cloud pyramid includes: Select different radii r1, r2, r3 (r1 < r2 < r3) respectively for screen space sputtering and point cloud reconstruction, and construct a multi-scale point cloud pyramid after coarse, medium, and fine three-layer point cloud sampling.
6. The method for controlling a robotic arm based on image point cloud cross-modal fusion and motion block transformer according to claim 1, characterized in that: Using a feature extraction network to extract features from the virtual views of multi-view fusion image-point cloud information includes: The input virtual view, which contains two-dimensional image information and three-dimensional coordinates, and processes the 2D and 3D information separately; 2D branch input: Merge RGB×3 and depth×1 into a 4-channel virtual view, and use a slightly modified Resnet-30 to extract two-dimensional features; 3D branch input: Merge XYZ×3 into a 3-channel virtual view, and use a feature extraction network similar to PointNet++ to extract 3D features; Cut the extracted features into N spatial tokens, where the tokens of the 3D features are added with position encoding, linearly mapped to d dimensions, and finally the tokens of the 2D and 3D features are concatenated and input into the Transformer decoder.
7. The method for controlling a robotic arm based on image point cloud cross-modal fusion and motion block transformer according to claim 1, characterized in that: Processing the joint states and historical action sequences involves mapping the joint states and historical action sequences into 512-dimensional features through a linear layer, where the historical action sequences also include positional encodings. The action sequence features, joint state features, and special tags [CLS] are then combined into a (k+1)×512-dimensional feature vector, which is then fed into the Transformer decoder. Four layers of self-attention and reparameterized sampling are used to generate the final style variable z, which is then fed into the Transformer decoder along with the virtual view features.
8. The method for controlling a robotic arm based on image point cloud cross-modal fusion and motion block transformer according to claim 1, characterized in that: The Transformer decoder decodes the encoded features, including: the Transformer decoder generates an action sequence through a cross-attention mechanism, with the encoder output as "key" and "value" and a fixed sine embedding as "query"; the continuous action sequence is divided into action blocks of fixed length, each block contains k time steps of action; prediction and execution are performed in units of blocks; time series integration is added when predicting the action of the block, that is, at each time t, a length k action sequence is requested from the strategy, and for the same time t, k "predicted actions for t" are collected from the past k frames, using an exponential decay weight w i =exp(-mi) Take the weighted average of k actions and get the final action to be executed.
9. A robotic arm control system based on image point cloud fusion and motion block transformer, the system being used to execute the robotic arm control method based on image point cloud cross-modal fusion and motion block transformer according to any one of claims 1 to 7, characterized in that: The system includes: a virtual viewpoint re-rendering unit, a feature encoding and decoding unit, and an action block prediction unit; The virtual viewpoint re-rendering unit is used to convert the monocular RGB-D image into a multi-view orthogonal projection point cloud, thereby reconstructing a globally consistent 3D scene representation under the virtual viewpoint; The feature encoding and decoding unit adopts the standard Transformer encoder-decoder architecture to encode the extracted 3D features, joint states and historical action sequences to obtain cross-modal global perception information; The action block prediction unit balances computational efficiency and action continuity by generating block actions, thereby ensuring the real-time and coherence of the predicted action sequence.
Citation Information
Patent Citations
An automated inspection robotic arm device, control system, and control method based on vision calibration.
CN108544531B
3D visual guidance mechanical arm grabbing method for irregular objects
CN119347785A
Material box type robot carrying system using 3D visual system and mechanical arm for automatic sorting
CN119551336A
Cited By
Robot action prediction method and system based on cross-modal feature enhancement
CN121105044A
Robot action sequence generation method and device, equipment and medium
CN121733579A
Robot cross-view-angle motion control method and system based on time-space view angle synthesis and robot
CN121893293A
Robot cross-view action control method and system based on spatiotemporal view synthesis, and robot
CN121893293B