Data processing method and system for robot simulation, computer device and medium
By acquiring video streams and forming 3D mesh data, the problem that 3D mesh data cannot be directly used for industrial-grade simulation in robot simulation scene construction has been solved. This method enables efficient simulation scene construction and seamless integration, thereby improving the efficiency of robot simulation development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING PHOENIX TECHNOLOGY CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies lack a meshing mechanism for transforming 3D point clouds into continuous topological structures in robot simulation scenario construction, making it impossible to output 3D mesh data that can be directly used in industrial-grade simulation environments and difficult to achieve seamless integration with existing robot simulation toolchains.
By acquiring video streams, extracting keyframe images, determining keyframes based on image perspective, texture information, and display area, and combining shooting pose and 3D point cloud to form 3D mesh data, using partial depth estimation and multi-frame fusion to generate global point cloud, optimizing the depth map through a loss function, and finally converting the 3D mesh data into a file format supported by the robot simulation environment.
It achieves an end-to-end closed loop from video stream to robot simulation scene, improving the speed and accuracy of simulation scene construction, supporting dynamic scene simulation, meeting the needs of industrial simulation and robot path planning, and eliminating the dependence on sparse registration.
Smart Images

Figure CN122134978A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and robotics, and in particular to a data processing method and system, computer equipment and media for robot simulation. Background Technology
[0002] In the development and training of robots, the construction of high-fidelity simulation scenarios in real-world environments has gradually become a crucial step. Simulation scenario construction refers to the use of computer vision and 3D reconstruction technology to transform physical scenes in the real world into virtual 3D models that can run on a computer. This process directly affects the training effect of robot algorithms and the efficiency and accuracy of transferring simulation data to reality. However, when constructing simulation scenarios, related technologies lack a meshing mechanism for converting 3D point clouds into continuous topological structures. This results in outputs that do not possess 3D mesh data directly usable in industrial-grade simulation environments, making seamless integration with existing robot simulation toolchains difficult. Summary of the Invention
[0003] This application provides a data processing method and system, computer equipment and medium for robot simulation, to solve or alleviate the problems described above.
[0004] In a first aspect, this application provides a data processing method for robot simulation, comprising the following steps: acquiring a video stream and extracting keyframe images from the video stream to obtain a keyframe image sequence; wherein the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area; determining the shooting pose and three-dimensional point cloud corresponding to the video stream based on the keyframe image sequence; and forming three-dimensional mesh data for adapting to the robot simulation environment based on the shooting pose and the three-dimensional point cloud.
[0005] Compared with related technologies, this data processing method for robot simulation has at least the following advantages: This method directly outputs 3D mesh data for robot simulation environments based on video streams, achieving an end-to-end closed loop from video stream to robot simulation scene construction. This not only solves or alleviates the problem of related technologies being unable to output 3D mesh data directly applicable to industrial-grade simulation environments when building simulation scenes, but also provides efficient and practical foundational support for the smooth migration of robots from simulation to reality. It better integrates seamlessly with existing robot simulation toolchains, eliminating reliance on traditional sparse registration and improving the efficiency of robot simulation development. Furthermore, this application generates 3D mesh data adapted to robot simulation environments based on shooting pose and 3D point clouds. It is compatible with static images and video sequence inputs, supports dynamic scenes, thus covering various robot simulation scenarios, improving the speed of simulation scene construction, and effectively meeting the application needs of industrial simulation, robot path planning, and other applications.
[0006] In a first possible implementation of the first aspect, the process of forming three-dimensional mesh data for adapting to a robot simulation environment based on the shooting pose and the three-dimensional point cloud includes: determining a three-dimensional Gaussian set with spatial attributes based on the three-dimensional point cloud, wherein each Gaussian point in the three-dimensional Gaussian set has feature information, the feature information including at least one of position, rotation, scale, and color; performing Gaussian reconstruction on keyframe images in the keyframe image sequence based on the shooting pose, the three-dimensional Gaussian set, and a preset image rendering mechanism, and compressing each Gaussian point into a geometric plane at the pixel level to render all... The keyframe images are defined by their unit normal vectors and the distance from the image capturing device to the geometric plane. The image capturing device is used to form the video stream. The pixel depths of all keyframe images are determined based on their unit normal vectors and the distance from the image capturing device to the geometric plane. Each keyframe image is converted into a depth map according to its pixel depth, and all depth maps are mapped to the same coordinate system and fused to form a new 3D point cloud, denoted as the global point cloud. Based on the global point cloud and the unit normal vectors of all keyframe images, 3D mesh data is formed to adapt to the robot simulation environment.
[0007] Therefore, this application generates a global point cloud by partial depth estimation and multi-frame fusion, and then performs normal estimation based on the unit normal vectors of the global point cloud and keyframe images to form a three-dimensional mesh data for adapting to the robot simulation environment. This can improve the realism of the robot simulation scene, enhance the ability to transfer from simulation to reality, and effectively meet the application needs of industrial simulation, robot path planning and other applications.
[0008] In a second possible implementation of the first aspect, the process of mapping all depth maps to the same coordinate system for fusion to form a global point cloud includes: determining a loss function to characterize the geometric consistency of different keyframe images based on the unit normal vector and pixel depth of the same 3D point in different keyframe images; wherein the 3D point is any point in the 3D point cloud corresponding to the video stream; optimizing the geometric plane parameters corresponding to each pixel in all depth maps through the loss function, and mapping the optimized depth maps to the same coordinate system for fusion to form the global point cloud.
[0009] Therefore, this application optimizes the depth map through a loss function, and then maps the optimized depth map to the same coordinate system for fusion to form a global point cloud, which can improve the accuracy and intelligence of the 3D point cloud spatial structure.
[0010] In a third possible implementation of the first aspect, the method further includes: designating the file format of the three-dimensional mesh data formed at the current moment as a first format; obtaining the file format supported by the robot simulation environment, designated as a second format; converting the three-dimensional mesh data of the first format into three-dimensional mesh data of the second format, and retaining parameter information during the conversion, to form three-dimensional mesh data adapted to the robot simulation environment; or, converting the three-dimensional mesh data of the first format into three-dimensional mesh data of the second format, and retaining parameter information during the conversion, and adding physical properties and coordinate systems adapted to the robot simulation environment, to form three-dimensional mesh data adapted to the robot simulation environment; wherein the parameter information includes at least one of vertex, normal, and texture coordinates, and the physical properties include friction coefficient and / or reflectivity.
[0011] Therefore, this application can automatically convert the final generated 3D mesh data from a standard format to a file format required or supported by the robot simulation environment. Moreover, the parameter information of the 3D mesh data can be preserved during the conversion process, so that the converted 3D mesh data can be directly imported into the industrial-grade simulation environment. This ensures that the output 3D mesh data can be correctly loaded by the simulation engine, and better supports subsequent physical simulation, rendering and intelligent agent interaction tasks.
[0012] In a fourth possible implementation of the first aspect, the process of determining the shooting pose and 3D point cloud corresponding to the video stream based on the keyframe image sequence includes: encoding the keyframe image sequence to obtain the 3D geometric features of the keyframe image sequence; extracting visual features from the keyframe image sequence to obtain the visual features of the keyframe image sequence; fusing the 3D geometric features and the visual features to obtain fused features; and determining the shooting pose and 3D point cloud corresponding to the video stream based on the fused features.
[0013] Therefore, this application obtains fused features by fusing three-dimensional geometric features and visual features. The fused features can simultaneously possess structural details and edge texture features from the coding branch, as well as semantic understanding capabilities and visual priors from the visual branch. At the same time, based on the fused features, the shooting pose and three-dimensional point cloud corresponding to the video stream can be determined, which can enable the point cloud generation process to have language guidance and semantic perception capabilities.
[0014] In a fifth possible implementation of the first aspect, the process of fusing the three-dimensional geometric features and the visual features to obtain fused features includes: enhancing the global and local information of the three-dimensional geometric features through a cross-attention mechanism to obtain enhanced three-dimensional geometric features; wherein the cross-attention mechanism is formed by alternating iterations of a first self-attention mechanism and a second self-attention mechanism a preset number of times, the first self-attention mechanism is used for a single keyframe image in the keyframe image sequence to enhance the local information of the three-dimensional geometric features, and the second self-attention mechanism is used for all keyframe images in the keyframe image sequence to enhance the global information of the three-dimensional geometric features; the enhanced three-dimensional geometric features and the visual features are then fused to obtain fused features.
[0015] Therefore, this application uses two self-attention mechanisms to alternately form a cross-attention mechanism, which can effectively fuse local and global information of three-dimensional geometric features. This enhances the collaborative expression of local and global information in keyframe image sequences when determining the shooting pose and three-dimensional point cloud corresponding to the video stream.
[0016] In a sixth possible implementation of the first aspect, the process of extracting keyframe images from the video stream includes: dividing the video stream into frames to obtain multiple frame images; comparing the image viewpoint of each frame image with a preset image viewpoint range; comparing the image texture information of each frame image with preset image texture information; and comparing the image display areas of two adjacent frame images; if the image viewpoint of a certain frame image is within the preset image viewpoint range, the similarity between the image texture information of the certain frame image and the preset image texture information is greater than or equal to a first threshold, and the overlap between the image display area of the certain frame image and the previous frame image is greater than a second threshold and less than a third threshold, then the certain frame image is taken as a keyframe image.
[0017] Therefore, it can be seen that by selecting key frame images through aspects such as image perspective, image texture information and image display area, this application can ensure that the selected frame images have a stable image perspective, no occlusion, rich texture information, and a large common viewing area between two adjacent frame images, thus fully displaying the structural features of the shooting scene.
[0018] Secondly, this application provides a data processing system for robot simulation, comprising: an image sequence module for acquiring a video stream and extracting keyframe images from the video stream to obtain a keyframe image sequence; wherein the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area; a pose point cloud module for determining the shooting pose and three-dimensional point cloud corresponding to the video stream based on the keyframe image sequence; and a data adaptation module for forming three-dimensional mesh data adapted to the robot simulation environment based on the shooting pose and the three-dimensional point cloud.
[0019] Compared with related technologies, this data processing system for robot simulation has at least the following advantages: This system directly outputs 3D mesh data for robot simulation environments based on video streams, achieving an end-to-end closed loop from video stream to robot simulation scene construction. This not only solves or alleviates the problem of related technologies being unable to output 3D mesh data directly applicable to industrial-grade simulation environments when building simulation scenes, but also provides efficient and practical foundational support for the smooth migration of robots from simulation to reality. It better integrates seamlessly with existing robot simulation toolchains, eliminating reliance on traditional sparse registration and improving the efficiency of robot simulation development. Furthermore, this application generates 3D mesh data adapted to robot simulation environments based on shooting poses and 3D point clouds. It is compatible with static image and video sequence inputs, supports dynamic scenes, thus covering various robot simulation scenarios, improving the speed of simulation scene construction, and effectively meeting the application needs of industrial simulation, robot path planning, and other applications.
[0020] Thirdly, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the data processing method for robot simulation described in any one of the above.
[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the data processing method for robot simulation described in any one of the above. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0023] In the attached diagram:
[0024] Figure 1 This is a schematic flowchart of a data processing method for robot simulation provided in one embodiment of this application; Figure 2 This is a schematic diagram of the hardware structure of a data processing system for robot simulation provided in an embodiment of this application. Figure 3 This is a schematic diagram of the architecture for building a simulation scene according to an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of a computer device suitable for implementing one or more embodiments of this application. Detailed Implementation
[0025] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0026] It is understood that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0027] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.
[0028] Figure 1 A flowchart illustrating a data processing method for robot simulation is shown. Specifically, in an exemplary embodiment, as... Figure 1 As shown, this embodiment provides a data processing method for robot simulation, including the following steps: S110, acquire the video stream and extract keyframe images from the video stream to obtain a keyframe image sequence; wherein, the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area. In some examples, the video stream is captured by an image capturing device, which includes, but is not limited to, a camera.
[0029] In some exemplary embodiments, the process of extracting keyframe images from a video stream may include: dividing the video stream into frames to obtain multiple frame images; comparing the image viewpoint of each frame image with a preset image viewpoint range; comparing the image texture information of each frame image with preset image texture information; and comparing the image display areas of two adjacent frame images; if the image viewpoint of a certain frame image is within the preset image viewpoint range, the similarity between the image texture information of a certain frame image and the preset image texture information is greater than or equal to a first threshold, and the overlap between the image display area of a certain frame image and the previous frame image is greater than a second threshold and less than a third threshold, then that frame image is used as a keyframe image. In some examples, if the image viewpoint of a certain frame image is not within the preset image viewpoint range, then the image viewpoint of that frame image is considered unstable, occluded, and unsuitable as a keyframe image. In some examples, if the similarity between the image texture information of a certain frame image and the preset image texture information is less than a first threshold, then the image texture information of that frame image is considered insufficient and unsuitable as a keyframe image. In some examples, if the overlap between the display area of a frame and the previous frame is less than a second threshold, the shared viewing area of the frame is considered small and unsuitable as a keyframe. In other examples, if the overlap between the display area of a frame and the previous frame is greater than a third threshold, the frame is considered too similar to the previous frame and is considered a redundant frame, unsuitable as a keyframe. In some examples, the specific values of the first, second, and third thresholds can be selected or set according to the actual situation, and no specific numerical limit is specified here. For example, the first, second, and third thresholds can be 85%, 50%, and 80%, respectively. Therefore, selecting keyframes based on dimensions such as image perspective, image texture information, and image display area ensures that the selected frame images have a stable image perspective, are unobstructed, have rich texture information, and have a large shared viewing area between adjacent frames, fully displaying the structural features of the shooting scene.
[0030] S120 determines the shooting pose and 3D point cloud corresponding to the video stream based on the keyframe image sequence. In some examples, the shooting pose can also be referred to as the camera pose.
[0031] In some exemplary embodiments, the process of determining the shooting pose and 3D point cloud corresponding to the video stream based on the keyframe image sequence may include: encoding the keyframe image sequence to obtain the 3D geometric features of the keyframe image sequence; extracting visual features from the keyframe image sequence to obtain the visual features of the keyframe image sequence; fusing the 3D geometric features and the visual features to obtain fused features; and determining the shooting pose and 3D point cloud corresponding to the video stream based on the fused features.
[0032] In some examples, keyframe image sequences can be encoded using a DINOv2 encoder. Encoding is performed to obtain a set of image tokens, which serve as the 3D geometric features of the keyframe image sequence; represented as... , The image tokens generated by the DINOv2 encoder do not store explicit geometric parameters (such as precise coordinates), but rather learn high-dimensional vectors whose numerical patterns implicitly incorporate visual content and its underlying geometric properties (such as relative position, scale features, and orientation) and possible semantic priors.
[0033] In some examples, the powerful capabilities of multimodal visual language models (such as Qwen3-VL) in visual semantic understanding can be fully utilized. These models can be introduced as a general source of multimodal knowledge injection. Only the visual branch (Vision Encoder) of the multimodal visual language model is used to extract mid-to-high-level semantic representations. After hierarchical Transformer encoding, this branch outputs visual feature maps or visual tokens containing rich semantic knowledge, serving as the visual features of the keyframe image sequence. This can be represented as: , This represents the set of key-value semantic tokens extracted from Qwen3-VL, used to carry the visual priors, object attributes, spatial relationships, and other knowledge learned by the multimodal visual language large model during pre-training. This represents the key and value in a key-value pair. This represents the key-value index in the key-value pair.
[0034] In some examples, the process of fusing 3D geometric features and visual features to obtain fused features can be as follows: , , , That is, the visual features obtained through the DINOv2 encoder are used as query vectors. Cross-modal fusion is performed with mid-to-high-level knowledge features from Qwen3-VL, through computation. and The similarity between the two means that the attention mechanism can be derived from the attention mechanism. Extracting corresponding visual information and performing semantic reconstruction and enhancement improves the fused features. It possesses both structural details and edge texture features from DINOv2 and semantic understanding capabilities and visual priors from Qwen3-VL.
[0035] In some examples, the process of determining the shooting pose and 3D point cloud corresponding to the video stream based on the fused features can be as follows: Introduce a Dense Prediction Transformer (DPT) to predict the pixel depth of the keyframe image, and then use the fused features obtained above... Inputting the data into the DPT network yields a high-precision depth map with accuracy exceeding the preset level, which is then combined with camera intrinsic parameters to estimate the dense point cloud. Furthermore, through viewpoint normalization encoding in the intermediate layers of the DPT, the pose (position + orientation) of the image capturing device in the current scene can be inferred, enabling simultaneous estimation of the shooting pose and ultimately obtaining the shooting pose and 3D point cloud of the entire scene.
[0036] Therefore, by fusing three-dimensional geometric features and visual features to obtain fused features, the fused features can simultaneously possess structural details and edge texture features from the coding branch, as well as semantic understanding capabilities and visual priors from the visual branch; at the same time, by determining the shooting pose and three-dimensional point cloud corresponding to the video stream based on the fused features, the point cloud generation process can possess language guidance and semantic perception capabilities.
[0037] In some exemplary embodiments, the process of fusing 3D geometric features and visual features to obtain fused features may include: enhancing the global and local information of the 3D geometric features through a cross-attention mechanism to obtain enhanced 3D geometric features; wherein the cross-attention mechanism is formed by alternating and iterating a first self-attention mechanism and a second self-attention mechanism a preset number of times, the first self-attention mechanism is used for a single keyframe image in the keyframe image sequence to enhance the local information of the 3D geometric features, facilitating the learning of spatial relationships and context in a single viewpoint; the second self-attention mechanism is used for all keyframe images in the keyframe image sequence to enhance the global information of the 3D geometric features, facilitating implicit reasoning of 3D geometry and integration of global consistency information; the enhanced 3D geometric features and visual features are then fused to obtain fused features. Therefore, by alternately forming a cross-attention mechanism using two self-attention mechanisms, the local and global information of the 3D geometric features can be effectively fused together, thereby enhancing the collaborative expressive ability of local and global information in the keyframe image sequence when determining the shooting pose and 3D point cloud corresponding to the video stream.
[0038] S130 generates 3D mesh data based on the shooting pose and 3D point cloud to adapt to the robot simulation environment. In some examples, the robot simulation environment includes, but is not limited to, the Isaac Sim environment.
[0039] In some exemplary embodiments, the process of forming 3D mesh data for adapting to a robot simulation environment based on the shooting pose and 3D point cloud may include: determining a 3D Gaussian set with spatial attributes based on the 3D point cloud, wherein each Gaussian point in the 3D Gaussian set has feature information, including at least one of position, rotation, scale, and color; performing Gaussian reconstruction on keyframe images in the keyframe image sequence based on the shooting pose, the 3D Gaussian set, and a preset image rendering mechanism, and compressing each Gaussian point into a geometric plane at the pixel level to render the unit normal vector of all keyframe images and the distance from the image capturing device to the geometric plane; wherein the image capturing device is used to form a video stream; determining the pixel depth of all keyframe images based on the unit normal vector of all keyframe images and the distance from the image capturing device to the geometric plane; converting the corresponding keyframe image into a depth map according to the pixel depth of each keyframe image, and mapping all depth maps to the same coordinate system for fusion to form a new 3D point cloud, denoted as the global point cloud; and forming 3D mesh data for adapting to a robot simulation environment based on the global point cloud and the unit normal vector of all keyframe images.
[0040] In some examples, the pixel depth of the keyframe image is denoted as... When determining the pixel depth of a keyframe image based on its unit normal vector and the distance from the image capturing device to the geometric plane, we have: In the formula, This represents the distance from the image capturing device to the geometric plane. This represents the unit normal vector of the keyframe image. This represents the direction of the pixel ray. When converting the corresponding keyframe image into a depth map according to the pixel depth of each keyframe image, a geometric regularization mechanism can be introduced into each keyframe image to improve the accuracy of local structure. The plane consistency is evaluated based on the normal vector between adjacent pixels and the depth change. For edge transition regions, image gradient constraints are introduced to reduce the regularization intensity and avoid excessive smoothing that leads to boundary blurring.
[0041] In some exemplary embodiments, the process of mapping all depth maps to the same coordinate system and fusing them to form a global point cloud includes: determining a loss function to characterize the geometric consistency of different keyframe images based on the unit normal vector and pixel depth of the same 3D point in different keyframe images; wherein the 3D point is any point in the 3D point cloud corresponding to the video stream; optimizing the geometric plane parameters corresponding to each pixel in all depth maps using the loss function, and mapping the optimized depth maps to the same coordinate system for fusion to form a global point cloud. In some examples, the loss function can be denoted as... Then we have: In the formula, and This represents a custom parameter. This indicates that a certain 3D point is at the th... Unit normal vector in each keyframe image This indicates that a certain 3D point is at the th... Unit normal vector in each keyframe image This indicates that a certain 3D point is at the th... Pixel depth in a keyframe image This indicates that a certain 3D point is at the th... The pixel depth in each keyframe image. In some examples, the loss function is used to jointly optimize the planar parameters corresponding to each pixel, improving geometric coherence across views. Therefore, this application optimizes the depth map using a loss function, and then maps the optimized depth map to the same coordinate system for fusion to form a global point cloud, which can improve the accuracy and intelligence of the 3D point cloud spatial structure.
[0042] In some exemplary embodiments, the process of forming 3D mesh data for adapting to the robot simulation environment based on the unit normal vectors of the global point cloud and all keyframe images can be as follows: The discrete unit normal vectors are processed into a continuous form using spatial interpolation. Then, based on the unit normal vectors of the global point cloud and all keyframe images, an initial 3D mesh data is generated using Poisson Surface Reconstruction or Delaunay triangulation algorithms. In the formula, This represents 3D mesh data used to adapt to robot simulation environments. Represents the global point cloud. This represents the normal field formed by the unit normal vectors of all keyframe images. Finally, the generated 3D mesh data can be lightweighted, smoothed, and redundant patches removed to obtain 3D mesh data with continuous structure, reasonable topology, and high geometric accuracy. Therefore, by generating a global point cloud through partial depth estimation and multi-frame fusion, and then estimating the normal vectors based on the global point cloud and the unit normal vectors of the keyframe images, 3D mesh data suitable for robot simulation environments can be formed. This can improve the realism of robot simulation scenes, enhance the transfer capability from simulation to reality, and effectively meet the application needs of industrial simulation, robot path planning, and other applications.
[0043] In some exemplary embodiments, after forming the 3D mesh data, the process may further include: designating the file format of the 3D mesh data formed at the current moment as a first format; obtaining a file format supported by the robot simulation environment, designated as a second format; converting the 3D mesh data in the first format to the 3D mesh data in the second format, and retaining parameter information during the conversion to form 3D mesh data adapted to the robot simulation environment; wherein the parameter information includes at least one of vertex, normal, and texture coordinates. In some examples, the first format includes, but is not limited to, the obj format, and the second format includes, but is not limited to, the usd (universal scene description) format. For example, a 3D modeling tool (such as Blender or MeshLab) loads the generated obj mesh file, and then exports the obj mesh file to usd format through a pre-provided USD plugin or script interface. During the conversion process, information such as the vertices, normals, and texture coordinates of the original mesh is retained to ensure that the generated 3D mesh data can be correctly loaded by the simulation engine and supports subsequent physical simulation, rendering, and intelligent agent interaction tasks.
[0044] In some exemplary embodiments, after forming the 3D mesh data, the process may further include: designating the file format of the 3D mesh data formed at the current moment as a first format; obtaining a file format supported by the robot simulation environment, designated as a second format; or converting the 3D mesh data in the first format to the 3D mesh data in the second format, retaining parameter information during the conversion, and adding physical properties and coordinate systems for adapting to the robot simulation environment to form 3D mesh data adapted to the robot simulation environment; wherein the parameter information includes at least one of vertex, normal, and texture coordinates, and the physical properties include friction coefficient and / or reflectivity. In some examples, the first format includes, but is not limited to, the obj format, and the second format includes, but is not limited to, the usd (universal scene description) format. For example, the generated obj mesh file is loaded using a 3D modeling tool (such as Blender or MeshLab), and then the obj mesh file is exported to the usd format through a pre-provided USD plugin or script interface. During the conversion process, the original mesh's vertex, normal, and texture coordinate information are preserved. At the same time, physical material properties (such as friction coefficient and reflectivity) and coordinate system information are optionally added to adapt to the Isaac Sim environment, ensuring that the generated 3D mesh data can be correctly loaded by the simulation engine and support subsequent physical simulation, rendering, and intelligent agent interaction tasks.
[0045] In summary, this application provides a data processing method for robot simulation, which acquires a video stream and extracts keyframe images from the video stream to obtain a keyframe image sequence. The keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area. The shooting pose and 3D point cloud corresponding to the video stream are determined based on the keyframe image sequence. Based on the shooting pose and 3D point cloud, 3D mesh data for adapting to the robot simulation environment is formed. Therefore, this method can directly output 3D mesh data for the robot simulation environment based on the video stream, realizing an end-to-end closed loop from video stream to robot simulation scene construction. This not only solves or alleviates the problem of related technologies being unable to output 3D mesh data directly applicable to industrial-grade simulation environments when building simulation scenes, but also provides efficient and practical basic support for the smooth migration of robots from simulation to reality. It better integrates seamlessly with existing robot simulation toolchains, eliminates the dependence on traditional sparse registration, and improves the efficiency of robot simulation development. Furthermore, this method generates 3D mesh data based on the shooting pose and 3D point cloud to adapt to the robot simulation environment. It is compatible with static image and video sequence input, supports dynamic scenes, and thus covers a variety of robot simulation scenarios, improves the speed of simulation scenario construction, and effectively meets the application needs of industrial simulation, robot path planning and other applications.
[0046] In an exemplary embodiment of this application, as Figure 2 As shown, a data processing system for robot simulation is provided, including: The image sequence module 210 is used to acquire a video stream and extract keyframe images from the video stream to obtain a keyframe image sequence; wherein the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area; The pose point cloud module 220 is used to determine the shooting pose and 3D point cloud corresponding to the video stream based on the key frame image sequence. The data adaptation module 230 is used to generate three-dimensional mesh data for adapting to the robot simulation environment based on the shooting pose and three-dimensional point cloud.
[0047] It is understood that the data processing system for robot simulation provided in the above embodiments and the data processing method for robot simulation provided in the above embodiments belong to the same concept. The specific way in which the data processing method for robot simulation performs operations has been described in detail in the above method embodiments and will not be repeated here. In practical applications, the data processing system for robot simulation provided in the above embodiments can allocate the above functions to different functional modules as needed. That is, the internal structure of the data processing system for robot simulation is divided into different functional modules, and then all or part of the functions of the corresponding functional modules are implemented by the data processing method for robot simulation described in the above embodiments. For example, all or part of the functions of the image sequence module 210 can be implemented by the relevant execution process of step S110, all or part of the functions of the pose point cloud module 220 can be implemented by the relevant execution process of step S120, and all or part of the functions of the data adaptation module 230 can be implemented by the relevant execution process of step S130. No specific limitations are imposed here.
[0048] In summary, this application provides a data processing system for robot simulation. This system can directly output 3D mesh data for robot simulation environments based on video streams, achieving an end-to-end closed loop from video stream to robot simulation scene construction. This not only solves or alleviates the problem of related technologies being unable to output 3D mesh data directly applicable to industrial-grade simulation environments during simulation scene construction, but also provides efficient and practical foundational support for the smooth migration of robots from simulation to reality. It better integrates seamlessly with existing robot simulation toolchains, eliminating reliance on traditional sparse registration and improving robot simulation development efficiency. Furthermore, this system generates 3D mesh data adapted to robot simulation environments based on shooting poses and 3D point clouds, compatible with static images and video sequence inputs, and supporting dynamic scenes, thus covering various robot simulation scenarios, improving simulation scene construction speed, and effectively meeting the application needs of industrial simulation, robot path planning, and other applications.
[0049] In an exemplary embodiment of this application, as Figure 3 As shown, a data processing method or system for robot simulation described in some of the above embodiments is provided for building a robot simulation scene. This includes: keyframe image extraction. In the keyframe image extraction stage, a Qwen3-vl-72B multimodal large model can be introduced, and relevant language instructions (Prompt) can be designed to guide the Qwen3-vl-72B multimodal large model to perform frame-by-frame analysis of the input video stream. Frames with sufficient structural display, no occlusion, rich texture, and a sufficient shared viewing area (overlapping area greater than 50%) with the previous frame are selected as keyframe images, resulting in a keyframe image sequence. Figure 3 The image sequence is described. The prompt can be designed as follows: "Please determine whether the current frame image is suitable for use as a keyframe image in 3D reconstruction. The criteria are as follows: 1. The current frame image has a stable viewpoint, no occlusion, and rich texture information; 2. The current frame image can completely display the scene's structural features; 3. The current frame image and the previous frame image have a large shared viewing area (visible area overlap greater than 50%); 4. The current frame image and the previous frame image should not be too similar (if the duplicate information exceeds 80%, it is considered a redundant frame and should not be retained)." By introducing the Qwen3-vl-72B multimodal large model and designing related prompts, significant improvements can be made in robustness, comprehension, and modeling effectiveness. Figure 3 In the text, the Qwen3-vl-72B multimodal large model represents a Qwen3-VL model with approximately 72 billion parameters.
[0050] Point cloud and pose reconstruction, including geometric position encoding, multimodal alignment and depth prediction modules.
[0051] Specifically, geometric position encoding: can be performed on keyframe image sequences using the DINOv2 encoder. Encoding is performed to obtain a set of image tokens, which serve as the 3D geometric features of the keyframe image sequence; represented as... , The image tokens generated by the DINO v2 image encoder do not store explicit geometric parameters (such as precise coordinates). Instead, they are learned high-dimensional vectors whose numerical patterns implicitly integrate visual content and its latent geometric attributes (such as relative position, scale features, and orientation) and possible semantic priors. Building upon this, a hierarchical spatiotemporal modeling mechanism is introduced to enhance the collaborative representation of local and global information in keyframe image sequences. First, an inter-frame self-attention mechanism is employed (learning spatial relationships and context from a single viewpoint), followed by a global self-attention mechanism (implicitly inferring 3D geometry and integrating global consistency information). These two self-attention mechanisms alternate, effectively fusing local and global information. This cross-attention mechanism is iterated L times.
[0052] Multimodal alignment: The powerful visual semantic understanding capabilities of the Qwen3-vl-72B multimodal large model can be fully utilized. The Qwen3-vl-72B multimodal large model is introduced as a general multimodal knowledge injection source. Only the visual branch (Vision Encoder) of the Qwen3-vl-72B multimodal large model is used to extract mid-to-high-level semantic representations. After hierarchical Transformer encoding, this branch outputs visual feature maps or visual tokens containing rich semantic knowledge, serving as the visual features of the keyframe image sequence; represented as: , This represents the set of key-value semantic tokens extracted from Qwen3-VL, used to carry the visual priors, object attributes, spatial relationships, and other knowledge learned by the multimodal visual language large model during pre-training. This represents the key and value in a key-value pair. This represents the key-value index in the key-value pair. To improve the accuracy of basic visual features (such as image tokens generated by the DINOv2 encoder) during the prediction stage, a cross-model cross-attention fusion module can be designed to inject high-level semantic knowledge from the Qwen3-vl-72B multimodal large model into the low-level details of DINOv2. Specifically, this can be represented as: , , , That is, the visual features obtained through the DINOv2 encoder are used as query vectors. Cross-modal fusion is performed with mid-to-high-level knowledge features from Qwen3-VL, through computation. and The similarity between the two means that the attention mechanism can be derived from the attention mechanism. Extracting corresponding visual information and performing semantic reconstruction and enhancement improves the fused features. It possesses both structural details and edge texture features from DINOv2 and semantic understanding capabilities and visual priors from Qwen3-VL.
[0053] Depth prediction module: This module performs pixel depth prediction on keyframe images by introducing DPT (Dense Prediction Transformer), and then uses the fused features obtained from the above-mentioned methods. Inputting the data into the DPT network yields a high-precision depth map with accuracy exceeding the preset level, which is then combined with camera intrinsic parameters to estimate the dense point cloud. Furthermore, through viewpoint normalization encoding in the intermediate layers of the DPT, the pose (position + orientation) of the image capturing device in the current scene can be inferred, enabling simultaneous estimation of the shooting pose and ultimately obtaining the shooting pose and 3D point cloud of the entire scene.
[0054] The PGSR network module is initialized with the 3D point cloud output from the DPT network, constructing an initial 3D Gaussian set with spatial attributes. Each Gaussian point contains information such as position, rotation, scale, and color. Subsequently, an alpha-blending rendering mechanism is used as the preset rendering mechanism to perform Gaussian cumulative reconstruction on all keyframe images in the keyframe image sequence. At the pixel level, each Gaussian point is compressed into a local plane representation, thereby rendering the normal vector and camera-to-plane distance corresponding to the keyframe image. Based on this, the unbiased depth value of each pixel is estimated according to the following formula: In the formula, This represents the distance from the image capturing device to the geometric plane. This represents the unit normal vector of the keyframe image. This represents the direction of pixel rays. To improve the accuracy of local structures, a geometric regularization mechanism can be introduced into each keyframe image to evaluate plane consistency based on the changes in normal vectors and depth between adjacent pixels. For edge transition regions, image gradient constraints are introduced to reduce the regularization intensity and avoid excessive smoothing that leads to boundary blurring. At the cross-viewpoint level of different keyframe images, a joint loss function based on multi-view geometric consistency can also be constructed, which includes: In the formula, and This represents a custom parameter. This indicates that a certain 3D point is at the th... Unit normal vector in each keyframe image This indicates that a certain 3D point is at the th... Unit normal vector in each keyframe image This indicates that a certain 3D point is at the th... Pixel depth in a keyframe image This indicates that a certain 3D point is at the th... The pixel depth in each keyframe image. The loss function is used to jointly optimize the planar parameters corresponding to each pixel, improving geometric coherence across views.
[0055] Finally, the depth maps corresponding to all keyframe images are mapped to the same coordinate system and fused to form a new 3D point cloud or a complete dense point cloud, denoted as the global point cloud. Then, spatial interpolation is used to make the discrete unit normal vectors continuous. Based on the global point cloud and the unit normal vectors of all keyframe images, Poisson Surface Reconstruction or Delaunay triangulation algorithms are used to generate initial 3D mesh data, resulting in: In the formula, This represents 3D mesh data used to adapt to robot simulation environments. Represents the global point cloud. This represents the normal field formed by the unit normal vectors of all keyframe images. Finally, the generated 3D mesh data can be lightweighted, smoothed, and redundant patches removed to obtain 3D mesh data with continuous structure, reasonable topology, and high geometric accuracy.
[0056] Format Conversion: Load the generated obj mesh file using a 3D modeling tool (such as Blender or MeshLab), and then export the obj mesh file to USD format via a pre-provided USD plugin or script interface. During the conversion process, information such as vertices, normals, and texture coordinates of the original mesh is preserved to ensure that the generated 3D mesh data can be correctly loaded by the simulation engine and supports subsequent physical simulation, rendering, and intelligent agent interaction tasks.
[0057] Simulation scene construction: Robot simulation scene is built using 3D mesh data in USD format in simulation environments such as Isaac Sim.
[0058] In an exemplary embodiment of this application, a computer device is also provided. The computer device may include a memory, a processor, and a computer program stored in the memory. The processor can execute the computer program to cause the computer device to perform actions such as... Figure 1 The steps of the data processing method for robot simulation are shown. Figure 4 A schematic diagram of the structure of a computer device 1000 is shown. (See attached diagram.) Figure 4 As shown, the computer device 1000 includes: a processor 1010, a memory 1020, a power supply 1030, a display unit 1040, and an input unit 1060.
[0059] The processor 1010 is the control center of the computer device 1000. It connects various components via interfaces and lines, and performs various functions of the computer device 1000 by running or executing computer programs / instructions stored in the memory 1020, thereby providing overall monitoring of the computer device 1000. In some embodiments, when the processor 1010 calls a computer program stored in the memory 1020, it can execute, for example... Figure 1 The steps of the data processing method for robot simulation are shown. Optionally, processor 1010 may include one or more processing units; preferably, processor 1010 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. In some embodiments, processor 1010 and memory 1020 may be implemented on a single chip; in other embodiments, they may be implemented on separate chips.
[0060] The memory 1020 mainly includes a program storage area and a data storage area. The program storage area can store the operating system, various applications, etc.; the data storage area can store instruction data created according to the use of the computer device 1000. In addition, the memory 1020 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0061] The computer device 1000 also includes a power supply 1030 (such as a battery) that supplies power to various components. The power supply can be logically connected to the processor 1010 through a power management system, thereby enabling the management of functions such as charging, discharging, and power consumption through the power management system.
[0062] The display unit 1040 can be used to display information input by the user or information provided to the user, and can also be used to display various menus of the computer device 1000, etc. In this embodiment, it is mainly used to display the display interface of each application in the computer device 1000, as well as text, pictures, and other objects displayed in the display interface. The display unit 1040 may include a display panel 1050. The display panel 1050 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0063] The input unit 1060 can be used to receive information such as numbers or characters input by the user. The input unit 1060 may include a touch panel 1070 and other input devices 1080. The touch panel 1070 can also be referred to as a touch screen, and the touch panel 1070 can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1070).
[0064] Specifically, the touch panel 1070 can detect user touch operations and the signals generated by these operations, convert these signals into touch point coordinates and send them to the processor 1010, and receive and execute commands transmitted by the processor 1010. Furthermore, the touch panel 1070 can employ various input methods such as resistive, capacitive, infrared, and surface acoustic waves to achieve interaction. Other input devices 1080 include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, and joystick.
[0065] Of course, the touch panel 1070 can also cover the display panel 1050. When the touch panel 1070 detects a touch operation on or near it, it can transmit the information to the processor 1010 to determine the type of touch event. Subsequently, the processor 1010 provides corresponding visual output on the display panel 1050 based on the type of touch event. Although in Figure 4 In this embodiment, the touch panel 1070 and the display panel 1050 are two separate components to realize the input and output functions of the computer device 1000. However, in some embodiments, the touch panel 1070 and the display panel 1050 can be integrated to realize the input and output functions of the computer device 1000.
[0066] The computer device 1000 may also include one or more sensors, such as pressure sensors, gravity acceleration sensors, proximity sensors, etc. Of course, depending on the specific application scenario, the computer device 1000 may also include other components such as cameras.
[0067] In an exemplary embodiment of this application, a computer-readable storage medium is also provided, which stores a computer program / instructions. When executed by a processor, the computer program / instructions enable the aforementioned computer device to perform the functions described in this application. Figure 1 The steps of the data processing method for robot simulation are shown.
[0068] It will be understood by those skilled in the art that Figure 4 This is merely an example of a computer device and does not constitute a limitation on the device. The device may include more or fewer components than illustrated, or a combination of certain components, or different components. For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.
[0069] Those skilled in the art will understand that this application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described in accordance with flowcharts and / or block diagrams of data processing methods for robot simulation, data processing systems for robot simulation, and computer program products, based on some embodiments. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be applied to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to create a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce implementations for the process... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0070] It is understood that although terms such as first, second, third, etc., may be used in this application to describe thresholds, these terms are only used to distinguish thresholds from each other. For example, without departing from the scope of embodiments of this application, a first threshold may also be referred to as a second threshold, and similarly, a second threshold may also be referred to as a first threshold.
[0071] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A data processing method for robot simulation, characterized in that, The method includes the following steps: A video stream is acquired, and keyframe images are extracted from the video stream to obtain a keyframe image sequence; wherein the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area; Based on the keyframe image sequence, the shooting pose and 3D point cloud corresponding to the video stream are determined; Based on the shooting pose and the 3D point cloud, 3D mesh data is formed to adapt to the robot simulation environment.
2. The data processing method for robot simulation according to claim 1, characterized in that, The process of generating 3D mesh data for adapting to the robot simulation environment based on the shooting pose and the 3D point cloud includes: A three-dimensional Gaussian set with spatial attributes is determined based on the three-dimensional point cloud. Each Gaussian point in the three-dimensional Gaussian set has feature information, which includes at least one of position, rotation, scale, and color. Based on the shooting pose, the three-dimensional Gaussian set, and the preset image rendering mechanism, Gaussian reconstruction is performed on the keyframe images in the keyframe image sequence, and each Gaussian point is compressed into a geometric plane at the pixel level to render the unit normal vector of all keyframe images and the distance from the image shooting device to the geometric plane; wherein, the image shooting device is used to form the video stream; The pixel depth of all keyframe images is determined based on the unit normal vector of all keyframe images and the distance from the image capturing device to the geometric plane. Based on the pixel depth of each keyframe image, the corresponding keyframe image is converted into a depth map, and all depth maps are mapped to the same coordinate system for fusion to form a new 3D point cloud, denoted as the global point cloud; Based on the global point cloud and the unit normal vectors of all keyframe images, a three-dimensional mesh data is formed to adapt to the robot simulation environment.
3. The data processing method for robot simulation according to claim 2, characterized in that, The process of mapping all depth maps to the same coordinate system and fusing them to form a global point cloud includes: Based on the unit normal vector and pixel depth of the same 3D point in different keyframe images, a loss function is determined to characterize the geometric consistency of different keyframe images; wherein, the 3D point is any point in the 3D point cloud corresponding to the video stream; The loss function is used to optimize the geometric plane parameters corresponding to each pixel in all depth maps, and the optimized depth maps are mapped to the same coordinate system for fusion to form the global point cloud.
4. The data processing method for robot simulation according to any one of claims 1 to 3, characterized in that, The method further includes: The file format of the 3D mesh data generated at the current moment is designated as the first format; Obtain the file format supported by the robot simulation environment, and denote it as the second format; The three-dimensional mesh data in a first format is converted into three-dimensional mesh data in a second format, while retaining parameter information during the conversion, to form three-dimensional mesh data adapted to the robot simulation environment; or, the three-dimensional mesh data in the first format is converted into three-dimensional mesh data in a second format, while retaining parameter information during the conversion, and physical properties and coordinate systems adapted to the robot simulation environment are added to form three-dimensional mesh data adapted to the robot simulation environment; wherein, the parameter information includes at least one of vertex, normal, and texture coordinates, and the physical properties include friction coefficient and / or reflectivity.
5. The data processing method for robot simulation according to any one of claims 1 to 3, characterized in that, The process of determining the shooting pose and 3D point cloud corresponding to the video stream based on the keyframe image sequence includes: The keyframe image sequence is encoded to obtain the three-dimensional geometric features of the keyframe image sequence; and visual features are extracted from the keyframe image sequence to obtain the visual features of the keyframe image sequence. The three-dimensional geometric features and the visual features are fused to obtain fused features; Based on the fusion features, the shooting pose and 3D point cloud corresponding to the video stream are determined.
6. The data processing method for robot simulation according to claim 5, characterized in that, The process of fusing the three-dimensional geometric features and the visual features to obtain the fused features includes: The global and local information of the three-dimensional geometric features are enhanced by a cross-attention mechanism to obtain enhanced three-dimensional geometric features. The cross-attention mechanism is formed by alternating and iterating a first self-attention mechanism and a second self-attention mechanism a preset number of times. The first self-attention mechanism is used for a single keyframe image in the keyframe image sequence to enhance the local information of the three-dimensional geometric features, and the second self-attention mechanism is used for all keyframe images in the keyframe image sequence to enhance the global information of the three-dimensional geometric features. The enhanced 3D geometric features and the visual features are fused to obtain the fused features.
7. The data processing method for robot simulation according to claim 1, characterized in that, The process of extracting keyframe images from the video stream includes: The video stream is divided into frames to obtain multiple frame images; The image viewpoint of each frame image is compared with the preset image viewpoint range, the image texture information of each frame image is compared with the preset image texture information, and the image display areas of two adjacent frame images are compared. If the image view of a certain frame is within the preset image view range, the similarity between the image texture information of the certain frame and the preset image texture information is greater than or equal to a first threshold, and the overlap between the image display area of the certain frame and the previous frame is greater than a second threshold and less than a third threshold, then the certain frame is taken as a key frame image.
8. A data processing system for robot simulation, characterized in that, The system includes: An image sequence module is used to acquire a video stream and extract keyframe images from the video stream to obtain a keyframe image sequence; wherein the keyframe images are determined based on at least one of image viewpoint, image texture information, and image display area; The pose point cloud module is used to determine the shooting pose and three-dimensional point cloud corresponding to the video stream based on the key frame image sequence. The data adaptation module is used to generate three-dimensional mesh data for adapting to the robot simulation environment based on the shooting pose and the three-dimensional point cloud.
9. A computer device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the data processing method for robot simulation as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the data processing method for robot simulation as described in any one of claims 1 to 7.