An image prediction method, a storage medium, an apparatus and a product

CN122473659BActive Publication Date: 2026-09-25SWANCOR ADVANCED MATERIALS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610932737.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25
Estimated Expiration
2046-06-26

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请实施例致力于提供一种图像预测方法、存储介质、设备及产品,以解决现有技术中在面对复杂的三维物理环境时,存在空间感知和理解能力差的问题

Benefits of technology

[0026]第五方面,本申请一实施例提供了一种计算机程序产品,该计算机程序产品包括指令,该指令在电子设备上执行时使电子设备实现第一方面所述的图像预测方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473659B_ABST
    Figure CN122473659B_ABST
Patent Text Reader

Abstract

The application provides an image prediction method, a storage medium, equipment and products, and relates to the technical field of robot control. The method comprises the following steps: acquiring a global scene representation of a target three-dimensional scene, wherein the global scene representation is a three-dimensional information representation constructed based on real-time sensing scene information; predicting a state change parameter of at least one entity in the target three-dimensional scene at a future time based on the global scene representation; transforming a scene representation corresponding to a current time in the global scene representation according to the state change parameter to obtain a predicted scene representation at the future time; and generating a predicted image at the future time based on the predicted scene representation. The embodiment of the application can solve the problem of poor spatial perception and understanding in the prior art when facing a complex three-dimensional physical environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, specifically to an image prediction method, storage medium, device, and product. Background Technology

[0002] With the continuous development of robotics technology, robots and other intelligent agents are widely used in various application scenarios to provide services to humans. In these application scenarios, intelligent agents typically need to have a certain level of environmental perception and semantic understanding capabilities.

[0003] Vision-Language-Action (VLA) models are commonly used in related technologies for controlling intelligent devices. However, existing VLA models suffer from poor spatial perception and understanding capabilities when faced with complex three-dimensional physical environments. Summary of the Invention

[0004] In view of this, the embodiments of this application aim to provide an image prediction method, storage medium, device and product to solve the problem of poor spatial perception and understanding ability in the face of complex three-dimensional physical environment in the prior art.

[0005] In a first aspect, one embodiment of this application provides an image prediction method, comprising: acquiring a global scene representation of a target three-dimensional scene, wherein the global scene representation is a three-dimensional information representation constructed based on real-time perceived scene information; predicting state change parameters of at least one entity in the target three-dimensional scene at a future time based on the global scene representation; transforming the scene representation in the global scene representation corresponding to the current time according to the state change parameters to obtain a predicted scene representation for the future time; and generating a predicted image for the future time based on the predicted scene representation.

[0006] In conjunction with the first aspect, in certain implementations of the first aspect, at least one entity includes a dynamic entity, and the state change parameters include pose change parameters and / or appearance change parameters; based on the state change parameters, the scene representation corresponding to the current moment in the global scene representation is transformed to obtain a predicted scene representation for the future moment, including: obtaining the scene representation corresponding to the current moment from the global scene representation to obtain the current scene representation; obtaining the entity representation corresponding to the dynamic entity from the current scene representation to obtain the dynamic entity representation; performing pose transformation processing on the dynamic entity representation based on the pose change parameters, and / or performing appearance transformation processing on the dynamic entity representation based on the appearance change parameters to obtain a predicted dynamic entity representation of the dynamic entity for the future moment; and determining the predicted scene representation for the future moment based on the predicted dynamic entity representation.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, at least one entity further includes a static entity, and the state change parameters further include illumination change parameters; after obtaining the scene representation corresponding to the current moment from the global scene representation to obtain the current scene representation, the method further includes: obtaining the entity representation corresponding to the static entity from the current scene representation to obtain the static entity representation; performing illumination transformation processing on the static entity representation according to the illumination change parameters to obtain the predicted static entity representation for the future moment; and determining the predicted scene representation for the future moment based on the predicted dynamic entity representation, including: determining the predicted scene representation for the future moment based on the predicted dynamic entity representation and the predicted static entity representation.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, the global scene representation includes multiple Gaussian parameters; based on the global scene representation, predicting the state change parameters of at least one entity in the target 3D scene at future moments includes: compressing and encoding the multiple Gaussian parameters in the global scene representation into discrete feature sequences; and using the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at future moments based on the feature sequences.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, multiple Gaussian parameters in the global scene representation are compressed and encoded into discrete feature sequences, including: performing entity-level semantic segmentation processing on the multiple Gaussian parameters in the global scene representation to obtain a set of Gaussian parameters corresponding to at least one entity; and encoding and quantizing the set of Gaussian parameters corresponding to at least one entity to obtain a discrete feature sequence.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, at least one entity includes a dynamic entity and a static entity; encoding and quantizing the Gaussian parameter sets corresponding to each of the at least one entity to obtain a discrete feature sequence includes: extracting features from the Gaussian parameter sets corresponding to the dynamic entity to obtain instance features, and extracting features from the Gaussian parameter sets corresponding to the static entity to obtain background features; fusing the instance features and background features to obtain continuous scene features; and quantizing the continuous scene features to obtain a discrete feature sequence.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, feature extraction is performed on the Gaussian parameter set corresponding to the dynamic entity to obtain instance features, including: mapping the Gaussian parameter set corresponding to the dynamic entity to point-level features; and aggregating the point-level features into entity-level features to obtain instance features.

[0012] In conjunction with the first aspect, in some implementations of the first aspect, the state change parameters of at least one entity in the target 3D scene at a future time are predicted based on the feature sequence using the world head in the visual language action model, including: using the visual language sub-model in the visual language action model to determine the visual language features at the current time based on the feature sequence; and using the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at a future time based on the feature sequence and the visual language features.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the visual language sub-model in the visual language action model is used to determine the visual language features at the current moment based on the feature sequence, including: using the visual language sub-model in the visual language action model, based on the feature sequence, to obtain rendered images under multiple viewpoint conditions and / or lighting conditions from the global scene representation; stitching the rendered images with the actual observed images at the current moment to obtain a stitched image; encoding the stitched image and then integrating it into the inference process of the visual language sub-model through a cross-attention mechanism to output the visual language features at the current moment.

[0014] In conjunction with the first aspect, in some implementations of the first aspect, multiple perspective conditions include at least two of the following: the current actual observation perspective condition, the bird's-eye view perspective condition, and the entity surround perspective condition.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, after generating a predicted image for a future time based on the predicted scene representation, the method further includes: obtaining the three-dimensional geometric features of the target entity from the global scene representation, wherein the target entity is the entity to be operated on in the target task; and generating an action sequence for performing the target task based on the three-dimensional geometric features and the predicted image.

[0016] In conjunction with the first aspect, in some implementations of the first aspect, obtaining the three-dimensional geometric features of the target entity from the global scene representation includes: extracting the entity representation corresponding to the target entity from the global scene representation to obtain the target entity representation; determining the surface normal vector, local occupancy rate, and centroid position of the target entity based on the target entity representation; and determining the three-dimensional geometric features of the target entity based on the surface normal vector, local occupancy rate, and centroid position.

[0017] In conjunction with the first aspect, in some implementations of the first aspect, generating an action sequence for performing a target task based on three-dimensional geometric features and a predicted image includes: generating candidate actions for performing the target task based on three-dimensional geometric features and a predicted image; determining the pose information of at least one probe point corresponding to the end effector based on the candidate actions; determining the occupancy value of at least one probe point in the global scene representation based on the pose information; correcting the action generation parameters based on the occupancy value of at least one probe point to obtain target action generation parameters; adjusting the candidate actions based on the target action generation parameters to finally obtain the target action, and the action sequence includes the target action.

[0018] In conjunction with the first aspect, in some implementations of the first aspect, after obtaining the global scene representation of the target 3D scene, the method further includes: determining the confidence level of the global scene representation; and predicting the state change parameters of at least one entity in the target 3D scene at a future time based on the global scene representation, including: if the confidence level is greater than a first confidence threshold, then predicting the state change parameters of at least one entity in the target 3D scene at a future time based on the global scene representation.

[0019] In conjunction with the first aspect, in some implementations of the first aspect, after determining the confidence level of the global scene representation, the method further includes: if the confidence level is less than or equal to a first confidence threshold and greater than a second confidence threshold, then obtaining a target rendering image under the target viewpoint condition from the global scene representation, and generating a prediction image for future moments based on the target rendering image; if the confidence level is less than or equal to the second confidence threshold, then generating a prediction image for future moments based on the actual observed image.

[0020] In conjunction with the first aspect, in some implementations of the first aspect, the state change parameters include pose change parameters, and the method further includes: obtaining the actual observed pose change parameters of at least one entity at a future time; determining the pose deviation value based on the deviation between the actual observed pose change parameters of at least one entity and the pose change parameters; if the pose deviation value is greater than a preset safety threshold, then performing target processing, wherein the target processing includes re-executing the image prediction process.

[0021] In conjunction with the first aspect, in some implementations of the first aspect, before obtaining the global scene representation of the target 3D scene, the method further includes: obtaining a trained visual language action model; inputting the visual language action model into multiple simulators; using the multiple simulators to evaluate the visual language action model based on a style-uniform virtual 3D scene, obtaining evaluation results output by the multiple simulators respectively; if the evaluation results output by the multiple simulators respectively meet a preset consistency condition, then deploying the visual language action model and using the visual language action model to perform an image prediction process.

[0022] In conjunction with the first aspect, in some implementations of the first aspect, before using multiple simulators to evaluate the visual language action model based on a virtual 3D scene with a unified style and obtaining the evaluation results output by the multiple simulators, the method further includes: acquiring 3D scene data corresponding to the real scene; constructing a target global scene representation sample based on the 3D scene data; and performing randomization editing on the entity representation samples in the target global scene representation sample to obtain global scene representation samples corresponding to the multiple virtual 3D scenes respectively.

[0023] Secondly, one embodiment of this application provides an image prediction device, comprising: an acquisition module, configured to acquire a global scene representation of a target three-dimensional scene, wherein the global scene representation is a three-dimensional information representation constructed based on real-time perceived scene information; a prediction module, configured to predict state change parameters of at least one entity in the target three-dimensional scene at a future time based on the global scene representation; a processing module, configured to transform the scene representation in the global scene representation corresponding to the current time according to the state change parameters to obtain a predicted scene representation for the future time; and a generation module, configured to generate a predicted image for the future time based on the predicted scene representation.

[0024] Thirdly, one embodiment of this application provides a computer-readable storage medium storing a computer program for performing the image prediction method described in the first aspect.

[0025] Fourthly, one embodiment of this application provides an electronic device, the electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being configured to perform the image prediction method described in the first aspect.

[0026] Fifthly, one embodiment of this application provides a computer program product including instructions that, when executed on an electronic device, cause the electronic device to implement the image prediction method described in the first aspect.

[0027] In this application, by obtaining a global scene representation of the target 3D scene constructed based on real-time perception, the state change parameters of at least one entity in the target 3D scene at future time are predicted. The object of prediction of the future state is transformed from two-dimensional pixels into the state change parameters of the entity in the 3D scene. Based on these parameters, the scene representation at the current time is transformed to obtain the predicted scene representation at the future time. Then, a predicted image is generated based on this. This makes the predicted image strictly subject to the constraints of 3D geometric transformation, thereby solving the geometric inconsistency problem caused by directly predicting images of two-dimensional pixels. It improves the geometric correctness of the predicted image in terms of spatial scale, occlusion relationship and object topology, and thus improves the spatial perception and understanding ability when facing complex 3D physical environments. Attached Figure Description

[0028] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0029] Figure 1 The diagram shown is a flowchart of an image prediction method provided in an embodiment of this application.

[0030] Figure 2 The diagram shown is a flowchart illustrating the entity state change parameter prediction process provided in an embodiment of this application.

[0031] Figure 3 The diagram shown is a flowchart illustrating the action sequence generation process provided in an embodiment of this application.

[0032] Figure 4 The diagram shown is a structural schematic of an image prediction device provided in an embodiment of this application.

[0033] Figure 5 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Detailed Implementation

[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0035] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented even without certain specific details. In some instances, methods and means well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0036] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0037] Furthermore, the terms “first,” “second,” “third,” and “fourth” are used only for distinguishing descriptions and should not be interpreted as indicating or implying relative importance.

[0038] In the field of robot control, the Visual-Language-Action (VLA) model has been widely used in recent years for task execution scenarios that generate operational actions based on visual and language instructions. By combining the Vision-Language Model (VLM) with robot action strategies, the VLA model significantly improves the semantic understanding and task execution capabilities of intelligent devices.

[0039] In existing technologies utilizing VLA models for scene perception and understanding, VLA models typically predict future RGB images directly. This purely two-dimensional pixel prediction method, lacking constraints on the explicit three-dimensional geometry of the scene during prediction, results in predictions that only involve texture and color interpolation. Especially in long-range tasks or complex scenes, the model's inability to "understand" the three-dimensional positions and spatial relationships between objects leads to predictions that often fail to guarantee geometric accuracy in spatial scale, occlusion relationships, and object topology. This often results in spatial scale distortion, incorrect occlusion relationships, and inconsistent object topology, causing VLA models to exhibit poor spatial perception and understanding capabilities when facing complex three-dimensional physical environments.

[0040] To address the aforementioned technical problems, embodiments of this application provide an image prediction method, storage medium, device, and product. The image prediction method includes: acquiring a global scene representation of a target three-dimensional scene, wherein the global scene representation is a three-dimensional information representation constructed based on real-time perceived scene information; predicting state change parameters of at least one entity in the target three-dimensional scene at a future time based on the global scene representation; transforming the scene representation in the global scene representation corresponding to the current time according to the state change parameters to obtain a predicted scene representation for the future time; and generating a predicted image for the future time based on the predicted scene representation. In this way, by acquiring the global scene representation of the target 3D scene constructed based on real-time perception, the state change parameters of at least one entity in the target 3D scene at future time are predicted. The object of prediction of the future state is transformed from two-dimensional pixels into the state change parameters of entities in the 3D scene. Based on these parameters, the scene representation at the current time is transformed to obtain the predicted scene representation at the future time. Then, a predicted image is generated based on this. This makes the predicted image strictly subject to the constraints of 3D geometric transformation, thereby solving the geometric inconsistency problem caused by directly predicting images of two-dimensional pixels. It improves the geometric correctness of the predicted image in terms of spatial scale, occlusion relationship and object topology, and thus improves the spatial perception and understanding ability when facing complex 3D physical environments.

[0041] The following is combined Figures 1 to 3 The image prediction method provided in this application is described in detail.

[0042] Figure 1The diagram shown is a flowchart illustrating an image prediction method according to an embodiment of this application. This method can be applied to electronic devices; exemplarily, the electronic device can be a mobile phone, computer, smart agent device, or other device with computing capabilities, wherein the smart agent device can be, for example, a robot. Figure 1 As shown, the image prediction method may include the following steps.

[0043] S110, Obtain the global scene representation of the target 3D scene. The global scene representation is a 3D information representation constructed based on the scene information perceived in real time.

[0044] In some examples, the target 3D scene can be any 3D physical environment in which the intelligent device is currently located, such as the indoor environment in which an indoor companion robot is located. The global scene representation can be any holistic information representation that can describe the spatial structure, geometry, appearance attributes, and semantic information of the target 3D scene in 3D form. For example, this global scene representation includes, but is not limited to: a 3D Gaussian Splatting (3DGS) display scene representation constructed based on real-time perceived scene information, a point cloud-based 3D representation, or a voxel-based 3D representation.

[0045] For example, an intelligent agent device can move in its environment by carrying an information acquisition device, and use the scene information collected in real time by the information acquisition device to incrementally construct a 3D information representation of the target 3D scene, thereby obtaining a global scene representation. The information acquisition device can be, for example, an RGB-D camera or a monocular camera combined with a depth estimation module. Scene information can include, for example, image information, depth information, point cloud information, etc.

[0046] In some specific examples, when a robot equipped with an RGB-D camera moves in an indoor environment, it uses the RGB-D camera to collect images, depth values, and other information about the current indoor environment in real time. Then, it uses incremental 3DGS (Simultaneous Localization and Mapping) technology to construct a Gaussian representation of the indoor 3D scene in real time. This yields a global scene representation of the indoor 3D scene where the robot is currently located. Gaussian representation... For example, it can be represented as shown in the following formula (1).

[0047] in, For a moment The number of Gaussian parameters. Each Gaussian parameter can contain 6 parameters, namely: center position... (World coordinate system); Rotational quaternion (Can be used to construct rotation matrices) ); Scale vector (Can be used to construct diagonal scaling matrices) ); covariance matrix ;color (Represented by spherical harmonic coefficients SH, supporting view-dependent lighting); Opacity .

[0048] S120, based on the global scene representation, predicts the state change parameters of at least one entity in the target 3D scene at future moments.

[0049] In some examples, an entity can refer to a distinguishable component of a scene that can be identified, tracked, or processed independently in some way. This includes, but is not limited to: dynamic entities (such as movable people, animals, objects, and / or robot-operated objects), static entities (such as building walls, fixed furniture, and other background objects), and scene regions divided according to preset semantic categories. State change parameters can be quantitative expressions describing the changes an entity is expected to undergo in terms of spatial pose, visual appearance, or ambient lighting from the current moment to a future moment. These include, but are not limited to: pose change parameters (such as rigid body rotation variables and translation vectors), appearance change parameters (such as residual vectors of spherical harmonic coefficients), lighting change parameters, or any combination thereof. Furthermore, a future moment can be, for example, the number of steps the intelligent agent device will execute (e.g., ...). The corresponding time after (step).

[0050] For example, a world head with 3D perception capabilities in a VLA model can be used to predict the state changes of at least one entity in a target 3D scene at future moments, such as changes in object pose, changes in object appearance, changes in scene lighting, etc., and output state change parameters that can measure the state change, such as pose change parameters of dynamic entities. (rotation vector) It can be mapped to a rotation matrix using Rodrigues' formula. Translation vector ), appearance change parameters of dynamic entities (Used to simulate changes in lighting, materials, and object states), scene-level lighting change parameters. .

[0051] S130, based on the state change parameters, transform the scene representation in the global scene representation corresponding to the current moment to obtain the predicted scene representation for the future moment.

[0052] In some examples, a predicted scene representation can refer to any three-dimensional information representation used to describe a scene at a future time, obtained by applying predicted state change parameters to the scene representation at the current time.

[0053] For example, the scene representation corresponding to the current moment can be obtained from the global scene representation, and then the state change parameters can be applied to the scene representation corresponding to the current moment. For example, the entity representation in the scene representation corresponding to the current moment can be transformed by rotating, translating, changing appearance, changing lighting, etc., according to the state change parameters, so as to obtain a predicted scene representation that can describe the scene state at future moments.

[0054] In some embodiments, where at least one entity includes a dynamic entity and the state change parameters include pose change parameters and / or appearance change parameters, step S130 may specifically include: obtaining a scene representation corresponding to the current moment from the global scene representation to obtain a current scene representation; obtaining an entity representation corresponding to the dynamic entity from the current scene representation to obtain a dynamic entity representation; performing pose transformation processing on the dynamic entity representation according to the pose change parameters, and / or performing appearance transformation processing on the dynamic entity representation according to the appearance change parameters to obtain a predicted dynamic entity representation of the dynamic entity at a future moment; and determining a predicted scene representation at a future moment based on the predicted dynamic entity representation.

[0055] In some examples, the dynamic entity representation can be a representation derived from the current scene representation that specifically describes the 3D structure and appearance information of dynamic entities in the scene. For example, it may include, but is not limited to: a 3D Gaussian cluster belonging to a specific dynamic object, an independently encoded 3D feature vector of the dynamic object, or a subset of objects segmented by semantic category from the scene representation, or any combination thereof.

[0056] For example, taking Gaussian representation as an example, it can be derived from the global scene representation. Query and extract the current scene representation. From the current scenario representation The index yields a dynamic entity representation. Rigid body pose transformation is performed on the dynamic entity representation based on the pose change parameters corresponding to each dynamic entity. For example, according to the following formula (2), the dynamic entity representation is transformed into a rigid body pose. central position Covariance Matrix Apply the SE(3) transform.

[0057] in, The transformed center position, The transformed covariance matrix is... The rotation matrix is ​​one of the pose change parameters. This is the translation vector in the pose change parameters. Additionally, the appearance transformation of each dynamic entity's representation can be performed based on its corresponding appearance change parameters. For example, according to the following formula (3), the residual update can be adjusted... The spherical harmonic coefficients.

[0058] in, The original color. The adjusted color, These are parameters related to appearance changes.

[0059] After the above pose transformation and / or appearance transformation processes, a predicted dynamic entity representation of the dynamic entity at a future time can be obtained. This can then be used to predict dynamic entity representations. Representation of the current scene The dynamic entity representation in the model is updated to obtain the predicted scene representation for future time moments.

[0060] In this way, by focusing the transformation process on the dynamic entity representation extracted independently from the global scene representation and applying pose and / or appearance transformation processing to it, the prediction of the future state of the dynamic entity is decoupled into the pose and appearance transformation process of an independent geometric structure. This ensures that the dynamic entity maintains geometric consistency when it moves, rotates or changes its appearance, thus solving the technical problems of overall object distortion and texture disorder in the existing image prediction process.

[0061] Furthermore, in some embodiments, where at least one entity further includes a static entity and the state change parameters also include illumination change parameters, after obtaining the scene representation corresponding to the current moment from the global scene representation to obtain the current scene representation, the method may further include: obtaining the entity representation corresponding to the static entity from the current scene representation to obtain the static entity representation; and performing illumination transformation processing on the static entity representation according to the illumination change parameters to obtain the predicted static entity representation for the future moment. Additionally, the above-mentioned determination of the predicted scene representation for the future moment based on the predicted dynamic entity representation may specifically include: determining the predicted scene representation for the future moment based on the predicted dynamic entity representation and the predicted static entity representation.

[0062] In some examples, a static entity representation can be a representation derived from the current scene representation that specifically describes the 3D structure and appearance information of relatively fixed static entities in the scene. For example, it may include, but is not limited to: a 3D Gaussian set belonging to a static background, dense point cloud blocks encoding fixed structures of the scene, or a subset of the background segmented by semantic categories from the scene representation, or any combination thereof.

[0063] For example, taking Gaussian representation as an example, from the current scene representation Static entity representation obtained from indexing The static entity representation is then subjected to illumination transformation based on the illumination change parameters. For example, the static entity representation is adjusted according to the following formula (4). The spherical harmonic coefficients.

[0064] in, For parameters related to illumination changes, The original color. The adjusted color.

[0065] After the above illumination transformation process, the predicted static entity representation of the static entity at future time can be obtained. Combining predictive dynamic entity representation And predicting static entity representation This allows us to combine the data to obtain a predicted scenario representation for future moments. .

[0066] In this way, by applying the illumination change parameters specifically to the independently extracted static entity representation and combining it with the dynamic entity representation after pose and / or appearance transformation, the prediction of the scene at future moments can accurately reflect the individual state changes of the dynamic foreground and uniformly simulate the impact of scene-level illumination changes on the static background, thereby achieving global consistency in lighting and atmosphere between the dynamic foreground and the static background in the final generated prediction image.

[0067] S140, based on the predicted scene representation, generates a predicted image for future moments.

[0068] For example, the generation process of the predicted image can be a process of projecting a 3D predicted scene representation onto a 2D imaging plane and calculating the color value and / or depth value of each pixel. In an embodiment employing Gaussian representation, this process is implemented through differentiable rasterization: based on preset camera intrinsics, volume rendering is performed on the Gaussian parameters in the predicted scene representation, outputting an RGB predicted image and / or the corresponding depth predicted image. This approach ensures that the generated predicted image strictly adheres to 3D geometric transformation constraints, avoiding common problems encountered when directly predicting from 2D images, such as spatial distortion, incorrect object occlusion relationships, and geometric inconsistencies.

[0069] In some specific examples, the rasterization process can be implemented using the following formula (5) to obtain future moments (such as future moments). The predicted image after each action step.

[0070] in, For camera internal parameters, For the future Predicted image after each action step.

[0071] In this embodiment, by acquiring a global scene representation of the target 3D scene constructed based on real-time perception, the state change parameters of at least one entity in the target 3D scene at future moments are predicted. The object of prediction of the future state is transformed from two-dimensional pixels into state change parameters of entities in the 3D scene. Based on these parameters, the scene representation at the current moment is transformed to obtain the predicted scene representation at the future moment. Then, a predicted image is generated based on this. This makes the predicted image strictly subject to the constraints of 3D geometric transformation, thereby solving the geometric inconsistency problem caused by directly predicting images of two-dimensional pixels. It improves the geometric correctness of the predicted image in terms of spatial scale, occlusion relationship and object topology, and thus improves the spatial perception and understanding ability when facing complex 3D physical environments.

[0072] Furthermore, explicit 3D representations such as 3DGS typically contain millions of Gaussian ellipsoids, with a parameter count far exceeding the context window capacity of the VLM in the VLA model. The lack of an effective mechanism in related technologies to compress large-scale 3D scene representations into compact feature representations that can be handled by VLM prevents the effective integration of explicit 3D representations like 3DGS with VLM.

[0073] To address the aforementioned problems, embodiments of this application also provide a compression encoding mechanism. The following, in conjunction with... Figure 2 Describe in detail the process of predicting entity state change parameters based on this compression coding mechanism.

[0074] Figure 2 The diagram shown is a flowchart illustrating the entity state change parameter prediction process provided in an embodiment of this application. Figure 1 Extending from the illustrated embodiment Figure 2 The illustrated embodiment will be described in detail below. Figure 2 The illustrated embodiments and Figure 1 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0075] like Figure 2 As shown, the global scene representation includes multiple Gaussian parameters. Accordingly, the above-mentioned prediction of the state change parameters of at least one entity in the target 3D scene at a future time based on the global scene representation (step S120) may specifically include the following steps.

[0076] S210 compresses and encodes multiple Gaussian parameters in the global scene representation into discrete feature sequences.

[0077] In some examples, a discrete feature sequence can refer to any sequence of information obtained by compressing and encoding continuous, high-dimensional 3D scene representation data into a finite number of quantized indices or discrete tokens. For example, the feature sequence could be a VLM-processable token sequence containing multiple discrete codebook indices. This discrete feature sequence allows for a data size much smaller than the number of parameters in the original 3D scene representation, thus fitting the model's context window constraints.

[0078] For example, the global scene representation can be obtained using the 3D scene tokenizer encoder according to the following formula (6). Multiple Gaussian parameters in the code are encoded into a compact 3D token sequence. .in, 3D Token Sequence The number of parameters in the text, ; This is the codebook size.

[0079] The aforementioned Tokenizer encoder can adopt a hierarchical vector quantization variational autoencoder (HierarchicalVQ-VAE) architecture.

[0080] In addition, in order to further improve coding efficiency and structural fidelity and solve the problem of effectively encoding a 3D scene representation containing multiple independent entities, in some embodiments, the above step S210 may specifically include: performing entity-level semantic segmentation processing on multiple Gaussian parameters in the global scene representation to obtain a Gaussian parameter set corresponding to at least one entity; and encoding and quantizing the Gaussian parameter set corresponding to at least one entity to obtain a discrete feature sequence.

[0081] In some examples, entity-level semantic segmentation processing can refer to using a semantic segmentation network to classify all Gaussian parameters in a scene according to the entities to which they belong. The semantic segmentation network can be, for example, a lightweight mobile segmentation anything model (MobileSAM) or an offline preprocessing segmentation anything model (SAM).

[0082] For example, when at least one entity includes both dynamic and static entities, a semantic segmentation network can decompose multiple Gaussian parameters in the global scene representation into Gaussian clusters of dynamic objects corresponding to the dynamic entities. ( A dynamic entity corresponds to a Gaussian cluster, and a static background Gaussian set corresponds to a static entity. The set of Gaussian parameters corresponding to a dynamic entity is called the Gaussian cluster of the dynamic object. The set of Gaussian parameters corresponding to the static entity is the set of Gaussian parameters for the static background. .

[0083] In addition, a solid-level rigid body transformation matrix can be added to the Gaussian parameter set corresponding to each dynamic entity. (in For rotation matrix, (Translation vector) and appearance vector Appearance Vector It can be used to encode lighting and material changes, etc. As shown in the following formula (7), the solid-level Gaussian parameter set can be obtained through the transformation matrix. Associated with local Gaussian parameters.

[0084] in, The center location of the local Gaussian parameters. The center location of the entity-level Gaussian parameter set. For rotation matrix, This is a translation vector. The hierarchical representation described above can reduce the dimensionality of independent transformations involving millions of Gaussian parameters to a finite number of entity-level transformations, providing a manageable parameter space for subsequent world head prediction.

[0085] In addition, for the non-metric scale problem of 3DGS, at least one of the following methods can be used for scale alignment and metric reconstruction.

[0086] Firstly, during model training, RGB-D depth maps are fused. As a supervisory signal for Gaussian position optimization. For example, the depth loss shown in Equation (8) is used. As part of the loss function.

[0087] in, Depth maps for 3DGS rendering (e.g., recording the nearest Gaussian depth for each pixel during rasterization). This is a true depth map.

[0088] Secondly, during scene initialization, the absolute scale factor is anchored using a calibration object of known size (such as an ArUco calibration board or standard-sized furniture). The center position of all Gaussian parameters is multiplied by this factor for scaling correction.

[0089] After obtaining the Gaussian parameter set corresponding to each entity, feature extraction can be performed on the Gaussian parameter set corresponding to each entity. The extracted features are then fused into scene-level continuous latent variables, which are then quantized and mapped to discrete codebook indices to finally form discrete feature sequences.

[0090] In this way, by first performing entity-level semantic segmentation on the Gaussian parameters and then encoding the Gaussian sets of each entity separately, the encoding process can perceive and retain the independent geometric structure of different objects in the scene, avoiding the problem of mixed features of different objects caused by indiscriminate encoding. This improves the ability of the compressed discrete feature sequence to represent the independent entity structure in the scene, and provides a more accurate feature basis for subsequent state prediction of specific entities.

[0091] In addition, in some embodiments, the at least one entity mentioned above includes a dynamic entity and a static entity. Accordingly, the above-mentioned encoding and quantization of the Gaussian parameter sets corresponding to each of the at least one entity to obtain a discrete feature sequence may specifically include: extracting features from the Gaussian parameter sets corresponding to the dynamic entity to obtain instance features, and extracting features from the Gaussian parameter sets corresponding to the static entity to obtain background features; fusing the instance features and background features to obtain continuous scene features; and quantizing the continuous scene features to obtain a discrete feature sequence.

[0092] For example, a set of Gaussian parameters (such as a Gaussian cluster of dynamic objects) can be used for each dynamic entity. Geometric and appearance features are extracted to obtain entity-level feature vectors, i.e., instance features. Simultaneously, the Gaussian parameter set corresponding to the static entities serving as the background (such as the static background Gaussian set) is also processed. Extract geometric and appearance features to obtain background features.

[0093] A specific method for fusing instance features corresponding to all dynamic entities and background features corresponding to static entities could be as follows: A Transformer encoder could aggregate instance features corresponding to all dynamic entities and background features corresponding to static entities to output continuous scene features. These continuous scene features could, for example, be scene-level continuous latent variables. .

[0094] Finally, the continuous features of the scene can be quantized position by position, and then quantized position by position into the discrete codebook index shown in the following formula (9). .

[0095] in, For learnable codebook vectors, For codebook size, For scene-level continuous latent variables, 3D Token Sequence The number of parameters in the code. Obtain the discrete codebook index corresponding to each entity. Then, by arranging them in order, a discrete feature sequence can be finally formed. .

[0096] In this way, by processing the Gaussian parameter sets corresponding to dynamic entities and static entities through two parallel feature extraction paths, and then fusing and quantizing the two, the motion information of dynamic entities and the structural information of static backgrounds can be explicitly decoupled and organically integrated in the feature space. This preserves the complete information of the scene in the generated discrete feature sequence, and provides a feature basis for the subsequent world head to discriminately predict changes in dynamic entities and static scenes.

[0097] In some embodiments, the above-mentioned feature extraction of the Gaussian parameter set corresponding to the dynamic entity to obtain instance features includes: mapping the Gaussian parameter set corresponding to the dynamic entity to point-level features; and aggregating the point-level features to entity-level features to obtain instance features.

[0098] For example, the set of Gaussian parameters corresponding to the dynamic entity can be... Gaussian parameters in The data is mapped to point-level features via a multilayer perceptron (MLP), and then aggregated into entity-level feature vectors through an attention mechanism. This yields the instance features.

[0099] In this way, by performing a two-stage operation of first mapping the Gaussian parameter set to a point-level feature space and then aggregating it, the unstructured and massive Gaussian spheres are transformed into a fixed-dimensional entity-level feature. This enables the efficient processing of arbitrary dynamic entities in subsequent calculations, overcoming the computational challenges posed by the disorder and massive number of the original Gaussian spheres.

[0100] S220 uses the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at a future time based on the feature sequence.

[0101] For example, the feature sequence can be The embedding process is performed according to the following formula (10) to obtain the embedding vector sequence. .

[0102] This embedding vector sequence As one of the model input parameters, it is input into the VLA model, and the world head in the VLA model is used to predict the state change parameters of at least one dynamic entity and / or static entity in the target 3D scene at future times.

[0103] In addition, in some embodiments, step S220 may specifically include: using the visual language sub-model in the visual language action model to determine the visual language features at the current moment based on the feature sequence; and using the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at a future moment based on the feature sequence and the visual language features.

[0104] In some examples, visual language features can be high-level semantic feature vectors formed in the hidden state by the visual language sub-model (VLM) after encoding and reasoning the input information, which integrate visual perception, language instructions and understanding of three-dimensional spatial structure.

[0105] Because existing Virtual Machines (VLMs) process two-dimensional image sequences and lack persistent three-dimensional spatial memory, they frequently suffer from "visual forgetting" in navigation and multi-step operation tasks. This leads to their inability to determine whether target objects are occluded, to infer detour paths, or to construct a global spatial map using historical observations. This limitation of two-dimensional perception inevitably causes VLMs to produce illusions in complex three-dimensional physical environments during reasoning, such as misjudging object depth and repeatedly exploring already visited areas.

[0106] To address the aforementioned issues, this embodiment represents the feature sequence corresponding to the global scene representation. It serves as an externally queryable memory for the VLM. Specifically, it can be used as a feature sequence. The corresponding embedding vector sequence The hidden state of the VLM is injected through the cross-attention mechanism as shown in the following formula (11).

[0107] This allows the VLM to directly access the scene's 3D topological information, rather than relying solely on the projected 2D image. Thus, the VLM can extract the visual language features of the current moment based on this externally queryable memory. .

[0108] Based on this, when predicting state change parameters, the visual language features of the scene at the current moment can be used as output by the VLM. The device state characteristics corresponding to the current device state Task instructions And the embedding vector sequence of the 3D Token sequence corresponding to the global scene representation at the current moment. Using information such as world head in VLA model as input, predict the state change parameters of at least one entity in the target 3D scene at future time according to the following formula (12).

[0109] in, The number of dynamic entities in the target 3D scene. For appearance change parameters, The rotation matrix is ​​one of the pose change parameters. The translation vector is one of the pose change parameters. These are parameters related to changes in illumination.

[0110] In this way, by using the feature sequence corresponding to the global scene representation as the external queryable memory of the visual language sub-model, the prediction of the future entity state is achieved based on the high-level semantic information extracted from the three-dimensional spatial memory. This breaks through the limitations of occlusion and forgetting in two-dimensional image sequences, enabling the visual language sub-model to have true three-dimensional spatial memory and reasoning capabilities.

[0111] In some embodiments, the above-mentioned determination of visual language features at the current moment based on feature sequences using the visual language sub-model in the visual language action model may specifically include: using the visual language sub-model in the visual language action model, based on feature sequences, obtaining rendered images under multiple viewpoint conditions and / or lighting conditions from the global scene representation; stitching the rendered images with the actual observed images at the current moment to obtain a stitched image; encoding the stitched image and then integrating it into the inference process of the visual language sub-model through a cross-attention mechanism to output the visual language features at the current moment.

[0112] For example, during the extraction of visual language features, VLM can trigger the Tokenizer decoder to decode using a special 3D query token, thereby enabling spatial memory lookup of the global scene representation. The 3D query token can be, for example, a... Specifically, VLM can query tokens via 3D, triggering the Tokenizer decoder to utilize feature sequences. Features that meet specific viewpoint and / or lighting conditions are queried in the global scene representation, decoded into the corresponding 3D scene representation, and then projected onto the 2D imaging plane to obtain the corresponding rendered image.

[0113] For example, VLM can query the query conditions in the 3D query token. Conditional rendering queries are implemented using the Tokenizer decoder according to the following formula (13).

[0114] in, The decoder adjusts the spherical harmonic coefficients (and lighting conditions) according to the query conditions. Related) and camera pose (and viewing angle conditions) The set of Gaussian parameters that can be rendered is output after (related to the above). Then, the image can be rendered using the following formula (14) to obtain the corresponding rendered image.

[0115] in, To render the image, This refers to the camera's internal parameters.

[0116] In some embodiments, the multiple view conditions include at least two of the following: the current actual observation view condition, the bird's-eye view view condition, and the entity surround view condition.

[0117] For example, before each inference step, VLM automatically queries the rendered images under the following multiple viewpoint conditions in the manner described above, specifically including: querying the rendered images under the current actual observation viewpoint conditions, such as the RGB image and depth image rendered under the current camera pose, for comparison and verification with the actual observation images obtained by the real camera; querying the rendered images under the bird's-eye view viewpoint conditions, such as the semantically annotated rendered images viewed from above the scene, for enhancing global spatial planning; and querying the rendered images under the entity surrounding viewpoint conditions, such as the rendered images of eight equally spaced surrounding viewpoints around the target object to be operated on, for verifying the relationship between occlusion and pose.

[0118] Of course, in addition to the above-mentioned perspective conditions, other perspective conditions may also be included, which are not limited here.

[0119] In addition, the rendered image can be compared with the actual observation image captured by the camera at the current moment. The images are stitched together along the sequence dimension, for example, according to the following formula (15), to obtain the stitched image. .

[0120] in, This is a rendered image based on the current actual observation perspective. For rendering images under bird's-eye view conditions, ( ) represents N rendered images under the condition of an entity's surrounding view.

[0121] For example, the above-mentioned stitched image can be used as an enhanced input and integrated into the inference process of VLM through a cross-attention mechanism, for example, by integrating it into the inference process of VLM according to the following formula (16), and finally outputting the visual language features at the current moment.

[0122] in, Query vectors generated for VLM For perspective conditions The corresponding position code, This is an intermediate result output by VLM based on the cross-attention mechanism.

[0123] In this way, by decoupling the rendering of multiple images from the 3D scene representation according to multiple viewpoints and lighting conditions as auxiliary views, and using them together with the actual observed images as input to the visual language sub-model, the model's reasoning process can explicitly perceive and cross-validate geometric information from multiple different viewpoints, thereby eliminating ambiguity and occlusion illusion when understanding 3D space from a single viewpoint, and outputting visual language features with more complete spatial perception capabilities.

[0124] In this embodiment, by compressing and encoding a large number of global Gaussian parameters into a low-dimensional discrete feature sequence, and then using it as the basis for world head prediction, the model does not need to directly process millions of original Gaussian parameters. This solves the scale gap problem where the number of explicit 3D representation parameters far exceeds the model context window limit, and enables the prediction of the future state of the scene with acceptable computational complexity.

[0125] In addition, in the process of generating action sequences using the action head in the VLA model, the action head is based on the action sequence predicted by two-dimensional visual features and lacks the guidance of explicit three-dimensional geometric features. This leads to insufficient fine contact point estimation, collision avoidance and occlusion reasoning capabilities, resulting in a significant decrease in the success rate of task execution in confined space operations and dense scenes.

[0126] To address the aforementioned problems, this application also provides a three-dimensional geometry-enhanced motion head. The following describes... Figure 3 Describe in detail the process of generating an action sequence based on this action head.

[0127] Figure 3 The diagram shown is a schematic flowchart of an action sequence generation process provided in an embodiment of this application. Figure 1 Extending from the illustrated embodiment Figure 3 The illustrated embodiment will be described in detail below. Figure 3 The illustrated embodiments and Figure 1 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0128] like Figure 3 As shown, after generating the predicted image of the future time based on the predicted scene representation (step S140), the image prediction method further includes the following steps.

[0129] S310: Obtain the three-dimensional geometric features of the target entity from the global scene representation.

[0130] In some examples, the target entity can be the entity to be operated on in the target task, and the target task can be the task currently to be executed. For example, if the target task is "pick up the water glass on the table", then the target entity can be the water glass to be picked up. Additionally, 3D geometric features can be features that represent the 3D geometric structure of the target entity, such as the surface normal vectors, local occupancy, and centroid position of the target entity.

[0131] For example, if the global scene representation has been compressed into discrete 3D token sequences Then, the Tokenizer decoder can be used to extract the 3D token sequence. The corresponding embedding vector sequence The local geometric features of the target entity are decoded, and then the three-dimensional geometric features of the target entity are obtained.

[0132] In some embodiments, step S310 may specifically include: extracting the entity representation corresponding to the target entity from the global scene representation to obtain the target entity representation; determining the surface normal vector, local occupancy rate, and centroid position of the target entity based on the target entity representation; and determining the three-dimensional geometric features of the target entity based on the surface normal vector, local occupancy rate, and centroid position.

[0133] For example, taking the global scene representation as a Gaussian representation, if the global scene representation has been compressed into a discrete 3D token sequence... Then the Tokenizer decoder can be used to represent the current scene. Extract the Gaussian cluster corresponding to the target entity. That is, the representation of the target entity is obtained.

[0134] In some specific examples, the process of determining the surface normal vectors of a target entity based on its representation may include: for the Gaussian cluster... Each Gaussian parameter in the matrix can be obtained through the covariance matrix. The eigenvector corresponding to the minimum eigenvalue is used to estimate the local surface normal vector. The minor axis direction of the flat Gaussian approximates the normal direction. The local surface normal vectors within the target region corresponding to the target entity are aggregated to obtain the surface normal vector of the target entity. Specifically, the aggregation process can be a weighted average, where the weights can be the opacity corresponding to each Gaussian parameter. .

[0135] Furthermore, based on the target entity representation, the specific process of determining the local occupancy rate of the target entity may include: for the Gaussian cluster For each Gaussian parameter in the formula, the local occupancy rate of the target entity is calculated using the following formula (17). .

[0136] in, For query point Gaussian parameters in the domain, For the first Opacity of a Gaussian parameter.

[0137] Additionally, due to the Gaussian cluster corresponding to the target entity It is a set of Gaussian parameters at the entity level; therefore, this Gaussian cluster... Corresponding center position This refers to the centroid location of the target entity.

[0138] Obtain the surface normal vector of the target entity Local occupancy rate and the position of the center of mass Then, the surface normal vector can be... Local occupancy rate and the position of the center of mass The set of features constituted Three-dimensional geometric features of the target entity .

[0139] In this way, by determining the surface normal vector, local occupancy rate, and centroid position of the target entity based on the target entity representation corresponding to the target entity to be operated, the complex, implicit 3D scene representation is transformed into a set of accurate, explicit, and easily usable 3D geometric features by subsequent action heads, providing clear key information for subsequent action planning, such as "where is the object to be operated", "which direction is it facing", and "whether the space is occupied".

[0140] S320 generates a sequence of actions to perform a target task based on three-dimensional geometric features and predicted images.

[0141] For example, three-dimensional geometric features and predicted images can be used as input information to the action head in the VLA model, and then the action head can be used to generate an action sequence to perform the target task.

[0142] In some specific examples, three-dimensional geometric features can be... As an additional condition, with the predicted image Visual language features of the scene at the current moment output by VLM The device state characteristics corresponding to the current device state Together, they are concatenated as input information according to the following formula (18) and input to the action head so that the action head can generate the action sequence for performing the target task based on the above input information by using stream matching denoising.

[0143] in, This is the input information corresponding to the action head.

[0144] Furthermore, during the flow matching and denoising process of the action head, a 3D occupancy penalty mechanism can be introduced to achieve collision detection during action planning. Based on this, in some embodiments, step S320 may specifically include: generating candidate actions for performing the target task based on 3D geometric features and a predicted image; determining the pose information of at least one probe point corresponding to the end effector based on the candidate actions; determining the occupancy value of at least one probe point in the global scene representation based on the pose information; correcting the action generation parameters based on the occupancy value of at least one probe point to obtain target action generation parameters; adjusting the candidate actions based on the target action generation parameters to finally obtain the target action, and the action sequence includes the target action.

[0145] In some examples, the target action can be the action corresponding to a target action step, where the target action step can be any action step in the action sequence. Candidate actions can be actions corresponding to intermediate states generated during flow matching denoising. The end effector can be a specific execution module on the intelligent agent device used to perform the target task, such as a hand, foot, or head task execution module. At least one probe point corresponding to the end effector can be a critical probe point that determines whether task execution is successful or whether a collision is likely to occur. For example, if the end effector is a hand, its corresponding at least one probe point may include the fingertip of the gripper, the center point of the wrist, or an approximate point of the elbow. Pose information can be position, orientation, and attitude information. Additionally, action generation parameters can be adjustable parameters that affect action generation, such as the velocity field used in flow matching denoising.

[0146] For example, for each action step in the action sequence, a stream matching mechanism can be used to iteratively denoise the initial noisy actions, obtaining the actions corresponding to each action step, and finally generating the action sequence. For instance, in the process of performing stream matching denoising for the target action step, let the intermediate denoising state be... After being decoded by the action decoder, it can be mapped to candidate actions. The candidate action Further mapping to a set of critical probe points for the end effector The pose information of the detector point set includes, for example, the left and right fingertips of the gripper, the center of the wrist, and an approximate point of the elbow. Based on the pose information of this detector point set, the global scene representation can be determined according to the following formula (19). At each detection point Occupancy value at the location .

[0147] in, For a moment Global scene representation Number of Gaussian parameters For the first The center position of each Gaussian parameter For the first The transparency of each Gaussian parameter. The total penalty value can be calculated based on the occupancy value of each probe point. For example, the total penalty value is determined by the sum of the occupancy values ​​of each probe point as shown in the following formula (20).

[0148] in, Candidate actions The corresponding key probe point set of the end effector This refers to the number of detection points in the key detection point set. This is the total penalty value. The velocity field used in the next iteration for denoising is corrected using this total penalty value. For example, the velocity field can be corrected using the following formula (21).

[0149] in, This represents the original velocity field corresponding to the next iteration step. For the adjusted velocity field, This is the correction factor. Based on this correction, the intermediate state of the velocity field after denoising is... The denoising process is repeated in the next iteration, and so on. This allows the denoising process to automatically avoid collisions with various Gaussian parameters in the global scene representation, and the gradient can be processed by the action decoder. Return to the intermediate state of noise reduction Finally, after multiple iterations of denoising, the target action corresponding to the target action step can be obtained.

[0150] In this way, by calculating the occupancy value of the end effector probe points bound to the candidate actions in the global scene representation, and converting the occupancy into the gradient or penalty term for correcting the action generation parameters, the action denoising process can be guided by explicit 3D collision information, thereby having a certain obstacle avoidance capability in the process of generating action sequences, and improving safety when operating in narrow or dense spaces.

[0151] In addition, during the flow matching and denoising process of the action head, inter-layer consistency checks can also be performed. At the same time, during the inter-layer consistency check, the three-dimensional geometric features of the target entity to be operated on are used as one of the condition inputs. For example, the similarity between the velocity fields corresponding to adjacent iteration steps is calculated according to the following formula (22).

[0152] in, For the first The denoising intermediate state corresponding to each iteration step 3D geometric features corresponding to the target entity The velocity field is given by the input conditions. For the first The denoising intermediate state corresponding to each iteration step 3D geometric features corresponding to the target entity The velocity field is given by the input conditions. For the first The iteration step and the first The similarity between the velocity fields corresponding to each iteration step. If If the value is 0.95, the iterative denoising process of the current action step can be stopped early, and the corresponding action can be output. Otherwise, the denoising process continues for the next iteration step until the preset maximum iteration step is reached.

[0153] In this embodiment, by explicitly extracting the three-dimensional geometric features of the target entity to be operated from the global scene representation and using them together with the predicted image as the basis for action generation, the action planning no longer relies solely on visual texture information, but simultaneously obtains guidance on the three-dimensional geometric structure of the object, thereby significantly improving the accuracy and success rate of intelligent devices in fine contact and obstacle avoidance actions.

[0154] Furthermore, to improve the accuracy and robustness of system decision-making, in some embodiments, after step S110 above, the method may further include: determining the confidence level of the global scene representation. Accordingly, step S120 above may specifically include: if the confidence level is greater than a first confidence threshold, then based on the global scene representation, predicting the state change parameters of at least one entity in the target 3D scene at future time points.

[0155] In some examples, the confidence level of the global scene representation can be represented by a learnable cue token, which can include at least two modes, such as an Active mode when the confidence level is greater than a first confidence threshold. This Active mode indicates that the global scene representation corresponding to the target 3D scene has been constructed with high confidence (stable number of Gaussian parameters and low reprojection error). In this Active mode, full-featured 3D querying and a 3D-aware world head can be enabled to perform operations based on the global scene representation to predict the state change parameters of at least one entity in the target 3D scene at future time points.

[0156] In this way, by introducing a quantitative evaluation of the confidence level of the global scene representation and using the evaluation result as a prerequisite for enabling image prediction based on the global scene representation, the system can automatically perceive the quality of its own construction of the global scene representation. This avoids forcing image prediction based on an inaccurate 3D scene map when the accuracy of the 3D scene map is insufficient, effectively preventing the generation of cascading errors and improving the accuracy and robustness of the system's decision-making.

[0157] In other embodiments, after determining the confidence level of the global scene representation, the method may further include: if the confidence level is less than or equal to a first confidence threshold and greater than a second confidence threshold, then obtaining a target rendering image under the target viewpoint conditions from the global scene representation, and generating a predicted image for future times based on the target rendering image; if the confidence level is less than or equal to the second confidence threshold, then generating a predicted image for future times based on the actual observed image.

[0158] In some examples, the target view condition can be, for example, a bird's-eye view condition, or other view conditions, such as an entity surround view condition for the target entity to be operated on.

[0159] For example, the cue token corresponding to the confidence level of the global scene representation may include not only the Active mode described above, but also other modes. For instance, a Passive mode may be used when the confidence level is less than or equal to a first confidence threshold and greater than a second confidence threshold. This mode indicates that the global scene representation corresponding to the target 3D scene is being built or has low confidence (e.g., there are dynamic scenes or areas with missing textures). In this Passive mode, only the target rendered image from the bird's-eye view rendered by the tokenizer can be used as input to the VLM, and the world head can revert to a 2D pixel-level image prediction method for image prediction. Additionally, the cue token may also include a Disabled mode when the confidence level is less than or equal to the second confidence threshold. This mode indicates that the global scene representation has not yet been built or lacks 3DGS capability (e.g., in the initialization phase of a new environment or in the case of pure visual input without a depth sensor). In this Disabled mode, the VLA model can completely revert to a pure 2D inference mode, for example, relying solely on currently observed actual images for image prediction.

[0160] In some specific examples, the above prompt token can be inserted into the model input information according to the following formula (23) and input into the VLA model to achieve adaptive switching of the above three modes.

[0161] in, For the above three modes, The specific type of the intelligent agent device. This is a stitched image obtained by combining a rendered image under multi-view conditions with an actual observed image. The current device status. This is a task instruction.

[0162] Additionally, mode switching can be automatically triggered by the scene confidence estimator in the VLA model: when the reprojection error of the global scene representation... Continue to exceed During a frame, the mode can be upgraded from Passive or Disabled to Active.

[0163] In this way, by setting up multiple operating modes on the confidence decay path, the system can still use the 3D scene map to generate auxiliary 3D views (such as bird's-eye views) for enhanced prediction when the 3D scene map is partially available, and seamlessly switch to a pure 2D prediction mode that does not rely on any 3D information when the 3D scene map is completely unavailable. This achieves adaptive switching between the three modes and ensures the continuous availability and functional continuity of the system under any environmental perception state.

[0164] In addition, to further improve the safety of action execution during model operation, in some embodiments, when the state change parameters include pose change parameters, the method further includes: obtaining the actual observed pose change parameters of at least one entity at a future time; determining the pose deviation value based on the deviation between the actual observed pose change parameters of at least one entity and the pose change parameters; if the pose deviation value is greater than a preset safety threshold, then performing target processing, wherein the target processing includes re-executing the image prediction process.

[0165] In some examples, the actual observed pose change parameters can be the pose change parameters of at least one entity at a future time, obtained from observations reconstructed based on the real-time 3DGS global scene representation. This at least one entity can be a dynamic entity.

[0166] For example, the pose change parameters predicted by the world head can be compared in real time. Compared with the actual observed pose change parameters obtained by real-time reconstruction of 3DGS (or 6D pose estimation network) The pose deviation between the two can be calculated according to the following formula (24). .

[0167] in, These are the initial pose parameters of the entity. For Lie algebra logarithmic mapping, used to... The deviation is converted to a 6-dimensional vector norm. If ( If a preset safety threshold is set, it can be determined that an unexpected physical event has occurred, such as an object sliding, a collision, or unmodeled deformation, thereby triggering replanning and / or emergency braking.

[0168] In this way, by accurately measuring the deviation between the real-time observed changes in the actual pose of the entity and the changes in the pose predicted by the model, the system can capture three-dimensional geometric anomalies that are difficult to reflect by traditional two-dimensional image errors. This enables early detection of physical accidents such as object sliding and non-rigid deformation, and automatically triggers safety response mechanisms such as replanning, significantly enhancing the operational safety of the system after deployment.

[0169] Furthermore, after model training and before model deployment, multiple simulators can be used to cross-validate the model. Based on this, in some embodiments, before step S110 above, the method may further include: acquiring a trained visual language action model; inputting the visual language action model into multiple simulators; using the multiple simulators to evaluate the visual language action model based on a style-consistent virtual 3D scene, obtaining evaluation results output by the multiple simulators respectively; if the evaluation results output by the multiple simulators meet a preset consistency condition, then deploying the visual language action model and using the visual language action model to perform an image prediction process.

[0170] In some examples, multiple simulators may include academic benchmark simulators and industrial-grade simulators. Academic benchmark simulators may include, for example, MuJoCo, while industrial-grade simulators may include, for example, Isaac Sim and Drake. A preset consistency condition may be that the consistency score among the evaluation results output by the multiple simulators is greater than or equal to a preset score threshold.

[0171] For example, the style of virtual 3D scenes in multiple simulators can be unified in advance. The specific process may include: extracting scene assets of the virtual 3D scene from each simulator, such as meshes, textures, material parameters, etc.; baking the mesh surface into a 3D Gaussian set through a Gaussian Scene Description File (GSDF) converter, for example, sampling the Gaussian center on the triangular facet by area weight, aligning the normal vector with the minor axis of covariance, and mapping the texture color to spherical harmonic coefficients, so as to obtain a virtual 3D scene with a unified style.

[0172] In some specific examples, the trained VLA model can be deployed to multiple simulators, and the policy performance of the model in different simulators can be compared under a unified 3DGS rendering to eliminate visual domain differences. Specifically, the policy performance of the VLA model in each simulator can be evaluated to obtain the evaluation results output by each simulator, and the consistency of the evaluation results output by different simulators can be judged. For example, the task execution success rate output by different simulators can be compared across simulators for consistency and scored to obtain a cross-simulator consistency score. The cross-simulator consistency score can be calculated, for example, using the following formula (25).

[0173] in, For consistency score, Action strategy for the model The simulator MuJoCo is in the target state. The success rate of task execution is as follows. Action strategy for the model When the simulator Isaac Sim is in the target state The success rate of task execution is as follows. Action strategy for the model When the simulator Drake is in the target state The success rate of task execution is as follows. This refers to the boundary region (i.e., the set of states corresponding to the preset risk level). for The number of states contained therein. Visual consistency score, The simulator MuJoCo is in the target state. The image obtained after 3DGS rendering. The simulator IsaacSim is in the target state. The image obtained after 3DGS rendering. It is the impact factor.

[0174] For example, if the evaluation results output by multiple simulators meet a preset consistency condition, such as a consistency score... ( If a preset score threshold is set, a significant gap across simulators can be identified, triggering domain randomization or data augmentation of physical parameters, requiring retraining until validation is passed. If the evaluation results output by multiple simulators do not meet preset consistency conditions, such as consistency scores... If so, it can be determined that the VLA model has been validated and can be deployed and used on simulators or other electronic devices (such as servers or smart agent devices).

[0175] In this way, by uniformly transforming the scene into a consistent 3D scene representation (such as 3DGS) for rendering and evaluation in multiple simulator evaluations, visual rendering is decoupled from physical simulation. This allows the evaluation results to purely reflect the model's performance in different physical environments, avoiding performance evaluation bias caused by differences in visual style, and thus effectively selecting models that truly possess cross-domain transfer robustness.

[0176] Furthermore, to reduce the cost of model training and evaluation, global scene representation samples for training and evaluation can be constructed based on virtual scene data assets reconstructed from real scenes. Based on this, in some embodiments, before evaluating the visual language action model using multiple simulators based on style-uniform virtual 3D scenes and obtaining the evaluation results output by each simulator, the method further includes: acquiring 3D scene data corresponding to the real scene; constructing target global scene representation samples based on the 3D scene data; and performing randomization editing on the entity representation samples in the target global scene representation samples to obtain global scene representation samples corresponding to the multiple virtual 3D scenes respectively.

[0177] In some examples, the target global scene representation sample can be virtual scene 3D data reconstructed from real scene 3D scene data, such as a Gaussian parameter set reconstructed using 3DGS.

[0178] For example, a 3D scene data of the real environment can be obtained by capturing images of the environment for 10-20 minutes using an image acquisition device (such as a mobile phone, tablet, or robot's built-in camera). Incremental 3DGS reconstruction can then be performed on this 3D scene data to obtain... This refers to the target global scene representation sample. After instance segmentation and physical attribute annotation (such as annotation quality, friction coefficient, and coefficient of restitution), the data can be imported into a simulator as scene data assets for training and evaluation after GSDF transformation. Furthermore, after obtaining the target global scene representation samples, entity-level randomization editing can be performed on the entity representation samples to generate diverse sample data. Specifically, randomization editing can include at least one of the following: dynamic entity replacement processing, pose randomization processing, lighting randomization processing, and viewpoint randomization processing.

[0179] In some specific examples, dynamic entity replacement processing may include, for instance, replacing the entity representation samples corresponding to at least one of the dynamic entities. The entity representation is replaced with another object (e.g., entities with similar semantics retrieved from a pre-defined 3D asset library) to maintain physical interaction relationships. Pose randomization processing may include, for example, applying a random rigid body transformation to at least one of the dynamic entities. This is done to generate novel spatial configurations. Illumination randomization processes may include, for example, adjusting the spherical harmonic coefficients of each Gaussian parameter. Scene-level lighting parameters are used to simulate lighting at different times of day. Viewpoint randomization may include, for example, rendering camera trajectory viewpoint images not captured during model training to enhance viewpoint generalization. In this way, scene representations of multiple virtual 3D scenes corresponding to multiple enhanced trajectories can be generated based on a single real scene, resulting in multiple global scene representation samples for pre-training and evaluation of the VLA model.

[0180] In this way, by reconstructing a basic target global scene representation sample from a real scene and randomly editing at least one entity representation sample, a massive number of diverse 3D scene representations with reasonable geometric structures are automatically generated from a real entity sample. This enables the construction of a large-scale, high-fidelity, and style-consistent cross-simulator evaluation and training scenario at extremely low cost, significantly improving the model's generalization ability to novel spatial configurations and training and evaluation coverage.

[0181] In addition, the embodiments of this application can use a four-stage progressive training strategy to train the VLA model. The data flow and model state of each stage are described in detail below.

[0182] The first stage is the pre-training of the 3D scene tokenizer: the training objective is to learn the compressed representation of the 3D scene representation from the 3DGS to the compact 3D token sequence; the training data can be global scene representation sample data constructed based on real scenes and synthetic 3D scene asset data; the training process can adopt supervised learning, and the loss function adopts rendering-based reconstruction loss (such as RGB L1 + depth L1 + LPIPS perceptual loss) and vector quantization (VQ) commitment loss; after training, the pre-trained tokenizer encoder, tokenizer decoder and corresponding codebook (i.e. feature sequence) can be obtained.

[0183] In some specific examples, the aforementioned Tokenizer can be trained using a rendering-based reconstruction loss, avoiding the difficulties associated with directly comparing a variable number of Gaussian parameters. For instance, the loss function shown in formula (26) can be used during training. .

[0184] in, The set of Gaussian parameters output by the Tokenizer decoder. For pixel coordinates The actual RGB color value at that location; For pixel coordinates Based on The rendered predicted RGB color values; For pixel coordinates The actual depth value at that location; For pixel coordinates Based on The rendered predicted depth value; For codebook The corresponding embedding vector sequence; To stop the gradient operation; Weights for depth loss; VQ commitment loss weights are applied. Reconstruction loss (including RGB color reconstruction loss and depth reconstruction loss) is calculated by decoding a Gaussian rasterized rendered image, ensuring that the tokenizer preserves geometric and appearance information.

[0185] The second stage is the pre-training of the 3D perception world head: the training objective is to learn to predict entity-level state transitions and appearance changes; training data can be demonstration data of intelligent agent devices (such as robots) with video frames and unlabeled human videos; the training process can adopt supervised learning, and the loss function is constructed by relying on 2D rendering loss and depth loss (without Gaussian ground truth). During the training process, the tokenizer can be frozen (or fine-tuned in low rank), and the world head applies transformation through the continuous Gaussian parameter set output by the decoder, and the gradient is backpropagated through differentiable rasterization; the training strategy includes freezing the VLM base parameters and fine-tuning the parameters corresponding to the world head (LoRA, rank=64).

[0186] In some specific examples, the training process of the world head does not require the true values ​​of the 3D Gaussian parameters, as shown in the following formula (27), and the loss function can be constructed based on the 2D rendering loss and depth loss. .

[0187] in, For real-time observation images of future moments, A predicted image of future moments obtained from the world's top predictions. For a true depth map of future moments, This is a predicted depth map of future moments obtained from the world's top predictions. For LPIPS loss weights, The depth loss weights are used. By using differentiable rasterization, the rendered loss gradient is backpropagated to the world head via Gaussian parameters and rigid body transformation, indirectly optimizing the prediction accuracy of entity-level state transformation parameters.

[0188] The third phase is VLA joint training: the training objective is to establish a vision-language-3D-action mapping; training data can utilize multi-task data from cross-agent devices and synthetic data generated by 3DGS; training strategies may include joint training of the tokenizer (e.g., low-rank update), world head, and action head, where the world head can employ an error-blocking action mask, and the action sequence length... It can be dynamically selected based on the task type, such as short-term tasks (e.g., grabbing, placing). Long-duration tasks (such as navigation, multi-step operations) The training process can use the loss function shown in the following formula (28). .

[0189] in, This is a loss to the world. For the loss of the action head, For stream matching loss, For collision penalty weights, This is a collision penalty.

[0190] The fourth stage involves task fine-tuning and alignment training with residual reinforcement learning (RL). The training objective is to optimize for specific scenarios and improve industrial-grade success rates. Training strategies include: performing LoRA low-rank fine-tuning (rank=64) on the dual-head decoder, with optional freezing or low-rank updates of the base model; validating the model using multiple simulators and 3DGS; freezing the main model and training the residual policy network. (Can contain 2 layers of MLP and have <1M parameters) Collect recovery behavior in real intelligent agent devices or high-fidelity simulators; collect residual success trajectories to perform light supervised fine-tuning (SFT) on the action head and internalize the recovery behavior into the VLA master model.

[0191] In addition, in some embodiments, the VLA master model can be deployed on a server, which can simultaneously serve multiple intelligent agent devices, thereby achieving model inference and maintenance of global scene representation through an asynchronous inference mechanism.

[0192] For example, on the intelligent agent device side, incremental updates of the local global scene representation (such as a 3DGS map) can be maintained, only updating the changed entity-level transformation parameters. With the addition of Gaussian parameters uploaded to the server, Gaussian parameters for static entities (such as the background) only need to be transmitted once, reducing bandwidth usage by over 80%. On the server side, inference can be performed based on a complete 3D token sequence, supporting multiple intelligent agent devices sharing the same scene memory. Furthermore, the server can automatically select the model size based on the construction status (or confidence level) of the global scene representation and task complexity. For example, in Active mode, a 7B model based on 3D perception can be prioritized, while in Disabled mode, a 0.5B edge model based on 2D can be used as a fallback.

[0193] Based on the above embodiments, a specific example will be given below in conjunction with the overall system architecture to better illustrate the entire solution.

[0194] In some specific examples, embodiments of this application provide an explicit spatial memory VLA control system based on a 3D Gaussian scene tokenizer. This system adopts a six-layer closed-loop architecture of "perception-encoding-memory-imagination-execution-verification," comprising the following functional levels from bottom to top.

[0195] Perception layer: Real-time acquisition of environmental observations through multi-view RGB-D cameras or monocular vision synchronous localization and mapping (such as SLAM) to construct or update explicit 3D Gaussian scene representation.

[0196] Encoding layer: By compressing millions of Gaussian parameters into a compact 3D token sequence through a 3D scene tokenizer, the scale gap between 3DGS and VLM is bridged.

[0197] Memory layer: The 3D Token sequence is used as a queryable and persistent spatial memory of the VLM, supporting rendering queries based on conditions such as viewpoint, lighting, and semantics.

[0198] Imaginary layer: The world head predicts the pose and appearance changes of the solid body at future moments, which are applied to the continuous Gaussian features output by the Tokenizer decoder and then rasterized by 3DGS to generate a geometrically consistent predicted image.

[0199] Execution layer: The action head generates a continuous action sequence based on adaptive truncated flow matching and integrates 3D geometric features from 3D token sequence decoding to assist in contact point estimation.

[0200] Verification layer: 3DGS high-fidelity rendering is introduced into the cross-validation of multiple simulators, and unexpected physical events are detected at runtime through 3D geometric consistency monitoring.

[0201] It should be noted that any parts not fully described in the above examples can be referred to the aforementioned embodiments, and will not be repeated here.

[0202] The above text combined Figures 1 to 3 The image prediction method embodiments of this application are described in detail below, in conjunction with... Figure 4 This application provides a detailed description of embodiments of the image prediction apparatus. It should be understood that the descriptions of the image prediction method embodiments correspond to the descriptions of the image prediction apparatus embodiments; therefore, any parts not described in detail can be found in the preceding method embodiments.

[0203] Figure 4 The diagram shown is a structural schematic of an image prediction device provided in an embodiment of this application. Figure 4 As shown, the image prediction device 400 provided in this application embodiment includes: The acquisition module 401 is used to acquire the global scene representation of the target 3D scene. The global scene representation is a 3D information representation constructed based on the scene information perceived in real time. Prediction module 402 is used to predict the state change parameters of at least one entity in the target 3D scene at future time based on the global scene representation; The processing module 403 is used to transform the scene representation corresponding to the current moment in the global scene representation according to the state change parameters, so as to obtain the predicted scene representation for the future moment. The generation module 404 is used to generate a predicted image for future times based on the predicted scene representation.

[0204] In one embodiment of this application, at least one entity includes a dynamic entity, and the state change parameters include pose change parameters and / or appearance change parameters. Accordingly, the processing module 403 is further configured to: obtain a scene representation corresponding to the current moment from the global scene representation to obtain a current scene representation; obtain an entity representation corresponding to the dynamic entity from the current scene representation to obtain a dynamic entity representation; perform pose transformation processing on the dynamic entity representation according to the pose change parameters, and / or perform appearance transformation processing on the dynamic entity representation according to the appearance change parameters to obtain a predicted dynamic entity representation of the dynamic entity at a future moment; and determine a predicted scene representation at a future moment based on the predicted dynamic entity representation.

[0205] In one embodiment of this application, at least one entity further includes a static entity, and the state change parameters further include illumination change parameters. Accordingly, the processing module 403 is further configured to: obtain the entity representation corresponding to the static entity from the current scene representation to obtain a static entity representation; perform illumination transformation processing on the static entity representation according to the illumination change parameters to obtain a predicted static entity representation for future times; and determine a predicted scene representation for future times based on the predicted dynamic entity representation and the predicted static entity representation.

[0206] In one embodiment of this application, the global scene representation includes multiple Gaussian parameters. Accordingly, the prediction module 402 is further configured to: compress and encode the multiple Gaussian parameters in the global scene representation into discrete feature sequences; and use the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at a future time based on the feature sequences.

[0207] In one embodiment of this application, the prediction module 402 is further configured to: perform entity-level semantic segmentation processing on multiple Gaussian parameters in the global scene representation to obtain a Gaussian parameter set corresponding to at least one entity; and encode and quantize the Gaussian parameter set corresponding to at least one entity to obtain a discrete feature sequence.

[0208] In one embodiment of this application, at least one entity includes a dynamic entity and a static entity. Accordingly, the prediction module 402 is further configured to: extract features from the Gaussian parameter set corresponding to the dynamic entity to obtain instance features, and extract features from the Gaussian parameter set corresponding to the static entity to obtain background features; fuse the instance features and background features to obtain continuous scene features; and quantize the continuous scene features to obtain a discrete feature sequence.

[0209] In one embodiment of this application, the prediction module 402 is further configured to: map the set of Gaussian parameters corresponding to the dynamic entity to point-level features; and aggregate the point-level features to entity-level features to obtain instance features.

[0210] In one embodiment of this application, the prediction module 402 is further configured to: determine the visual language features at the current moment based on the feature sequence using the visual language sub-model in the visual language action model; and predict the state change parameters of at least one entity in the target 3D scene at a future moment based on the feature sequence and the visual language features using the world head in the visual language action model.

[0211] In one embodiment of this application, the prediction module 402 is further configured to: utilize the visual language sub-model in the visual language action model to obtain rendered images under multiple viewpoint conditions and / or lighting conditions from the global scene representation based on feature sequences; stitch the rendered images with the actual observed images at the current moment to obtain a stitched image; encode the stitched image and integrate it into the inference process of the visual language sub-model through a cross-attention mechanism to output the visual language features at the current moment.

[0212] In one embodiment of this application, the multiple perspective conditions include at least two of the following: the current actual observation perspective condition, the bird's-eye view perspective condition, and the entity surround perspective condition.

[0213] In one embodiment of this application, the generation module 404 is further configured to: obtain the three-dimensional geometric features of the target entity from the global scene representation, wherein the target entity is the entity to be operated on in the target task; and generate an action sequence for performing the target task based on the three-dimensional geometric features and the predicted image.

[0214] In one embodiment of this application, the generation module 404 is further configured to: extract the entity representation corresponding to the target entity from the global scene representation to obtain the target entity representation; determine the surface normal vector, local occupancy rate and centroid position of the target entity based on the target entity representation; and determine the three-dimensional geometric features of the target entity based on the surface normal vector, local occupancy rate and centroid position.

[0215] In one embodiment of this application, the generation module 404 is further configured to: generate candidate actions for performing a target task based on three-dimensional geometric features and a predicted image; determine the pose information of at least one probe point corresponding to the end effector based on the candidate actions; determine the occupancy value of at least one probe point in the global scene representation based on the pose information; correct the action generation parameters based on the occupancy value of at least one probe point to obtain target action generation parameters; adjust the candidate actions based on the target action generation parameters to finally obtain the target action, wherein the action sequence includes the target action.

[0216] In one embodiment of this application, the acquisition module 401 is further configured to: determine the confidence level of the global scene representation. Correspondingly, the prediction module 402 is further configured to: if the confidence level is greater than a first confidence threshold, predict the state change parameters of at least one entity in the target 3D scene at a future time based on the global scene representation.

[0217] In one embodiment of this application, the acquisition module 401 is further configured to: if the confidence level is less than or equal to a first confidence threshold and greater than a second confidence threshold, acquire a target rendering image under the target viewpoint condition from the global scene representation, and generate a prediction image for future moments based on the target rendering image; if the confidence level is less than or equal to the second confidence threshold, generate a prediction image for future moments based on the actual observed image.

[0218] In one embodiment of this application, the state change parameters include pose change parameters, and the prediction module 402 is further configured to: obtain the actual observed pose change parameters of at least one entity at a future time; determine the pose deviation value based on the deviation between the actual observed pose change parameters of at least one entity and the pose change parameters; if the pose deviation value is greater than a preset safety threshold, then perform target processing, wherein the target processing includes re-executing the image prediction process.

[0219] In one embodiment of this application, the acquisition module 401 is further configured to: acquire a trained visual language action model; input the visual language action model into multiple simulators; use the multiple simulators to evaluate the visual language action model based on a virtual 3D scene with a unified style, and obtain evaluation results output by the multiple simulators respectively; if the evaluation results output by the multiple simulators respectively meet a preset consistency condition, then deploy the visual language action model and use the visual language action model to perform an image prediction process.

[0220] In one embodiment of this application, the acquisition module 401 is further configured to: acquire three-dimensional scene data corresponding to the real scene; construct a target global scene representation sample based on the three-dimensional scene data; and perform randomization editing on the entity representation samples in the target global scene representation sample to obtain global scene representation samples corresponding to multiple virtual three-dimensional scenes respectively.

[0221] Below, for reference Figure 5 This describes an electronic device according to embodiments of the present application. Figure 5 The diagram shown is a structural schematic of an electronic device provided in an exemplary embodiment of this application.

[0222] like Figure 5 As shown, the electronic device 500 includes one or more processors 501 and memory 502.

[0223] The processor 501 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.

[0224] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the image prediction methods of the various embodiments of this application described above and / or other desired functions. The computer-readable storage medium may also store various contents such as global scene representation, state change parameters, predicted scene representation, predicted image, etc.

[0225] In one example, the electronic device 500 may also include an input device 503 and an output device 504, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0226] The input device 503 may include, for example, a keyboard, a mouse, etc.

[0227] The output device 504 can output various information to the outside, including global scene representation, state change parameters, predicted scene representation, predicted image, etc. The output device 504 may include, for example, a display, speaker, printer, and communication network and its connected remote output devices, etc.

[0228] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device 500 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 500 may include any other suitable components depending on the specific application.

[0229] In addition to the methods and apparatus described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image prediction methods according to various embodiments of this application described above.

[0230] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0231] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the image prediction methods according to various embodiments of this application described above.

[0232] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0233] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0234] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0235] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0236] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0237] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. An image prediction method, characterized in that, include: A global scene representation of the target 3D scene is obtained. The global scene representation is a 3D Gaussian representation of the target 3D scene constructed in real time based on incremental 3D Gaussian sputtering real-time localization and mapping technology. The global scene representation includes multiple Gaussian parameters. The multiple Gaussian parameters in the global scene representation are subjected to entity-level semantic segmentation to obtain a set of Gaussian parameters corresponding to at least one entity in the target 3D scene. Encode and quantize the Gaussian parameter set corresponding to each of the at least one entity to obtain a discrete feature sequence; Using the world head in the visual language action model, the state change parameters of at least one entity in the target 3D scene are predicted at a future time based on the feature sequence. The at least one entity includes dynamic entities and static entities. The state change parameters include at least one of pose change parameters, appearance change parameters, and illumination change parameters. Based on the state change parameters, the scene representation corresponding to the current moment in the global scene representation is transformed to obtain the predicted scene representation for the future moment; Based on the predicted scene representation, a predicted image for the future time is generated.

2. The method according to claim 1, characterized in that, The step of transforming the scene representation corresponding to the current time in the global scene representation according to the state change parameters to obtain the predicted scene representation for the future time includes: Obtain the scene representation corresponding to the current moment from the global scene representation to obtain the current scene representation; Obtain the entity representation corresponding to the dynamic entity from the current scene representation to obtain the dynamic entity representation; The dynamic entity representation is subjected to pose transformation processing based on the pose change parameters, and / or the dynamic entity representation is subjected to appearance transformation processing based on the appearance change parameters, to obtain the predicted dynamic entity representation of the dynamic entity at the future time. Based on the predicted dynamic entity representation, the predicted scene representation for the future time is determined.

3. The method according to claim 2, characterized in that, After obtaining the scene representation corresponding to the current time from the global scene representation to obtain the current scene representation, the method further includes: Obtain the entity representation corresponding to the static entity from the current scene representation to obtain the static entity representation; The static entity representation is subjected to illumination transformation processing based on the illumination change parameters to obtain the predicted static entity representation for the future time. The step of determining the predicted scene representation at the future time based on the predicted dynamic entity representation includes: Based on the predicted dynamic entity representation and the predicted static entity representation, the predicted scene representation for the future time is determined.

4. The method according to claim 1, characterized in that, The process of encoding and quantizing the Gaussian parameter sets corresponding to each of the at least one entity to obtain discrete feature sequences includes: Feature extraction is performed on the Gaussian parameter set corresponding to the dynamic entity to obtain instance features, and feature extraction is performed on the Gaussian parameter set corresponding to the static entity to obtain background features; The instance features and the background features are fused to obtain continuous scene features; The continuous features of the scene are quantized to obtain a discrete feature sequence.

5. The method according to claim 4, characterized in that, The step of extracting features from the Gaussian parameter set corresponding to the dynamic entity to obtain instance features includes: Map the set of Gaussian parameters corresponding to the dynamic entity to point-level features; The point-level features are aggregated into entity-level features to obtain instance features.

6. The method according to claim 1, characterized in that, The method of using the world head in the visual language action model to predict the state change parameters of at least one entity in the target 3D scene at a future time based on the feature sequence includes: Using the visual language sub-model in the visual language action model, the visual language features at the current moment are determined based on the feature sequence; Using the world head in the visual language action model, the state change parameters of at least one entity in the target 3D scene at future moments are predicted based on the feature sequence and the visual language features.

7. The method according to claim 6, characterized in that, The step of determining the visual language features at the current moment based on the feature sequence using the visual language sub-model in the visual language action model includes: Using the visual language sub-model in the visual language action model, based on the feature sequence, render images under multiple viewpoint conditions and / or lighting conditions are obtained from the global scene representation; The rendered image is stitched together with the actual observed image at the current moment to obtain a stitched image; The stitched image is encoded and then incorporated into the inference process of the visual language sub-model through a cross-attention mechanism to output the visual language features at the current moment.

8. The method according to claim 7, characterized in that, The multiple perspective conditions include at least two of the following: the current actual observation perspective condition, the bird's-eye view perspective condition, and the entity surround perspective condition.

9. The method according to claim 1, characterized in that, After generating the predicted image for the future time based on the predicted scene representation, the method further includes: The three-dimensional geometric features of the target entity are obtained from the global scene representation, where the target entity is the entity to be operated on in the target task; Based on the three-dimensional geometric features and the predicted image, an action sequence for performing the target task is generated.

10. The method according to claim 9, characterized in that, The step of obtaining the three-dimensional geometric features of the target entity from the global scene representation includes: Extract the entity representation corresponding to the target entity from the global scene representation to obtain the target entity representation; Based on the target entity representation, the surface normal vector, local occupancy rate, and centroid position of the target entity are determined; The three-dimensional geometric features of the target entity are determined based on the surface normal vector, the local occupancy rate, and the centroid position.

11. The method according to claim 9, characterized in that, The step of generating an action sequence to perform the target task based on the three-dimensional geometric features and the predicted image includes: Based on the three-dimensional geometric features and the predicted image, candidate actions for performing the target task are generated; Based on the candidate actions, determine the pose information of at least one probe point corresponding to the end effector; Based on the pose information, determine the occupancy value of at least one probe point in the global scene representation; The action generation parameters are corrected based on the occupancy value of the at least one detection point to obtain the target action generation parameters; The candidate actions are adjusted based on the target action generation parameters to finally obtain the target action, and the action sequence includes the target action.

12. The method according to claim 1, characterized in that, After obtaining the global scene representation of the target 3D scene, the method further includes: Determine the confidence level of the global scene representation; The step of predicting the state change parameters of at least one entity in the target 3D scene at future times based on the global scene representation includes: If the confidence level is greater than the first confidence level threshold, then based on the global scene representation, predict the state change parameters of at least one entity in the target 3D scene at future time.

13. The method according to claim 12, characterized in that, After determining the confidence level of the global scene representation, the method further includes: If the confidence level is less than or equal to the first confidence threshold and greater than the second confidence threshold, then the target rendering image under the target view condition is obtained from the global scene representation, and the predicted image for the future time is generated based on the target rendering image; If the confidence level is less than or equal to the second confidence level threshold, then a predicted image for the future time is generated based on the actual observed image.

14. The method according to claim 1, characterized in that, Also includes: Obtain the actual observed pose change parameters of the at least one entity at the future time; Based on the deviation between the actual observed pose change parameters of the at least one entity and the pose change parameters, the pose deviation value is determined. If the pose deviation value is greater than a preset safety threshold, then target processing is performed, wherein the target processing includes re-executing the image prediction process.

15. The method according to claim 1, characterized in that, Before obtaining the global scene representation of the target 3D scene, the method further includes: Obtain a trained visual language action model; The visual language action model is input into multiple simulators; Using the multiple simulators, the visual language action model is evaluated based on a virtual 3D scene with a unified style, and the evaluation results output by the multiple simulators are obtained respectively. If the evaluation results output by the multiple simulators meet the preset consistency conditions, then the visual language action model is deployed, and the image prediction process is performed using the visual language action model.

16. The method according to claim 15, characterized in that, Before evaluating the visual language action model using the multiple simulators based on a style-uniform virtual 3D scene and obtaining the evaluation results output by the multiple simulators, the method further includes: Obtain 3D scene data corresponding to the real scene; Construct a target global scene representation sample based on the aforementioned 3D scene data; The entity representation samples in the target global scene representation sample are randomized and edited to obtain global scene representation samples corresponding to the multiple virtual 3D scenes respectively.

17. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the image prediction method according to any one of claims 1 to 16.

18. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the image prediction method according to any one of claims 1 to 16.

19. A computer program product, characterized in that, The computer program product includes instructions that, when executed on an electronic device, cause the electronic device to implement the image prediction method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method and computer product for view prediction

    CN116630366A

  • Grabbing method and system based on Gaussian spatter, robot and storage medium

    CN118769237A

  • Multi-modal three-dimensional instance segmentation method based on three-dimensional Gaussian splashing

    CN119296104A