Point cloud guidance-based cross-modal image generation method and device, equipment and medium
Patent Information
- Application Number
- CN202610849239.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]目前,相关的跨模态图像生成过程中,生成结果往往更容易受到图像外观信息的影响,而对目标场景中的空间几何关系约束不足,在复杂场景下,若生成过程缺少稳定的几何约束,容易导致生成图像中的物体轮廓、深度层次、遮挡关系或者空间位置关系与目标场景不一致,从而影响生成图像在自动驾驶仿真、机器人视觉训练等场景中的可用性;同时,在多模态数据参与图像生成的情况下,不同模态数据之间通常存在表达形式、空间分布和特征尺度等方面的差异,如果不同模态信息之间的对应关系不够准确,则可能导致生成图像在结构区域、边缘区域或遮挡区域出现局部错位,使生成结果难以同时兼顾场景结构稳定性和图像视觉连续性
本公开的示例实施例中的基于点云引导的跨模态图像生成方法,一方面,通过获取目标场景对应的激光雷达点云数据和图像数据,并根据激光雷达点云数据确定与图像数据对应的几何监督信息,使后续模型训练过程能够在图像外观信息之外引入目标场景的空间几何约束,从而减少仅依赖图像数据进行场景表达时容易出现的深度层次不清、物体轮廓偏移或者空间位置关系不稳定等问题,提高四维高斯散射场景模型对目标场景真实空间结构的表征准确性;另一方面,通过图像数据和几何监督信息对预设的初始四维高斯散射场景模型进行训练,得到目标四维高斯散射场景模型,使四维高斯散射场景模型在学习图像纹理、颜色等视觉外观信息的同时,能够受到几何监督信息对场景几何结构的约束,从而使训练得到的目标四维高斯散射场景模型能够更准确地保持目标场景中道路、车辆、行人、建筑物等对象之间的空间位置关系、深度关系和遮挡关系,降低复杂场景下生成图像与目标场景真实空间结构不一致的风险,提高生成图像在自动驾驶仿真、机器人视觉训练等场景中的可用性;再一方面,通过根据目标四维高斯散射场景模型生成与目标场景对应的几何条件信息,并结合几何条件信息和图像数据进行跨模态特征融合,使图像生成过程能够在融合特征中同时保留来自图像数据的视觉外观信息以及来自目标四维高斯散射场景模型的几何结构信息,从而降低不同模态数据之间因表达形式、空间分布和特征尺度差异导致的对应不准确问题,减少生成图像在结构区域、边缘区域或遮挡区域出现局部错位的情况,提高生成图像的结构稳定性、视觉连续性以及与目标场景空间结构之间的一致性。
Smart Images

Figure CN122737352A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of four-dimensional Gaussian scattering cross-modal image generation technology, and more specifically, to a point cloud-guided cross-modal image generation method, apparatus, device, and medium. Background Technology
[0002] With the development of technologies such as autonomous driving, intelligent robots, and virtual scene construction, the image generation quality for complex target scenes has gradually become an important factor affecting model training, simulation verification, and data augmentation. Target scenes typically include roads, vehicles, pedestrians, buildings, and other environmental objects. There are complex spatial relationships, depth relationships, and occlusion relationships between different objects. Therefore, the generated scene images not only need to have a good visual appearance, but also need to match the real spatial structure of the target scene.
[0003] Currently, in related cross-modal image generation processes, the generated results are often more susceptible to the influence of image appearance information, while the constraints on the spatial geometric relationships in the target scene are insufficient. In complex scenes, if the generation process lacks stable geometric constraints, it is easy for the object contours, depth levels, occlusion relationships, or spatial positional relationships in the generated image to be inconsistent with the target scene, thereby affecting the usability of the generated image in scenarios such as autonomous driving simulation and robot vision training. At the same time, when multimodal data participates in image generation, there are usually differences in expression, spatial distribution, and feature scale between different modal data. If the correspondence between different modal information is not accurate enough, it may lead to local misalignment in structural regions, edge regions, or occlusion regions of the generated image, making it difficult for the generated result to simultaneously take into account the stability of scene structure and the visual continuity of the image.
[0004] Therefore, how to improve the consistency between the generated image and the spatial structure of the target scene during cross-modal image generation, and reduce the impact of inaccurate correspondence between different modal information on the generation results, has become an urgent technical problem to be solved.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a method, apparatus, device, and medium for cross-modal image generation based on point cloud guidance, thereby improving the scene structure stability and visual continuity of the generated image.
[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0008] According to a first aspect of the present disclosure, a point cloud-guided cross-modal image generation method is provided, comprising: Acquire lidar point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the lidar point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projection based on the lidar point cloud data. The preset initial four-dimensional Gaussian scattering scene model is trained using the image data and the geometric supervision information to obtain the target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. Based on the target four-dimensional Gaussian scattering scene model, geometric condition information corresponding to the target scene is generated, and feature extraction, spatial alignment and fusion processing are performed on the geometric condition information and the image data to obtain fused features; Image generation is performed based on the fusion features to obtain a target image corresponding to the target scene.
[0009] In some example embodiments of this disclosure, based on the foregoing scheme, determining the geometric supervision information corresponding to the image data according to the lidar point cloud data includes: The lidar point cloud data is preprocessed to obtain preprocessed lidar point cloud data. The point cloud preprocessing includes at least one of noise filtering, outlier removal, and coordinate system one. Based on the calibration relationship between the lidar point cloud data and the image data, the preprocessed lidar point cloud data is mapped to the coordinate space corresponding to the image data; Based on the mapped lidar point cloud data, pixel-level depth supervision information corresponding to the image data is determined, and the pixel-level depth supervision information is used as the geometric supervision information.
[0010] In some exemplary embodiments of this disclosure, based on the foregoing scheme, the step of training a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model includes: Determine the image reconstruction loss of the initial four-dimensional Gaussian scattering scene model based on the image data; The geometric constraint loss of the initial four-dimensional Gaussian scattering scene model is determined based on the geometric supervision information. Based on the image reconstruction loss and the geometric constraint loss, the model parameters of the initial four-dimensional Gaussian scattering scene model are updated to obtain the target four-dimensional Gaussian scattering scene model.
[0011] In some exemplary embodiments of this disclosure, based on the foregoing scheme, updating the model parameters of the initial four-dimensional Gaussian scattering scene model according to the image reconstruction loss and the geometric constraint loss includes: Based on the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model, determine the sparse constraint loss; The model training loss is determined based on the image reconstruction loss, the geometric constraint loss, and the sparse constraint loss. The Gaussian unit parameters in the initial four-dimensional Gaussian scattering scene model are optimized based on the model training loss, so that the optimized four-dimensional Gaussian scattering scene model can represent the spatial geometry of the target scene.
[0012] In some example embodiments of this disclosure, based on the foregoing scheme, generating geometric condition information corresponding to the target scene according to the target four-dimensional Gaussian scattering scene model includes: Conditional rendering is performed on the target four-dimensional Gaussian scattering scene model to obtain a geometric condition map; Based on the geometric condition diagram, generate multi-channel geometric condition information; The geometric condition map includes at least one of a depth condition map, a density condition map, and a reflectivity fusion condition map.
[0013] In some exemplary embodiments of this disclosure, based on the foregoing scheme, the step of performing feature extraction, spatial alignment, and fusion processing on the geometric condition information and the image data to obtain fused features includes: The geometric condition information is used to extract features to obtain geometric features; Feature extraction is performed on the image data to obtain image features; Based on the spatial cross-attention mechanism, the geometric features and the image features are spatially aligned to obtain aligned geometric features and aligned image features; Obtain the weather state information corresponding to the target scene, and determine the weather gating parameter based on the weather state information. The weather gating parameter is used to adjust the fusion ratio of the aligned geometric features and the aligned image features under different weather conditions. Based on the weather gating parameters, the aligned geometric features and the aligned image features are fused to obtain the fused features.
[0014] In some example embodiments of this disclosure, based on the foregoing scheme, the step of fusing the aligned geometric features and the aligned image features based on the weather gating parameters to obtain the fused features includes: Based on the density information in the geometric condition information, determine the structural weight information corresponding to different regions in the target scene; Based on the structural weight information and the weather gating parameters, determine the geometric fusion weight and image fusion weight corresponding to different regions in the target scene; Based on the geometric fusion weights and the image fusion weights, the aligned geometric features and the aligned image features are adaptively weighted and fused to obtain initial fused features; The initial fusion features are subjected to residual refinement to obtain the fusion features.
[0015] According to a second aspect of the present disclosure, a point cloud-guided cross-modal image generation apparatus is provided, comprising: The information determination module is used to acquire lidar point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the lidar point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projection based on the lidar point cloud data. The model training module is used to train a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. The feature fusion module is used to generate geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and to perform feature extraction, spatial alignment and fusion processing on the geometric condition information and the image data to obtain fused features; An image generation module is used to generate an image based on the fusion features to obtain a target image corresponding to the target scene.
[0016] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions that, when executed by the processor, implement the point cloud-guided cross-modal image generation method of the first aspect.
[0017] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the point cloud-guided cross-modal image generation method of the first aspect.
[0018] The technical solutions provided in this disclosure may have the following beneficial effects: The point cloud-guided cross-modal image generation method in the exemplary embodiments of this disclosure, on the one hand, acquires LiDAR point cloud data and image data corresponding to the target scene, and determines geometric supervision information corresponding to the image data based on the LiDAR point cloud data. This enables the subsequent model training process to introduce spatial geometric constraints of the target scene in addition to image appearance information, thereby reducing problems such as unclear depth levels, object contour offsets, or unstable spatial positional relationships that are prone to occur when relying solely on image data for scene representation, and improving the accuracy of the four-dimensional Gaussian scattering scene model in representing the real spatial structure of the target scene. On the other hand, by training a preset initial four-dimensional Gaussian scattering scene model with image data and geometric supervision information, a target four-dimensional Gaussian scattering scene model is obtained. This allows the four-dimensional Gaussian scattering scene model to learn visual appearance information such as image texture and color while being constrained by geometric supervision information on the scene's geometric structure, thus enabling the trained target four-dimensional Gaussian scattering scene model to... To more accurately maintain the spatial positional, depth, and occlusion relationships between objects such as roads, vehicles, pedestrians, and buildings in the target scene, the risk of inconsistencies between the generated image and the real spatial structure of the target scene in complex scenarios is reduced, improving the usability of the generated image in scenarios such as autonomous driving simulation and robot vision training. On the other hand, by generating geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and combining the geometric condition information with image data for cross-modal feature fusion, the image generation process can simultaneously retain the visual appearance information from the image data and the geometric structure information from the target four-dimensional Gaussian scattering scene model in the fused features. This reduces the inaccuracy of correspondence between different modal data due to differences in expression, spatial distribution, and feature scale, reduces the occurrence of local misalignment in structural, edge, or occluded regions of the generated image, and improves the structural stability, visual continuity, and consistency with the spatial structure of the target scene of the generated image.
[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0021] Figure 1 The illustration shows a schematic flowchart of a point cloud-guided cross-modal image generation method according to some embodiments of the present disclosure.
[0022] Figure 2 The schematic diagram illustrates a process for training a target four-dimensional Gaussian scattering scene model according to some embodiments of the present disclosure.
[0023] Figure 3 The illustration schematically shows a process diagram for constructing fusion features according to some embodiments of the present disclosure.
[0024] Figure 4 The illustration shows a schematic diagram of the composition of a point cloud-guided cross-modal image generation apparatus according to some embodiments of the present disclosure.
[0025] Figure 5 The schematic diagram illustrates the structural schematic of a computer system of an electronic device according to some embodiments of the present disclosure.
[0026] Figure 6 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is shown.
[0027] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0029] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0031] Furthermore, the accompanying drawings are for illustrative purposes only and are not necessarily drawn to scale. The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0032] In this example embodiment, a point cloud-guided cross-modal image generation method is first provided. This point cloud-guided cross-modal image generation method can be applied to terminal devices, such as mobile phones, computers and other electronic devices, or to servers. This embodiment does not make any special limitations on this, and the following description will take the server executing the method as an example. Figure 1 The illustration schematically shows a flowchart of a point cloud-guided cross-modal image generation method according to some embodiments of the present disclosure. Reference Figure 1 As shown, the point cloud-guided cross-modal image generation method may include the following steps: Step S110: Obtain LiDAR point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the LiDAR point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projection based on the LiDAR point cloud data. Step S120: Train the preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain the target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. Step S130: Generate geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and perform feature extraction, spatial alignment and fusion processing on the geometric condition information and the image data to obtain fused features; Step S140: Based on the fusion features, an image is generated to obtain a target image corresponding to the target scene.
[0033] According to the point cloud-guided cross-modal image generation method in this example embodiment, on the one hand, by acquiring the LiDAR point cloud data and image data corresponding to the target scene, and determining the geometric supervision information corresponding to the image data based on the LiDAR point cloud data, the subsequent model training process can introduce the spatial geometric constraints of the target scene in addition to the image appearance information. This reduces problems such as unclear depth levels, object contour offsets, or unstable spatial positional relationships that are prone to occur when relying solely on image data for scene representation, thereby improving the accuracy of the four-dimensional Gaussian scattering scene model in representing the real spatial structure of the target scene. On the other hand, by training the preset initial four-dimensional Gaussian scattering scene model with image data and geometric supervision information, the target four-dimensional Gaussian scattering scene model is obtained. This allows the four-dimensional Gaussian scattering scene model to be constrained by the geometric supervision information on the scene's geometric structure while learning visual appearance information such as image texture and color, thus enabling the trained target four-dimensional Gaussian scattering scene model to be more... Accurately maintaining the spatial positional, depth, and occlusion relationships between objects such as roads, vehicles, pedestrians, and buildings in the target scene reduces the risk of inconsistencies between the generated image and the real spatial structure of the target scene in complex scenarios, improving the usability of the generated image in scenarios such as autonomous driving simulation and robot vision training. On the other hand, by generating geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and combining the geometric condition information with image data for cross-modal feature fusion, the image generation process can simultaneously retain the visual appearance information from the image data and the geometric structure information from the target four-dimensional Gaussian scattering scene model in the fused features. This reduces the inaccuracy of correspondence between different modal data due to differences in expression, spatial distribution, and feature scale, reduces the occurrence of local misalignment in structural, edge, or occluded regions of the generated image, and improves the structural stability, visual continuity, and consistency with the spatial structure of the target scene of the generated image.
[0034] The point cloud-guided cross-modal image generation method in this example embodiment will be further explained below.
[0035] In step S110, LiDAR point cloud data and image data corresponding to the target scene are acquired, and geometric supervision information corresponding to the image data is determined based on the LiDAR point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projecting the LiDAR point cloud data.
[0036] In one example embodiment of this disclosure, the target scene refers to the scene object or scene area that needs to be generated across modalities. For example, the target scene may be a road scene, vehicle driving scene, or pedestrian crossing scene in an autonomous driving environment, or an indoor scene, warehouse scene, or outdoor operation scene in which an intelligent robot performs visual perception tasks. It may also be a spatial scene used for virtual scene construction, digital twin modeling, or simulation data generation. This embodiment does not impose any special limitations on the specific type of the target scene, as long as the target scene can be described by LiDAR point cloud data and image data and can be used as an object for cross-modal image generation.
[0037] LiDAR point cloud data refers to the set of points obtained after scanning a target scene using LiDAR (Light Detection and Ranging). Each point may include three-dimensional spatial coordinates, and may further include reflection intensity, scanning time, scan line number, or other point attribute information related to spatial structure. Image data refers to data acquired by an image acquisition device to characterize the visual appearance of a target scene. For example, image data may be a red-green-blue (RGB) image, or a series of consecutive images or image sequences. This embodiment does not impose any special limitations on this.
[0038] In some optional implementations, LiDAR point cloud data and image data can be acquired synchronously by LiDAR and cameras on the same acquisition platform, or they can be read from a pre-established multi-sensor dataset. To ensure that LiDAR point cloud data corresponds to image data, they can be matched based on the acquisition timestamp, sensor calibration information, scene number, or frame sequence number. For example, in autonomous driving scenarios, LiDAR point cloud data and image data that meet a preset threshold at the same time or time difference can be identified as corresponding data. When the target scene contains moving objects, the LiDAR point cloud data can also be temporally corrected based on vehicle pose, sensor pose, or motion compensation information to reduce spatial correspondence deviations caused by the movement of dynamic objects. The above acquisition methods are merely examples; any method that can obtain LiDAR point cloud data and image data corresponding to the same target scene is acceptable, and this embodiment does not impose any special limitations on this.
[0039] When determining the geometric supervision information corresponding to the image data based on LiDAR point cloud data, necessary data processing can be performed on the LiDAR point cloud data first to ensure that the LiDAR point cloud data can stably represent the spatial geometric relationships of the target scene. For example, noise filtering can be performed on the LiDAR point cloud data to remove outliers caused by sensor ranging errors, environmental reflection interference, or invalid echoes; outlier removal can be performed on the LiDAR point cloud data to delete isolated points that are significantly inconsistent with the distribution of neighboring points and cannot effectively represent the structure of the target scene; coordinate system unification can also be performed on the LiDAR point cloud data to transform data from different acquisition coordinate systems to the same spatial coordinate system. Of course, other preprocessing methods can also be used for LiDAR point cloud data depending on the actual acquisition quality, and this embodiment does not impose any special limitations on this.
[0040] After completing the above processing, geometric supervision information corresponding to the image data can be determined based on the spatial correspondence between the LiDAR point cloud data and the image data. Geometric supervision information refers to information used to constrain the scene's geometric structure during subsequent model training. Geometric supervision information can include pixel-level depth supervision information. For example, the LiDAR point cloud data mapped to the image coordinate space can be used to generate sparse or dense depth information according to pixel positions, and this depth information can be used as the geometric supervision information corresponding to the image data. Geometric supervision information can also include point cloud projection constraint information, depth mask information, or spatial structure constraint information, as long as it can be used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. This spatial correspondence can be determined by the calibration relationship between the LiDAR and the image acquisition device. The calibration relationship can include extrinsic and intrinsic parameters. Extrinsic parameters represent the rotation and translation relationships between the LiDAR coordinate system and the camera coordinate system, while intrinsic parameters represent the imaging parameters of the image acquisition device. Based on this calibration relationship, the LiDAR point cloud data can be mapped to the coordinate space corresponding to the image data, and the spatial depth, structural boundaries, or geometric distribution information of the corresponding positions in the image data can be determined based on the mapped point cloud data, thereby forming geometric supervision information.
[0041] In step S120, a preset initial four-dimensional Gaussian scattering scene model is trained using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model.
[0042] In one example embodiment of this disclosure, a four-dimensional Gaussian scattering scene model refers to a model used to characterize the scene structure and appearance information of a target scene in both spatial and temporal dimensions using Gaussian units. A Gaussian unit can be understood as a basic representation unit used to describe a local spatial region of the target scene. Each Gaussian unit may include at least one of a position parameter, a scale parameter, a rotation parameter, an opacity parameter, a color parameter, and a time-dependent parameter. The position parameter can be used to represent the spatial position of the Gaussian unit in the target scene; the scale and rotation parameters can be used to represent the spatial shape of the Gaussian unit; the opacity parameter can be used to represent the contribution of the Gaussian unit in the rendering process; the color parameter can be used to represent the appearance information corresponding to the Gaussian unit; and the time-dependent parameter can be used to represent the changes in the state of objects or the scene in the target scene over time. Of course, different four-dimensional Gaussian scattering scene models can adopt different parameter structures depending on the specific application scenario, and this embodiment does not impose any special limitations on this.
[0043] The preset initial four-dimensional Gaussian scattering scene model refers to the initial model used to characterize the target scene before training begins. This initial four-dimensional Gaussian scattering scene model can be initialized based on the initial point distribution of the target scene, the initial three-dimensional reconstruction results corresponding to the image data, or a preset Gaussian unit distribution method. This embodiment does not impose any special limitations on this. For target scenes containing dynamic objects, the initial four-dimensional Gaussian scattering scene model can also be set with time-related state parameters to characterize the spatial structure and visual appearance of the target scene as it changes over time during subsequent training.
[0044] When training a pre-defined initial four-dimensional Gaussian scattering scene model using image data and geometric supervision information, the image data can be used as the basis for appearance supervision, and the geometric supervision information can be used as the basis for geometric structure supervision. This allows the initial four-dimensional Gaussian scattering scene model to be constrained by both image appearance and spatial structure during training. For example, the image reconstruction loss can be determined based on the difference between the rendered image generated by the initial four-dimensional Gaussian scattering scene model under the current parameter state and the image data; the geometric constraint loss can be determined based on the difference between the spatial depth, structural distribution, or geometric position represented by the initial four-dimensional Gaussian scattering scene model under the current parameter state and the geometric supervision information; and the model parameters of the initial four-dimensional Gaussian scattering scene model can be updated based on the image reconstruction loss and the geometric constraint loss to obtain the target four-dimensional Gaussian scattering scene model.
[0045] During training, auxiliary constraints can be determined based on the distribution of Gaussian units in the initial four-dimensional Gaussian scattering scene model. For example, sparse constraint loss can be determined based on the spatial density, scale, opacity distribution, or quantity distribution of Gaussian units to reduce the impact of invalid stacking, redundant distribution, or unreasonable drift of Gaussian units on the representation of scene geometry. By using image reconstruction loss, geometric constraint loss, and sparse constraint loss together for model training, the optimization process of model parameters can maintain the consistency of image appearance, enhance the spatial structure representation of the target scene, and make the distribution of Gaussian units more consistent with the geometric features of the target scene.
[0046] Through the above training method, the target four-dimensional Gaussian scattering scene model can form a scene representation of the target scene under the combined effect of image data and geometric supervision information. Since the geometric supervision information comes from LiDAR point cloud data, it can reflect the depth relationship, spatial position relationship and structural contour in the target scene. Therefore, introducing geometric supervision information when training the initial four-dimensional Gaussian scattering scene model can reduce the structural deviation caused by the model only relying on the appearance fitting of image data, so that the trained target four-dimensional Gaussian scattering scene model can more stably represent the scene geometry of the target scene.
[0047] In step S130, geometric condition information corresponding to the target scene is generated based on the target four-dimensional Gaussian scattering scene model, and feature extraction, spatial alignment and fusion processing are performed on the geometric condition information and the image data to obtain fused features.
[0048] In one exemplary embodiment of this disclosure, geometric condition information refers to conditional information generated based on a target four-dimensional Gaussian scattering scene model, used to characterize the spatial geometry of the target scene and capable of participating in subsequent image generation. Since the target four-dimensional Gaussian scattering scene model is trained under the joint constraints of image data and geometric supervision information, the geometric condition information generated based on the target four-dimensional Gaussian scattering scene model can carry information related to the spatial structure of the target scene, such as depth levels, structural density, spatial boundaries, or reflection properties. Geometric condition information can exist in the form of geometric condition maps, multi-channel feature maps, conditional tensors, or other data forms; this embodiment does not impose any special limitations on this.
[0049] Cross-modal feature fusion refers to the process of corresponding, combining, and uniformly expressing feature information from different data modalities. In this embodiment, the data involved in cross-modal feature fusion includes geometric condition information and image data. The geometric condition information comes from the target four-dimensional Gaussian scattering scene model and is mainly used to characterize the spatial depth relationship, structural distribution relationship, and geometrically related constraints in the target scene. The image data is mainly used to characterize the color, texture, brightness, edges, and visual appearance information in the target scene. Since geometric condition information and image data belong to the geometric modality and image modality, respectively, they differ in data form, spatial scale, channel meaning, and feature expression method. If they are directly used as inputs for subsequent image generation, it is easy to lead to a lack of stable correspondence between geometric structure information and image appearance information. Therefore, this embodiment uses cross-modal feature fusion to uniformly express the spatial structural constraints contained in the geometric condition information and the visual appearance information contained in the image data, so that the subsequent image generation process can utilize both the geometric structure information and image appearance information of the target scene based on the same feature expression.
[0050] In some optional implementations, cross-modal feature fusion may include processes such as feature extraction, spatial alignment, and fusion processing. For example, geometric features can be extracted from geometric condition information to obtain geometric features, and image features can be extracted from image data to obtain image features. Geometric features can be used to characterize depth levels, structural boundaries, spatial density, and occlusion relationships in the target scene, while image features can be used to characterize color, texture, local edges, and visual semantic information in the target scene. Since there may be spatial positional deviations or scale differences between geometric features and image features, spatial alignment can be further performed to establish a correspondence between them at pixel positions, local regions, or feature scales. After spatial alignment, weighted fusion, stitching fusion, attention fusion, gated fusion, or other fusion methods can be used to process the aligned geometric features and image features to obtain fused features. This embodiment does not impose special limitations on the specific network structure or fusion operator for cross-modal feature fusion.
[0051] Fusion features refer to a unified feature representation obtained through cross-modal feature fusion for subsequent image generation. This fusion feature is neither a single-source image appearance feature nor a single-source geometric structure feature, but rather simultaneously includes spatial geometric constraints provided by geometric condition information and visual appearance information provided by the image data. Specifically, fusion features can include geometric representations constraining object contours, depth levels, occlusion relationships, and spatial positional relationships in the target image, as well as image representations of color, texture, brightness variations, and local visual details in the generated target image. Therefore, fusion features can serve as an intermediate bridge between geometric structure information and image appearance information, enabling subsequent image generation to no longer rely solely on appearance information from the image data, nor be limited to structured generation based solely on geometric condition information.
[0052] Optionally, the fusion features can be represented as feature maps, feature tensors, conditional coding vectors, or multi-scale feature sets. For example, when both geometric condition information and image data are converted into two-dimensional feature maps, the fusion features can be two-dimensional fusion feature maps corresponding to the image data; when the subsequent image generation process requires multi-scale input, the fusion features can also include multi-scale fusion features at different spatial resolutions; when the subsequent image generation model uses conditional coding input, the fusion features can be further encoded into conditional vectors or conditional feature sequences.
[0053] By combining geometric condition information and image data to perform cross-modal feature fusion, fused features are obtained. This allows the geometric condition information generated by the target four-dimensional Gaussian scattering scene model to no longer remain at the level of independent structural constraints, but to establish a correlation with the visual appearance information in the image data at the feature level. Since the fused features simultaneously contain spatial geometric constraints and image appearance expressions, when generating images based on the fused features, the target image can better maintain the object contours, depth levels, occlusion relationships, and spatial positional relationships in the target scene while generating color, texture, and local details. This reduces the impact of inaccurate correspondence of different modal information on the generated image and improves the structural stability and visual continuity of the target image.
[0054] In step S140, an image is generated based on the fusion features to obtain a target image corresponding to the target scene.
[0055] In one example embodiment of this disclosure, the fusion feature can be a feature representation obtained by fusing geometric condition information and image data through cross-modal feature fusion. Since the fusion feature contains both the geometric structure information and image appearance information of the target scene, the fusion feature can be used as input conditions or guiding information for the image generation process to generate a target image corresponding to the target scene.
[0056] The target image can be a scene image used for autonomous driving simulation, a data-enhanced image used for intelligent robot vision training, or an image used for virtual scene construction or other cross-modal image generation tasks. This embodiment does not impose any special limitations on the specific use of the target image.
[0057] In some optional implementations, the fused features can be input into an image generation model, and the target image can be generated by the image generation model. For example, the image generation model can be a conditional generation model, a diffusion generation model, an encoder-decoder generation model, or other models capable of generating images based on input features. A diffusion generation model refers to a model that generates images through a stepwise denoising process; an encoder-decoder generation model refers to a model that generates images by encoding, transforming, and decoding input features. This embodiment does not specifically limit the specific type of image generation model.
[0058] In image generation based on fusion features, the geometric structure information in the fusion features can be used to constrain the object contours, depth levels, occlusion relationships, and spatial positional relationships in the target image, while the image appearance information in the fusion features can guide the color, texture, and visual details in the target image. Since the fusion features are obtained after fusing geometric condition information and image data, the generation process of the target image no longer relies solely on single-modal information, but can maintain constraints on the spatial structure of the target scene while generating the image's appearance content. This approach reduces the probability of inconsistencies between the object contours, depth levels, occlusion relationships, or spatial positional relationships in the generated image and the target scene, improving the structural stability and visual continuity of the target image.
[0059] After the target image is generated, it can be output to a simulation system, training system, or display system, or stored in a dataset as a data sample for subsequent model training or algorithm verification. Since the target image is generated based on fused features, which include geometric information from the target's four-dimensional Gaussian scattering scene model and appearance information corresponding to the image data, the target image can better match the spatial structure of the target scene while maintaining its visual appearance, thereby improving its usability in complex target scenes.
[0060] The contents of steps S110 to S140 will be described in detail below.
[0061] In one example embodiment of this disclosure, the geometric supervision information corresponding to image data can be determined based on lidar point cloud data through the following steps, which may specifically include: Point cloud preprocessing can be performed on LiDAR point cloud data to obtain preprocessed LiDAR point cloud data. Point cloud preprocessing includes at least one of noise filtering, outlier removal, and coordinate system one. Based on the calibration relationship between LiDAR point cloud data and image data, the preprocessed LiDAR point cloud data is mapped to the coordinate space corresponding to the image data. Then, based on the mapped LiDAR point cloud data, pixel-level depth supervision information corresponding to the image data is determined, and the pixel-level depth supervision information is used as geometric supervision information.
[0062] Each point in the spatial point set corresponding to the LiDAR point cloud data can include three-dimensional spatial coordinates, and may further include reflection intensity, scanning time, scan line number, or other point attribute information that can characterize the spatial structure of the target scene. Because LiDAR point cloud data may be affected by factors such as sensor ranging errors, target surface reflection characteristics, ambient lighting, rain and fog obstruction, motion jitter, and multipath reflection during acquisition, the original LiDAR point cloud data may contain noisy points, isolated points, duplicate points, points with abnormal coordinates, or data with inconsistent coordinate systems. Therefore, before determining the geometric supervision information corresponding to the image data based on the LiDAR point cloud data, point cloud preprocessing can be performed on the LiDAR point cloud data to improve the stability of the LiDAR point cloud data in representing the spatial structure of the target scene.
[0063] Noise filtering refers to the process of filtering out abnormal points in LiDAR point cloud data caused by measurement errors, invalid echoes, or environmental interference. Noise filtering can be based on information such as the spatial location of a point, reflection intensity, depth range, neighborhood distribution, or acquisition time. For example, an effective spatial range corresponding to the target scene can be set, and points outside this range can be deleted; points with reflection intensities below a preset threshold or above an abnormal intensity threshold can be identified as noise points based on the reflection intensity of each point in the LiDAR point cloud data; and the distance distribution between a point and its neighbors can be used to determine whether a point belongs to measurement noise. For autonomous driving or robot vision scenarios, noise filtering can reduce abnormal points caused by invalid ranging, glass reflections, strongly reflective objects, or sparse echoes from a distance, thus allowing the retained LiDAR point cloud data to more accurately represent the true spatial distribution of roads, vehicles, pedestrians, buildings, and other environmental objects in the target scene.
[0064] Outlier removal refers to the deletion of points in LiDAR point cloud data that are significantly inconsistent with the spatial distribution of neighboring points and are difficult to represent the true structure of the target scene. Outliers and ordinary noise points overlap to some extent, but outlier identification is more focused on the perspective of local spatial distribution. For example, for a point in the target scene, multiple neighboring points in 3D space can be identified, and the distance distribution, density distribution, or local geometric consistency between the point and its neighbors can be calculated. When the distance between the point and its neighbors is significantly greater than the local average distance, or when the number of neighboring points within a preset range around the point is less than a preset number, the point can be identified as an outlier and removed. By removing outliers, the probability of a small number of isolated points forming erroneous depth supervision information when subsequently mapped to the coordinate space corresponding to the image data can be reduced, avoiding interference from isolated outliers in the representation of object boundaries, depth levels, or spatial relationships in the target scene.
[0065] Coordinate system one refers to transforming the points in the LiDAR point cloud data from the original acquisition coordinate system to a target coordinate system used to establish a correspondence with the image data. Since LiDAR point cloud data is typically acquired in the LiDAR coordinate system, while image data usually corresponds to the camera coordinate system or image coordinate space of the image acquisition device, it is necessary to ensure that the LiDAR point cloud data and image data have a mappable basis in terms of coordinate representation before determining the geometric supervision information corresponding to the image data. Coordinate system one can include transforming points in the LiDAR coordinate system to the camera coordinate system, vehicle coordinate system, world coordinate system, or other unified coordinate system. The transformation process can be based on rotation and translation parameters, or it can be combined with the pose of the acquisition platform, sensor installation relationships, or multi-sensor calibration results.
[0066] Calibration relations refer to a set of parameters used to describe the spatial correspondence between LiDAR point cloud data and image data. Since LiDAR point cloud data is typically located in a three-dimensional LiDAR coordinate system, while image data is typically located in a two-dimensional image coordinate space, they represent data from different sensor modes. Therefore, a spatial mapping between them needs to be established through calibration relations. Calibration relations can include external parameters between the LiDAR and the image acquisition device, as well as internal parameters of the image acquisition device. External parameters can represent the rotation and translation relationships between the LiDAR coordinate system and the camera coordinate system; internal parameters can represent imaging attributes of the image acquisition device such as focal length, principal point position, pixel scale, and distortion parameters. Based on the internal and external parameters, three-dimensional spatial points can be projected onto the image plane. Relevant camera projection models typically use camera intrinsic and extrinsic parameters to map three-dimensional points onto the two-dimensional image plane.
[0067] The preprocessed LiDAR point cloud data can be mapped to the coordinate space corresponding to the image data. This can include transforming the 3D points in the preprocessed LiDAR point cloud data from the LiDAR coordinate system to the camera coordinate system, and projecting the transformed 3D points onto the image coordinate space where the image data resides according to the imaging parameters of the image acquisition device. The coordinate space corresponding to the image data can be a pixel coordinate space, a normalized image plane coordinate space, or a feature coordinate space with the same spatial resolution as the image data. This embodiment does not impose any special limitations on this, as long as the mapped LiDAR point cloud data can correspond to the positions in the image data.
[0068] When an image pixel location corresponds to multiple mapping points, one or more points can be selected for subsequent determination of geometric supervision information based on depth values, visibility rules, or occlusion relationships. For example, a mapping point closer to the camera can be selected as the effective point corresponding to the pixel location to reflect the occlusion relationship between foreground objects and background objects; a depth reference value for the pixel location can also be determined based on the depth statistics of multiple mapping points; or the depth levels corresponding to multiple mapping points can be retained for geometric supervision of complex occluded areas. This embodiment does not impose special limitations on this, as long as the mapping result can be used to determine the geometric supervision information corresponding to the image data.
[0069] Pixel-level depth supervision information refers to depth supervision data corresponding to pixel positions or pixel regions in image data. Since the mapped LiDAR point cloud data is already located in the coordinate space corresponding to the image data, the depth supervision information of the corresponding pixel position in the image data can be determined based on the position of the mapped LiDAR point cloud data in the image coordinate space and the corresponding depth value. This pixel-level depth supervision information can be represented in the form of depth map, sparse depth map, depth label, depth mask, or depth supervision matrix, etc. This embodiment does not impose any special limitations on this, as long as it can be used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model during subsequent training.
[0070] Optionally, pixel-level depth supervision information can also include validity markers to indicate which pixel locations in the image data have reliable depth supervision values. Since LiDAR point cloud data is typically sparse, especially in distant regions, areas with weak reflection, or occluded areas, it may only cover a portion of pixel locations after mapping to the coordinate space corresponding to the image data. Therefore, by setting validity markers, geometric constraints can be calculated using only locations with valid depth supervision values during subsequent training, avoiding interference from invalid or missing depth locations on the training of the initial four-dimensional Gaussian scattering scene model. Validity markers can form pixel-level depth supervision information together with depth supervision values, or they can be used as independent supervision masks in determining subsequent geometric constraint losses.
[0071] In this embodiment, using pixel-level depth supervision information as geometric supervision information means that when training a preset initial four-dimensional Gaussian scattering scene model using image data and geometric supervision information, the pixel-level depth supervision information serves as the supervisory basis for constraining the scene's geometric structure. Pixel-level depth supervision information can transform the three-dimensional spatial depth relationships contained in the LiDAR point cloud data into a supervisory representation corresponding to the image data, enabling the training process to perceive the spatial distance relationships and structural contours of the target scene at the image pixel locations. Since the pixel-level depth supervision information corresponds spatially to the image data, it can work in conjunction with image reconstruction-related supervision on the initial four-dimensional Gaussian scattering scene model, allowing the model to learn image appearance information while being constrained by real spatial depth relationships.
[0072] By preprocessing the LiDAR point cloud data and combining the calibration relationship between the LiDAR point cloud data and image data, the preprocessed LiDAR point cloud data is mapped to the coordinate space corresponding to the image data. Then, pixel-level depth supervision information is determined based on the mapped LiDAR point cloud data. This allows the three-dimensional spatial structure represented by the LiDAR point cloud data to be transformed into geometric supervision information with pixel position correspondence with the image data. This reduces the interference of noise points, outliers, and coordinate inconsistencies on the geometric supervision information, lowers the probability of incorrect geometric constraints caused by spatial mismatch between point cloud data and image data, and improves the accuracy and reliability of geometric supervision information in the subsequent training of the four-dimensional Gaussian scattering scene model.
[0073] In one example embodiment of this disclosure, it can be achieved through Figure 2 The steps described in the document involve training a pre-defined initial four-dimensional Gaussian scattering scene model using image data and geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model. (Refer to...) Figure 2 As shown, it can specifically include: Step S210: Determine the image reconstruction loss of the initial four-dimensional Gaussian scattering scene model based on the image data; Step S220: Determine the geometric constraint loss of the initial four-dimensional Gaussian scattering scene model based on the geometric supervision information; Step S230: Update the model parameters of the initial four-dimensional Gaussian scattering scene model according to the image reconstruction loss and the geometric constraint loss to obtain the target four-dimensional Gaussian scattering scene model.
[0074] Image data refers to image information corresponding to the target scene and used to characterize the visual appearance of the target scene. For example, image data may include color distribution, brightness variations, texture details, object edges, and local appearance structures in the target scene. The initial four-dimensional Gaussian scattering scene model refers to the four-dimensional Gaussian scattering (4DGS) scene model used to characterize the target scene at the beginning of training. It can spatially represent the target scene through multiple Gaussian units and generate rendering results corresponding to the target scene at a given viewing angle or time state.
[0075] Image reconstruction loss refers to the loss information used to measure the difference between the image result generated by the initial four-dimensional Gaussian scattering scene model based on the current model parameters and the image data. It is used to constrain the learning process of the initial four-dimensional Gaussian scattering scene model on the visual appearance of the target scene.
[0076] When determining the image reconstruction loss, the target scene can first be rendered based on the model parameters of the initial four-dimensional Gaussian scattering scene model during the current training phase to obtain a model-rendered image corresponding to the image data. The model-rendered image can be understood as the image representation of the target scene by the initial four-dimensional Gaussian scattering scene model under the current parameter state; it can have the same or corresponding viewpoint, image size, and pixel position relationship as the image data. To ensure that the image reconstruction loss accurately reflects the appearance difference between the model-rendered image and the image data, the model-rendered image and the image data can be standardized in terms of pixel coordinates, image resolution, number of channels, or brightness range. For example, when the resolution of the model-rendered image and the image data is inconsistent, the scale of either the model-rendered image or the image data can be adjusted; when the color channel ranges of the two are inconsistent, pixel value normalization can be performed; when the image data includes multiple frames, the model-rendered image corresponding to each frame of image data can be determined separately, and the image reconstruction loss can be determined based on the correspondence between the frames. This embodiment does not specifically limit the specific method of standardization processing, as long as it enables a comparable relationship between the model-rendered image and the image data.
[0077] Image reconstruction loss can be determined based on the pixel-dimensional differences between the model-rendered image and the image data. For example, the difference between each pixel value in the model-rendered image and the corresponding pixel value in the image data can be calculated, and the image reconstruction loss can be obtained based on the absolute value, squared value, or weighted result of each pixel difference. Pixel values can include red, green, and blue color channel values, as well as brightness values, grayscale values, or other image appearance channel values. When there are image regions of different importance in the target scene, pixel weights can be determined based on the importance of the image regions, the effective pixel region, or the foreground region, and the differences between different pixel positions can be weighted based on the pixel weights to make the image reconstruction loss focus more on the regions related to the expression of the target scene. Of course, the image reconstruction loss is not limited to being determined based on pixel value differences; it can also be determined based on structural differences, perceptual feature differences, or multi-scale image differences between the image data and the model-rendered image. This embodiment does not impose any special limitations on this.
[0078] In some optional implementations, the image reconstruction loss may include at least one of color reconstruction loss, brightness reconstruction loss, structural consistency loss, or perceptual consistency loss. Color reconstruction loss can be used to measure the difference in color channels between the model-rendered image and the image data; brightness reconstruction loss can be used to measure the difference in brightness distribution between the two; structural consistency loss can measure the degree of consistency between the two in local brightness, contrast, and structural information through methods such as Structural Similarity Index Measure (SSIM); perceptual consistency loss can be determined based on the differences between image features extracted by a preset feature extraction network, used to constrain the model-rendered image to approximate the image data in higher-level visual features. These various losses can be used individually or in combination according to preset weights to form the image reconstruction loss. This embodiment does not specifically limit the specific composition of the image reconstruction loss.
[0079] Geometric constraint loss refers to the loss information used to measure the difference between the scene geometry represented by the initial four-dimensional Gaussian scattering scene model under the current model parameters and the geometric supervision information. Unlike image reconstruction loss, which mainly focuses on differences in image appearance, geometric constraint loss is mainly used to constrain the scene geometry represented by the initial four-dimensional Gaussian scattering scene model during training, so that the model avoids optimizing only around appearance information such as color and texture while learning image data.
[0080] When determining the geometric constraint loss, we can first determine the model's geometric representation corresponding to the geometric supervision information based on the target scene represented by the initial four-dimensional Gaussian scattering scene model under the current model parameters. The model's geometric representation can be depth information obtained from model rendering, or it can be the geometric structure information formed by the model at corresponding pixel positions, spatial positions, or viewing angles. For example, when the geometric supervision information is pixel-level depth supervision information, we can compare the model depth map or model depth value generated by the initial four-dimensional Gaussian scattering scene model at the corresponding viewpoint with the pixel-level depth supervision information to determine the geometric constraint loss. When the geometric supervision information is expressed as spatial point constraints or projection constraints, we can determine the geometric constraint loss based on the positional or depth deviation between the geometric representation of the corresponding spatial position of the initial four-dimensional Gaussian scattering scene model and the geometric supervision information.
[0081] Optionally, the geometric constraint loss can be determined based on at least one of depth value difference, depth gradient difference, structural boundary difference, or spatial distribution difference. Depth value difference can be used to constrain the depth results generated by the model to be consistent with the depth supervision values in the geometric supervision information; depth gradient difference can be used to constrain regions in the target scene with significant depth changes, making the model more stable in maintaining object boundaries and spatial hierarchy; structural boundary difference can be used to constrain the geometric boundaries represented by the model to match the structural boundaries reflected in the geometric supervision information; spatial distribution difference can be used to constrain the correspondence between the Gaussian units or local structures represented by the model and the real spatial distribution of the target scene. The above methods for determining the geometric constraint loss can be used individually or in combination, and this embodiment does not impose any special limitations on this.
[0082] In an optional implementation, to enhance the adaptability of geometric constraint loss to different regions of the target scene, the geometric loss weights at different locations can be determined based on the confidence level, depth range, point cloud density, or image region type of the geometric supervision information. For example, a higher geometric loss weight can be set for regions with densely distributed LiDAR point cloud data and reliable depth supervision values; a lower geometric loss weight can be set for distant sparse regions, occluded regions, or regions with uncertain depth supervision; and the geometric loss weight can be appropriately increased for structural boundary regions, object contour regions, or regions with drastic depth changes in the target scene to strengthen the model's constraints on key geometric structures. This embodiment does not impose any special limitations on the specific method for determining the geometric loss weights.
[0083] Model parameters refer to the trainable parameters used to characterize the target scene in the initial four-dimensional Gaussian scattering scene model. For example, model parameters may include at least one of the following: spatial position parameters, scale parameters, rotation parameters, opacity parameters, color parameters, and time-related parameters related to Gaussian units. They may also include other trainable parameters used to control model rendering, dynamic change expression, or feature mapping. This embodiment is not limited to these.
[0084] When updating model parameters based on image reconstruction loss and geometric constraint loss, the combined loss used to train the initial four-dimensional Gaussian scattering scene model can be determined first based on these two losses. The combined loss can be obtained by combining the image reconstruction loss and geometric constraint loss with preset weights, or the weights can be dynamically adjusted based on the training stage, the complexity of the target scene, the reliability of geometric supervision information, or the quality of the image data. For example, in the early stages of training, the weight of the geometric constraint loss can be appropriately increased to enable the model to quickly establish a geometric representation that matches the spatial structure of the target scene; in the later stages of training, the weight of the image reconstruction loss can be appropriately increased to further optimize the image appearance details of the target scene; for regions with sparse geometric supervision information, the influence of the geometric constraint loss can be reduced; for regions with complex structures or obvious occlusion relationships, the influence of the geometric constraint loss can be increased.
[0085] Optionally, the model parameters of the initial four-dimensional Gaussian scattering scene model can be iteratively updated using gradient descent, stochastic gradient descent, adaptive moment estimation optimization, or other parameter optimization methods. Each iteration of training may include: generating a model rendering image and a model geometric representation based on the current model parameters; determining the image reconstruction loss based on the image data; determining the geometric constraint loss based on geometric supervision information; determining the comprehensive loss based on the image reconstruction loss and the geometric constraint loss; calculating the gradient of the model parameters based on the comprehensive loss, and updating the model parameters along the direction that reduces the comprehensive loss. After multiple iterations, when the comprehensive loss meets the preset convergence condition, the number of training iterations reaches the preset number, or the model rendering result meets the preset quality requirements, the updated four-dimensional Gaussian scattering scene model can be determined as the target four-dimensional Gaussian scattering scene model.
[0086] It is understandable that updating model parameters is not simply about making the rendered image pixel-wise closer to the image data, but rather about adapting the model parameters to both the image appearance and spatial geometry of the target scene through the combined effects of image reconstruction loss and geometric constraint loss. Specifically, image reconstruction loss can guide color parameters, opacity parameters, or parameters related to appearance representation to optimize the color, texture, and brightness distribution presented by the image data; geometric constraint loss can guide spatial position parameters, scale parameters, rotation parameters, or parameters related to geometric representation to optimize the depth relationships, spatial distribution, and structural contours reflected by the geometric supervision information. Of course, different model parameters may influence each other during actual training. This embodiment does not limit a certain loss to only acting on a specific parameter, as long as the image reconstruction loss and geometric constraint loss can jointly participate in the model parameter update.
[0087] By determining the image reconstruction loss of the initial four-dimensional Gaussian scattering scene model based on image data and the geometric constraint loss of the initial four-dimensional Gaussian scattering scene model based on geometric supervision information, and then updating the model parameters of the initial four-dimensional Gaussian scattering scene model based on the image reconstruction loss and the geometric constraint loss, the model training process can simultaneously take into account the consistency of image appearance and the constraints of scene geometry. This reduces the problems that are prone to occur when training based solely on image data, such as unstable depth relationships, spatial structure offset, or inaccurate representation of object contours, and improves the comprehensive representation ability of the target four-dimensional Gaussian scattering scene model of the target scene's visual appearance and spatial geometry.
[0088] In one example embodiment of this disclosure, the model parameters of an initial four-dimensional Gaussian scattering scene model can be updated based on image reconstruction loss and geometric constraint loss through the following steps, specifically including: The sparse constraint loss can be determined based on the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model; the model training loss can be determined based on the image reconstruction loss, geometric constraint loss, and sparse constraint loss; and the Gaussian unit parameters in the initial four-dimensional Gaussian scattering scene model can be optimized based on the model training loss so that the optimized four-dimensional Gaussian scattering scene model represents the spatial geometric structure of the target scene.
[0089] In the initial four-dimensional Gaussian scattering scene model, the Gaussian unit refers to the basic model unit used to characterize the local spatial structure and local appearance attributes of the target scene. Each Gaussian unit can correspond to a local spatial region in the target scene and participate in the expression of the target scene through its own spatial position, scale, orientation, opacity, color or other model parameters.
[0090] The distribution state of Gaussian units refers to the overall or local distribution of Gaussian units in the initial four-dimensional Gaussian scattering scene model. For example, the distribution state of Gaussian units can include the spatial density, number distribution, scale distribution, opacity distribution, distance relationship between adjacent Gaussian units, and the degree of aggregation of Gaussian units in different regions of the target scene.
[0091] Sparse constraint loss refers to the loss information used to constrain the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model. Its purpose is to make the spatial distribution of Gaussian units participating in the representation of the target scene more reasonable, reducing redundant stacking, aberrant clustering, or ineffective drift of Gaussian units. Since the initial four-dimensional Gaussian scattering scene model continuously adjusts the Gaussian unit parameters during training based on image reconstruction loss and geometric constraint loss, without constraints on the distribution state of Gaussian units, some Gaussian units may over-cluster in local regions or exhibit redundant distribution in regions of the target scene lacking effective structural contributions, leading to unstable spatial representations in local areas. Therefore, sparse constraint loss can be determined based on the distribution state of Gaussian units, enabling the model training process to suppress unnecessary Gaussian unit distributions while ensuring the expressive power of the target scene.
[0092] In some optional implementations, the sparsity constraint loss can be determined based on the spatial density of Gaussian units. Specifically, the spatial region corresponding to the target scene can be divided into multiple local spatial regions, and the number, density, or distribution intensity of Gaussian units in each local spatial region can be counted. When the number of Gaussian units in a certain local spatial region is too large, or the distance between Gaussian units is too small, resulting in an overly concentrated distribution of Gaussian units, the corresponding sparsity constraint component can be determined based on the Gaussian unit density of that local spatial region. In this way, the situation where the initial four-dimensional Gaussian scattering scene model uses too many Gaussian units to repeatedly express the same spatial structure in local regions can be reduced, making the distribution of Gaussian units more balanced and reducing the interference of redundant Gaussian units on subsequent model parameter optimization.
[0093] Optionally, when determining the sparsity constraint loss, the spatial density, scale distribution, and opacity distribution of Gaussian cells can also be considered comprehensively. For example, the Gaussian cell density in a local region can be determined first based on the spatial location of the Gaussian cells in the target scene, and then the scale and opacity parameters of each Gaussian cell can be combined to determine whether there is redundant representation in the local region. If a local region has a large number of Gaussian cells, a small scale, and some Gaussian cells contribute little opacity, then the local region can be determined to have a larger sparsity constraint loss. If a local region is a structural boundary region or a region with significant depth changes in the target scene, then a relatively high Gaussian cell distribution density can be allowed in the local region to avoid excessive sparsification leading to a loss of structural details in the target scene. The above processing methods can be selected according to the spatial structural complexity of the target scene, and this embodiment does not impose any special limitations on this.
[0094] Model training loss refers to the comprehensive loss used to uniformly guide the optimization of parameters in the initial four-dimensional Gaussian scattering scene model. It can comprehensively reflect the consistency of image appearance, the consistency of scene geometry, and the rationality of Gaussian cell distribution. By incorporating different types of losses into the model training loss, the training process of the initial four-dimensional Gaussian scattering scene model can no longer optimize only a single objective, but rather achieve a balance among multiple constraint objectives.
[0095] In some optional implementations, the image reconstruction loss, geometric constraint loss, and sparse constraint loss can be weighted and fused according to preset weights to obtain the model training loss. The preset weights can represent the degree of contribution of different losses during model training. For example, the weight corresponding to the image reconstruction loss can be used to control the learning intensity of color, texture, and visual appearance information in image data; the weight corresponding to the geometric constraint loss can be used to control the constraint strength of the model on the spatial structural relationships reflected by the geometric supervision information; and the weight corresponding to the sparse constraint loss can be used to control the constraint strength of the model on the rationality of the Gaussian unit distribution. By setting different loss weights, the model training loss can be adapted to different training needs in different application scenarios or different training stages.
[0096] Optionally, the weights corresponding to image reconstruction loss, geometric constraint loss, and sparse constraint loss can be fixed or dynamically adjusted. For example, in the early stages of training, the weight corresponding to geometric constraint loss can be appropriately increased to allow the initial four-dimensional Gaussian scattering scene model to prioritize establishing a model representation that matches the spatial geometry of the target scene. In the later stages of training, the weight corresponding to image reconstruction loss can be appropriately increased to allow the model to further optimize the texture, color, and visual details of the target scene. When excessive clustering or redundant distribution of Gaussian units is detected in local regions during training, the weight corresponding to sparse constraint loss can be appropriately increased to enhance the constraint on the distribution state of Gaussian units. The above weight adjustment methods can be determined based on the loss change trend, training rounds, target scene complexity, or the reliability of geometric supervision information; this embodiment does not impose any special limitations on this.
[0097] The model training loss can be used to simultaneously constrain the image representation, geometric representation, and Gaussian unit distribution of the initial four-dimensional Gaussian scattering scene model. Specifically, the image reconstruction loss can make the model training loss focus on the visual appearance in the image data, enabling the target four-dimensional Gaussian scattering scene model to maintain the color and texture representation of the target scene; the geometric constraint loss can make the model training loss focus on the spatial structure reflected by the geometric supervision information, enabling the target four-dimensional Gaussian scattering scene model to maintain the depth hierarchy, spatial positional relationships, and structural contours of the target scene; the sparsity constraint loss can make the model training loss focus on the distribution state of Gaussian units, enabling the target four-dimensional Gaussian scattering scene model to reduce redundant representations and maintain a more stable spatial modeling structure. The combined effect of these three can avoid the problem of the model training process being biased only towards image appearance, only towards geometric constraints, or neglecting the rationality of the Gaussian unit distribution.
[0098] Gaussian unit parameters refer to the parameters associated with Gaussian units in the initial four-dimensional Gaussian scattering scene model that can be updated through training. Gaussian unit parameters can include at least one of position, scale, rotation, opacity, and color parameters, and may also include parameters characterizing the temporal changes of the Gaussian unit. Position parameters determine the spatial location of the Gaussian unit in the target scene; scale parameters determine the spatial coverage of the Gaussian unit in different directions; rotation parameters determine the spatial pose of the Gaussian unit; opacity parameters determine the contribution of the Gaussian unit to scene rendering and structural representation; and color parameters determine the corresponding appearance of the Gaussian unit. Different Gaussian unit parameters collectively determine how the initial four-dimensional Gaussian scattering scene model represents the spatial geometry and image appearance of the target scene.
[0099] When optimizing Gaussian unit parameters based on model training loss, the optimization direction corresponding to the Gaussian unit parameters can be calculated based on the model training loss, and the Gaussian unit parameters can be iteratively adjusted according to this optimization direction. For example, the spatial position of the Gaussian unit in the target scene can be adjusted according to the influence of the model training loss on the position parameter, so that the Gaussian unit is closer to the real structural region in the target scene; the spatial shape of the Gaussian unit can be adjusted according to the influence of the model training loss on the scale parameter and the rotation parameter, so that the Gaussian unit can better cover the local structure in the target scene; the influence of the model training loss on the opacity parameter can be used to reduce the influence of Gaussian units that contribute weakly to the expression of the target scene, and to enhance Gaussian units that contribute to key structural regions; the image appearance expression corresponding to the Gaussian unit can also be optimized according to the influence of the model training loss on the color parameter. This embodiment does not make special limitations on the specific optimization methods for various Gaussian unit parameters.
[0100] The optimized four-dimensional Gaussian scattering scene model can characterize the spatial geometry of the target scene. This is demonstrated by the matching of the spatial position, scale, and distribution of Gaussian units within the target scene with the object contours, depth levels, spatial boundaries, and occlusion relationships. For example, when the target scene contains objects such as roads, vehicles, pedestrians, or buildings, the optimized four-dimensional Gaussian scattering scene model can express the spatial boundaries of these objects through the position and distribution of Gaussian units, express the local structural morphology through the scale and orientation of Gaussian units, and express the structural contribution of different regions through the opacity and density distribution of Gaussian units. Because geometric constraint loss is involved in determining the model training loss, the optimized four-dimensional Gaussian scattering scene model can better maintain a spatial structure consistent with geometric supervision information; because image reconstruction loss is involved in determining the model training loss, the optimized four-dimensional Gaussian scattering scene model can maintain an appearance consistent with the image data; and because sparse constraint loss is involved in determining the model training loss, the optimized four-dimensional Gaussian scattering scene model can reduce the interference of invalid or redundant Gaussian units on the spatial structure representation.
[0101] By determining the sparse constraint loss based on the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model, and combining the image reconstruction loss, geometric constraint loss, and sparse constraint loss to determine the model training loss, the Gaussian unit parameters in the initial four-dimensional Gaussian scattering scene model are optimized based on the model training loss. This ensures that the parameter updates of Gaussian units are not only constrained by the image appearance and geometric structure, but also by the rationality of the Gaussian unit distribution. This reduces the problems of redundant stacking, abnormal aggregation, or invalid drift of Gaussian units in local regions, and improves the stability of the optimized four-dimensional Gaussian scattering scene model in representing the spatial geometric structure of the target scene.
[0102] In one example embodiment of this disclosure, the generation of geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model can be achieved through the following steps, specifically including: The system can perform conditional rendering on a target four-dimensional Gaussian scattering scene model to obtain a geometric condition map; based on the geometric condition map, multi-channel geometric condition information can be generated; wherein, the geometric condition map can include at least one of a depth condition map, a density condition map, and a reflectivity fusion condition map.
[0103] Conditional rendering, in particular, refers to rendering not merely to generate ordinary color images, but to extract conditional rendering results that reflect the geometric structure of the target scene from the target four-dimensional Gaussian scattering scene model. In other words, conditional rendering focuses more on information related to spatial depth, structural density, or reflection properties in the target four-dimensional Gaussian scattering scene model, so that the rendering results can serve as the geometric condition basis required for subsequent cross-modal feature fusion.
[0104] When performing conditional rendering on a target four-dimensional Gaussian scattering scene model, Gaussian units in the model can be projected and accumulated based on the target viewpoint, the target image plane, or imaging parameters corresponding to the image data. Gaussian units can carry information such as spatial location, scale, orientation, opacity, color, or reflection-related attributes. During conditional rendering, the corresponding Gaussian unit attributes can be selected for rendering based on different rendering objectives. For example, when a depth condition map is needed, rendering can be performed based on the spatial distance or depth value of the Gaussian unit relative to the rendering viewpoint; when a density condition map is needed, rendering can be performed based on the density of Gaussian units in the corresponding spatial region or image location; when a reflectivity fusion condition map is needed, fusion rendering can be performed based on the reflection attributes of the corresponding region of the Gaussian unit or the reflection intensity information related to the LiDAR point cloud data.
[0105] A geometric condition map is an information map rendered from a target four-dimensional Gaussian scattering scene model, capable of visually representing the geometric conditions of the target scene. Geometric condition maps are not limited to traditional visual images but can be used to characterize geometrically related information such as depth relationships, structural density, or reflectance properties of the target scene. Specifically, a geometric condition map can include at least one of a depth condition map, a density condition map, and a reflectance fusion condition map. The depth condition map can represent the spatial distance relationships between different image positions in the target scene; the density condition map can represent the distribution of Gaussian units in different image positions or spatial regions within the target four-dimensional Gaussian scattering scene model; and the reflectance fusion condition map can represent conditional information related to reflectance properties in the target scene, supplementing surface response differences that are difficult to stably express solely by color or texture.
[0106] Multi-channel geometric condition information refers to a geometric condition expression with multiple information channels, formed by combining at least one geometric condition graph. Multi-channel geometric condition information can be used to uniformly represent different types of geometric conditions of a target scene, such as spatial depth conditions, structural density conditions, and reflection attribute conditions. Compared to a single geometric condition graph, multi-channel geometric condition information can describe the geometric structure of the target scene from multiple dimensions. This allows for cross-modal feature fusion by combining geometric condition information and image data, enabling the utilization not only of a single geometric constraint but also the comprehensive utilization of complementary relationships between multiple types of geometric conditions.
[0107] In some optional implementations, if the geometric condition map includes a depth condition map, a density condition map, and a reflectance fusion condition map, then the depth condition map, density condition map, and reflectance fusion condition map can be combined as different channels to generate multi-channel geometric condition information. For example, the depth condition map can be used as the first geometric channel, the density condition map as the second geometric channel, and the reflectance fusion condition map as the third geometric channel, and stacked according to a preset channel order to obtain multi-channel geometric condition information. If only one or two geometric condition maps are generated, multi-channel geometric condition information with the corresponding number of channels can also be constructed based on the generated geometric condition maps; if it needs to be adapted to the subsequent feature extraction structure, channel mapping, channel expansion, or feature encoding processing can also be performed on the existing channels. This embodiment does not impose any special limitations on this.
[0108] Multi-channel geometric condition information can be represented in tensor form, where each channel can correspond to a geometric condition map or an encoded geometric condition feature. This multi-channel geometric condition information can have the same or corresponding spatial dimensions as the image data, or it can match the intermediate feature scale in the subsequent feature extraction process. When the multi-channel geometric condition information has the same spatial dimensions as the image data, it can establish a correspondence with the image data at the pixel level or local region level; when the multi-channel geometric condition information matches the intermediate feature scale, it can directly participate in feature extraction or spatial alignment in subsequent cross-modal feature fusion.
[0109] A depth condition map is a geometric condition map used to express the spatial depth relationships of a target scene. It is obtained by conditionally rendering a four-dimensional Gaussian scattering scene model from the target's perspective, and is used to characterize the spatial distance or depth hierarchy corresponding to different image positions in the target scene. The depth values in the depth condition map can come from the depth contributions of the Gaussian units involved in rendering that image position, or from the weighted result of the depth values of multiple Gaussian units. Through the depth condition map, subsequent processing can obtain the distance relationships, front-back hierarchy, and occlusion basis between objects in the target scene, thus providing structural constraints of the depth dimension for cross-modal feature fusion.
[0110] A density condition map is a geometrical condition map used to represent the structural density or Gaussian cell distribution density of a target scene. It is generated based on the cumulative distribution of Gaussian cells in a target four-dimensional Gaussian scattering scene model across spatial regions or image locations. For example, in areas of the target scene with complex structures or concentrated object boundaries, the contribution of Gaussian cells projected or accumulated to the corresponding image location can be higher, resulting in a more pronounced density response in the density condition map; conversely, in areas with smooth structures or open spaces, the corresponding density response can be lower. The density condition map enables subsequent processing to identify regions in the target scene with strong structural information, significant spatial variations, or where maintaining structural stability is crucial.
[0111] A reflectivity fusion condition map is a geometric condition map used to express information related to the reflectivity attributes of a target scene. It can reflect the differences in reflectivity response exhibited by different objects or regions within the target scene during LiDAR scanning or scene modeling, and can also reflect the reflectivity attribute conditions obtained by fusing a four-dimensional Gaussian scattering scene model of the target. Since different materials, surface states, or spatial objects may have different reflectivity attributes, the reflectivity fusion condition map can provide auxiliary condition information at the surface response level for the target scene, in addition to depth and density condition maps. Through the reflectivity fusion condition map, subsequent processing can obtain supplementary constraints related to the surface attributes of objects in the target scene, thus making the expression of multi-channel geometric condition information more complete.
[0112] The geometric condition map can include any one, two, or three of the following: a depth condition map, a density condition map, and a reflectance fusion condition map. When the geometric condition map includes only a depth condition map, the multi-channel geometric condition information can primarily express the spatial distance relationships of the target scene. When the geometric condition map includes both a depth condition map and a density condition map, the multi-channel geometric condition information can simultaneously express the depth hierarchy and structural distribution of the target scene. When the geometric condition map includes a depth condition map, a density condition map, and a reflectance fusion condition map, the multi-channel geometric condition information can simultaneously express the depth relationships, structural density, and reflectance properties of the target scene. The specific choice of which geometric condition map(s) to use depends on the complexity of the target scene, the requirements of the image generation task, and the condition information provided by the target four-dimensional Gaussian scattering scene model. This embodiment does not impose any special limitations on this.
[0113] By performing conditional rendering on the target four-dimensional Gaussian scattering scene model to obtain a geometric condition map, and generating multi-channel geometric condition information based on the geometric condition map, the spatial geometric structure represented in the target four-dimensional Gaussian scattering scene model can be expressed in at least one of the forms of depth condition map, density condition map, and reflectivity fusion condition map. This transforms the spatial distance relationships, structural distribution state, and reflectivity attribute responses within the model into condition information that can be used for subsequent cross-modal feature fusion, thereby improving the integrity and usability of the geometric structure information transferred to the image generation stage.
[0114] In one example embodiment of this disclosure, it can be achieved through Figure 3 The steps described above enable cross-modal feature fusion by combining geometric condition information and image data to obtain fused features. (Refer to...) Figure 3 As shown, it can specifically include: Step S310: Extract features from the geometric condition information to obtain geometric features; Step S320: Extract features from the image data to obtain image features; Step S330: Based on the spatial cross-attention mechanism, the geometric features and the image features are spatially aligned to obtain aligned geometric features and aligned image features; Step S340: Obtain the weather state information corresponding to the target scene, and determine the weather gating parameter based on the weather state information. The weather gating parameter is used to adjust the fusion ratio of the aligned geometric features and the aligned image features under different weather conditions. Step S350: Based on the weather gating parameters, the aligned geometric features and the aligned image features are fused to obtain the fused features.
[0115] Geometric features refer to the feature representations extracted from geometric condition information. They are used to convert the more direct conditional values, channel responses, or image-based geometric information in the geometric condition information into feature representations suitable for subsequent spatial alignment and fusion processing. In other words, geometric features are not simply equivalent to the original geometric condition information, but rather a structured representation formed after feature encoding of the geometric condition information. It can more comprehensively reflect the depth changes, structural boundaries, spatial density changes, and local geometric relationships in the target scene.
[0116] When extracting features from geometric condition information, the corresponding feature extraction method can be selected according to the data format of the geometric condition information. For example, when the geometric condition information is a multi-channel geometric condition map, the multi-channel geometric condition map can be used as input, and the geometric conditions in different channels can be jointly encoded through a feature extraction network to obtain geometric features; when the geometric condition information is a geometric condition tensor, channel mapping, scale transformation, and spatial feature extraction can be performed on it through a tensor encoding structure to obtain geometric features; when the geometric condition information includes geometric condition expressions at multiple scales, features can be extracted from the geometric condition information at different scales separately, and the features at different scales can be aggregated to obtain geometric features used to characterize the spatial structure of the target scene. This embodiment does not impose any special limitations on the specific input format and feature extraction structure of the geometric condition information.
[0117] In some optional implementations, feature extraction of the geometric condition information can be achieved using a Convolutional Neural Network (CNN). Specifically, receptive fields of local spatial regions in the geometric condition information can be extracted using convolutional layers, and the stability of feature representation can be enhanced using normalization layers or activation layers. Furthermore, feature information from shallow geometric responses to deep geometric structures can be extracted progressively through multi-layer convolutional structures. For example, shallow convolutional features can be used to preserve depth edges, local density abrupt changes, and reflection attribute variations in the target scene; deep convolutional features can be used to express a larger range of spatial structural layouts, occlusion relationships, and regional distributions in the target scene. For scenes requiring the preservation of fine-grained spatial positional correspondences, convolutional structures with smaller strides or those preserving spatial resolution can be used to avoid losing structural boundary information during geometric feature extraction. For scenes requiring enhanced global geometric understanding, different levels of geometric features can also be extracted using multi-scale convolutions or feature pyramid structures.
[0118] In some optional implementations, feature extraction of geometric condition information can also be achieved using an attention-based coding structure. For example, the geometric condition information can be divided into multiple spatial regions or feature blocks, and the geometric relationships between different regions can be modeled through a self-attention mechanism. This allows the obtained geometric features to not only contain local depth, density, or reflection attribute information, but also the spatial relationships between different regions. For cases where the target scene contains complex structures such as vehicles, pedestrians, road boundaries, and building outlines, the attention-based coding structure can establish relationships between structural regions over a larger scope, thereby enhancing the expressive power of geometric features for the spatial layout of the target scene. Of course, the geometric feature extraction method can also combine convolutional structures with attention structures; this embodiment does not impose any special limitations on this approach.
[0119] Image features refer to the feature representation obtained after feature encoding of image data. They are used to convert pixel-level appearance information in image data into feature representations suitable for subsequent spatial alignment and fusion processing. Unlike geometric features, which mainly focus on the spatial structure of the target scene, image features mainly focus on the visual appearance representation of the target scene, such as road textures, vehicle colors, pedestrian appearances, building facade textures, and visual details of other environmental objects.
[0120] Optionally, when extracting features from image data, the image data can be input into a preset image feature extraction network, and image features can be extracted through the image feature extraction network. The image feature extraction network may include a convolutional neural network, an encoder network, a feature pyramid network, or other network structures capable of extracting image appearance features. Among them, a convolutional neural network can be used to extract local texture, edges, and color distribution in image data; an encoder network can be used to convert image data into a higher-dimensional feature representation; a feature pyramid network can be used to extract image features from multiple scales, so that the image features simultaneously contain local details and global structural information. This embodiment does not impose any special limitations on the specific structure of the image feature extraction network.
[0121] Spatial cross-attention is an attention processing mechanism used to establish spatial correspondences between features of different modalities. Geometric features originate from geometric condition information and mainly represent the spatial structure of the target scene; image features originate from image data and mainly represent the visual appearance of the target scene. Since geometric features and image features belong to different modalities, they may differ in their representation, spatial distribution, feature scale, and channel semantics. Directly fusing them can easily lead to inaccurate feature correspondences in structural regions, edge regions, or occluded regions. Therefore, spatial cross-attention can be used to spatially align geometric features and image features, enabling a more accurate correspondence between them in spatial location or local regions.
[0122] For example, the spatial cross-attention mechanism can use geometric features as query features and image features as key and value features. By calculating the correlation between query features and key features, it determines the spatial correspondence between geometric features and image features, and then performs weighted aggregation of image features based on this correspondence to obtain an image feature representation that matches the geometric structure. Conversely, it can also use image features as query features and geometric features as key and value features, performing reverse attention on geometric features through image features to obtain a geometric feature representation that matches the image's appearance region.
[0123] Spatial alignment can be performed on geometric features and image features, meaning that spatial locations, structural boundaries, local regions, or feature responses in geometric features and image features can be aligned. Specifically, when a geometric feature represents a significant abrupt change in depth or a structural boundary at a certain spatial location, the corresponding edge, texture change, or object boundary region in the image feature can be focused on using a spatial cross-attention mechanism; when an image feature exhibits strong appearance texture or object semantic response in a certain region, the corresponding depth level, density distribution, or reflectance attribute region in the geometric feature can be focused on using a spatial cross-attention mechanism.
[0124] After obtaining the aligned geometric features and aligned image features, weather state information corresponding to the target scene can be acquired, and weather gating parameters can be determined based on the weather state information. Weather state information refers to information characterizing the weather environment state at the time of target scene acquisition or target image generation. For example, weather state information may include at least one of the following environmental states: sunny, cloudy, rainy, foggy, snowy, dusty, low illumination, or strong backlight. It may also include quantitative information related to the weather environment, such as rainfall intensity, fog concentration, visibility, ambient brightness, humidity, road surface reflectivity, or weather category confidence level. This embodiment does not impose any special limitations on this, as long as it can reflect the impact of the weather environment on the reliability of image data and geometric condition information.
[0125] Weather gating parameters are parameters determined based on weather conditions and used to adjust the contribution ratios of aligned geometric features and aligned image features during the fusion process. Different weather conditions have varying impacts on image data and geometric information. For example, under clear or normal lighting conditions, image data typically provides clear texture, color, and edge information; under rainy, foggy, snowy, or low-light conditions, image data may exhibit blurred textures, reduced contrast, weakened edges, or localized occlusion. Meanwhile, the spatial depth relationships and structural distributions contained in the geometric information provide more stable structural constraints for the target scene. Therefore, weather gating parameters can be determined based on weather conditions, enabling the cross-modal fusion process under different weather conditions to dynamically adjust the contribution levels of geometric and image features.
[0126] Weather gating parameters can be determined based on weather condition information and a preset gating mapping relationship. For example, the preset gating mapping relationship can include the correspondence between weather categories and geometric feature weights and image feature weights. In sunny or well-lit conditions, higher image feature weights can be determined to allow the fused features to retain more texture, color, and visual details in the image data. In foggy, rainy, snowy, or low-light conditions, higher geometric feature weights can be determined to allow the fused features to inherit more depth, spatial structure, and object contour constraints from the geometric condition information.
[0127] By determining weather gating parameters based on weather conditions and performing gated fusion processing on aligned geometric and image features based on these parameters, the cross-modal feature fusion process can move beyond relying solely on fixed fusion methods or single structural weights. Instead, it can dynamically adjust the contribution ratio of geometric structural information and image appearance information according to the weather environment of the target scene. Thus, when image data suffers from texture degradation, edge blurring, or visual noise due to weather conditions such as rain, fog, snow, or low illumination, the fused features can better utilize the spatial structural constraints provided by geometric conditions. Conversely, when the image data quality is high, the fused features can fully preserve the texture and color details in the image data. This reduces structural misalignment or visual discontinuities caused by instability in cross-modal fusion under different weather conditions, improving the adaptability of the fused features to complex environments.
[0128] Aligned geometric features and aligned image features have established a spatial correspondence through a spatial cross-attention mechanism, but they still represent information content from different modalities. Specifically, aligned geometric features mainly include the depth hierarchy, structural density, boundary relationships, and spatial positional relationships of the target scene; aligned image features mainly include the color, texture, visual edges, and appearance semantic information of the target scene. Fusion processing refers to jointly encoding or combining the aligned geometric features and aligned image features based on weather gating parameters to obtain fused features that simultaneously contain geometric structure information and image appearance information.
[0129] Optionally, the fusion process may include at least one of feature concatenation, weighted fusion, attention fusion, gating fusion, or convolutional fusion. For example, aligned geometric features and aligned image features can be concatenated along the channel dimension based on weather gating parameters, and the concatenated features can be compressed and integrated through convolutional layers to obtain fused features; alternatively, different fusion weights can be assigned to the aligned geometric features and aligned image features according to the structural or appearance importance of different spatial regions, and weighted fusion can be performed according to the fusion weights; alternatively, the contribution level of geometric features and image features in different regions can be adaptively selected through a gating structure. This embodiment does not specifically limit the specific method of fusion processing.
[0130] By extracting features from geometric condition information and image data respectively, geometric features and image features are obtained. Based on the spatial cross-attention mechanism, the geometric features and image features are spatially aligned. Then, the aligned geometric features and aligned image features are fused. This allows the geometric structure information and image appearance information to be fused after establishing a spatial correspondence. This reduces the impact of differences in the expression form, spatial distribution and feature scale of different modal data on the fusion process, reduces the probability of feature mismatch in structural regions, edge regions or occluded regions, and improves the consistency between geometric constraints and visual appearance in the fused features.
[0131] In one example embodiment of this disclosure, the fusion of aligned geometric features and aligned image features to obtain fused features can be achieved through the following steps, specifically including: Based on the density information in the geometric condition information, the structural weight information corresponding to different regions in the target scene can be determined; based on the structural weight information and weather gating parameters, the geometric fusion weight and image fusion weight corresponding to different regions in the target scene can be determined; based on the geometric fusion weight and image fusion weight, the aligned geometric features and aligned image features are adaptively weighted and fused to obtain the initial fused features; the initial fused features are subjected to residual refinement processing to obtain the fused features.
[0132] Density information within the geometric condition information refers to information generated by the target four-dimensional Gaussian scattering scene model, used to characterize the strength of structural distribution in different spatial or image regions of the target scene. This density information can originate from the density condition map in the geometric condition information, or from channel information related to Gaussian cell distribution, spatial structural strength, or local geometric complexity in the multi-channel geometric condition information. Since the target four-dimensional Gaussian scattering scene model represents the target scene through Gaussian cells, the number, cumulative contribution, opacity response, or spatial coverage intensity of Gaussian cells in different regions can reflect the structural density of that region; correspondingly, density information can be used to characterize which regions in the target scene have strong structural relationships and which regions have relatively gentle structural changes.
[0133] Different regions in the target scene can be different pixel regions in the image coordinate space, different feature regions in the feature space, or local spatial regions divided based on geometric condition information. For example, the image plane corresponding to the geometric condition information can be divided into multiple grid regions, and the density information in each grid region can be statistically analyzed; the density response can also be determined directly based on pixel position or feature point position; adaptive region division can also be performed based on structural boundaries, occluded regions, or depth variation regions in the target scene, so that the structural weight information can more accurately reflect the importance of the geometric structure of different regions in the target scene.
[0134] Structural weight information refers to information used to indicate the degree to which different regions in the target scene depend on geometric constraints during fusion processing. Regions with high density information typically indicate a concentrated distribution of Gaussian units, significant structural changes, or complex spatial boundaries, such as vehicle edges, pedestrian outlines, building boundaries, road edges, occlusion boundaries, or regions with abrupt depth changes. In these regions, if the fusion process relies excessively on image appearance features, structural misalignment may occur due to texture interference, edge blurring, or modal correspondence deviations. Therefore, higher structural weights can be assigned to these regions to enhance the role of aligned geometric features in the fusion process. Conversely, regions with low density information typically indicate gentler structural changes or weaker geometric constraints, such as road surfaces, sky areas, or large flat areas of walls. In these regions, structural weights can be appropriately reduced to allow the aligned image features to retain more color, texture, and visual continuity.
[0135] Weather gating parameters are gate parameters determined based on the weather state information corresponding to the target scene. They are used to characterize the reliability differences between geometric features and image features during the fusion process under different weather conditions. Weather state information can include weather environmental conditions such as sunny, cloudy, rainy, foggy, snowy, dusty, low light, or strong backlight, as well as quantitative information such as rainfall intensity, fog concentration, visibility, ambient brightness, humidity, road surface reflectivity, or weather category confidence. Weather state information can be obtained through weather recognition from image data, acquired through environmental sensors on the acquisition platform, or read from external meteorological data, acquisition logs, scene metadata, or dataset annotation information.
[0136] Weather gating parameters can be represented as scalar gating parameters, region gating maps, channel gating vectors, or multi-scale gating tensors. When weather state information is used to adjust the overall fusion ratio, a global weather gating parameter can be determined. When different areas in the target scene are affected by weather to varying degrees—for example, there is fog above the image, rain reflections on the road surface, or snow occlusion in local areas—a region gating map with corresponding spatial dimensions to the aligned geometric and image features can be generated. When different feature channels have different sensitivities to weather changes, channel gating vectors corresponding to different feature channels can also be generated. All of these different forms of weather gating parameters can be used to adjust the contribution ratio of aligned geometric and image features in the fusion process.
[0137] In some optional implementations, weather gating parameters can be determined based on weather state information and a preset gating mapping relationship. The preset gating mapping relationship can include the correspondence between weather categories and geometric feature weights and image feature weights. For example, in sunny or well-lit conditions, higher image feature weights can be determined to ensure that the fused features retain more texture, color, and visual details in the image data; in foggy, rainy, snowy, or low-light conditions, higher geometric feature weights can be determined to ensure that the fused features inherit more depth, spatial structure, and object contour constraints from the geometric condition information. Optionally, weather state information can also be input into a gating parameter generation network, which outputs weather gating parameters. The gating parameter generation network can include at least one of fully connected layers, convolutional layers, normalization layers, or activation functions; this embodiment does not impose any special limitations on this.
[0138] The geometric fusion weights and image fusion weights for different regions in a target scene can be determined based on structural weight information and weather gating parameters. This means that structural weight information, reflecting the importance of regional structures, and weather gating parameters, reflecting differences in modal reliability under weather conditions, are used together to determine the fusion ratio. Geometric fusion weights represent the contribution of aligned geometric features to the fusion process, while image fusion weights represent the contribution of aligned image features. Structural weight information can be used to increase geometric fusion weights in structurally complex regions, while weather gating parameters can be used to further increase geometric fusion weights under degraded weather conditions or to increase image fusion weights under high-quality weather conditions.
[0139] For example, for areas with high structure density and significant weather degradation, such as vehicle edge areas in foggy weather, road boundary areas in rainy weather, or pedestrian outline areas in snowy weather, a higher geometric fusion weight can be determined by simultaneously using higher structure weight information and weather gating parameters that favor geometric features. This allows the initial fused features to inherit more of the depth, structural boundaries, and spatial relationships of the aligned geometric features in this area. Conversely, for areas with low structure density and good image quality, such as road texture areas, building surface areas, or sky areas in sunny weather, a higher image fusion weight can be determined by using lower structure weight information and weather gating parameters that favor image features. This allows the initial fused features to retain more of the texture, color, and visual details of the aligned image features in this area.
[0140] In some optional implementations, structural weight information and weather gating parameters can be input into a weight mapping module to output geometric fusion weights and image fusion weights. The weight mapping module may include convolutional layers, normalization layers, activation functions, gating units, or other network structures capable of converting structural weight information and weather gating parameters into fusion weights. Optionally, the geometric fusion weights and image fusion weights can satisfy a complementary relationship, for example, their sum being a preset value; alternatively, they can be determined independently based on the structural weight information and weather gating parameters to allow certain regions to simultaneously retain high geometric structural constraints and high image appearance contributions. In this way, the fusion weights can vary not only with changes in the structural state of the target scene region but also with the impact of the target scene's weather conditions on the reliability of image data and geometric condition information.
[0141] When adaptively weighting and fusing aligned geometric features and aligned image features based on geometric fusion weights and image fusion weights, the geometric fusion weights and image fusion weights can first be adjusted to match the spatial scale and channel format of the aligned geometric features and aligned image features. For example, when the spatial resolution of the geometric fusion weights or image fusion weights is inconsistent with the aligned geometric features and aligned image features, scale matching can be performed through upsampling, downsampling, interpolation, or pooling; when the number of channels of the geometric fusion weights or image fusion weights is inconsistent with the number of feature channels, channel adaptation can be performed through channel duplication, linear mapping, or convolutional mapping. Through the above processing, the geometric fusion weights and image fusion weights can be applied to the aligned geometric features and aligned image features pixel-wise, region-wise, or channel-wise.
[0142] The aligned geometric features can be weighted according to the geometric fusion weight, and the aligned image features can be weighted according to the image fusion weight. Then, the weighted geometric features and the weighted image features can be summed, stitched together and mapped, or otherwise combined to obtain the initial fused features.
[0143] In some optional implementations, adaptive weighted fusion can further adjust the fusion ratio by incorporating local region feature responses. For example, even if a region has high density information, if the aligned image features of that region exhibit clear and stable edge responses, and the weather gating parameter indicates high reliability of the image data, a certain image fusion weight can be retained while maintaining a high geometric fusion weight to balance structural stability and visual detail. If a region has low density information, but the weather gating parameter indicates that the current weather causes significant noise, fogging, raindrop occlusion, or unstable texture in the image features, the geometric fusion weight can be appropriately increased to reduce the interference of abnormal appearance features on the initial fused features. This embodiment does not impose special limitations on the weight adjustment strategy in adaptive weighted fusion, as long as the aligned geometric features and aligned image features can be fused according to the structural weight information and the weather gating parameter.
[0144] The initial fusion feature refers to the intermediate fusion result obtained by adaptively weighting and fusing aligned geometric features and aligned image features according to geometric fusion weights and image fusion weights. The initial fusion feature can simultaneously contain both geometric structure information and image appearance information of the target scene, and the contribution of these two types of information in different regions can be dynamically adjusted based on structural weight information and weather gating parameters. For example, in densely structured areas of the target scene or areas where the image is significantly affected by weather degradation, the initial fusion feature can retain more geometric structure constraints to reduce object boundary misalignment, confused depth relationships, or unreasonable occlusion logic in subsequent image generation. In textured areas of the target scene or areas with good image quality, the initial fusion feature can retain more image appearance information to improve the visual naturalness and continuity of the subsequent target image.
[0145] The initial fusion feature is an intermediate feature obtained by adaptively weighting and fusing the aligned geometric features and the aligned image features based on structural weight information. Although the initial fusion feature already contains both geometric structure information and image appearance information, since the aligned geometric features and the aligned image features originate from different modalities, there may still be subtle differences between them in local edges, occluded regions, densely structured regions, or complex texture regions. Therefore, the initial fusion feature may have problems such as edge discontinuities, inconsistent local responses, feature conflicts, or insufficient detail representation. Residual refinement refers to further modifying and enhancing the initial fusion feature by learning or calculating compensating features, while retaining the main information of the initial fusion feature, to obtain a more stable fusion feature.
[0146] For example, residual refinement can be achieved through a residual network structure. Specifically, the initial fusion features can be input into the residual refinement module, which generates residual correction features based on the initial fusion features. The residual correction features are then added to the initial fusion features or subjected to other residual combination processing to obtain the fusion features. The residual correction features can be used to compensate for local structural errors, edge discontinuities, or texture loss that still exist after adaptive weighted fusion. The residual structure can refine the parts that need correction without completely reconstructing the initial fusion features, thus helping to preserve the geometric constraints and image appearance information already formed in the initial fusion features.
[0147] In some optional implementations, residual refinement may include at least one of local convolutional refinement, multi-scale refinement, or boundary region refinement. Local convolutional refinement integrates neighborhood information in the initial fused features through convolution operations to reduce local noise and abrupt changes in local features; multi-scale refinement processes the initial fused features at different scales to simultaneously improve local details and global structural consistency; boundary region refinement enhances structural boundaries, depth abrupt change regions, or occlusion boundary regions to reduce edge breaks or structural misalignments. The above residual refinement methods can be used individually or in combination, and this embodiment does not impose any special limitations on this.
[0148] The fusion feature is the final fusion representation obtained after residual refinement, used for subsequent image generation based on the fusion feature. Compared to the initial fusion feature, the fusion feature can have better local continuity, more stable edge representation, and a more harmonious relationship between geometric and appearance information. Since the fusion feature is obtained by adaptively weighting and refining aligned geometric features and aligned image features under the combined effect of structural weight information and weather gating parameters, it can more reasonably balance geometric constraints and image appearance representation in different structural regions and under different weather conditions, thereby reducing local misalignment, feature conflict, or visual discontinuity caused by weather degradation that may occur during cross-modal fusion.
[0149] By determining the structural weight information corresponding to different regions in the target scene based on the density information in the geometric condition information, and then performing adaptive weighted fusion of aligned geometric features and aligned image features based on the structural weight information, the resulting initial fused features are further refined using residual processing. This allows the fusion process to dynamically adjust the contribution ratio of geometric features and image features according to the structural density of different regions, thereby reducing the problems of insufficient geometric constraints in structurally complex regions and limited appearance details in textured regions caused by fixed fusion methods. It also reduces the impact of edge discontinuities, local noise, or feature conflicts in the initial fused features on subsequent image generation, and improves the stability of the fused features in structural regions, edge regions, and occluded regions.
[0150] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0151] Furthermore, in this example embodiment, a point cloud-guided cross-modal image generation apparatus is also provided. (Refer to...) Figure 4 As shown, the point cloud-guided cross-modal image generation device 400 includes: an information determination module 410, a model training module 420, a feature fusion module 430, and an image generation module 440. Wherein: The information determination module 410 is used to acquire lidar point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the lidar point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projection based on the lidar point cloud data. The model training module 420 is used to train a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. The feature fusion module 430 is used to generate geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and to perform feature extraction, spatial alignment and fusion processing on the geometric condition information and the image data to obtain fused features; The image generation module 440 is used to generate an image based on the fusion features to obtain a target image corresponding to the target scene.
[0152] In some example embodiments of this disclosure, based on the foregoing scheme, determining the geometric supervision information corresponding to the image data according to the lidar point cloud data includes: The lidar point cloud data is preprocessed to obtain preprocessed lidar point cloud data. The point cloud preprocessing includes at least one of noise filtering, outlier removal, and coordinate system one. Based on the calibration relationship between the lidar point cloud data and the image data, the preprocessed lidar point cloud data is mapped to the coordinate space corresponding to the image data; Based on the mapped lidar point cloud data, pixel-level depth supervision information corresponding to the image data is determined, and the pixel-level depth supervision information is used as the geometric supervision information.
[0153] In some exemplary embodiments of this disclosure, based on the foregoing scheme, the step of training a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model includes: Determine the image reconstruction loss of the initial four-dimensional Gaussian scattering scene model based on the image data; The geometric constraint loss of the initial four-dimensional Gaussian scattering scene model is determined based on the geometric supervision information. Based on the image reconstruction loss and the geometric constraint loss, the model parameters of the initial four-dimensional Gaussian scattering scene model are updated to obtain the target four-dimensional Gaussian scattering scene model.
[0154] In some exemplary embodiments of this disclosure, based on the foregoing scheme, updating the model parameters of the initial four-dimensional Gaussian scattering scene model according to the image reconstruction loss and the geometric constraint loss includes: Based on the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model, determine the sparse constraint loss; The model training loss is determined based on the image reconstruction loss, the geometric constraint loss, and the sparse constraint loss. The Gaussian unit parameters in the initial four-dimensional Gaussian scattering scene model are optimized based on the model training loss, so that the optimized four-dimensional Gaussian scattering scene model can represent the spatial geometry of the target scene.
[0155] In some example embodiments of this disclosure, based on the foregoing scheme, generating geometric condition information corresponding to the target scene according to the target four-dimensional Gaussian scattering scene model includes: Conditional rendering is performed on the target four-dimensional Gaussian scattering scene model to obtain a geometric condition map; Based on the geometric condition diagram, generate multi-channel geometric condition information; The geometric condition map includes at least one of a depth condition map, a density condition map, and a reflectivity fusion condition map.
[0156] In some exemplary embodiments of this disclosure, based on the foregoing scheme, the step of performing feature extraction, spatial alignment, and fusion processing on the geometric condition information and the image data to obtain fused features includes: The geometric condition information is used to extract features to obtain geometric features; Feature extraction is performed on the image data to obtain image features; Based on the spatial cross-attention mechanism, the geometric features and the image features are spatially aligned to obtain aligned geometric features and aligned image features; Obtain the weather state information corresponding to the target scene, and determine the weather gating parameter based on the weather state information. The weather gating parameter is used to adjust the fusion ratio of the aligned geometric features and the aligned image features under different weather conditions. Based on the weather gating parameters, the aligned geometric features and the aligned image features are fused to obtain the fused features.
[0157] In some example embodiments of this disclosure, based on the foregoing scheme, the step of fusing the aligned geometric features and the aligned image features based on the weather gating parameters to obtain the fused features includes: Based on the density information in the geometric condition information, determine the structural weight information corresponding to different regions in the target scene; Based on the structural weight information and the weather gating parameters, determine the geometric fusion weight and image fusion weight corresponding to different regions in the target scene; Based on the geometric fusion weights and the image fusion weights, the aligned geometric features and the aligned image features are adaptively weighted and fused to obtain initial fused features; The initial fusion features are subjected to residual refinement to obtain the fusion features.
[0158] The specific details of each module of the point cloud-guided cross-modal image generation device have been described in detail in the corresponding point cloud-guided cross-modal image generation method, so they will not be repeated here.
[0159] It should be noted that although several modules or units of the point cloud-guided cross-modal image generation apparatus have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0160] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described point cloud-guided cross-modal image generation method is also provided.
[0161] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be embodied in the following forms: a completely hardware embodiment, a completely software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0162] The following reference Figure 5 To describe an electronic device 500 according to such an embodiment of the present disclosure. Figure 5 The electronic device 500 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0163] like Figure 5 As shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include, but are not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including storage unit 520 and processing unit 510), and a display unit 540.
[0164] The storage unit stores program code that can be executed by the processing unit 510, causing the processing unit 510 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 510 can perform actions such as... Figure 1Step S110 involves acquiring lidar point cloud data and image data corresponding to the target scene, and determining geometric supervision information corresponding to the image data based on the lidar point cloud data; Step S120 involves training a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model, wherein the geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model; Step S130 involves generating geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and performing cross-modal feature fusion based on the geometric condition information and the image data to obtain fused features; Step S140 involves generating an image based on the fused features to obtain a target image corresponding to the target scene.
[0165] Storage unit 520 may include readable media in the form of volatile storage units, such as random access memory (RAM) 521 and / or cache memory (Cache) 522, and may further include read-only memory (ROM) 523.
[0166] Storage unit 520 may also include a program / utility 524 having a set (at least one) program module 525, such program module 525 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0167] Bus 530 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0168] Electronic device 500 can also communicate with one or more external devices 570 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 500, and / or with any device that enables electronic device 500 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 550. Furthermore, electronic device 500 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 560. As shown, network adapter 560 communicates with other modules of electronic device 500 via bus 530. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0169] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0170] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0171] refer to Figure 6 As shown, a program product 600 for implementing the above-described point cloud-guided cross-modal image generation method according to embodiments of the present disclosure is described. It may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0172] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0173] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0174] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0175] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0176] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0177] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0178] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This specification is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. It is understood that the specification and embodiments are to be considered exemplary only and are intended to illustrate the true scope and technical concepts of this disclosure.
[0179] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A point cloud-guided cross-modal image generation method, characterized in that, include: Acquire lidar point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the lidar point cloud data. The geometric supervision information includes pixel-level depth supervision information obtained by projection based on the lidar point cloud data. The preset initial four-dimensional Gaussian scattering scene model is trained using the image data and the geometric supervision information to obtain the target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. Based on the target four-dimensional Gaussian scattering scene model, geometric condition information corresponding to the target scene is generated, and feature extraction, spatial alignment and fusion processing are performed on the geometric condition information and the image data to obtain fused features; Image generation is performed based on the fusion features to obtain a target image corresponding to the target scene.
2. The method according to claim 1, characterized in that, The step of determining the geometric supervision information corresponding to the image data based on the lidar point cloud data includes: The lidar point cloud data is preprocessed to obtain preprocessed lidar point cloud data. The point cloud preprocessing includes at least one of noise filtering, outlier removal, and coordinate system one. Based on the calibration relationship between the lidar point cloud data and the image data, the preprocessed lidar point cloud data is mapped to the coordinate space corresponding to the image data; Based on the mapped lidar point cloud data, pixel-level depth supervision information corresponding to the image data is determined, and the pixel-level depth supervision information is used as the geometric supervision information.
3. The method according to claim 1, characterized in that, The step of training a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model includes: Determine the image reconstruction loss of the initial four-dimensional Gaussian scattering scene model based on the image data; The geometric constraint loss of the initial four-dimensional Gaussian scattering scene model is determined based on the geometric supervision information. Based on the image reconstruction loss and the geometric constraint loss, the model parameters of the initial four-dimensional Gaussian scattering scene model are updated to obtain the target four-dimensional Gaussian scattering scene model.
4. The method according to claim 3, characterized in that, The step of updating the model parameters of the initial four-dimensional Gaussian scattering scene model based on the image reconstruction loss and the geometric constraint loss includes: Based on the distribution state of Gaussian units in the initial four-dimensional Gaussian scattering scene model, determine the sparse constraint loss; The model training loss is determined based on the image reconstruction loss, the geometric constraint loss, and the sparse constraint loss. The Gaussian unit parameters in the initial four-dimensional Gaussian scattering scene model are optimized based on the model training loss, so that the optimized four-dimensional Gaussian scattering scene model can represent the spatial geometry of the target scene.
5. The method according to claim 1, characterized in that, The step of generating geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model includes: Conditional rendering is performed on the target four-dimensional Gaussian scattering scene model to obtain a geometric condition map; Based on the geometric condition diagram, generate multi-channel geometric condition information; The geometric condition map includes at least one of a depth condition map, a density condition map, and a reflectivity fusion condition map.
6. The method according to claim 1, characterized in that, The step of performing feature extraction, spatial alignment, and fusion processing on the geometric condition information and the image data to obtain fused features includes: The geometric condition information is used to extract features to obtain geometric features; Feature extraction is performed on the image data to obtain image features; Based on the spatial cross-attention mechanism, the geometric features and the image features are spatially aligned to obtain aligned geometric features and aligned image features; Obtain the weather state information corresponding to the target scene, and determine the weather gating parameter based on the weather state information. The weather gating parameter is used to adjust the fusion ratio of the aligned geometric features and the aligned image features under different weather conditions. Based on the weather gating parameters, the aligned geometric features and the aligned image features are fused to obtain the fused features.
7. The method according to claim 6, characterized in that, The process of fusing the aligned geometric features and the aligned image features based on the weather gating parameters to obtain the fused features includes: Based on the density information in the geometric condition information, determine the structural weight information corresponding to different regions in the target scene; Based on the structural weight information and the weather gating parameters, determine the geometric fusion weight and image fusion weight corresponding to different regions in the target scene; Based on the geometric fusion weights and the image fusion weights, the aligned geometric features and the aligned image features are adaptively weighted and fused to obtain initial fused features; The initial fusion features are subjected to residual refinement to obtain the fusion features.
8. A point cloud-guided cross-modal image generation device, characterized in that, include: The information determination module is used to acquire lidar point cloud data and image data corresponding to the target scene, and determine geometric supervision information corresponding to the image data based on the lidar point cloud data. The model training module is used to train a preset initial four-dimensional Gaussian scattering scene model using the image data and the geometric supervision information to obtain a target four-dimensional Gaussian scattering scene model. The geometric supervision information is used to constrain the scene geometry represented by the four-dimensional Gaussian scattering scene model. The feature fusion module is used to generate geometric condition information corresponding to the target scene based on the target four-dimensional Gaussian scattering scene model, and to perform cross-modal feature fusion by combining the geometric condition information and the image data to obtain fused features; An image generation module is used to generate an image based on the fusion features to obtain a target image corresponding to the target scene.
9. An electronic device, characterized in that, include: processor; as well as A memory storing computer-readable instructions that, when executed by the processor, implement the point cloud-guided cross-modal image generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the point cloud-guided cross-modal image generation method as described in any one of claims 1 to 7.