A viewgraph generation method of cross-modal fusion and multi-frequency coding in an ultra-low orbit scenario

By employing a viewpoint image generation method combining cross-modal fusion and multi-frequency coding in the ultra-low orbit scenario of aircraft, and utilizing CLIP encoder and lightweight cross-modal attention module combined with multi-frequency coding technology, the problem of visual coherence and geometric consistency of aircraft targets under large-scale viewpoint and scale changes is solved, and stable reconstruction of thin deployable components and image generation stability are achieved.

CN121259221BActive Publication Date: 2026-04-17XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
Filing Date
2025-12-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In the ultra-low orbit scenario of aircraft, existing viewpoint condition generation methods based on diffusion priors are insufficient in terms of condition information fusion capability, limited radiation or illumination representation and insufficient cross-viewpoint geometric constraints. They are unable to obtain new viewpoint images with visual coherence and geometric consistency of texture details under a wide range of viewpoint and scale changes. Especially for thin, deployable components such as solar panels and other small-scale parts, distortions such as missing parts, breaks, jitter or mis-adhesion are prone to occur, making it difficult to stably reconstruct their continuous and complete geometry and appearance.

Method used

A viewpoint image generation method based on cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios is adopted. The high-dimensional semantic features of the main viewpoint image are extracted by CLIP encoder and concatenated with camera pose parameters. The high-dimensional conditional vector of the main viewpoint is generated by inputting it into a lightweight cross-modal attention module. The secondary viewpoint image is then subjected to multi-frequency coding of ray direction vector, ray origin coordinates and camera pose parameters to construct a latent diffusion model, which guides the denoising and decoding process of the latent representation tensor to generate a new viewpoint image.

Benefits of technology

It improves the texture detail coherence and geometric consistency of aircraft target images under wide range of viewing angles and scale changes, ensures stable reconstruction of thin deployable components under new viewing angles, and generates images with straight structures, sharp edges, clear occlusion boundaries, reduced background ghosting, and more stable and natural overall generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121259221B_ABST
    Figure CN121259221B_ABST
Patent Text Reader

Abstract

The application discloses a kind of viewgraph generation methods of cross-modal fusion and multi-frequency coding under ultra-low orbit scene, solve the ultra-low orbit scene of aircraft target, based on the generation of view condition and single picture three-dimensional reconstruction method of priori diffusion, because of insufficient condition information fusion ability, radiation or illumination representation is limited and cross-view geometric constraint is insufficient, it is difficult to obtain the problem that visual coherence and geometric consistency of new view image with texture details under large range of view and scale change simultaneously;The application effectively captures the multi-order angle information of direction distribution to the perception coding of light direction vector, then encodes the position of light starting coordinate, finally encodes the pose of camera pose parameter, can depict the influence of relative attitude on depth perception and perspective distortion in image generation process;The above three encoding results are used in the decoding stage of latent diffusion model, effectively improve the geometric consistency of detail recovery under new view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision and image generation methods, specifically to a viewpoint map generation method using cross-modal fusion and multi-frequency coding in ultra-low orbit scenes. Background Technology

[0002] Humans can infer the three-dimensional structure and appearance of objects based on a single image, relying on long-established visual experience and semantic-geometric priors, even if the form has not been directly observed in reality. This ability is fundamental to tasks such as navigation, object manipulation, and artistic creation. In contrast, traditional 3D reconstruction methods (such as Structure-Self-Motion (SfM), Multi-View Stereo (MVS), and volume rendering-based implicit representation methods (such as NeRF)) typically rely on sufficient multi-view observations, CAD annotations, or category-level priors, limiting their applicability in open-world tasks.

[0003] In recent years, image-based pre-trained diffusion models (such as Stable Diffusion) based on large-scale data have been shown to contain implicit geometric information (such as depth, normals, and foreground / background segmentation). Researchers have used this to inject 2D image priors into NeRF or point cloud networks, achieving 3D reconstruction based on single images through mechanisms such as score distillation sampling (SDS). However, these methods still generally face problems such as cross-viewpoint inconsistency, texture blurring, and loss of detail.

[0004] Against this backdrop, Zero123 (Zero1to3) proposes to jointly inject the input RGB image with the relative camera pose (rotation, translation) into a pre-trained U-Net network to construct a latent diffusion model with viewpoint conditions. This model can generate new images given the target viewpoint and optimize NeRF accordingly to achieve zero-shot 3D reconstruction. Although Zero123 is mainly trained on synthetic multi-view data (such as rendering data based on Objaverse), it shows good generalization ability on real images and artistic images. Nevertheless, when there are large viewpoint changes (such as large-angle rotation or significant side / pitch changes), geometric inconsistencies and texture drift are still prone to occur, and these are more pronounced on targets with complex structures or special appearances (such as insects, people, and transparent objects).

[0005] To address geometric inconsistencies and texture drift, various strategies have been proposed. For example, Cascade-Zero123 constructs a two-stage cascaded generation process: first, a base model generates neighboring views, then these views are input together with the source image into a refined model to obtain the target view, thereby improving geometric consistency and texture detail quality. SyncDreamer generates multiple views simultaneously during the inversion stage, promoting information sharing between intermediate states through cross-view feature association in 3D perception to obtain consistent multi-view results. Consistent-1-to-3 (also known as Consistent123) introduces guided attention based on geometric / epochal constraints and a multi-view attention aggregation mechanism to enhance geometric consistency and visual coherence of texture details between different views. Zero-to-Hero filters and reweights the attention map during the inference stage, alleviating deformation and misalignment problems to some extent without retraining.

[0006] While the aforementioned methods have achieved significant progress in scenarios with conventional viewpoint changes, coverage remains insufficient for observations of aircraft operating in very low Earth orbit (LEO) environments using imaging payloads. An imaging payload is an imaging device mounted on an observation aircraft, with another aircraft as the target. Under this condition of "both the imaging platform and the target are aircraft," there is often high-speed relative motion, rapid changes in the line of sight and baseline, significant fluctuations in the target's scale within the field of view with relative altitude and attitude, and complex and variable background interference. Affected by these factors, existing methods for generating viewpoint conditions based on diffusion priors and single-viewpoint... Figure 3 Existing reconstruction methods suffer from insufficient conditional information fusion capabilities, limited representation of radiation or illumination, and inadequate cross-viewpoint geometric constraints. This makes it difficult to simultaneously obtain visually coherent and geometrically consistent new-viewpoint images with varying texture details across a wide range of viewpoints and scales. This is particularly true for small-scale components on aircraft, such as thin, deployable parts (e.g., solar panels), which are characterized by their extremely small thickness, high aspect ratio, and susceptibility to perspective shortening and high-light reflection. Furthermore, these components have a low pixel count in imaging, and their boundaries are easily confused with the background. Existing methods are prone to distortions such as missing parts, breaks, jitter, or mis-adhesion, making it difficult to stably reconstruct their continuous and complete geometry and appearance. Summary of the Invention

[0007] The purpose of this invention is to address the problem that existing viewpoint condition generation methods based on diffusion priors in ultra-low orbit scenarios of aircraft are unable to obtain visually coherent and geometrically consistent new viewpoint images with texture details under large-scale viewpoint and scale changes due to insufficient condition information fusion capabilities, limited radiation or illumination representation, and insufficient cross-viewpoint geometric constraints. This invention provides a viewpoint map generation method based on cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A viewpoint map generation method based on cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios is characterized by the following steps:

[0010] Step 1: In the ultra-low orbit scenario, acquire multi-view sequence images and corresponding camera extrinsic matrix for the aircraft target, and calculate the corresponding camera pose parameters according to the width, height and set field of view of the image. Take one of the view sequence images as the main view image and the remaining view sequence images as the auxiliary view images.

[0011] Step 2: Use the CLIP encoder to extract high-dimensional semantic features from the main view image, concatenate the high-dimensional semantic features with the corresponding camera pose parameters of the main view image, and input them into a lightweight cross-modal attention module for cross-modal attention fusion to generate a high-dimensional conditional vector of the main view.

[0012] Step 3: For each auxiliary viewpoint image, a spherical harmonic function embedding model based on multi-frequency coding is used to perceptually encode the light direction vector, and a multi-frequency sine and cosine function expansion is used to perform Fourier position coding on the coordinates of the light source, and the corresponding camera pose parameters are also coded. The above coding results are spliced, fused and compressed according to the dimensions to obtain the corresponding conditional features of each auxiliary viewpoint image.

[0013] Step 4: Construct a latent diffusion model consisting of a U-Net denoising network, a variational autoencoder, and a decoder; input the multi-view sequence images of the aircraft target into the variational autoencoder to obtain clean latent representation tensors. The latent representation tensor is obtained by forward diffusion in the latent space. , as input to the U-Net denoising network;

[0014] Step 5: Input the high-dimensional conditional vector of the main viewpoint and the corresponding conditional features of each secondary viewpoint image into the U-Net denoising network to guide the latent representation tensor. Iterative denoising is performed to restore the latent representation under the target's perspective, and after being mapped by the decoder, a new perspective image of the ultra-low orbit vehicle target is obtained. For each perspective, the lightweight cross-modal attention module, the spherical harmonic function embedding model, and the U-Net denoising network are trained based on the difference between the new perspective image of the ultra-low orbit vehicle target and the image of that perspective in the multi-view sequence image.

[0015] Step 6: Input the image of the target aircraft under test from any perspective in the ultra-low orbit scenario and the extrinsic parameter matrix of the camera from the new perspective into the trained lightweight cross-modal attention module, spherical harmonic function embedding model and U-Net denoising network to obtain the new perspective image of the target aircraft under test.

[0016] Further, in step 1, the camera pose parameters include position parameters representing the camera's position in the world coordinate system, attitude parameters representing the camera's orientation in the world coordinate system, and frustum parameters representing the main view direction vector and the scale information of the imaging plane.

[0017] Furthermore, step 2 specifically involves:

[0018] Step 2.1: Extract high-dimensional semantic features from the main viewpoint image using the CLIP encoder;

[0019] Step 2.2: Sequentially concatenate the high-dimensional semantic features of the main view image with the corresponding camera pose parameters of the main view image, perform a one-time linear mapping, and apply layer normalization to form a cross-modal initial feature representation. ;

[0020] Step 2.3: Represent the initial features across modalities. The input is a lightweight cross-modal attention module consisting of two Transformer encoder layers, which learns the nonlinear coupling relationship between image semantics and geometric pose to obtain conditional feature vectors with cross-modal correlation. ;

[0021] Step 2.4: Conditional feature vectors with cross-modal correlation After projection transformation into a cross-attention conditional vector, a high-dimensional conditional vector of the main viewpoint is obtained.

[0022] Furthermore, step 2.2 specifically includes:

[0023] Step 2.2.1: Extract high-dimensional semantic features from the main viewpoint image. With camera pose parameters Concatenation is performed along the feature dimension, expressed as follows:

[0024] ;

[0025] This represents the concatenated vector. Indicates dimension;

[0026] Step 2.2.2: Process the concatenated vector A one-time linear mapping is performed, followed by layer normalization, to form an initial feature representation across modalities. The expression is:

[0027] ;

[0028] in, This indicates a normalization operation. The weight matrix represents the linear mapping. , The bias vector represents the linear mapping. .

[0029] Furthermore, in step 2.3, both layers of the Transformer encoder include a multi-head self-attention mechanism and a feedforward network, with the following expressions:

[0030] ;

[0031] ;

[0032] in, This represents the conditional feature vector of the first-layer Transformer encoder; This represents the conditional feature vector of the second-layer Transformer encoder, which is a conditional feature vector with cross-modal correlation.

[0033] Furthermore, step 2.4 specifically involves:

[0034] Conditional feature vectors with cross-modal correlation Linearly map back to the semantic dimension of the original CLIP encoder to generate cross-attention conditional vectors. This yields the high-dimensional conditional vector from the main perspective; the expression is:

[0035] ;

[0036] In the formula, Represents the linear mapping weight matrix. , Represents the bias matrix of the linear mapping. .

[0037] Furthermore, step 3 specifically involves:

[0038] Step 3.1: Employ a spherical harmonic function embedding model based on multi-frequency coding for each ray direction vector in each secondary viewpoint image. After normalization, the ray direction distribution is obtained. A spherical harmonic function is then used to perform perceptual encoding on the ray direction distribution, and the order is selected. ,get 3D directional encoding feature vector ;

[0039] Step 3.2: Set the coordinates of the ray's origin. Fourier position coding is performed using a multi-frequency sine and cosine function expansion based on neural radiation fields, specifically as follows:

[0040] Define six frequency bands , Then each coordinate dimension Mapped to:

[0041] ;

[0042] In the formula, Represents coordinate dimension The mapping result, j represents the component of the coordinate direction of the starting point of the ray, with values ​​of x, y, and z;

[0043] The coordinates of the starting point of the ray are: ;

[0044] The coordinates of the ray's origin are linearly transformed to 32 dimensions to obtain the position code of the ray's origin coordinates. ;

[0045] Step 3.3: Input the camera pose parameters corresponding to each auxiliary viewpoint image into the pose encoder, and map them into a 46-dimensional pose encoding feature vector through a two-layer feedforward network. ;

[0046] Step 3.4: Encode the direction feature vector Position encoding of the starting point coordinates of the light ray Pose encoding feature vector Concatenate the dimensional features into a unified ray condition characteristic. The expression is:

[0047] ;

[0048] Unify the characteristics of radiation conditions The input is fused and compressed to 78 dimensions into a two-layer linear network to obtain the conditional features corresponding to each secondary viewpoint image. .

[0049] Furthermore, in step 4, the multi-view sequence images of the aircraft target are input into a variational autoencoder to obtain a clean latent representation tensor. The expression for adding noise in the latent space via forward diffusion is:

[0050] ;

[0051] in, For the potential representation tensor, Let be the cumulative noise modulation coefficient at time step t, i be the view index, and t be the diffusion time step. To and Standard Gaussian noise of the same shape express The covariance.

[0052] Furthermore, in step 5, the high-dimensional conditional vector of the main viewpoint is input into the cross-attention layer of the U-Net denoising network, and the high-dimensional conditional vector of the main viewpoint is used as a key value to modulate the latent representation tensor. semantic distribution;

[0053] The conditional features corresponding to each secondary viewpoint image are input into the convolutional branch of the U-Net denoising network. The conditional features corresponding to each secondary viewpoint image are then processed by the convolutional branch and combined with the latent representation tensor. Point-by-point fusion is used to provide geometric priors.

[0054] The beneficial effects of this invention are:

[0055] 1. This invention provides a viewpoint image generation method based on cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios. The method calculates camera pose parameters based on the main viewpoint image of the target aircraft and its corresponding extrinsic parameter matrix, constructs the geometric mapping relationship of the target viewpoint, and then introduces a lightweight cross-modal attention module for fusion to generate a high-dimensional conditional vector of the main viewpoint used to guide the diffusion process. This vector is used in the decoding stage of the potential diffusion model and can guide the feature focusing and weighting at different spatial locations, thereby improving the texture detail coherence and geometric consistency of the image generated from the free viewpoint.

[0056] 2. This invention provides a viewpoint image generation method based on cross-modal fusion and multi-frequency encoding in ultra-low orbit scenarios. By perceptually encoding the ray direction vector in the auxiliary viewpoint image of the aircraft target, it can effectively capture multi-order angular information of the direction distribution, enhancing the network's ability to model the continuity of details under complex angular changes. Furthermore, the position encoding of the ray origin coordinates helps the network distinguish the relative layout of viewpoints in different spatial locations. Finally, the pose encoding of the camera pose parameters can characterize the influence of relative pose on depth perception and perspective distortion during image generation. These three encoding results are used in the decoding stage of the latent diffusion model, effectively improving the geometric consistency of detail recovery under new viewpoints. Attached Figure Description

[0057] Figure 1 This is a flowchart of an embodiment of the viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to the present invention;

[0058] Figure 2 This is a flowchart of step 3 of an embodiment of the viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios of the present invention;

[0059] Figure 3 This is a schematic diagram of the ray direction vector and ray origin coordinates in an embodiment of the viewpoint map generation method of cross-modal fusion and multi-frequency coding in an ultra-low orbit scene of the present invention;

[0060] Figure 4This is an example of a new perspective image generated in an embodiment of the perspective image generation method of cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to the present invention. Detailed Implementation

[0061] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] This invention provides a method for generating viewpoint maps in ultra-low orbit scenarios through cross-modal fusion and multi-frequency coding, such as... Figure 1 As shown, it includes the following steps:

[0063] Step 1: In the ultra-low orbit scenario, acquire multi-view image sequences and corresponding camera extrinsic matrices for the spacecraft target. And based on the image's width, height, and a set field of view, calculate the camera pose parameters, including:

[0064] (1) Translation Parameters:

[0065] The position of the camera in the world coordinate system is denoted as:

[0066] ;

[0067] These represent the camera's coordinates on the x, y, and z axes in the world coordinate system, respectively.

[0068] From the extrinsic matrix The translation component is obtained.

[0069] (2) Rotation Parameters:

[0070] The orientation of the camera in the world coordinate system can be represented by the rotation component of the extrinsic parameter matrix. extract.

[0071] To facilitate network processing, this embodiment converts the rotation matrix 𝑅 into a compact form of three-dimensional Euler angles (yaw, pitch, roll).

[0072] (3) Intrinsic Projection Parameters:

[0073] The main view direction vector and imaging plane scale information are calculated based on the image width, height, and set field of view (FOV).

[0074] One view sequence image from the multi-view image sequence of the aircraft target is used as the primary view image, and the remaining view sequence images are used as secondary view images.

[0075] Step 2: Extract high-dimensional semantic features from the main viewpoint image using a CLIP encoder, concatenate these features with the camera pose parameters, and input them into a lightweight cross-modal attention module (i.e., ...). Figure 1 Cross-modal attention fusion is performed using CCSF (Conceptual Data Set Instructions) to generate a high-dimensional conditional vector from the main perspective; specifically:

[0076] Step 2.1: Extract high-dimensional semantic features from the main viewpoint image using the CLIP encoder;

[0077] Step 2.2: Sequentially concatenate the high-dimensional semantic features of the main view image with the corresponding camera pose parameters of the main view image, perform a one-time linear mapping, and apply layer normalization to form a cross-modal initial feature representation. Specifically:

[0078] Step 2.2.1: Extract high-dimensional semantic features from the main viewpoint image. With camera pose parameters Concatenation is performed along the feature dimension, expressed as follows:

[0079] ;

[0080] This represents the concatenated vector. Indicates dimension;

[0081] Step 2.2.2: Process the concatenated vector A one-time linear mapping is performed, followed by layer normalization, to form an initial feature representation across modalities. The 772-dimensional coarsely concatenated features are reduced to 512 dimensions, and subsequent self-attention learning is stabilized through normalization. The expression is:

[0082] ;

[0083] in, This indicates a normalization operation. The weight matrix represents the linear mapping. , The bias vector represents the linear mapping. .

[0084] Step 2.3: Represent the initial features across modalities. Considered as a token sequence of length 1, the input is a lightweight cross-modal attention module consisting of two Transformer encoder layers. Through multi-head self-attention and a feedforward network, the nonlinear coupling relationship between image semantics and geometric pose is learned, resulting in a conditional feature vector with cross-modal correlation. ;

[0085] Both Transformer encoder layers include a multi-head self-attention mechanism and a feedforward network, with the following expressions:

[0086] ;

[0087] ;

[0088] in, This represents the conditional feature vector of the first-layer Transformer encoder; This represents the conditional feature vector of the second-layer Transformer encoder, which is a conditional feature vector with cross-modal correlation.

[0089] In this embodiment, a multi-head self-attention mechanism is used to capture deep dependencies within the feature dimension, and a feed-forward network is used to enrich the expressive power.

[0090] Step 2.4: Conditional feature vectors with cross-modal correlation Linearly map back to the semantic dimension of the original CLIP encoder to generate cross-attention conditional vectors. This yields the high-dimensional conditional vector from the main perspective; the expression is:

[0091] ;

[0092] In the formula, Represents the linear mapping weight matrix. , Represents the bias matrix of the linear mapping. .

[0093] By injecting the high-dimensional conditional vector of the main perspective into the cross-attention layer of the U-Net denoising network, the denoising process can simultaneously rely on image semantics and camera geometric information, achieving content consistency and structural constraints during the synthesis of new perspectives.

[0094] Step 3: To address the complex depth variations and occlusion issues in ultra-low orbit scenarios, this embodiment proposes a light-based geometric perception condition construction method, specifically as follows: Figure 2 and Figure 3 As shown, for each secondary viewpoint image, a spherical harmonic function embedding model based on multi-frequency coding is introduced ( Figure 1 The PASHE algorithm employs spherical harmonic functions for perceptual encoding of ray direction vectors, enabling the extraction of multi-order angular features to characterize direction dependence. It uses multi-frequency sine and cosine function expansions for Fourier position encoding of the ray origin coordinates, capturing high-frequency details in the spatial distribution. Furthermore, it performs pose encoding on the corresponding camera pose parameters, resulting in a compact pose embedding vector. These encoding results are then concatenated, fused, and compressed according to dimension to obtain the corresponding conditional features for each secondary viewpoint image. Specifically:

[0095] Step 3.1: For each secondary viewpoint image, for each ray direction vector Normalization is performed, and the ray direction distribution is encoded using a spherical harmonic function, with the order selected. ,get 3D directional encoding feature vector ;

[0096] Step 3.2: Determine the starting coordinates of each ray direction vector. Fourier position coding is performed using a multi-frequency sine and cosine function expansion based on neural radiation fields, specifically as follows:

[0097] Define six frequency bands , Then each coordinate dimension Mapped to:

[0098] ;

[0099] In the formula, Represents coordinate dimension The mapping result, j represents the component of the coordinate direction of the starting point of the ray, with values ​​of x, y, and z;

[0100] The coordinates of the starting point of the ray are: ;

[0101] The coordinates of the ray's origin are linearly transformed to 32 dimensions to obtain the position code of the ray's origin coordinates. ;

[0102] Step 3.3: Input the camera pose parameters corresponding to each auxiliary viewpoint image into the pose encoder, and map them into a 46-dimensional pose encoding feature vector through a two-layer feedforward network. ;

[0103] Step 3.4: Encode the direction feature vector Position encoding of the starting point coordinates of the light ray Pose encoding feature vector Concatenate the dimensional features into a unified ray condition characteristic. The expression is:

[0104] ;

[0105] Unify the characteristics of radiation conditions The input is fused and compressed to 78 dimensions into a two-layer linear network to obtain the conditional features corresponding to each secondary viewpoint image. .

[0106] At each iteration step of the diffusion sampling, the corresponding conditional features of each secondary viewpoint image The convolutional branches injected into the U-Net denoising network serve as geometric modulation signals parallel to the latent variables, thereby finely guiding the generator of the U-Net backbone to model occlusion boundaries, depth abrupt changes, and scale variations.

[0107] Step 4: Construct a latent diffusion model consisting of a U-Net denoising network, a variational autoencoder, and a decoder. Input the multi-view sequence images of the aircraft target into the variational autoencoder to obtain the corresponding clean latent variables. Subsequently, noise is added to the latent space via forward diffusion to obtain the latent representation tensor. The expression is:

[0108]

[0109] in, is the cumulative noise modulation coefficient for time step t; i is the view index, and t is the diffusion time step. To and Standard Gaussian noise of the same shape (mean 0, covariance 1), noise samples at different views and time steps are independent and typically resampled stepwise; the resulting latent representation tensor The result will be used as input to the U-Net denoising network; the denoised potential result will then be restored to the pixel domain by the decoder to generate the target viewpoint image.

[0110] Step 5: Input the high-dimensional conditional vector from the main viewpoint into the cross-attention layer of the U-Net denoising network. The high-dimensional conditional vector from the main viewpoint is used as a key value to modulate the latent representation tensor. The semantic distribution of each secondary viewpoint image is input into the convolutional branch of the U-Net denoising network. The conditional features of each secondary viewpoint image are then processed by the convolutional branch and combined with the latent representation tensor. Point-by-point fusion is used to provide geometric priors; the latent representation tensor is guided by the high-dimensional conditional vector of the main viewpoint and the corresponding conditional features of each secondary viewpoint image. Perform iterative denoising to restore the latent representation from the target's perspective. After being decoded and mapped, a new perspective image of the ultra-low orbit spacecraft target was obtained. For each viewpoint, the lightweight cross-modal attention module (CCSF), the spherical harmonic function embedding model (PASHE), and the U-Net denoising network are trained based on the differences between the new viewpoint image of the ultra-low orbit vehicle target and the image of that viewpoint in the multi-view sequence image.

[0111] Step 6: Input the image of the target aircraft under test from any perspective in the ultra-low orbit scenario and the extrinsic parameter matrix of the camera from the new perspective into the trained lightweight cross-modal attention module, spherical harmonic function embedding model and U-Net denoising network to obtain the new perspective image of the target aircraft under test.

[0112] The model is tested on images from new perspectives. Input is a single-view image and the corresponding camera's extrinsic parameter matrix, where the extrinsic parameter matrix specifies the desired target viewpoint and pose. The model uses a trained lightweight cross-modal attention module and a spherical harmonic function embedding module for forward inference to generate images from the corresponding target viewpoint. During testing, the model weights are frozen, and only forward generation is performed without gradient updates.

[0113] Figure 4 This embodiment presents a new perspective image generated using a cross-modal fusion and multi-frequency coding perspective image generation method in an ultra-low Earth orbit (UEO) scenario. The first column contains the original image of the satellite's solar panels in the UEO scenario observed by the simulated spacecraft housing the imaging payload. The second to fifth columns contain other new perspective images generated from the first column of input images, rotating around a fixed center with constant pitch and roll angles and yaw angles at 60-degree intervals. These images are used to verify the model's performance in perspective transfer and multi-viewpoint mapping. Figure 1 Performance in terms of consistency.

[0114] according to Figure 4 It is clearly visible that the satellite solar panels maintain a straight structure, sharp edges, and no obvious stretching or breakage even under large angle changes; thin components such as antennas remain intact and visible during rotation, with almost no occlusion, flipping, or deformation artifacts. A correspondence between pixels and viewing angles is explicitly established through perceptual encoding of light direction vectors, Fourier position encoding of light origin coordinates, and pose encoding of camera pose parameters, giving the new perspective image stronger geometric robustness at the detail level. Meanwhile, the lightweight cross-modal attention module ensures consistency in appearance and lighting across different viewpoints through the fusion of semantics and camera pose parameters, making metallic reflections and highlight areas appear natural and continuous across multiple views. The synergistic effect of these two components results in a higher degree of realism and consistency in the structure of the generated satellite solar panels, with clean occlusion boundaries, reduced background ghosting, and a more stable and natural overall generation.

[0115] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating viewpoint maps through cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios, characterized in that, Includes the following steps: Step 1: In the ultra-low orbit scenario, acquire multi-view sequence images and corresponding camera extrinsic matrix for the aircraft target, and calculate the corresponding camera pose parameters according to the width, height and set field of view of the image. Take one of the view sequence images as the main view image and the remaining view sequence images as the auxiliary view images. Step 2: Use the CLIP encoder to extract high-dimensional semantic features from the main view image, concatenate the high-dimensional semantic features with the corresponding camera pose parameters of the main view image, and input them into a lightweight cross-modal attention module for cross-modal attention fusion to generate a high-dimensional conditional vector of the main view. Step 3: For each auxiliary viewpoint image, a spherical harmonic function embedding model based on multi-frequency coding is used to perceptually encode the light direction vector, and a multi-frequency sine and cosine function expansion is used to perform Fourier position coding on the coordinates of the light source, and the corresponding camera pose parameters are also coded. The above coding results are spliced, fused and compressed according to the dimensions to obtain the corresponding conditional features of each auxiliary viewpoint image. Step 4: Construct a latent diffusion model consisting of a U-Net denoising network, a variational autoencoder, and a decoder; input the multi-view sequence images of the aircraft target into the variational autoencoder to obtain clean latent representation tensors. The latent representation tensor is obtained by forward diffusion in the latent space. , as input to the U-Net denoising network; Step 5: Input the high-dimensional conditional vector of the main viewpoint and the corresponding conditional features of each secondary viewpoint image into the U-Net denoising network to guide the latent representation tensor. Iterative denoising is performed to restore the latent representation under the target's perspective, and after being mapped by the decoder, a new perspective image of the ultra-low orbit target is obtained. For each perspective, based on the difference between the new perspective image of the ultra-low orbit vehicle target obtained and the image of that perspective in the multi-view sequence image, the lightweight cross-modal attention module, the spherical harmonic function embedding model, and the U-Net denoising network are trained. Step 6: Input the image of the target aircraft under test from any perspective in the ultra-low orbit scenario and the extrinsic parameter matrix of the camera from the new perspective into the trained lightweight cross-modal attention module, spherical harmonic function embedding model and U-Net denoising network to obtain the new perspective image of the target aircraft under test.

2. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 1, characterized in that, In step 1, the camera pose parameters include position parameters representing the camera's position in the world coordinate system, attitude parameters representing the camera's orientation in the world coordinate system, and frustum parameters representing the main view direction vector and the scale information of the imaging plane.

3. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 1, characterized in that, Step 2 is as follows: Step 2.1: Extract high-dimensional semantic features from the main viewpoint image using the CLIP encoder; Step 2.2: Sequentially concatenate the high-dimensional semantic features of the main view image with the corresponding camera pose parameters of the main view image, perform a one-time linear mapping, and apply layer normalization to form a cross-modal initial feature representation. ; Step 2.3: Represent the initial features across modalities. The input is a lightweight cross-modal attention module consisting of two Transformer encoder layers, which learns the nonlinear coupling relationship between image semantics and geometric pose to obtain conditional feature vectors with cross-modal correlation. ; Step 2.4: Conditional feature vectors with cross-modal correlation After projection transformation into a cross-attention conditional vector, a high-dimensional conditional vector of the main viewpoint is obtained.

4. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 3, characterized in that, Step 2.2 specifically involves: Step 2.2.1: Extract high-dimensional semantic features from the main viewpoint image. With camera pose parameters Concatenation is performed along the feature dimension, expressed as follows: ; This represents the concatenated vector. Indicates dimension; Step 2.2.2: Process the concatenated vector A one-time linear mapping is performed, followed by layer normalization, to form an initial feature representation across modalities. The expression is: ; in, This indicates a normalization operation. The weight matrix represents the linear mapping. , The bias vector represents the linear mapping. .

5. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 3, characterized in that, In step 2.3, both layers of the Transformer encoder include a multi-head self-attention mechanism and a feedforward network, with the following expressions: ; ; in, This represents the conditional feature vector of the first-layer Transformer encoder; This represents the conditional feature vector of the second-layer Transformer encoder, which is a conditional feature vector with cross-modal correlation.

6. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 5, characterized in that, Step 2.4 specifically involves: Conditional feature vectors with cross-modal correlation Linearly map back to the semantic dimension of the original CLIP encoder to generate cross-attention conditional vectors. This yields the high-dimensional conditional vector from the main perspective; the expression is: ; In the formula, Represents the linear mapping weight matrix. , Represents the bias matrix of the linear mapping. .

7. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 5, characterized in that, Step 3 specifically involves: Step 3.1: Employ a spherical harmonic function embedding model based on multi-frequency coding for each ray direction vector in each secondary viewpoint image. After normalization, the ray direction distribution is obtained. A spherical harmonic function is then used to perform perceptual encoding on the ray direction distribution, and the order is selected. ,get 3D directional encoding feature vector ; Step 3.2: Set the coordinates of the ray's origin. Fourier position coding is performed using a multi-frequency sine and cosine function expansion based on neural radiation fields, specifically as follows: Define six frequency bands , Then each coordinate dimension Mapped to: ; In the formula, Represents coordinate dimension The mapping result, j represents the component of the coordinate direction of the starting point of the ray, with values ​​of x, y, and z; The coordinates of the starting point of the ray are: ; The coordinates of the ray's origin are linearly transformed to 32 dimensions to obtain the position code of the ray's origin coordinates. ; Step 3.3: Input the camera pose parameters corresponding to each auxiliary viewpoint image into the pose encoder, and map them into a 46-dimensional pose encoding feature vector through a two-layer feedforward network. ; Step 3.4: Encode the direction feature vector Position encoding of the starting coordinates of the light ray Pose encoding feature vector Concatenate the dimensional features into a unified ray condition characteristic. The expression is: ; Unify the characteristics of radiation conditions The input is fused and compressed to 78 dimensions into a two-layer linear network to obtain the conditional features corresponding to each secondary viewpoint image. .

8. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 1, characterized in that, In step 4, the multi-view sequence images of the aircraft target are input into a variational autoencoder to obtain a clean latent representation tensor. The expression for adding noise in the latent space via forward diffusion is: ; in, For the potential representation tensor, Let be the cumulative noise modulation coefficient at time step t, i be the view index, and t be the diffusion time step. To and Standard Gaussian noise of the same shape express The covariance.

9. The viewpoint map generation method for cross-modal fusion and multi-frequency coding in ultra-low orbit scenarios according to claim 1, characterized in that, In step 5, the high-dimensional conditional vector of the main viewpoint is input into the cross-attention layer of the U-Net denoising network, and the high-dimensional conditional vector of the main viewpoint is used as a key value to modulate the latent representation tensor. semantic distribution; The conditional features corresponding to each secondary viewpoint image are input into the convolutional branch of the U-Net denoising network. The conditional features corresponding to each secondary viewpoint image are then processed by the convolutional branch and combined with the latent representation tensor. Point-by-point fusion is used to provide geometric priors.

Citation Information

Patent Citations

  • Scene space three-dimensional model dynamic modeling method based on multi-modal data

    CN119339008A

  • Image rendering method and system fusing three-dimensional perception

    CN121074238A