Multi-source image time domain super-division method and system for giant constellation

By using the neural Schrödinger bridge model and dual-path semantic constraint technology, the problems of insufficient temporal resolution and multi-source data fusion in space-based remote sensing systems were solved, achieving high-timeliness and high-precision remote sensing image conversion and improving dynamic situational awareness capabilities.

CN121481841APending Publication Date: 2026-02-06PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV

Patent Information

Application Number
CN202511504065.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Current space-based remote sensing systems face bottlenecks in temporal resolution and multi-source remote sensing data fusion, resulting in insufficient temporal resolution of remote sensing images, making it difficult to achieve timely monitoring of border and conflict areas. Furthermore, the asynchronous nature of multi-source data makes it difficult to form a spatiotemporally consistent semantic expression.

Method used

A multi-source image temporal super-resolution method for giant constellations is adopted. Cross-modal feature alignment is performed through a neural Schrödinger bridge model. Combined with adversarial learning and phased optimization strategies, dual-path semantic constraints are introduced to achieve high-fidelity and high-precision conversion of SAR/infrared images to visible light modalities.

Benefits of technology

It effectively overcomes the "pairing difficulty" problem caused by differences in time phase and perspective of multi-source remote sensing data, and achieves high-time-efficiency and high-precision remote sensing image time series density, improves the integrity and accuracy of dynamic situational awareness, and reduces computational complexity and training difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481841A_ABST
    Figure CN121481841A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing and computer vision, and discloses a multi-source image time domain super-division method and system for giant constellations, and the method comprises the steps: obtaining multi-source remote sensing image data which comprises an SAR image, an infrared image and a visible light image; constructing a cross-modal feature space, mapping multi-source remote sensing image data to a unified semantic space, and realizing cross-modal feature alignment; the SAR image or the infrared image is converted into a visible light modal image based on a neural Schrodinger bridge model, and the neural Schrodinger bridge model is decomposed into a plurality of Markov chain sub-problems and solved through adversarial learning and a staged optimization strategy; two-way semantic constraints are introduced into the neural Schrodinger bridge model, and the two-way semantic constraints comprise visual feature matching constraints and text guiding constraints, so that semantic consistency of the generated image in a CLIP feature space is optimized; and the visible light image sequence after time domain super-division is output, and the time resolution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing and computer vision, and particularly relates to a multi-source image time domain super-resolution method and system for a giant constellation. BACKGROUND

[0002] Current space-based remote sensing systems face serious technical bottlenecks in practical applications, mainly in (1) insufficient time resolution of remote sensing images due to satellite revisit period limitation; (2) difficulty in fusion of multi-source remote sensing data due to different acquisition time phases.

[0003] A single satellite usually carries only 1-2 imaging sensors (such as optical / SAR), making it difficult to simultaneously acquire multi-modal data. The revisit period of low-orbit satellites for a fixed area is usually 1-3 days, and even if a constellation is formed, it is difficult to achieve hourly observation. Optical sensors are severely affected by clouds, and the actual effective data acquisition rate is less than 30%. The lack of data supply directly leads to various application problems: more than 60% of mobile targets are lost due to data intervals of several hours; discrete snapshot observations are difficult to reconstruct the dynamic evolution process of the battlefield; it is impossible to obtain multi-modal verification data after an attack in real time; the monitoring blind period between revisits has a missing detection rate of up to 40% for illegal border crossings; it is difficult to discover sudden events such as pipeline damage and illegal construction that damage basic military facilities in a timely manner. The lack of time resolution directly affects the integrity and accuracy of dynamic situation awareness. Discrete time point "snapshot" information is difficult to construct a continuous and complete situation evolution map. In critical applications such as crisis warning and conflict escalation assessment, data updates may lead to misjudgment or delay. Therefore, relying on a giant constellation combining high, medium and low orbits, constructing a multi-source image time domain super-resolution technology for a giant constellation, converting modalities for multi-source satellite multi-modal image data at different times, and finally forming an image mainly in the visible light modality, increasing the time series density of remote sensing images in the visible light modality in sensitive areas, and achieving significant super-resolution in time resolution, can truly meet the monitoring needs of high-time-efficiency border and conflict areas.

[0004] In the field of national security, the fusion of multi-source data (such as optical, SAR, and infrared) is the key to improving the situational awareness capability. However, due to the differences in orbit period, sensor configuration, and observation conditions of different satellites, there is significant temporal asynchrony in multi-source data, which poses a fundamental challenge to cross-modal data fusion. The essence of this problem lies in the decoupling of the time dimension and the spatial dimension, making it difficult for multi-source data to form a spatio-temporally consistent semantic representation. The observation time interval of different satellites for the same region can range from a few minutes (such as low-orbit satellite networking) to several hours (such as medium-high-orbit satellites), during which the position of dynamic targets (such as moving vehicles, ships, and personnel) can change by several kilometers. The development process of sudden events (such as fires, explosions, and crowd gatherings) has minute-level change characteristics. Due to the temporal differences, traditional multi-source data can only provide discrete "snapshots" and cannot reconstruct the complete chain of event evolution. SUMMARY

[0005] To achieve the purpose of the present application, a multi-source image time domain super-resolution method for a mega constellation is provided, comprising: Step S1: acquiring multi-source remote sensing image data, the multi-source remote sensing image data comprising SAR images, infrared images, and visible light images; Step S2: constructing a cross-modal feature space, mapping the multi-source remote sensing image data to a unified semantic space, and realizing cross-modal feature alignment; Step S3: converting the SAR images or infrared images to visible light modality images based on a neural Schrödinger bridge model, wherein the neural Schrödinger bridge model is decomposed into multiple Markov chain sub-problems and solved through an adversarial learning and staged optimization strategy; Step S4: introducing a two-way semantic constraint in the neural Schrödinger bridge model, the two-way semantic constraint comprising a visual feature matching constraint and a text-guided constraint to optimize the semantic consistency of the generated images in the CLIP feature space; Step S5: outputting the time domain super-resolved visible light image sequence to improve the temporal resolution.

[0006] In some embodiments, step S2 comprises: Step S21: using a CLIP model to establish a joint embedding space, mapping the SAR images, infrared images, and visible light images to the unified semantic space; Step S22: suppressing the speckle noise of the SAR images through a spectral normalization convolution layer; Step S23: using adaptive instance normalization to dynamically adjust the feature distribution, realizing feature alignment from SAR or infrared modality to visible light modality; Step S24: Constructing a multi-scale feature pyramid and applying a gated attention mechanism for feature selection.

[0007] In some embodiments, step S3 comprises: Step S31: decomposing the neural Schrödinger bridge process into a series of Markov chains, each corresponding to a local neural Schrödinger bridge sub-problem, learning the mapping from the source domain to the target domain; Step S32: for each sub-problem, training the conditional generator according to the optimization objective, wherein the optimization objective includes: minimizing the transport cost, implementing entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

[0008] In some embodiments, in step S3, the staged optimization strategy comprises: Generating intermediate samples through interpolation and noise addition, and gradually optimizing the generation result; Repeating the generation process until the final visible light modality image is obtained.

[0009] In some embodiments, in step S4, the dual-path semantic constraint comprises: a visual feature matching loss to constrain the similarity of the generated image in the CLIP visual feature space with the target domain visible light image, and the relative similarity with the source domain SAR or infrared image; a text-guided loss to constrain the matching degree of the generated image in the CLIP text semantic space with the target description.

[0010] In some embodiments, the visual feature matching loss is determined according to the following formula: wherein, L feat represents the visual feature matching loss value; N represents the number of samples in a training batch; represents the loss calculated for all N samples in the batch and summed up; x i represents the i-th input source domain image; G represents the generator in the neural Schrödinger bridge model; G x i represents the visible light modality image generated by the generator according to the source domain image x i ; E vis represents the image feature extractor; E opt represents the target domain feature extractor;​E sar This represents the source domain image feature extractor.

[0011] In some specific embodiments, the text guidance loss is determined according to the following formula: in, L text Indicates the text-guided loss value; Indicates the distribution of data from the source domain. p x All input images sampled in the middle x Seeking expectations; E vis Indicates an image feature extractor; G ( x This indicates that the generator is based on the source domain image. x The generated visible light modal image; E t This represents the text encoder of the CLIP model; prompt represents the target text prompt; exp( ) represents an exponential function; K This represents the number of negative sample text descriptions used for contrastive learning; p k Indicates the first k A negative sample text description Indicates all K The similarity of the text descriptions of each negative sample is summed.

[0012] To achieve the same inventive objective, this application also provides a multi-source image temporal super-resolution system for giant constellations, comprising: Data acquisition module: used to acquire multi-source remote sensing image data, including SAR images, infrared images and visible light images; Feature space construction module: used to construct a cross-modal feature space, mapping the multi-source remote sensing image data to a unified semantic space to achieve cross-modal feature alignment; Image conversion module: used to convert the SAR image or infrared image into a visible light modal image based on the neural Schrödinger bridge model, wherein the neural Schrödinger bridge model is decomposed into multiple Markov chain problems and solved through adversarial learning and phased optimization strategies; Semantic constraint module: used to introduce dual-path semantic constraints into the neural Schrödinger bridge model, the dual-path semantic constraints including visual feature matching constraints and text guidance constraints, to optimize the semantic consistency of the generated image in the CLIP feature space; The result output module is used to output the visible light image sequence after time-domain super-resolution, thereby improving the temporal resolution.

[0013] In some embodiments, the feature space construction module is configured to perform the following steps: Step S21: Establish a joint embedding space using a CLIP model to map the SAR image, infrared image and visible light image to the unified semantic space; Step S22: Suppress speckle noise of the SAR image through a spectral normalization convolution layer; Step S23: Dynamically adjust the feature distribution using adaptive instance normalization to achieve feature alignment from SAR or infrared modal to visible light modal; Step S24: Construct a multi-scale feature pyramid and apply a gated attention mechanism for feature selection.

[0014] In some embodiments, the image conversion module is configured to perform the following steps: Step S31: Decompose the neural Schrödinger bridge process into a series of Markov chains, each corresponding to a local neural Schrödinger bridge sub-problem, to learn the mapping from the source domain to the target domain; Step S32: For each sub-problem, train a conditional generator according to an optimization objective, wherein the optimization objective includes: minimizing the transport cost, implementing entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

[0015] The above technical solutions have the following advantages: The application provides a multi-source image time domain super-resolution method and system for a giant constellation, introduces a neural Schrodinger bridge model based on cross-modal feature space matching, combines adversarial learning and a staged optimization strategy, can learn the complex nonlinear mapping relationship between SAR / infrared and visible light image data without relying on strict pairing, effectively overcomes the 'pairing difficulty' problem of multi-source remote sensing data caused by time phase and viewing angle differences, and realizes high-fidelity and high-precision conversion from the SAR / infrared mode to the visible light mode. Moreover, the application innovatively introduces a double-path semantic constraint mechanism, which simultaneously constrains the absolute similarity of the generated image and the target domain, the relative similarity with the source domain and the semantic matching degree with the text description in the CLIP joint feature space, ensures that the generated image not only looks like visible light, but also is loyal to the source image in semantic content and structure information and meets the task expectation, greatly reduces semantic distortion and false generation in complex scenes. In addition, the application adopts spectral normalization convolution, adaptive instance normalization (AdaIN) and gated attention feature pyramid components in the feature extraction stage, effectively suppresses the inherent speckle noise interference of the SAR image, realizes adaptive alignment and fusion of the feature distribution between different modes. At the same time, the complex neural Schrodinger bridge problem is decomposed into a series of Markov chain sub-problems and solved through a staged strategy, which reduces the computational complexity and training difficulty of the model, makes it more suitable for processing high-resolution remote sensing image data, and improves the practicability and scalability of the method. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 A flowchart of a multi-source image time domain super-resolution method for a giant constellation provided by an embodiment of the present application is shown in the figure. Figure 2 A structural diagram of a multi-source image time domain super-resolution system for a giant constellation provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments.

[0019] Examples of the described embodiments are illustrated in the accompanying drawings, throughout which like reference characters represent like elements or elements with similar functions. The embodiments described below are exemplary and intended to be illustrative of the present invention, and are not to be understood as limiting the present invention.

[0020] Embodiment one One embodiment of the present application provides a multi-source image time domain super-resolution method for mega constellation, referring to Figure 1 As shown in the figure, comprising: Step S1: acquiring multi-source remote sensing image data, the multi-source remote sensing image data including SAR image, infrared image and visible light image; Step S2: constructing a cross-modal feature space, mapping the multi-source remote sensing image data to a unified semantic space, and realizing cross-modal feature alignment; In one specific embodiment of the present application, step S2 comprises: Step S21: using a CLIP model to establish a joint embedding space, mapping the SAR image, infrared image and visible light image to the unified semantic space; Step S22: suppressing the speckle noise of the SAR image through a spectral normalization convolution layer; Step S23: using adaptive instance normalization to dynamically adjust the feature distribution, realizing feature alignment from SAR or infrared modal to visible light modal; Step S24: constructing a multi-scale feature pyramid and applying a gated attention mechanism for feature selection.

[0021] Specifically, a joint embedding space based on CLIP is established to map SAR / infrared modal images and visible light images to a unified semantic space. By using the contrast learning characteristics of CLIP, cross-modal feature alignment is realized under the condition of no pairing.

[0022] (1) Spectral-spatial joint adaptation Spectral normalization convolution: introduce a spectral normalization convolution layer in the input stage, and the formula is represented as: Effectively suppress the pollution of speckle noise of SAR image to CLIP feature space (2) Adaptive instance normalization AdaIN ( x ) gamma ( sigma ( x ) x mu ( x ) β wherein, gamma , β The parameters are dynamically generated by a modal classifier to realize alignment of feature distribution from SAR to visible light.

[0023] (3) Multi-scale pyramid fusion A three-scale feature pyramid is constructed, and a gated attention mechanism is used to realize feature selection: G = sigma ( Conv ( x ))⊙ x wherein, σ is a sigmoid function, and ⊙ is element-wise multiplication.

[0024] (4) Improved InfoNCE loss of contrast learning optimization design Difficult negative samples are introduced to mine, and the similarity Top30% negative samples are selected from the batch. The temperature coefficient τ adopts an adaptive strategy: tau =0.1+0.05×epoch / 100.

[0025] Step S3: converting the SAR image or infrared image into a visible light modal image based on a neural Schrödinger bridge model, wherein the neural Schrödinger bridge model is decomposed into a plurality of Markov chain sub-problems, and is solved through an adversarial learning and a staged optimization strategy; In one specific embodiment of the present application, step S3 comprises: Step S31: decomposing the neural Schrödinger bridge process into a series of Markov chains, each Markov chain corresponding to a local neural Schrödinger bridge sub-problem, and learning a mapping from a source domain to a target domain; Step S32: for each sub-problem, generating a condition generator according to an optimization target, wherein the optimization target comprises: minimizing the transmission cost, realizing entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

[0026] In one specific embodiment of the present application, in step S3, the staged optimization strategy comprises: Intermediate samples are generated by interpolation and noise addition, and the generated results are gradually optimized; The generation process is repeated until the final visible light modal image is obtained.

[0027] The method solves the SB problem through adversarial learning and stage-by-stage optimization strategy, and can learn meaningful mapping of high-dimensional data through adversarial learning and regularization. Compared with traditional SB methods, the method has low computational cost and is suitable for high-resolution images. It can be combined with various GAN techniques and regularization methods to adapt to different tasks. The specific steps are as follows: (1) Decompose SB into multiple sub-problems Decompose the SB process into a series of Markov chains, namely: Wherein: (source domain), (target domain), each is a local SB, learning the mapping from to .

[0028] (2) Model each sub-problem as adversarial learning For each time step , the method trains a conditional generator , whose goal is: 1) Minimize the transport cost. Denoted as: Wherein, the term encourages and to approach, realizing optimal transport. The term implements entropy regularization to avoid overfitting.

[0029] 2) Constrain the generated distribution to match the target distribution: According to the following formula: Wherein, is the KL divergence, is the marginal distribution of the target domain samples, is the image data distribution. Make the generated distribution match the target distribution to ensure matches the target distribution .

[0030] 3) Regularization: Wherein, R is the contrastive loss, which ensures that the generated image retains the structural information of the input.

[0031] Finally, the optimization goal of the method is: Wherein, is the expected loss of CLIP, is the main loss of SB, is the weight coefficient of is a hyperparameter for the conditional overall weight.

[0032] (3) Sequential Refinement The method adopts a multi-step generation strategy, specifically: given , predict . Through interpolation and noise addition, intermediate samples are generated: wherein, wherein, , is the data distribution, is the adjustment coefficient, is the time step, is the data distribution. is used to control the interpolation weight. Repeat the process to gradually optimize the generation result.

[0033] Step S4: introducing a two-way semantic constraint in the neural Schrödinger bridge model, the two-way semantic constraint includes a visual feature matching constraint and a text guided constraint, to optimize the semantic consistency of the generated image in the CLIP feature space; In one specific embodiment of the present application, in step S4, the two-way semantic constraint includes: a visual feature matching loss to constrain the similarity of the generated image in the CLIP visual feature space with the target domain visible light image, and the relative similarity with the source domain SAR or infrared image; a text guided loss to constrain the matching degree of the generated image in the CLIP text semantic space with the target description.

[0034] Specifically, the present application innovatively adopts a two-way constraint to constrain both the absolute distance of the generated feature and the target domain, and maintain the relative similarity with the source domain.

[0035] Improved loss function: wherein, the CLIP semantic loss term: wherein, λsb represents the weight of Lsb, Lsb represents the loss of SB, λclip represents the weight of L clip , L clip represents the loss of CLIP. α represents the weight of L feat , β represents the weight of Ltext, and Ltext represents the loss of text.

[0036] In one embodiment of the present application, the visual feature matching loss is determined according to the following formula: wherein, L feat represents the visual feature matching loss value; N represents the number of samples in a training batch; represents the loss is calculated for all N samples in the batch and summed up; x i represents the i-th input source domain image; G represents the generator in the neural Schrödinger bridge model; G x i represents the visible light modality image generated by the generator according to the source domain image x i ; E vis represents the image feature extractor; E opt represents the target domain feature extractor; E sar represents the source domain image feature extractor.

[0037] In one embodiment of the present application, the text-guided loss is determined according to the following formula: wherein, L text represents the text-guided loss value; represents the expectation of all input images sampled from the source domain data distribution p x ; x E vis represents the image feature extractor; G x ) represents the visible light modality image generated by the generator according to the source domain image x ; E t represents the text encoder of the CLIP model; prompt represents the target text prompt; exp( ) represents the exponential function; K represents the number of negative sample text descriptions for contrastive learning; p k represents the i-th negative sample text description, k represents the sum of similarities for all K negative sample text descriptions.​​​​

[0038] Step S5: Output the visible light image sequence after temporal super-resolution to improve temporal resolution.

[0039] This application proposes a remote sensing image multimodal transformation method based on unpaired neural Schrödinger bridges. This method learns the optimal transport mapping between two different modal image distributions through adversarial learning and sequential refinement, thus overcoming the limitations of traditional diffusion models on high-dimensional data. Addressing the insufficient semantic alignment in cross-modal transformation by traditional unpaired NSB methods, this application proposes the CLIP multimodal feature matching mechanism to innovatively construct cross-modal semantic constraints.

[0040] Example 2 One embodiment of the present invention provides a multi-source image temporal super-resolution system for giant constellations, referring to... Figure 2 As shown, it includes: Data acquisition module 10: used to acquire multi-source remote sensing image data, including SAR images, infrared images and visible light images; Feature space construction module 20: used to construct a cross-modal feature space, mapping the multi-source remote sensing image data to a unified semantic space to achieve cross-modal feature alignment; Image conversion module 30: used to convert the SAR image or infrared image into a visible light modal image based on the neural Schrödinger bridge model, wherein the neural Schrödinger bridge model is decomposed into multiple Markov chain problems and solved through adversarial learning and phased optimization strategies; Semantic constraint module 40: used to introduce dual-path semantic constraints in the neural Schrödinger bridge model, the dual-path semantic constraints including visual feature matching constraints and text guidance constraints, to optimize the semantic consistency of the generated image in the CLIP feature space; Result output module 50: Used to output the visible light image sequence after time-domain super-resolution, thereby improving the temporal resolution.

[0041] In one specific embodiment of the present invention, the feature space construction module 20 is used to perform the following steps: Step S21: Use the CLIP model to establish a joint embedding space, and map the SAR image, infrared image and visible light image to the unified semantic space; Step S22: Suppress speckle noise in the SAR image using a spectrum-normalized convolutional layer; Step S23: Adaptive instance normalization is used to dynamically adjust the feature distribution to achieve feature alignment from SAR or infrared modes to visible light modes; Step S24: Construct a multi-scale feature pyramid and apply a gated attention mechanism for feature selection.

[0042] In one embodiment of the present application, the image conversion module 30 is configured to perform the following steps: Step S31: decompose the neural Schrödinger bridge process into a series of Markov chains, each corresponding to a local neural Schrödinger bridge subproblem, and learn the mapping from the source domain to the target domain; Step S32: for each subproblem, train the conditional generator according to the optimization objective, wherein the optimization objective includes: minimizing the transport cost, implementing entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

[0043] The multi-source image time domain super-resolution technology for giant constellation provided in the present application breaks through the time resolution bottleneck and multi-source data fusion barrier of traditional remote sensing systems, and provides high-time-efficiency and high-precision dynamic monitoring capabilities for national security key fields. The technology takes the neural Schrödinger bridge as the core framework, combines cross-modal semantic constraints and a phased optimization strategy, creates a multi-modal joint embedding space based on CLIP, solves the time-phase asynchrony problem of SAR / infrared and visible light images through spectral normalization convolution, adaptive instance normalization (AdaIN), and a gated attention mechanism. And introduce double-channel semantic constraints (visual feature matching + text guidance), realize cross-modal semantic alignment in the CLIP feature space, and realize time domain super-resolution.

[0044] The above describes only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0045] Each embodiment in the present specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between each embodiment can be referred to each other. The embodiments of the present application are described with reference to flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of the flows and / or blocks in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device realize a function implemented in the flowchart and / or block diagram. Figure 1 one flow or multiple flows and / or blocks Figure 1apparatuses that carry out functions specified in one or more of the blocks. Such computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 one or more of the processes and / or blocks Figure 1 one or more of the blocks. Such computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the processes Figure 1 one or more of the processes and / or blocks Figure 1 one or more of the blocks. Such computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the processes

[0046] The above detailed description has set forth various embodiments of the methods and devices provided by the application. The description of the application has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the application as claimed. Many modifications and variations will occur to those of ordinary skill in the art, in light of the above teachings. Accordingly, it should be noted that the above description is intended to be illustrative only and not limiting of the application. The scope of the application should be determined with reference to the appended claims.

[0047] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", "one specific embodiment" or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0048] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art will understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not drive the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A temporal super-resolution method for multi-source images of giant constellations, characterized in that, include: Step S1: Acquire multi-source remote sensing image data, which includes SAR images, infrared images, and visible light images; Step S2: Construct a cross-modal feature space and map the multi-source remote sensing image data to a unified semantic space to achieve cross-modal feature alignment; Step S3: Based on the neural Schrödinger bridge model, the SAR image or infrared image is converted into a visible light modal image. The neural Schrödinger bridge model is decomposed into multiple Markov chain problems and solved through adversarial learning and phased optimization strategies. Step S4: Introduce dual-path semantic constraints into the neural Schrödinger bridge model. The dual-path semantic constraints include visual feature matching constraints and text guidance constraints to optimize the semantic consistency of the generated image in the CLIP feature space. Step S5: Output the visible light image sequence after temporal super-resolution to improve temporal resolution.

2. The multi-source image temporal super-resolution method for giant constellations according to claim 1, characterized in that, Step S2 includes: Step S21: Use the CLIP model to establish a joint embedding space, and map the SAR image, infrared image and visible light image to the unified semantic space; Step S22: Suppress speckle noise in the SAR image using a spectrum-normalized convolutional layer; Step S23: Adaptive instance normalization is used to dynamically adjust the feature distribution to achieve feature alignment from SAR or infrared modes to visible light modes; Step S24: Construct a multi-scale feature pyramid and apply a gated attention mechanism for feature selection.

3. The multi-source image temporal super-resolution method for giant constellations according to claim 1, characterized in that, Step S3 includes: Step S31: Decompose the neural Schrödinger bridging process into a series of Markov chains, each Markov chain corresponding to a local neural Schrödinger bridging subproblem, and learn the mapping from the source domain to the target domain; Step S32: For each subproblem, train a condition generator according to the optimization objective, wherein the optimization objective includes: minimizing transmission cost, achieving entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

4. The multi-source image temporal super-resolution method for giant constellations according to claim 1, characterized in that, In step S3, the phased optimization strategy includes: Intermediate samples are generated by interpolation and noise addition, and the generated results are gradually optimized. Repeat the generation process until the final visible light modal image is obtained.

5. The multi-source image temporal super-resolution method for giant constellations according to claim 1, characterized in that, In step S4, the dual-path semantic constraints include: Visual feature matching loss is used to constrain the similarity of the generated image to the target domain visible light image in the CLIP visual feature space, as well as its relative similarity to the source domain SAR or infrared image. Text-guided loss is used to constrain the matching degree between the generated image and the target description in the CLIP text semantic space.

6. The temporal super-resolution method for multi-source images of giant constellations according to claim 5, characterized in that, The visual feature matching loss is determined according to the following formula: in, L feat This represents the visual feature matching loss value; N This indicates the number of samples in a training batch. This indicates that all items in the batch... N Calculate the loss for each sample and sum them. x i This represents the i-th input source domain image; G This represents the generator in the neural Schrödinger bridge model; G ( x i This indicates that the generator is based on the source domain image. x i The generated visible light modal image; E vis Indicates an image feature extractor; E opt This represents the target domain feature extractor; E sar This represents the source domain image feature extractor.

7. The multi-source image temporal super-resolution method for giant constellations according to claim 5, characterized in that, The text guidance loss is determined according to the following formula: in, L text Indicates the text-guided loss value; Indicates the distribution of data from the source domain. p x All input images sampled in the middle x Seeking expectations; E vis Indicates an image feature extractor; G ( x This indicates that the generator is based on the source domain image. x The generated visible light modal image; E t This represents the text encoder of the CLIP model; prompt represents the target text prompt; exp( ) represents an exponential function; K This represents the number of negative sample text descriptions used for contrastive learning; p k Indicates the first k A negative sample text description Indicates all K The similarity of the text descriptions of each negative sample is summed.

8. A multi-source image temporal super-resolution system for giant constellations, characterized in that, include: Data acquisition module: used to acquire multi-source remote sensing image data, including SAR images, infrared images and visible light images; Feature space construction module: used to construct a cross-modal feature space, mapping the multi-source remote sensing image data to a unified semantic space to achieve cross-modal feature alignment; Image conversion module: used to convert the SAR image or infrared image into a visible light modal image based on the neural Schrödinger bridge model, wherein the neural Schrödinger bridge model is decomposed into multiple Markov chain problems and solved through adversarial learning and phased optimization strategies; Semantic constraint module: used to introduce dual-path semantic constraints into the neural Schrödinger bridge model, the dual-path semantic constraints including visual feature matching constraints and text guidance constraints, to optimize the semantic consistency of the generated image in the CLIP feature space; The result output module is used to output the visible light image sequence after time-domain super-resolution, thereby improving the temporal resolution.

9. The multi-source image temporal super-resolution system for giant constellations according to claim 8, characterized in that, The feature space construction module is used to perform the following steps: Step S21: Use the CLIP model to establish a joint embedding space, and map the SAR image, infrared image and visible light image to the unified semantic space; Step S22: Suppress speckle noise in the SAR image using a spectrum-normalized convolutional layer; Step S23: Adaptive instance normalization is used to dynamically adjust the feature distribution to achieve feature alignment from SAR or infrared modes to visible light modes; Step S24: Construct a multi-scale feature pyramid and apply a gated attention mechanism for feature selection.

10. The multi-source image temporal super-resolution system for giant constellations according to claim 8, characterized in that, The image conversion module is used to perform the following steps: Step S31: Decompose the neural Schrödinger bridging process into a series of Markov chains, each Markov chain corresponding to a local neural Schrödinger bridging subproblem, and learn the mapping from the source domain to the target domain; Step S32: For each subproblem, train a condition generator according to the optimization objective, wherein the optimization objective includes: minimizing transmission cost, achieving entropy regularization, and applying contrast constraints to preserve the structural information of the input image.

Citation Information

Patent Citations

  • Wetland classification method based on multi-source images

    CN111652193A

  • Multi-source multi-mode remote sensing image air-sea interface target cooperative detection and identification method

    CN116486248A

  • Single image super-resolution reconstruction method based on improved diffusion model

    CN117575907A

  • Remote sensing image super-resolution method and product based on diffusion model and multi-modal large language model

    CN119722462A

  • Visual large model construction method and device based on multi-source remote sensing image, equipment and medium

    CN120612601A

Cited By

  • Electroencephalogram-visual cross-modal alignment method, device and equipment

    CN122388960A