A remote sensing image change reasoning method, device, equipment and storage medium

CN122840244APending Publication Date: 2026-09-29XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610990270.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,这种生成模式存在固有缺陷,推理阶段前序生成的错误会不断累积放大,导致长文本描述质量严重下降

Benefits of technology

[0012]本申请实施例提供的计算机设备,包括存储器和处理器,所述存储器存储有可在处理器上运行的计算机程序,所述处理器执行所述程序时实现本申请实施例所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840244A_ABST
    Figure CN122840244A_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing image change reasoning method and device, equipment and a storage medium. The method comprises the following steps: first, a multi-modal dialogue input containing a pre-disaster remote sensing image and a post-disaster remote sensing image, a task instruction and a true description is constructed, a unique diffusible answer area is divided, and pure mask discrete diffusion noise is executed; then, through a double mask training process, the first round completes initial prediction, position alignment and confidence evaluation, the second round performs repair prediction based on the prediction draft and the re-masked position, both rounds take the initial random mask set as the fixed supervision range, and the model parameters are updated in combination with the cross-entropy loss and the entropy regularization; finally, the remote sensing image change description text is generated through multi-step iterative denoising. The application realizes the deep alignment of training and reasoning, focuses on disaster entities and change semantics, significantly improves the accuracy of remote sensing image change description under zero sample, and can be widely applied to emergency rescue, disaster assessment and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent analysis technology for remote sensing images, and in particular to a reasoning method, apparatus, device, and storage medium for remote sensing image changes. Background Technology

[0002] With the rapid development of remote sensing technology, describing changes in pre- and post-disaster remote sensing images has become a core technical means in fields such as emergency rescue, disaster assessment, and land monitoring. This task requires the model to jointly analyze remote sensing images of the same geographic area from different time phases, identify changes in objects such as buildings, roads, bridges, water bodies, vegetation, and disaster-affected areas, and generate natural language descriptions that include temporal relationships, changed objects, changing actions, and disaster semantics. Unlike ordinary image descriptions, remote sensing change descriptions not only require the model to understand the semantics of a single image but also to accurately align the differences between images from two different time phases and express them semantically, placing extremely high demands on the model's multimodal understanding and generation capabilities.

[0003] Current methods for describing changes in remote sensing images mainly fall into two categories: one is based on autoregressive multimodal language models, which generates text token by token from left to right. However, this generation method has inherent flaws; errors generated in the preceding stages of the inference phase accumulate and amplify, leading to a severe decline in the quality of long text descriptions. Another type of text generation method based on discrete diffusion provides an iterative repair generation path, but existing solutions still have many shortcomings: First, the training and inference behaviors are misaligned, only using single random noise addition, single forward pass, and global supervision, failing to simulate the core behaviors of "retaining high-confidence tokens, backing up low-confidence tokens, and iterative repair" in the inference stage; Second, the diffusion region is chaotic, and the diffusion process is prone to erroneously destroying conditional information such as image tokens and task instructions, causing the model to face both missing conditions and missing answers simultaneously, resulting in mixed training objectives; Third, the training signal is unbalanced, with low-semantic whitespace, punctuation, and function word tokens having completely consistent weights with key entity words and disaster words, and the model tends to output conservative high-frequency tokens in the high-noise stage, weakening the learning of the semantics of changing objects and disasters; Fourth, there is unnecessary module redundancy, using the end-to-end length predictor as the core training component, which not only does not conform to the data flow where the actual answer length is known in the training stage, but also introduces additional errors.

[0004] Therefore, how to achieve deep alignment between training and inference behavior of discrete diffusion models, accurately control the diffusion destruction area, focus training signals on the core semantics of the task, and at the same time simplify the model structure and reduce the cost of engineering implementation are the technical problems that urgently need to be solved in the current field. Summary of the Invention

[0005] In view of this, the inference method, apparatus, device, and storage medium for remote sensing image changes provided in this application can achieve deep alignment between the training and inference behavior of discrete diffusion models, accurately control the diffusion destruction area, focus the training signal on the core semantics of the task, and simplify the model structure while reducing the cost of engineering implementation. The inference method, apparatus, device, and storage medium for remote sensing image changes provided in this application are implemented as follows: This application provides an embodiment of a method for inferring changes in remote sensing images, including: The training samples are processed to construct multimodal conversational input and divide the answer region to obtain a complete input sequence with the answer region boundary marked. The training samples include pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions and real change description text. The complete input sequence marked with the answer region boundary is subjected to pure mask-style discrete diffusion noise addition processing to obtain the first noise-added sequence and the initial random mask set; The first noisy sequence and the initial random mask set are subjected to a first round of forward prediction, prediction position alignment and confidence evaluation to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization. The first round vocabulary output and the first double mask position are processed by constructing a draft answer, repairing the double mask, and performing a second round of forward prediction to obtain the second cross-entropy loss and the second entropy regularization. The first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization are weighted and combined and backpropagated to obtain the trained multimodal discrete diffusion language model. The remote sensing image to be analyzed and the fixed task instructions are input into the trained multimodal discrete diffusion language model to perform multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

[0006] In some embodiments, the process of constructing multimodal conversational input and segmenting answer regions on training samples to obtain a complete input sequence labeled with answer region boundaries includes: The user-side information and assistant-side information in the training samples are standardized for conversational input format to obtain a basic input structure that conforms to the target multimodal processor input specification. The user-side information includes fixed task instructions, pre-disaster remote sensing images and post-disaster remote sensing images, and the assistant-side information includes text describing real changes. The target multimodal processor is used to encode the basic input structure to obtain a complete input sequence; The length of the user-side conditional sequence in the complete input sequence is statistically processed to obtain the user-side conditional input boundary length; Using the user-side conditional input boundary length as the boundary, the complete input sequence is processed to mark the answer region, resulting in a complete input sequence marked with the answer region boundary.

[0007] In some embodiments, performing pure mask-based discrete diffusion noise addition processing on the complete input sequence marked with the answer region boundaries to obtain a first noisy sequence and an initial random mask set includes: The preset diffusion training parameters are sampled using a course-based time step method to obtain the current diffusion time step. The current diffusion time step is compared with the preset total number of diffusion steps to obtain the current mask ratio. The complete input sequence marked with the answer region boundary is subjected to maskable filtering processing to obtain a candidate mask set that conforms to the preset maskable rules; Random mask replacement, empty set verification, and forced mask processing are performed on the candidate mask set and the current mask ratio to obtain the initial random mask set and the first noise-added sequence.

[0008] In some embodiments, the first round of forward prediction, prediction position alignment, and confidence evaluation processing on the first noisy sequence and the initial random mask set to obtain the first round vocabulary output, the first mask position, the first cross-entropy loss, and the first entropy regularization includes: The first noisy sequence is input into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the first round of vocabulary output; The initial random mask set is aligned using a prediction mechanism to obtain an aligned supervisory mask set. The first cross-entropy loss is calculated by performing a first cross-entropy loss process on the first round vocabulary output, the aligned supervision mask set, and the real change description text. The first entropy regularization process is performed on the first round vocabulary output, the complete input sequence marked with the answer region boundary, and the initial random mask set to obtain the first entropy regularization; The first round word list output, the initial random mask set, and the preset dynamic temperature confidence parameter are subjected to low confidence filtering to obtain the first mask position.

[0009] In some embodiments, the step of constructing a draft answer, repairing the overmask, and performing a second round of forward prediction on the first round vocabulary output and the first overmask position to obtain a second cross-entropy loss and a second entropy regularization includes: The first round of vocabulary output is processed by maximum probability prediction extraction to obtain the answer draft sequence; The second-stage input sequence is obtained by replacing the remask markers at the first remask position with the answer draft sequence. The second-stage input sequence is fed into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the second-round vocabulary output. The second round vocabulary output, the aligned supervision mask set, and the real change description text are processed by the second cross-entropy loss calculation to obtain the second cross-entropy loss. The second entropy regularization process is performed on the output of the second round vocabulary, the complete input sequence marked with the boundaries of the answer region, and the initial random mask set to obtain the second entropy regularization.

[0010] In some embodiments, the weighted combination and backpropagation of the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization to obtain the trained multimodal discrete diffusion language model includes: The first cross-entropy loss and the first entropy regularization are subjected to a first round of loss weighting calculation to obtain the first round of loss components; The second round of loss weighting is performed on the second cross-entropy loss and the second entropy regularization to obtain the second round of loss components. The total loss of the model training is obtained by performing a weighted summation of the loss components from the first round and the loss components from the second round. The total training loss of the model and the multimodal discrete diffusion language model to be trained are backpropagated and the parameters are iteratively updated to obtain the trained multimodal discrete diffusion language model.

[0011] This application provides an inference device for remote sensing image changes, comprising: The processing module is used to construct multimodal conversational input and divide the answer region into training samples to obtain a complete input sequence with the answer region boundary marked. The training samples include pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions and real change description text. The processing module is also used to perform pure mask-style discrete diffusion noise addition processing on the complete input sequence marked with the answer region boundary to obtain the first noise-added sequence and the initial random mask set; The processing module is further configured to perform a first round of forward prediction, prediction position alignment and confidence evaluation on the first noisy sequence and the initial random mask set to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization; The processing module is also used to construct a draft answer, repair the overmask, and perform a second round of forward prediction processing on the first round vocabulary output and the first overmask position to obtain the second cross-entropy loss and the second entropy regularization. The propagation module is used to perform weighted combination and backpropagation processing on the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss and the second entropy regularization to obtain the trained multimodal discrete diffusion language model. The generation module is used to input the remote sensing image to be analyzed and the fixed task instructions into the trained multimodal discrete diffusion language model to perform multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

[0012] The computer device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.

[0013] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method described in this application embodiment.

[0014] This application provides a method, apparatus, device, and storage medium for inferring changes in remote sensing images. It first constructs a multimodal dialogic input containing pre-disaster and post-disaster remote sensing images, task instructions, and a realistic description. A unique, diffusible answer region is defined, and pure mask-based discrete diffusion noise addition is performed. Subsequently, a dual-mask training process is used. The first round completes initial prediction, location alignment, and confidence evaluation. The second round performs repair prediction based on the prediction draft and re-masked locations. Both rounds use an initial random mask set as a fixed supervision range, combining cross-entropy loss and entropy regularization to update model parameters. Finally, multi-step iterative denoising generates descriptive text of remote sensing image changes. This achieves deep alignment between discrete diffusion model training and inference behavior, precisely controls the diffusion damage area, focuses the training signal on the core semantics of the task, simplifies the model structure, reduces engineering implementation costs, and solves the technical problems mentioned in the background art. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1A schematic diagram illustrating the implementation flow of a remote sensing image change inference method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating an implementation process for obtaining a complete input sequence, provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a remote sensing image change inference device provided in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0018] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0019] Figure 1 This is a schematic flowchart illustrating the implementation of a remote sensing image change inference method provided in this application embodiment, including steps 101 to 106. Wherein, Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order for a remote sensing image change inference method. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0020] Step 101: Perform multimodal conversational input construction and answer region segmentation on the training samples to obtain a complete input sequence with the answer region boundaries marked.

[0021] In this embodiment, training samples are first obtained. Each training sample contains pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions, and text describing real changes. The user-side information and assistant-side information in the training samples are standardized according to the dialogue format of a multimodal large model to obtain a basic input structure conforming to the target multimodal processor input specification. The user-side information includes fixed task instructions, pre-disaster remote sensing images, and post-disaster remote sensing images, while the assistant-side information is text describing real changes. The target multimodal processor is used to encode the basic input structure to obtain a complete input sequence containing user-side conditional sequences and assistant-side answer sequences. The length of the user-side conditional sequence in the complete input sequence is calculated to obtain the user-side conditional input boundary length. Using this boundary length as the boundary, the complete input sequence is marked with answer regions, resulting in a complete input sequence marked with answer region boundaries. The assistant answer token region after the boundary length is the only region where diffusion destruction is allowed; image tokens, task instruction tokens, and user-side conditional prompt tokens cannot be diffusion destroyed.

[0022] The masking rules are as follows: only tokens within the answer area can participate in masking; hint tokens, image condition tokens, task instruction tokens, and user-side condition information cannot participate in masking; special tokens declared by the tokenizer generally cannot participate in masking, but dedicated diffusion mask markers are reserved as writable mask markers; "think start and end protection markers" and preset protected tokens do not participate in masking; blank tokens can participate in masking, but their weight is set to 0.01 in subsequent cross-entropy calculations.

[0023] Step 102: Perform pure mask-style discrete diffusion noise addition processing on the complete input sequence marked with the boundary of the answer region to obtain the first noise-added sequence and the initial random mask set.

[0024] In this embodiment, course-based time-step sampling is first performed according to preset diffusion training parameters to obtain the current diffusion time step. The core parameters of the diffusion training are preset, including total diffusion steps T=64, step size Δt=8, course starting level C=3, course ramp-up ratio ρ=0.1, and high-noise bias maximum exponent α_max=2.0. In the early stages of training, the upper bound of sampling is:

[0025] The current diffusion time step t is sampled within the integer closed interval [8, 24], and t can take any integer value within this interval, not limited to multiples of 8, to avoid the model directly encountering highly destructive samples in the initial stage. As the training rounds progress, the upper bound of sampling increases linearly to 64 according to the course progress, and the sampling range expands to the integer closed interval [8, 64]. Simultaneously, the high noise bias exponent gradually increases from 0 to 2.0. Approximately uniform sampling, when Time sampling probability and The probabilities are directly proportional, and the following relative probabilities exist between t=64 and t=8:

[0026] Where P(t=64) represents the sampling probability of sampling at diffusion time step 64 in the course-style time step sampling, and P(t=8) represents the sampling probability of sampling at diffusion time step 8; t represents the current diffusion time step; 64 and 8 represent the time step values ​​in the high noise stage and the low noise stage, respectively; the exponent 2 is the high noise bias exponent α reached in the later stage of training, which is used to make the sampling probability proportional to t raised to the power of α.

[0027] This significantly increases the sampling probability of high-noise samples in the later stages of training, enhancing the training model's ability to recover text from a near-full-mask state.

[0028] Calculate the ratio of the current diffusion time step to the total number of diffusion steps to obtain the current mask ratio. The formula for calculating the mask ratio is:

[0029] Where r represents the current mask ratio; t represents the current diffusion time step; T represents the preset total diffusion steps; and r is used to determine the proportion of tokens extracted from the candidate mask set and replaced with dedicated diffusion mask markers in this noise addition.

[0030] This ratio represents the proportion of the number of candidate tokens that need to be masked during this noise-adding process to the total number of candidate mask tokens, and serves as the basis for the subsequent random selection of mask positions.

[0031] The complete input sequence marked with the answer region boundaries is subjected to maskable filtering to obtain a candidate mask set that conforms to preset maskable rules. A random mask replacement is performed on the candidate mask set and the current mask ratio. The selected candidate mask token is replaced with a dedicated diffusion mask marker, while the unselected candidate mask tokens remain visible. Random token replacement is not used throughout the noise addition process. An empty set check is performed on the obtained initial random mask set. If the diffusion time step t > 0 and the initial random mask set is empty, at least one candidate mask token is forcibly selected for replacement, resulting in the final initial random mask set and the first noise-added sequence.

[0032] Step 103: Perform the first round of forward prediction, prediction position alignment and confidence evaluation on the first noisy sequence and the initial random mask set to obtain the first round vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization.

[0033] In this embodiment, the first noisy sequence is input into the multimodal discrete diffusion language model to be trained for forward prediction, resulting in the first round of vocabulary output. This model is built upon the Qwen3.5-9B multimodal visual language model / Qwen3.5-VL style multimodal backbone. Before training, the causal attention restrictions in the text configuration and full attention module are disabled, allowing tokens within the answer region to access contextual information from both sides simultaneously, thus meeting the requirements of discrete diffusion fill-in-the-blank tasks. After performing multimodal joint inference on the input sequence, the model outputs the vocabulary probability distribution corresponding to each answer position, i.e., the first round of vocabulary output. Since the multimodal language model uses a next-token prediction mechanism—that is, the probability distribution of the i-th position output by the model is used to predict the token at the (i+1)-th position in the input sequence—the initial random mask set needs to be shifted left to ensure that the positions of the supervision masks correspond one-to-one with the predicted positions output by the model. After the left shift, an aligned supervision mask set is obtained, which is used to limit the calculation range of the subsequent cross-entropy loss. Using the real tokens describing the actual changes in the text as the supervision target, the first cross-entropy loss is calculated by combining the aligned supervision mask set with the first round vocabulary output. The formula for calculating the first cross-entropy loss is as follows:

[0034] in, For real tokens, For the corresponding token type weight, The first round of vocabulary outputs the result at position i. Represents the token-by-token cross-entropy. The size of the initial random mask set.

[0035] Differentiated token type weights are used in the calculation process, and these weights only apply to the cross-entropy loss, not to entropy regularization. Specifically, special tokens, pure blank tokens, and pure punctuation tokens have a weight of 0.01, function words have a weight of 0.1, number tokens have a weight of 0.5, and other content tokens including nouns, adjectives, disaster words, and entity words have a weight of 1.0. The system decodes each token number individually and classifies them according to the decoded string. When the same token number appears multiple times in the training set, the median of the classification weights is used. The cross-entropy for each token is first multiplied by the corresponding weight, and then the average is taken according to the number of tokens involved in the calculation, without using a denominator that is normalized to the sum of weights, to obtain the first cross-entropy loss. The first entropy regularization is calculated on the first round of vocabulary output, the complete input sequence marked with the answer region boundaries, and the initial random mask set. Entropy regularization applies to positions in the answer region that satisfy the maskable rules and do not belong to the initial random mask set, i.e.:

[0036] Where E represents the set of positions where entropy regularization applies; A represents the set of positions in the answer region that satisfy the maskable rule; M represents the initial random mask set; A M represents the set of visible real token positions obtained after excluding the initial random mask positions from the maskable positions in the answer region.

[0037] This set corresponds to the visible real tokens in the answer region that were not initially masked randomly, and does not participate in the cross-entropy calculation. The above set operations are performed in the same coordinate system; if calculated in the logits coordinates of the model vocabulary output, then both A and M must be aligned to their predicted positions according to the same shift rule. Shannon entropy is defined as:

[0038] Where H(p_{k,i}) represents the Shannon entropy of the k-th round vocabulary output at position i; p_{k,i}(v) represents the probability that the token is predicted as vocabulary element v at position i of the k-th round vocabulary output; v represents any token in the vocabulary; the summation range is the entire vocabulary; and log represents the logarithmic function.

[0039] The formula for calculating the first entropy regularization Ent1 is:

[0040] Where Ent_k represents the entropy regularization of the k-th round, k∈{1,2} corresponds to the first and second rounds of forward prediction, respectively; |A M| represents the number of positions in the entropy regularization set; i represents the set A. Any token position in M; H(p_{k,i}) represents the Shannon entropy of the probability distribution of the vocabulary corresponding to that position; A is the set of positions in the answer region that satisfy the maskable rule; M is the initial random mask set. This represents the location of the visible real token in the answer region that is not covered by the initial mask. If A If M is empty, the entropy regularization term for this round is set to 0, or its calculation is skipped in the implementation to avoid undefined cases caused by the mean of an empty set. Entropy regularization does not use the weights of the real label and token type; instead, it is added to the total loss as a mean. By adding a positive entropy regularization term to the total loss, it penalizes high-entropy distributions and encourages the model to form low-entropy, high-confidence predictions for visible real tokens. This avoids using visible answer tokens as cross-entropy supervision targets for simple copy-and-paste training; instead, it imposes low-entropy confidence constraints on the prediction distribution through entropy regularization. Based on the preset dynamic temperature confidence parameter, the first round of vocabulary output is filtered for low confidence within the initial random mask set to obtain the first mask position. The dynamic temperature calculation formula is:

[0041] Where τ(t′) represents the dynamic temperature of the current remasking stage; τ_0 represents the preset initial temperature or base temperature coefficient; t′ represents the subsequent time step corresponding to the current remasking; T represents the total number of diffusion steps; 0.5+0.5t′ / T in parentheses is used to linearly adjust the temperature with subsequent time steps, so that the larger the time step, the higher the temperature and the more thorough the remasking screening.

[0042] in This represents the subsequent time step corresponding to the current remasking, where T is the total number of diffusion steps. Confidence assessment can employ at least one of the following strategies: entropy strategy, top-k marginal probability strategy, or maximum probability strategy: when using the entropy strategy, a higher entropy in the output distribution indicates lower confidence; when using the top-k marginal probability strategy, a smaller interval between the two highest probabilities indicates lower confidence; when using the maximum probability strategy, a lower maximum probability value indicates lower confidence. The remasking ratio is determined by... The remasking candidate range is controlled and limited to the initial random mask set, and is not allowed to extend to the visible real token area that was not initially masked. The positions of low-confidence tokens are selected based on the evaluation results, ultimately yielding the first remasking position.

[0043] Step 104: Construct a draft answer, repair the overmask, and perform a second round of forward prediction processing on the first round vocabulary output and the first overmask position to obtain the second cross-entropy loss and the second entropy regularization.

[0044] In this embodiment, the maximum probability predicted token for each position in the first-round vocabulary output is extracted and concatenated to obtain the answer draft sequence. The tokens corresponding to the first mask positions in the answer draft sequence are replaced with dedicated diffusion mask markers to obtain the second-stage input sequence. This process does not call the forward noise function or perform any additional random noise addition. The second-stage input sequence is then fed into the same multimodal discrete diffusion language model to be trained for forward prediction, resulting in the second-round vocabulary output, which can be used to determine the second mask set again.

[0045] Using the real tokens describing the actual changes in the text as the sole supervision target, the second cross-entropy loss is calculated by combining the aligned supervision mask set with the output of the second-round vocabulary. The supervision scope of the second cross-entropy loss remains the initial random mask set and does not change due to changes in the position of the first mask or the second mask set. The formula for calculating the second cross-entropy loss is:

[0046] in, This is the result of the second-round vocabulary output at position i, with all other parameters being exactly the same as the first cross-entropy loss. The calculation uses the same differentiated token type weighting rule as the first cross-entropy loss.

[0047] The second entropy regularization is calculated on the second-round vocabulary output, the complete input sequence marked with the answer region boundaries, and the initial random mask set. The scope of the second entropy regularization is exactly the same as that of the first entropy regularization, which is still the positions in the answer region that satisfy the maskable rules and do not belong to the initial random mask set. The formula for calculating the second entropy regularization is:

[0048] The parameter definition is exactly the same as the first entropy regularization, and it also does not use real labels and token type weights.

[0049] Step 105: Perform weighted combination and backpropagation processing on the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization to obtain the trained multimodal discrete diffusion language model.

[0050] In this embodiment, the first cross-entropy loss and the first entropy regularization are weighted according to preset weights to obtain the first-round loss component; the second cross-entropy loss and the second entropy regularization are weighted according to the same weights to obtain the second-round loss component. The first-round loss component and the second-round loss component are weighted and summed to obtain the total model training loss. The formula for calculating the total loss is as follows:

[0051] Among them, the weighting coefficient , Entropy canonical equilibrium coefficient .

[0052] The total training loss of the model is input into the backpropagation algorithm to calculate the gradient of the total loss with respect to the trainable parameters of the model. A preset optimizer is used to iteratively update the low-rank adaptation parameters of the model based on the gradient. During training, only the low-rank adaptation parameters are updated, without modifying the parameters of the backbone network, which significantly reduces training cost and memory usage. All training steps are repeated until the total training loss of the model converges, the preset maximum number of training epochs is reached, or the remote sensing image change description index on the validation set no longer improves. Training is then stopped and the model parameters are saved, resulting in a fully trained multimodal discrete diffusion language model. Causal attention constraints are kept off throughout the training process, allowing tokens in the answer region to access contextual information on both sides, adapting to the requirements of discrete diffusion fill-in-the-blank tasks.

[0053] Step 106: Input the remote sensing image to be analyzed and the fixed task instructions into the trained multimodal discrete diffusion language model to perform multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

[0054] In this embodiment, pre-disaster remote sensing images, post-disaster remote sensing images, and fixed task instructions are acquired and their inference input format is standardized to obtain an inference infrastructure conforming to the input specifications of the target multimodal processor. The target multimodal processor is used to encode the inference infrastructure to obtain an inference multimodal conditional sequence. An initial answer cache space is determined based on a preset length, external length prediction results, or a length given by the application scenario. An initial full-mask answer sequence is constructed using the same number of dedicated diffusion mask markers. This length is only used to construct the initial answer space and is not part of the cross-entropy loss, entropy regularization, or remasking loss during the training phase. The inference multimodal conditional sequence is concatenated with the initial full-mask answer sequence to obtain the initial inference input sequence. The initial inference input sequence is input into the trained multimodal discrete diffusion language model for multi-step iterative denoising. Starting from the initial time step t0, each iteration sequentially performs forward prediction, dynamic temperature confidence assessment, high-confidence token retention, low-confidence token remasking, and time step decrementing operations. Repeat the above iterative process until the time step decreases to 0 or a preset stopping condition is met, to obtain the complete answer sequence. Decode the complete answer sequence to obtain the corresponding remote sensing image change description text.

[0055] This application's embodiments explicitly simulate the iterative repair process during the inference phase through a double-pass remasking training mechanism, achieving deep alignment between training and inference. A pure mask-based discrete diffusion noise addition method is employed, strictly limiting diffusion destruction to the answer region, avoiding erroneous destruction of multimodal conditional information, and focusing the model training objective on answer repair and generation, rather than conditional understanding tasks. A fixed initial random mask set is used as the cross-entropy supervision range, combined with the complementary constraints of cross-entropy loss and entropy regularization, ensuring the stability of the training objective while avoiding simple copying of the visible context, thus improving the semantic accuracy of the generated text. During the inference phase, multi-step iterative denoising and text generation continuously corrects low-confidence prediction results, significantly improving the quality of long text change descriptions, and can be directly applied to practical scenarios such as emergency relief, disaster assessment, and land monitoring.

[0056] In the above Figure 1 Based on the above, this application embodiment also provides a schematic diagram of the implementation process for obtaining a complete input sequence. For example... Figure 2 As shown, steps 201 to 204 are included: Step 201: perform conversational input format standardization processing on user-side information and assistant-side information in training samples, to obtain a basic input structure conforming to the input specification of the target multimodal processor.

[0057] In the embodiment of the present application, user-side information and assistant-side information are first extracted from training samples, wherein the user-side information consists of a fixed task instruction, pre-disaster remote sensing images and post-disaster remote sensing images, and the assistant-side information is a corresponding real change description text. According to the dialogue template supported by the target multimodal processor, the user-side information is encapsulated into a single-round user message, and the assistant-side information is encapsulated into a corresponding single-round assistant message, so as to form a standardized dialogue structure. The fixed task instruction adopts a unified expression, which is used to clearly inform the model that it needs to compare two remote sensing images and generate a change description in natural language form; the pre-disaster and post-disaster remote sensing images are preprocessed according to the resolution and number of channels required by the processor and then embedded into the user message, ensuring that the image information can be correctly recognized. Finally, the basic input structure conforming to the input specification of the target multimodal processor is obtained, eliminating format differences between different samples.

[0058] Step 202: encode the basic input structure by using the target multimodal processor to obtain a complete input sequence.

[0059] In the embodiment of the present application, the target multimodal processor matched with the target multimodal language model is adopted to perform joint encoding on the standardized basic input structure. The encoding process processes text information and image information simultaneously: word segmentation and tokenization conversion are performed on text content such as the fixed task instruction in the user message, visual feature extraction is performed on pre-disaster and post-disaster remote sensing images and converted into an image token sequence through serialization, and special token markers required by the dialogue format are automatically added at the same time. After encoding is completed, a complete input sequence arranged in the order of "user-side condition sequence + assistant-side answer sequence" is obtained, wherein the user-side condition sequence includes all tokens corresponding to user messages, and the assistant-side answer sequence includes all tokens corresponding to the real change description text.

[0060] Step 203: perform length statistics processing on the user-side condition sequence in the complete input sequence to obtain a user-side condition input分界 length.

[0061] In the embodiment of the present application, the complete input sequence obtained through encoding is traversed, and the total number of all tokens contained in the user-side condition sequence is counted, and this number is the user-side condition input分界 length. The statistical range covers all content on the user side, including text tokens of task instructions, image tokens of two remote sensing images, and special tokens for dialogue format, ensuring that the分界 length can accurately separate the user-side conditions and the assistant-side answers.

[0062] Step 204: Using the user-side conditional input boundary length as the boundary, perform answer region marking processing on the complete input sequence to obtain the complete input sequence marked with the answer region boundary.

[0063] In this embodiment, the complete input sequence is divided and labeled using the statistically obtained user-side conditional input boundary length as the boundary: all tokens in the complete input sequence with indices less than the boundary length are marked as non-diffusible conditional regions, and all tokens with indices greater than or equal to the boundary length are marked as diffuseable answer regions. After labeling, a complete input sequence marked with answer region boundaries is obtained. Subsequent discrete diffusion disruption and loss calculation operations can only be applied to tokens within the answer region. Image tokens, task instruction tokens, and user-side conditional prompt tokens within the conditional region remain intact and will not be modified by the diffusion process.

[0064] This application's embodiments eliminate format differences between different training samples by constructing a standardized conversational input format, ensuring that the multimodal processor can uniformly and accurately parse text and image information, thus improving the consistency of input data. Based on the user-side conditional sequence length, the answer region boundaries are precisely delineated, clarifying the unique legitimate scope of diffusion disruption. This fundamentally prevents key conditional information such as image tokens and task instruction tokens from being modified during the diffusion process, solving the problem of mixed training objectives in traditional diffusion models. Joint encoding using a processor compatible with the multimodal language model achieves seamless fusion of text and image information, ensuring that the model can fully utilize the visual features of two-phase remote sensing images and the semantic information of task instructions for change analysis.

[0065] In some embodiments, a pure mask-style discrete diffusion noise addition process is performed on the complete input sequence marked with the boundary of the answer region to obtain a first noise-added sequence and an initial random mask set, including: performing a course-style time-step sampling process on the preset diffusion training parameters to obtain the current diffusion time step.

[0066] Specifically, core parameters for diffusion training are pre-defined, including the total number of diffusion steps, step size, starting course level, course ramp-up ratio, and maximum high-noise bias exponent. In the early stages of training, the upper bound of the sampling for each diffusion time step is set to the product of the starting course level and the step size. This ensures that the current diffusion time step randomly samples within an integer range not lower than the step size and not higher than this upper bound, preventing the model from directly encountering highly destructive samples in the initial stages. As training progresses, the upper bound of the sampling increases linearly according to the course ramp-up ratio up to the total number of diffusion steps, expanding the sampling range to an integer range not lower than the step size and not higher than the total number of diffusion steps. Simultaneously, the high-noise bias exponent gradually increases from 0 to its maximum value. When the bias exponent is 0, the sampling probability at different time steps is approximately uniform; when the bias exponent reaches its maximum value, the sampling probability is proportional to the square of the time step, significantly increasing the sampling probability of high-noise samples in the later stages of training, thus enhancing the model's ability to recover text from a near-fully masked state.

[0067] Furthermore, the ratio of the current diffusion time step to the preset total diffusion steps is calculated to obtain the current mask ratio.

[0068] Specifically, the current diffusion time step obtained by sampling is divided by the preset total number of diffusion steps to obtain the current mask ratio. This ratio represents the ratio of the number of candidate tokens that need to be masked in this noise addition process to the total number of candidate mask tokens, and serves as the basis for subsequent random selection of mask positions.

[0069] Furthermore, the complete input sequence marked with the boundaries of the answer region is subjected to maskable filtering to obtain a set of candidate masks that conform to the preset maskable rules.

[0070] Specifically, based on preset masking rules, the complete input sequence marked with the answer region boundaries is filtered token by token to obtain a candidate mask set. The masking rules are as follows: only tokens within the answer region can participate in masking; hint tokens, image condition tokens, task instruction tokens, and user-side condition information cannot participate in masking; special tokens declared by the tokenizer generally cannot participate in masking, but dedicated diffusion mask markers are reserved as writable mask markers; preset protected tokens cannot participate in masking; blank tokens can participate in masking, but their weights are set to lower values ​​in subsequent cross-entropy calculations. The filtering process excludes all tokens that do not meet the rules, retaining only tokens within the answer region that meet the requirements as candidate masks.

[0071] Furthermore, random mask replacement, empty set verification, and forced mask processing are performed on the candidate mask set and the current mask ratio to obtain the initial random mask set and the first noisy sequence.

[0072] Specifically, based on the calculated current mask ratio, a corresponding number of tokens are randomly selected from the obtained candidate mask set. These selected tokens are then uniformly replaced with dedicated diffusion mask markers. Unselected candidate mask tokens retain their original true values. Random token replacement is not used throughout the noise addition process. After the initial replacement, an empty set check is performed on the obtained initial random mask set: if the current diffusion time step is greater than 0, but the initial random mask set is empty, at least one token is forcibly randomly selected from the candidate mask set and replaced with a dedicated diffusion mask marker, ensuring that each training sample has a valid cross-entropy supervision position. Finally, the final initial random mask set and the corresponding first noise-added sequence are obtained.

[0073] This application employs a strategy combining curriculum-based time-step sampling with high-noise bias. In the early stages of training, the diffusion intensity is limited to reduce the difficulty of cold starts, while in the later stages, the sampling probability of high-noise samples is increased. This balances the stability of the training process with the model's ability to generate complete text from a near-full-mask state during cold starts. Only pure mask noise is used for noise addition, without random token replacement, avoiding the introduction of additional irrelevant noise that could interfere with model learning. This allows the model to focus on learning its ability to recover true semantics from masked tokens. A strict masking selection rule limits the candidate mask range, while an empty set verification and forced masking mechanism ensure that each training sample has an effective cross-entropy supervision position, avoiding the decrease in training efficiency caused by unsupervised samples.

[0074] In some embodiments, the first noisy sequence and the initial random mask set are subjected to a first round of forward prediction, prediction position alignment and confidence evaluation to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization, including: inputting the first noisy sequence into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the first round of vocabulary output.

[0075] Specifically, the first noisy sequence is input into the multimodal discrete diffusion language model to be trained. This model is built on a general multimodal visual language model, and the causal attention restrictions in the text configuration and full attention module have been disabled before training, allowing tokens within the answer region to access contextual information from both sides simultaneously, adapting to the requirements of discrete diffusion fill-in-the-blank tasks. After performing multimodal joint inference on the input sequence, the model outputs the vocabulary probability distribution corresponding to each answer position, i.e., the first round of vocabulary output.

[0076] Furthermore, the initial random mask set is aligned using a prediction mechanism to obtain an aligned supervised mask set.

[0077] Specifically, since the multimodal language model to be trained uses a next-token prediction mechanism—that is, the probability distribution of the i-th position in the model output is used to predict the token at the (i+1)-th position in the input sequence—the initial random mask set needs to be shifted to the left to ensure that the positions of the supervision masks correspond one-to-one with the predicted positions in the model output. After the left shift, an aligned supervision mask set is obtained, which is used to limit the calculation range of the subsequent cross-entropy loss.

[0078] Furthermore, the first cross-entropy loss is calculated by performing a first cross-entropy loss process on the first round vocabulary output, the aligned supervision mask set, and the actual change description text.

[0079] Specifically, using the real-world change descriptions in the training samples as the supervised target, and combining the aligned supervised mask set with the first-round vocabulary output, the cross-entropy loss is calculated token-by-token. Differentiated token type weights are used in the calculation, and these weights only affect the cross-entropy loss, not entropy regularization. Specifically, special tokens, pure whitespace tokens, and pure punctuation tokens have a weight of 0.01, function words have a weight of 0.1, number tokens have a weight of 0.5, and other content tokens including nouns, adjectives, disaster words, and entity words have a weight of 1.0. The token-by-token cross-entropy is first multiplied by the corresponding weight, and then averaged over the number of tokens involved in the calculation to obtain the first cross-entropy loss.

[0080] Furthermore, the first entropy regularization process is performed on the first round vocabulary output, the complete input sequence marked with the answer region boundaries, and the initial random mask set to obtain the first entropy regularization.

[0081] Specifically, the scope of entropy regularization is first determined: positions in the answer region that satisfy the maskable rule and are not part of the initial random mask set, i.e., visible real token positions in the answer region not covered by the initial mask. For each position within this scope, the Shannon entropy of its probability distribution is calculated based on the output of the first round of vocabulary. Then, the average of the entropy values ​​of all positions is taken to obtain the first entropy regularization. Entropy regularization does not use the real labels. By adding a positive entropy regularization term to the total loss, it penalizes high-entropy distributions and encourages the model to form low-entropy, high-confidence predictions for visible real tokens. This avoids using visible answer tokens as cross-entropy supervision targets for simple copy-and-paste training, but instead imposes low-entropy confidence constraints on their prediction distribution through entropy regularization.

[0082] Furthermore, the first round of vocabulary output, the initial random mask set, and the preset dynamic temperature confidence parameter are subjected to low-confidence filtering to obtain the position of the first layer of mask.

[0083] Specifically, based on a preset dynamic temperature confidence parameter, the confidence of the first round of vocabulary output is evaluated within the initial random mask set. The dynamic temperature is dynamically adjusted with each subsequent time step corresponding to the current re-mask; the larger the time step, the higher the dynamic temperature. Confidence evaluation can employ any of the following strategies: entropy strategy, top-k marginal probability strategy, or maximum probability strategy: when using the entropy strategy, a higher entropy in the output distribution indicates lower confidence; when using the top-k marginal probability strategy, a smaller interval between the two highest probabilities indicates lower confidence; when using the maximum probability strategy, a lower maximum probability value indicates lower confidence. The positions of low-confidence tokens are selected based on the evaluation results, and selection is only allowed within the initial random mask set, not extending to the initially unmasked visible real token area, ultimately yielding the first mask position.

[0084] This application's embodiments employ a prediction mechanism to align the initial random mask set with the next token prediction paradigm of the multimodal language model, ensuring the accuracy of the cross-entropy loss calculation and avoiding training bias caused by supervision misalignment. It achieves a mutually exclusive division of the scope of cross-entropy loss and entropy regularization: cross-entropy only supervises the mask tokens that need to be recovered, while entropy regularization only constrains the confidence of visible real tokens. This ensures the model's ability to learn and recover semantics while encouraging the model to form high-confidence predictions for the visible context, thus improving the model's semantic modeling capabilities. By filtering low-confidence tokens only within the initial random mask set, it avoids extending remasking to the originally reliable visible real token region, preventing noise propagation and providing accurate target locations for subsequent second-round repair training.

[0085] In some embodiments, the first round vocabulary output and the first re-mask position are processed to construct a draft answer, repair the re-mask, and perform a second round of forward prediction to obtain a second cross-entropy loss and a second entropy regularization, including: performing maximum probability prediction extraction processing on the first round vocabulary output to obtain a draft answer sequence.

[0086] Specifically, based on the first-round vocabulary output obtained from the first round of forward pass, for each position within the answer region, the token with the highest probability value in the vocabulary probability distribution at that position is extracted as the initial prediction result for that position. The highest probability prediction tokens for all positions are then concatenated according to the order of the answer region to obtain a complete draft answer sequence. This draft answer sequence represents the model's first complete prediction result for the initial noisy sequence, reflecting the model's initial generation capability under the initial masking conditions.

[0087] Furthermore, the answer draft sequence and the first mask position are replaced by a remasking mark to obtain the second-stage input sequence.

[0088] Specifically, based on the draft answer sequence, the predicted tokens corresponding to the first mask positions are uniformly replaced with dedicated diffusion mask markers, while the predicted tokens at other positions remain unchanged. The entire replacement process does not call any forward noise functions or perform any additional random noise addition operations; it only performs backtracking processing on the low-confidence positions identified in the first round. This process completely replicates the core operational logic of the inference stage, namely, retaining predictions that the model considers reliable and only re-masking uncertain positions for secondary repair. After the replacement is completed, the second-stage input sequence is obtained.

[0089] Furthermore, the second-stage input sequence is fed into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the second-round vocabulary output.

[0090] Specifically, the second-stage input sequence is fed into the same multimodal discrete diffusion language model to be trained as the first round, and the second round of forward inference is performed. The model retains the setting of disabling causal attention, allowing tokens within the answer region to access contextual information from both sides simultaneously. After performing multimodal joint inference on the second-stage input sequence, the model outputs the vocabulary probability distribution corresponding to each answer position, i.e., the second-round vocabulary output.

[0091] Furthermore, the second round of vocabulary output, the aligned supervision mask set, and the actual change description text are processed by calculating the second cross-entropy loss to obtain the second cross-entropy loss.

[0092] Specifically, the text describing real changes in the training samples is used as the sole supervised target. Combining the aligned supervised mask set with the second-round vocabulary output obtained in this round, cross-entropy loss is calculated token-by-token. The supervised scope of the second cross-entropy loss is completely consistent with that of the first cross-entropy loss, always encompassing all tokens in the initial random mask set. Regardless of whether a token is re-masked in the second round, as long as it belongs to the initial random mask set, it must participate in the calculation of the second cross-entropy loss. The calculation process uses the same differentiated token type weighting rules as the first cross-entropy loss: special tokens, pure blank tokens, and pure punctuation tokens have a weight of 0.01, function words have a weight of 0.1, number tokens have a weight of 0.5, and content tokens have a weight of 1.0. The token-by-token cross-entropy is first multiplied by the corresponding weight, and then averaged according to the number of tokens involved in the calculation to obtain the second cross-entropy loss. This design avoids the dynamic drift of the cross-entropy supervised set with the re-masking position, ensuring the consistency of the training objectives in both rounds.

[0093] Furthermore, the second-round vocabulary output, the complete input sequence marked with the answer region boundaries, and the initial random mask set are subjected to second-entropy regularization calculation to obtain the second-entropy regularization.

[0094] Specifically, the scope of the second entropy regularization is exactly the same as that of the first entropy regularization: it covers the locations in the answer region that satisfy the maskable rule and are not part of the initial random mask set—that is, the visible real token locations in the answer region that are not covered by the initial mask. For each location within this scope, the Shannon entropy of its probability distribution is calculated based on the output of the second-round vocabulary. Then, the average entropy values ​​of all locations are taken to obtain the second entropy regularization. The second entropy regularization also does not use real labels or any token type weights. By adding a positive entropy regularization term to the total loss, it continues to penalize high-entropy distributions, encouraging the model to make low-entropy, high-confidence predictions for visible real tokens, thus forming a consistent constraint with the first-round entropy regularization.

[0095] This embodiment eliminates random noise addition in the second round of forward pass, performing remasking repair solely based on the first round's prediction draft and low-confidence positions. This completely replicates the core behavior of the inference phase: "retaining high-confidence tokens, backing up low-confidence tokens, and repairing again." This achieves behavioral isomorphism between training and inference from a training mechanism perspective, resolving the training-inference misalignment problem of traditional diffusion models. The supervision scope of the second cross-entropy loss remains the initial random mask set, unaffected by changes in the first mask position, preventing the cross-entropy supervision set from dynamically drifting with remasking and ensuring the consistency and stability of the training objectives in both rounds. The two rounds of entropy regularization employ identical scope and calculation rules, forming a unified confidence constraint, further enhancing the model's understanding of contextual semantics and its confidence assessment capabilities.

[0096] In some embodiments, the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization are weighted and combined and backpropagated to obtain a trained multimodal discrete diffusion language model, including: performing a first round of loss weighting calculation on the first cross-entropy loss and the first entropy regularization to obtain the first round of loss components.

[0097] Specifically, the weight coefficient for the first-round loss is pre-set to 1.0, and the balance coefficient for entropy regularization is set to 0.1. The first cross-entropy loss is added to the result of multiplying the first entropy regularization by the balance coefficient, and then multiplied by the weight coefficient of the first-round loss to obtain the first-round loss component. The balance coefficient is used to adjust the contribution ratio of cross-entropy loss and entropy regularization, avoiding excessive dominance of the entropy regularization term in the training process, and ensuring that the model can learn to recover the mask token while also modeling the confidence of the visible context.

[0098] Furthermore, a second round of loss weighting is performed on the second cross-entropy loss and the second entropy regularization to obtain the second round of loss components.

[0099] Specifically, the same weight parameters as the first round are used, i.e., the weight coefficient of the second round loss is 1.0, and the balance coefficient of the entropy regularization remains 0.1. The second cross-entropy loss is added to the result of multiplying the second entropy regularization by the balance coefficient, and then multiplied by the weight coefficient of the second round loss to obtain the second round loss component. The two rounds use completely identical weight settings to ensure a balance between the training contributions of the initial prediction stage and the remasking repair stage, avoiding model bias towards one stage of learning.

[0100] Furthermore, the first-round loss component and the second-round loss component are weighted and summed to obtain the total loss for model training.

[0101] Specifically, the first-round loss component obtained in the first step is directly added to the second-round loss component obtained in the second step to obtain the total training loss of the model. Since the weight coefficients of the two rounds of loss are both 1.0, the total loss is the arithmetic sum of the two rounds of loss components. This design embodies the core idea of ​​dual-pass training: the initial prediction and the repair prediction are equally important and together constitute the optimization objective of the model.

[0102] Furthermore, backpropagation and parameter iterative update are performed on the total training loss of the model and the multimodal discrete diffusion language model to be trained to obtain the trained multimodal discrete diffusion language model.

[0103] Specifically, this method employs low-rank adaptation to fine-tune a general multimodal visual language model. During training, only the low-rank adaptation parameters are updated, without modifying the parameters of the model's backbone network, significantly reducing training costs and memory usage. The total training loss of the model is input into the backpropagation algorithm to calculate the gradient of the total loss with respect to the trainable parameters of the model. A preset optimizer is then used to iteratively update the model parameters based on the gradient. All training steps are repeated until the total training loss of the model converges, the preset maximum number of training epochs is reached, or the descriptive index of remote sensing image changes on the validation set no longer improves. At this point, training is stopped, and the model parameters are saved, ultimately yielding a trained multimodal discrete diffusion language model.

[0104] This application's embodiments employ a weighted combination of two rounds of equally weighted losses to balance the training contributions of the initial prediction phase and the remasking repair phase, preventing the model from favoring one phase and ensuring the effectiveness of the dual-pass training mechanism. By adjusting the contribution ratio of cross-entropy loss and entropy regularization through a preset balance coefficient, the core task of prioritizing the learning of the recovered mask token is ensured while also considering the need for visible context confidence modeling, achieving a balance in multi-objective optimization. The use of low-rank adaptation for parameter updates trains only a small number of low-rank parameters without modifying the model's backbone network, significantly reducing the memory usage and computational cost required for training, and improving the model's practicality and training efficiency.

[0105] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0106] like Figure 3 As shown in the illustration, this application also provides an inference device 300 for remote sensing image changes. The device includes: The processing module 301 is used to construct multimodal conversational input and divide the answer region into the training samples to obtain a complete input sequence with the answer region boundary marked. The training samples include pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions and real change description text.

[0107] The processing module 301 is also used to perform pure mask-style discrete diffusion noise addition processing on the complete input sequence marked with the boundary of the answer region to obtain the first noise-added sequence and the initial random mask set.

[0108] The processing module 301 is also used to perform a first round of forward prediction, prediction position alignment and confidence evaluation on the first noisy sequence and the initial random mask set to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization.

[0109] The processing module 301 is also used to construct a draft answer, repair the mask, and perform a second round of forward prediction processing on the first round vocabulary output and the first mask position to obtain the second cross-entropy loss and the second entropy regularization.

[0110] The propagation module 302 is used to perform weighted combination and backpropagation processing on the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization to obtain the trained multimodal discrete diffusion language model.

[0111] The generation module 303 is used to input the remote sensing image to be analyzed and the fixed task instructions into the trained multimodal discrete diffusion language model to carry out multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

[0112] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0113] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0114] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module that implements the method or a structure within a hardware component.

[0115] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.

[0116] This application also provides a non-volatile computer-readable storage medium storing a computer program or instructions thereon, which, when executed, enables the method described in this application embodiment to be implemented.

[0117] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.

[0118] The aforementioned storage media include, but are not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions.

[0119] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0120] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0121] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A method for inferring changes in remote sensing images, characterized in that, include: The training samples are processed to construct multimodal conversational input and divide the answer region to obtain a complete input sequence with the answer region boundary marked. The training samples include pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions and real change description text. The complete input sequence marked with the answer region boundary is subjected to pure mask-style discrete diffusion noise addition processing to obtain the first noise-added sequence and the initial random mask set; The first noisy sequence and the initial random mask set are subjected to a first round of forward prediction, prediction position alignment and confidence evaluation to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization. The first round vocabulary output and the first double mask position are processed by constructing a draft answer, repairing the double mask, and performing a second round of forward prediction to obtain the second cross-entropy loss and the second entropy regularization. The first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization are weighted and combined and backpropagated to obtain the trained multimodal discrete diffusion language model. The remote sensing image to be analyzed and the fixed task instructions are input into the trained multimodal discrete diffusion language model to perform multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

2. The method according to claim 1, characterized in that, The process of constructing multimodal conversational input and dividing the answer region from the training samples yields a complete input sequence labeled with the boundaries of the answer region, including: The user-side information and assistant-side information in the training samples are standardized for conversational input format to obtain a basic input structure that conforms to the target multimodal processor input specification. The user-side information includes fixed task instructions, pre-disaster remote sensing images and post-disaster remote sensing images, and the assistant-side information includes text describing real changes. The target multimodal processor is used to encode the basic input structure to obtain a complete input sequence; The length of the user-side conditional sequence in the complete input sequence is statistically processed to obtain the user-side conditional input boundary length; Using the user-side conditional input boundary length as the boundary, the complete input sequence is processed by answer region marking to obtain a complete input sequence marked with the answer region boundary.

3. The method according to claim 1, characterized in that, The process of performing pure mask-based discrete diffusion noise addition on the complete input sequence marked with the answer region boundaries yields a first noisy sequence and an initial random mask set, including: The preset diffusion training parameters are sampled using a course-based time step method to obtain the current diffusion time step. The current diffusion time step is compared with the preset total number of diffusion steps to obtain the current mask ratio. The complete input sequence marked with the answer region boundary is subjected to maskable filtering processing to obtain a candidate mask set that conforms to the preset maskable rules; Random mask replacement, empty set verification, and forced mask processing are performed on the candidate mask set and the current mask ratio to obtain the initial random mask set and the first noise-added sequence.

4. The method according to claim 1, characterized in that, The first round of forward prediction, prediction position alignment, and confidence evaluation processing on the first noisy sequence and the initial random mask set to obtain the first round vocabulary output, the first mask position, the first cross-entropy loss, and the first entropy regularization includes: The first noisy sequence is input into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the first round of vocabulary output; The initial random mask set is aligned using a prediction mechanism to obtain an aligned supervisory mask set. The first cross-entropy loss is calculated by performing a first cross-entropy loss process on the first round vocabulary output, the aligned supervision mask set, and the real change description text. The first entropy regularization process is performed on the first round vocabulary output, the complete input sequence marked with the answer region boundary, and the initial random mask set to obtain the first entropy regularization; The first round word list output, the initial random mask set, and the preset dynamic temperature confidence parameter are subjected to low confidence filtering to obtain the first mask position.

5. The method according to claim 4, characterized in that, The process of constructing a draft answer, repairing the overmask, and performing a second round of forward prediction on the first round vocabulary output and the first overmask position to obtain a second cross-entropy loss and a second entropy regularization includes: The first round of vocabulary output is processed by maximum probability prediction extraction to obtain the answer draft sequence; The second-stage input sequence is obtained by replacing the remask markers at the first remask position with the answer draft sequence. The second-stage input sequence is fed into the multimodal discrete diffusion language model to be trained for forward prediction processing to obtain the second-round vocabulary output. The second round vocabulary output, the aligned supervision mask set, and the real change description text are processed by the second cross-entropy loss calculation to obtain the second cross-entropy loss. The second entropy regularization process is performed on the output of the second round vocabulary, the complete input sequence marked with the boundaries of the answer region, and the initial random mask set to obtain the second entropy regularization.

6. The method according to claim 1, characterized in that, The weighted combination and backpropagation of the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss, and the second entropy regularization yields the trained multimodal discrete diffusion language model, including: The first cross-entropy loss and the first entropy regularization are subjected to a first round of loss weighting calculation to obtain the first round of loss components; The second round of loss weighting is performed on the second cross-entropy loss and the second entropy regularization to obtain the second round of loss components. The total loss of the model training is obtained by performing a weighted summation of the loss components from the first round and the loss components from the second round. The total training loss of the model and the multimodal discrete diffusion language model to be trained are backpropagated and the parameters are iteratively updated to obtain the trained multimodal discrete diffusion language model.

7. A reasoning device for remote sensing image changes, characterized in that, include: The processing module is used to construct multimodal conversational input and divide the answer region into training samples to obtain a complete input sequence with the answer region boundary marked. The training samples include pre-disaster remote sensing images, post-disaster remote sensing images, fixed task instructions and real change description text. The processing module is also used to perform pure mask-style discrete diffusion noise addition processing on the complete input sequence marked with the answer region boundary to obtain the first noise-added sequence and the initial random mask set; The processing module is further configured to perform a first round of forward prediction, prediction position alignment and confidence evaluation on the first noisy sequence and the initial random mask set to obtain the first round of vocabulary output, the first mask position, the first cross-entropy loss and the first entropy regularization; The processing module is also used to construct a draft answer, repair the overmask, and perform a second round of forward prediction processing on the first round vocabulary output and the first overmask position to obtain the second cross-entropy loss and the second entropy regularization. The propagation module is used to perform weighted combination and backpropagation processing on the first cross-entropy loss, the first entropy regularization, the second cross-entropy loss and the second entropy regularization to obtain the trained multimodal discrete diffusion language model. The generation module is used to input the remote sensing image to be analyzed and the fixed task instructions into the trained multimodal discrete diffusion language model to perform multimodal condition construction and multi-step iterative denoising generation processing to obtain the remote sensing image change description text.

8. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.