Heterogeneous task migration-oriented multi-modal large model collaborative optimization method

Through technical means such as cross-modal completion and mapping correction, the decision-making efficiency and reliability problems caused by the reduction of perceptual dimensions in heterogeneous task migration of multimodal large models are solved, and the agent is quickly adapted and optimized in a perceptually restricted environment is achieved, and task execution efficiency and stability are improved.

CN120373415AActive Publication Date: 2025-07-25NO 15 INST OF CHINA ELECTRONICS TECH GRP

Patent Information

Application Number
CN202510854571.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

During the heterogeneous task migration process, the multimodal large model has a decrease in decision efficiency and task execution reliability due to the sudden decrease in perceptual dimensions. The mapping between the original sub-strategy and abstract features is misaligned, affecting the agent's stable decision-making and path evasion in the early stages.

Method used

Through the five-ring closed-loop closed-loop time-driven across mode completion, mapping correction, hierarchical tuning, gradient weighting and fast adaptation, the triple high-frequency feedback mechanism of self-consistent information flow, decision flow, and reward flow is built, and the missing features are reconstructed using the cross-mode feature autoregressive complement, the dual-map relocator correction subtask triggers the boundaries, the hierarchical weight tuner adjusts the action advantages, and the rapid adaptation is achieved through consistency constraints and meta-strategy buffers.

Benefits of technology

Achieve rapid adaptation and optimization of agents in perceptually restricted environments, improve path smoothness and energy consumption convergence in the initial stage of migration, maintain security and redundancy, and improve heterogeneous task migration efficiency, decision-making reliability and long-term robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373415A_ABST
    Figure CN120373415A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous task migration-oriented multi-modal large model collaborative optimization method, particularly relates to the field of simulated confrontation, and is used for solving the problem that the decision-making efficiency and the task execution reliability of a multi-modal large model in heterogeneous task migration are reduced due to sudden reduction of sensing dimensions. According to the method, a triple high-frequency feedback mechanism of self-consistent information flow, decision flow and reward flow is constructed through full-link series cross-modal completion, mapping correction, hierarchical tuning, gradient weight and rapid adaptive five-loop closed-loop time-dependent driving; after multi-source observation passes through the implicit vector pool, triggering a boundary to be mapped to a target domain in real time, and self-correcting action advantages along with environment evolution; gradient consistency constraint anchors turn-back and dissipation errors in a convergence domain, and a reward signal is subjected to buffer scheduling to catalyze an experience subtask cluster to perform high-optimal activation and maintain activity; the global collaboration enables the initial migration section to generate a smooth path, the roundabout distance and the energy consumption peak value are converged remarkably, and the heterogeneous task migration efficiency, the decision reliability and the long-term robustness are increased synchronously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of simulated confrontation, and more specifically, to a multi-modal large model collaborative optimization method for heterogeneous task migration. Background Art

[0002] In the adversarial simulation environment, the agent first completed training in a complex scenario containing optoelectronic observation and inertial navigation data to form a hierarchical decision chain based on multimodal perception. Subsequently, the model was migrated to a simplified scenario that only retained radar echoes and communication telemetry. The sudden decrease in perception latitude made the originally rich environmental representation suddenly sparse, and the feature extraction paths that each sub-strategy within the hierarchical decision chain relied on were compressed, resulting in the previously constructed two types of sub-task trigger conditions of staying away from threats and approaching targets being difficult to match. The agent's response to environmental changes was delayed, and path avoidance and target approach actions frequently corrected.

[0003] The core technical difficulty caused by the above situation is that the multimodal large model loses key visual and inertial information during migration, and the mapping between the original sub-strategy and the abstract features is misaligned. The mapping misalignment first produces a signal gap in the perception layer, and then propagates to the upper level along the decision-making chain, weakening the command of the high-level planning module over the low-level actions. The result is a circuitous track, delayed triggering of risk avoidance actions, and exhaustion of safety redundancy. It is difficult for the intelligent agent to stably call key subtasks in the early stages, and the overall collaborative efficiency and task reliability are simultaneously attenuated.

[0004] In order to solve the above problems, a technical solution is now provided. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a multimodal large model collaborative optimization method for heterogeneous task migration, which constructs a self-consistent information flow, decision flow, and reward flow triple high-frequency feedback mechanism through full-link serial cross-modal completion, mapping correction, hierarchical tuning, gradient reweighting and rapid adaptation five-loop closed-loop time-driven; after multi-source observations are connected through the latent vector pool, the trigger boundary is mapped to the target domain in real time, and the action advantage is self-calibrated as the environment evolves; the gradient consistency constraint anchors the return and dissipation errors in the convergence domain, and the reward signal catalyzes the high-priority activation and maintains the activity of the empirical subtask cluster through buffer scheduling; global collaboration enables the generation of a smooth path at the beginning of the migration, and the detour distance and energy consumption peak converge significantly, and the efficiency, decision reliability and long-term robustness of heterogeneous task migration are simultaneously improved to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions: S1: Inject a cross-modal feature autoregressive completer into the data stream containing only radar and communication telemetry, fuse the reconstructed optoelectronic and inertial navigation gap features with the original observation sequence, and write them into a unified latent vector pool; S2: Drive the dual - mapping relocator using the latent vector pool, compare the source task hierarchical strategy trigger features in real - time, redraw the sub - task trigger boundaries away from threats and towards goals, and output the mapping correction matrix; S3: Run the hierarchical weight tuner according to the mapping correction matrix, re - distribute the action advantages and generate the corrected action probability stream, and send it together with the hidden state to the consistency constraint; S4: The consistency constraint calculates the fitting bias coefficient based on the cross - domain gradient backpropagation degree and the distribution distance dissipation rate, then re - weights the action probability stream, and encapsulates the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and writes it into the meta - policy buffer; S5: When the meta - policy buffer detects that the adjusted reward stream reaches the sparse threshold, activate the fast adaptation scheduler, sort and enable the sub - task clusters that retain the source experience weights according to the fitting bias coefficient and push them to the executor queue to complete path reshaping and safety margin recovery in the migration scenario.

[0007] In a preferred embodiment, step S1 includes the following: Use the cross - modal feature autoregressive completer to reconstruct the missing optical and inertial navigation features according to the current observation sequence and the historical observation sequence including radar echoes and telemetry data; fuse the reconstructed optical and inertial navigation features with the current observation vector to generate a complete feature vector; use an encoder to convert the complete feature vector into a latent vector; store the latent vector in a unified latent vector pool.

[0008] In a preferred embodiment, step S2 includes the following: Extract the latent vector at the current moment from the latent vector pool; use the dual - mapping relocator to generate the mapping value of the sub - task away from threats by calculating the feature direction consistency and feature amplitude difference between the latent vector and the trigger feature vector of the sub - task away from threats in the source task; generate the mapping value of the sub - task towards goals by calculating the feature direction consistency and feature amplitude difference between the latent vector and the trigger feature vector of the sub - task towards goals in the source task; determine whether to trigger the sub - task away from threats according to the comparison between the mapping value of the sub - task away from threats and the preset threshold.

[0009] In a preferred embodiment, step S2 further includes the following: Determine whether to trigger the sub - task towards goals according to the comparison between the mapping value of the sub - task towards goals and the preset threshold; generate the correction sub - matrix of the sub - task away from threats through optimization calculation to minimize the difference between the corrected trigger feature vector of the sub - task away from threats and the latent vector; generate the correction sub - matrix of the sub - task towards goals through optimization calculation to minimize the difference between the corrected trigger feature vector of the sub - task towards goals and the latent vector; combine the correction sub - matrix of the sub - task away from threats and the correction sub - matrix of the sub - task towards goals into a mapping correction matrix.

[0010] In a preferred embodiment, step S3 includes the following: Extract the correction parameters at the current moment from the mapping correction matrix; use the hierarchical weight tuner to multiply the mapping correction matrix by the source action advantage vector through matrix multiplication to generate the action advantage vector of the target task; use the softmax function to process the action advantage vector of the target task to generate the corrected action probability stream; extract the hidden state at the current moment from the hidden vector pool; send the corrected action probability stream and the hidden state to the consistency constraint for processing.

[0011] In a preferred embodiment, step S4 includes the following: Calculate the cross-domain gradient return degree and the distribution distance dissipation rate through the consistency constraint, where the cross-domain gradient return degree is obtained by decomposing the source task policy gradient and the target task policy gradient into orthogonal components and coincident components and performing a time-domain convolution operation on the orthogonal components, and the distribution distance dissipation rate is obtained by performing a fast Fourier transform on the source task feature distribution and the target task feature distribution to generate a frequency-domain spectrum and calculating the energy difference between the two frequency-domain spectra.

[0012] In a preferred embodiment, step S4 further includes the following: Use the parser to perform a weighted sum operation on the cross-domain gradient return degree and the distribution distance dissipation rate and map it through the sigmoid function to generate the fitting bias coefficient; perform a linear combination on the action probability stream according to the fitting bias coefficient to generate the weighted action probability stream.

[0013] In a preferred embodiment, step S4 further includes the following: Multiply the original reward stream by the adjustment factor based on the fitting bias coefficient to generate the adjusted reward stream; encapsulate the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and write it into the meta-policy buffer.

[0014] In a preferred embodiment, step S5 includes the following: The meta-policy buffer determines the sparsity by calculating the proportion of the number of non-zero reward values in the adjusted reward stream to the total number of time steps; compare the sparsity with the preset sparsity threshold; when the sparsity is lower than the sparsity threshold, the fast adaptation scheduler is activated; the fast adaptation scheduler sorts the sub-task clusters that retain the source experience weight according to the fitting bias coefficient.

[0015] In a preferred embodiment, step S5 further includes the following: Push the sorted sub-task clusters to the executor queue; the executor executes actions according to the order of the executor queue to achieve path reshaping and safety margin recovery.

[0016] Technical effects and advantages of the multi-modal large model collaborative optimization method for heterogeneous task migration of the present invention: Through the time-driven five-ring closed-loop of cross-modal completion, mapping correction, hierarchical tuning, gradient reweighting, and rapid adaptation in series throughout the entire link, the present invention constructs a triple high-frequency feedback mechanism of self-consistent information flow, decision flow, and reward flow. After multi-source observations penetrate through the hidden vector pool, the trigger boundary is mapped to the target domain in real time, and the action advantage is self-calibrated with the evolution of the environment; the gradient consistency constraint anchors the return and dissipation errors within the convergence domain, and the reward signal catalyzes the highly optimized activation and maintains the activity of the empirical sub-task cluster through buffer scheduling; global collaboration generates a smooth path at the initial stage of migration, significantly converging the detour distance and peak energy consumption, compressing the training step size, and maintaining the safety redundancy persistently, simultaneously improving the efficiency of heterogeneous task migration, decision reliability, and long-term robustness. Description of the Drawings

[0017] Figure 1 It is a flowchart of the multi-modal large model collaborative optimization method for heterogeneous task migration of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1: Figure 1 The multi-modal large model collaborative optimization method for heterogeneous task migration of the present invention is given, including: S1: Inject a cross-modal feature autoregressive completer at the data stream entrance containing only radar and communication telemetry, fuse the reconstructed optoelectronic and inertial navigation gap features with the original observation sequence, and then write them into the unified hidden vector pool.

[0020] S2: Use the hidden vector pool to drive the double-mapping relocator, compare the source task hierarchical strategy trigger features in real time, redraw the trigger boundaries of sub-tasks far from threats and approaching targets, and output the mapping correction matrix.

[0021] S3: Run the hierarchical weight tuner according to the mapping correction matrix, reallocate the action advantage, generate the corrected action probability flow, and send it together with the hidden state to the consistency constraint.

[0022] S4: The consistency constraint calculates the fitting bias coefficient according to the cross-domain gradient return degree and the distribution distance dissipation rate, then reweights the action probability flow, and encapsulates the fitting bias coefficient and the adjusted reward flow into a convergence signaling packet and writes it into the meta-policy buffer.

[0023] S5: When the meta - policy buffer detects that the adjusted reward stream reaches the sparse threshold, it activates the fast adaptation scheduler, sorts and enables the subtask clusters that retain the source experience weights according to the fitting bias coefficient, and pushes them to the executor queue, thus completing the path reshaping and safety margin recovery in the migration scenario.

[0024] In the adversarial simulation environment, agents usually complete training in complex scenarios with rich perceptual information. For example, relying on multi - modal data such as optoelectronic observations, inertial navigation data, radar echoes, and communication telemetry, a stable hierarchical decision - making chain is formed.

[0025] However, when the agent migrates to a simplified scenario with limited perception, such as only retaining radar echo and communication telemetry data, the sudden reduction in the perception dimension will lead to the sparsification of the environmental representation. The sub - policies in the original decision - making chain become invalid due to the compression of the feature extraction path, manifested as problems such as path avoidance hysteresis and frequent corrective actions for approaching the target. This adaptability challenge brought by heterogeneous task migration not only reduces the task execution efficiency of the agent but also threatens its safety.

[0026] To solve this problem, the invention proposes a multi - modal large - model collaborative optimization method for heterogeneous task migration. Through steps such as cross - modal feature completion, decision - boundary correction, and action probability tuning, the agent can achieve fast adaptation and optimization in a perception - limited environment. Step S1, as the starting point of the entire solution, aims to fill the perception gap through a cross - modal feature autoregressive completer, construct a unified hidden vector pool, and provide a complete feature representation basis for subsequent steps.

[0027] Step S1 includes the following content: S1.1, Definition of the data - flow entry: In the simplified scenario of the target task, the perceptual data received by the agent is limited to radar echo and communication telemetry data, excluding optoelectronic observation data and inertial navigation data in the source task. These perceptual data are organized in a time - series form, called the original observation sequence. The original observation sequence consists of observation vectors at multiple moments, and each observation vector contains two parts: the radar - echo component and the communication - telemetry component. Compared with the observation vectors of the source task, the observation vectors of the target task are significantly reduced in dimension, resulting in a sparse environmental representation. This sparse environmental representation reduces the agent's ability to analyze and make decisions because key environmental information is not fully captured.

[0028] S1.2, Function and construction of the cross - modal feature autoregressive completer: To solve the shortage of missing optoelectronic observation data and inertial navigation data in the target task, a cross - modal feature autoregressive completer is introduced. The function of the cross - modal feature autoregressive completer is to predict and reconstruct the missing optoelectronic features and inertial navigation features by analyzing the historical data of the original observation sequence and the observation vector at the current moment. These reconstructed parts are collectively referred to as missing features.

[0029] The prediction process relies on an autoregressive mechanism, which uses the observation vectors at multiple historical moments in the original observation sequence and the observation vector at the current moment to calculate the estimated value of the missing feature. The parameters of the autoregressive model are first pre-trained on the complete dataset of the source task and then fine-tuned in the simplified scenario of the target task to adapt to the characteristics of radar echo and communication telemetry data.

[0030] The autoregressive model can identify long-term dependencies in time series data, thereby generating accurate estimated values of missing features. By reconstructing the missing features, the agent obtains complete perceptual information close to the source task, enhancing its decision-making ability in the simplified scenario.

[0031] S1.3, Feature Fusion: The missing features generated by the cross-modal feature autoregressive completer are fused with the observation vector at the current moment to generate a complete feature vector.

[0032] The specific fusion method is to perform a vector concatenation operation on the missing features and the observation vector at the current moment in a predetermined order to form a complete feature vector containing radar echo components, communication telemetry components, as well as reconstructed optoelectronic features and inertial navigation features. This vector concatenation operation ensures that the complete feature vector not only retains the direct data of the original observation sequence but also integrates the reconstructed missing features, constituting a comprehensive cross-modal environmental representation.

[0033] Selecting the vector concatenation operation can maintain the independence of the original data and the reconstructed data while providing a unified feature representation for subsequent processing. The complete feature vector provides the agent with comprehensive environmental information, reducing decision-making errors caused by data missing.

[0034] S1.4, Latent Vector Encoding: The complete feature vector is input into an encoder, which is processed to generate a latent vector.

[0035] The encoder is a neural network whose structure and parameters are pre-trained on the complete dataset of the source task. The role of the encoder is to map the high-dimensional complete feature vector into a low-dimensional latent vector. Through multiple layers of calculations of the neural network, the key information in the complete feature vector is extracted and compressed into a compact representation form. The neural network can learn the complex non-linear relationships between features and generate a potential representation with high discrimination for decision-making. The latent vector retains the core information of the complete feature vector in a low-dimensional form, facilitating the agent's use in subsequent processing while reducing the computational complexity.

[0036] S1.5, Latent Vector Pool Writing: The generated latent vectors are written into a unified storage structure called the latent vector pool. The latent vector pool is responsible for storing the latent vectors generated at all times, forming a time series set containing cross-modal information. It records the latent vectors at each time step in chronological order to ensure that subsequent steps can access the complete latent vector sequence.

[0037] The technical feature of step S1 is that the missing optoelectronic features and inertial navigation features are reconstructed by a cross-modal feature autoregressive completer, and then fused with the observation vector at the current time to generate a complete feature vector. Subsequently, the encoder is used to convert it into a latent vector and store it in the latent vector pool. Reconstructing the missing features can make up for the deficiency of the perceptual data in the target task, enabling the agent to still obtain a comprehensive environmental representation in a simplified scenario and avoiding decision-making errors caused by incomplete information.

[0038] Step S1 reconstructs the optoelectronic and inertial navigation features through a cross-modal feature autoregressive completer and generates a latent vector pool, providing a complete cross-modal representation for subsequent processing. However, simply completing the features is not sufficient to restore the decision-making logic of the source task. It is necessary to further correct the sub-task trigger boundaries to adapt to the simplified environment of the target task, which is the core task of step S2.

[0039] Step S2 includes the following: S2.1, Latent vector pool call: Extract the latent vector at the current time from the latent vector pool.

[0040] The latent vector pool is a storage structure that contains the latent vectors at all times. Each latent vector is generated by fusing the original observation data and the reconstructed features in the target task. The original observation data includes radar echoes and communication telemetry data, and the reconstructed features are supplementary information extracted from optoelectronic and inertial navigation data by a cross-modal feature autoregressive completer. The latent vector at the current time is a high-dimensional vector that fully represents the agent's cross-modal perception of the environment at the current time point. The purpose of extracting the latent vector at the current time is to provide comprehensive environmental information for the subsequent redrawing of the trigger boundary, ensuring that the processing is based on accurate cross-modal data.

[0041] S2.2, Definition of the source task hierarchical strategy trigger feature: In the source task, the agent's decision-making depends on multi-modal perception data, forming trigger conditions for two types of subtasks, namely the "away from threat" subtask and the "towards target" subtask. These trigger conditions are composed of specific feature combinations, collectively referred to as the source task trigger feature set. The source task trigger feature set is divided into two parts: one is the trigger feature vector of the "away from threat" subtask, and the other is the trigger feature vector of the "towards target" subtask. The trigger feature vector of the "away from threat" subtask consists of features such as threat distance in electro-optical observation data or threat intensity in radar echo data; the trigger feature vector of the "towards target" subtask consists of features such as target direction in inertial navigation data or target signal intensity in communication telemetry data. The purpose of defining these trigger feature vectors is to provide a benchmark for subsequent mapping and repositioning, ensuring the accurate identification and correction of subtask trigger conditions in the target task.

[0042] S2.3, Construction and Function of the Dual Mapping Repositioner: The dual mapping repositioner contains two independent mapping functions, corresponding to the "away from threat" subtask and the "towards target" subtask respectively. Each mapping function generates a mapping value by comparing the hidden vector at the current moment with the trigger feature vector of the corresponding subtask in the source task. The calculation process of the mapping value includes two parts: one is to calculate the feature direction consistency between the hidden vector and the trigger feature vector, represented by the cosine value of the vector angle; the other is to calculate the feature amplitude difference between the hidden vector and the trigger feature vector, represented by the normalized value of the distance between vectors. The mapping value is generated by weighted combination of the feature direction consistency and the feature amplitude difference, and the weighting parameter is adjusted according to the sparsity of the feature distribution in the target task, and the adjustment parameter is determined by the initial test of the target task.

[0043] The feature direction consistency reflects the similarity degree between the hidden vector and the trigger feature vector, while the feature amplitude difference reflects the proximity degree between the two. The combination of the two can comprehensively evaluate the matching degree. The mapping value provides a quantitative comparison result, ensuring the accuracy and adaptability of the redrawing of the trigger conditions.

[0044] S2.4, Redrawing of the Trigger Boundary: According to the mapping value generated by the dual mapping repositioner, determine the trigger boundaries of the "away from threat" subtask and the "towards target" subtask in the target task. The determination of the trigger boundary depends on a preset threshold: when the mapping value of the "away from threat" subtask exceeds its corresponding threshold, trigger this subtask; when the mapping value of the "towards target" subtask exceeds its corresponding threshold, trigger this subtask. The threshold is calibrated through the empirical data of the source task and the initial test of the target task, ensuring a match with the dynamic characteristics of the simplified environment of the target task.

[0045] The reduction of perceptual data in the target task leads to a change in the feature distribution, and the original triggering conditions of the source task are no longer applicable, so adjustments need to be made according to the new environmental representation. The agent can accurately identify threats and targets in a simplified environment, improving the response speed and execution efficiency of decision-making.

[0046] S2.5, Generation of the mapping correction matrix: Based on the redrawn triggering boundary, generate a mapping correction matrix to adjust the adaptability of the source task sub-strategy in the target task. The mapping correction matrix consists of two sub-matrices, corresponding to the "away from threat" sub-task and the "towards target" sub-task respectively. Each sub-matrix is generated through optimization calculations, and the optimization goal is to minimize the difference between the corrected source task triggering feature vector and the current moment hidden vector. The optimization process adopts an iterative adjustment method, aligning the corrected features with the hidden vector of the target task by gradually reducing the difference.

[0047] Map the triggering features of the source task to the feature space of the target task through matrix transformation to solve the mapping misalignment problem caused by the reduction of the perception dimension. The mapping correction matrix ensures that the sub-task execution is consistent with the target task environment, improving the accuracy and reliability of decision-making.

[0048] The technical feature of step S2 is to use the cross-modal representation in the hidden vector pool to drive the double-mapping relocator, compare the hierarchical strategy triggering features in the source task in real time, redraw the triggering boundaries of the "away from threat" sub-task and the "towards target" sub-task, and generate a mapping correction matrix. Due to the sparsification of the feature distribution caused by the reduction of the perception dimension in the target task, the triggering conditions of the source task are no longer directly applicable. Through mapping relocation and boundary redrawing, the requirements of the simplified environment can be adapted. This enables the agent to accurately identify threats and targets even when the perceptual data is reduced, improving the efficiency and stability of task execution.

[0049] In the simplified scenario of the target task, the reduction of perceptual data makes the action selection strategy need to adapt to the environmental changes. Step S3 adjusts the action advantage by calling the mapping correction matrix generated in step S2, uses the hierarchical weight tuner and the softmax function to generate the corrected action probability flow, and extracts the hidden state from the hidden vector pool in step S1 to provide environmental information. This process forms a complete technical chain, which is closely connected with the previous steps and provides input data for the strategy flow adjustment in the subsequent step S4, supporting the efficient decision-making of the agent in the sparsified environment.

[0050] Step S3 includes the following: S3.1, Call of the mapping correction matrix: Extract the correction parameters at the current moment from the mapping correction matrix generated in step S2. The mapping correction matrix is a structure containing two sub-matrices. One sub-matrix corresponds to the correction parameters for the "away from threat" subtask, and the other sub-matrix corresponds to the correction parameters for the "towards target" subtask. These two sub-matrices are used to adjust the action advantage to ensure that the action selection adapts to the feature distribution of the target task. The purpose of calling the mapping correction matrix is to provide a correction basis for the subsequent adjustment of action advantage, so that the action selection strategy can match the dynamic environment of the target task, thereby enhancing the adaptability of decision-making.

[0051] S3.2, Construction and Function of the Hierarchical Weight Tuner: The hierarchical weight tuner is a dynamic adjustment module responsible for reallocating the action advantage according to the mapping correction matrix. Action advantage represents the tendency of the agent towards each action in a specific state, reflecting the priority of the action. In the source task, the action advantage is generated by the source policy network and is a vector containing multiple action advantage values. The hierarchical weight tuner generates the action advantage vector in the target task by performing a linear transformation on the source action advantage vector. The specific way of linear transformation is to perform a matrix multiplication operation on the mapping correction matrix and the source action advantage vector, that is, applying the correction parameters to the source action advantage vector through matrix multiplication to obtain the adjusted action advantage vector.

[0052] Using matrix multiplication enables the mapping correction matrix to directly act on the action advantage, making it adapt to the feature distribution after the reduction of the perception dimension in the target task. The adjustment process of the action advantage is based on the correction parameters of the mapping correction matrix, which can ensure the alignment of the action selection strategy with the target task environment, thereby improving the accuracy of decision-making.

[0053] S3.3, Generation of the Action Probability Flow: Generate the action probability flow based on the corrected action advantage vector. The action probability flow represents the probability distribution of the agent selecting each action in the current state. The method for generating the action probability flow is to process the corrected action advantage vector using the softmax function. The specific calculation process is as follows: First, calculate the exponential value of each advantage value in the corrected action advantage vector, and then divide the exponential value of each advantage value by the sum of the exponential values of all advantage values to obtain the probability value of each action, thus forming the action probability flow. The softmax function can normalize the advantage values in the action advantage vector into a probability distribution, ensuring that the sum of the probability values of all actions is 1, which is convenient for the agent to select actions according to the probability distribution.

[0054] The softmax function has a wide range of applications in the field of reinforcement learning and can effectively balance the exploration and exploitation of action selection, thereby enhancing the stability of the strategy.

[0055] S3.4, Extraction of the Hidden State: Extract the hidden state at the current moment from the latent vector pool generated in step S1. The hidden state is a vector representing the abstraction of the current environmental representation, which integrates the reconstructed optoelectronic features, inertial features, as well as the original radar data and telemetry data. The purpose of extracting the hidden state is to provide environmental information for subsequent consistency constraints, ensuring that the policy adjustment is consistent with the environmental dynamics. The extraction process of the hidden state directly depends on the latent vector pool and is completed by selecting the vector corresponding to the current moment from the latent vector pool, ensuring the integrity and accuracy of the environmental representation.

[0056] S3.5, Feed into the consistency constraint: Take the corrected action probability flow and the hidden state as inputs and feed them into the consistency constraint. The consistency constraint will calculate the fitting bias coefficient based on the action probability flow and the hidden state in subsequent steps, and further adjust the policy flow to suppress cross-domain gradient drift. The purpose of feeding into the consistency constraint is to provide the necessary data support for the final adjustment of the policy flow, ensuring that the decisions of the agent are consistent with the environmental characteristics of the target task, thereby optimizing the generation process of the policy flow.

[0057] Use the mapping correction matrix to drive the hierarchical weight tuner, adjust the source action advantage vector through matrix multiplication to generate the action advantage vector of the target task, combine with the softmax function to generate the corrected action probability flow, and feed the action probability flow and the hidden state into the consistency constraint together. The action probability flow can be adaptively adjusted according to the dynamic environment of the target task, thereby improving the decision-making accuracy and environmental adaptability of the agent in heterogeneous task migration.

[0058] In the scenario where the perception dimension is reduced, the action probability flow and the hidden state generated in step S3 may become invalid due to the misalignment of environmental characteristics. Step S4 uses the consistency constraint to calculate the cross-domain gradient foldback degree and the distribution distance dissipation rate based on the action probability flow and the hidden state in step S3, generates the fitting bias coefficient and re-weights the action probability flow, and at the same time encapsulates the convergence signaling packet and writes it into the meta-policy buffer. This process effectively corrects the deviation between the policy flow and the environmental characteristics, provides accurate input data for the fast adaptation scheduling in subsequent step S5, and supports the path reshaping and decision optimization of the agent in the scenario of sudden reduction of the perception dimension.

[0059] The consistency constraint ensures the consistency and adaptability of the model's decisions between the source task and the target task by calculating the cross - domain gradient foldback degree and the distribution distance dissipation rate. The specific process starts from the output of the previous stage. The hierarchical weight tuner adjusts the action advantage according to the mapping correction matrix, generating the corrected action probability flow and hidden state, which provide the basic data for the calculation of the consistency constraint. The consistency constraint first analyzes the policy gradients of the source task and the target task using the corrected action probability flow and hidden state. By capturing the reverse sweep amplitude of the gradient during the transfer process, it calculates the cross - domain gradient foldback degree, which reflects the differences and potential inconsistencies in policy adjustments between tasks. At the same time, the consistency constraint calculates the distribution distance dissipation rate by comparing the frequency - domain spectra of the feature distributions of the two tasks, quantifying the dissipation degree of the environmental representation, and thus revealing the evolution of the feature space between tasks. Subsequently, these two metrics are input into the parser to generate the fitting bias coefficient. This coefficient adjusts the action probability flow by linearly combining the policy features of the source task and the target task, making it both consistent with the source task policy and adaptable to the new environmental features in the target task. Finally, the fitting bias coefficient and the adjusted reward flow are encapsulated into a convergence signaling packet and written into the meta - policy buffer, providing a signal for subsequent policy adjustments. Through this process, the consistency constraint effectively connects the output of the previous stage, achieves the consistency between the policy flow and the dynamic features of the target task environment, and improves the decision - making robustness and adaptability of the model in scenarios with a sudden reduction in the perception dimension.

[0060] Step S4 includes the following: S4.1, Calculation of the cross - domain gradient foldback degree: The calculation of the cross - domain gradient foldback degree aims to quantify the difference in the reverse sweep amplitude of the policy gradients between the source task and the target task, reflecting the intensity of gradient drift. The calculation process is divided into two stages: First, the policy gradients of the source task and the target task are decomposed into orthogonal components and coincident components. The orthogonal components represent the differences in the directions of the policy gradients of the two tasks, while the coincident components represent the commonalities in the directions of the policy gradients of the two tasks.

[0061] Second, within a preset time window, a time - domain convolution operation is performed on the orthogonal components of the source task and the orthogonal components of the target task. By integrating the product of the orthogonal components of the two tasks changing over time, the cross - domain gradient foldback degree is obtained.

[0062] The reason for using time - domain convolution is that this method can capture the dynamic change trend of the gradient in the time dimension, thus accurately reflecting the degree of gradient drift. The calculation result of the cross - domain gradient foldback degree provides a quantitative basis for the difference in policy optimization between the source task and the target task for subsequent steps.

[0063] S4.2, Calculation of the distribution distance dissipation rate: The calculation of the distribution spacing dissipation rate is used to measure the dissipation degree of the source task and the target task in the feature distribution, so as to reflect the difference in the environmental representation of the two tasks. The calculation process includes two stages: First, perform fast Fourier transforms on the feature distributions of the source task and the target task respectively to generate their respective frequency domain spectra, where the frequency domain spectra represent the distributions of the feature distributions on different frequency components.

[0064] Second, calculate the energy difference between the frequency domain spectrum of the source task and the frequency domain spectrum of the target task. The specific method is to divide the norm of the difference between the two frequency domain spectra by the sum of the norms of the two frequency domain spectra, and the obtained ratio is the distribution spacing dissipation rate.

[0065] The reason for choosing frequency domain analysis is that this method can reveal the deep differences in the feature distribution in the frequency dimension, and the ratio of the energy differences provides a standardized quantification result of the dissipation degree. The calculation result of the distribution spacing dissipation rate provides a key quantitative index for adjusting the strategy to adapt to the feature of the target task environment.

[0066] S4.3, Calculation of the fitting bias coefficient: The calculation of the fitting bias coefficient generates a bias parameter for policy fusion by inputting the cross-domain gradient reversal degree and the distribution spacing dissipation rate into the parser. The processing process of the parser is as follows: First, perform a weighted sum operation on the cross-domain gradient reversal degree and the distribution spacing dissipation rate, where the weights are determined by preset parameters; Then, map the result of the weighted sum to the interval from 0 to 1 through the sigmoid function, and the obtained value is the fitting bias coefficient. The fitting bias coefficient represents the bias degree when the source task policy and the target task policy are fused.

[0067] The non-linear mapping of the sigmoid function can smoothly convert the result of the weighted sum into a fusion weight, which is convenient for precisely controlling the amplitude of the policy adjustment. The generation of the fitting bias coefficient can adaptively adjust the policy fusion ratio according to the values of the cross-domain gradient reversal degree and the distribution spacing dissipation rate, ensuring the flexibility and accuracy of the adjustment process.

[0068] S4.4, Re-weight the action probability flow: The processing of re-weighting the action probability flow uses the fitting bias coefficient to adjust the action probability flow generated in step S3 to generate a weighted action probability flow adapted to the target task. The specific method is as follows: According to the value of the fitting bias coefficient, perform a linear combination of the action probability flow of the source task and the action probability flow of the target task, where the fitting bias coefficient determines the weight ratio of the two. When the fitting bias coefficient is close to 1, the weighted action probability flow is more inclined to retain the policy characteristics of the source task; when the fitting bias coefficient is close to 0, it is more inclined to the policy characteristics of the target task.

[0069] For example, the reweighting of the action probability flow can be as follows: Using the fitting bias coefficient Perform weighted adjustment on the action probability flow generated in step S3 to generate a weighted action probability flow :

[0070] where is the action probability flow in the source task and serves as a reference benchmark.

[0071] where is the corrected action probability flow output by step S3; Control the fusion ratio. When is close to 1, it tends to retain the source task policy; when is close to 0, it tends to the target task policy.

[0072] Using linear combination can achieve smooth fusion of the source task policy and the target task policy, while the fitting bias coefficient precisely controls the degree of fusion. The generation of the weighted action probability flow can adapt to the environmental changes of the target task on the basis of retaining the source task experience, thereby enhancing the adaptability of the policy.

[0073] S4.5, encapsulate the convergence signaling packet: The process of encapsulating the convergence signaling packet is to integrate the fitting bias coefficient and the adjusted reward flow into a data packet and write it into the meta-policy buffer. The adjusted reward flow is generated by multiplying the original reward flow by a modulation factor based on the fitting bias coefficient, where the modulation factor is determined by a function of the fitting bias coefficient. The encapsulation process integrates the fitting bias coefficient and the adjusted reward flow as key information into the convergence signaling packet, which is used to provide a signal for policy adjustment in subsequent steps.

[0074] Calculate the cross-domain gradient return degree and the distribution distance dissipation rate through the consistency constraint, generate the fitting bias coefficient, reweight the action probability flow using the fitting bias coefficient, and at the same time encapsulate the fitting bias coefficient and the adjusted reward flow as a convergence signaling packet and write it into the meta-policy buffer.

[0075] The cross-domain gradient return degree and the distribution distance dissipation rate respectively quantify the differences between the source task and the target task from the perspectives of policy gradient and feature distribution. The fitting bias coefficient adaptively adjusts the policy fusion ratio based on these differences, and the encapsulation of the convergence signaling packet ensures the effective transmission of the adjustment information. By suppressing the influence of cross-domain gradient drift and feature distribution dissipation, the weighted action probability flow generated in step S4 can be consistent with the dynamic characteristics of the target task environment, thereby enhancing the robustness and adaptability of the agent's decision-making.

[0076] Steps S1 to S4 gradually optimize the feature representation, trigger boundary, action probability flow, and policy flow through cross-modal feature completion, mapping correction, hierarchical tuning, and consistency constraints, generating a convergent signaling packet containing a fitting bias coefficient and an adjusted reward flow. However, the sudden reduction in the perception dimension may still cause the agent to be unable to quickly adapt to environmental changes in the target task, manifested as lagging path reshaping and insufficient safety margin. To address this issue, step S5 introduces a meta-policy buffer and a fast adaptation scheduler, using the output of step S4 to achieve path reshaping and safety margin recovery, enhancing decision-making robustness and task reliability in migration scenarios.

[0077] Step S5 directly depends on the adjusted reward flow and fitting bias coefficient in the convergent signaling packet generated in the previous stage. The meta-policy buffer calculates the sparsity using the adjusted reward flow and compares the result with a preset sparsity threshold to trigger the fast adaptation scheduler. The fast adaptation scheduler then sorts and enables sub-task clusters based on the fitting bias coefficient, and finally drives action execution through the actuator queue. This process forms a closed-loop feedback from reward flow monitoring to action execution, ensuring path reshaping and decision optimization for the agent in scenarios with sudden reduction in the perception dimension.

[0078] Step S5 includes the following: S5.1, Monitoring of the meta-policy buffer: The meta-policy buffer is responsible for receiving and continuously monitoring the adjusted reward flow in the convergent signaling packet generated in the previous stage. The adjusted reward flow is a time series that records the sequence of reward values obtained by the agent after performing actions in the target task. To judge the sparsity degree of the reward signal, the meta-policy buffer calculates the sparsity of the adjusted reward flow and compares it with a preset sparsity threshold. The sparsity threshold is a standard value determined according to the density distribution of the reward signal in the target task, used to identify whether the reward signal is too sparse. If the sparsity is lower than the sparsity threshold, it indicates that subsequent adaptation mechanisms need to be triggered. The meta-policy buffer ensures the real-time and accuracy of sparsity calculation through continuous monitoring, providing a basis for subsequent decision-making.

[0079] S5.2, Detection of reward flow sparsity: The calculation method of the sparsity of the adjusted reward flow is to count the proportion of the number of non-zero reward values in the total number of time steps. Specifically, for an adjusted reward flow containing multiple time steps, first count the number of non-zero reward values, and then divide this number by the total number of time steps to obtain the sparsity. The sparsity represents the density of the reward signal, and the lower the value, the fewer non-zero reward values and the sparser the reward signal. When the sparsity is lower than the preset sparsity threshold, it indicates that it is difficult for the agent to effectively adapt to the environment through conventional learning methods, and additional processing mechanisms are needed to improve the adaptation efficiency. This detection process ensures an accurate judgment of the sparsity degree of the reward signal.

[0080] S5.3, Activation of the Fast Adaptation Scheduler and Enabling of Sub - task Clusters: When the sparsity of the adjusted reward stream is lower than the preset sparsity threshold, the fast adaptation scheduler is activated. The fast adaptation scheduler sorts the sub - task clusters that retain the source experience weights according to the fitting bias coefficient generated in the previous stage. A sub - task cluster is a predefined set of sub - tasks, such as "away from threat" and "towards target". Each sub - task cluster retains the experience weights in the source task, which are used to guide action selection in the target task. The fitting bias coefficient represents the degree of bias in the integration of the source task and target task policies. The larger the value, the higher the adaptability of the sub - task cluster to the target task. The fast adaptation scheduler arranges all sub - task clusters in descending order of the fitting bias coefficient and pushes the sorted sub - task clusters to the executor queue. The executor sequentially calls the sub - task clusters and their corresponding experience weights in the queue order to execute actions. This process realizes the efficient screening and priority allocation of sub - task clusters.

[0081] S5.4, Path Remodeling and Safety Margin Recovery: The executor executes actions according to the order of sub - task clusters in the executor queue. The agent realizes path remodeling and safety margin recovery by adjusting path planning and action selection. Specifically, the executor preferentially calls sub - task clusters with higher fitting bias coefficients. These sub - task clusters integrate the experience weights of the source task, enabling the agent to generate a smooth path, reducing detours and frequent corrections. At the same time, the source experience weights retained by the sub - task clusters ensure that the agent remains sensitive to potential threats, thereby maintaining safety redundancy. In this way, the agent can quickly adapt to environmental changes in the target task, improving the stability and safety of path planning.

[0082] The sparsity of the adjusted reward stream is monitored through the meta - policy buffer, and the fast adaptation scheduler is activated when the sparsity is lower than the preset sparsity threshold. Sub - task clusters that retain the source experience weights are enabled based on the fitting bias coefficient, ultimately realizing path remodeling and safety margin recovery.

[0083] In scenarios where the perception dimension suddenly decreases, sparse reward signals make it difficult for the agent to quickly adapt to the environment through conventional learning methods. Monitoring sparsity and combining source task experience can effectively make up for this deficiency. It enables the agent to quickly generate a smooth path in the target task, reducing detours and energy consumption, while remaining sensitive to threats to maintain safety redundancy, thereby improving the efficiency and reliability of task execution.

[0084] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0085] It should be noted that the system of the present invention can be deployed on the device itself to implement embedded applications, or can also run on a PC or other terminal with a user interface, so as to meet various hardware environments and usage requirements.

[0086] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0087] It should be noted that in this text, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.

[0088] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A multi-modal large model collaborative optimization method for heterogeneous task migration, characterized in that Including the steps: S1: Inject a cross-modal feature autoregressive completer at the data stream entrance containing only radar and communication telemetry. After fusing the reconstructed electro-optical and inertial navigation gap features with the original observation sequence, write them into a unified latent vector pool; S2: Use the latent vector pool to drive a double-mapping relocalizer, compare the source task hierarchical strategy trigger features in real time, redraw the trigger boundaries of the sub-tasks of moving away from threats and approaching targets, and output a mapping correction matrix; S3: Run a hierarchical weight tuner according to the mapping correction matrix, reallocate the action advantages and generate a corrected action probability stream, and send it together with the hidden state to a consistency constraint; S4: The consistency constraint calculates the fitting bias coefficient based on the cross-domain gradient backpropagation degree and the distribution spacing dissipation rate, then re-weights the action probability stream, and encapsulates the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and writes it into the meta-policy buffer; S5: When the meta-policy buffer detects that the adjusted reward stream reaches the sparse threshold, activate the fast adaptation scheduler, enable the sub-task clusters that retain the source experience weights sorted by the fitting bias coefficient and push them to the executor queue to complete path reshaping and safety margin recovery in the migration scenario.

2. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 1, wherein, Step S1 includes the following: Use a cross-modal feature autoregressive completer to reconstruct the missing optical and inertial navigation features according to the current observation sequence and historical observation sequence containing radar echoes and telemetry data; Fuse the reconstructed optical and inertial navigation features with the current observation vector to generate a complete feature vector; use an encoder to convert the complete feature vector into a latent vector; store the latent vector in a unified latent vector pool.

3. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 2, wherein Step S2 includes the following: Extract the latent vector at the current moment from the latent vector pool; use a double-mapping relocalizer to generate a mapping value for the sub-task of moving away from threats by calculating the feature direction consistency and feature amplitude difference between the latent vector and the trigger feature vector of the sub-task of moving away from threats in the source task; generate a mapping value for the sub-task of approaching targets by calculating the feature direction consistency and feature amplitude difference between the latent vector and the trigger feature vector of the sub-task of approaching targets in the source task; determine whether to trigger the sub-task of moving away from threats according to the comparison between the mapping value of the sub-task of moving away from threats and a preset threshold.

4. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 3, wherein, Step S2 also includes the following: Determine whether to trigger the sub-task of approaching targets according to the comparison between the mapping value of the sub-task of approaching targets and a preset threshold; generate a correction sub-matrix for the sub-task of moving away from threats through optimization calculation to minimize the difference between the corrected trigger feature vector of the sub-task of moving away from threats and the latent vector; generate a correction sub-matrix for the sub-task of approaching targets through optimization calculation to minimize the difference between the corrected trigger feature vector of the sub-task of approaching targets and the latent vector; combine the correction sub-matrix of the sub-task of moving away from threats and the correction sub-matrix of the sub-task of approaching targets into a mapping correction matrix.

5. The collaborative optimization method for multimodal large models for heterogeneous task migration according to claim 4, wherein, Step S3 includes the following: Extract the calibration parameters at the current moment from the mapping calibration matrix; use the hierarchical weight tuner to multiply the mapping calibration matrix by the source action advantage vector through matrix multiplication to generate the action advantage vector of the target task; use the softmax function to process the action advantage vector of the target task to generate the calibrated action probability stream; extract the hidden state at the current moment from the hidden vector pool; send the calibrated action probability stream and the hidden state into the consistency constraint for processing.

6. The collaborative optimization method for multimodal large models for heterogeneous task migration according to claim 5, characterized in that, Step S4 includes the following: Calculate the cross-domain gradient reversal degree and the distribution distance dissipation rate through the consistency constraint. The cross-domain gradient reversal degree is obtained by decomposing the source task policy gradient and the target task policy gradient into orthogonal components and coincident components and performing a time-domain convolution operation on the orthogonal components. The distribution distance dissipation rate is obtained by performing a fast Fourier transform on the source task feature distribution and the target task feature distribution to generate frequency-domain spectra and calculating the energy difference between the two frequency-domain spectra.

7. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 6, wherein Step S4 also includes the following: Use the parser to perform a weighted sum operation on the cross-domain gradient reversal degree and the distribution distance dissipation rate and map it through the sigmoid function to generate the fitting bias coefficient; linearly combine the action probability stream according to the fitting bias coefficient to generate the weighted action probability stream.

8. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 7, characterized in that Step S4 also includes the following: Multiply the original reward stream by the adjustment factor based on the fitting bias coefficient to generate the adjusted reward stream; encapsulate the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and write it into the meta-policy buffer.

9. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 8, characterized in that, Step S5 includes the following: The meta-policy buffer determines the sparsity by calculating the proportion of the number of non-zero reward values in the adjusted reward stream to the total number of time steps; compare the sparsity with the preset sparsity threshold; when the sparsity is lower than the sparsity threshold, the fast adaptation scheduler is activated; the fast adaptation scheduler sorts the sub-task clusters that retain the source experience weight according to the fitting bias coefficient.

10. The multi-modal large model collaborative optimization method for heterogeneous task migration according to claim 9, wherein, Step S5 also includes the following: Push the sorted sub-task clusters to the executor queue; the executor executes actions according to the order of the executor queue to achieve path reshaping and safety margin recovery.

Citation Information

Patent Citations

  • Application fusion system oriented to big data analysis

    CN117331995A

  • Decision-making large model-oriented multi-level heterogeneous memory collaborative scheduling method

    CN119576555A

  • DQN-based distributed computing network coordinate flow scheduling system and method

    US20240129236A1

Cited By

  • Cross-scene hidden danger migration identification method based on large model hidden space mapping

    CN121093168A

  • Multi-target layered depth optimization method for illumination control

    CN121388405A

  • A multi-objective hierarchical depth optimization method for lighting control

    CN121388405B

  • Intelligent agent collaborative navigation method based on communication opinions

    CN121594894A