Multimodal large model collaborative optimization method for heterogeneous task migration

Through cross-modal completion and mapping correction methods, the mapping misalignment problem caused by the lack of visual and inertial information during the migration process of multimodal large models is solved, which improves the migration efficiency and decision-making reliability of the agent in a perceptually restricted environment, and realizes the recovery of smooth paths and safety margins.

CN120373415BActive Publication Date: 2025-09-02NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510854571.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-02
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In an adversarial simulation environment, when migrating to a simplified scenario where only radar echo and communication telemetry are retained, the multimodal large model loses key visual and inertial information, resulting in mapping dislocation between the original sub-strategy and abstract features, affecting the agent's path evasion and target approaching actions, and reducing coordination efficiency and task reliability.

Method used

Through the five-ring closed-loop mechanism of cross-modal completion, mapping correction, hierarchical tuning, gradient weighting and fast adaptation, a triple high-frequency feedback mechanism of self-consistent information flow, decision flow, and reward flow is built. The missing features are reconstructed using the cross-modal feature autoregression complement, the dual-map relocator corrects the trigger boundary, the hierarchical weight tuner adjusts the action advantages, the consistency constraint calculates the fit bias coefficient, and the meta-strategy buffer activates the sub-task cluster to realize path reshaping and safe margin recovery.

Benefits of technology

In heterogeneous task migration, migration efficiency, decision-making reliability and long-term robustness are significantly improved, ensuring that the agent quickly adapts to and generates smooth paths in perceptually restricted environments, reducing detour distances and energy consumption peaks, and maintaining safety and redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373415B_ABST
    Figure CN120373415B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal large model collaborative optimization method for heterogeneous task migration, which specifically relates to the field of simulated confrontation. It is used to solve the problem of decreased decision-making efficiency and task execution reliability of multimodal large models in heterogeneous task migration caused by a sudden decrease in perception dimensions. It is driven in time through a five-loop closed-loop of cross-modal completion, mapping correction, hierarchical tuning, gradient weighting and rapid adaptation connected in series throughout the entire link, to construct a triple high-frequency feedback mechanism of self-consistent information flow, decision flow and reward flow; after multi-source observations are connected through a latent vector pool, the trigger boundary is mapped to the target domain in real time, and the action advantage is self-calibrated as the environment evolves; the gradient consistency constraint anchors the return and dissipation errors in the convergence domain, and the reward signal catalyzes the high-quality activation and maintenance of the experience subtask cluster through buffer scheduling; global collaboration enables the generation of a smooth path at the initial stage of migration, and the detour distance and energy consumption peak converge significantly, and the efficiency, decision reliability and long-term robustness of heterogeneous task migration are simultaneously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of simulated confrontation, and more specifically, to a multimodal large model collaborative optimization method for heterogeneous task migration. Background Art

[0002] In an adversarial simulation environment, the agent was first trained in a complex scenario containing both electro-optical observation and inertial navigation data, forming a hierarchical decision chain based on multimodal perception. This model was then migrated to a simplified scenario retaining only radar echoes and communication telemetry. This sudden reduction in perception depth resulted in a sparser representation of the previously rich environment. This compressed the feature extraction paths that each sub-strategy within the hierarchical decision chain relied on, making it difficult to match the trigger conditions for the previously constructed subtasks of moving away from the threat and approaching the target. This resulted in a delayed response to environmental changes, and frequent corrections were required for path avoidance and target approach.

[0003] The core technical difficulty arising from this situation lies in the loss of critical visual and inertial information during migration of the multimodal large model, leading to a misalignment in the mapping between the original sub-strategies and the abstract features. This misalignment first creates a signal gap at the perception layer, which then propagates upwards along the decision-making chain, weakening the high-level planning module's ability to control lower-level actions. The result is circuitous trajectory, delayed triggering of evasive maneuvers, and the exhaustion of safety margins. This makes it difficult for the intelligent agent to reliably invoke key sub-tasks in the early stages, leading to a simultaneous decline in overall collaborative efficiency and task reliability.

[0004] In order to solve the above problems, a technical solution is now provided. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a multimodal large-model collaborative optimization method for heterogeneous task migration. Through the full-link serial cross-modal completion, mapping correction, hierarchical tuning, gradient reweighting and rapid adaptation five-loop closed-loop time-driven, a self-consistent information flow, decision flow, and reward flow triple high-frequency feedback mechanism is constructed; after the multi-source observation is connected through the latent vector pool, the trigger boundary is mapped to the target domain in real time, and the action advantage is self-calibrated as the environment evolves; the gradient consistency constraint anchors the return and dissipation errors in the convergence domain, and the reward signal catalyzes the high-quality activation and maintenance of the experience subtask cluster through buffer scheduling; global collaboration enables the generation of a smooth path at the beginning of the migration, and the detour distance and energy consumption peak converge significantly, and the efficiency, decision reliability and long-term robustness of heterogeneous task migration are simultaneously improved to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] S1: Inject a cross-modal feature autoregressive completer into the data stream containing only radar and communication telemetry, fuse the reconstructed optoelectronic and inertial navigation gap features with the original observation sequence, and write them into a unified latent vector pool;

[0008] S2: Use the latent vector pool to drive the dual-mapping relocalizer, compare the source task hierarchical strategy triggering features in real time, redraw the triggering boundaries of the away-threat and toward-target subtasks, and output the mapping correction matrix;

[0009] S3: Run the layer weight tuner according to the mapping correction matrix to redistribute the action advantages and generate a corrected action probability stream, which is fed into the consistency constraint together with the hidden state;

[0010] S4: The consistency constraint calculates the fitting bias coefficient based on the cross-domain gradient foldback and the distribution spacing dissipation rate, then reweights the action probability stream, and encapsulates the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and writes it to the meta-policy buffer;

[0011] S5: When the meta-policy buffer detects that the adjusted reward stream reaches the sparse threshold, the fast adaptation scheduler is activated. Subtask clusters that retain the source experience weights are enabled by sorting the fitting bias coefficients and pushed to the executor queue to complete path reshaping and safety margin recovery in the migration scenario.

[0012] In a preferred embodiment, step S1 includes the following contents:

[0013] A cross-modal feature autoregressive completer is used to reconstruct the missing optical and inertial navigation features based on the current and historical observation sequences containing radar echoes and telemetry data. The reconstructed optical and inertial navigation features are fused with the current observation vector to generate a complete feature vector. An encoder is used to convert the complete feature vector into a latent vector, which is then stored in a unified latent vector pool.

[0014] In a preferred embodiment, step S2 includes the following contents:

[0015] The latent vector at the current moment is extracted from the latent vector pool; using the dual mapping relocator, the mapping value of the away-threat subtask is generated by calculating the feature direction consistency and feature amplitude difference between the latent vector and the feature vector triggering the away-threat subtask in the source task; the mapping value of the toward-target subtask is generated by calculating the feature direction consistency and feature amplitude difference between the latent vector and the feature vector triggering the toward-target subtask in the source task; and the mapping value of the away-threat subtask is compared with a preset threshold to determine whether to trigger the away-threat subtask.

[0016] In a preferred embodiment, step S2 further includes the following:

[0017] The mapping value of the approaching target subtask is compared with a preset threshold to determine whether to trigger the approaching target subtask; the correction submatrix of the staying away from threat subtask is generated through optimization calculation, so that the difference between the corrected triggering eigenvector and the latent vector of the staying away from threat subtask is minimized; the correction submatrix of the approaching target subtask is generated through optimization calculation, so that the difference between the corrected triggering eigenvector and the latent vector of the staying away from threat subtask is minimized; the correction submatrix of the staying away from threat subtask and the correction submatrix of the approaching target subtask are combined into a mapping correction matrix.

[0018] In a preferred embodiment, step S3 includes the following contents:

[0019] Extract the correction parameters at the current moment from the mapping correction matrix; use the hierarchical weight tuner to multiply the mapping correction matrix with the source action advantage vector through matrix multiplication to generate the action advantage vector of the target task; use the soft maximization function to process the action advantage vector of the target task to generate a corrected action probability stream; extract the hidden state at the current moment from the latent vector pool; send the corrected action probability stream and hidden state to the consistency constraint for processing.

[0020] In a preferred embodiment, step S4 includes the following contents:

[0021] The cross-domain gradient reflux and distribution spacing dissipation rate are calculated through the consistency constraint. The cross-domain gradient reflux is obtained by decomposing the source task policy gradient and the target task policy gradient into orthogonal components and overlapping components, and performing time domain convolution operations on the orthogonal components. The distribution spacing dissipation rate is obtained by performing fast Fourier transform on the source task feature distribution and the target task feature distribution to generate frequency domain spectra, and calculating the energy difference between the two frequency domain spectra.

[0022] In a preferred embodiment, step S4 further includes the following:

[0023] The parser is used to perform weighted sum operations on the cross-domain gradient reflux and the distribution spacing dissipation rate, and the fitting bias coefficient is generated through sigmoid function mapping. The action probability flow is linearly combined according to the fitting bias coefficient to generate a weighted action probability flow.

[0024] In a preferred embodiment, step S4 further includes the following:

[0025] The original reward stream is multiplied by the adjustment factor based on the fitting bias coefficient to generate an adjusted reward stream; the fitting bias coefficient and the adjusted reward stream are encapsulated into a convergence signaling packet and written into the meta-policy buffer.

[0026] In a preferred embodiment, step S5 includes the following contents:

[0027] The meta-policy buffer determines the sparsity by calculating the ratio of the number of non-zero reward values ​​in the adjusted reward stream to the total number of time steps; the sparsity is compared with a preset sparsity threshold; when the sparsity is lower than the sparsity threshold, the fast adaptation scheduler is activated; the fast adaptation scheduler sorts the subtask clusters that retain the source experience weights according to the fitting bias coefficient.

[0028] In a preferred embodiment, step S5 further includes the following:

[0029] The sorted subtask clusters are pushed to the executor queue; the executor performs actions according to the order of the executor queue to achieve path reshaping and safety margin recovery.

[0030] The technical effects and advantages of the multimodal large model collaborative optimization method for heterogeneous task migration of the present invention are as follows:

[0031] The present invention uses a five-loop closed-loop system of cross-modal completion, mapping correction, hierarchical tuning, gradient reweighting, and rapid adaptation to drive the system in real time, creating a self-consistent triple high-frequency feedback mechanism for information flow, decision flow, and reward flow. Multi-source observations are connected through a latent vector pool, triggering boundaries to be mapped to the target domain in real time, and action advantages are self-calibrated as the environment evolves. Gradient consistency constraints anchor foldback and dissipation errors within the convergence domain, and reward signals, through buffered scheduling, catalyze the high-quality activation and maintenance of experiential subtask clusters. Global coordination enables the generation of smooth paths at the beginning of migration, significantly reducing detour distances and peak energy consumption, compressing training step sizes, and maintaining long-term safety redundancy. This results in a simultaneous increase in heterogeneous task migration efficiency, decision reliability, and long-term robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flowchart of the multimodal large model collaborative optimization method for heterogeneous task migration of the present invention. DETAILED DESCRIPTION

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] Example 1: Figure 1 The present invention provides a multimodal large model collaborative optimization method for heterogeneous task migration, including:

[0035] S1: Inject a cross-modal feature autoregressive completer into the data stream containing only radar and communication telemetry, fuse the reconstructed optoelectronic and inertial navigation gap features with the original observation sequence, and write them into a unified latent vector pool.

[0036] S2: Use the latent vector pool to drive the dual mapping relocalizer, compare the source task hierarchical strategy triggering features in real time, redraw the triggering boundaries of the away-threat and toward-target subtasks, and output the mapping correction matrix.

[0037] S3: Run the layer weight tuner according to the mapping correction matrix to redistribute the action advantages and generate a corrected action probability stream, which is fed into the consistency constraint together with the hidden state.

[0038] S4: The consistency constraint calculates the fitting bias coefficient based on the cross-domain gradient retracement and the distribution spacing dissipation rate, then reweights the action probability stream, and encapsulates the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and writes it into the meta-strategy buffer.

[0039] S5: When the meta-policy buffer detects that the adjusted reward stream reaches the sparse threshold, the fast adaptation scheduler is activated. Subtask clusters that retain the source experience weights are enabled by sorting the fitting bias coefficients and pushed to the executor queue, thereby completing path reshaping and safety margin recovery in the migration scenario.

[0040] In adversarial simulation environments, intelligent agents are usually trained in complex scenarios with rich perceptual information, such as relying on multimodal data such as optoelectronic observations, inertial navigation data, radar echoes, and communication telemetry to form a stable hierarchical decision chain.

[0041] However, when the agent migrates to a simplified scenario with limited perception, such as retaining only radar echoes and communication telemetry data, the sudden reduction in perception latitude leads to a sparse representation of the environment. Sub-strategies in the original decision chain become ineffective due to feature extraction path compression, resulting in problems such as delayed path avoidance and frequent corrections during target approach. This adaptability challenge presented by heterogeneous task migration not only reduces the agent's task execution efficiency but also threatens its safety.

[0042] To address this issue, we propose a multimodal large-model collaborative optimization method for heterogeneous task migration. Through steps such as cross-modal feature completion, decision boundary correction, and action probability tuning, we enable rapid adaptation and optimization of intelligent agents in perceptually constrained environments. Step S1, the starting point of the entire solution, aims to fill perceptual gaps through a cross-modal feature autoregressive completer, constructing a unified latent vector pool and providing a complete feature representation foundation for subsequent steps.

[0043] Step S1 includes the following contents:

[0044] S1.1, Data flow entry definition:

[0045] In the simplified scenario of the target task, the perceptual data received by the agent is limited to radar echoes and communication telemetry data, excluding the optoelectronic observation data and inertial navigation data from the source task. This perceptual data is organized in a time series, referred to as a raw observation sequence. A raw observation sequence consists of observation vectors at multiple moments, each of which contains two components: a radar echo component and a communication telemetry component. Compared to the observation vectors of the source task, the observation vectors of the target task are significantly reduced in dimensionality, resulting in a sparse representation of the environment. This sparse representation of the environment reduces the agent's ability to analyze and make decisions because critical environmental information is not fully captured.

[0046] S1.2, the role and construction of cross-modal feature autoregressive complement:

[0047] To address the lack of electro-optical observation data and inertial navigation data in the target task, a cross-modal feature autoregressive complementor is introduced. The function of the cross-modal feature autoregressive complementor is to predict and reconstruct the missing electro-optical and inertial navigation features by analyzing the historical data of the original observation sequence and the observation vector at the current moment. These reconstructed features are collectively referred to as missing features.

[0048] The prediction process relies on an autoregressive mechanism, which uses observation vectors from multiple historical moments in the original observation sequence and the observation vector at the current moment to calculate estimates of missing features. The parameters of the autoregressive model are first pre-trained on the complete dataset of the source task and then fine-tuned in a simplified scenario of the target task to adapt to the characteristics of radar echoes and communication telemetry data.

[0049] Autoregressive models can identify long-term dependencies in time series data, generating accurate estimates of missing features. By reconstructing missing features, the agent obtains complete perceptual information close to the source task, enhancing its decision-making ability in simplified scenarios.

[0050] S1.3, Feature Fusion:

[0051] The missing features generated by the cross-modal feature autoregressive completer are fused with the observation vector at the current moment to generate a complete feature vector.

[0052] The fusion method is to perform a vector concatenation operation on the missing features and the current observation vector in a predetermined order, forming a complete feature vector that includes the radar echo component, the communication telemetry component, and the reconstructed electro-optical and inertial navigation features. This vector concatenation operation ensures that the complete feature vector retains the direct data of the original observation sequence while integrating the reconstructed missing features, forming a comprehensive cross-modal environmental representation.

[0053] The vector concatenation operation maintains the independence of the original and reconstructed data while providing a unified feature representation for easier processing. The complete feature vector provides the agent with comprehensive environmental information, reducing decision errors caused by missing data.

[0054] S1.4, Hidden Vector Encoding:

[0055] The complete feature vector is input into an encoder and processed to generate a latent vector.

[0056] The encoder is a neural network whose structure and parameters are pre-trained using the complete dataset of the source task. The encoder maps the high-dimensional full feature vector to a low-dimensional latent vector. Through multi-layered computations of the neural network, it extracts key information from the full feature vector and compresses it into a compact representation. Neural networks are capable of learning complex nonlinear relationships between features, generating latent representations that are highly discriminative for decision making. The latent vector retains the core information of the full feature vector in a low-dimensional form, making it easier for the agent to use in subsequent processing while reducing computational complexity.

[0057] S1.5, hidden vector pool write:

[0058] The generated latent vectors are written to a centrally managed storage structure called the latent vector pool. This pool stores all latent vectors generated at all times, forming a time series containing cross-modal information. This records the latent vectors at each moment in chronological order, ensuring that subsequent steps can access the complete latent vector sequence.

[0059] The technical feature of step S1 is that it reconstructs the missing optoelectronic and inertial navigation features through a cross-modal feature autoregressive completer, fuses them with the current observation vector to generate a complete feature vector, and then uses an encoder to convert this into a latent vector and store it in a latent vector pool. Reconstructing missing features can compensate for the lack of perceptual data in the target task, allowing the agent to obtain a comprehensive environmental representation in simplified scenarios and avoid decision-making errors caused by incomplete information.

[0060] Step S1 reconstructs the optoelectronic and inertial navigation features using a cross-modal feature autoregressive complement and generates a latent vector pool, providing a complete cross-modal representation for subsequent processing. However, feature complementation alone is insufficient to recover the decision logic of the source task. Further correction of the subtask triggering boundaries is required to adapt to the simplified environment of the target task. This is the core task of step S2.

[0061] Step S2 includes the following contents:

[0062] S2.1, hidden vector pool call:

[0063] Extract the current hidden vector from the hidden vector pool.

[0064] The latent vector pool is a storage structure that contains latent vectors at all times. Each latent vector is generated by fusing the original observation data and reconstructed features from the target task. The original observation data includes radar echoes and communication telemetry data, while the reconstructed features are supplementary information extracted from electro-optical and inertial navigation data using a cross-modal feature autoregressive completer. The latent vector at the current moment is a high-dimensional vector that fully represents the agent's cross-modal perception of the environment at that point in time. The purpose of extracting the latent vector at the current moment is to provide comprehensive environmental information for the subsequent redrawing of the trigger boundary, ensuring that the processing process is based on accurate cross-modal data.

[0065] S2.2, Definition of triggering features of source task layering strategy:

[0066] In the source task, the agent's decision-making relies on multimodal perception data, forming trigger conditions for two subtasks: the "avoid the threat" subtask and the "approach the target" subtask. These trigger conditions are composed of specific feature combinations, collectively referred to as the source task trigger feature set. The source task trigger feature set consists of two parts: the trigger feature vector for the "avoid the threat" subtask and the trigger feature vector for the "approach the target" subtask. The trigger feature vector for the "avoid the threat" subtask is composed of features such as threat range from electro-optical observation data or threat intensity from radar echo data; the trigger feature vector for the "approach the target" subtask is composed of features such as target direction from inertial navigation data or target signal strength from communication telemetry data. The purpose of defining these trigger feature vectors is to provide a benchmark for subsequent mapping and relocalization, ensuring accurate identification and correction of the subtask trigger conditions in the target task.

[0067] S2.3, Construction and Function of Dual Mapping Relocator:

[0068] The dual-mapping relocalizer consists of two independent mapping functions, one for the "moving away from the threat" subtask and the other for the "moving toward the target" subtask. Each mapping function generates a mapping value by comparing the current latent vector with the triggering feature vector of the corresponding subtask in the source task. The mapping value calculation process consists of two steps: first, calculating the feature direction consistency between the latent vector and the triggering feature vector, represented by the cosine of the angle between the vectors; and second, calculating the feature amplitude difference between the latent vector and the triggering feature vector, represented by the normalized value of the distance between the vectors. The mapping value is generated by a weighted combination of the feature direction consistency and the feature amplitude difference. The weighting parameter is adjusted based on the sparsity of the feature distribution in the target task and is determined through initial testing of the target task.

[0069] Feature direction consistency reflects the similarity between the latent vector and the trigger feature vector, while feature amplitude difference reflects their proximity. The combination of the two provides a comprehensive assessment of the match. The mapping value provides a quantitative comparison result, ensuring the accuracy and adaptability of the trigger condition remapping.

[0070] S2.4, triggering the redrawing of the border:

[0071] Based on the mapping values ​​generated by the dual-mapping relocalizer, the triggering boundaries for the "avoiding threat" and "approaching target" subtasks within the target task are determined. These triggering boundaries rely on preset thresholds: the "avoiding threat" subtask is triggered when its mapping value exceeds its corresponding threshold; the "approaching target" subtask is triggered when its mapping value exceeds its corresponding threshold. These thresholds are calibrated using empirical data from the source task and initial testing of the target task to ensure they match the dynamic characteristics of the target task's simplified environment.

[0072] The reduction in sensory data in the target task results in a shift in feature distribution, making the original trigger conditions of the source task inapplicable and requiring adjustments based on the new environmental representation. The agent can accurately identify threats and targets in a simplified environment, improving decision-making responsiveness and execution efficiency.

[0073] S2.5, Generation of mapping correction matrix:

[0074] Based on the redrawn trigger boundaries, a mapping correction matrix is ​​generated to adjust the adaptability of the source task sub-policy in the target task. This mapping correction matrix consists of two sub-matrices, one for the "avoid the threat" sub-task and the other for the "approach the target" sub-task. Each sub-matrix is ​​generated through an optimization calculation, the goal of which is to minimize the difference between the corrected source task trigger feature vector and the current latent vector. The optimization process uses an iterative adjustment method to align the corrected feature vector with the target task latent vector by gradually reducing the difference.

[0075] Matrix transformation maps the triggering features of the source task to the feature space of the target task, resolving the mapping misalignment problem caused by the reduction of the perceptual dimension. The mapping correction matrix ensures that the subtask execution is consistent with the target task environment, improving the accuracy and reliability of decision making.

[0076] The technical feature of step S2 is that it uses the cross-modal representations in the latent vector pool to drive a dual-mapping relocalizer, which compares the hierarchical policy triggering features in the source task in real time, redraws the triggering boundaries of the "avoiding threats" subtask and the "approaching target" subtask, and generates a mapping correction matrix. Because the reduced perceptual dimensions of the target task lead to a sparse feature distribution, the triggering conditions of the source task cannot be directly applied. Through mapping relocalization and boundary redrawing, the requirements of a simplified environment can be adapted. This enables the intelligent agent to accurately identify threats and targets even with reduced perceptual data, improving the efficiency and stability of task execution.

[0077] In the simplified scenario of the target task, the reduction in perceptual data requires the action selection strategy to adapt to environmental changes. Step S3 adjusts the action advantages by calling the mapping correction matrix generated in step S2. A layered weight tuner and soft maximization function are used to generate a corrected action probability stream. Meanwhile, hidden states are extracted from the latent vector pool in step S1 to provide environmental information. This process forms a complete technical chain, closely connected with the previous steps, and provides input data for the subsequent policy stream adjustment in step S4, supporting efficient decision-making in sparse environments.

[0078] Step S3 includes the following contents:

[0079] S3.1, calling of mapping correction matrix:

[0080] Extract the current correction parameters from the mapping correction matrix generated in step S2. The mapping correction matrix is ​​a structure consisting of two submatrices: one corresponding to the correction parameters for the "avoid the threat" subtask, and the other corresponding to the correction parameters for the "approach the target" subtask. These two submatrices are used to adjust the action advantage, ensuring that the action selection is consistent with the characteristic distribution of the target task. The purpose of calling the mapping correction matrix is ​​to provide a correction basis for subsequent adjustments to the action advantage, enabling the action selection strategy to match the dynamic environment of the target task, thereby improving the adaptability of the decision.

[0081] S3.2, Construction and Function of Hierarchical Weight Tuner:

[0082] The Hierarchical Weight Tuner is a dynamic adjustment module responsible for redistributing action advantages based on the mapping correction matrix. Action advantages represent the agent's preference for various actions in a specific state and reflect the priority of the actions. In the source task, action advantages are generated by the source policy network and are a vector containing multiple action advantage values. The Hierarchical Weight Tuner generates the action advantage vector in the target task by performing a linear transformation on the source action advantage vector. The specific method of linear transformation is to perform a matrix multiplication operation on the mapping correction matrix and the source action advantage vector. That is, the correction parameters are applied to the source action advantage vector through matrix multiplication to obtain the adjusted action advantage vector.

[0083] Using matrix multiplication, the mapping correction matrix directly applies to the action advantage, adapting it to the reduced perceptual dimensionality of the target task. The adjustment of the action advantage, based on the correction parameters of the mapping correction matrix, ensures that the action selection strategy is aligned with the target task environment, thereby improving decision accuracy.

[0084] S3.3, Generation of action probability stream:

[0085] Based on the corrected action advantage vector, an action probability stream is generated. The action probability stream represents the probability distribution of each action selected by the agent in the current state. The action probability stream is generated by processing the corrected action advantage vector using a soft maximization function. The specific calculation process is as follows: first, the exponential value of each advantage value in the corrected action advantage vector is calculated. Then, the exponential value of each advantage value is divided by the sum of the exponential values ​​of all advantage values ​​to obtain the probability value of each action, thus forming the action probability stream. The soft maximization function can normalize the advantage values ​​in the action advantage vector into a probability distribution, ensuring that the sum of the probability values ​​of all actions is 1, which facilitates the agent to select actions based on the probability distribution.

[0086] Soft maximization functions are widely used in the field of reinforcement learning. They can effectively balance the exploration and exploitation of action selection, thereby improving the stability of the strategy.

[0087] S3.4, extraction of hidden state:

[0088] The current hidden state is extracted from the latent vector pool generated in step S1. The hidden state is a vector that represents an abstract representation of the current environment, integrating reconstructed electro-optical and inertial navigation features with raw radar and telemetry data. The purpose of extracting the hidden state is to provide environmental information for subsequent consistency constraints, ensuring that policy adjustments are consistent with environmental dynamics. The hidden state extraction process directly relies on the latent vector pool and is completed by selecting the vector corresponding to the current moment from the latent vector pool, ensuring the completeness and accuracy of the environmental representation.

[0089] S3.5, input consistency constraint:

[0090] The corrected action probability stream and hidden state are fed into the consistency constraint as input. In subsequent steps, the consistency constraint will calculate the fitting bias coefficient based on the action probability stream and hidden state, further adjusting the policy stream to suppress cross-domain gradient drift. The purpose of feeding the consistency constraint is to provide the necessary data support for the final adjustment of the policy stream, ensuring that the agent's decisions are consistent with the environmental characteristics of the target task, thereby optimizing the generation process of the policy stream.

[0091] A hierarchical weight tuner is driven by a mapping correction matrix. The source action advantage vector is adjusted through matrix multiplication to generate the target task action advantage vector. A soft maximization function is then used to generate a corrected action probability stream, which is then fed into a consistency constraint along with the hidden state. This action probability stream can be adaptively adjusted based on the dynamic environment of the target task, thereby improving the agent's decision-making accuracy and environmental adaptability in heterogeneous task transfer.

[0092] In scenarios where the perceptual dimension is reduced, the action probability stream and hidden state generated in step S3 may become invalid due to misalignment of environmental features. Step S4, using a consistency constraint, calculates the cross-domain gradient foldback and distribution spacing dissipation rate based on the action probability stream and hidden state in step S3, generates a fitting bias coefficient, and reweights the action probability stream. It also encapsulates the convergence signaling packet and writes it to the meta-policy buffer. This process effectively corrects the deviation between the policy stream and environmental features, providing accurate input data for the subsequent rapid adaptive scheduling in step S5, supporting the agent's path reshaping and decision optimization in scenarios where the perceptual dimension is suddenly reduced.

[0093] The consistency constraint ensures consistency and adaptability of model decisions between the source and target tasks by calculating the cross-domain gradient reflux and the distribution gap dissipation rate. Starting from the output of the previous stage, the layer-wise weight tuner adjusts the action advantages according to the mapping correction matrix, generating a rectified action probability stream and hidden state. These outputs provide the basis for the calculation of the consistency constraint. The consistency constraint first uses the rectified action probability stream and hidden state to analyze the policy gradients of the source and target tasks. By capturing the magnitude of the gradient's reverse retracement during the transfer process, it calculates the cross-domain gradient reflux. This metric reflects the differences and potential inconsistencies in policy adjustments between tasks. Furthermore, the consistency constraint quantifies the degree of dissipation in the environment representation by comparing the frequency domain spectra of the feature distributions of the two tasks, thereby calculating the distribution gap dissipation rate, which reveals the evolution of the feature space between tasks. These two metrics are then input into the parser to generate the fitting bias coefficient. This coefficient linearly combines the policy features of the source and target tasks to adjust the action probability stream so that it maintains consistency with the source task policy in the target task while adapting to new environmental characteristics. Finally, the fitted bias coefficients and the adjusted reward stream are packaged into a convergence signaling packet and written to the meta-policy buffer, providing a signal for subsequent policy adjustments. Through this process, the consistency constraint effectively connects the output of the previous stage, ensuring consistency between the policy stream and the dynamic characteristics of the target task environment, improving the model's decision-making robustness and adaptability in scenarios with a sudden decrease in perceptual dimensionality.

[0094] Step S4 includes the following contents:

[0095] S4.1, calculation of cross-domain gradient reentry:

[0096] The calculation of cross-domain gradient retracement aims to quantify the difference in the reverse retrace amplitude of the policy gradient between the source task and the target task, which is used to reflect the intensity of gradient drift. The calculation process is divided into two stages:

[0097] First, the policy gradient of the source task and the policy gradient of the target task are decomposed into orthogonal components and overlapping components, where the orthogonal components represent the differences in the directions of the two task policy gradients, while the overlapping components represent the commonalities in the directions of the two task policy gradients.

[0098] Secondly, within a preset time window, a time-domain convolution operation is performed on the orthogonal components of the source task and the orthogonal components of the target task, and the cross-domain gradient retrieval degree is obtained by integrating the product of the orthogonal components of the two over time.

[0099] Temporal convolution is used because it can capture the dynamic changes in gradients over time, accurately reflecting the degree of gradient drift. The calculation of cross-domain gradient retracement provides a quantitative basis for the optimization differences between the source and target task strategies in subsequent steps.

[0100] S4.2, Calculation of the Dissipation Rate for Distribution Spacing:

[0101] The calculation of the distribution gap dissipation rate is used to measure the degree of dispersion in the feature distribution of the source task and the target task, reflecting the difference in the environmental representation of the two tasks. The calculation process includes two stages:

[0102] First, the feature distribution of the source task and the feature distribution of the target task are respectively subjected to fast Fourier transform to generate their respective frequency domain spectra, where the frequency domain spectrum represents the distribution of the feature distribution on different frequency components.

[0103] Secondly, the energy difference between the frequency domain spectrum of the source task and the frequency domain spectrum of the target task is calculated. The specific method is to divide the norm of the difference between the two frequency domain spectra by the sum of the norms of the two frequency domain spectra. The resulting ratio is the distribution spacing dissipation rate.

[0104] Frequency domain analysis was chosen because it reveals deep differences in the frequency dimension of feature distributions, while the ratio of energy differences provides a standardized quantification of the degree of dissipation. The calculated dissipation ratio of the distribution spacing provides a key quantitative indicator for adjusting strategies to the characteristics of the target mission environment.

[0105] S4.3, calculation of fitting bias coefficient:

[0106] The calculation of the fitting bias coefficient is to input the cross-domain gradient foldback and distribution spacing dissipation rate into the parser to generate the bias parameter for policy fusion. The parser process is as follows:

[0107] First, the cross-domain gradient reentry degree and the distribution spacing dissipation rate are weighted and calculated, where the weight is determined by the preset parameters;

[0108] The weighted sum is then mapped to a range of 0 to 1 using a sigmoid function. The resulting value is the fitting bias coefficient, which indicates the degree of bias between the source and target task policies when they are fused.

[0109] The nonlinear mapping of the sigmoid function smoothly converts the weighted sum into fusion weights, facilitating precise control of the magnitude of policy adjustments. The generation of the fitted bias coefficients adaptively adjusts the policy fusion ratio based on the values ​​of the cross-domain gradient foldback and the distribution spacing dissipation rate, ensuring flexibility and accuracy in the adjustment process.

[0110] S4.4, reweighted action probability flow:

[0111] The process of re-weighting the action probability stream is to adjust the action probability stream generated in step S3 using the fitting bias coefficient to generate a weighted action probability stream that is suitable for the target task. The specific method is:

[0112] Based on the value of the fitting bias coefficient, a linear combination of the source task's action probability stream and the target task's action probability stream is performed. The fitting bias coefficient determines the weight ratio between the two. When the fitting bias coefficient is close to 1, the weighted action probability stream is more likely to retain the policy characteristics of the source task; when the fitting bias coefficient is close to 0, it is more likely to retain the policy characteristics of the target task.

[0113] For example, the reweighted action probability stream can be as follows:

[0114] Using the fitted bias coefficient The action probability stream generated in step S3 Perform weighted adjustment to generate weighted action probability flow :

[0115]

[0116] in, is the action probability stream in the source task, which serves as a reference benchmark.

[0117] in is the corrected action probability stream output from step S3; Control the fusion ratio, when When it is close to 1, it tends to retain the source task strategy; when When it is close to 0, it tends to be a target task strategy.

[0118] The use of linear combinations enables smooth fusion of the source and target task policies, while the fitting bias coefficient precisely controls the degree of fusion. The generation of a weighted action probability stream can adapt to environmental changes in the target task while retaining the experience of the source task, thereby improving the adaptability of the policy.

[0119] S4.5, encapsulates the convergence signaling packet:

[0120] The process of encapsulating the convergence signaling packet combines the fitted bias coefficients and the adjusted reward stream into a single packet and writes it to the meta-policy buffer. The adjusted reward stream is generated by multiplying the original reward stream by an adjustment factor based on the fitted bias coefficients, where the adjustment factor is determined as a function of the fitted bias coefficients. The encapsulation process integrates the fitted bias coefficients and the adjusted reward stream as key information into the convergence signaling packet, which provides a signal for policy adjustments in subsequent steps.

[0121] The cross-domain gradient reflux and distribution spacing dissipation rate are calculated through the consistency constraint to generate the fitting bias coefficient, and the fitting bias coefficient is used to reweight the action probability flow. At the same time, the fitting bias coefficient and the adjusted reward flow are encapsulated as a convergence signaling packet and written into the meta-strategy buffer.

[0122] The cross-domain gradient retracement and distribution gap dissipation rate quantify the differences between the source and target tasks from the perspectives of policy gradients and feature distribution, respectively. The fitting bias coefficient adaptively adjusts the policy fusion ratio based on these differences, while the encapsulation of the convergence signaling packet ensures the effective transmission of this adjustment information. By suppressing the effects of cross-domain gradient drift and feature distribution dissipation, the weighted action probability stream generated in step S4 is able to align with the dynamic characteristics of the target environment, thereby improving the robustness and adaptability of the agent's decision-making.

[0123] Steps S1 to S4 progressively optimize the feature representation, triggering boundaries, action probability stream, and policy stream through cross-modal feature completion, mapping correction, hierarchical tuning, and consistency constraints, generating a converged signaling packet containing fitted bias coefficients and an adjusted reward stream. However, the sudden reduction in perceptual dimensionality may still prevent the agent from rapidly adapting to environmental changes during the target task, manifesting as delayed path reshaping and insufficient safety margin. To address this issue, step S5 introduces a meta-policy buffer and a fast adaptation scheduler, leveraging the output of step S4 to implement path reshaping and safety margin recovery, improving decision robustness and task reliability in transfer scenarios.

[0124] Step S5 directly relies on the adjusted reward stream and fitting bias coefficients in the convergence signaling packet generated in the previous stage. The meta-policy buffer uses the adjusted reward stream to calculate sparsity and compares the result with a preset sparsity threshold to trigger the fast adaptation scheduler. The fast adaptation scheduler then sorts and activates subtask clusters based on the fitting bias coefficients, ultimately driving action execution through the executor queue. This process forms a closed-loop feedback loop from reward stream monitoring to action execution, ensuring the agent's path reshaping and decision optimization in scenarios with a sudden decrease in perceptual dimensionality.

[0125] Step S5 includes the following contents:

[0126] S5.1, Monitoring of the Meta-Policy Buffer:

[0127] The meta-policy buffer is responsible for receiving and continuously monitoring the adjusted reward stream from the convergence signaling packets generated in the previous stage. The adjusted reward stream is a time series that records the sequence of reward values ​​obtained by the agent after performing actions in the target task. To determine the sparsity of the reward signal, the meta-policy buffer calculates the sparsity of the adjusted reward stream and compares it with a preset sparsity threshold. The sparsity threshold is a standard value determined based on the density distribution of the reward signal in the target task and is used to identify whether the reward signal is too sparse. If the sparsity falls below the sparsity threshold, it indicates that the subsequent adaptation mechanism should be triggered. Through continuous monitoring, the meta-policy buffer ensures the real-time and accuracy of the sparsity calculation, providing a basis for subsequent decision-making.

[0128] S5.2, reward flow sparsity detection:

[0129] The sparsity of the adjusted reward stream is calculated by counting the proportion of non-zero reward values ​​in the stream to the total number of time steps. Specifically, for an adjusted reward stream containing multiple time steps, the number of non-zero reward values ​​is first counted, and then this number is divided by the total number of time steps to determine the sparsity. Sparsity indicates the density of the reward signal; lower values ​​indicate fewer non-zero reward values ​​and a sparser reward signal. When the sparsity falls below a preset sparsity threshold, it indicates that the agent is struggling to effectively adapt to the environment through conventional learning methods and requires additional processing mechanisms to improve adaptation efficiency. This detection process ensures accurate judgment of the sparsity of the reward signal.

[0130] S5.3, Activation of the fast adaptation scheduler and enabling of subtask clusters:

[0131] When the sparsity of the adjusted reward stream falls below a preset sparsity threshold, the fast adaptive scheduler is activated. The fast adaptive scheduler sorts subtask clusters that retain the source experience weights based on the fitting bias coefficient generated in the previous stage. Subtask clusters are predefined sets of subtasks, such as "avoid threat" and "approach target." Each subtask cluster retains the experience weights from the source task to guide action selection in the target task. The fitting bias coefficient indicates the degree of bias in the convergence of the source and target task policies. A larger value indicates a more compatible subtask cluster with the target task. The fast adaptive scheduler sorts all subtask clusters in descending order of their fitting bias coefficients and pushes the sorted subtask clusters to the executor queue. The executor sequentially calls the subtask clusters and their corresponding experience weights according to the queue order to execute actions. This process enables efficient screening and priority assignment of subtask clusters.

[0132] S5.4, Path Reshaping and Safety Margin Recovery:

[0133] The executor executes actions sequentially based on the subtask clusters in the executor queue. The agent adjusts its path planning and action selection to achieve path reshaping and restore safety margins. Specifically, the executor prioritizes subtask clusters with higher fitting bias coefficients. These subtask clusters incorporate the empirical weights of the source tasks, enabling the agent to generate smooth paths, reduce detours, and reduce frequent deviations. Furthermore, the retained source empirical weights of the subtask clusters ensure that the agent remains sensitive to potential threats, thereby maintaining safety redundancy. In this way, the agent can quickly adapt to environmental changes during the target task, improving the stability and safety of path planning.

[0134] The sparsity of the adjusted reward stream is monitored through the meta-policy buffer, and the fast adaptation scheduler is activated when the sparsity is lower than the preset sparsity threshold. The subtask clusters that retain the source experience weights are enabled according to the fitting bias coefficient, ultimately achieving path reshaping and safety margin recovery.

[0135] In scenarios where the perceptual dimensionality decreases dramatically, the sparseness of reward signals makes it difficult for the agent to quickly adapt to the environment through conventional learning methods. However, monitoring sparsity and integrating it with the source task experience can effectively compensate for this deficiency. This allows the agent to quickly generate a smooth path in the target task, reducing detours and energy consumption, while maintaining threat sensitivity to maintain safety redundancy, thereby improving the efficiency and reliability of task execution.

[0136] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0137] It should be noted that the system of the present invention can be deployed on the device itself to realize embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting a variety of hardware environments and usage requirements.

[0138] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.

[0139] It should be noted that, in this document, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "includes" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0140] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A multimodal large model collaborative optimization method for heterogeneous task migration, characterized by: Including steps: S1: Inject a cross-modal feature autoregressive completer into the data stream containing only radar and communication telemetry, fuse the reconstructed optoelectronic and inertial navigation gap features with the original observation sequence, and write them into a unified latent vector pool; S2: Use the latent vector pool to drive the dual-mapping relocalizer, compare the source task hierarchical strategy triggering features in real time, redraw the triggering boundaries of the away-threat and toward-target subtasks, and output the mapping correction matrix; Step S2 includes the following contents: Extract the current hidden vector from the hidden vector pool; use the dual mapping relocator to generate the mapping value of the away-threat subtask by calculating the feature direction consistency and feature amplitude difference between the hidden vector and the feature vector triggered by the away-threat subtask in the source task; and generate the mapping value of the toward-target subtask by calculating the feature direction consistency and feature amplitude difference between the latent vector and the feature vector triggered by the toward-target subtask in the source task. Determine whether to trigger the stay-away-threat subtask based on the comparison of the mapping value of the stay-away-threat subtask with a preset threshold; Determine whether to trigger the target subtask based on the comparison of the mapping value of the target subtask with the preset threshold; The correction matrix of the threat-avoidance subtask is generated by optimizing calculation, so that the difference between the corrected eigenvector and the latent vector of the threat-avoidance subtask is minimized; the correction matrix of the target-approaching subtask is generated by optimizing calculation, so that the difference between the corrected eigenvector and the latent vector of the target-approaching subtask is minimized; the correction matrix of the threat-avoidance subtask and the correction matrix of the target-approaching subtask are combined into a mapping correction matrix; S3: Run the layer weight tuner according to the mapping correction matrix to redistribute the action advantages and generate a corrected action probability stream, which is fed into the consistency constraint together with the hidden state; S4: The consistency constraint calculates the fitting bias coefficient based on the cross-domain gradient foldback and the distribution spacing dissipation rate, then reweights the action probability stream, and encapsulates the fitting bias coefficient and the adjusted reward stream into a convergence signaling packet and writes it to the meta-policy buffer; S5: When the meta-policy buffer detects that the adjusted reward stream reaches the sparse threshold, the fast adaptation scheduler is activated. Subtask clusters that retain the source experience weights are enabled by sorting the fitting bias coefficients and pushed to the executor queue to complete path reshaping and safety margin recovery in the migration scenario.

2. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 1 is characterized in that: Step S1 includes the following contents: Using a cross-modal feature autoregressive completer, missing optical and inertial navigation features are reconstructed from the current and historical observation sequences containing radar echoes and telemetry data. The reconstructed optical and inertial navigation features are fused with the current observation vector to generate a complete feature vector; an encoder is used to convert the complete feature vector into a latent vector; and the latent vector is stored in a unified latent vector pool.

3. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 1 is characterized in that: Step S3 includes the following contents: Extract the correction parameters at the current moment from the mapping correction matrix; use the hierarchical weight tuner to multiply the mapping correction matrix with the source action advantage vector through matrix multiplication to generate the action advantage vector of the target task; use the soft maximization function to process the action advantage vector of the target task to generate a corrected action probability stream; extract the hidden state at the current moment from the latent vector pool; send the corrected action probability stream and hidden state to the consistency constraint for processing.

4. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 3 is characterized in that: Step S4 includes the following contents: The cross-domain gradient reflux and distribution spacing dissipation rate are calculated through the consistency constraint. The cross-domain gradient reflux is obtained by decomposing the source task policy gradient and the target task policy gradient into orthogonal components and overlapping components, and performing time domain convolution operations on the orthogonal components. The distribution spacing dissipation rate is obtained by performing fast Fourier transform on the source task feature distribution and the target task feature distribution to generate frequency domain spectra, and calculating the energy difference between the two frequency domain spectra.

5. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 4 is characterized in that: Step S4 also includes the following: The parser is used to perform weighted sum operations on the cross-domain gradient reflux and the distribution spacing dissipation rate, and the fitting bias coefficient is generated through sigmoid function mapping. The action probability flow is linearly combined according to the fitting bias coefficient to generate a weighted action probability flow.

6. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 5 is characterized in that: Step S4 also includes the following: The original reward stream is multiplied by the adjustment factor based on the fitting bias coefficient to generate an adjusted reward stream; the fitting bias coefficient and the adjusted reward stream are encapsulated into a convergence signaling packet and written into the meta-policy buffer.

7. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 6 is characterized in that: Step S5 includes the following contents: The meta-policy buffer determines the sparsity by calculating the ratio of the number of non-zero reward values ​​in the adjusted reward stream to the total number of time steps; the sparsity is compared with a preset sparsity threshold; when the sparsity is lower than the sparsity threshold, the fast adaptation scheduler is activated; the fast adaptation scheduler sorts the subtask clusters that retain the source experience weights according to the fitting bias coefficient.

8. The multimodal large model collaborative optimization method for heterogeneous task migration according to claim 7 is characterized in that: Step S5 also includes the following: The sorted subtask clusters are pushed to the executor queue; the executor performs actions according to the order of the executor queue to achieve path reshaping and safety margin recovery.

Citation Information

Patent Citations

  • Application fusion system oriented to big data analysis

    CN117331995A

  • Decision-making large model-oriented multi-level heterogeneous memory collaborative scheduling method

    CN119576555A