A multi-modal large model optimization method, device and medium for complex tasks
Patent Information
- Application Number
- CN202511338790.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-09-18
AI Technical Summary
[0005]本说明书一个或多个实施例提供了一种用于复杂任务的多模态大模型优化方法、设备及介质,用于解决如下技术问题:因此,在现有多模态大模型执行复杂任务的过程中,因模型结构与参数配置静态固化以及训练策略对数据质量与任务难度动态变化响应滞后,存在计算效率与任务性能难以兼顾以及复杂场景下语义理解与推理精度不足的问题
[0013] The above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: By using the technical solutions of the embodiments of this specification, by analyzing multi-dimensional complexity indicators such as image resolution, number of objects, spatial layout, text length, semantic ambiguity, audio duration and event density in real time, and combining them with specific task intentions, such as visual question answering requiring fine spatial understanding and audio generation requiring high-fidelity temporal modeling, dynamic parameter adjustment strategies are generated. Based on the dynamic parameter adjustment strategies, the model parameters are adjusted, and a significant improvement in overall throughput and response speed is achieved under given hardware resources, enabling the model to actively adapt to various complex and changing input scenarios.
Smart Images

Figure CN121390274B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of multimodal large model technology, and in particular to a method, device and medium for multimodal large model optimization for complex tasks. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence technology, multimodal large models have shown broad application prospects in many fields due to their ability to simultaneously process and understand information from multiple modalities such as images, text, and audio. Existing multimodal models typically employ fixed network architectures and parameter configurations for training and inference, and their general framework generally consists of a modal encoder, an input projection layer, and a large language model (LLM). Although these models perform well on routine tasks, their performance remains significantly insufficient when handling complex multimodal tasks (such as scenarios requiring fine spatial relation reasoning, long-term action understanding, or cross-modal semantic disambiguation).
[0003] Existing technologies suffer from several prominent problems. First, the model structure lacks adaptability, failing to dynamically adjust according to the inherent complexity of the input data and specific task requirements. This leads to wasted computational resources on simple tasks and performance degradation on complex tasks due to insufficient capabilities. Second, in the multimodal feature fusion stage, existing methods often employ static or late-stage fusion strategies, making it difficult to achieve deep, adaptive interaction between modalities. This can easily result in modal bias (such as one modality excessively dominating decision-making) or loss of fusion information, hindering the model's accurate capture and utilization of cross-modal correlations. Furthermore, traditional training methods face the challenge of scarce and unevenly distributed high-quality, diverse multimodal data, and fail to effectively combine phased capability enhancement with dynamic course learning mechanisms, resulting in poor model generalization ability and robustness in complex scenarios.
[0004] Therefore, in the process of performing complex tasks using existing multimodal large models, due to the static and fixed model structure and parameter configuration, as well as the lag in the response of training strategies to dynamic changes in data quality and task difficulty, there are problems such as difficulty in balancing computational efficiency and task performance, and insufficient semantic understanding and reasoning accuracy in complex scenarios. Summary of the Invention
[0005] This specification provides one or more embodiments of a multimodal large model optimization method, device, and medium for complex tasks, which are used to solve the following technical problems: Therefore, in the process of existing multimodal large models performing complex tasks, due to the static fixation of model structure and parameter configuration and the lag in the response of training strategies to dynamic changes in data quality and task difficulty, there are problems such as difficulty in balancing computational efficiency and task performance, as well as insufficient semantic understanding and inference accuracy in complex scenarios.
[0006] One or more embodiments of this specification employ the following technical solutions:
[0007] This specification provides one or more embodiments of a multimodal large model optimization method for complex tasks. The method includes: receiving multimodal input data; extracting modality-specific features from the multimodal input data to generate an initial feature representation for each modality, wherein the multimodal input data includes image data, text data, and audio data; generating a dynamic parameter adjustment strategy based on a pre-acquired task type and input data complexity; adjusting the computational parameters of the multimodal encoder in the multimodal large model according to the dynamic parameter adjustment strategy to determine the corresponding optimized modality processing submodule; performing cross-modal fusion through the optimized modality processing submodule and the initial feature representation to generate a fused multimodal feature representation; and using a pre-trained multimodal large model to perform semantic parsing and inference on the multimodal feature representation to generate task output results.
[0008] This specification provides one or more embodiments of a multimodal large model optimization apparatus for complex tasks, comprising:
[0009] At least one processor; and,
[0010] A memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described method.
[0012] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions configured to perform the above-described method.
[0013] The above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: By using the technical solutions of the embodiments of this specification, by analyzing multi-dimensional complexity indicators such as image resolution, number of objects, spatial layout, text length, semantic ambiguity, audio duration and event density in real time, and combining them with specific task intentions, such as visual question answering requiring fine spatial understanding and audio generation requiring high-fidelity temporal modeling, dynamic parameter adjustment strategies are generated. Based on the dynamic parameter adjustment strategies, the model parameters are adjusted, and a significant improvement in overall throughput and response speed is achieved under given hardware resources, enabling the model to actively adapt to various complex and changing input scenarios. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0015] Figure 1 A flowchart illustrating a multimodal large model optimization method for complex tasks, provided as an embodiment of this specification;
[0016] Figure 2 This is a schematic diagram of the structure of a multimodal large model optimization device for complex tasks, provided as an embodiment of this specification. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0018] This specification provides a multimodal large model optimization method for complex tasks. It should be noted that the execution entity in this specification can be a server or any device with data processing capabilities. Figure 1 This specification provides a flowchart illustrating a multimodal large model optimization method for complex tasks, as illustrated in the embodiments of this specification. Figure 1 As shown, the main steps include the following:
[0019] Step S101: Receive multimodal input data, extract modality-specific features from the multimodal input data, and generate an initial feature representation for each modality.
[0020] The multimodal input data includes image data, text data, and audio data;
[0021] Modality-specific feature extraction is performed on the multimodal input data to generate an initial feature representation for each modality. Specifically, this includes: segmenting the image data into multiple image blocks, extracting the visual feature vector of each image block, and introducing enhanced two-dimensional rotational position coding to embed spatial position information to determine the initial feature representation corresponding to the image modality; performing word segmentation and embedding on the text data, using a knowledge-enhanced semantic decoder for context encoding, and calling image or audio features for disambiguation when ambiguity is detected to determine the initial feature representation corresponding to the text modality; and performing wavelet transform on the audio data to generate a time-frequency map, using temporal attention to capture temporal dynamic features, using spectral attention to focus on key frequency bands, and generating dual-channel feature representations of voiceprint and semantics to determine the initial feature representation corresponding to the audio modality.
[0022] In one embodiment of this specification, compared to traditional multimodal general frameworks, which include modal encoders, input projection layers, and large language models (LLMs), the multimodal large model framework in this embodiment mainly consists of two parts: a multimodal encoder and an LLM. The multimodal encoder encodes multimodal data into a vector feature space and maps its output to the LLM input feature space. The LLM primarily uses Qwen-3 to initialize model parameters. In this embodiment, the language model uses Qwen3 for parameter initialization, preserving its powerful language understanding and generation capabilities. Through feature mapping by the multimodal encoder, the conversion from multimodal input to the language space is achieved, supporting tasks such as cross-modal dialogue and reasoning.
[0023] A unified modal encoder module was designed to simultaneously process data from multiple modalities, including images, text, and audio. The unified modal encoder employs a hierarchical structure, with each layer performing feature extraction and fusion for data from different modalities.
[0024] An image modality processing branch, or image modality processing submodule, is constructed using an improved VisionTransformer (ViT) structure. Based on ViT, a dynamic block-splitting strategy is implemented. Specifically, the size and number of blocks are dynamically adjusted according to the resolution and content complexity of the input image. For example, for high-resolution and content-rich images, the block size is appropriately reduced while the number of blocks is increased to improve fine-grained feature extraction capabilities; for low-resolution or simple images, the block size is increased while the number of blocks is reduced to decrease computational redundancy. Simultaneously, an attention mechanism is introduced into the inter-layer connections of ViT to enhance the model's focus on key regions of the image.
[0025] In the image modality processing submodule, ViT is upgraded to a cross-modal guided dynamic block segmentation network. It not only segments images based on their own features but also receives guidance signals from text keywords or audio semantics in real time. It prioritizes ultra-fine segmentation of associated regions, with a minimum granularity of 4×4 pixels, while adaptively merging unassociated regions, achieving an intelligent balance between accuracy and efficiency. Furthermore, a dynamic resolution strategy is introduced, employing NaViT's Patch n'Pack technique to package multiple patches from different images into a single sequence, preserving the variable resolution of different images. Simultaneously, multiple images can be processed within a single subsequence computation, improving the model's computational throughput. Compressed codes are generated for non-critical regions, and valid patches are marked using a mask matrix, skipping invalid computations during inference. A multi-scale feature cache pool is designed to reuse repeated features across frames / images. Enhanced 2D Rotated Position Vectors (RoPE) are introduced in the height and width dimensions, a relative position attenuation factor is introduced to reduce the association weight of distant regions, and spatial semantic encoding (binding orientation labels) is added to strengthen spatial relationship understanding. For extreme scenarios such as occlusion and similar objects, an edge enhancement module is added. The contour features of the object are extracted through Canny edge detection, and then concatenated with ViT features and input into the fusion layer to improve the target discrimination.
[0026] A text modality processing submodule is constructed using a Transformer-based bidirectional encoder. More hidden layers and attention heads are added to the encoder to enhance the understanding of semantic information and the ability to capture context. For example, the traditional 6-layer Transformer encoder is expanded to 12 layers, and the number of attention heads is increased from 8 to 16. In this way, the model can better handle long texts and complex semantic relationships. The text modality processing submodule introduces a knowledge-enhanced semantic decoder on top of the Transformer, embedding a lightweight knowledge graph to bind abstract concepts in the text with entity knowledge; it also adds a cross-modal disambiguation unit, which automatically calls image or audio features to assist in calibrating the semantic direction when there is ambiguity in the text (e.g., whether "apple" refers to a fruit or a brand).
[0027] An audio modality processing submodule is established, employing a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). First, CNNs are used to perform preliminary feature extraction on the audio signal, capturing local features. For example, multiple convolutional kernels of different sizes are designed to extract features from different frequency ranges of the audio. Then, the features extracted by the CNN are input into an RNN, such as a Long Short-Term Memory (LSTM) network, utilizing the RNN's time-series processing capabilities to extract temporal features of the audio. Finally, the audio features are fused with image and text features in an attention-based fusion layer. In the fusion layer, attention weights are calculated between features of different modalities, and the features are weighted and summed according to these weights to achieve effective integration of multimodal features.
[0028] The audio modality processing submodule abandons traditional convolutional and recurrent architectures, employing a temporal spectral dual-attention network. It first uses wavelet transform to decompose the audio into multi-scale time-frequency maps, then focuses on key frequency bands through spectral attention, capturing emotional cues such as pauses and intonation changes. Simultaneously, it incorporates dual-channel encoding of voiceprint semantics, preserving original audio features while generating abstract representations aligned with text semantics. Furthermore, the temporal spectral attention network incorporates adaptive environmental noise filtering, using voiceprint features to separate the target audio from background noise, and uses a sentiment polarity analyzer to quantify audio emotions (such as anger and joy) into vectors aligned with text sentiment labels.
[0029] Step S102: Based on the pre-acquired task type and input data complexity, a dynamic parameter adjustment strategy is generated to adjust the calculation parameters of the multimodal encoder in the multimodal large model according to the dynamic parameter adjustment strategy, so as to determine the corresponding optimized modal processing submodule.
[0030] Based on the pre-acquired task type and input data complexity, a dynamic parameter adjustment strategy is generated. Specifically, this includes: analyzing the resolution of the image data, the number of target objects in the scene, and the spatial layout complexity to determine the image complexity index; analyzing the length of the text data, the semantic structure complexity, and the existence of ambiguous entities to determine the text complexity index; analyzing the duration, spectral variation range, and event density of the audio data to determine the audio complexity index; receiving the user-specified task type, which includes visual question answering, audio-text generation, or cross-modal reasoning tasks; and generating a dynamic parameter adjustment strategy based on the image complexity index, text complexity index, audio complexity index, and task type using a pre-defined strategy generator. This dynamic parameter adjustment strategy includes adjustment instructions for the computational accuracy, block granularity, context length, and fusion weights of each modal processing submodule in the multimodal encoder. According to the dynamic parameter adjustment strategy, the computational parameters of each modality processing submodule in the multimodal encoder are adjusted. Specifically, in the image processing submodule, image blocks are dynamically divided according to the image block granularity adjustment instruction in the dynamic parameter adjustment strategy, and non-critical regions are merged. The floating-point precision mode is switched according to the computational precision adjustment instruction. In the text processing submodule, the maximum sequence length of the Transformer encoder is adjusted according to the context length adjustment instruction in the dynamic parameter adjustment strategy. The knowledge enhancement decoder is activated or disabled according to the semantic complexity instruction. In the audio processing submodule, the scale parameter of the wavelet transform is adjusted according to the spectral analysis granularity instruction in the dynamic parameter adjustment strategy. The environmental noise filtering module is enabled or disabled according to the event density instruction. In the fusion module, the weight ratio of each modality in the attention mechanism is dynamically set according to the fusion weight adjustment instruction in the dynamic parameter adjustment strategy.
[0031] In one embodiment of this specification, a dynamic parameter adjustment module is introduced into the multimodal large model. This module can automatically adjust the model parameters based on the complexity of the input data and the difficulty of the task. For example, when processing simple image recognition tasks, the computational load of parameters in certain complex layers of the model is reduced to improve inference speed; when facing complex multimodal inference tasks, such as those requiring the integration of image, text, and audio information for decision-making, the number of parameters and computational accuracy of relevant layers are increased to enhance the model's processing capabilities. Through this dynamic adjustment mechanism, the adaptability and efficiency of the model in different scenarios are improved while ensuring model performance. Specifically, a parameter adjustment control unit is designed, which adjusts the model parameters based on the complexity of the input data and the difficulty index of the task. The complexity and difficulty index are calculated as follows: for image data, factors such as image resolution, number of objects, and scene complexity are analyzed; for text data, text length, vocabulary complexity, and semantic structure complexity are considered; for audio data, the duration of the audio, frequency variation range, and complexity of audio events are evaluated. For example, when the image resolution is higher than a certain threshold and the number of objects is large, it is judged as a complex image. The model parameters are dynamically adjusted based on the calculated indexes. When the task is simple, such as a simple image classification task, the amount of parameter computation can be reduced by decreasing the number of neurons in certain layers or reducing the computational precision. For example, the number of neurons in the fully connected layers of the model can be reduced by half, and the computational precision can be reduced from 32-bit floating-point numbers to 16-bit floating-point numbers. When the task is complex, such as a multimodal inference task, the number of parameters and computational precision of relevant layers can be increased. For example, in the multimodal fusion layer, additional neurons and weight matrices can be added, while the computational precision can be increased to 64-bit floating-point numbers to improve the model's processing power.
[0032] In real-world applications, the data composition and computational requirements of multimodal tasks are highly dynamic. The computational resources and processing granularity required for a low-resolution image with a simple background are drastically different from those required for a high-resolution image containing numerous fine-grained targets and occlusion relationships. Similarly, the demands on model analysis and filtering capabilities differ significantly between a short, clear speech command and a long audio stream filled with environmental noise and multiple semantic events. Using models with fixed parameters and structures to handle the most complex scenarios results in massive waste of computational resources for simple tasks, while small models optimized for efficiency are inadequate for challenging tasks. This severely limits the practical deployment efficiency and widespread application of these models.
[0033] In one embodiment of this specification, the received image, text, and audio data are first parsed in parallel to extract key complexity features. For image data, its original resolution is analyzed; higher resolution generally means more pixel information and richer details, requiring more refined processing. Simultaneously, an object detection algorithm is run to count the number of identifiable objects in the scene; a larger number indicates more complex spatial relationships and context. Furthermore, the spatial layout complexity is comprehensively evaluated by analyzing the relative positions between the bounding boxes of detected objects, their overlap relationships (such as the degree of occlusion), and the regularity or disorder of the overall scene layout, ultimately generating a comprehensive image complexity index. For text data, the length of its token sequence is first counted; long texts typically require a longer context window to maintain semantic coherence. Then, natural language understanding technology is used to analyze the nesting depth of syntactic structures, the number of clauses, and the degree of rhetorical device usage to determine semantic structure complexity. In addition, by querying the built-in knowledge base or entity linking tool, the presence of ambiguous entities in the text, such as polysemous words, ambiguous pronoun references, or domain-specific terms, is detected. These entities, if not disambiguated, will seriously affect subsequent understanding, thus generating a text complexity index. For audio data, calculate its total duration, as longer audio contains more information. By performing short-time Fourier transform or similar time-frequency analysis on the audio signal, observe the intensity and range of its spectrum changes on the time axis. The greater and more frequent the changes, the higher the analysis difficulty. At the same time, use an audio event detection model to identify and count the density of distinct events (such as speech conversion, musical fluctuations, and sound effects) in the audio stream. The higher the event density, the more complex the audio content. Finally, synthesize an audio complexity index.
[0034] While performing a quantitative assessment of the complexity of multimodal data, the system receives task types explicitly specified by the user or preset by the upstream application system based on the scenario. For example, visual question answering tasks require deep fusion of image content and question text, audio-text generation tasks require accurate conversion of audio information into text, and cross-modal reasoning tasks require logical inference by integrating information from all modalities. Different tasks have significantly different emphases on modalities and resource allocation requirements. Subsequently, all indicators are input into a preset policy generator, which is essentially a lightweight decision model trained offline. Based on the "data feature-task type-optimal configuration" mapping relationship learned from a large amount of historical training data, a comprehensive decision is made, and a dynamic parameter adjustment strategy is output. This strategy includes a series of executable adjustment instructions. For the image processing submodule, the instructions specify its block granularity (e.g., using finer-grained blocks for complex images to preserve details, and coarser-grained blocks for simple images to improve efficiency) and computational precision (e.g., switching between FP32, FP16, or BFLOAT16 floating-point formats). For the text processing submodule, the instructions specify the maximum sequence length that the Transformer encoder must process (to adapt to long texts) and whether to activate the knowledge-enhancing decoder (activating only when ambiguous entities are detected to save computation). For the audio processing submodule, the instructions specify the wavelet transform scaling parameters (to control the granularity of time-frequency analysis) and the on / off state of the environmental noise filtering module (enabled in high-noise scenarios). For the fusion module, the instructions explicitly set the initial weight ratios of image, text, and audio features in the attention mechanism to guide the model to focus on the key modalities of the task.
[0035] Finally, the above strategy is translated into specific adjustments to the computational parameters of each submodule of the multimodal encoder. The image processing submodule dynamically calls image segmentation algorithms to perform non-uniform grid partitioning of the input image based on block granularity instructions, further subdividing related regions (such as regions guided by text keywords) into ultra-small blocks, merging non-critical regions, and switching the numerical precision mode of the underlying tensor operations according to precision instructions. The text processing submodule allocates corresponding memory buffers and sets attention masks according to context length instructions to ensure effective processing of long sequences, and bypasses or connects to the knowledge-enhancing decoder subnetwork in the inference pipeline according to semantic complexity instructions. The audio processing submodule configures the wavelet transform parameters to generate time-frequency maps of corresponding precision based on spectral analysis granularity instructions, and controls whether the noise filtering module preprocesses the original audio signal according to event density instructions. The fusion module reads the fusion weight instructions and uses them as the initial bias when weighting and fusing query, key, and value vectors in attention calculation, thereby dynamically adjusting the contribution of each modality in the information flow. Through the above operations, the internal computing resource configuration of the model was precisely aligned with the specific data and task requirements, laying a solid foundation for subsequent efficient and high-precision cross-modal fusion and inference.
[0036] Traditional models must process all inputs according to their design limits, regardless of their simplicity. This leads to severe computational redundancy and energy waste on many simple tasks. Through dynamic perception and adjustment, the model automatically reduces computational intensity and adopts more economical processing strategies when facing simple data and undemanding tasks, thereby significantly saving computing resources and shortening response time. When faced with highly complex data and challenging tasks, it can instantly mobilize sufficient computing power to ensure processing accuracy and reliability. Dynamic block partitioning and fine-grained processing enhance the model's ability to understand occlusion, small objects, and complex spatial relationships in images. By activating the knowledge-enhanced decoder and cross-modal disambiguation, text ambiguity problems are effectively solved. Environmental noise filtering and dual attention mechanisms strengthen the model's ability to extract audio information in noisy environments, enabling the model to cope with complex and ever-changing application scenarios.
[0037] Step S103: By optimizing the modality processing submodule and the initial feature representation, cross-modal fusion is performed to generate a fused multimodal feature representation. Then, a pre-trained multimodal large model is used to perform semantic parsing and reasoning on the multimodal feature representation to generate the task output result.
[0038] The optimized modality processing submodule and the initial feature representation are used to perform cross-modal fusion to generate a fused multimodal feature representation. Specifically, this includes: using a modality probe unit to generate an association anchor point for the initial feature representation of each modality, wherein the association anchor point is used to identify the semantic correlation between features; using a contrastive learning fusion layer, the similarity between different modal features is calculated using the association anchor point, and based on the similarity, similar semantic features are clustered in the vector space; a dynamic weighted attention mechanism is used to perform weighted fusion according to the pre-determined weights of the modal features to generate an initial fused multimodal feature representation; and the initial fused multimodal feature representation is adversarially trained using a modality discriminator to eliminate modality bias and retain modality specificity, and the fused multimodal feature representation is output.
[0039] In one embodiment of this specification, the fusion mechanism employs an innovative progressive inter-modal fusion module instead of a single high-level fusion. A modality probe unit is set at the bottom layer, with associated anchor points attached to each modality feature during generation. Then, through contrastive learning of the fusion layer, similar semantics of different modalities are forced to cluster in the feature space. Finally, dynamic weighted attention is used to adjust the weights of each modality in real time according to task requirements, ultimately outputting fused features that possess both modality specificity and global relevance. Furthermore, based on the progressive inter-modal fusion, a cross-modal adversarial calibration module is added, introducing a modality discriminator. Through adversarial training, it forces each modality feature to retain specificity while eliminating modality bias during fusion (e.g., avoiding excessive dominance of text modality in image semantic judgment).
[0040] The method further includes: aggregating raw data from an open-source multimodal dataset, wherein the raw data includes image-text pairs, audio-text pairs, and multimodal interleaved data; calculating image-text pair similarity using a cross-modal similarity model and retaining samples higher than a preset threshold; oversampling low-frequency concepts and undersampling high-frequency concepts to construct a training dataset; and using the training dataset to perform multi-stage training on a large multimodal model to obtain a large multimodal model that meets preset requirements, wherein the multi-stage training includes a pre-training stage and a post-training stage.
[0041] In one embodiment of this specification, a rigorous data cleaning and filtering process is constructed when determining the training data. For multimodal interleaved data, such as web page data, the original content from large-scale open-source datasets is first aggregated, followed by multi-stage cleaning and filtering. Image classifiers and semantic analysis tools are used to discard images irrelevant to the semantics of the article content and remove noisy elements such as advertisements and QR codes. For image-text pairing data, advanced similarity calculation tools such as the CLIP model are used to calculate the similarity between images and text, retaining only pairings with similarity higher than a specific threshold (e.g., 0.3). Simultaneously, a concept-balanced resampling strategy is employed to resample the data, reducing bias and ensuring data balance and diversity. For image-text pairing data, the CLIP model is used to calculate the similarity between images and text. Specifically, images and text are input into the image encoder and text encoder of the CLIP model, respectively, to obtain corresponding feature vectors, and then the cosine similarity between the two feature vectors is calculated. Only pairings with similarity higher than 0.3 are retained. Simultaneously, a concept-balanced resampling strategy is employed to oversample concepts with low frequency in the dataset and undersample concepts with excessive frequency. For example, by statistically analyzing the frequency of each concept in the dataset, for concepts whose frequency is below a certain threshold, their corresponding image-text pairing data is copied to increase their proportion in the dataset; for concepts whose frequency is above a certain threshold, some pairing data is randomly deleted to reduce bias in the data and ensure data balance and diversity.
[0042] Targeted data augmentation is performed for data of different modalities. For image data, traditional data augmentation methods such as rotation, scaling, cropping, and adding noise are used, while Generative Adversarial Networks (GANs) are introduced to generate new image data that is similar to the original data but has certain differences, thus expanding the scale and diversity of the image dataset. Specifically, traditional data augmentation methods such as rotation, scaling, cropping, and adding noise are used. For example, images are randomly rotated from 0 to 360 degrees, scaled between 0.5 and 1.5, different regions of the image are cropped, and Gaussian noise is added to the image. Simultaneously, GAN technology is used to construct a generator and a discriminator. The generator is given random noise as input to generate new images similar to the original image, while the discriminator determines whether the generated image is real or generated. By continuously training the generator and discriminator, the generator can generate high-quality new image data, expanding the scale and diversity of the image dataset.
[0043] For text data, data augmentation methods such as synonym replacement, sentence structure transformation, and text summarization are used to increase the variability of the text data. Specifically, synonym replacement tools, such as WordNet, are used to replace words in the text with synonyms. For example, "beautiful" can be replaced with synonyms such as "pretty" or "charming." Simultaneously, sentence structure can be changed, such as changing active sentences to passive sentences, or text summarization can be performed to generate different versions of the text data, increasing its variability. For audio data, new audio samples are generated by changing parameters such as volume, speech rate, and pitch, enriching the audio dataset. Specifically, data augmentation is performed by changing parameters such as volume, speech rate, and pitch. For example, the audio volume can be randomly increased or decreased within a certain range, the speech rate can be adjusted between 0.8 and 1.2 times, and the pitch can be raised or lowered within a certain frequency range. These operations generate new audio samples, enriching the audio dataset. Through data augmentation and expansion, the generalization ability of the model is improved, enabling it to better cope with various data variations in real-world applications.
[0044] The pre-training phase includes a ViT training phase, a joint pre-training phase, and a joint long context activation phase. The ViT training phase includes a first training phase, a second training phase, and a third training phase. In the first training phase, the parameters in the multimodal encoder, except for the adapter layer, are frozen, and the adapter layer is trained using clean multimodal data to initially align modal representations. In the second training phase, all model parameters are unfrozen, noisy data and multiple loss functions are introduced, and the large multimodal model is reinforced through training on large-scale multimodal data. In the third training phase, data from video, programming, and 3D understanding domains are introduced, and the sequence length is expanded. Cross-domain capability transfer is performed through the transfer adapter module. In the joint pre-training phase, the proportion of multimodal data in the training data is increased, and the feature space mapping is strengthened through the multimodal alignment loss function. In the joint long context activation phase, the model context length is expanded in two stages, long and short sequence data are mixed for training, and multimodal question-answer pairs are synthesized to accelerate long context learning. The post-training phase includes a joint supervised fine-tuning SFT sub-phase, a long thought chain SFT sub-phase, and a reinforcement learning and curriculum sampling combined sub-phase. The joint supervised fine-tuning SFT sub-phase uses a mixture of plain text SFT data and visual-language SFT data to simultaneously fine-tune the model, employing a dynamic weight loss function to balance language loss and multimodal alignment loss. The long thought chain SFT sub-phase constructs a high-quality long inference path dataset containing planning, evaluation, reflection, and exploration steps, freezes the preset parameters of the multimodal encoder and language model, fine-tunes only the adapter layer and multimodal fusion layer of the language model, and uses logical coherence loss to constrain the semantic associations between inference steps. The reinforcement learning and curriculum sampling combined sub-phase constructs a task-related reward function, calculates reward values based on model output, and updates model parameters using a policy gradient method to maximize cumulative rewards. The difficulty of the training data is dynamically adjusted based on the model training progress, gradually transitioning from simple to complex samples according to the curriculum sampling strategy.
[0045] In one embodiment of this specification, the training method is improved, mainly including a pre-training stage and a post-training stage. The pre-training stage begins with the ViT training stage. The ViT training stage employs a phased, progressive training method. In the first stage, a dynamic module selector is introduced to automatically filter training modules based on the strength of text-image semantic association. Accurate alignment from global to local levels is achieved through coarse-grained and fine-grained dual-track comparative learning. Simultaneously, a hard negative sample generator is enabled to enhance alignment stability. After starting with clean data, noisy data is gradually mixed in as model accuracy improves, initially aligning the representations of different modalities, enabling the model to initially understand the correspondence between multimodal data. In the second stage, a knowledge distillation mechanism is embedded when unfreezing all parameters. The module-specialized model trained in the first stage serves as the teacher to constrain the key layer outputs. Simultaneously, a capability node unlocking mechanism is adopted (e.g., OCR capability is decomposed into a hierarchical target of text detection → font recognition → semantic error correction). On large-scale multimodal data, the model's basic capabilities such as accumulation of various multimodal knowledge, visual localization, and OCR are specifically strengthened. In the third stage, new domain data such as video, programming, and 3D understanding are added to a more balanced data mix. Dedicated transfer adapters (such as dynamic temporal convolution modules for video) are provided for cross-domain learning, and a progressive sequence expansion strategy (from 512...) is employed. → The model uses 4096 tokens and a sparse attention + memory caching mechanism to process long sequences. It also incorporates a task self-decomposer to improve the planning ability in complex scenarios. At the same time, it optimizes the training weights through an inter-stage capability calibration feedback loop to further improve the model's processing ability in complex tasks.
[0046] The next step is the joint pre-training phase. The multimodal encoder is jointly trained with the LLM initialized with Qwen-3. This achieves accurate mapping between multimodal features and the LLM input feature space while preserving the LLM's native language capabilities, thus enhancing cross-modal semantic conversion capabilities. A hybrid strategy of plain text and multimodal data is employed. Plain text data is sampled from corpora consistent with the Qwen-3 training distribution (e.g., books, web page text, dialogue data); multimodal data encompasses image-text pairings, audio-text associations, and image-audio-text interleaved data. The learning rate scheduler of the original LLM checkpoint (e.g., cosine annealing scheduler) is used, accumulating an additional 1.4T tokens. In the initial training phase (the first 20% of steps), only plain text data is input to maintain the basic capabilities of the language model; subsequently, the proportion of multimodal data is gradually increased from 10% to 60%, strengthening cross-modal fusion through multimodal alignment loss (constraining the feature similarity between the encoder output and the LLM input).
[0047] Finally, in the joint long context activation stage, during the final pre-training phase, the model context length is expanded from 8K to 128K to improve the understanding of long texts and long multimodal data, while maintaining the accuracy of short context processing. The inverse frequency of RoPE (Rotation Position Encoding) vectorization is reset from 50,000 to 800,000 to enhance the accuracy of long-distance position information encoding; the context length is expanded in two stages, the first stage from 8K... → 32K, Phase Two from 32K → The dataset is 128K in length, with each stage achieving a smooth transition through a fourfold length expansion. Each sub-stage mixes 25% long data and 75% short data. Long data includes long text (e.g., reports of tens of thousands of words), long multimodal data (long interleaved data such as multi-page illustrated manuals, long videos such as 10-minute instructional video frame sequences, and long documents such as PDF papers); short data reuses the 8K / 32K length data from the previous stage. A small number of multimodal QA pairs (e.g., temporal reasoning questions based on long videos, summary question-and-answer questions based on long documents) are synthesized simultaneously to accelerate long-context learning.
[0048] After the pre-training phase, the post-training phase is conducted, which includes a joint supervised fine-tuning SFT sub-stage, a long thought chain SFT sub-stage, and a reinforcement learning and curriculum sampling combination sub-stage. In the joint supervised fine-tuning SFT sub-stage, the base model is optimized through instruction-based fine-tuning to enhance its ability to follow instructions and participate in dialogue, while balancing the collaborative performance of the language model and the multimodal encoder. A hybrid set of plain text SFT data and visual-language SFT data is used. The plain text SFT data includes instruction-following data (e.g., "Write a short essay about AI") and dialogue data (e.g., small talk, knowledge-based question answering); the visual-language SFT data covers multimodal instructions, such as "Describe the spatial relationships of objects in a picture" and "Judge emotion based on audio and generate a response." The LLM and multimodal encoder parameters are simultaneously fine-tuned using the AdamW optimizer (learning rate 5e-5). A dynamic weight loss function is used, with language loss accounting for 40% and multimodal alignment loss accounting for 60%, ensuring stable model responses in both plain text dialogues (e.g., small talk) and multimodal scenarios (e.g., image-based question answering).
[0049] The Long-CoT SFT sub-stage enhances the model's ability to reason about long logical chains in complex tasks, enabling it to generate detailed and coherent reasoning paths, such as multi-step spatial relationship reasoning and long video action temporal analysis. A high-quality dataset of 10K characters is constructed through cue word engineering, containing validation reasoning paths for text input (e.g., the steps to solve Huarong Road puzzles) and image input (e.g., analyzing object occlusion relationships in an image). Each path encompasses core human cognitive processes: planning (e.g., "first step: locate the target position"), evaluation (e.g., "whether this step ignored occlusions"), reflection (e.g., "adjust the step order to avoid conflicts"), and exploration (e.g., "try alternative solutions"). The reasoning paths are manually labeled and cross-modal consistent, such as text reasoning and image feature matching. A lightweight SFT is employed, freezing most parameters of the multimodal encoder and LLM, with only fine-tuning of the LLM's adapter layer and multimodal fusion layer (learning rate 1e-5). By constraining the semantic relationships between reasoning steps through logical coherence loss, the Long-CoT reasoning process generated by the model is made closer to human thinking, improving the performance of complex tasks such as maze navigation and long video action prediction.
[0050] The reinforcement learning and curriculum sampling sub-stage introduces a reinforcement learning mechanism combined with a curriculum sampling strategy during training. In the reinforcement learning phase, a task-related reward system is constructed, where the model receives rewards or penalties based on its output performance in the actual task. For example, in a visual localization task, a higher reward is given if the model accurately identifies the location of a target object and completes the corresponding operation; a penalty is given if the identification is incorrect or the operation fails. By continuously adjusting model parameters to maximize rewards, the model's performance in complex tasks is improved. The curriculum sampling strategy dynamically adjusts the difficulty of the training data based on the model's training progress and capability improvement. In the early stages of training, simple, basic multimodal data samples are provided. As training progresses, more complex and challenging data samples are gradually introduced, enabling the model to gradually adapt to tasks of varying difficulty, improving training stability and efficiency. The deep integration of reinforcement learning and curriculum sampling achieves a more intelligent training loop.
[0051] Conventional training methods typically employ a static, one-size-fits-all strategy, mixing all data for end-to-end training. This lacks fine-grained stage divisions and targeted capability enhancement, resulting in low training efficiency and uneven model capability development. The embodiments in this specification effectively solve these problems through a phased, progressive training strategy. The first stage focuses on modality alignment, avoiding premature introduction of noise interference. The second stage unfreezes all parameters and introduces large-scale data and multiple loss functions, enabling the model to systematically consolidate and expand its basic capabilities, avoiding capability biases that may result from a single loss function. The third stage introduces new domain data and long sequence processing mechanisms, equipped with a dedicated adapter, allowing the model's capabilities to be safely and efficiently extended to more complex domains, avoiding catastrophic forgetting and training instability. This progressive training model ensures solid, comprehensive, and balanced development of model capabilities, significantly improving the stability of the training process and the maturity of the final model.
[0052] Conventional methods often begin training directly with a high proportion of multimodal data, which can easily compromise the inherent strong language capabilities of the language model. The embodiments in this specification, by maintaining pure text training in the early stages, solidify the model's language core. Subsequently, multimodal data is gradually introduced and increased, supplemented by cross-modal alignment loss, ensuring that multimodal features can be smoothly and accurately mapped to the language space. This not only preserves the original advantages of LLM to the greatest extent but also endows it with powerful cross-modal semantic transformation capabilities, achieving an effective combination of strong language foundation and strong multimodal understanding.
[0053] Conventional models typically have a fixed context length, making them ineffective for handling real-world applications such as long documents and videos. The embodiments in this specification expand the context length gradually in two stages, employing a strategy of mixing long and short data. This allows the model to smoothly adapt to longer sequences, avoiding training instability and positional encoding distortion problems that can arise from direct expansion. Measures such as resetting the RoPE frequency and synthesizing multimodal QA pairs further enhance the model's ability to accurately model long-distance dependencies and its practicality for reasoning with long contexts, enabling it to handle complex tasks such as long video summarization and academic literature understanding.
[0054] Finally, the integration of joint SFT, Long-CoT SFT, reinforcement learning, and curriculum sampling in the post-training phase significantly enhances the model's instruction compliance, complex reasoning ability, and generalization ability. Joint SFT, through a dynamic weight loss function, cleverly balances the model's performance under plain text and multimodal instructions, making it a versatile dialogue assistant. Long-CoT SFT, through high-quality reasoning path data and logical coherence loss, specifically refines the model's thought process for long-chain, multi-step reasoning, making its reasoning more closely resemble human logic and significantly improving its performance on complex planning and reasoning tasks. The combination of reinforcement learning and curriculum sampling introduces a dynamic and adaptive training environment, enabling the model to learn from easy to difficult according to its own learning progress and continuously optimize its behavioral strategies on the target task using rewards as signals, ultimately achieving excellent generalization and robustness in complex real-world scenarios.
[0055] The technical solutions in the embodiments of this specification analyze multi-dimensional complexity indicators such as image resolution, number of objects, spatial layout, text length, semantic ambiguity, audio duration, and event density in real time, and combine them with specific task intentions, such as visual question answering requiring fine spatial understanding and audio generation requiring high-fidelity temporal modeling. Dynamic parameter adjustment strategies are generated, and model parameters are adjusted based on dynamic parameter adjustment strategies. Under given hardware resources, a significant improvement in overall throughput and response speed is achieved, enabling the model to actively adapt to various complex and changing input scenarios.
[0056] This specification also provides an embodiment of a multimodal large model optimization device for complex tasks, such as... Figure 2 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method.
[0057] This specification also provides a non-volatile computer storage medium storing computer-executable instructions configured to perform the above-described method.
[0058] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0059] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0060] The devices, media, and methods provided in the embodiments of this specification are one-to-one correspondences. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0061] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0062] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for optimizing a multimodal large model for complex tasks, characterized in that, The method includes: Receive multimodal input data, extract modality-specific features from the multimodal input data, and generate an initial feature representation for each modality, wherein the multimodal input data includes image data, text data, and audio data; Based on the pre-acquired task type and input data complexity, a dynamic parameter adjustment strategy is generated to adjust the calculation parameters of the multimodal encoder in the multimodal large model according to the dynamic parameter adjustment strategy, so as to determine the corresponding optimized modal processing submodule. Through the optimized modality processing submodule and the initial feature representation, cross-modal fusion is performed to generate a fused multimodal feature representation. Then, a pre-trained multimodal large model is used to perform semantic parsing and inference on the multimodal feature representation to generate task output results. Based on the pre-obtained task type and input data complexity, a dynamic parameter adjustment strategy is generated, which includes: The resolution, number of target objects in the scene, and spatial layout complexity of the image data are analyzed to determine the image complexity index. Analyze the length, semantic structure complexity, and presence of ambiguous entities in the text data to determine the text complexity index; Analyze the duration, spectral variation range, and event density of the audio data to determine the audio complexity index; Receive a task type specified by the user, wherein the task type includes a visual question answering task, an audio text generation task, or a cross-modal reasoning task; Based on the image complexity index, the text complexity index, the audio complexity index, and the task type, a dynamic parameter adjustment strategy is generated through a preset strategy generator. This dynamic parameter adjustment strategy includes adjustment instructions for the computational accuracy, block granularity, context length, and fusion weights of each modal processing submodule in the multimodal encoder. According to the dynamic parameter adjustment strategy, the calculation parameters of each modal processing submodule in the multimodal encoder are adjusted, specifically including: In the image processing submodule, the image blocks are dynamically divided according to the image block granularity adjustment instruction in the dynamic parameter adjustment strategy, and non-critical areas are merged. The floating-point precision mode is switched according to the calculation precision adjustment instruction. In the text processing submodule, the maximum sequence length of the Transformer encoder is adjusted according to the context length adjustment instruction in the dynamic parameter adjustment strategy, and the knowledge enhancement decoder is activated or disabled according to the semantic complexity instruction. In the audio processing submodule, the scale parameter of wavelet transform is adjusted according to the spectral analysis granularity instruction in the dynamic parameter adjustment strategy, and the environmental noise filtering module is enabled or disabled according to the event density instruction. In the fusion module, the weight ratio of each modality in the attention mechanism is dynamically set according to the fusion weight adjustment instruction in the dynamic parameter adjustment strategy.
2. The multimodal large model optimization method for complex tasks according to claim 1, characterized in that, Cross-modal fusion is performed using the optimized modality processing submodule and the initial feature representation to generate a fused multimodal feature representation, specifically including: The modality probe unit generates association anchors for the initial feature representation of each modality, wherein the association anchors are used to identify semantic correlations between features; By contrastive learning fusion layer, the similarity between different modal features is calculated using the associated anchor points, and based on the similarity, similar semantic features are clustered in vector space; A dynamic weighted attention mechanism is adopted to perform weighted fusion based on the pre-determined weights of modal features to generate an initial fused multimodal feature representation; The initial fused multimodal feature representation is adversarially trained using a modality discriminator to eliminate modality bias and retain modality specificity, and the fused multimodal feature representation is then output.
3. The multimodal large model optimization method for complex tasks according to claim 1, characterized in that, Modality-specific feature extraction is performed on the multimodal input data to generate an initial feature representation for each modality, specifically including: The image data is divided into blocks to obtain multiple image blocks. The visual feature vector of each image block is extracted, and an enhanced two-dimensional rotational position code is introduced to embed spatial position information to determine the initial feature representation corresponding to the image modality. The text data is segmented and embedded, and context encoding is performed using a knowledge-enhanced semantic decoder. When ambiguity is detected, image features or audio features are called for disambiguation to determine the initial feature representation corresponding to the text modality. The audio data is subjected to wavelet transform to generate a time-frequency map. Temporal attention is used to capture time dynamic features, spectral attention is used to focus on key frequency bands, and dual-channel feature representations of voiceprint and semantics are generated to determine the initial feature representation corresponding to the audio modality.
4. The multimodal large model optimization method for complex tasks according to claim 1, characterized in that, The method further includes: The raw data is aggregated from an open-source multimodal dataset, wherein the raw data includes image-text pairs, audio-text pairs, and multimodal interleaved data; The cross-modal similarity model is used to calculate the similarity between image and text pairs, and samples with similarity above a preset threshold are retained. Low-frequency concepts are oversampled, and high-frequency concepts are undersampled to construct the training dataset; The multimodal large model is trained in multiple stages using the training dataset to obtain a multimodal large model that meets the preset requirements. The multi-stage training includes a pre-training stage and a post-training stage.
5. The multimodal large model optimization method for complex tasks according to claim 4, characterized in that, The pre-training phase includes a ViT training phase, a joint pre-training phase, and a joint long context activation phase. The ViT training phase includes a first training phase, a second training phase, and a third training phase. In the first training phase, the parameters in the multimodal encoder except for the adapter layer are frozen, and the adapter layer is trained using clean multimodal data to initially align the modal representations. In the second training phase, all model parameters are unfrozen, noisy data and multiple loss functions are introduced, and the multimodal large model is reinforced and trained on large-scale multimodal data. In the third training phase, data from the video, programming, and 3D understanding domains are introduced, and the sequence length is extended to enable cross-domain capability transfer through the transfer adapter module. In the joint pre-training phase, the proportion of multimodal data in the training data is increased, and the feature space mapping is enhanced through a multimodal alignment loss function; In the joint long context activation phase, the model context length is expanded in two stages, long and short sequence data are mixed for training, and multimodal question-answer pairs are synthesized to accelerate long context learning.
6. The multimodal large model optimization method for complex tasks according to claim 5, characterized in that, The post-training phase includes a joint supervised fine-tuning SFT sub-phase, a long mind chain SFT sub-phase, and a sub-phase combining reinforcement learning and curriculum sampling. The joint supervised fine-tuning SFT sub-stage uses mixed plain text SFT data and visual-language SFT data to simultaneously fine-tune the model, and adopts a dynamic weight loss function to balance language loss and multimodal alignment loss. The long thought chain SFT sub-stage constructs a high-quality long reasoning path dataset that includes planning, evaluation, reflection, and exploration steps, freezes the preset parameters of the multimodal encoder and language model, fine-tunes only the adapter layer and multimodal fusion layer of the language model, and uses logical coherence loss to constrain the semantic association between reasoning steps. The reinforcement learning and course sampling sub-stages construct a task-related reward function, calculate the reward value based on the model output, and update the model parameters through the policy gradient method to maximize the cumulative reward. The difficulty of the training data is dynamically adjusted based on the model training progress, and the course sampling strategy is used to gradually transition from simple samples to complex samples.
7. A multimodal large model optimization device for complex tasks, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-6.
8. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are configured to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-modal task processing and dialogue task processing method, system and equipment
CN117033585A
Method and device for dynamically generating multi-modal hybrid expert model
CN118865409A