Mixed precision model fusion training method and system for multi-modal large model
Patent Information
- Application Number
- CN202611285947.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]因此,本发明提供了一种面向多模态大模型的混合精度模型融合训练方法解决机器学习领域多模态大模型混合精度训练中分区精度门控难以自适应的问题
[0063] The beneficial effects of this invention are as follows: By normalizing training state parameters and comparing accuracy sensitivity, the activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value are incorporated into the same accuracy sensitivity determination process. This enables modal coding partitions, cross-modal projection partitions, and fusion training partitions to be configured with corresponding computational accuracy according to the training state, enhancing the partition adaptability, training state correlation, and accuracy control stability of the mixed accuracy configuration of multimodal large models. By forming a unified gating basis through partition comparison results, bias gain gating values, fusion weights, and accuracy backoff markers, high-precision calibration bias, feature output ratio, and computational accuracy backoff path are collaboratively involved in the control within the same machine learning training chain, improving the consistency of feature expression, parameter update stability, and accuracy maintenance capability of multimodal large model fusion training.
Smart Images

Figure CN122778060A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a method and system for training mixed-precision models for multimodal large models. Background Technology
[0002] With the development of machine learning, large-scale pre-trained models, and multimodal representation learning, large multimodal models are gradually taking on complex tasks such as cross-modal retrieval, visual question answering, speech understanding, and video semantic analysis. Related training techniques typically revolve around the unified encoding, cross-modal projection, and fusion prediction of text, image, audio, and video samples. Mixed-precision training, which combines low-bit-width computation with high-precision calibration, can reduce training memory usage, communication overhead, and computational load, and has become an important supporting method for large-scale model training.
[0003] From the perspective of engineering implementation and training mechanism, there is still room for improvement in the mixed-precision training technology for multimodal large models, mainly in two aspects: First, the configuration of computational precision largely relies on fixed-precision strategies, hierarchical empirical rules, and unified operator templates, making it difficult to combine activation overflow, cross-modal alignment, loss variation, gradient fluctuation, and parameter update offset to reflect the precision sensitivity differences of different training partitions; Second, the mixed-precision output lacks bias evaluation and benefit gating based on high-precision calibration, resulting in a lack of unified and coordinated constraints on feature fusion weights, precision backoff marking, and parameter update processes, which affects the stability of fusion training. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a mixed-precision model fusion training method for multimodal large models to solve the problem of difficult adaptive partition precision gating in mixed-precision training of multimodal large models in the field of machine learning.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a hybrid precision model fusion training method for multimodal large models, comprising,
[0008] Obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partition training inputs;
[0009] Short-range pre-training is performed on the partitioned training input, training state parameters are statistically analyzed, and after calculating the precision sensitivity, it is compared with a preset sensitivity threshold to select the calculation precision and obtain the mixed precision configuration.
[0010] Perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration comparison to obtain the partition comparison results;
[0011] Compare the partition comparison results and calculate the deviation benefit threshold value. Match the deviation benefit threshold value with the preset weight range to obtain the fusion weight. Compare the deviation benefit threshold value with the preset backoff threshold to obtain the accuracy backoff mark and form the gated fusion parameters.
[0012] The feature outputs of each training partition are fused according to the gating fusion parameters. The fusion training loss is calculated and the parameters of the multimodal large model are updated. The training partition corresponding to the accuracy back-down mark is configured to a higher level of calculation accuracy. The training is iterated until the stopping condition is met, and the multimodal large model with fusion training is obtained.
[0013] As the mixed-precision model fusion training method for multimodal large models described in this invention, its characteristic is that: the pairing of multimodal samples of the same training task with task supervision labels specifically includes,
[0014] Obtain a multimodal training sample set and extract text samples, image samples, audio samples, video samples, and task supervision labels from the same training task;
[0015] Assign training task numbers to the same training task, and pair text samples, image samples, audio samples, video samples, and task supervision labels according to the training task numbers to obtain task sample groups.
[0016] As the mixed-precision model fusion training method for multimodal large models described in this invention, its characteristic is that: obtaining the partitioned training input specifically includes,
[0017] Based on modal identification information, text samples, image samples, audio samples, and video samples are mapped to modal coding partitions respectively, and modal feature encoding processing is performed to obtain modal coding partitions;
[0018] Cross-modal feature mapping and alignment are performed on the modality-coded partitions to construct cross-modal semantic associations and obtain cross-modal projected partitions;
[0019] Multimodal feature fusion calculation is performed on the cross-modal projection partition to generate fused training representations. The fused training subspace is then divided based on the feature dimensions of the fused training representations and the mapping relationship between the training paths to form fused training partitions.
[0020] The task sample groups are arranged according to the training paths corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition training input.
[0021] As the mixed-precision model fusion training method for multimodal large models described in this invention, its characteristic is that the statistical training state parameters specifically include:
[0022] According to the preset short-range rounds, the multimodal samples in the partition training input are forward propagated in the modality coding partition, and the layer output activation values of the modality coding partition are recorded.
[0023] Modal features are obtained by feature aggregation of the layer output activation values;
[0024] Modal features are input into cross-modal projection partitions for dimensionality mapping, and the dimensionality mapping results are then subjected to semantic space mapping to obtain projected feature values.
[0025] The projected feature values are input into the fusion training partition for feature fusion, the fusion prediction result is calculated, and the loss is calculated by combining the fusion prediction result with the task supervision label to obtain the loss value;
[0026] Backpropagation is performed on the loss value to obtain the gradient value of each training partition, and the parameters are updated according to the gradient value to obtain the updated parameter value.
[0027] The training state parameters are obtained by continuously recording the layer output activation values, projected feature values, loss values, gradient values and parameter update values within a preset short-range round. The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value are calculated.
[0028] As the mixed-precision model fusion training method for multimodal large models described in this invention, its characteristic is that obtaining the mixed-precision configuration specifically includes,
[0029] The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value in the training state parameters are normalized to obtain the state normalization group.
[0030] The maximum normalized value in the state normalization group is selected as the accuracy sensitivity of the corresponding training partition;
[0031] The accuracy sensitivity is compared with a preset sensitivity threshold, and high sensitivity level, medium sensitivity level, and low sensitivity level are marked.
[0032] The calculation precision is configured according to the sensitivity level labeled in each training partition. The high sensitivity level, medium sensitivity level, and low sensitivity level correspond to the first calculation precision, the second calculation precision, and the third calculation precision, respectively, resulting in a mixed precision configuration.
[0033] As the mixed-precision model fusion training method for multimodal large models described in this invention, its characteristic is that obtaining the partition comparison results specifically includes,
[0034] The computational precision in the mixed precision configuration is mapped to the modality coding partition, cross-modality projection partition, and fusion training partition, and arranged in the execution order of the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition precision execution path;
[0035] Perform mixed-precision forward computation on the partition training input along the partition precision execution path to extract modal features of modal coding partitions, projection features of cross-modal projection partitions, and fusion features of fusion training partitions, and determine the feature output of each training partition;
[0036] The feature outputs of each training partition are arranged according to the partition precision execution path to obtain the precision calculation results;
[0037] Perform high-precision calibration operations on the same partition training input according to the first calculation precision, extract the reference modal features of the modal coding partition, the reference projection features of the cross-modal projection partition, and the reference fusion features of the fusion training partition to obtain the high-precision calibration results;
[0038] The precision calculation results and high-precision calibration results are aligned by partition, and the modal features are paired with the reference modal features, the projection features with the reference projection features, and the fusion features with the reference fusion features according to the same training partition to obtain the partition comparison results.
[0039] As the hybrid precision model fusion training method for multimodal large models described in this invention, its characteristic is that: the comparison of partitioning results and calculation of the bias gain threshold specifically includes,
[0040] The partition comparison results are analyzed, and the difference amplitudes between modal features and reference modal features, projection features and reference projection features, and fusion features and reference fusion features are calculated to obtain feature deviation groups.
[0041] The feature bias group is normalized to determine the accuracy bias values corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition;
[0042] The accuracy benefit value corresponding to the accuracy deviation value is determined based on a preset benefit benchmark value.
[0043] By matching a pre-defined gating mapping relationship between the combination of accuracy deviation value and accuracy gain value, the deviation gain gating value of each training partition is obtained.
[0044] As the hybrid precision model fusion training method for multimodal large models described in this invention, its characteristic is that: the formation of gated fusion parameters specifically includes,
[0045] Arrange the bias gain gate values of each training partition according to the partition precision execution path to obtain the gate value sequence;
[0046] The deviation payoff gate values in the gate value sequence are matched with preset weight intervals, and the fusion weights are configured to decrease from low to high according to the interval values to obtain the fusion weight sequence.
[0047] Traverse the fusion weight sequence, summarize the fusion weights of each training partition, obtain the weight baseline value, and use the weight baseline value as the proportional benchmark to determine the weight ratio of each training partition, thus obtaining the weight ratio sequence.
[0048] Compare the bias gain gate value in the gate value sequence with a preset back-off threshold, and label the training partition that reaches the preset back-off threshold with an accuracy back-off flag.
[0049] The gating fusion parameters are formed by combining the weight ratio sequence of the execution path according to the partition precision and the precision backoff flag.
[0050] As the hybrid precision model fusion training method for multimodal large models described in this invention, its characteristic is that: the obtained multimodal large model after fusion training specifically includes,
[0051] Based on the weight ratio sequence in the gated fusion parameters, the feature output in the precision calculation result is configured with weights to obtain the partitioned weighted features;
[0052] The partition-weighted features are input into the fusion training partition for feature fusion, and the fusion prediction result is calculated.
[0053] The fusion prediction results and task supervision labels are used to calculate the loss, resulting in the fusion training loss;
[0054] Perform backpropagation on the fusion training loss to update the multimodal large model parameters corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition;
[0055] Configure the training partition with the precision fallback flag in the gated fusion parameters to a higher level of computational precision, and update the mixed precision configuration.
[0056] Repeatedly perform mixed-precision training operations, partition comparison, gating fusion parameter formation, and model parameter updates until the preset stopping condition is met, to obtain a multimodal large model that has completed fusion training.
[0057] Secondly, this invention provides a hybrid precision model fusion training system for multimodal large models, comprising:
[0058] The sample partitioning module is used to obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partitioned training inputs.
[0059] The precision selection module is used to perform short-range pre-training on the partitioned training input, statistically analyze the training state parameters, and compare the calculated precision sensitivity with a preset sensitivity threshold to select the calculation precision and obtain a mixed precision configuration.
[0060] The calibration and comparison module is used to perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration and comparison to obtain the partition comparison results.
[0061] The gating generation module is used to compare the partition comparison results and calculate the deviation benefit gating value. It matches the deviation benefit gating value with the preset weight range to obtain the fusion weight, and compares the deviation benefit gating value with the preset backoff threshold to obtain the accuracy backoff mark, thus forming the gating fusion parameters.
[0062] The fusion training module is used to fuse the feature outputs of each training partition according to the gating fusion parameters, calculate the fusion training loss and update the parameters of the multimodal large model, configure the training partition corresponding to the accuracy back-down marker to a higher level of calculation accuracy, iterate training until the stopping condition is met, and obtain the multimodal large model that has completed fusion training.
[0063] The beneficial effects of this invention are as follows: By normalizing training state parameters and comparing accuracy sensitivity, the activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value are incorporated into the same accuracy sensitivity determination process. This enables modal coding partitions, cross-modal projection partitions, and fusion training partitions to be configured with corresponding computational accuracy according to the training state, enhancing the partition adaptability, training state correlation, and accuracy control stability of the mixed accuracy configuration of multimodal large models. By forming a unified gating basis through partition comparison results, bias gain gating values, fusion weights, and accuracy backoff markers, high-precision calibration bias, feature output ratio, and computational accuracy backoff path are collaboratively involved in the control within the same machine learning training chain, improving the consistency of feature expression, parameter update stability, and accuracy maintenance capability of multimodal large model fusion training. Attached Figure Description
[0064] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 This is a flowchart of a hybrid precision model fusion training method for large multimodal models.
[0066] Figure 2 This is a system module diagram for hybrid precision model fusion training for large multimodal models.
[0067] Figure 3 A flowchart for configuring mixed precision.
[0068] Figure 4 A flowchart for generating gating fusion parameters. Detailed Implementation
[0069] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0070] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0071] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0072] Reference Figures 1-4 This is one embodiment of the present invention, which provides a hybrid precision model fusion training method for multimodal large models, including the following steps:
[0073] S1: Obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partition training inputs.
[0074] S1.1: Obtain a multimodal training sample set and extract text samples, image samples, audio samples, video samples, and task supervision labels from the same training task.
[0075] It should be noted that sample files are retrieved from the multimodal training corpus that has been collected, cleaned, and labeled. Sample files belonging to the same training task are filtered according to the training task name, training task number, sample source identifier, and label fields. Text descriptions, question-and-answer content, and subtitle content are converted into text samples; image files are converted into image samples; video frame-by-frame images are converted into image samples, and the source video number and frame extraction time position are recorded; audio files and audio track segments are converted into audio samples; video files are converted into video samples according to frame sequence and time interval, and the source video number and frame sequence range are recorded; the sample label field, task answer field, and category label field are converted into task supervision labels. Image samples with the same source video number are associated with video samples and are processed as image modality input and video modality input respectively in the subsequent encoding stage. Image modality input is used for local frame-level feature learning, and video modality input is used for temporal dynamic feature learning to achieve complementary modality joint modeling, and task supervision labels are not generated repeatedly. Text samples, image samples, audio samples, video samples, and task supervision labels with consistent training task names, complete sample source identifiers, and complete task supervision labels are retained.
[0076] S1.2: Assign training task numbers to the same training task, and pair text samples, image samples, audio samples, video samples and task supervision labels according to the training task numbers to obtain task sample groups.
[0077] It should be noted that the retained text samples, image samples, audio samples, video samples, and task supervision labels are categorized according to the training task name. A unique training task number is assigned to the same training task name, and the training task number is filled into the record fields of the text samples, image samples, audio samples, video samples, and task supervision labels respectively. Text samples, image samples, audio samples, video samples, and task supervision labels are retrieved according to the training task number. Text samples, image samples, audio samples, video samples, and task supervision labels with the same training task number, the same training task name, and complete task supervision labels are placed into the same pairing set. Samples with inconsistent training task numbers and samples with missing task supervision labels are not included in the pairing. After the pairing is completed, the task sample group is obtained.
[0078] S1.3 performs modal parsing on the task sample group to generate corresponding modal identification information.
[0079] It should be noted that text samples, image samples, audio samples, and video samples are read one by one from the task sample group, and modality discrimination processing is performed on each sample. Semantic type analysis is performed on text samples to extract text semantic features and label the text modality type. Image feature extraction is performed on image samples to identify image content attributes and label the image modality type. Mel spectrum analysis and speech event recognition are performed on audio samples to label the audio modality type. Temporal segmentation and keyframe extraction are performed on continuous frame sequences of video samples, and temporal semantic features are extracted to label the video modality type. The modality type labeling results corresponding to each sample are written into the sample record field to generate modality identification information.
[0080] S1.4: Based on modal identification information, text samples, image samples, audio samples and video samples are mapped to modal coding partitions respectively, and modal feature encoding processing is performed to obtain modal coding partitions.
[0081] It should be noted that, based on modal identification information, modal routing allocation processing is performed on text samples, image samples, audio samples, and video samples in the task sample group respectively. According to the modal identification information, text samples are routed to the text modal coding submodule, image samples are routed to the image modal coding submodule, audio samples are routed to the audio modal coding submodule, and video samples are routed to the video modal coding submodule. The corresponding modal coding network is then called to perform feature encoding processing to generate text encoding features, image encoding features, audio encoding features, and video encoding features. The modal coding features are then aggregated into modal coding partitions to form modal coding partitions.
[0082] S1.5: Perform cross-modal feature mapping and alignment processing on the modal coding partitions, construct cross-modal semantic associations, and obtain cross-modal projection partitions.
[0083] It should be noted that a unified feature space mapping process is performed on the text coding features, image coding features, audio coding features, and video coding features in the modality coding partition, respectively, to transform the different modality features into a shared semantic space through vector projection. Similarity calculation and semantic alignment optimization are performed on the projected multimodal features to establish semantic association matching relationships between text features, image features, audio features, and video features. The cross-modal features that have completed semantic alignment are grouped and aggregated according to the projection results to form cross-modal projection partitions.
[0084] S1.6: Perform multimodal feature fusion calculation on cross-modal projection partitions to generate fused training representations, and divide the fused training subspace based on the feature dimensions of the fused training representations and the mapping relationship between the training paths to form fused training partitions.
[0085] It should be noted that multimodal feature splicing and nonlinear fusion calculations are performed on the alignment features in the cross-modal projection partition to generate a fused training representation; the fused training representation is subjected to feature dimension analysis to determine the dimensional distribution of different modal features in the fused vector; a mapping relationship is established based on the feature dimension distribution and the preset training path mapping rules, and the fused training representation is subjected to subspace partitioning to divide the fused training representation into fused training subspaces corresponding to different training paths; each fused training subspace is aggregated according to the path to form a fused training partition.
[0086] To further explain, the training path mapping rules are set based on the differences in the semantic contribution of multimodal features in the fusion training representation and the computational objectives of different training stages. They are used to divide the feature dimensions in the fusion training representation into three types of computational paths: modality encoding, cross-modal interaction, and fusion optimization.
[0087] S1.7: Arrange task sample groups according to the training paths corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition training input.
[0088] It should be noted that the text samples, image samples, audio samples, and video samples in the task sample group are arranged to the input terminals of the text encoder, image encoder, audio encoder, and video encoder of the modality coding partition, respectively. The text features, image features, audio features, and video features output from the modality coding partition are arranged to the input terminal of the cross-modality projection partition. The text projection features, image projection features, audio projection features, and video projection features output from the cross-modality projection partition are arranged to the input terminal of the fusion training partition. The task supervision label is arranged to the loss calculation node of the fusion training partition. The training task number, sample input position, partition input position, and task supervision label position are recorded according to the execution order of the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition training input.
[0089] S2: Perform short-range pre-training on the partitioned training input, statistically analyze the training state parameters, and compare the calculated precision sensitivity with the preset sensitivity threshold to select the calculation precision and obtain the mixed precision configuration.
[0090] S2.1: According to the preset short-range rounds, the multimodal samples in the partition training input are forward propagated in the modality coding partition, and the layer output activation value of the modality coding partition is recorded.
[0091] It should be noted that the training input for each partition is fed into the modal coding partition round by round according to the training task number, so that text samples enter the text encoder, image samples enter the image encoder, audio samples enter the audio encoder, and video samples enter the video encoder. The text encoder, image encoder, audio encoder, and video encoder output the intermediate features of the corresponding layer according to the forward propagation direction. The layer name, layer number, output tensor range, and activation distribution information of the text encoder, image encoder, audio encoder, and video encoder in each round are recorded, and the recorded content is used as the layer output activation value of the modal coding partition.
[0092] To further explain, the number of short-range rounds is set according to the needs of short-range pre-training, and is used to limit the number of forward propagation records of the partition training input within the modality coding partition. The initial computational precision of short-range pre-training adopts a unified baseline precision configuration, which is set to 32-bit floating-point computational precision to ensure the stability and comparability of the modality coding partition during the short-range forward propagation stage. The forward propagation recording process of the preset number of short-range rounds is executed under the baseline precision configuration, and stability statistical analysis is performed based on the loss change value, gradient fluctuation amplitude, and parameter update offset of each round. The number of short-range rounds is 3 to 5 rounds. By comprehensively judging whether the loss decrease trend gradually flattens out, whether the gradient fluctuation amplitude gradually stabilizes, and whether the parameter update amplitude gradually becomes consistent, it is determined that the short-range pre-training process has covered the initial feature learning stage of the model.
[0093] S2.2: Perform feature aggregation on the layer output activation values to obtain modal features.
[0094] It should be noted that the activation values of the modality coding layer outputs are categorized and arranged according to text encoder, image encoder, audio encoder, and video encoder, preserving the differences in activation distribution among different layers within the same encoder. The activation values of the text encoder layer outputs are processed by attention weight convergence and sequence compression according to the word position dimension to generate text modality features. The activation values of the image encoder layer outputs are processed by two-dimensional feature aggregation and attention convergence according to the spatial position of image blocks to generate image modality features. The activation values of the audio encoder layer outputs are processed by temporal aggregation and attention convergence according to the audio frame time sequence to generate audio modality features. The activation values of the video encoder layer outputs are processed by temporal segment aggregation and attention convergence according to the temporal structure of video frames to generate video modality features. The text modality features, image modality features, audio modality features, and video modality features are combined to form modality features.
[0095] S2.3: Input the modal features into the cross-modal projection partition for dimensional mapping, and perform semantic space mapping on the dimensional mapping results to obtain the projected feature values.
[0096] It should be noted that text modal features, image modal features, audio modal features, and video modal features are respectively input into the projection layer in the cross-modal projection partition. The feature length, channel dimension, and feature arrangement are processed according to the input dimension requirements of the projection layer to form text dimension mapping results, image dimension mapping results, audio dimension mapping results, and video dimension mapping results with the same feature dimension. The text dimension mapping results, image dimension mapping results, audio dimension mapping results, and video dimension mapping results are mapped to a unified semantic space to eliminate the differences in dimensional scale and expression position of different modal features, and the projected feature values are obtained.
[0097] S2.4: Input the projected feature values into the fusion training partition to perform feature fusion, calculate the fusion prediction result, and calculate the loss value by combining the fusion prediction result with the task supervision label.
[0098] It should be noted that the projection feature values are input into the fusion training partition. The fusion training partition is associated with the task supervision label according to the training task number. The text projection features, image projection features, audio projection features, and video projection features are fused in the same semantic space to form fused features. The fusion training partition sends the fused features to the training layer used to output the task results. The training layer performs output transformation on the fused features according to the format requirements of the task supervision label to generate a fusion prediction result consistent with the format of the task supervision label. The fusion prediction result is compared with the task supervision label under the same training task number, and the loss value is formed according to the difference between the fusion prediction result and the task supervision label.
[0099] S2.5: Perform backpropagation on the loss value to obtain the gradient value of each training partition, and perform parameter update according to the gradient value to obtain the parameter update value.
[0100] It should be noted that the loss value is propagated layer by layer along the backpropagation direction of the fusion training partition, the cross-modal projection partition, and the modal coding partition. The gradient values corresponding to the training layer parameters in the fusion training partition, the gradient values corresponding to the projection layer parameters in the cross-modal projection partition, and the gradient values corresponding to the text encoder, image encoder, audio encoder, and video encoder parameters in the modal coding partition are determined respectively. Based on the gradient values corresponding to the trainable parameters in each training partition, and combined with the optimizer rules to determine the parameter update direction and parameter update step size, the trainable parameters in the fusion training partition, the cross-modal projection partition, and the modal coding partition are updated, and the change in parameters before and after the update is recorded as the parameter update value.
[0101] To further explain, the optimizer rules include a learning rate adjustment mechanism, a momentum accumulation mechanism, and a gradient constraint handling mechanism. These are set based on the gradient distribution characteristics of each training partition and are used to constrain and correct the gradient update process to control the stability and convergence of parameter updates. The learning rate adjustment mechanism is used to uniformly scale the parameter update step size according to the training state of the training partition to avoid excessively large or small parameter update amplitudes. The momentum accumulation mechanism is used to accumulate and fuse the current gradient with historical gradients to enhance the stability of the parameter update direction and suppress gradient oscillations. The gradient constraint handling mechanism is used to limit the magnitude of gradient values to prevent abnormal gradients from interfering with the parameter update process.
[0102] S2.6: Continuously record the layer output activation value, projected feature value, loss value, gradient value and parameter update value within a preset short-range round, calculate the activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value to obtain the training state parameters.
[0103] It should be noted that, according to the training epochs, the layer output activation values, projected feature values, loss values, gradient values, and parameter update values within a preset short-range epoch are saved. The layer output activation values are compared with the numerical expression range of each candidate computational precision in the preset detection precision group. Simultaneously, the numerical underflow of gradient values during backpropagation, the cumulative error during parameter updates, and the numerical stability after loss scaling are statistically analyzed. The proportion exceeding the numerical expression range of each candidate computational precision and the proportion exhibiting numerical instability are calculated to form the activation overflow ratio. The text projection feature values, image projection feature values, audio projection feature values, and video projection feature values are analyzed. The eigenvalues are calculated using cosine similarity to determine cross-modal semantic consistency, and the average of all similarities is used to form the cross-modal alignment distance. The loss values of adjacent training epochs are interpolated and their absolute values are taken to form the loss change value. The gradient values of the same training partition in adjacent training epochs are interpolated and their absolute values are taken to form the gradient fluctuation value. The parameter update values of the same training partition in adjacent training epochs are interpolated and their absolute values are taken to form the parameter update offset value. The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value are combined to obtain the training state parameters.
[0104] To further explain, the detection precision group is based on the mixed precision training candidate computation precision settings and is used to statistically analyze the overflow of layer output activation values under different candidate computation precisions. For example, the detection precision group includes a 32-bit floating-point format, a 16-bit floating-point format, and an 8-bit quantization format constructed based on the quantization scale and zero-point parameters. The 32-bit floating-point format, the 16-bit floating-point format, and the 8-bit quantization format constructed based on the quantization scale and zero-point parameters have corresponding upper and lower limits, respectively. When the layer output activation value exceeds the corresponding upper and lower limits, it is recorded as an overflow.
[0105] S2.7: Normalize the activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value in the training state parameters to obtain the state normalization group.
[0106] It should be noted that activation overflow ratio, cross-modal alignment distance, loss change, gradient fluctuation, and parameter update offset are processed separately according to the training partition. The cross-modal alignment distance is calculated by the cross-modal projection partition based on text projection feature values, image projection feature values, audio projection feature values, and video projection feature values. Specifically, it includes: mapping each modal projection feature value to a unified semantic space, calculating the distance between any two modalities, constructing a cross-modal distance set, and performing an arithmetic mean on the distance values in the cross-modal distance set to obtain the cross-modal alignment distance; the cross-modal alignment distance calculation result is synchronously mapped to the modality coding partition and the fusion training partition for comparison with other training state parameters of each training partition. Unified normalization processing is performed; the minimum recorded value of the same training state parameter within the same training partition is used as the scale starting point, and the maximum recorded value of the same training state parameter within the same training partition is used as the scale ending point. The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value are converted into dimensionless normalized values. For training state parameters with the same maximum and minimum recorded values, the normalized value is set to zero to avoid losing the basis for comparison during the scale transformation. The normalized values of activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value are combined to obtain the state normalization group.
[0107] S2.8: Select the maximum normalized value in the state normalization group as the accuracy sensitivity of the corresponding training partition.
[0108] It should be noted that in the state normalization group, the normalized values of activation overflow ratio, cross-modal alignment distance, loss change, gradient fluctuation, and parameter update offset are arranged according to modality coding partition, cross-modal projection partition, and fusion training partition, respectively. The five normalized values within the same training partition are compared, and the normalized value with the largest value is selected as the accuracy sensitivity of the same training partition. When two or more normalized values within the same training partition reach their maximum values at the same time, the set of index types that reach the maximum normalized value is recorded, and the maximum normalized value and the set of index types are used together as the accuracy sensitivity output of the same training partition.
[0109] S2.9: Compare the accuracy sensitivity with the preset sensitivity threshold and label it as high sensitivity level, medium sensitivity level and low sensitivity level.
[0110] It should be noted that the accuracy sensitivity corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition is compared with a preset sensitivity threshold. If the accuracy sensitivity reaches the high-order threshold of the preset sensitivity threshold, the corresponding training partition is marked as high sensitivity level. If the accuracy sensitivity is lower than the high-order threshold but reaches the low-order threshold, the corresponding training partition is marked as medium sensitivity level. If the accuracy sensitivity is lower than the low-order threshold, the corresponding training partition is marked as low sensitivity level.
[0111] To further explain, the sensitivity threshold is set based on the numerical distribution of the state normalization group and the switching of computational precision, and is used to distinguish the sensitivity of the training partition to changes in computational precision; the sensitivity threshold is divided into low-order boundary value and high-order boundary value, with the low-order boundary value set at 0.30 and the high-order boundary value set at 0.70.
[0112] S2.10: Configure the calculation precision according to the sensitivity level of each training partition. High sensitivity level, medium sensitivity level and low sensitivity level correspond to the first calculation precision, the second calculation precision and the third calculation precision, respectively, to obtain the mixed precision configuration.
[0113] It should be noted that the calculation precision is configured according to the high sensitivity level, medium sensitivity level, and low sensitivity level of the modality coding partition, cross-modality projection partition, and fusion training partition. The training partition with a sensitivity level of high sensitivity is configured with the first calculation precision, the training partition with a sensitivity level of medium sensitivity is configured with the second calculation precision, and the training partition with a sensitivity level of low sensitivity is configured with the third calculation precision. The calculation precision configuration results of the modality coding partition, cross-modality projection partition, and fusion training partition are combined to obtain the mixed precision configuration.
[0114] To further explain, the first, second, and third computational precisions are set based on the numerical formats supported by the training device and the training stability requirements, and are used to control the training computation precision of training partitions with different sensitivity levels; for example, the first computational precision is 32-bit floating-point precision, the second computational precision is 16-bit floating-point precision, and the third computational precision is 8-bit fixed-point precision. The first computational precision is higher than the second computational precision, and the second computational precision is higher than the third computational precision.
[0115] S3: Perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration comparison to obtain the partition comparison results.
[0116] S3.1: Map the computational precision in the mixed precision configuration to the modality coding partition, cross-modality projection partition, and fusion training partition, and arrange them in the execution order of the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition precision execution path.
[0117] It should be noted that, based on the mixed precision configuration, the computational precision configuration results of the modal coding partition, cross-modal projection partition, and fusion training partition are found. The modal coding partition is associated with the corresponding first, second, or third computational precision, the cross-modal projection partition is associated with the corresponding first, second, or third computational precision, and the fusion training partition is associated with the corresponding first, second, or third computational precision. According to the execution order of the partition training input in the multimodal large model, the modal coding partition and its corresponding computational precision are arranged first, the cross-modal projection partition and its corresponding computational precision are arranged in the middle, and the fusion training partition and its corresponding computational precision are arranged last, thus obtaining the partition precision execution path.
[0118] S3.2: Perform mixed-precision forward computation on the partition training input along the partition precision execution path to extract the modal features of the modal coding partition, the projection features of the cross-modal projection partition, and the fusion features of the fusion training partition, and determine the feature output of each training partition.
[0119] It should be noted that, according to the training partition order arranged in the partition precision execution path, the partition training input is fed into the multimodal large model; the modality coding partition encodes text samples, image samples, audio samples, and video samples according to the associated computational precision, and extracts text modality features, image modality features, audio modality features, and video modality features as the modality features of the modality coding partition; the cross-modality projection partition performs dimensional mapping and semantic space mapping on the modality features according to the associated computational precision, and extracts projection features; the fusion training partition performs fusion processing on the projection features according to the associated computational precision, and extracts fusion features; the modality features, projection features, and fusion features are respectively determined as the feature outputs of each training partition.
[0120] S3.3: Arrange the feature outputs of each training partition according to the partition precision execution path to obtain the precision calculation results.
[0121] It should be noted that the partition precision execution path is used as the basis for arrangement. Modal features, projection features, and fusion features generated by the same partition training input are searched according to the training task number. The modal features are matched with the computational precision configuration results of the modal coding partition, the projection features are matched with the computational precision configuration results of the cross-modal projection partition, and the fusion features are matched with the computational precision configuration results of the fusion training partition. The modal features, projection features, and fusion features are arranged according to the order of the modal coding partition, the cross-modal projection partition, and the fusion training partition in the partition precision execution path to obtain the precision calculation results.
[0122] S3.4: Perform high-precision calibration operation on the same partition training input according to the first calculation precision, extract the reference modal features of the modal coding partition, the reference projection features of the cross-modal projection partition and the reference fusion features of the fusion training partition, and obtain the high-precision calibration result.
[0123] It should be noted that the same partition training input as the mixed-precision training operation is selected, and the arrangement of training task numbers, text samples, image samples, audio samples, video samples, and task supervision labels remains unchanged. The modality coding partition, cross-modal projection partition, and fusion training partition are all configured with the first computational precision, and the trainable parameters in each training partition are frozen. Backpropagation and parameter update operations are not performed. The modality coding partition encodes text samples, image samples, audio samples, and video samples at the first computational precision to extract reference modality features. The cross-modal projection partition performs dimensionality mapping and semantic space mapping on the reference modality features at the first computational precision to extract reference projection features. The fusion training partition fuses the reference projection features at the first computational precision to extract reference fusion features. The reference modality features, reference projection features, and reference fusion features are arranged according to the execution order of the training partitions to obtain high-precision calibration results.
[0124] S3.5: Align the precision calculation results and high-precision calibration results by partition, and pair the modal features with the reference modal features, the projection features with the reference projection features, and the fusion features with the reference fusion features according to the same training partition to obtain the partition comparison results.
[0125] It should be noted that, according to the training task number and the partition precision execution path, the modal features, projection features, and fusion features in the precision calculation results are partitioned and aligned with the reference modal features, reference projection features, and reference fusion features in the high-precision calibration results. Modal features belonging to the modal coding partition under the same training task number are paired with the reference modal features; projection features belonging to the cross-modal projection partition under the same training task number are paired with the reference projection features; and fusion features belonging to the fusion training partition under the same training task number are paired with the reference fusion features. During the alignment process, the high-precision calibration results are used as a unified reference benchmark, and consistency constraints are applied to the training partition names and execution order. However, it is not required that the calculation precision configuration of the mixed precision calculation results be consistent with that of the high-precision calibration results, thus obtaining the partition comparison results.
[0126] S4: Compare the partition comparison results and calculate the deviation benefit threshold value. Match the deviation benefit threshold value with the preset weight range to obtain the fusion weight. Compare the deviation benefit threshold value with the preset backoff threshold to obtain the accuracy backoff mark and form the gated fusion parameters.
[0127] S4.1: Analyze the partition comparison results, calculate the difference amplitude between modal features and reference modal features, projection features and reference projection features, and fusion features and reference fusion features, and obtain the feature deviation group.
[0128] It should be explained that, according to the training task number and training partition name, modal features and reference modal features in the modal coding partition, projection features and reference projection features in the cross-modal projection partition, and fusion features and reference fusion features in the fusion training partition are located from the partition comparison results. Under the same training task number and the same training partition name, modal features and reference modal features are matched according to the same feature position. The modal feature values at the same feature position are then compared with the reference modal feature values. The absolute value of the difference is taken to obtain the modal feature difference amplitude. The average modal feature difference amplitude is calculated according to all feature positions to obtain the modal feature difference mean. The average reference modal feature absolute value is calculated according to all feature positions to obtain the modal reference feature scale value. The modal feature difference mean and modal reference feature scale value are proportionally calibrated to obtain the modal feature bias. Projection features and reference projection features are matched according to the same feature position. The projection feature values at the same feature position are then compared with the reference projection feature values. The modal feature difference is processed by taking the absolute value of the difference to obtain the projection feature difference amplitude. The average projection feature difference amplitude across all feature locations is then calculated to obtain the projection feature difference mean. The average projection feature scale value is obtained by taking the absolute value of the reference projection feature across all feature locations. The projection feature deviation is obtained by proportionally calibrating the projection feature difference mean and the projection feature scale value. The fused feature is then matched with the reference fused feature at the same feature location. The fused feature value at the same feature location is then compared with the reference fused feature value, and the absolute value of the difference is taken to obtain the fused feature difference amplitude. The average fused feature difference amplitude across all feature locations is then calculated to obtain the fused feature difference mean. The average fused feature scale value is obtained by taking the absolute value of the reference fused feature across all feature locations. The fused feature deviation is obtained by proportionally calibrating the fused feature scale value. Finally, the modal feature deviation, projection feature deviation, and fused feature deviation are combined to obtain the feature deviation group.
[0129] S4.2: Normalize the feature bias group to determine the accuracy bias value corresponding to the modality coding partition, cross-modality projection partition and fusion training partition.
[0130] It should be noted that the modal feature bias, projection feature bias, and fusion feature bias in the feature bias group are normalized separately, with the minimum feature bias within their respective training partitions as the normalization starting point and the maximum feature bias within their respective training partitions as the normalization ending point. For training partitions where the maximum and minimum feature biases are the same, the normalized value is recorded as 0. The normalized value of the modal feature bias is determined as the modal coding partition accuracy bias value, the normalized value of the projection feature bias is determined as the cross-modal projection partition accuracy bias value, and the normalized value of the fusion feature bias is determined as the fusion training partition accuracy bias value.
[0131] S4.3: Use the preset benefit benchmark value as a benchmark to calibrate the accuracy benefit value corresponding to the accuracy deviation value.
[0132] It should be noted that a preset revenue benchmark value is used as the calibration benchmark for the accuracy deviation value. The accuracy deviation values corresponding to the modal coding partition, cross-modal projection partition, and fusion training partition are calibrated proportionally. The ratio of the accuracy deviation value to the preset revenue benchmark value is determined as the normalized deviation value, and the result of subtracting 1 from the normalized deviation value is used as the accuracy revenue value. When the accuracy deviation value reaches the preset revenue benchmark value, the accuracy revenue value is recorded as 0, indicating that the corresponding training partition has reached the lower limit of revenue. When the accuracy deviation value is 0, the accuracy revenue value is recorded as 1, indicating that the corresponding training partition has reached the upper limit of revenue. The accuracy revenue value is calibrated separately for the modal coding partition, cross-modal projection partition, and fusion training partition.
[0133] To further explain, the benchmark value for returns is set based on the normalized distribution of the characteristic deviation group and the adjustment of the mixed precision configuration. It is used to determine whether the precision deviation value meets the return calibration condition that requires participation in the gating judgment, and the value is 0.50.
[0134] S4.4: Match the preset gating mapping relationship according to the combination interval of the accuracy deviation value and the accuracy gain value to obtain the deviation gain gating value of each training partition.
[0135] It should be explained that the precision deviation values and precision gain values corresponding to the modal coding partition, cross-modal projection partition, and fusion training partition are arranged in pairs to form a deviation-gain combination. The deviation-gain combination is matched with the combination interval in the preset gating mapping relationship, which divides the combination interval according to the first deviation threshold, the second deviation threshold, and the gain threshold. If the precision deviation value is lower than the first deviation threshold and the precision gain value reaches the gain threshold, a low deviation combination interval is matched, and the deviation-gain gating value is configured as a low gating value. If the precision deviation value reaches the first deviation threshold but is lower than the second deviation threshold, a medium deviation combination interval is matched, and the deviation-gain gating value is configured as a medium gating value. If the precision deviation value reaches the second deviation threshold, a high deviation combination interval is matched, and the deviation-gain gating value is configured as a high gating value. The deviation-gain gating values are output for the modal coding partition, cross-modal projection partition, and fusion training partition respectively, to obtain the deviation-gain gating values for the modal coding partition, cross-modal projection partition, and fusion training partition.
[0136] To further explain, the gating mapping relationship is based on the accuracy deviation value, accuracy gain value, and gating fusion parameters, and needs to be set to convert the combination relationship of accuracy deviation value and accuracy gain value into a deviation gain gating value; for example, the first deviation threshold is 0.30, the second deviation threshold is 0.70, the gain threshold is 0.50, the low gating value is 0.20, the medium gating value is 0.50, and the high gating value is 0.90.
[0137] S4.5: Arrange the bias gain gate values of each training partition according to the partition precision execution path to obtain the gate value sequence.
[0138] It should be noted that, according to the arrangement of training partitions in the partition precision execution path, the bias gain gate values already obtained for the modality coding partition, cross-modality projection partition, and fusion training partition are found. The bias gain gate values of the modality coding partition are arranged in the modality coding partition position, the bias gain gate values of the cross-modality projection partition are arranged in the cross-modality projection partition position, and the bias gain gate values of the fusion training partition are arranged in the fusion training partition position. During the arrangement process, the training task number and training partition name are retained to ensure that the bias gain gate values under the same training task number are consistent with the order of training partitions in the partition precision execution path, thus obtaining a gate value sequence.
[0139] S4.6: Match the deviation revenue gate values in the gate value sequence to preset weight intervals, and configure decreasing fusion weights according to the interval values from low to high to obtain a fusion weight sequence.
[0140] It should be noted that the deviation reward gate values of the modal coding partition, cross-modal projection partition, and fusion training partition are read item by item according to the gate value sequence. The deviation reward gate values are matched with the preset weight intervals to determine the preset weight intervals in which the deviation reward gate values fall. The interval values of the preset weight intervals are arranged from low to high. The first fusion weight is configured for the low interval values, the second fusion weight is configured for the middle interval values, and the third fusion weight is configured for the high interval values. The fusion weights corresponding to the modal coding partition, cross-modal projection partition, and fusion training partition are arranged according to the gate value sequence to obtain the fusion weight sequence.
[0141] To further explain, the preset weight range is set based on the value range of the deviation revenue threshold and the configuration of the fusion weight, and is used to convert the deviation revenue threshold into the corresponding fusion weight; for example, the low range value corresponds to the first fusion weight ratio of 0.50, the medium range value corresponds to the second fusion weight ratio of 0.30, and the high range value corresponds to the third fusion weight ratio of 0.20.
[0142] S4.7: Traverse the fusion weight sequence, summarize the fusion weights of each training partition, obtain the weight baseline value, and use the weight baseline value as the proportional benchmark to determine the weight ratio of each training partition, thus obtaining the weight ratio sequence.
[0143] It should be noted that the fusion weights of the modal coding partition, the cross-modal projection partition, and the fusion training partition are read according to the order of the fusion weight sequence. The fusion weights of the modal coding partition, the cross-modal projection partition, and the fusion training partition are summarized to obtain the weight baseline value. Using the weight baseline value as a unified proportional benchmark, the fusion weights of the modal coding partition, the cross-modal projection partition, and the fusion training partition are proportionally calibrated to obtain the weight proportions of the modal coding partition, the cross-modal projection partition, and the fusion training partition. The weight proportions of the modal coding partition, the cross-modal projection partition, and the fusion training partition are combined according to the order of the fusion weight sequence to obtain the weight proportion sequence.
[0144] S4.8: Compare the bias gain gate value in the gate value sequence with the preset backoff threshold, and label the training partition that has reached the preset backoff threshold with the accuracy backoff mark.
[0145] It should be noted that the bias gain gate values of the modal coding partition, cross-modal projection partition, and fusion training partition are read item by item according to the gate value sequence, and the bias gain gate values are compared with preset backoff thresholds respectively. If the bias gain gate value reaches the preset backoff threshold, the training partition to which the bias gain gate value belongs is marked with a precision backoff mark. If the bias gain gate value is lower than the preset backoff threshold, the training partition to which the bias gain gate value belongs is not marked with a precision backoff mark. The precision backoff marks of the modal coding partition, cross-modal projection partition, and fusion training partition are retained according to the gate value sequence.
[0146] To further explain, the backoff threshold needs to be set based on the range of the bias gain gate value and the calculation precision. It is used to determine whether the training partition needs to be configured with a higher calculation precision in subsequent training, and the value is 0.80.
[0147] S4.9: Combine the weight ratio sequence and precision backoff flag according to the partition precision execution path to form the gating fusion parameters.
[0148] It should be noted that, according to the arrangement of modal coding partitions, cross-modal projection partitions, and fusion training partitions in the partition precision execution path, the weight proportions of the corresponding training partitions in the weight proportion sequence are read, and the annotation results of the corresponding training partitions in the precision fallback markers are read; the weight proportions of modal coding partitions are paired with the precision fallback markers of modal coding partitions, the weight proportions of cross-modal projection partitions are paired with the precision fallback markers of cross-modal projection partitions, and the weight proportions of fusion training partitions are paired with the precision fallback markers of fusion training partitions; the corresponding weight proportions of training partitions without precision fallback markers are retained, and the corresponding weight proportions of training partitions with precision fallback markers are retained and the higher-level computational precision configuration requirements are recorded. The pairing results are arranged according to the partition precision execution path to form the gating fusion parameters.
[0149] S5: Fuse the feature outputs of each training partition according to the gating fusion parameters, calculate the fusion training loss and update the parameters of the multimodal large model, configure the training partition corresponding to the accuracy back-down mark to a higher level of calculation accuracy, iterate training until the stopping condition is met, and obtain the multimodal large model with completed fusion training.
[0150] S5.1: Configure weights for the feature output in the precision calculation result based on the weight ratio sequence in the gated fusion parameters to obtain the partitioned weighted features.
[0151] It should be noted that, according to the partition precision execution path, the weight ratio sequence is read from the gated fusion parameters, and the modal features, projection features, and fusion features under the same training task number are found in the precision calculation results; the weight ratio of the modal coding partition is configured to the modal features, the weight ratio of the cross-modal projection partition is configured to the projection features, and the weight ratio of the fusion training partition is configured to the fusion features; the modal features, projection features, and fusion features are weighted according to the corresponding weight ratio. The weighting does not change the original feature dimensions of the modal features, projection features, and fusion features. The partition weighted features are arranged according to the source of the training partition to obtain the partition weighted features.
[0152] S5.2: Input the partition-weighted features into the fusion training partition for feature fusion and calculate the fusion prediction result.
[0153] It should be noted that the partition weighted features are input into the fusion training partition according to the training task number. The fusion training partition performs feature fusion on the partition weighted features corresponding to the modal features and the partition weighted features corresponding to the projection features to generate fusion weighted features. The fusion training partition inputs the fusion weighted features into the output mapping module, and performs mapping transformation on the fusion weighted features according to the format requirements of the task supervision label to generate fusion prediction results.
[0154] S5.3: Calculate the loss by combining the fused prediction results with the task supervision labels to obtain the fused training loss.
[0155] It should be noted that the fusion prediction results are matched with the task supervision labels under the same training task number to verify that the output format of the fusion prediction results is consistent with the label format of the task supervision labels. The loss calculation method is selected according to the field where the task supervision labels are located. The task supervision labels are recorded in the category label field. The fusion training partition uses the cross-entropy loss calculation method to calculate the loss between the category score and the category label field in the fusion prediction results. The task supervision labels are recorded in the task answer field. The fusion training partition uses the sequence cross-entropy loss calculation method to calculate the loss between the answer sequence and the task answer field in the fusion prediction results to obtain the fusion training loss.
[0156] S5.4: Perform backpropagation on the fusion training loss to update the multimodal large model parameters corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition.
[0157] It should be noted that the fusion training loss is propagated along the backpropagation directions of the fusion training partition, the cross-modal projection partition, and the modality coding partition, respectively, to obtain the gradient values of the training layer parameters in the fusion training partition, the gradient values of the projection layer parameters in the cross-modal projection partition, and the gradient values of the text encoder parameters, image encoder parameters, audio encoder parameters, and video encoder parameters in the modality coding partition. The training layer parameters are updated according to the gradient values of the training layer parameters in the fusion training partition, the projection layer parameters are updated according to the gradient values of the projection layer parameters in the cross-modal projection partition, and the text encoder parameters, image encoder parameters, audio encoder parameters, and video encoder parameters are updated according to the gradient values of the text encoder parameters, image encoder parameters, audio encoder parameters, and video encoder parameters in the modality coding partition, thus completing the update of the multimodal large model parameters corresponding to the modality coding partition, the cross-modal projection partition, and the fusion training partition.
[0158] S5.5: Configure the training partition with the precision fallback flag in the gated fusion parameters to a higher calculation precision and update the mixed precision configuration.
[0159] It should be noted that the precision fallback markers of the modality coding partition, cross-modal projection partition, and fusion training partition in the gated fusion parameters are read. The training partitions marked with precision fallback markers are located, and the calculation precision is adjusted according to the hierarchical relationship of the first, second, and third calculation precision. The training partitions configured with the third calculation precision are adjusted to the second calculation precision, the training partitions configured with the second calculation precision are adjusted to the first calculation precision, and the training partitions configured with the first calculation precision are kept at the first calculation precision. The training partitions without precision fallback markers are kept at their original calculation precision. The original mixed precision configuration is replaced with the adjusted calculation precision configuration results of the modality coding partition, cross-modal projection partition, and fusion training partition, and the mixed precision configuration is updated.
[0160] S5.6: Repeatedly perform mixed-precision training operations, partition comparison, gating fusion parameter formation and model parameter update until the preset stopping condition is met, and obtain a multimodal large model with completed fusion training.
[0161] It should be noted that the fusion training rounds are numbered, and the mixed precision training operation, partition comparison, gated fusion parameter formation, and multimodal large model parameter update are executed sequentially according to the current mixed precision configuration. The fusion training loss, precision fallback flag, and mixed precision configuration change results obtained in each training round are recorded. After each training round, the change in fusion training loss, fusion training round number, and precision fallback flag are compared with the preset stopping conditions. If the preset stopping conditions are not met, the fusion training round number is increased by one round, and the mixed precision training operation, partition comparison, gated fusion parameter formation, and multimodal large model parameter update are continued. When the preset stopping conditions are met, the training loop is stopped, the current multimodal large model parameters are fixed, and the multimodal large model with fusion training is obtained.
[0162] To further explain, the stopping condition is set based on the change magnitude of the fusion training loss, the maximum number of training epochs, and the change in the accuracy backoff marker. It is used to determine whether the multimodal large model should end the fusion training. For example, the training loop is stopped when the change magnitude of the fusion training loss is less than 0.01 for three consecutive epochs and no new accuracy backoff markers are added to the modality encoding partition, cross-modality projection partition, and fusion training partition. The maximum number of training epochs is set to 100 epochs as a mandatory stopping condition.
[0163] This embodiment also provides a hybrid precision model fusion training system for multimodal large models, including:
[0164] The sample partitioning module is used to obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partitioned training inputs.
[0165] The precision selection module is used to perform short-range pre-training on the partitioned training input, statistically analyze the training state parameters, and compare the calculated precision sensitivity with a preset sensitivity threshold to select the calculation precision and obtain a mixed precision configuration.
[0166] The calibration and comparison module is used to perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration and comparison to obtain the partition comparison results.
[0167] The gating generation module is used to compare the partition comparison results and calculate the deviation benefit gating value. It matches the deviation benefit gating value with the preset weight range to obtain the fusion weight, and compares the deviation benefit gating value with the preset backoff threshold to obtain the accuracy backoff mark, thus forming the gating fusion parameters.
[0168] The fusion training module is used to fuse the feature outputs of each training partition according to the gating fusion parameters, calculate the fusion training loss and update the parameters of the multimodal large model, configure the training partition corresponding to the accuracy back-down marker to a higher level of calculation accuracy, iterate training until the stopping condition is met, and obtain the multimodal large model that has completed fusion training.
[0169] In summary, this invention incorporates activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value, and parameter update offset value into the same precision sensitivity determination process by normalizing training state parameters and comparing precision sensitivity. This allows modal coding partitions, cross-modal projection partitions, and fusion training partitions to be configured with corresponding computational precision according to the training state, enhancing the partition adaptability, training state correlation, and precision control stability of multimodal large model mixed precision configuration. By forming a unified gating basis through partition comparison results, bias gain gating value, fusion weight, and precision backoff marker, high-precision calibration bias, feature output ratio, and computational precision backoff path are collaboratively controlled in the same machine learning training chain, improving the consistency of feature expression, parameter update stability, and precision preservation capability of multimodal large model fusion training.
[0170] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A mixed-precision model fusion training method for a multi-modal large model, characterized in that: include, Obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partition training inputs; Short-range pre-training is performed on the partitioned training input, training state parameters are statistically analyzed, and after calculating the precision sensitivity, it is compared with a preset sensitivity threshold to select the calculation precision and obtain the mixed precision configuration. Perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration comparison to obtain the partition comparison results; Calculate the deviation benefit threshold value, match the deviation benefit threshold value with the preset weight range to obtain the fusion weight, compare the deviation benefit threshold value with the preset backoff threshold to obtain the accuracy backoff mark, and form the gated fusion parameters; The feature outputs of each training partition are fused according to the gating fusion parameters. The fusion training loss is calculated and the parameters of the multimodal large model are updated. The training partition corresponding to the accuracy back-down mark is configured to a higher level of calculation accuracy. The training is iterated until the stopping condition is met, and the multimodal large model with fusion training is obtained.
2. The hybrid precision model fusion training method for multimodal large models as described in claim 1, characterized in that: The process of pairing multimodal samples for the same training task with task supervision labels specifically includes: Obtain a multimodal training sample set and extract text samples, image samples, audio samples, video samples, and task supervision labels from the same training task; Assign training task numbers to the same training task, and pair text samples, image samples, audio samples, video samples, and task supervision labels according to the training task numbers to obtain task sample groups.
3. The hybrid precision model fusion training method for multimodal large models as described in claim 2, characterized in that: The obtained partition training input specifically includes, Modal parsing is performed on the task sample group to generate corresponding modal identification information; Based on modal identification information, text samples, image samples, audio samples, and video samples are mapped to modal coding partitions respectively, and modal feature encoding processing is performed to obtain modal coding partitions; Cross-modal feature mapping and alignment are performed on the modality-coded partitions to construct cross-modal semantic associations and obtain cross-modal projected partitions; Multimodal feature fusion calculation is performed on the cross-modal projection partition to generate fused training representations. The fused training subspace is then divided based on the feature dimensions of the fused training representations and the mapping relationship between the training paths to form fused training partitions. The task sample groups are arranged according to the training paths corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition training input.
4. The hybrid precision model fusion training method for multimodal large models as described in claim 3, characterized in that: The statistical training state parameters specifically include, According to the preset short-range rounds, the multimodal samples in the partition training input are forward propagated in the modality coding partition, and the layer output activation values of the modality coding partition are recorded. Modal features are obtained by feature aggregation of the layer output activation values; Modal features are input into cross-modal projection partitions for dimensionality mapping, and the dimensionality mapping results are then subjected to semantic space mapping to obtain projected feature values. The projected feature values are input into the fusion training partition for feature fusion, the fusion prediction result is calculated, and the loss is calculated by combining the fusion prediction result with the task supervision label to obtain the loss value; Backpropagation is performed on the loss value to obtain the gradient value of each training partition, and the parameters are updated according to the gradient value to obtain the updated parameter value. The training state parameters are obtained by continuously recording the layer output activation values, projected feature values, loss values, gradient values and parameter update values within a preset short-range round. The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value are calculated.
5. The hybrid precision model fusion training method for multimodal large models as described in claim 4, characterized in that: The obtained mixed precision configuration specifically includes, The activation overflow ratio, cross-modal alignment distance, loss change value, gradient fluctuation value and parameter update offset value in the training state parameters are normalized to obtain the state normalization group. The maximum normalized value in the state normalization group is selected as the accuracy sensitivity of the corresponding training partition; The accuracy sensitivity is compared with a preset sensitivity threshold, and high sensitivity level, medium sensitivity level, and low sensitivity level are marked. The calculation precision is configured according to the sensitivity level labeled in each training partition. The high sensitivity level, medium sensitivity level, and low sensitivity level correspond to the first calculation precision, the second calculation precision, and the third calculation precision, respectively, resulting in a mixed precision configuration.
6. The hybrid precision model fusion training method for multimodal large models as described in claim 5, characterized in that: The obtained partition comparison results specifically include, The computational precision in the mixed precision configuration is mapped to the modality coding partition, cross-modality projection partition, and fusion training partition, and arranged in the execution order of the modality coding partition, cross-modality projection partition, and fusion training partition to obtain the partition precision execution path; Perform mixed-precision forward computation on the partition training input along the partition precision execution path to extract modal features of modal coding partitions, projection features of cross-modal projection partitions, and fusion features of fusion training partitions, and determine the feature output of each training partition; The feature outputs of each training partition are arranged according to the partition precision execution path to obtain the precision calculation results; Perform high-precision calibration operations on the same partition training input according to the first calculation precision, extract the reference modal features of the modal coding partition, the reference projection features of the cross-modal projection partition, and the reference fusion features of the fusion training partition to obtain the high-precision calibration results; The precision calculation results and high-precision calibration results are aligned by partition, and the modal features are paired with the reference modal features, the projection features with the reference projection features, and the fusion features with the reference fusion features according to the same training partition to obtain the partition comparison results.
7. The hybrid precision model fusion training method for multimodal large models as described in claim 6, characterized in that: The calculated deviation payoff threshold specifically includes, The partition comparison results are analyzed, and the difference amplitudes between modal features and reference modal features, projection features and reference projection features, and fusion features and reference fusion features are calculated to obtain feature deviation groups. The feature bias group is normalized to determine the accuracy bias values corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition; The accuracy benefit value corresponding to the accuracy deviation value is determined based on a preset benefit benchmark value. By matching a pre-defined gating mapping relationship between the combination of accuracy deviation value and accuracy gain value, the deviation gain gating value of each training partition is obtained.
8. The hybrid precision model fusion training method for multimodal large models as described in claim 7, characterized in that: The formation of the gating fusion parameters specifically includes, Arrange the bias gain gate values of each training partition according to the partition precision execution path to obtain the gate value sequence; The deviation payoff gate values in the gate value sequence are matched with preset weight intervals, and the fusion weights are configured to decrease from low to high according to the interval values to obtain the fusion weight sequence. Traverse the fusion weight sequence, summarize the fusion weights of each training partition, obtain the weight baseline value, and use the weight baseline value as the proportional benchmark to determine the weight ratio of each training partition, thus obtaining the weight ratio sequence. Compare the bias gain gate value in the gate value sequence with a preset back-off threshold, and label the training partition that reaches the preset back-off threshold with an accuracy back-off flag. The gating fusion parameters are formed by combining the weight ratio sequence of the execution path according to the partition precision and the precision backoff flag.
9. The hybrid precision model fusion training method for multimodal large models as described in claim 8, characterized in that: The obtained multimodal large model after fusion training specifically includes, Based on the weight ratio sequence in the gated fusion parameters, the feature output in the precision calculation result is configured with weights to obtain the partitioned weighted features; The partition-weighted features are input into the fusion training partition for feature fusion, and the fusion prediction result is calculated. The fusion prediction results and task supervision labels are used to calculate the loss, resulting in the fusion training loss; Perform backpropagation on the fusion training loss to update the multimodal large model parameters corresponding to the modality coding partition, cross-modality projection partition, and fusion training partition; Configure the training partition with the precision fallback flag in the gated fusion parameters to a higher level of computational precision, and update the mixed precision configuration. Repeatedly perform mixed-precision training operations, partition comparison, gating fusion parameter formation, and model parameter updates until the preset stopping condition is met, to obtain a multimodal large model that has completed fusion training.
10. A hybrid precision model fusion training system for multimodal large models, based on the hybrid precision model fusion training method for multimodal large models as described in any one of claims 1 to 9, characterized in that: include, The sample partitioning module is used to obtain a multimodal training sample set, pair multimodal samples of the same training task with task supervision labels, and divide the large multimodal model into modality coding partitions, cross-modality projection partitions, and fusion training partitions to obtain partitioned training inputs; The precision selection module is used to perform short-range pre-training on the partitioned training input, statistically analyze the training state parameters, and compare the calculated precision sensitivity with a preset sensitivity threshold to select the calculation precision and obtain a mixed precision configuration. The calibration and comparison module is used to perform mixed-precision training operations according to the mixed-precision configuration, extract the feature output of each training partition, and perform high-precision calibration and comparison to obtain the partition comparison results. The gating generation module is used to compare the partition comparison results and calculate the deviation benefit gating value. It matches the deviation benefit gating value with the preset weight range to obtain the fusion weight, and compares the deviation benefit gating value with the preset backoff threshold to obtain the accuracy backoff mark, thus forming the gating fusion parameters. The fusion training module is used to fuse the feature outputs of each training partition according to the gating fusion parameters, calculate the fusion training loss and update the parameters of the multimodal large model, configure the training partition corresponding to the accuracy back-down marker to a higher level of calculation accuracy, iterate training until the stopping condition is met, and obtain the multimodal large model that has completed fusion training.