Multi-modal large language model establishment method, medium and system

By introducing alignment function and modal contribution allocation model in multimodal feature fusion, the problems of insufficient feature alignment and neglected modal interaction are solved, and higher quality feature fusion and model performance improvement are achieved.

CN120123981AActive Publication Date: 2025-06-10WUHAN TECHN COLLEGE OF COMM

Patent Information

Application Number
CN202510218709.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing multimodal feature fusion technology has problems such as insufficient feature alignment and ignoring modal interaction influence and timing effects, resulting in poor feature fusion effect.

Method used

By introducing alignment function and split index, the alignment effect of feature extraction is evaluated, and a modal contribution allocation model is designed, including feature contribution function, interaction influence function and timing attenuation function, to optimize the feature fusion process.

Benefits of technology

More accurate feature alignment evaluation and optimization is achieved, making full use of the interactive information and timing relationship between modes, and improving the accuracy of feature fusion and the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123981A_ABST
    Figure CN120123981A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal big language model establishing method, medium and system, and belongs to the technical field of computer models.The multi-modal big language model establishing method comprises the steps that firstly, standardized preprocessing is conducted on text, image and audio data; then, constructing a feature extraction structure comprising a bidirectional long-short-term memory network, a residual network and a one-dimensional convolutional network; and evaluating a feature extraction effect through an alignment degree function, calculating a splitting index based on an alignment component matrix, and constructing a modal contribution distribution model to evaluate feature importance. Finally, feature fusion is achieved through an attention mechanism, and model training is completed through a classifier layer and loss function optimization. According to the method, efficient multi-modal feature fusion is realized through feature alignment optimization and contribution value distribution, and the technical problem of poor feature fusion effect caused by insufficient multi-modal feature alignment in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer models, and in particular, relates to a method, medium and system for establishing a multimodal large language model. Background Art

[0002] Multimodal large language models are an important research direction in the field of artificial intelligence. They are mainly used to process the fusion analysis tasks of multiple modal data such as text, images, and audio. Traditional multimodal feature fusion methods mainly include simple concatenation, weighted averaging, and attention mechanisms. The simple concatenation method directly connects the feature vectors of different modalities into a longer vector. Although it is simple to implement, it ignores the correlation between modalities. The weighted average method weightedly combines different modal features by setting fixed weights, but it is difficult to adapt to the dynamic changes in the importance of modalities in different scenarios. Although the attention mechanism-based method can adaptively adjust the feature weights, it is easy to cause information loss in the feature fusion process due to the lack of a clear measurement of the degree of feature alignment.

[0003] The existing multimodal feature fusion technology has the following main problems: First, different modal data have different feature distributions and statistical characteristics, and direct feature fusion can easily lead to information misalignment between modalities. Second, existing methods often ignore the interaction between modal features and cannot fully utilize the complementary information of multimodal data. Third, for the processing of time series data, there is a lack of effective mechanisms to model the impact of historical features on current predictions. In addition, during the feature extraction process, there is also a lack of effective evaluation methods for the contribution of features at different levels to the final prediction results.

[0004] In response to these problems, a method is needed that can effectively measure and optimize the degree of feature alignment, and on this basis, achieve better feature fusion. Especially when processing long-sequence, multi-scale multimodal data, how to ensure the effectiveness of feature extraction and the accuracy of feature fusion is a key issue that needs to be solved urgently. However, due to the lack of effective measurement and optimization mechanism for feature alignment, the existing technology is difficult to achieve high-quality feature fusion, resulting in limited model performance. In other words, there is a technical problem in the existing technology that insufficient multimodal feature alignment leads to poor feature fusion effect. Summary of the invention

[0005] In view of this, the present invention provides a method, medium and system for establishing a multimodal large language model, which can solve the technical problem in the prior art that insufficient multimodal feature alignment leads to poor feature fusion effect.

[0006] The present invention is implemented as follows: A method for establishing a multi-modal large language model provided by the first aspect of the present invention includes the following steps: obtaining text data, image data, and audio data as training input data and performing preprocessing to obtain standardized training data; constructing a multi-layer convolutional neural network structure, setting a text feature extraction layer, an image feature extraction layer, and an audio feature extraction layer; calculating the alignment function between the convolutional kernel parameters of each layer and the standardized training data, constructing an alignment component matrix and a non-alignment component matrix, and calculating the splitting index; constructing a modal contribution allocation model, including a feature contribution function, an interaction influence function, and a temporal decay function, where the modal contribution allocation model is used to calculate the contribution value of each modal feature to the model prediction result; calculating the attention weight based on the contribution value and performing feature fusion to obtain a trained multi-modal large language model.

[0007] Among them, the step of preprocessing the training input data is specifically: performing word segmentation on the text data and converting words into 768-dimensional vector representations through word embedding technology; uniformly adjusting the image data to a size of 224×224 pixels and performing normalization processing to scale the pixel values to between 0 and 1; resampling the audio data to uniformly adjust the sampling rate to 16000 Hz and extracting Mel spectrogram features.

[0008] Among them, in the multi-layer convolutional neural network structure, the text feature extraction layer adopts 3 layers of bidirectional long short-term memory layers, and the number of hidden units in each layer is 512; the image feature extraction layer adopts a 152-layer residual network, including 4 residual blocks; the audio feature extraction layer adopts 6 convolutional layers, and the size of the convolutional kernel in each layer is 3, and the number of channels increases layer by layer from 64 to 512.

[0009] Among them, the step of calculating the alignment function is specifically: calculating the mutual information between the feature map output by each layer of the network and the input data; calculating the cosine similarity matrix between the feature maps; combining the mutual information and the similarity matrix to obtain the alignment function value.

[0010] Among them, in the modal contribution allocation model, the feature contribution function adopts a weighted attention mechanism, and weights the feature importance according to the historical prediction accuracy; the interaction influence function quantifies the interaction between modalities by calculating the mutual information and conditional mutual information between features; the temporal decay function adopts an exponential decay form, and the decay rate is adaptively adjusted according to the time interval.

[0011] Among them, the calculation of the attention weight adopts an 8-head attention mechanism, calculates the attention score through dot-product attention, and uses the softmax function to normalize to obtain the final attention weight.

[0012] Among them, it also includes the step of constructing a classifier layer, and the classifier layer includes 3 fully connected layers, and the dimensions of the hidden layers are 768, 384, and 192 in sequence. After each fully connected layer, a batch normalization layer and a ReLU activation function are connected.

[0013] Among them, during the training process, a combination form of cross-entropy loss and auxiliary loss is adopted, and the auxiliary loss includes an L2 regularization term and an adversarial training loss. When the value of the loss function is less than 0.01 or the accuracy of the validation set reaches 95%, the training is terminated.

[0014] The second aspect of the present invention provides a computer-readable storage medium, and program instructions are stored in the computer-readable storage medium. When the program instructions run on a computer, they are used to execute the above-mentioned method for establishing a multi-modal large language model.

[0015] The third aspect of the present invention provides a system for establishing a multi-modal large language model, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is arranged inside the system.

[0016] Compared with the prior art, the present invention provides a method, medium, and system for establishing a multi-modal large language model. The present invention proposes a multi-modal feature fusion method based on an alignment degree function and modal contribution allocation. This method quantifies the effect of feature extraction by introducing an alignment degree function, constructs an alignment component matrix and a non-alignment component matrix, and evaluates the alignment degree of features through a split exponent. At the same time, a modal contribution allocation model including a feature contribution function, an interaction influence function, and a temporal decay function is designed to achieve an accurate evaluation of the importance of different modal features.

[0017] By introducing the alignment degree function and the split exponent, the present invention can accurately evaluate and optimize the alignment effect in the feature extraction process, and effectively solve the problem of feature misalignment. The modal contribution allocation model realizes a comprehensive evaluation of the importance of features by considering the self-contribution, interaction influence, and temporal effect of features, overcoming the defect of ignoring feature interaction and temporal influence in the existing methods. In addition, the attention weight calculation mechanism based on the contribution value makes the feature fusion process more accurate and interpretable.

[0018] The solution of the present invention effectively solves the problem of poor feature fusion effect caused by insufficient multimodal feature alignment through innovative designs such as feature alignment measurement and optimization, modal contribution allocation, and attention fusion based on contribution values. This solution not only improves the quality of feature extraction but also enhances the accuracy of feature fusion, enabling the model to better handle complex multimodal data analysis tasks. In summary, the present invention solves the technical problem of poor feature fusion effect caused by insufficient multimodal feature alignment in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0021] As Figure 1 shown, it is a flowchart of a method for establishing a multimodal large language model provided in the first aspect of the present invention. This method includes the following steps:

[0022] S01. Obtain text data, image data, and audio data as training input data, and preprocess the training input data to obtain standardized training data;

[0023] S02. Construct a multi-layer convolutional neural network structure, and set a text feature extraction layer, an image feature extraction layer, and an audio feature extraction layer in the multi-layer convolutional neural network structure. Among them, the text feature extraction layer adopts a bidirectional long short-term memory network structure, the image feature extraction layer adopts a residual neural network structure, and the audio feature extraction layer adopts a one-dimensional convolutional neural network structure;

[0024] S03. Calculate the alignment function between the convolutional kernel parameters of each layer in the text feature extraction layer, the image feature extraction layer, and the audio feature extraction layer and the standardized training data;

[0025] S04. Based on the alignment function, construct an alignment component matrix and a non-alignment component matrix, and calculate the splitting index between the alignment component matrix and the non-alignment component matrix;

[0026] S05. Establish a mapping relationship between the splitting index and the network parameters of each layer in the multi-layer convolutional neural network structure;

[0027] S06. Construct a modal contribution allocation model between the multi-layer convolutional neural network structure and the attention mechanism layer. The modal contribution allocation model includes a feature contribution function, an interaction influence function, and a temporal decay function;

[0028] S07. Optimize the parameters of the multi-layer convolutional neural network structure according to the mapping relationship to obtain optimized network parameters;

[0029] S08. Extract features from the standardized training data using the optimized network parameters to obtain multi-modal features;

[0030] S09. Input the multi-modal features into the modal contribution allocation model to calculate the contribution value of each modal feature to the model prediction result;

[0031] S10. Construct an attention mechanism layer and calculate the attention weights between different modal features based on the contribution values;

[0032] S11. Perform weighted fusion on the multi-modal features based on the attention weights to obtain a fused feature vector;

[0033] S12. Construct a classifier layer and input the fused feature vector into the classifier layer for training;

[0034] S13. Calculate the loss function value between the model output result and the labeled data;

[0035] S14. Optimize and update the parameters of the multi-layer convolutional neural network structure, the modal contribution allocation model, the attention mechanism layer, and the classifier layer according to the loss function value;

[0036] S15. Repeat steps S08 to S14 until the loss function value is less than a preset threshold to obtain a trained multi-modal large language model;

[0037] Among them, the specific structure of the modal contribution allocation model is:

[0038] The feature contribution function is used to calculate the self-contribution value of a single modal feature. The input includes a single modal feature in the multi-modal features and the historical prediction accuracy of the single modal feature, and the output is the feature contribution score of the single modal feature;

[0039] The interaction influence function is used to calculate the degree of mutual influence between different modal features. The input includes any two modal features and the joint prediction accuracy of the two modal features, and the output is the interaction contribution score between the two modal features;

[0040] The time-series decay function is used to calculate the influence degree of historical features on the current prediction. The input includes the modal feature at the current moment and the sequence of modal features at historical moments, and the output is the feature influence decay coefficient based on the time distance;

[0041] The modal contribution distribution model comprehensively calculates the final modal feature contribution value by combining the feature contribution score, the interaction contribution score, and the feature influence attenuation coefficient.

[0042] The following is a detailed description of the specific implementation manners of the above steps.

[0043] The specific implementation manner of step S01 is as follows: Standardize and preprocess three different types of training input data. First, perform word segmentation on the text data. Use the Jieba word segmentation algorithm to split the text into a sequence of words, and convert each word into a vector representation of a fixed dimension through word embedding technology. The word embedding dimension is set to 768 dimensions. At the same time, remove stop words and perform lemmatization; for image data, uniformly adjust all images to a size of 224×224 pixels, perform normalization processing to scale the pixel values to between 0 and 1, and perform data augmentation including random cropping, horizontal flipping, brightness and contrast adjustment. The enhanced images are made to satisfy a distribution with a mean of 0 and a variance of 1 through normalization processing; for audio data, first perform resampling to uniformly adjust the sampling rate to 16,000 Hz, then extract Mel spectrogram features, set the frame length to 25 milliseconds, the frame shift to 10 milliseconds, and the number of Mel filter banks to 80. Finally, perform normalization processing on the features. Through the above preprocessing steps, the original data of different modalities are converted into standardized feature representations, laying a foundation for subsequent feature extraction and model training.

[0044] The specific implementation manner of step S02 is as follows: Construct a multi-layer convolutional neural network structure, including three parallel feature extraction branches. The text feature extraction layer uses a bidirectional long short-term memory network structure, including 3 layers of bidirectional long short-term memory layers, with 512 hidden units in each layer. Add residual connections between each layer to alleviate the problem of gradient disappearance, and add layer normalization and dropout layers after each layer. The dropout rate is set to 0.1; the image feature extraction layer uses a residual neural network structure, specifically using a 152-layer residual network, including 4 residual blocks, each residual block containing multiple 3×3 convolutional layers and 1×1 bottleneck layers, and achieving residual learning through shortcut connections; the audio feature extraction layer uses a one-dimensional convolutional neural network structure, including 6 convolutional layers, with a convolutional kernel size of 3 and a stride of 1 in each layer, and the number of channels increasing layer by layer from 64 to 512. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected, and max pooling layers are used for downsampling. This multi-branch network structure can effectively extract the feature representations of different modality data.

[0045] The specific implementation of step S03 is as follows: Calculate the alignment degree between each layer of network parameters and the training data, and measure the effectiveness of feature extraction by introducing an alignment degree function. First, calculate the mutual information between the feature map output by each layer of the network and the input data, using a mutual information calculation method based on kernel density estimation; then calculate the similarity matrix between the feature maps, and use cosine similarity to measure the relationship between features; finally, combine the mutual information and the similarity matrix to obtain the alignment degree function value. The larger this function value is, the more effective the feature extraction is. For the text feature extraction layer, calculate the alignment degree between the word vector sequence and the output of the long short-term memory network; for the image feature extraction layer, calculate the alignment degree between the feature map output by each residual block and the original image; for the audio feature extraction layer, calculate the alignment degree between the feature map output by each convolutional layer and the Mel spectrum feature.

[0046] The specific implementation of step S04 is as follows: Based on the alignment degree function value, construct an alignment component matrix and a non-alignment component matrix. The alignment component matrix contains the features with the alignment degree function value greater than the threshold, and the threshold is set to 0.7. The non-alignment component matrix contains other features. Use the singular value decomposition method to analyze the feature spaces of these two matrices, calculate the singular values corresponding to the main eigenvectors, and define the splitting index based on the ratio of the singular values. Specifically, take the ratio of the sum of the first k largest singular values of the alignment component matrix to the sum of the first k largest singular values of the non-alignment component matrix as the splitting index, where k is set to 20% of the matrix rank. The larger the splitting index is, the better the alignment effect of the feature extraction is.

[0047] The specific implementation of step S05 is as follows: Establish the mapping relationship between the splitting index and the network parameters, and adopt an optimization method based on gradients. First, calculate the gradients of the splitting index with respect to the network parameters of each layer, and then adjust the parameters according to the gradient direction and magnitude. Specifically, use the adaptive moment estimation algorithm for parameter update. For the text feature extraction layer, mainly optimize the weight matrix and bias term in the long short-term memory network; for the image feature extraction layer, optimize the convolution kernel parameters in the residual block; for the audio feature extraction layer, optimize the convolution kernel parameters of the one-dimensional convolutional layer. In this way, the optimization goal of the splitting index is transformed into an optimization problem of the network parameters.

[0048] The specific implementation of step S06 is as follows: construct a modal contribution allocation model, which includes three core functions. The feature contribution function adopts a weighted attention mechanism, weights the feature importance according to the historical prediction accuracy, and the accuracy is calculated using exponential moving average with a smoothing factor set to 0.9; the interaction function quantifies the interaction between modalities by calculating the mutual information and conditional mutual information between features, and uses a non-parametric method based on kernel density estimation to estimate the mutual information; the temporal decay function adopts an exponential decay form, and the decay rate is adaptively adjusted according to the time interval, with a longer interval resulting in faster decay, specifically using the negative power form of the exponential function. The outputs of these three functions are combined through weighted summation to obtain the final modal feature contribution value.

[0049] The specific implementation of step S07 is as follows: optimize the network parameters according to the established mapping relationship, and adopt an iterative optimization method based on gradients. First, calculate the split index under the current parameters, then calculate the gradients of the split index with respect to each layer's parameters, and use the adaptive moment estimation algorithm to update the parameters. The initial learning rate is set to 0.001, and the cosine annealing strategy is used to adjust the learning rate. During the optimization process, the regularization constraints of the parameters are also considered, and L2 regularization is adopted to prevent overfitting, with the regularization coefficient set to 0.0001. For different feature extraction layers, different learning rates and regularization coefficients can be set to adapt to the characteristics of different modal data.

[0050] The specific implementation of step S08 is as follows: use the optimized network parameters to extract features from the standardized training data to obtain multi-modal feature representations. The text features are extracted through a bidirectional long short-term memory network to obtain sequence features with a dimension of 512; the image features are extracted through a residual network to obtain spatial features with a dimension of 2048; the audio features are extracted through a one-dimensional convolutional network to obtain time-frequency features with a dimension of 1024. To unify the dimensions of different modal features, a fully connected layer is used to map all features to the same feature space, which is uniformly set to 768 dimensions. During the feature extraction process, the batch processing method is used to improve the calculation efficiency, and the batch size is set to 32.

[0051] The specific implementation of step S09 is as follows: input the multi-modal features into the modal contribution allocation model to calculate the contribution values of each modality to the prediction result. First, calculate the self-contribution scores of each modal feature through the feature contribution function, considering the discriminative ability and historical performance of the features; then calculate the synergistic effect between modalities through the interaction function to quantify the mutual enhancement or inhibition effect between features; finally, consider the influence of temporal information through the temporal decay function to decay the contribution of historical features. The outputs of the three functions are averaged through weighting to obtain the final contribution value, and the weights are determined through grid search and are set to 0.4, 0.4, and 0.2 respectively.

[0052] The specific implementation of step S10 is as follows: Construct an attention mechanism layer and calculate the attention weights based on the modal contribution values. The multi-head attention mechanism is adopted, with the number of heads set to 8, and each attention head independently learns the correlation between different modal features. The calculation of the attention weights considers three components: the query vector, the key vector, and the value vector. Among them, the query vector and the key vector are generated from the features at the current moment, and the value vector contains historical information. The attention scores are calculated through dot-product attention, and the final attention weights are obtained by normalizing with the softmax function.

[0053] The specific implementation of step S11 is as follows: Weighted fusion of multi-modal features is performed based on the attention weights to obtain a fused feature vector. First, the features of different modalities are projected into the same feature space through linear transformation, and then the weighted sum is calculated according to the attention weights. To enhance the expressive ability of the features, a gating mechanism is added during the fusion process, and the sigmoid function is used to control the information flow. Finally, the final fused feature vector is obtained through residual connection and layer normalization, and the dimension is maintained at 768.

[0054] The specific implementation of step S12 is as follows: Construct a classifier layer, adopting a multi-layer perceptron structure. The classifier contains 3 fully connected layers, with the hidden layer dimensions being 768, 384, and 192 in sequence, and the output dimension of the last layer is equal to the number of classes. A batch normalization layer and a ReLU activation function are connected after each fully connected layer, and dropout is used to prevent overfitting, with the dropout rate set to 0.1. During the training process, label smoothing technology is adopted, and the smoothing factor is set to 0.1 to improve the generalization ability of the model.

[0055] The specific implementation of step S13 is as follows: Calculate the loss function value between the model output result and the labeled data. The loss function adopts a combination form of cross-entropy loss and auxiliary loss. Among them, the cross-entropy loss is used to measure the classification accuracy, and the auxiliary loss includes an L2 regularization term and an adversarial training loss. Adversarial training improves the robustness of the model by adding perturbations, and the size of the perturbations is controlled by the parameter ε, which is set to 0.01. The weights of each part in the total loss function are determined through grid search on the validation set.

[0056] The specific implementation of step S14 is as follows: Optimize the parameters of each component of the model according to the loss function value. The adaptive moment estimation algorithm is used for backpropagation, and the initial learning rate is set to 0.001. To handle the optimization requirements of different components, different learning rates are adopted for different parts. The learning rate of the feature extraction layer is 0.1 times the base learning rate, and the attention layer and the classifier layer use the base learning rate. At the same time, gradient clipping is used to prevent gradient explosion, and the clipping threshold is set to 1.0. After each epoch, the model performance is evaluated on the validation set. If the performance does not improve for 5 consecutive epochs, the learning rate is decreased, and the decrease factor is 0.1.

[0057] The specific implementation of step S15 is as follows: The processes of feature extraction, contribution value calculation, attention fusion, classification prediction, and parameter optimization are repeatedly executed until the loss function value is less than a preset threshold. The preset threshold is set to 0.01, or the training can also be terminated early when the accuracy on the validation set reaches 95%. During the training process, the order of the training data is randomly shuffled in each epoch, and the model checkpoints are saved regularly. When the model converges, the checkpoint with the best performance on the validation set is selected as the final model. The trained multi-modal large language model can effectively integrate information from different modalities and achieve accurate reasoning.

[0058] The equations or mathematical models involved in the present invention will be explained in detail below.

[0059] The calculation of the word embedding vector is represented as follows:

[0060] E = WV + b;

[0061] In the formula, E is the word embedding vector; W is the weight matrix; V is the one-hot encoded vector; b is the bias term. Among them, the dimension of W is 768 × the size of the vocabulary, which is obtained through pre-training; the dimension of b is 768 × 1, and it is initialized to 0.

[0062] The principle of establishing the above word embedding vector calculation equation is as follows: First, a vocabulary mapping matrix is constructed to represent each word as a one-hot vector with a dimension equal to the size of the vocabulary; then, the weight matrix W is introduced for linear transformation to map the high-dimensional sparse vector to a low-dimensional dense space; finally, the bias term b is added to increase the flexibility of the model. The weight matrix W is obtained through pre-training on a large-scale corpus, using algorithms such as word2vec or GloVe, and is optimized by minimizing the prediction loss of the context words. The effect of this equation is to encode the semantic information of words into a vector space with a fixed dimension, so that words with similar semantics are closer in the vector space.

[0063] The calculation of image normalization is represented as follows:

[0064]

[0065] In the formula, I norm is the normalized image; I is the original image; μ is the image mean; σ is the standard deviation; ∈ is a small value to prevent division by zero, taking 1e-5.

[0066] The principle of establishing the above image normalization equation is as follows: Based on the standardization idea in statistics, first calculate the mean μ and standard deviation σ of the image pixel values, where the mean reflects the overall brightness level and the standard deviation reflects the contrast level. By subtracting the mean and dividing by the standard deviation, the data distribution is adjusted to a standard distribution with a mean of 0 and a variance of 1. Adding the ε term is to prevent numerical instability problems when the standard deviation approaches 0. The effect of this equation is to eliminate the brightness and contrast differences of the image, making the feature distributions of different images more consistent.

[0067] The calculation of audio Mel-spectrum feature extraction is expressed as follows:

[0068] M = H · |STFT(x)| 2 ;

[0069] In the formula, M is the Mel-spectrum feature; H is the Mel filter bank matrix; STFT(x) is the short-time Fourier transform; and x is the audio signal.

[0070] The principle of establishing the above audio Mel-spectrum feature extraction equation is as follows: First, perform a short-time Fourier transform (STFT) on the audio signal to obtain a time-frequency representation; then square the spectrum to get the power spectrum; finally, perform weighted summation through the Mel filter bank H. The Mel filter bank is designed based on the human ear's auditory characteristics, with higher resolution in the low-frequency region and lower resolution in the high-frequency region. The effect of this equation is to convert the audio signal into spectrum features that are more in line with the human ear's perception characteristics.

[0071] The forward calculation of the bidirectional long short-term memory network is expressed as follows:

[0072] f t = σ(W f · [h t-1 , x t + b f );

[0073] i t = σ(W i · [h t-1 , x t + b i );

[0074]

[0075] o t = σ(W o · [h t-1 , x t + b o );

[0076] h t = o t · tanh(C t );

[0077] In the formula, f t is the forget gate; i t is the input gate; o t is the output gate; C t is the cell state; h t is the hidden state; W f , W i , W C , W o are weight matrices; b f , b i , b C , b o are bias terms; σ is the sigmoid function; x t is the input vector.

[0078] The establishment principle of the above bidirectional long short-term memory network equations is as follows: Based on the structure of the recurrent neural network, three gating mechanisms, namely the forget gate, the input gate, and the output gate, are introduced. The forget gate controls the degree of forgetting of historical information, the input gate controls the degree of input of new information, and the output gate controls the degree of output of information. Each gate maps the output to between 0 and 1 through the sigmoid function to achieve soft gating. The cell state C is updated through the combination of the forget gate and the input gate, and the hidden state h is generated through the combination of the output gate and the cell state. The weight matrices W and the bias terms b are optimized through the backpropagation algorithm. The effect of this set of equations is to effectively process long sequence data and capture long-term dependencies.

[0079] The calculation representation of the residual block in the residual network is as follows:

[0080] y = F(x, W i ) + x;

[0081] In the formula, y is the output of the residual block; F(x, W i ) is the residual mapping function; x is the input; W i is the weight matrix of the i-th layer.

[0082] The establishment principle of the residual block equation in the above residual network is as follows: Based on the traditional convolutional network, a shortcut connection is introduced, and the input is directly added to the output of the non-linear transformation F(x). F(x) includes a combination of multiple convolutional layers, batch normalization layers, and activation functions. The effect of this equation is to alleviate the gradient vanishing problem of deep networks, making it easier to train deep networks.

[0083] The calculation representation of the one-dimensional convolutional layer is as follows:

[0084]

[0085] In the formula, y j is the j-th output feature map; xi is the i-th input feature map; k i,j is the convolution kernel; b j is the bias term; M is the number of input channels; * represents the convolution operation.

[0086] The principle of establishing the above one-dimensional convolutional layer equation is as follows: Based on the definition of convolution operation, a sliding window operation is performed on the input feature map, and the inner product of the local region and the convolution kernel is calculated at each position, and the bias term is added. The convolution kernel parameters are optimized through backpropagation, and the bias term is used to adjust the overall offset of the feature map. The effect of this equation is to extract local feature patterns and has translational invariance.

[0087] The calculation of the alignment degree function is expressed as follows:

[0088] A(F, D) = αMI(F, D) + (1 - α)S(F);

[0089] In the formula, V(F, D) is the value of the alignment degree function; F is the feature map; D is the input data; MI(F, D) is the mutual information; S(F) is the feature similarity; α is the weight coefficient, and its value range is from 0 to 1.

[0090] The calculation of the mutual information is expressed as follows:

[0091]

[0092] In the formula, p(x, y) is the joint probability distribution; p(x) and p(y) are the marginal probability distributions.

[0093] The calculation of the feature similarity is expressed as follows:

[0094]

[0095] In the formula, F is the feature matrix; ||·|| represents the Frobenius norm of the matrix.

[0096] The principle of establishing the above alignment degree function equation is as follows: Combining the two indicators of mutual information and feature similarity, a weighted average is performed through the weight α. The mutual information measures the statistical correlation between the feature and the input data, and the feature similarity measures the structural consistency of the feature space. α is determined by grid search on the validation set. The effect of this equation is to comprehensively evaluate the effectiveness of feature extraction.

[0097] The construction of the aligned component matrix and the unaligned component matrix is expressed as follows:

[0098] M aligned = [F i |A(F i , D)>θ];

[0099] M unaligned = [Fi |A(F i , D) ≤ θ];

[0100] Wherein, M aligned is the alignment component matrix; M unaligned is the non-alignment component matrix; θ is a threshold value, set to 0.7.

[0101] The calculation of the splitting index is expressed as follows:

[0102]

[0103] Wherein, SI is the splitting index; is the i-th singular value of the alignment component matrix; is the i-th singular value of the non-alignment component matrix; k is the number of singular values taking the top 20%.

[0104] The calculation of the feature contribution function is expressed as follows:

[0105] C f (x) = w · x · acc hist ;

[0106] Wherein, C f (x) is the feature contribution score; x is the feature vector; w is the attention weight; acc hist is the historical prediction accuracy rate, calculated by exponential moving average.

[0107] The principle of establishing the above feature contribution function equation is: multiplying the feature vector by the attention weight and the historical accuracy rate, the attention weight is calculated by the softmax function, and the historical accuracy rate is updated by exponential moving average. The effect of this equation is to dynamically adjust the contribution according to the importance of the feature and its historical performance.

[0108] The calculation of the interaction influence function is expressed as follows:

[0109] C i (x 1 , x 2 ) = MI(x 1 , x 2 |y);

[0110] Wherein, C i (x 1 , x 2 ) is the interaction contribution score; MI(x 1 , x 2 |y) is the conditional mutual information; y is the prediction label.

[0111] The principle of establishing the above interaction influence function equation is as follows: The conditional mutual information is used to measure the interaction between different modal features, and the conditional probability distribution of the predicted labels is considered. The effect of this equation is to quantify the synergistic effect between features.

[0112] The calculation of the time series decay function is expressed as follows:

[0113] D(t) = e -λt ;

[0114] In the formula, D(t) is the decay coefficient; λ is the decay rate, which is positively correlated with the time interval t.

[0115] The principle of establishing the above time series decay function equation is as follows: The exponential function form is adopted, and the decay rate λ is positively correlated with the time interval. The effect of this equation is to simulate the natural decay process of feature influence over time.

[0116] The calculation of the final modal contribution value is expressed as follows:

[0117] C total = β 1 C f + β 2 C i + β 3 D;

[0118] In the formula, C total is the final contribution value; β 1 , β 2 , β 3 are the combination weights, which are determined by grid search.

[0119] The calculation of the multi-head attention mechanism is expressed as follows:

[0120]

[0121] In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; d k is the dimension of the key vector.

[0122] The principle of establishing the above multi-head attention mechanism equation is as follows: The scaled dot-product attention calculation is performed on the three matrices of query, key, and value, and the attention weights are obtained by normalizing through the softmax function. The scaling factor is used to prevent the gradient of the softmax function from vanishing due to the excessive dot-product result. The effect of this equation is to learn the dynamic correlation relationship between features.

[0123] The calculation of the total loss function is expressed as follows:

[0124] L total = L CE + λ 1 L L2 + λ2 L adv ;

[0125] Wherein, L total is the total loss; L CE is the cross - entropy loss; L L2 is the L2 regularization loss; L adv is the adversarial training loss; λ 1 , λ 2 is the weight coefficient.

[0126] The calculation of the adversarial training loss is shown as follows:

[0127] L adv = max ||δ||≤∈ L(x + δ, y);

[0128] Wherein, δ is the adversarial perturbation; ∈ is the upper limit of the perturbation size, set to 0.01; x is the input sample; y is the true label.

[0129] The establishment principle of the above - mentioned total loss function equation is: combining the cross - entropy loss, L2 regularization loss and adversarial training loss, and the weight coefficient is adjusted through the performance of the validation set. The cross - entropy loss is used to optimize the classification accuracy, L2 regularization is used to prevent overfitting, and adversarial training is used to improve the robustness of the model. The effect of this equation is to achieve multi - objective optimization and balance the performance indicators of the model.

[0130] The construction of the above - mentioned formulas fully considers the characteristics and requirements of each processing link. The power relationship is used to reflect the non - linear mapping of features, the division relationship is used for normalization processing, the exponential relationship is used to express the attenuation effect, and the summation relationship is used to integrate the contributions of multiple features. Using the logarithmic relationship in the mutual information calculation conforms to the basic principles of information theory, and using the dot product and softmax in the attention mechanism conforms to the probability characteristics of the attention distribution. Each formula contains adjustable parameters and necessary constraint terms to improve the adaptability and robustness of the model. Compared with the existing technology, this solution has innovations in feature alignment measurement, modal contribution allocation and loss function design, etc., and can better handle the fusion problem of multi - modal data.

[0131] The second aspect of the present invention provides a computer - readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above - mentioned method for establishing a multi - modal large - language model.

[0132] The third aspect of the present invention provides a system for establishing a multi - modal large - language model, including the above - mentioned computer - readable storage medium. The system can be any one of a computer, a server, and a single - chip microcomputer. The computer - readable storage medium is set inside the system, and a micro - processor for executing the program instructions stored in the computer - readable storage medium is set inside the system.

[0133] Specifically, the principle of the present invention is as follows: The technical principle of the present invention is mainly based on the following aspects. First, an alignment function is introduced as a quantitative index to evaluate the effect of feature extraction. This function combines two dimensions of mutual information and feature similarity, and can comprehensively reflect the correspondence between features and the original data. By calculating the mutual information between the feature map and the input data, the degree to which the features retain the original information can be measured; by calculating the similarity matrix between the feature maps, the structural consistency of the feature space can be evaluated.

[0134] Secondly, based on the alignment component matrix and the non-alignment component matrix constructed by the alignment function, the splitting index is calculated by singular value decomposition, providing a clear optimization goal for the feature extraction process. The larger the splitting index, the better the alignment effect, which provides a clear direction for optimizing the network parameters.

[0135] Thirdly, the design of the modal contribution allocation model fully considers multiple influencing dimensions of features. The feature contribution function realizes the dynamic evaluation of feature importance by combining attention weights and historical performance; the interaction influence function describes the interaction between features through conditional mutual information; the temporal decay function simulates the natural decay law of feature influence over time. This multi-dimensional contribution evaluation mechanism makes the feature fusion process more accurate and reasonable.

[0136] The following provides a specific Embodiment 1 of the present invention. The specific implementation of each step in this Embodiment 1 is described in detail as follows.

[0137] The specific implementation of step S01 is: perform standardized preprocessing on three different types of training input data. Among them, text data is represented by word embedding vectors, and the specific calculation process is as follows:

[0138] E = WV + b;

[0139] In the formula, E is the word embedding vector with a dimension of 768; W is the weight matrix with a dimension of 768×vocabulary size, which is pre-trained on a large-scale corpus through the word2vec algorithm; V is the one-hot encoded vector with a dimension of vocabulary size×1; b is the bias term with a dimension of 768×1, initialized to 0. For the standardized processing of image data, the following calculation is adopted:

[0140]

[0141] In the formula, I norm is the normalized image; I is the original image; μ is the image mean; σ is the standard deviation; ∈ is a small value to prevent division by zero, taking 1e-5. For audio data, Mel spectrogram feature extraction is calculated as follows:

[0142] M = H · |STFT(x)| 2 ;

[0143] Where M is the Mel spectrum feature; H is the Mel filter bank matrix; STFT(x) is the short-time Fourier transform; and x is the audio signal. During the audio feature extraction process, first, the sampling rate is uniformly adjusted to 16,000 Hz, then the frame length is set to 25 milliseconds, the frame shift is set to 10 milliseconds, and the number of Mel filter banks is 80. The main function of this step is to convert the original data of different modalities into a standardized feature representation, laying a foundation for subsequent feature extraction and model training.

[0144] The specific implementation of step S02 is: construct a multi-layer convolutional neural network structure, and the text feature extraction layer uses a bidirectional long short-term memory network structure. Its calculation process is as follows:

[0145] f t = σ(W f · [h t-1 , x t + b f );

[0146] i t = σ(W i · [h t-1 , x t + b i );

[0147]

[0148] o t = σ(W o · [h t-1 , x t + b o );

[0149] h t = o t · tanh(C t );

[0150] Where f t is the forget gate; i t is the input gate; o t is the output gate; C t is the cell state; h t is the hidden state; W f , W i , W C , W o are weight matrices; b f , b i , b C , b o are bias terms; σ is the sigmoid function; xt is the input vector. The image feature extraction layer adopts a residual network structure, and the calculation process of the residual block is as follows:

[0151] y = F(x, W i ) + x;

[0152] where y is the output of the residual block; F(x, W i ) is the residual mapping function; x is the input; W i is the weight matrix of the i-th layer. The audio feature extraction layer adopts a one-dimensional convolutional neural network structure, and its calculation process is as follows:

[0153]

[0154] where y j is the j-th output feature map; x i is the i-th input feature map; k i,j is the convolution kernel; b j is the bias term; M is the number of input channels; * represents the convolution operation. The main function of this step is to extract the deep feature representations of each modality data through a deep neural network.

[0155] The specific implementation of step S03 is: calculate the alignment function between the convolution kernel parameters of each layer in the feature extraction layer and the standardized training data, and its calculation process is as follows:

[0156] A(F, D) = αMI(F, D) + (1 - α)S(F);

[0157] where A(F, D) is the alignment function value; F is the feature map; D is the input data; MI(F, D) is the mutual information; S(F) is the feature similarity; α is the weight coefficient, and its value range is from 0 to 1. The specific calculation of the mutual information is:

[0158]

[0159] where p(x, y) is the joint probability distribution; p(x) and p(y) are the marginal probability distributions. The calculation of the feature similarity is:

[0160]

[0161] where F is the feature matrix; ||·|| represents the Frobenius norm of the matrix. The main function of this step is to evaluate the effectiveness of feature extraction and ensure that the extracted features can retain the key information of the original data.

[0162] The specific implementation of step S04 is: construct the alignment component matrix and the non-alignment component matrix based on the alignment function, and its construction process is as follows:

[0163] Maligned = [F i |A(F i , D)>θ];

[0164] M unaligned = [F i |A(F i , D)≤θ];

[0165] Where, M aligned is the alignment component matrix; M unaligned is the non - alignment component matrix; θ is the threshold, set to 0.7. Then calculate the splitting index:

[0166]

[0167] Where, SI is the splitting index; is the i - th singular value of the alignment component matrix; is the i - th singular value of the non - alignment component matrix; k is the number of singular values taking the top 20%. The main function of this step is to quantify the alignment degree of features and provide a basis for subsequent parameter optimization.

[0168] The specific implementation of step S05 is: establish a mapping relationship between the splitting index and network parameters, and adopt a gradient - based optimization method. The process of establishing the mapping relationship includes calculating the gradient of the splitting index with respect to each layer parameter:

[0169]

[0170] Where, W is the network parameter. Parameter update adopts the adaptive moment estimation algorithm:

[0171] m t = β 1 m t-1 +(1 - β 1 )g t ;

[0172]

[0173] Where, m t and v t are the first - order and second - order momenta; β 1 and β 2 are the momentum decay rates; η is the learning rate; ∈ is the numerical stability constant. The main function of this step is to improve the alignment effect of feature extraction by optimizing network parameters.

[0174] The specific implementation of step S06 is: construct a modal contribution allocation model, including a feature contribution function, an interaction influence function, and a temporal decay function. The calculation process of the feature contribution function is:

[0175] Cf C(x) = w·x·acc hist ;

[0176] Wherein, C f C(x) is the feature contribution score; x is the feature vector; w is the attention weight; acc hist is the historical prediction accuracy. The historical prediction accuracy is calculated using exponential moving average:

[0177] acc hist = βacc hist-1 + (1 - β)acc t ;

[0178] Wherein, β is the smoothing factor, set to 0.9; acc t is the prediction accuracy at the current moment. The calculation process of the interaction influence function is:

[0179] C i (x 1 , x 2 ) = MI(x 1 , x 2 |y);

[0180] Wherein, C i (x 1 , x 2 ) is the interaction contribution score; MI(x 1 , x 2 |y) is the conditional mutual information; y is the prediction label. The calculation process of the time series decay function is:

[0181] D(t) = e -λt ;

[0182] Wherein, D(t) is the decay coefficient; λ is the decay rate, which is positively correlated with the time interval t. The final modal contribution value is calculated as:

[0183] C total = β 1 C f + β 2 C i + β 3 D;

[0184] Wherein, C total is the final contribution value; β 1 , β 2 , β 3 are the combined weights, determined by grid search. The main function of this step is to evaluate the contribution degree of different modal features to the prediction result.

[0185] The specific implementation of step S07 is as follows: optimize the parameters of the multi-layer convolutional neural network structure according to the established mapping relationship, and adopt a gradient-based iterative optimization method. The objective function for parameter optimization is:

[0186] L opt = I SI + γL reg ;

[0187] In the formula, L opt is the optimization objective function; I SI is the loss based on the split exponent; L reg is the regularization term; γ is the regularization coefficient, set to 0.0001. The parameter update adopts an adaptive learning rate adjustment strategy:

[0188]

[0189] In the formula, η t is the learning rate at time t; η 0 is the initial learning rate, set to 0.001; T is the total number of iterations. The main function of this step is to optimize the network parameters and improve the effect of feature extraction.

[0190] The specific implementation of step S08 is as follows: use the optimized network parameters to extract features from the standardized training data to obtain multi-modal features. The calculation process of text feature extraction is:

[0191]

[0192] In the formula, is the text feature; E t is the word embedding sequence; W lstm is the long short-term memory network parameter. The calculation process of image feature extraction is:

[0193]

[0194] In the formula, is the image feature; I t is the image input; W res is the residual network parameter. The calculation process of audio feature extraction is:

[0195]

[0196] In the formula, is the audio feature; M t is the Mel spectrogram feature; W conv is the convolutional network parameter. The main function of this step is to obtain the feature representations of each modality.

[0197] The specific implementation of step S09 is as follows: Input the multi-modal features into the modal contribution distribution model to calculate the feature contribution value. The calculation of the contribution value includes three aspects: self-contribution, interaction contribution, and temporal influence. The calculation process of self-contribution is as follows:

[0198]

[0199] In the formula, C self is the self-contribution value; n is the number of features; w i is the feature weight. The calculation process of interaction contribution is as follows:

[0200]

[0201] In the formula, C inter is the interaction contribution value. The calculation process of temporal influence is as follows:

[0202]

[0203] In the formula, C temp is the temporal contribution value; T is the size of the historical time window. The main function of this step is to quantify the influence degree of each modal feature on the prediction result.

[0204] The specific implementation of step S10 is as follows: Construct an attention mechanism layer and calculate the attention weight based on the contribution value. The calculation process of the multi-head attention mechanism is as follows:

[0205]

[0206] In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; d k is the dimension of the key vector. The calculation process of the attention weight is as follows:

[0207]

[0208] In the formula, α ij is the attention weight; e ij is the attention score. The integration process of the multi-head attention is as follows:

[0209] MultiHead(Q, K, V) = Concat(head 1 , …, head h )W O ;

[0210] In the formula, head i is the output of the i-th attention head; W O is the output mapping matrix. The main function of this step is to learn the dynamic correlation relationship between different modal features.

[0211] The specific implementation of step S11 is as follows: based on the attention weights, the multi-modal features are weighted and fused to obtain a fused feature vector. The calculation process of feature weighted fusion is as follows:

[0212]

[0213] In the formula, h fusion is the fused feature vector; α i is the attention weight of the i-th feature; h i is the i-th modal feature. To enhance the expressive ability of the features, a gating mechanism is introduced:

[0214] g i = σ(W g h i + b g );

[0215] h gated = g i ⊙ h i ;

[0216] In the formula, g i is the gating value; W g is the gating weight matrix; b g is the bias term; ⊙ represents element-wise multiplication. Finally, through residual connection and layer normalization:

[0217] h final = LayerNorm(h fusion + h gated );

[0218] In the formula, h final is the final fused feature vector. The main function of this step is to effectively combine the features of different modalities into a unified representation.

[0219] The specific implementation of step S12 is as follows: construct a classifier layer, adopting a multi-layer perceptron structure. The forward propagation calculation process of the classifier is as follows:

[0220] z 1 = ReLU(W 1 h final + b 1 );

[0221] z 2 = ReLU(W 2 z 1 + b 2 );

[0222] z 3 = W 3 z 2 + b 3 ;

[0223] where z 1 , z 2 , z 3 are the outputs of each layer; W 1 , W 2 , W 3 are the weight matrices; b 1 , b 2 , b 3 are the bias terms. The calculation process of label smoothing is:

[0224] y smooth = (1 - ∈)y + ∈ / K;

[0225] where y smooth is the smoothed label; y is the original label; ∈ is the smoothing factor, set to 0.1; K is the number of classes. The main function of this step is to map the fused features to the class space to achieve the final classification prediction.

[0226] The specific implementation of step S13 is: calculate the loss function value between the model output result and the labeled data. The calculation process of cross-entropy loss is:

[0227]

[0228] where y i is the true label; is the predicted probability. The calculation process of L2 regularization loss is:

[0229]

[0230] where λ is the regularization coefficient; N is the number of parameters. The calculation process of adversarial training loss is:

[0231] L adv = max ||δ||≤∈ L(x + δ, y);

[0232] where δ is the adversarial perturbation; ∈ is the upper limit of the perturbation size, set to 0.01. The calculation of the total loss function is:

[0233] L total = L CE + λ 1 L L2 + λ 2 L adv ;

[0234] where λ 1 , λ 2 are the weight coefficients. The main function of this step is to evaluate the training effect of the model and provide guidance for parameter optimization.

[0235] The specific implementation of step S14 is as follows: Optimize and update the parameters of each component of the model according to the loss function value. The calculation process of the parameter gradient is as follows:

[0236]

[0237] In the formula, g t is the gradient at time t; W t is the model parameter. The calculation process of gradient clipping is as follows:

[0238]

[0239] In the formula, θ is the clipping threshold, set to 1.0. The calculation process of learning rate adjustment is as follows:

[0240]

[0241] In the formula, η 0 is the initial learning rate; γ is the decay factor, set to 0.1; s is the decay step. The main function of this step is to optimize the model parameters through gradient descent.

[0242] The specific implementation of step S15 is as follows: Repeat the processes of feature extraction, contribution value calculation, attention fusion, classification prediction, and parameter optimization until convergence. The calculation process of convergence judgment is as follows:

[0243] ΔL = |L t -L t-1 | < ∈;

[0244] In the formula, ΔL is the change in the loss function; ∈ is the convergence threshold, set to 0.01. The calculation process of validation set performance evaluation is as follows:

[0245]

[0246] In the formula, Acc val is the validation set accuracy; N is the number of samples; I(·) is the indicator function. The judgment condition for model saving is:

[0247]

[0248] In the formula, is the validation set accuracy at time t. The main function of this step is to obtain the optimal model parameters through iterative training.

[0249] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: During the development of a new generation of multilingual interaction assistants, a research team adopted the method for establishing a multimodal large language model of the present invention for technical implementation. This assistant needs to simultaneously process text instructions, image scenes, and voice conversations input by users. To achieve efficient multimodal information understanding and response, the research team collected a training set containing 100,000 pieces of multimodal interaction data. Among them, the text data includes user instructions and system responses, with an average length of 50 words; the image data is scene pictures, uniformly adjusted to a size of 224×224 pixels; the audio data is user voices, with a sampling rate of 16,000 Hz.

[0250] The research team first preprocesses the training data. For the text data, the Jieba word segmentation tool is used for word segmentation, and each word is converted into a 768-dimensional vector representation through a pre-trained word embedding model. For the image data, normalization processing is performed to scale the pixel values to between 0 and 1, and data augmentation methods such as random cropping and horizontal flipping are adopted to expand the training samples. For the audio data, the Mel spectrum features of 80 Mel filter banks are extracted, the frame length is set to 25 milliseconds, and the frame shift is 10 milliseconds.

[0251] In the construction of the feature extraction network, the text feature extraction layer adopts a 3-layer bidirectional long short-term memory network, with the number of hidden units in each layer being 512 and the dropout rate set to 0.1. The image feature extraction layer adopts a 152-layer residual network, including 4 residual blocks, and each residual block contains multiple 3×3 convolutional layers and 1×1 bottleneck layers. The audio feature extraction layer adopts a 6-layer one-dimensional convolutional network, with the convolutional kernel size being 3 and the number of channels increasing layer by layer from 64 to 512.

[0252] The research team evaluates the feature extraction effect and calculates the alignment function values shown in Table 1 below:

[0253] Table 1 Alignment function values

[0254] Modal type The first layer The second layer The third layer The fourth layer The fifth layer The sixth layer Text 0.82 0.79 0.75 - - - Image 0.85 0.81 0.78 0.73 - - Audio 0.83 0.80 0.76 0.72 0.69 0.65

[0255] Based on the alignment function values, an alignment component matrix and a non-alignment component matrix are constructed, and the alignment threshold is set to 0.7. The calculated splitting index of the text features is 1.85, the splitting index of the image features is 1.92, and the splitting index of the audio features is 1.78.

[0256] In the modal contribution allocation model, the feature contribution function adopts a weighted attention mechanism, and the smoothing factor of the historical prediction accuracy is set to 0.9. The interaction influence function quantifies the interaction between modalities by calculating the conditional mutual information, and the decay rate of the temporal decay function is set to 0.2. Finally, the contribution value distributions of the three modalities are shown in Table 2 below:

[0257] Table 2 Contribution Value Distribution

[0258] Interaction scenario Text contribution value Image contribution value Audio contribution value Instruction understanding 0.65 0.20 0.15 Scene description 0.25 0.60 0.15 Voice dialogue 0.30 0.15 0.55

[0259] The attention mechanism layer adopts an 8-head attention structure, and the dimension of each attention head is 96. During the feature fusion process, a gating mechanism is used to control the information flow, and the hidden layer dimension of the gating network is set to 384. The classifier layer contains 3 fully connected layers with dimensions of 768, 384, and 192 in sequence. After each layer, a batch normalization layer and a ReLU activation function are connected, and the dropout rate is set to 0.1.

[0260] During the training process, a batch processing method with a batch size of 32 is adopted, the initial learning rate is set to 0.001, and the cosine annealing strategy is used to adjust the learning rate. In the loss function, the cross-entropy loss weight is 1.0, the L2 regularization coefficient is 0.0001, and the perturbation size of adversarial training is set to 0.01. The training process data on the validation set is shown in Table 3 below:

[0261] Table 3 Training Data Table

[0262] Number of training rounds Loss value Accuracy 10 0.42 0.85 20 0.35 0.89 30 0.28 0.92 40 0.22 0.94 50 0.18 0.95

[0263] When the accuracy rate of the validation set reaches 95%, the model training is completed. Compared with the traditional multi-modal fusion method, the method previously adopted by the research team was a simple feature splicing and weighted average method with fixed weights, and it could only reach an accuracy rate of 88% on the same validation set. The traditional method mainly has the following problems: the lack of alignment evaluation in the feature extraction process leads to uneven feature quality of different modalities; the fixed weights are used in the feature fusion process, which cannot adapt to the dynamic changes of modality importance in different scenarios; the interaction and temporal effects between features are not considered, resulting in insufficient information utilization.

[0264] The present invention realizes the quality evaluation and optimization of the feature extraction process by introducing an alignment degree function and a splitting exponent; realizes the dynamic evaluation of feature importance through a modality contribution allocation model; and realizes more accurate feature fusion through a multi-head attention mechanism and a gating mechanism. These innovative designs effectively solve the problem of poor fusion effect caused by insufficient feature alignment, and enable the model to achieve a significant improvement in the multi-modal interaction understanding task.

[0265] In practical applications, this multi-modal interaction assistant can accurately understand the user's text instructions, image scenes, and voice conversations, and make reasonable responses according to the importance of each modality in different scenarios. Especially in complex multi-modal interaction scenarios, the model shows strong robustness and adaptability. The improvement of this effect benefits from the innovative designs of the present invention in feature alignment optimization and modality contribution allocation.

[0266] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Table 4 below.

[0267] Table 4 Variable Explanation Table

[0268]

[0269]

[0270] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A method for establishing a multimodal large language model, characterized in that: The following steps are involved: Acquire text data, image data and audio data as training input data and perform preprocessing to obtain standardized training data; construct a multi-layer convolutional neural network structure, and set a text feature extraction layer, an image feature extraction layer and an audio feature extraction layer; Calculate the alignment function between the convolution kernel parameters of each layer and the standardized training data, construct the aligned component matrix and the non-aligned component matrix and calculate the splitting index; construct a modal contribution allocation model, including a feature contribution function, an interaction influence function and a time series attenuation function, and the modal contribution allocation model is used to calculate the contribution value of each modal feature to the model prediction result; Attention weights are calculated based on the contribution values ​​and feature fusion is performed to obtain a trained multimodal large language model.

2. The method for establishing a multimodal large language model according to claim 1, characterized in that: The steps for preprocessing the training input data are as follows: segment the text data and convert the words into 768-dimensional vector representations through word embedding technology; adjust the image data to a uniform size of 224×224 pixels and normalize the pixel values ​​to scale between 0 and 1; resample the audio data to adjust the sampling rate to 16000 Hz and extract the Mel spectrum features.

3. The method for establishing a multimodal large language model according to claim 1, characterized in that: In the multi-layer convolutional neural network structure, the text feature extraction layer adopts 3 layers of bidirectional long short-term memory layers, and the number of hidden units in each layer is 512; the image feature extraction layer adopts a 152-layer residual network, including 4 residual blocks; The audio feature extraction layer uses 6 convolutional layers, the convolution kernel size of each layer is 3, and the number of channels increases from 64 to 512 layer by layer.

4. The method for establishing a multimodal large language model according to claim 1, characterized in that: The steps for calculating the alignment function are as follows: calculating the mutual information between the feature map output by each layer of the network and the input data; calculating the cosine similarity matrix between the feature maps; and combining the mutual information and the similarity matrix to obtain the alignment function value.

5. The method for establishing a multimodal large language model according to claim 1, characterized in that: In the modal contribution allocation model, the feature contribution function adopts a weighted attention mechanism to weight the feature importance according to the historical prediction accuracy; the interaction influence function quantifies the interaction between modalities by calculating the mutual information and conditional mutual information between features; the time series decay function adopts an exponential decay form, and the decay rate is adaptively adjusted according to the time interval.

6. The method for establishing a multimodal large language model according to claim 1, characterized in that: The attention weight is calculated using an 8-head attention mechanism. The attention score is obtained by dot product attention calculation, and the final attention weight is obtained by normalization using the softmax function.

7. The method for establishing a multimodal large language model according to claim 1, characterized in that: The method also includes the step of constructing a classifier layer, wherein the classifier layer includes three fully connected layers, and the hidden layer dimensions are 768, 384, and 192, respectively, and each fully connected layer is followed by a batch normalization layer and a ReLU activation function.

8. The method for establishing a multimodal large language model according to claim 1, characterized in that: A combination of cross entropy loss and auxiliary loss is used in the training process, where the auxiliary loss includes L2 regularization term and adversarial training loss. The training is terminated when the loss function value is less than 0.01 or the verification set accuracy reaches 95%.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed in a computer, they are used to execute the method for establishing a multimodal large language model according to any one of claims 1 to 8.

10. A multimodal large language model building system, characterized in that: The system comprises the computer-readable storage medium as claimed in claim 9, wherein the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.

Citation Information

Patent Citations

  • Target-oriented multi-modal sentiment classification method

    CN113065577A

  • Speech synthesis method and speech synthesis model training method and device

    CN113112987A

  • Multi-modal data fusion method based on compound collaborative structure feature recombination network

    CN113378989A

  • Multi-mode classroom teacher speech behavior analysis method, system and equipment based on text main drive

    CN117009580A

  • Multi-modal sentiment analysis method combining pre-training model and self-attention block

    CN118898046A

Cited By

  • Adaptive text-guided fiber bundle feature fusion method and system based on large language model

    CN121030694A

  • An Adaptive Text-Guided Fiber Bundle Feature Fusion Method and System Based on a Large Language Model

    CN121030694B

  • Deep sea geology mapping method and system based on multi-modal large model

    CN122223151A