Parameter-extensible multi-modal task continuous learning method and device
Through the multimodal task continuous learning method, the problems of feature fusion and parameter optimization in visual language processing are solved, and the stable and accurate prediction of the model in diversified tasks is achieved, which improves the adaptability and robustness of visual language tasks.
Patent Information
- Application Number
- CN202510542452.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing visual language processing methods are difficult to balance the contribution of different modal features during the feature processing stage. Parameter optimization lacks accurate evaluation, knowledge transfer lacks temperature adaptation mechanism, model lacks processing mechanism for uncertain samples, and task adaptation is not flexible enough, resulting in unstable performance of the model in diversified tasks.
By receiving the input data of the system, preprocessing and encoding, the adaptive feature matrix is calculated, attention calculation and parameter optimization are performed, feature projection and dynamic alignment are performed, knowledge extraction and error detection are performed, task feature extraction and decision-making processing are realized, and feature heterogeneity, semantic aligning and knowledge transfer problems in visual language tasks are solved.
The quality of feature fusion and model generalization capabilities are improved, the accuracy and reliability of model prediction are ensured, and the continuous improvement and stability of model performance are achieved, and the needs of different tasks are adapted to.
Smart Images

Figure CN120338045A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning, and in particular, to a method and device for continuous learning of multi-modal tasks with scalable parameters. Background Art
[0002] Visual language tasks are a key research direction in the field of artificial intelligence, which aims to achieve in-depth understanding and collaborative processing of visual and language information by computers. Such tasks not only need to process visual features in images, but also need to understand the associated text descriptions, and establish accurate semantic associations between the two modalities. With the rapid development of fields such as autonomous driving, intelligent healthcare, and intelligent manufacturing, the demand for visual language processing systems is increasing day by day. Especially in industrial scenarios, it is required that the system can accurately understand complex visual scenes and generate professional text descriptions, or accurately identify corresponding visual targets according to text instructions, which is of great significance for improving production efficiency and safety.
[0003] Current research on visual language processing mainly focuses on two aspects: feature extraction and modality fusion. In terms of feature extraction, pre-trained convolutional neural networks are generally used to extract visual features, and word embeddings and position encodings are used to process text features. In terms of modality fusion, mainstream methods include simple feature concatenation, attention mechanisms, and cross-encoders, etc. At the same time, some research has improved the generalization ability of the model by introducing pre-training techniques and contrastive learning. However, these methods often adopt fixed network structures and static parameter configurations, and it is difficult to adapt to the task requirements of different scales and types. In terms of knowledge transfer, existing methods mainly rely on simple knowledge distillation or fine-tuning strategies, lacking in-depth analysis and optimization of the importance of model parameters.
[0004] In practical applications, there are several key technical problems in existing visual language processing methods: First, in the feature processing stage, the dimensions and scales of different modality features vary greatly, and existing feature normalization methods are difficult to effectively balance the contribution degrees of different modalities, resulting in poor feature fusion effects. Second, in the parameter optimization process, there is a lack of an accurate evaluation mechanism for parameter sensitivity, which is easy to fall into local optimal solutions and it is difficult to determine which parameters are the most critical for the current task. Third, in the knowledge transfer process, due to the lack of an effective temperature adaptation mechanism, it is difficult to balance the ratio of soft labels and hard labels, affecting the effect of knowledge transfer. Fourth, in the model optimization stage, existing methods lack a systematic processing mechanism for uncertain samples, and it is difficult to accurately identify and correct high-risk prediction results. Finally, in terms of task adaptation, existing methods often adopt a unified feature extraction strategy without considering the specific requirements of different tasks, resulting in unstable performance of the model when processing diverse tasks. These technical problems severely restrict the application effect and promotion value of visual language processing systems in actual scenarios. Summary of the Invention
[0005] Objective of the invention: To provide a method and device for continuous learning of multi-modal tasks with expandable parameters, aiming to solve at least one technical problem existing in the prior art.
[0006] Technical solution: The method for continuous learning of multi-modal tasks with expandable parameters includes the following steps:
[0007] S1. Receive system input data and perform preprocessing encoding to generate a preprocessing feature matrix of the input data; wherein the system input data includes original visual data and original text data;
[0008] S2. Based on the preprocessing feature matrix of the input data, calculate the feature position information and semantic information respectively to obtain an adaptive feature matrix of the input data;
[0009] S3. Perform attention calculation and parameter optimization according to the adaptive feature matrix of the input data and the initial parameter matrix pre-stored in the system to obtain an optimized parameter matrix of the input data;
[0010] S4. Perform projection processing and dynamic alignment on the optimized parameter matrix to obtain an aligned feature matrix of the input data;
[0011] S5. Perform knowledge extraction and importance evaluation based on the aligned feature matrix to obtain a distilled feature matrix of the input data;
[0012] S6. Perform error detection and compensation processing on the distilled feature matrix to obtain an error-corrected feature matrix of the input data;
[0013] S7. Perform task feature extraction and decision-making processing on the error-corrected feature matrix to obtain a prediction result matrix of the input data.
[0014] The device for continuous learning of multi-modal tasks with expandable parameters includes:
[0015] At least one processor; and,
[0016] A memory communicatively connected to at least one of the processors; wherein,
[0017] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method for continuous learning of multi-modal tasks with expandable parameters.
[0018] Beneficial effects: Through dual-modal feature extraction and encoding, the present invention ensures a high-quality representation of the input data; through positional encoding and semantic encoding, the expressive power of the features is enhanced; through multi-layer parameter optimization and attention calculation, effective interaction between features is achieved; through projection alignment and knowledge distillation, the quality of feature fusion and the generalization ability of the model are improved; through error detection and task adaptation, the accuracy and reliability of model prediction are ensured; not only does it solve key problems such as feature heterogeneity, semantic misalignment, and knowledge transfer in visual language tasks, but it also continuously improves the model performance through the organic cooperation of each module, while maintaining the stability and controllability of the optimization process, providing a reliable technical foundation for the further development of visual language tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the method of the present invention.
[0020] Figure 2 It is a flowchart of step S1 of the present invention.
[0021] Figure 3 It is a flowchart of step S2 of the present invention.
[0022] Figure 4 It is a flowchart of step S3 of the present invention.
[0023] Figure 5 It is a flowchart of step S4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following describes the present application in more detail in combination with specific embodiments. As Figure 1 shown, the present application proposes a parameter-scalable multi-modal task continual learning method, including the following steps:
[0025] S1. Receive system input data, including original visual data and original text data; perform preprocessing encoding on the original visual data and the original text data respectively to generate initial feature matrices; perform normalization processing on the initial feature matrices to obtain preprocessing feature matrices;
[0026] S2. Based on the preprocessing feature matrices, calculate the feature position information and semantic information respectively to generate a position encoding matrix and a semantic encoding matrix; perform a fusion operation on the preprocessing feature matrices, the position encoding matrix, and the semantic encoding matrix to obtain an adaptive feature matrix;
[0027] S3. Perform hierarchical processing on the pre-stored initial parameter matrix to obtain multi-layer parameter matrices; perform attention calculation on the adaptive feature matrix and the multi-layer parameter matrices to obtain an attention feature matrix; perform parameter optimization on the attention feature matrix to obtain an optimized parameter matrix;
[0028] S4. Project the features in the optimized parameter matrix into a unified dimensional space to obtain a projected feature matrix; perform a dynamic alignment operation on the projected feature matrix and the pre-stored alignment parameter matrix to obtain an aligned feature matrix;
[0029] S5. Perform knowledge extraction on the aligned feature matrix and the pre-stored historical parameter matrix to obtain a knowledge feature matrix; evaluate the importance of the parameters in the knowledge feature matrix to obtain a distilled feature matrix;
[0030] S6. Based on the pre-stored verification data matrix, perform error detection on the distilled feature matrix to obtain an error metric matrix; perform compensation processing on the error metric matrix to obtain an error-corrected feature matrix;
[0031] S7. Based on the pre-stored task data matrix, perform task feature extraction on the error-corrected feature matrix to obtain a task feature matrix; perform decision-making processing on the task feature matrix to obtain a prediction result matrix.
[0032] Specifically, the system input data includes original visual data and original text data. The original visual data can be data from images or videos, such as photos, surveillance videos, medical images (such as CT or MRI), or satellite images, etc., and can also be visual content such as paintings or scanned handwritten texts. The original text data can be natural language texts, such as articles, books, conversation logs, or social media content, and can also be encoded or structured text data, such as XML files, JSON data, or text responses to questionnaires.
[0033] As Figure 2 shown, according to one aspect of the present application, step S1 is further as follows:
[0034] S11. Receive the original visual data, use a block processing algorithm to divide the original visual data into primitive blocks of a predetermined size; perform feature extraction on each primitive block to generate an initial visual feature sequence; retrieve visual feature items that match the initial visual feature sequence from the pre-stored visual feature dictionary, and combine the retrieved visual feature items into a visual feature matrix; input the visual feature matrix into a pre-trained encoding model to generate an encoded visual feature matrix;
[0035] S12. Receive the original text data, perform word segmentation processing on the original text data using a predefined word segmentation rule to generate an initial token sequence; retrieve text feature items that match the initial token sequence from the pre-stored text feature dictionary, and combine the retrieved text feature items into a text feature matrix; input the text feature matrix into a pre-trained encoding model to generate an encoded text feature matrix;
[0036] S13. Concatenate the encoded visual feature matrix and the encoded text feature matrix to generate an initial feature matrix; calculate the mean and standard deviation of each feature in the initial feature matrix; and perform normalization processing on the initial feature matrix according to the mean and standard deviation to obtain a preprocessed feature matrix.
[0037] In an embodiment of the present application, the text tokenization and feature encoding method: W(t) = softmax(QK T / sqrt(d))V; where Q = XW_Q is the query matrix; K = XW_K is the key matrix; V = XW_V is the value matrix; X is the input text feature matrix; W_Q, W_K, and W_V are learnable parameter matrices; d is the feature dimension; t is the time step; and W(t) is the text feature encoding result at time step t.
[0038] In this embodiment, through a dual feature encoding and normalization processing mechanism, efficient preprocessing of visual and text data is achieved. The features of the two modalities are concatenated after being mapped by the pre-trained encoding model, and then normalized by the mean and standard deviation, effectively eliminating the scale differences of data in different modalities, enabling the subsequent feature fusion and optimization processes to be carried out in a unified feature space, improving the compatibility and processability of multi-modal features, and laying a solid foundation for subsequent parameter optimization.
[0039] According to one aspect of the present application, step S11 is further as follows:
[0040] S111. Receive the original visual data, perform bilateral filtering operation on the original visual data to obtain filtered image data; perform adaptive contrast enhancement processing on the filtered image data to obtain enhanced image data; and convert the enhanced image data into a tensor format to obtain image tensor data.
[0041] S112. Use the sliding window algorithm to divide the image tensor data into adjacent overlapping image blocks to generate an overlapping image block sequence; perform boundary compensation processing on the overlapping image block sequence to generate a compensated image block sequence; and perform bilinear interpolation resampling on the compensated image block sequence according to a preset size to obtain a unified primitive block sequence.
[0042] S113. Perform multi-scale pyramid transformation on each primitive block in the unified primitive block sequence to generate a pyramid feature sequence; apply a local descriptor extraction algorithm to the pyramid feature sequence to generate a local feature sequence; and perform feature aggregation operation on the local feature sequence to obtain an initial visual feature sequence.
[0043] S114. Perform local sensitive hashing encoding on the initial visual feature sequence to generate a feature hash sequence; perform nearest neighbor search in the pre-stored visual feature dictionary using the feature hash sequence to generate a matching feature sequence; perform similarity weight calculation on the matching feature sequence to obtain a weighted feature sequence;
[0044] S115. Rearrange the weighted feature sequence into a matrix form to generate a feature rearrangement matrix; perform sparsification processing on the feature rearrangement matrix to generate a sparse feature matrix; convert the sparse feature matrix into a dense representation to obtain a visual feature matrix;
[0045] S116. Input the visual feature matrix into the pre-trained encoding model to generate an initial encoding matrix; perform residual connection operation on the initial encoding matrix to generate a residual feature matrix; perform layer normalization processing on the residual feature matrix to obtain an encoded visual feature matrix.
[0046] In an embodiment of the present application, the image bilateral filtering and feature extraction method: F(x, y) = k -1( x, y)∫∫f(ξ, η)·c(ξ, η, x, y)·s(f(ξ, η), f(x, y))dξdη; where k(x, y) = ∫∫c(ξ, η, x, y)·s(f(ξ, η), f(x, y))dξdη is the normalization factor; c(ξ, η, x, y) = exp(-(x - ξ) 2 +(y - η) 2 ) / (2σd 2 )) is the spatial weight function; s(f(ξ, η), f(x, y)) = exp(-|f(ξ, η) - f(x, y)| 2 / (2σr 2 )) is the pixel value similarity function; σd is the spatial standard deviation; σr is the range standard deviation; (x, y) is the target pixel coordinate; (ξ, η) is the neighborhood pixel coordinate; f is the input image; F(x, y) is the pixel value of the output image after filtering at the coordinate (x, y).
[0047] This embodiment realizes the efficient processing and feature expression of visual data through a multi-level visual feature extraction and enhancement mechanism. This embodiment not only ensures the integrity and stability of feature extraction, but also retains the original feature information through residual connection, improving the feature expression ability and robustness.
[0048] According to one aspect of the present application, step S12 is further:
[0049] S121. Receive the original text data, perform character-level normalization processing on the original text data to generate normalized text data; filter out noise symbols from the normalized text data to obtain filtered text data; convert the filtered text data into Unicode encoding format to obtain encoded text data.
[0050] S122. Receive the encoded text data, use regular expressions to perform boundary recognition on the encoded text data to generate a text boundary sequence; perform n-gram tokenization processing on the text boundary sequence to generate an n-gram sequence; use a conditional random field model to predict the token boundaries of the n-gram sequence to obtain an initial tokenization sequence.
[0051] S123. Receive the initial tokenization sequence, perform morphological analysis on the initial tokenization sequence to generate a morphological feature sequence; match the morphological feature sequence with a pre-stored root and affix table to generate a morpheme analysis sequence; perform token normalization processing on the morpheme analysis sequence to obtain an initial token sequence.
[0052] S124. Receive the initial token sequence and a pre-stored text feature dictionary, calculate sub-word encodings for the initial token sequence to generate a sub-word feature sequence; use the sub-word feature sequence to perform fuzzy matching search in the text feature dictionary to generate a matching feature sequence; perform context relevance calculation on the matching feature sequence to obtain a context feature sequence.
[0053] S125. Receive the context feature sequence, reorganize the context feature sequence into a matrix form to generate a feature organization matrix; perform sparse representation transformation on the feature organization matrix to generate a sparse representation matrix; convert the sparse representation matrix into a dense representation to obtain a text feature matrix.
[0054] S126. Receive the text feature matrix, input the text feature matrix into a pre-trained encoding model to generate an initial encoding matrix; perform multi-head attention operation on the initial encoding matrix to generate an attention feature matrix; perform layer normalization processing on the attention feature matrix to obtain an encoded text feature matrix.
[0055] This embodiment realizes the efficient representation and semantic understanding of text data through a fine-grained text processing and feature encoding mechanism. This embodiment not only improves the accuracy of tokenization, but also enhances semantic expression through context relevance calculation. At the same time, the use of multi-head attention captures the complex dependencies between tokens.
[0056] According to one aspect of the present application, step S13 is further as follows:
[0057] S131. Receive the encoded visual feature matrix and the encoded text feature matrix, calculate the dimension parameters of the two feature matrices, and generate a dimension mapping vector; align the feature dimensions according to the dimension mapping vector to generate an aligned feature sequence; splice the aligned feature sequence along the feature dimension direction to obtain an initial feature matrix.
[0058] S132. Receive the initial feature matrix, calculate the mean of each feature along the feature dimension to generate a feature mean vector; subtract the corresponding mean from each feature in the initial feature matrix to generate a centralized feature matrix; calculate the variance of each feature in the centralized feature matrix to generate a feature variance vector; perform a square root operation on the feature variance vector to obtain a standard deviation vector.
[0059] S133. Receive the centralized feature matrix and the standard deviation vector, divide each feature in the centralized feature matrix by the corresponding standard deviation to generate a normalized feature matrix; perform a numerical stability check on the normalized feature matrix to generate a stability index vector; adjust the outliers according to the stability index vector to obtain a preprocessed feature matrix.
[0060] In this embodiment, by performing dimension alignment and splicing on the visual feature matrix and the text feature matrix, the generated initial feature matrix can more accurately reflect the comprehensive features of multimodal data, improving the accuracy and consistency of feature representation. This embodiment can more accurately process and analyze multimodal data, improving the overall performance and decision-making ability of the system, which is of great significance for application scenarios that need to process complex multimodal data, such as intelligent recommendation systems and multimodal information retrieval.
[0061] As Figure 3 shown, according to one aspect of the present application, step S2 is further as follows:
[0062] S21. Receive the visual features in the preprocessed feature matrix; calculate the sine position encoding and cosine position encoding according to the position index of the visual features to generate a visual position encoding matrix; retrieve the corresponding semantic information according to the visual features based on the pre-stored visual semantic dictionary to generate a visual semantic encoding matrix; perform weighted superposition of the visual features with the visual position encoding matrix and the visual semantic encoding matrix to obtain an adaptive visual feature matrix;
[0063] S22. Receive the text features in the preprocessed feature matrix; calculate the sine position encoding and cosine position encoding according to the position index of the text features to generate a text position encoding matrix; retrieve the corresponding semantic information according to the text features based on the pre-stored text semantic dictionary to generate a text semantic encoding matrix; perform weighted superposition of the text features with the text position encoding matrix and the text semantic encoding matrix to obtain an adaptive text feature matrix;
[0064] S23. Concatenate the adaptive visual feature matrix and the adaptive text feature matrix to generate a concatenated feature matrix; perform feature normalization on the concatenated feature matrix to obtain an adaptive feature matrix.
[0065] In this embodiment, through the dual information enhancement of position encoding and semantic encoding, the expression ability of features is improved. It not only enhances the expression ability of features, but also ensures the numerical stability of the fused features through feature normalization processing; it can dynamically adjust the importance degree of features according to the requirements of different tasks, and improves the adaptability of the model to different types of inputs.
[0066] According to one aspect of the present application, step S21 is further as follows:
[0067] S211. Receive the visual feature part in the preprocessed feature matrix, construct a feature position index table, and generate a position index vector; calculate the frequency parameter of the position encoding according to the position index vector to generate a frequency parameter matrix; map the position index vector to the sine space to generate a sine encoding matrix; map the position index vector to the cosine space to generate a cosine encoding matrix; concatenate the sine encoding matrix and the cosine encoding matrix in the feature dimension to obtain a visual position encoding matrix.
[0068] S212. Receive the pre-stored visual semantic dictionary and the visual features in the preprocessed feature matrix, calculate the similarity between the features and the dictionary items to generate a similarity score matrix; perform Top-K screening on the similarity score matrix to generate a candidate semantic sequence; calculate the weights of each semantic item in the candidate semantic sequence to generate a semantic weight vector; perform weighted combination on the candidate semantic sequence according to the semantic weight vector to obtain a visual semantic encoding matrix.
[0069] S213. Receive the visual feature part in the preprocessed feature matrix, the visual position encoding matrix, and the visual semantic encoding matrix, construct a feature fusion network to generate a fusion weight matrix; calculate the weight coefficient of the position encoding according to the fusion weight matrix to generate a position weight vector; calculate the weight coefficient of the semantic encoding according to the fusion weight matrix to generate a semantic weight vector; perform weighted superposition of the visual features, position encoding, and semantic encoding according to the corresponding weights to obtain an adaptive visual feature matrix.
[0070] In this embodiment, through the refined position encoding and semantic encoding mechanisms, the expression ability of visual features is enhanced. It not only improves the expression ability of visual features, but also realizes the dynamic balance of position information and semantic information through the adaptive weight mechanism.
[0071] According to one aspect of the present application, step S22 is further as follows:
[0072] S221. Receive the text feature part in the preprocessed feature matrix, generate position indices in sequence order to generate a position sequence vector; calculate the position encoding frequencies of different dimensions based on the position sequence vector to generate an encoding frequency matrix; perform a sine function mapping on the position sequence vector and the encoding frequency matrix to generate a sine encoding sequence; perform a cosine function mapping on the position sequence vector and the encoding frequency matrix to generate a cosine encoding sequence; perform dimensional interleaving and combination on the sine encoding sequence and the cosine encoding sequence to obtain a text position encoding matrix.
[0073] S222. Receive the pre-stored text semantic dictionary and the text features in the preprocessed feature matrix, construct a fast retrieval index using locality-sensitive hashing to generate a retrieval index matrix; perform a nearest neighbor search in the text semantic dictionary according to the retrieval index matrix to generate a candidate semantic matrix; calculate the relevance of each semantic item in the candidate semantic matrix to the query feature to generate a relevance vector; perform importance ranking on the candidate semantic matrix according to the relevance vector to generate a ranked semantic matrix; perform a semantic fusion operation on the ranked semantic matrix to obtain a text semantic encoding matrix.
[0074] S223. Receive the text feature part in the preprocessed feature matrix, the text position encoding matrix, and the text semantic encoding matrix, calculate the attention scores of the three features to generate an attention score matrix; perform softmax normalization on the attention score matrix to generate a normalized weight matrix; calculate the contribution degree of the position encoding according to the normalized weight matrix to generate a position contribution vector; calculate the contribution degree of the semantic encoding according to the normalized weight matrix to generate a semantic contribution vector; perform dynamic weighted combination of the text features, position encoding, and semantic encoding according to their respective contribution degrees to obtain an adaptive text feature matrix.
[0075] In this embodiment, through a hierarchical position and semantic encoding strategy, the depth enhancement of text features is realized. It not only improves the expression ability of text features, but also ensures the effective utilization of different types of information through dynamic weight allocation, while reducing feature redundancy.
[0076] According to one aspect of the present application, step S23 is further as follows:
[0077] S231. Receive the adaptive visual feature matrix and the adaptive text feature matrix, calculate the statistical distributions of the two feature matrices to generate a distribution feature vector; perform a distribution alignment operation according to the distribution feature vector to generate an alignment coefficient matrix; apply the alignment coefficient matrix to the two feature matrices to generate a calibrated visual matrix and a calibrated text matrix; splice the calibrated visual matrix and the calibrated text matrix in the feature dimension to obtain a spliced feature matrix.
[0078] S232. Receive the spliced feature matrix, calculate local response normalization along the feature dimension to generate a local response matrix; perform batch normalization operation on the local response matrix to generate a batch normalization matrix; calculate the exponential moving average of the batch normalization matrix to generate a smoothed feature matrix; perform layer normalization on the smoothed feature matrix to obtain a normalized feature matrix.
[0079] S233. Receive the normalized feature matrix, construct a feature correlation graph to generate a correlation matrix; perform graph attention operation based on the correlation matrix to generate an attention feature matrix; apply residual connection to the attention feature matrix to generate a residual feature matrix; perform final feature calibration on the residual feature matrix to obtain an adaptive feature matrix.
[0080] In this embodiment, through multi-level feature splicing and normalization processing, the deep fusion of visual and text features is achieved. It not only improves the accuracy of feature fusion but also enhances the generalization ability of the model through the multiple normalization mechanism.
[0081] As Figure 4 shown, according to one aspect of the present application, step S3 is further as follows:
[0082] S31. Receive the pre-stored initial parameter matrix, divide the initial parameter matrix into N sub-matrices according to the preset number of levels N to generate a sequence of parameter sub-matrices; perform independent initialization operations on each sub-matrix in the sequence of parameter sub-matrices to generate a sequence of initialized parameters; reorganize the parameters in the sequence of initialized parameters in the hierarchical order to obtain a multi-level parameter matrix; where N is a natural number greater than 1;
[0083] S32. Based on the parameters in the multi-level parameter matrix, generate a query matrix, a key matrix, and a value matrix to obtain a group of parameter transformation matrices; perform matrix multiplication operations on the adaptive feature matrix with the query matrix, the key matrix, and the value matrix in the group of parameter transformation matrices respectively to generate a group of feature transformation matrices; perform attention calculation on the group of feature transformation matrices to obtain an attention feature matrix;
[0084] S33. Calculate the gradient value of the attention feature matrix with respect to the current task to generate a gradient matrix; input the gradient matrix into a preset optimizer to generate a parameter update matrix; perform parameter update operation on the parameter update matrix and the multi-level parameter matrix to obtain an optimized parameter matrix.
[0085] In an embodiment of the present application, the parameter matrix hierarchical algorithm: L(θ) = ∑(i = 1 to N) α_i||θ_i - θ_(i - 1)|| 2 + β_i||▽θ_i|| 2; where θ_i is the parameter matrix of the i-th layer; α_i is the inter-layer continuity constraint coefficient; β_i is the gradient regularization coefficient; N is the total number of layers; ▽θ_i is the parameter gradient; L(θ) is the multi-layer parameter optimization objective function.
[0086] Attention calculation method: A(X) = softmax((XW_Q)(XW_K) T / sqrt(d) + M)XW_V; where X is the input feature matrix; W_Q, W_K, W_V are learnable parameter matrices; M is the position encoding matrix; d is the feature dimension; A(X) is the attention-weighted feature matrix, T is the transpose.
[0087] In this embodiment, through the hierarchical optimization of the multi-layer parameter matrix and the introduction of the attention mechanism, efficient learning and optimization of the parameters are achieved. It not only improves the expression ability of the model but also enhances the model's perception ability of important features through the introduction of the attention mechanism. At the same time, the use of the optimizer ensures the stability and convergence of parameter updates.
[0088] According to one aspect of the present application, step S31 is further as follows:
[0089] S311. Receive the pre-stored initial parameter matrix, calculate the feature distribution statistics of the initial parameter matrix, and generate a parameter statistics matrix; perform spectral analysis on the parameter statistics matrix to generate a spectral decomposition matrix; obtain a hierarchical division vector according to the singular value distribution of the spectral decomposition matrix.
[0090] S312. Receive the initial parameter matrix and the hierarchical division vector, perform block diagonalization decomposition on the initial parameter matrix according to the hierarchical division vector to generate a sequence of block diagonal matrices; perform parameter normalization processing on the sequence of block diagonal matrices to generate a sequence of normalized blocks; reorganize the sequence of normalized blocks according to the hierarchical information to obtain a sequence of parameter sub-matrices.
[0091] S313. Receive the sequence of parameter sub-matrices, calculate the nuclear norm for each sub-matrix in the sequence of parameter sub-matrices to generate a sequence of norms; determine the scaling factor for each sub-matrix according to the sequence of norms to generate a sequence of scaling factors; apply the sequence of scaling factors to the sequence of parameter sub-matrices to obtain a sequence of scaled parameters.
[0092] S314. Receive the sequence of scaled parameters, perform orthogonal initialization on each matrix in the sequence of scaled parameters to generate a sequence of orthogonal parameters; apply adaptive learning rate adjustment to the sequence of orthogonal parameters to generate a sequence of adjusted parameters; perform gradient normalization processing on the sequence of adjusted parameters to obtain a sequence of initialized parameters.
[0093] S315. Receive the initialization parameter sequence, construct a parameter dependency graph structure, and generate a parameter dependency matrix; perform a topological sort on the initialization parameter sequence according to the parameter dependency matrix to generate a sorted parameter sequence; perform sparse compression on the sorted parameter sequence to obtain a compressed parameter sequence.
[0094] S316. Receive the compressed parameter sequence, reorganize the compressed parameter sequence into a hierarchical structure to generate a hierarchical parameter matrix; perform cross-layer connection optimization on the hierarchical parameter matrix to generate an optimized parameter matrix; perform structure normalization processing on the optimized parameter matrix to obtain a multi-layer parameter matrix.
[0095] In this embodiment, through the hierarchical decomposition and initialization mechanism of the parameters, the efficient organization and optimization of the model parameters are realized. It not only improves the expression ability of the parameters, but also avoids the problem of gradient disappearance through orthogonal initialization. At the same time, the construction of the dependency graph ensures the orderliness of parameter updates, providing good initial conditions for subsequent parameter optimization.
[0096] According to one aspect of the present application, step S32 is further as follows:
[0097] S321. Receive the multi-layer parameter matrix and the adaptive feature matrix, perform a grouped linear transformation on the multi-layer parameter matrix to generate a linear transformation sequence; apply an adaptive activation function to the linear transformation sequence to generate an activation transformation sequence; construct a parameter projection space according to the activation transformation sequence to obtain a projection parameter sequence.
[0098] S322. Receive the projection parameter sequence, divide the projection parameter sequence into three parts: query, key, and value, to generate an initial QKV sequence; perform a parameter decoupling operation on the initial QKV sequence to generate a decoupled QKV sequence; perform scaling normalization processing on the decoupled QKV sequence to obtain a group of parameter transformation matrices.
[0099] S323. Receive the group of parameter transformation matrices and the adaptive feature matrix, perform a multi-head splitting operation on the adaptive feature matrix to generate a multi-head feature sequence; perform a batch matrix multiplication operation on the multi-head feature sequence and the group of parameter transformation matrices to generate a transformed feature sequence; perform a head recombination operation on the transformed feature sequence to obtain a group of feature transformation matrices.
[0100] S324. Receive the group of feature transformation matrices, calculate a query-key similarity matrix to generate a similarity score matrix; apply temperature scaling and masking processing to the similarity score matrix to generate a masked score matrix; perform a softmax normalization operation on the masked score matrix to obtain an attention weight matrix.
[0101] S325. Receive the value matrices in the attention weight matrix and the feature transformation matrix group, perform weighted aggregation operations to generate an aggregated feature matrix; apply a residual connection to the aggregated feature matrix to generate a residual feature matrix; perform layer normalization processing on the residual feature matrix to obtain an intermediate feature matrix.
[0102] S326. Receive the intermediate feature matrix, input the intermediate feature matrix into a feed-forward neural network to generate a feed-forward feature matrix; perform dropout regularization on the feed-forward feature matrix to generate a regularized feature matrix; apply a residual connection and layer normalization to the regularized feature matrix to obtain an attention feature matrix.
[0103] In this embodiment, through the refined design and optimization of the attention mechanism, effective modeling of complex relationships between features is achieved. It can not only capture long-term and short-term dependencies between features, but also realize multi-angle feature correlation analysis through the multi-head mechanism. At the same time, the use of the masking and temperature mechanisms improves the accuracy and controllability of attention calculation.
[0104] As Figure 5 shown, according to one aspect of the present application, step S4 is further as follows:
[0105] S41. Divide the optimization parameter matrix into a visual part and a text part to obtain a visual parameter matrix and a text parameter matrix; perform matrix multiplication on the visual parameter matrix and a pre-stored visual projection matrix to generate a visual projection feature; perform matrix multiplication on the text parameter matrix and a pre-stored text projection matrix to generate a text projection feature; splice the visual projection feature and the text projection feature in the feature dimension to obtain a projection feature matrix;
[0106] S42. Calculate the similarity scores between the features in the projection feature matrix to generate a similarity matrix; perform matrix multiplication on the similarity matrix and a pre-stored alignment parameter matrix to generate an alignment weight matrix; perform weighted summation on the alignment weight matrix and the projection feature matrix to obtain a primary alignment feature matrix;
[0107] S43. Calculate the statistical distribution information of the features in the primary alignment feature matrix to generate a distribution statistics matrix; perform normalization processing on the primary alignment feature matrix according to the distribution statistics matrix to generate a normalized feature matrix; perform a residual connection operation on the normalized feature matrix to obtain the final alignment feature matrix.
[0108] In an embodiment of the present application, the feature projection alignment algorithm: D(X, Y) = ||XP_x - YP_y||_F +λ·tr(P_x T P_xP_y TP_y); where X is the visual feature matrix; Y is the text feature matrix; P_x and P_y are the projection matrices; λ is the regularization coefficient; tr() is the trace of the matrix; ||·||_F is the Frobenius norm; D(X, Y) is the feature alignment loss.
[0109] In this embodiment, through the feature projection and dynamic alignment mechanism, the problems of dimension inconsistency and semantic misalignment in multimodal feature fusion are solved. It can not only handle the scale differences of different modal features, but also adaptively adjust the alignment weights, improving the accuracy and robustness of feature fusion.
[0110] According to one aspect of the present application, step S41 is further as follows:
[0111] S411. Receive the optimization parameter matrix, calculate the feature distribution statistics of the optimization parameter matrix, and generate a modal statistical matrix; perform spectral clustering analysis based on the modal statistical matrix to generate a modal classification vector; segment the optimization parameter matrix according to the modal classification vector to obtain a visual parameter matrix and a text parameter matrix.
[0112] S412. Receive the visual parameter matrix and the pre-stored visual projection matrix, perform singular value decomposition on the visual projection matrix to generate a sequence of visual basis vectors; construct an orthogonal projection operator according to the sequence of visual basis vectors to generate a visual projection operator; perform a tensor product operation on the visual parameter matrix and the visual projection operator to obtain visual transformed features.
[0113] S413. Receive the text parameter matrix and the pre-stored text projection matrix, perform singular value decomposition on the text projection matrix to generate a sequence of text basis vectors; construct an orthogonal projection operator according to the sequence of text basis vectors to generate a text projection operator; perform a tensor product operation on the text parameter matrix and the text projection operator to obtain text transformed features.
[0114] S414. Receive the visual transformed features, perform dimension normalization processing on the visual transformed features to generate normalized visual features; apply an adaptive scaling factor to the normalized visual features to generate scaled visual features; perform feature selection on the scaled visual features to obtain visual projection features.
[0115] S415. Receive the text transformed features, perform dimension normalization processing on the text transformed features to generate normalized text features; apply an adaptive scaling factor to the normalized text features to generate scaled text features; perform feature selection on the scaled text features to obtain text projection features.
[0116] S416. Receive the visual projection features and text projection features, calculate the mutual information score between the visual projection features and the text projection features, generate a mutual information matrix; perform feature alignment based on the mutual information matrix to generate an aligned feature matrix; perform dimensionality reorganization on the aligned feature matrix to obtain a projection feature matrix.
[0117] In this embodiment, through the modal separation and projection alignment mechanism of features, the difficult problems of multi-modal feature fusion are solved. It not only solves the problem of inconsistent scales of different modal features, but also ensures the maximum retention of feature information through orthogonality constraints. At the same time, the use of adaptive scaling improves the adaptability and robustness of feature fusion.
[0118] According to one aspect of the present application, step S42 is further as follows:
[0119] S421. Receive the projection feature matrix, construct a feature pair similarity metric space to generate a metric space matrix; calculate the Euclidean distance between features according to the metric space matrix to generate a distance matrix; convert the distance matrix into Gaussian similarity to generate a Gaussian similarity matrix; perform local sensitive hashing encoding on the Gaussian similarity matrix to generate an encoded similarity matrix; screen high-similarity feature pairs according to the encoded similarity matrix to obtain a similarity matrix.
[0120] S422. Receive the similarity matrix and the pre-stored alignment parameter matrix, perform singular value decomposition on the alignment parameter matrix to generate a decomposition parameter sequence; construct an orthogonal transformation matrix according to the decomposition parameter sequence to generate a transformation matrix; perform matrix block multiplication on the similarity matrix and the transformation matrix to generate a block weight matrix; perform sparsification processing on the block weight matrix to generate a sparse weight matrix; apply temperature scaling to the sparse weight matrix to obtain an alignment weight matrix.
[0121] S423. Receive the alignment weight matrix and the projection feature matrix, calculate the self-attention score according to the alignment weight matrix to generate an attention score matrix; perform row normalization on the attention score matrix to generate a normalized weight matrix; perform feature reorganization on the projection feature matrix according to the normalized weight matrix to generate a reorganized feature matrix; perform layer normalization operation on the reorganized feature matrix to generate a normalized feature matrix; pass the normalized feature matrix through a residual connection process to obtain a primary aligned feature matrix.
[0122] In this embodiment, through the refined similarity calculation and alignment weight generation mechanism, high-quality feature alignment is achieved. It not only improves the accuracy of feature alignment, but also reduces the computational complexity through matrix decomposition and sparsification processing.
[0123] According to one aspect of the present application, step S43 is further as follows:
[0124] S431. Receive the aligned feature matrix, calculate the first moment of each feature dimension to generate a mean vector; calculate the second moment of each feature dimension to generate a variance vector; calculate the covariance relationship between feature dimensions to generate a covariance matrix; construct a multi-dimensional Gaussian distribution model using the mean vector, variance vector, and covariance matrix to generate a Gaussian distribution matrix; extract distribution parameters according to the Gaussian distribution matrix to obtain a distribution statistics matrix.
[0125] S432. Receive the aligned feature matrix and the distribution statistics matrix, calculate the whitening transformation parameters according to the distribution statistics matrix to generate a whitening parameter matrix; perform ZCA whitening processing on the aligned feature matrix to generate a whitened feature matrix; calculate the batch statistics of the whitened feature matrix to generate a batch statistics vector; perform batch normalization according to the batch statistics vector to generate a batch feature matrix; apply layer normalization to the batch feature matrix to obtain a normalized feature matrix.
[0126] S433. Receive the normalized feature matrix, construct a skip connection channel to generate a skip connection matrix; generate a mapped feature matrix by feature mapping of the original aligned feature matrix; perform element-wise addition of the normalized feature matrix and the mapped feature matrix to generate a fused feature matrix; perform scale adjustment on the fused feature matrix to generate a scaled feature matrix; perform a gating operation on the scaled feature matrix to obtain the final aligned feature matrix.
[0127] In this embodiment, through statistical distribution modeling and multi-level normalization processing, stable expression of features is achieved. It not only improves the robustness of feature representation but also enhances the model's ability to process features with different distributions through an adaptive mechanism.
[0128] According to one aspect of the present application, step S5 is further as follows:
[0129] S51. Based on the pre-stored historical parameter matrix, input the aligned feature matrix into a pre-configured teacher model and student model to obtain a teacher output matrix and a student output matrix respectively; apply a temperature coefficient to scale the teacher output matrix and the student output matrix to generate a scaled feature matrix; calculate the cross-entropy between the outputs of the teacher model and the student model in the scaled feature matrix to obtain a knowledge feature matrix.
[0130] S52. Calculate the influence degree of each parameter in the knowledge feature matrix on the model output to generate a parameter sensitivity matrix; perform normalization processing on the values in the parameter sensitivity matrix to generate a parameter importance matrix; screen and update the parameters in the knowledge feature matrix according to the parameter importance matrix to obtain a primary distilled feature matrix.
[0131] S53. Sort the parameters in the primary distillation feature matrix according to their importance to generate a sorted parameter matrix; perform a threshold filtering operation on the sorted parameter matrix to generate a filtered parameter matrix; merge the parameters of the filtered parameter matrix with the pre-stored historical parameter matrix, and update to obtain the final distillation feature matrix.
[0132] In one embodiment of the present application, the knowledge distillation algorithm: L_KD = T 2 ·KL(softmax(z_t / T), softmax(z_s / T)) + α·L_CE(z_s, y); where z_t is the output of the teacher model; z_s is the output of the student model; T is the temperature coefficient; KL() is the KL divergence; L_CE is the cross-entropy loss; α is the balance coefficient; y is the true label.
[0133] In this embodiment, through the knowledge distillation and parameter importance evaluation mechanism, the efficient transfer and compression of model knowledge are achieved. It not only reduces the complexity of the model but also retains the key performance of the model. By adaptively adjusting the temperature coefficient, the effectiveness of knowledge transfer is ensured, and at the same time, the accuracy of the distillation process is improved through threshold filtering.
[0134] According to one aspect of the present application, step S51 is further as follows:
[0135] S511. Receive the alignment feature matrix and the historical parameter matrix, perform feature sampling on the alignment feature matrix to generate a sampled feature sequence; input the sampled feature sequence into the teacher model to generate a teacher feature sequence; perform batch normalization processing on the teacher feature sequence to obtain a teacher output matrix.
[0136] S512. Receive the alignment feature matrix and the historical parameter matrix, perform feature enhancement on the alignment feature matrix to generate an enhanced feature sequence; input the enhanced feature sequence into the student model to generate a student feature sequence; perform batch normalization processing on the student feature sequence to obtain a student output matrix.
[0137] S513. Receive the teacher output matrix, calculate the temperature adaptation factor to generate a teacher temperature vector; apply the teacher temperature vector to the teacher output matrix to generate a teacher scaling matrix; perform probability calibration on the teacher scaling matrix to obtain a teacher probability matrix.
[0138] S514. Receive the student output matrix, calculate the temperature adaptation factor to generate a student temperature vector; apply the student temperature vector to the student output matrix to generate a student scaling matrix; perform probability calibration on the student scaling matrix to obtain a student probability matrix.
[0139] S515. Receive the teacher probability matrix and the student probability matrix, calculate the local neighborhood similarity, and generate a similarity weight matrix; perform adaptive smoothing processing on the similarity weight matrix to generate a smoothed weight matrix; calculate the weighted cross-entropy according to the smoothed weight matrix to obtain a cross-entropy matrix.
[0140] S516. Receive the cross-entropy matrix, perform gradient clipping operation to generate a clipped gradient matrix; update the knowledge representation according to the clipped gradient matrix to generate an updated feature matrix; perform noise suppression processing on the updated feature matrix to obtain a knowledge feature matrix.
[0141] In this embodiment, through the refined design of knowledge distillation and the temperature adaptive mechanism, the efficient transfer of model knowledge is achieved. It not only improves the accuracy of knowledge transfer, but also balances the softness and hardness of knowledge through the temperature adaptive mechanism. At the same time, the consideration of the local neighborhood enhances the robustness of knowledge transfer.
[0142] According to one aspect of the present application, step S6 is further as follows:
[0143] S61. Based on the pre-stored verification data matrix, extract label information; based on the label information, calculate the prediction error of the distilled feature matrix to generate a prediction error matrix; calculate the temporal change amount of the parameters in the distilled feature matrix to generate a parameter change matrix; perform weighted combination on the prediction error matrix and the parameter change matrix to obtain an error metric matrix.
[0144] S62. According to the values in the error metric matrix, calculate a compensation coefficient to generate a compensation coefficient matrix; perform matrix multiplication operation on the compensation coefficient matrix and the distilled feature matrix to generate a compensated feature matrix; perform normalization processing on the compensated feature matrix to obtain a normalized compensation matrix.
[0145] S63. Perform residual connection operation on the normalized compensation matrix and the distilled feature matrix to generate a residual feature matrix; apply an activation function to the residual feature matrix for non-linear transformation to generate an activated feature matrix; perform verification operation on the activated feature matrix and the pre-stored verification data matrix to obtain an error correction feature matrix.
[0146] In an embodiment of the present application, the error detection and compensation algorithm: E(t) = γ·||y – y*|| 2 +(1-γ)·||θ_t - θ_(t-1)|| 2 ; where y is the true label; y* is the model prediction value; θ_t is the parameter at time t; γ is the balance coefficient; E(t) is the error metric at time t.
[0147] In this embodiment, through the error detection and compensation mechanism, the generalization ability and stability of the model are improved. This embodiment can timely detect and correct the deviations in the model prediction, improve the robustness of the model to noise and abnormal data through the dynamic adjustment of the compensation coefficient, and at the same time, the use of residual connections ensures the retention of the original effective features, realizing the steady improvement of the model performance.
[0148] According to one aspect of the present application, step S61 is further as follows:
[0149] S611. Receive the distilled feature matrix and the verification data matrix, extract the label information from the verification data matrix to generate a label sequence; input the distilled feature matrix into the prediction module to generate a prediction sequence; perform sample-level alignment on the prediction sequence and the label sequence to obtain an aligned prediction sequence.
[0150] S612. Receive the aligned prediction sequence and the label sequence, calculate the predicted probability distribution of each sample to generate a probability distribution matrix; perform information entropy calculation on the probability distribution matrix to generate an entropy value sequence; identify high-uncertainty samples according to the entropy value sequence to obtain an uncertainty matrix.
[0151] S613. Receive the uncertainty matrix, calculate the confidence interval for each prediction result to generate a confidence interval matrix; compare the confidence interval matrix with a preset threshold to generate an error flag sequence; calculate the sample weights according to the error flag sequence to obtain a prediction error matrix.
[0152] S614. Receive the distilled feature matrix, calculate the parameter differences between adjacent time steps to generate a difference sequence; perform exponential moving average on the difference sequence to generate a smoothed difference sequence; perform normalization processing on the smoothed difference sequence to obtain a change rate matrix.
[0153] S615. Receive the change rate matrix, apply an adaptive threshold to the change rate matrix to generate a threshold screening matrix; calculate the cumulative effect of parameter changes to generate a cumulative change matrix; combine the threshold screening matrix and the cumulative change matrix to obtain a parameter change matrix.
[0154] S616. Receive the prediction error matrix and the parameter change matrix, calculate the adaptive weight coefficient to generate a weight coefficient vector; perform weighted fusion according to the weight coefficient vector to generate a fusion index matrix; perform smoothing processing on the fusion index matrix to obtain an error index matrix.
[0155] In this embodiment, through the fine-grained error detection and index construction mechanism, the accurate evaluation and correction of the model prediction results are realized. It can not only accurately identify the uncertainty and errors in the model prediction, but also realize the dynamic adjustment of the error index through the introduction of adaptive weights. At the same time, the consideration of the cumulative effect ensures the stability of the long-term prediction performance, improving the reliability and accuracy of the model.
[0156] According to one aspect of the present application, step S62 is further as follows:
[0157] S621. Receive an error metric matrix, calculate an anomaly score for each error metric to generate an anomaly score vector; construct an exponential decay function based on the anomaly score vector to generate a decay function matrix; map the error metric matrix through the decay function matrix to generate a decayed error matrix; perform threshold segmentation on the decayed error matrix to generate a segmentation metric matrix; construct an adaptive compensation function based on the segmentation metric matrix to obtain a compensation coefficient matrix.
[0158] S622. Receive the compensation coefficient matrix and the distilled feature matrix, perform singular value decomposition on the compensation coefficient matrix to generate a decomposition coefficient sequence; construct a block diagonal compensation matrix using the decomposition coefficient sequence to generate a block compensation matrix; perform block matrix multiplication on the block compensation matrix and the distilled feature matrix to generate a product feature matrix; apply an attention mechanism to the product feature matrix to generate an attention feature matrix; fuse the attention feature matrix with the original features to obtain a compensated feature matrix.
[0159] S623. Receive the compensated feature matrix, calculate the local response normalization of the features to generate a local response matrix; perform batch normalization on the local response matrix to generate a batch feature matrix; calculate the layer statistics of the batch feature matrix to generate a layer statistic vector; perform layer normalization based on the layer statistic vector to generate a layer feature matrix; apply instance normalization to the layer feature matrix to obtain a normalized compensation matrix.
[0160] In this embodiment, through the adaptive compensation mechanism and multi-level normalization processing, accurate correction of features is achieved. It not only improves the accuracy of error correction, but also realizes the adaptive adjustment of the compensation intensity through the attention mechanism, while maintaining the stability of the feature distribution.
[0161] According to one aspect of the present application, step S7 is further as follows:
[0162] S71. Group the error correction feature matrix according to the task type to generate a grouped feature matrix; apply the corresponding task embedding vector to each group of features in the grouped feature matrix to generate an embedded feature matrix; perform feature fusion operation on the embedded feature matrix and the pre-stored task data matrix to obtain a task feature matrix;
[0163] S72. Input the task feature matrix into the pre-stored decision parameter matrix for linear transformation to generate a linear transformation matrix; apply a normalization function to the linear transformation matrix for numerical normalization to generate a normalized decision matrix; perform loss calculation on the normalized decision matrix and the pre-stored task label matrix to obtain a primary prediction result matrix;
[0164] S73. Calculate the confidence of each predicted value in the primary prediction result matrix to generate a confidence matrix; apply threshold filtering to the confidence matrix to generate a filtered result matrix; combine and optimize the filtered result matrix with the primary prediction result matrix to update and obtain the final prediction result matrix.
[0165] In an embodiment of the present application, the task feature extraction algorithm: F(X, T) = σ(W_2·ReLU(W_1X + E_T)); where X is the input feature; T is the task encoding; E_T is the task embedding vector; W_1 and W_2 are learnable parameter matrices; σ is the activation function; ReLU is the rectified linear unit; F(X, T) is the task-related feature representation.
[0166] This embodiment realizes the precise adaptation of the model to different tasks through the task feature extraction and decision-making processing mechanism. It not only improves the adaptability of the model to different tasks, but also enhances the reliability of the prediction results through confidence threshold filtering. At the same time, the use of combined optimization ensures the global optimality of the prediction results.
[0167] According to one aspect of the present application, step S71 is further as follows:
[0168] S711. Receive the error correction feature matrix and the pre-stored task data matrix, extract the task description information in the task data matrix to generate a task description sequence; perform semantic parsing on the task description sequence to generate a task semantic vector; perform clustering analysis on the error correction feature matrix according to the task semantic vector to obtain a grouped feature matrix.
[0169] S712. Receive the grouped feature matrix, calculate the similarity distribution of the intra-group features to generate a similarity distribution matrix; perform adaptive clustering based on the similarity distribution matrix to generate a clustering center sequence; calculate the representative features for each clustering center to obtain a representative feature sequence.
[0170] S713. Receive the representative feature sequence, retrieve the relevant task embedding vectors from the pre-stored task embedding library to generate a candidate embedding sequence; perform attention weighting on the candidate embedding sequence to generate a weighted embedding sequence; perform feature mapping on the weighted embedding sequence and the representative feature sequence to obtain a mapped feature matrix.
[0171] S714. Receive the mapped feature matrix, construct a feature dependency graph to generate a dependency relationship matrix; perform graph convolutional operations based on the dependency relationship matrix to generate a graph feature sequence; apply adaptive pooling to the graph feature sequence to obtain a pooled feature matrix.
[0172] S715. Receive the pooled feature matrix and the task data matrix, calculate the feature mutual information, and generate a mutual information matrix; select the optimal feature combination according to the mutual information matrix to generate a combined feature sequence; perform feature enhancement on the combined feature sequence to obtain an enhanced feature matrix.
[0173] S716. Receive the enhanced feature matrix, perform multi-scale feature fusion to generate a multi-scale feature matrix; apply the self-attention mechanism to the multi-scale feature matrix to generate an attention feature matrix; perform residual connection and layer normalization on the attention feature matrix to obtain a task feature matrix.
[0174] In this embodiment, through the task-aware feature extraction and fusion mechanism, the model realizes the precise adaptation to different tasks. It not only improves the relevance between features and tasks, but also captures the complex dependence relationships between features through graph structure modeling. At the same time, the use of multi-scale fusion ensures the integrity and richness of feature expression.
[0175] In another embodiment of the present application, S32 can also be:
[0176] S32. Receive the multi-layer parameter matrix and the adaptive feature matrix, perform sequence length analysis on the adaptive feature matrix to obtain a sequence length vector, and generate an attention mask matrix according to the sequence length vector; perform a masking operation on the attention mask matrix and the attention calculation result to generate a masked attention matrix, and then perform normalization processing on the masked attention matrix to obtain an attention feature matrix.
[0177] S321. Receive the adaptive feature matrix, calculate the valid length marker of each element in the sequence to generate a length marker vector; perform cumulative statistics on the length marker vector to generate a cumulative length vector; calculate the actual length of each sequence position according to the cumulative length vector to obtain a sequence length vector.
[0178] S322. Receive the sequence length vector, construct a two-dimensional index matrix to generate an index mapping matrix; calculate the valid position marker according to the index mapping matrix and the sequence length vector to generate a position marker matrix; convert the position marker matrix to a boolean type to obtain an attention mask matrix.
[0179] S323. Receive the attention mask matrix and the intermediate result of the attention calculation, expand the attention mask matrix to the same dimension as the attention score to generate an expanded mask matrix; perform a bitwise multiplication operation on the expanded mask matrix and the original attention score to generate a masked attention matrix; apply softmax normalization to the masked attention matrix to obtain an attention feature matrix.
[0180] In this embodiment, by calculating the effective length markers and cumulative length vectors of each element in the sequence, the accuracy and consistency of the feature representation are ensured, and the generated sequence length vector can better reflect the actual length of the data. This embodiment can process and analyze multimodal data more accurately, improving the overall performance and decision-making ability of the system.
[0181] In another embodiment of the present application, step S51 can also be:
[0182] S51. Receive the alignment feature matrix and the historical parameter matrix, calculate the output distribution entropy values of the teacher model and the student model, and generate a distribution entropy value matrix; adaptively adjust the temperature coefficient according to the distribution entropy value matrix to generate a temperature coefficient matrix; apply the temperature coefficient matrix to the model output to obtain a scaled feature matrix.
[0183] S511. Receive the model output distribution, calculate the probability distribution of each category to generate a probability distribution matrix; calculate the Shannon entropy for the probability distribution matrix to generate an entropy value vector; perform a sliding window average on the entropy value vector to obtain a distribution entropy value matrix.
[0184] S512. Receive the distribution entropy value matrix, calculate the entropy ratio of the teacher model and the student model to generate an entropy ratio vector; design a piecewise function mapping according to the entropy ratio vector to generate a mapping coefficient vector; convert the mapping coefficient vector into a temperature parameter to obtain a temperature coefficient matrix.
[0185] S513. Receive the temperature coefficient matrix and the original model output, perform a temperature scaling transformation on the original output to generate a scaled output matrix; perform a normalization process on the scaled output matrix to generate a normalized feature matrix; combine the normalized feature matrix with the residual information to obtain a scaled feature matrix; calculate the cross-entropy between the outputs of the teacher model and the student model in the scaled feature matrix to obtain a knowledge feature matrix.
[0186] In this embodiment, by calculating the probability distribution and Shannon entropy of each category, the generated distribution entropy value matrix can better reflect the distribution of the model output, improving the accuracy and stability of the model output. This embodiment can process and analyze the model output more accurately, improving the overall performance and decision-making ability of the system.
[0187] Case 1. A method for continuous learning of multimodal tasks with scalable parameters, including the following steps:
[0188] Step 1. Text modality adaptation.
[0189] In the first stage, train the text modality separately so that the model can effectively learn the representation of text data.
[0190] 1), Token decomposition: The input text data passes through the token decomposition layer, which converts the input original text into a series of "tokens". The token decomposition layer relies on an external token table, splits the original text into the smallest grammatical units, and maps these smallest grammatical units to token table IDs (unique identifiers). The basic steps for constructing the token table are as follows: At initialization, each token is treated as a sequence of characters; calculate the occurrence frequency of each pair of adjacent characters; merge the character pair with the highest frequency to generate a new token; continue iterating until a token table of the predefined length is obtained. When the token decomposition layer works, it decomposes the input text into the closest tokens. Some words in the text will be directly mapped to the tokens in the token table, while for words not in the token table, they are further split until each part can be matched to a token in the token table.
[0191] 2), Token position embedding: In the model architecture based on the self-attention mechanism, since the self-attention mechanism does not contain position information, the position embedding layer is a necessary component. Position embedding ensures that the model can perceive the order of each token in the input sequence by giving each position a unique representation. The token position embedding layer represents each position of the token with a fixed vector, using sine and cosine functions to generate the embedding for each position: PE i,2k =sin(i / 10000 2k / d_text ); PE i,2k+1 =cos(i / 10000 2k / d_text ); where i is the position index of the token, k is the dimension index of the embedding, starting from 0; d_text is the dimension of the token embedding. This design of position embedding ensures that the representation of each position is unique, and for long sequences, the periodic characteristics of sine and cosine help to introduce different position information into the model.
[0192] 3), Token semantic embedding: The task of the token semantic embedding layer is to map each token to a high-dimensional semantic space so that tokens that are semantically similar are close in the vector space. The semantic embedding layer uses an embedding matrix W_text, where each row corresponds to the embedding vector of a token in the token table. For each token T_i in the text input, its semantic embedding representation is: E(T_i) = W_text [T_i], T_i ∈ V; where E(T_i) represents the embedding vector of the i-th token T_i, and V is the token table. The embedding matrix W_text is optimized through an optimization objective function during training, with a shape of V × d_text, that is, each token has a d_text-dimensional embedding vector.
[0193] 4), Final token embedding: Before feeding the data into the attention module, the positional embedding and the semantic embedding are added together so that the model can consider both the semantic information and the sequential information of the tokens. Given the i-th token in an input sequence, the embedding vector X_i finally output by the text data embedding module text is: X_i text = E(T_i) + PE_i.
[0194] 5), Calculating attention: The token embeddings of the text are fed into attention modules with the same structure that are repeatedly stacked for processing. Each attention module consists of multiple layers of parametric attention layers and data attention layers. The calculation of each layer uses the self-attention mechanism to capture the dependencies between the data.
[0195] Among them, for each attention module:
[0196] ① Layer normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer of calculation, the representation of each input is normalized so that its mean is 0 and its variance is 1. For the input vector X_text, it is updated to: X_text = LN(X_text) = (X - μ) / (σ + ψ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and ψ is a small constant to prevent the denominator from being 0.
[0197] ② Q, K, V matrix calculation: Dynamically adjust the representation of the model by taking the parameters as part of the self-attention formula calculation. The parametric attention layer manages the interaction between the input data and a set of optimizable text data adaptation parameters through the cross-attention mechanism, uses a set of optimizable data as model parameters, and processes the interaction between the input data and these text adaptation parameter data through cross-attention. Considering the input data X_text and a set of optimizable text adaptation parameter data K_P and V_P, the basic calculation formula of the parametric attention layer is: F(X_text, K_P, V_P) = Λ(X_text · K_P T)·V_P; where the dimension of the input data X_text is T×d_text, T represents the sequence length of the input data, i.e., the number of tokens, and each token is a d_text-dimensional vector. The dimensions of the parameter data K_P and V_P are n×d_text, n represents the number of parameter data, and each parameter is a d_text-dimensional vector. In the standard Transformer model architecture, the attention scores are calculated by the softmax function. However, the exponential form of the softmax function may cause some values in the attention score matrix to become extremely high (or low), which may lead to vanishing gradients when calculating the gradients and affect the stability of training. Therefore, the calculation formula of the Λ function is defined as: Λ(X_text·K_P T )=f((X_text·K_P T )) i,j ·τ·1 / sqrt(∑ k=1 n |(X_text·K_P T )) i,k | 2 ), for all i∈1...T, j∈1...n; where T represents the sequence length of the input data and n represents the number of parameter data. The subscripts i, j, and k represent the matrix element indices. τ is a scaling factor, and the default value is sqrt(n). f is a non-linear activation function: f(x)=0.5x·(1+(2 / sqrt(π)) ∫0 x / sqrt(2) e -t2 dt). The design of the f activation function improves the gradient flow, especially in the small value and negative value regions, and avoids the problem of vanishing gradients that may occur in Softmax. And combined with the L2 regularization operation, f can generate a relatively uniform attention distribution and reduce the overflow problem in numerical calculations.
[0198] The parameter attention layer is further divided into Q-parameter attention, K-parameter attention, V-parameter attention, O-parameter attention, and FFN-parameter attention. Among them, the Q-parameter attention is used to calculate the query matrix Q: Q=F(X_text, K_P Q , V_P Q ); where X_text is the input data, K_P Q and V_P Q are the corresponding parameter matrices with dimensions of n×d_text. Similarly, calculate the key matrix K and the value matrix V: K=F(X_text, K_P K , V_P K ); V=F(X_text, K_P V , V_P V ).
[0199] ③ Calculate self-attention scores: The data attention layer calculates the result based on the Q, K, and V matrices of the parameter attention layer, and applies the standard self-attention calculation formula to calculate the X_attn matrix: X_attn = softmax[(Q·K T ) / sqrt(d_text)]·V.
[0200] ④ Calculate the self-attention output matrix: Input the X_attn matrix into the O parameter attention calculation module to calculate the self-attention output matrix O_attn: O_attn = F(X_attn, K_P O , V_P O ).
[0201] ⑤ Perform residual processing: Add the self-attention output matrix O_attn to the layer-normalized matrix: X_inter = LN(X) + O_attn; Perform layer normalization again: X_ffn = LN(X_inter).
[0202] ⑥ Calculate the attention module output matrix: Input X_ffn into the FFN parameter attention calculation module to calculate the attention module output matrix O_ffn: O_ffn = F(X_ffn, K_P O , V_P O ).
[0203] 6) Task decision: The main function of the decision layer is to convert the joint representation of the multi-modal input (the fused representation from vision and text) into a prediction for a specific task. In this embodiment, the decision layer consists of a series of fully connected layers, aiming to further transform and abstract the fused features to generate a representation suitable for a specific task. Convert the representation of the text input into a prediction for a specific task. Input the feature representation O_ffn after the FFN parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers: Z = FFN(O_ffn); where FFN represents a feed-forward neural network, which consists of multiple linear transformations and activation functions; Z is the feature representation obtained through the decision layer, providing a high-level representation related to the task.
[0204] 7) Calculate task-related losses: Depending on the task, the training objectives and loss functions of the text data will be different. During training, the AdamW optimizer is used to minimize the loss function. Specifically: Calculate the loss through the output of the model and the target labels; Calculate the gradient through the backpropagation algorithm; Use the optimization algorithm AdamW to update the parameters of the model.
[0205] Step 2: Visual modality adaptation.
[0206] In the second stage, train the text modality separately so that the model can effectively learn the representation of visual data.
[0207] 1). Primitive Decomposition: The input image is sliced into a series of non - overlapping blocks. Each block serves as the smallest visual unit and is flattened to form a primitive. The specific steps are as follows:
[0208] ①Slice the input image \(I\in\mathbb{R}\) H×W×C (with image size \(H\times W\) and \(C\) channels) into \(P\times P\) small blocks.
[0209] ②Flatten each \(P\times P\) small block to obtain a vector of shape \(P\) 2 \(\times C\). Each vector is a primitive.
[0210] 2). Primitive Position Embedding: Similar to token position embedding, the position embedding of each primitive is calculated through sine and cosine functions: \(PE\) i,2k \(=\sin(i / 10000\) 2k / d_ image \()\); \(PE\) i,2k+1 \(=\cos(i / 10000\) 2k / d_ image ) where \(i\) is the position index of the primitive, \(k\) is the dimension index of the embedding, starting from 0. \(d_{image}\) is the dimension of the primitive embedding space. The position embedding provides position information for each primitive, ensuring that the model can perceive the position of the primitive in the image. Moreover, the position embedding generated by sine and cosine functions has good performance in modeling long image sequences, avoiding the limitations that traditional position encodings may encounter.
[0211] 3). Primitive Semantic Embedding: The semantic embedding is mapped to a fixed dimension \(d_{image}\) through a trainable projection matrix \(W_{patch}\): \(E(V_i)=V_i\times W_{patch}\); where \(V_i\) represents the primitive with position index \(i\).
[0212] 4). Final Primitive Embedding: Before feeding the data into the attention module, the position embedding and the semantic embedding are added together so that the model can consider both the semantic information and the sequential information of the primitive. Given the \(i\) - th primitive in an input image, the embedding vector \(X_i\) image output by the image data embedding module is: \(X_i\) image \(=E(V_i)+PE_i\).
[0213] 5). Calculate Attention: The primitive embeddings of the picture are fed into repeatedly stacked attention modules with the same structure. Each attention module consists of multiple layers of parametric attention layers and data attention layers. The calculation of each layer captures the dependencies between data through the self - attention mechanism.
[0214] Among them, for each attention module:
[0215] ① Layer normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer's calculation, the representation of each input is normalized so that its mean is 0 and variance is 1. For the input vector X_image, it is updated as: X_image = LN(X_image) = (X - μ) / (σ + ψ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and ψ is a small constant to prevent the denominator from being 0.
[0216] ② Calculation of Q, K, V matrices: Add a set of visual adaptation parameters K_P new and V_P new , whose number of parameters is the same as that of text adaptation parameters, initialized as a zero matrix, and the new parameters are updated as: K_P = [K_P, K_P new , V_P = [V_P, V_P new ; Considering the input data X_image and a set of optimizable parameter data K_P and V_P, the basic calculation formula of the parameter attention layer is: F(X_image, K_P, V_P) = Λ(X_image · K_P T ) · V_P; where the dimension of the input data X_image is T × d_image, T represents the length of the primitive sequence, that is, the number of primitives, and each metadata is a vector of d_image dimensions. The dimensions of the parameter data K_P and V_P are 2n × d_image, 2n represents the number of parameter data, and each parameter is a vector of d_image dimensions. The Λ function and f function are the same as above. Calculate the query matrix Q: Q = F(X_image, K_P Q , V_P Q ); where X_image is the input data, K_P Q and V_P Q are the corresponding parameter matrices with dimensions of 2n × d_image. Similarly, calculate the key matrix K and the value matrix V: K = F(X_image, K_P K , V_P K ); V = F(X_image, K_P V , V_P V ).
[0217] ③ Calculation of self-attention scores: The data attention layer calculates the X_attn matrix using the standard self-attention calculation formula based on the calculation results of the Q, K, V matrices of the parameter attention layer.
[0218] ④ Calculation of self-attention output matrix: Input X_attn into the O parameter attention calculation module to calculate the self-attention output matrix O_attn.
[0219] ⑤Perform residual processing: Add the self-attention output matrix \(O_{attn}\) to the layer-normalized matrix, and then perform layer normalization again.
[0220] ⑥Calculate the output matrix of the attention module: Pass \(X_{ffn}\) into the FFN parameter attention calculation module to calculate the output matrix \(O_{ffn}\) of the attention module.
[0221] 6) Task decision-making: Convert the representation of the visual input into a prediction for a specific task. Input the feature representation \(O_{ffn}\) after the FFN parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers.
[0222] 7) Calculate task-related losses: Depending on the task, the training objectives and loss functions for visual data will vary. During training, the AdamW optimizer is used to minimize the loss function. Specifically: Calculate the loss through the output of the model and the target labels; Calculate the gradients through the backpropagation algorithm; Update the parameters of the model using the optimization algorithm AdamW.
[0223] Step 3: Visual-language modality alignment.
[0224] 1) Align vector dimensions.
[0225] Apply visual-text fine-tuning data. For each "visual-text" data pair, calculate the visual embedding and the text embedding respectively. The linear projection layer aligns the dimensions of the visual features and the text features into a unified dimensional space through a trainable linear transformation: \(V_{aligned}=X_{image}\cdot W_{image}\), \(W_{image}\in R\) d_image×d_text ; where \(W_{image}\) is the trainable linear transformation matrix used to project the image data embedding vector into the same dimension as the text data embedding vector, \(X_{image}\) is the image data embedding vector, and \(V_{aligned}\) is the aligned image data embedding vector.
[0226] 2) Vector concatenation: When the dimensions of the visual and text features are aligned, the next step is to concatenate their embedding vectors into a long vector to further fuse the information of the two modalities. The vector concatenation layer concatenates the already aligned visual and text features into a unified input vector so that they can participate in the calculation together in the subsequent attention mechanism: \(X = [X_{text}, X_{aligned}]\). Here, \(X_{aligned}\) is the visual embedding vector processed by the linear projection layer.
[0227] 3). Attention calculation: Send X to the attention modules with the same structure that are repeatedly stacked for processing. Each attention module consists of multiple layers of parametric attention layers and data attention layers. The calculation of each layer captures the dependencies between data through the self-attention mechanism.
[0228] Among them, for each attention module:
[0229] ① Layer normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer of calculation, the representation of each input is normalized so that its mean is 0 and variance is 1. For the input vector X, it is updated as: X = LN(X) = (X - μ) / (σ + ψ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and ψ is a small constant to prevent the denominator from being 0.
[0230] ② Q, K, V matrix calculation: Add a set of alignment parameters K_P align and V_P align , with the number of parameters being m, m < n, initialized as a zero matrix, and the new parameters are updated as: K_P = [K_P, K_P align , V_P = [V_P, V_P align ; Considering the input data X and a set of optimizable parameter data K_P and V_P, the basic calculation formula of the parametric attention layer is: F(X, K_P, V_P) = Λ(X · K_P T ) · V_P; where the dimension of the input data X is T × d_image, T represents the length of the concatenated sequence, that is, the number of input metadata, and each metadata is a d_image-dimensional vector. The dimensions of the parameter data K_P and V_P are (2n + m) × d_image, and (2n + m) represents the number of parameter data, and each parameter is a d_image-dimensional vector. The Λ function and f function are the same as above. Calculate the query matrix Q: Q = F(X, K_P Q , V_P Q ), where X is the input data, K_P Q and V_P Q are the corresponding parameter matrices with dimensions of (2n + m) × d_image. Similarly, calculate the key matrix K and the value matrix V: K = F(X, K_P K , V_P K ); V = F(X, K_P V , V_P V ).
[0231] ③ Self-attention score calculation: The data attention layer calculates the X_attn matrix based on the calculation results of the Q, K, V matrices of the parametric attention layer using the standard self-attention calculation formula.
[0232] ④ Self-attention output matrix calculation: Input X_attn into the O-parameter attention calculation module to calculate the self-attention output matrix O_attn.
[0233] ⑤ Residual processing: Add the self-attention output matrix O_attn to the layer-normalized matrix, and perform layer normalization again.
[0234] ⑥ Attention module output matrix calculation: Input X_ffn into the FFN-parameter attention calculation module to calculate the attention module output matrix O_ffn.
[0235] 4) Task decision-making: Convert the joint representation of the multi-modal input (the fused representation from vision and text) into a prediction for a specific task. Input the feature representation O_ffn after the FFN-parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers.
[0236] 5) Task-related loss calculation: Depending on the task, the training objective and loss function will vary. During training, the AdamW optimizer is used to minimize the loss function. Specifically: Calculate the loss through the output of the model and the target labels; Calculate the gradient through the backpropagation algorithm; Update the parameters of the model using the optimization algorithm AdamW.
[0237] In a further embodiment, a parameter-expandable multi-modal task continual learning device includes a text data embedding module, a visual data embedding module, a vector dimension alignment module, an attention module, and a task decision module. Among them, the text data embedding module includes a token decomposition layer, a token position embedding layer, and a token semantic embedding layer. The visual data embedding module includes a graphic element decomposition layer, a graphic element position embedding layer, and a graphic element semantic embedding layer. The vector dimension alignment module includes a linear projection layer and a vector concatenation layer. The attention module includes a layer normalization layer, a parameter attention layer, and a data attention layer. The task decision module includes a decision layer and an output layer.
[0238] The text data embedding module is mainly used to convert the input natural language text into a high-dimensional vector representation for subsequent processing in the model. This module includes three key sub-modules: the token decomposition layer, the token position embedding layer, and the token semantic embedding layer. The token decomposition layer converts the input original text (sentences, paragraphs, etc.) into a series of "tokens". The goal of token decomposition is to convert the text data into a format suitable for processing by a neural network model. In the inference stage, this layer relies on an external token table to split the original text into the smallest grammatical units and map these smallest grammatical units to token table IDs (unique identifiers). In a model architecture based on the self-attention mechanism, since the self-attention mechanism does not contain position information, the position embedding layer is a necessity. The position embedding ensures that the model can perceive the order of each token in the input sequence by giving each position a unique representation. The task of the token semantic embedding layer is to map each token to a high-dimensional semantic space so that tokens that are semantically similar are closer in the vector space.
[0239] The purpose of the visual data embedding module is to convert the input image into a high-dimensional vector representation suitable for model processing for joint modeling with text data. In this process, the image obtains a semantic embedding vector through patch decomposition and linear projection, and then combines with the position information to form the final visual data representation. The role of the patch decomposition layer is to divide the input image into several blocks, each block serving as the smallest visual unit, and then convert these visual units into vectors of a fixed size (patches) for subsequent processing. Since the self-attention mechanism of ViT does not have built-in sequential information, the task of the patch position embedding layer is to add position information to each image patch so that the model can understand the position of each image patch in the image. The role of the patch semantic embedding layer is to provide a high-dimensional semantic representation for each image patch so that the model can capture the semantic information of each patch in the image through these representations.
[0240] Due to the dimensional differences in the representations of visual data and text data, this module is needed to align their dimensions, so as to ensure that two different modalities of data can effectively interact in the same space. The core goal of this process is to enable the vectors of visual data and text data to share representations in the same dimensional space, thus facilitating subsequent cross-modal learning and the inference of multi-modal tasks. Since visual data and text data have different dimensions in the original embedding space, it is not appropriate to directly put them into the same model for calculation. Therefore, the goal of the linear projection layer is to align the dimensions of visual features and text features to a unified dimensional space through a trainable linear transformation. After the dimensions of visual and text features are aligned, the next step is to concatenate their embedding vectors into a long vector to further fuse the information of the two modalities. The vector concatenation layer concatenates the already aligned visual and text features into a unified input vector, enabling them to jointly participate in the calculation in the subsequent attention mechanism.
[0241] The layer normalization layer in the attention module is a commonly used normalization technique in the Transformer architecture, which is used to accelerate training and improve model stability. By introducing the parameter attention layer, each input data can not only interact with other input data, but also interact with the learnable parameters of the model. Different from the traditional attention mechanism (which only supports the interaction between input data and input data), in this embodiment, the parameters are used as part of the self-attention formula calculation to dynamically adjust the representation of the model. Specifically, the parameter attention layer manages the interaction between input data and a set of optimizable parameter data through the cross-attention mechanism. The data attention layer enables different modality data and same modality data to perform mutual attention in the same space. This layer is the interaction between input data and input data, through which the model can improve its understanding ability of input data in multi-modal tasks.
[0242] The main function of the decision layer of the task decision module is to convert the joint representation of the multi-modal input (the fused representation from vision and text) into a prediction for a specific task. The decision layer of this embodiment consists of a series of fully connected layers, aiming to further transform and abstract the fused features to generate a representation suitable for a specific task. The output layer is the last part of the task decision module and also the last part of the entire framework, responsible for mapping the result processed by the decision layer to a specific output space. Depending on the task, the structure and type of the output layer will also vary. For classification tasks, such as text classification, image classification, etc., the output layer uses a softmax function to convert the output of the decision layer into a probability distribution. For generation tasks, such as image caption generation, etc., the calculation of the output layer gradually generates text through the Transformer decoder. For regression tasks, such as picture scores, object positions, etc., the output layer directly generates a numerical output.
[0243] Case 2. A multi-modal task continual learning device with scalable parameters, comprising a multi-modal data embedding module, a vector dimension alignment module, an attention optimization module, a task decision module, and a dynamic parameter adaptation and continual optimization mechanism. Taking the three modalities of text, image, and audio as an example, the multi-modal data embedding module includes ① a text data embedding layer; ② a visual data embedding layer; ③ an audio data embedding layer. The vector dimension alignment module includes ① a linear projection layer; ② a vector concatenation layer. The attention optimization module includes ① a layer normalization layer; ② a parameter attention layer; ③ a data attention layer. The task decision module includes ① a decision layer; ② an output layer.
[0244] (1) The multi-modal data embedding module.
[0245] The multi-modal data embedding module is mainly used to convert data of different modalities (text, image, audio) into high-dimensional vector representations for subsequent processing in the model. This module includes three key sub-modules: text data embedding, visual data embedding, and audio data embedding.
[0246] ① The text data embedding layer. Token decomposition: The input text is split into the smallest grammatical units, each unit is called a token, and indexing processing is performed; Token position embedding: The word order information of the tokens after the input text is decomposed is recorded through sine / cosine encoding to obtain the position embedding PE_i of the i-th token; Token semantic embedding: Each token is mapped to a high-dimensional semantic space d_text to ensure that synonyms and similar sentences are close in the semantic vector space, obtaining the semantic embedding WE_i of the i-th token; Before the data is input into the subsequent module, the position embedding and the semantic embedding need to be added so that the model can consider both the semantic information and the order information of the tokens. Given the i-th token in an input sequence, the embedding vector X_i finally output by the text data embedding module text is: X_i text = WE_i + PE_i.
[0247] ② Visual data embedding layer. Primitive decomposition: Use a Vision Transformer (ViT) to divide the image into multiple visual patches, each patch is called a primitive, and perform indexing. Primitive position embedding: Add position information to each image patch so that the model can understand the position of each image patch in the image. Similar to token position embedding, in this solution, the position embedding of each primitive is calculated through sine and cosine functions, and the position embedding PE_i of the i-th primitive is obtained. Primitive semantic embedding: Provide a high-dimensional semantic representation for each image patch so that the model can capture the semantic information of each primitive in the image through these representations. In this solution, the semantic embedding is mapped to a fixed dimension d_vision through a trainable projection matrix W_patch: VE_i = V_i × W_patch; where V_i represents the primitive with position index i. Before feeding the data into the subsequent module, the position embedding and the semantic embedding need to be added so that the model can consider both the semantic information and the order information of the primitives. Given the i-th primitive in an input sequence, the embedding vector X_i finally output by the visual data embedding module is vision as: X_i vision = VE_i + PE_i.
[0248] ③ Audio data embedding layer. Time-frequency conversion layer: Convert the speech signal into Mel spectrogram or MFCC features. Audio block decomposition layer: Divide the audio into frames of fixed length for feature extraction and perform indexing. Audio embedding layer: Use the pre-trained HuBERT for feature vectorization. Specifically, a convolutional feature extractor is used to extract low-level features; context modeling is performed through Transformer layers; a fixed-length audio embedding vector is generated, denoted as X audio .
[0249] (2) Vector dimension alignment module.
[0250] In this embodiment, the vector dimension alignment module is a crucial step. Since there are dimensional differences in the representations of different modalities, it is necessary to align their dimensions through this module to ensure that data of several different modalities can effectively interact in the same space. The core goal of this process is to enable vectors of different modalities to share representations in the same dimensional space, thus facilitating subsequent cross-modal learning and inference of multi-modal tasks.
[0251] ① Linear projection layer. Since data of different modalities have different dimensions in the original embedding space, it is not appropriate to directly put them into the same model for calculation. Therefore, the goal of the linear projection layer is to align the dimensions of visual features, text features, and audio features into a unified dimensional space through a trainable linear transformation. This embodiment is completed by three linear transformation matrices, which map the embedding vectors of the visual modality to the same dimension as the embedding vectors of the text modality: For the input data of the text modality: V_aligned text = X text · W_1 + b_1, W_1 ∈ R d_text×d_aligned ; For the input data of the visual modality: V_aligned vision = X vision · W_2 + b_2, W_2 ∈ R d_vision×d_aligned ; For the input data of the audio modality: V_aligned audio = X audio · W_3 + b_3, W_3 ∈ R d_audio×d_aligned ; where W_1, W_2, W_3 are trainable linear transformation matrices for projecting the data embeddings of different modalities into a common dimension. b_1, b_2, b_3 are trainable bias terms.
[0252] ② Vector concatenation layer. When the dimensions of data of different modalities are aligned, the next step is to concatenate their embedding vectors into a long vector to further fuse the information of different modalities. The vector concatenation layer concatenates the already aligned features into a unified input vector so that they can jointly participate in the calculation in the subsequent attention mechanism: X = [V_aligned text , V_aligned vision , V_aligned audio ].
[0253] (3) Attention module.
[0254] (4) Task decision module.
[0255] (5) Dynamic parameter adaptation and continuous optimization mechanism.
[0256] This embodiment supports the dynamic expansion of model parameters. When the model needs to adapt to new tasks or modalities, only the relevant parameters need to be supplemented, without modifying the entire model architecture or training from scratch. The old parameter matrices are denoted as K_P, V_P, and the newly added parameter matrices are denoted as K_P new 、V_P new , then the existing parameter matrices are directly concatenated to get: K' = [K_P, K_P new ]; V' = [V_P, V_P new ; The newly added parameter K_Pnew , V_P new Initialized to 0, substituting K' and V' into the attention calculation formula will neither affect the performance of the model on the original task nor the optimization of the new task. Therefore, this mechanism supports the dynamic parameter adaptation of the model and the continuous optimization of tasks, which is an ability not possessed by all previous multi-modal models.
[0257] Case 3: A method for continuous learning of multi-modal tasks with scalable parameters, specifically including:
[0258] Step 1: Data preparation and preprocessing.
[0259] 1.1 Collect natural language text data: Prioritize publicly available text datasets with high quality and high diversity. Typical sources include general text data, domain-specific text data, dialogue text data, and cross-lingual text data, etc. Uniformly convert the encoding of the text data to UTF-8, remove special characters, and remove infrequent symbols. Use language tools to correct grammar errors in the text data. Store the preprocessed data and perform version control using DVC (Data Version Control).
[0260] 1.2 Collect visual image data: Prioritize general datasets and domain-specific datasets with high quality and complete annotations. Use tools such as histogram analysis to batch detect common problems, including blurring, overexposure / underexposure, and duplication, etc. Conduct expert annotation review for complex problems. Store the preprocessed data in Parquet format and achieve efficient reading and writing in combination with cloud storage AWS S3. Perform version control using DVC (Data Version Control).
[0261] 1.3 Collect multi-modal data: Prioritize publicly available multi-modal datasets with high quality and high diversity. Typical sources include image-text pair data, web graphic-text pairs, visual question-answering data, etc. Use the CLIP pre-trained model to calculate the similarity between images and texts, and eliminate samples with low correlation. Remove duplicate samples, low-resolution images, and invalid texts. Unify the resolution and RGB format of the image data. Unify the UTF-8 encoding of the text data, remove special characters, and remove infrequent symbols. Store the preprocessed data in Parquet format and achieve efficient reading and writing in combination with cloud storage AWS S3. Perform version control using DVC (Data Version Control).
[0262] Step 2: Text modality adaptation.
[0263] Train the text modality separately so that the model can effectively learn the representation of text data.
[0264] 2.1 Token Decomposition: The input text data passes through the token decomposition layer, which converts the input original text into a series of "tokens". The token decomposition layer relies on an external token table to split the original text into the smallest grammatical units and maps these smallest grammatical units to token table IDs (unique identifiers). When the token decomposition layer works, it decomposes the input text into the closest tokens. Some words in the text will be directly mapped to the tokens in the token table, while for words not in the token table, they will be further split until each part can be matched to a token in the token table.
[0265] 2.2 Token Position Embedding: By giving each position a unique representation, it ensures that the model can perceive the order of each token in the input sequence. The token position embedding layer represents each position of the token with a fixed vector, using sine and cosine functions to generate the embedding for each position: PE_(i, 2k) = sin(i / 10000 2k / d_text ) ; PE_(i, 2k + 1) = cos(i / 10000 2k / d_text ) ; where, i is the position index of the token, and k is the embedding dimension index, starting from 0.
[0266] 2.3 Token Semantic Embedding: The semantic embedding layer uses the embedding matrix W_text, where each row corresponds to the embedding vector of a token in the token table. For each token T_i in the text input, its semantic embedding representation is: E(T_i) = W_text[T_i], T_i ∈ V; where, E(T_i) represents the embedding vector of the i-th token T_i, and V is the token table. The embedding matrix W_text is optimized through the optimization objective function during the training process, with a shape of V × d_text, that is, each token has an embedding vector of d_text dimensions.
[0267] 2.4 Final Token Embedding: Before feeding the data into the attention module, the position embedding and the semantic embedding are added together so that the model can consider both the semantic information and the order information of the tokens simultaneously. Given the i-th token in an input sequence, the embedding vector X_i text output by the text data embedding module finally is: X_i text = E(T_i) + PE_i.
[0268] 2.5 Attention Calculation: The token embeddings of the text are fed into repeatedly stacked attention modules with the same structure for processing. Each attention module consists of multiple layers of parametric attention layers and data attention layers. The calculation of each layer captures the dependencies between the data through the self-attention mechanism.
[0269] Among them, for each attention module:
[0270] 2.5.1 Layer Normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer's calculation, the representation of each input is normalized so that its mean is 0 and variance is 1. For the input vector X_text, it is updated as: X_text = LN(X_text) = (X - μ) / (σ + δ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and δ is a small constant to prevent the denominator from being 0.
[0271] 2.5.2 Q, K, V Matrix Calculation: By taking the parameters as part of the self-attention formula calculation, the model's representation is dynamically adjusted. The parametric attention layer manages the interaction between the input data and a set of optimizable text data adaptation parameters through the cross-attention mechanism, uses a set of optimizable data as model parameters, and processes the interaction between the input data and these text adaptation parameter data through cross-attention. Considering the input data X_text and a set of optimizable text adaptation parameter data K_P and V_P, the basic calculation formula of the parametric attention layer is: F(X_text, K_P, V_P) = Ξ(X_text · K_P T ) · V_P; where, T is the transpose, the dimension of the input data X_text is T × d_text, T represents the sequence length of the input data, that is, the number of tokens, and each metadata is a d_text-dimensional vector. The dimensions of the parameter data K_P and V_P are n × d_text, n represents the number of parameter data, and each parameter is a d_text-dimensional vector. Define the calculation formula of the Ξ function: Ξ(X_text · K_P T ) = f((X_text · K_P T ) i,j ·τ·1 / sqrt(∑ k=1 n | (X_text · K_P T ) i,k | 2 )), for all i ∈ 1...T, j ∈ 1...n; where T represents the sequence length of the input data and n represents the number of parameter data. The subscripts i, j, and k represent matrix element indices. τ is a scaling factor, with a default value of sqrt(n). f is a non-linear activation function: f(x) = 0.5x · (1 + 2 / sqrt(π) ∫0 x / sqrt(2) e -t2 dt); Calculate the query matrix Q: Q = F(X_text, K_P Q , V_P Q ); where, X_text is the input data, K_P Q and V_PQ is the corresponding parameter matrix with dimensions n×d_text. Similarly, calculate the key matrix K and the value matrix V: K = F(X_text, K_P K , V_P K ); V = F(X_text, K_P V , V_P V ); where F is the basic calculation function of the parameter attention layer.
[0272] 2.5.3 Self-attention score calculation: The data attention layer calculates the X_attn matrix by applying the standard self-attention calculation formula based on the calculation results of the Q, K, and V matrices of the parameter attention layer: X_attn = softmax[(Q·K T ) / sqrt(d_text)]·V.
[0273] 2.5.4 O_attn matrix calculation: Input X_attn into the O parameter attention calculation module to calculate the self-attention output matrix O_attn: O_attn = F(X_attn, K_P O , V_P O ).
[0274] 2.5.5 Residual processing: Add the self-attention output matrix O_attn to the layer-normalized matrix: X_inter = LN(X) + O_attn; perform layer normalization again: X_ffn = LN(X_inter).
[0275] 2.5.6 Attention module output matrix calculation: Input X_ffn into the FFN parameter attention calculation module to calculate the attention module output matrix O_ffn: O_ffn = F(X_ffn, K_P O , V_P O ).
[0276] 2.6 Task decision-making: Convert the representation of the text input into a prediction for a specific task. Input the feature representation O_ffn after the FFN parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers: Z = FFN(O_ffn), where FFN represents a feed-forward neural network composed of multiple linear transformations and activation functions. Z is the feature representation obtained through the decision layer, providing a high-level representation related to the task.
[0277] 2.7 Task-related loss calculation: Calculate the cross-entropy loss through the output of the model and the target labels; calculate the gradient through the backpropagation algorithm; update the parameters of the model using the optimization algorithm AdamW. Repeat the above steps.
[0278] Step Three: Visual modality adaptation.
[0279] In the second stage, the text modality is trained separately so that the model can effectively learn the representation of visual data.
[0280] 3.1 Primitive Decomposition: The input image is sliced into a series of non-overlapping blocks. Each block serves as the smallest visual unit and will be flattened to form a primitive. The specific steps are as follows:
[0281] 3.1.1 The input image \(I\in\mathbb{R}\) H×W×C (with image size \(H\times W\) and \(C\) channels) is sliced into \(P\times P\) small blocks.
[0282] 3.1.2 Each \(P\times P\) small block is flattened to obtain a vector of shape \(P\) 2 \(\times C\). Each vector is a primitive.
[0283] 3.2 Primitive Position Embedding: Similar to token position embedding, the position embedding of each primitive is calculated through sine and cosine functions: \(PE_{(i,2k)} = \sin(i / 10000\) 2k / d_image ) ; \(PE_{(i,2k + 1)}=\cos(i / 10000\) 2k / d_image ) ; where \(i\) is the position index of the primitive, \(k\) is the dimension index of the embedding, starting from 0; \(d_{image}\) is the dimension of the embedding space.
[0284] 3.3 Primitive Semantic Embedding: The semantic embedding is mapped to a fixed dimension \(d_{image}\) through a trainable projection matrix \(W_{patch}\): \(E(V_i)=V_i\times W_{patch}\); where \(V_i\) represents the primitive with position index \(i\).
[0285] 3.4 Final Primitive Embedding: Before feeding the data into the attention module, the position embedding and semantic embedding are added so that the model can consider both the semantic information and sequential information of the primitives. Given the \(i\)-th primitive in an input image, the embedding vector \(X_i\) image finally output by the image data embedding module is: \(X_i\) image \(=E(V_i)+PE_i\).
[0286] 3.5 Attention Calculation: The primitive embeddings of the picture are fed into repeatedly stacked attention modules with the same structure for processing. Each attention module consists of multiple layers of parametric attention layers and data attention layers. The calculation of each layer captures the dependencies between data through the self-attention mechanism.
[0287] Among them, for each attention module:
[0288] 3.5.1 Layer Normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer's calculation, the representation of each input is normalized so that its mean is 0 and its variance is 1. For the input vector X_image, it is updated as: X_image = LN(X_image) = (X - μ) / (σ + δ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and δ is a small constant to prevent the denominator from being 0.
[0289] 3.5.2 Q, K, V Matrix Calculation: Add a set of visual adaptation parameters K_P new and V_P new , whose number of parameters is the same as that of text adaptation parameters, initialized as a zero matrix, and the new parameters are updated as: K_P = [K_P, K_P new , V_P = [V_P, V_P new ; Considering the input data X_image and a set of optimizable parameter data K_P and V_P, the basic calculation formula of the parameter attention layer is: F(X_image, K_P, V_P) = Ξ(X_image · K_P T ) · V_P; where the dimension of the input data X_image is T × d_image, T represents the length of the primitive sequence, that is, the number of primitives, and each metadata is a vector of d_image dimensions. The dimensions of the parameter data K_P and V_P are 2n × d_image, 2n represents the number of parameter data, and each parameter is a vector of d_image dimensions. Ξ and f are the same as the previous formulas. Calculate the query matrix Q: Q = F(X_image, K_P Q , V_P Q ); where X_image is the input data, K_P Q and V_P Q are the corresponding parameter matrices with dimensions of 2n × d_image. Similarly, calculate the key matrix K and the value matrix V: K = F(X_image, K_P K , V_P K ); V = F(X_image, K_P V , V_P V ).
[0290] 3.5.3 Self-Attention Score Calculation: Based on the calculation results of the Q, K, V matrices of the parameter attention layer, the data attention layer applies the standard self-attention calculation formula to calculate the X_attn matrix, which is the same as the previous formula.
[0291] 3.5.4 O_attn Matrix Calculation: Input X_attn into the O-parameter attention calculation module to calculate the self-attention output matrix O_attn, using the same formula as before.
[0292] 3.5.5 Residual Processing: Add the self-attention output matrix O_attn to the layer-normalized matrix and perform layer normalization again, using the same formula as before.
[0293] 3.5.6 Attention Module Output Matrix Calculation: Input X_ffn into the FFN-parameter attention calculation module to calculate the attention module output matrix O_ffn, using the same formula as before.
[0294] 3.6 Task Decision: Convert the representation of the visual input into predictions for specific tasks. Input the feature representation O_ffn after the FFN-parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers, using the same formula as before.
[0295] 3.7 Task-related Loss Calculation: Calculate the cross-entropy loss through the output of the model and the target labels; calculate the gradients through the backpropagation algorithm; use the optimization algorithm AdamW to update the parameters of the model. Repeat the above steps.
[0296] Step Four: Visual-Linguistic Modal Alignment.
[0297] 4.1 Vector Dimension Alignment: Apply visual-text data. For each "visual-text" data pair, calculate the visual embedding and text embedding respectively, following the same steps as in Step One and Step Two. The linear projection layer aligns the dimensions of the visual and text features into a unified dimensional space through a trainable linear transformation: V_aligned = X_image · W_image, where W_image ∈ R d_image×d_text ; where W_image is the trainable linear transformation matrix used to project the image data embedding vector into the same dimension as the text data embedding vector. X_image is the image data embedding vector, and V_aligned is the aligned image data embedding vector.
[0298] 4.2 Vector Concatenation: When the dimensions of the visual and text features are aligned, concatenate their embedding vectors into a long vector to further fuse the information of the two modalities. The vector concatenation layer concatenates the already aligned visual and text features into a unified input vector so that they can participate in the calculation together in the subsequent attention mechanism: X = [X_text, X_aligned].
[0299] 4.3 Attention Calculation: Send X into attention modules with the same structure that are repeatedly stacked for processing. Each attention module consists of multiple layers of parameter attention layers and data attention layers. The calculation of each layer captures the dependencies between data through the self-attention mechanism.
[0300] Among them, for each attention module:
[0301] 4.3.1 Layer Normalization: Layer normalization is applied to each attention layer to ensure the stability and efficiency of the model during training. After each layer of calculation, the representation of each input is normalized so that its mean is 0 and its variance is 1. For the input vector X, it is updated as: X = LN(X) = (X - μ) / (σ + δ) · γ + β; where μ and σ are the mean and standard deviation of X respectively, γ and β are learnable scaling and offset parameters, and δ is a small constant to prevent the denominator from being 0.
[0302] 4.3.2 Q, K, V Matrix Calculation: Add a set of alignment parameters K_P align and V_P align , with the number of parameters being m, m < n, initialized as a zero matrix, and the new parameters are updated as: K_P = [K_P, K_P align , V_P = [V_P, V_P align . Considering the input data X and a set of optimizable parameter data K_P and V_P, the basic calculation formula of the parameter attention layer is: F(X, K_P, V_P) = Ξ(X · K_P T ) · V_P; where the dimension of the input data X is T × d_image, T represents the length of the concatenated sequence, that is, the number of word tokens plus graph tokens, and each metadata is a d_image-dimensional vector. The dimensions of the parameter data K_P and V_P are (2n + m) × d_image, and (2n + m) represents the number of parameter data, and each parameter is a d_image-dimensional vector. Calculate the query matrix Q: Q = F(X, K_P Q , V_P Q ); where X is the input data, K_P Q and V_P Q are the corresponding parameter matrices with dimensions of (2n + m) × d_image. Similarly, calculate the key matrix K and the value matrix V: K = F(X, K_P K , V_P K ); V = F(X, K_P V , V_P V ).
[0303] 4.3.3 Self-attention score calculation: The data attention layer calculates the X_attn matrix by applying the standard self-attention calculation formula based on the calculation results of the Q, K, and V matrices of the parameter attention layer.
[0304] 4.3.4 O_attn matrix calculation: Input X_attn into the O parameter attention calculation module to calculate the self-attention output matrix O_attn.
[0305] 4.4.5 Residual processing: Add the self-attention output matrix O_attn to the layer-normalized matrix. Perform layer normalization again.
[0306] 4.3.5 Attention module output matrix calculation: Input X_ffn into the FFN parameter attention calculation module to calculate the attention module output matrix O_ffn.
[0307] 4.4 Task decision-making: Convert the joint representation of the multi-modal input (the fused representation from vision and text) into a prediction for a specific task. Input the feature representation O_ffn after the FFN parameter attention fusion of the last attention module, and the decision layer processes this representation through a set of fully connected layers.
[0308] 4.5 Task-related loss calculation: Through contrastive learning, encourage positive image-text pairs to be close in the feature space and negative pairs to be far apart, aligning the feature spaces of the vision and text transformers; calculate the gradient through the backpropagation algorithm; use the optimization algorithm AdamW to update the parameters of the model. Repeat the above steps.
[0309] Step Five: Evaluation metrics.
[0310] 5.1 Text modality evaluation metrics: Use the machine translation evaluation metric BLEU based on n-gram matching, the generation metric ROUGE that focuses on recall, and the perplexity metric that measures the probability distribution of the predicted text by the language model.
[0311] 5.2 Visual modality evaluation metrics: Top-1 / Top-5 accuracy: Whether the highest probability category in the prediction result is correct (Top-1) or whether the correct category is included in the top five (Top-5).
[0312] 5.3 Multi-modal evaluation metrics: Cross-modal retrieval mAP: Calculate the retrieval accuracy of image-text matching, and calculate the average precision after sorting by Hamming distance; CLIP-I / CLIP-T: The CLIP feature similarity between the generated image and text, measuring the quality of modality alignment.
[0313] According to one aspect of the present application, a parameter-expandable multi-modal task continuous learning device includes:
[0314] At least one processor; and,
[0315] A memory communicatively connected to the at least one processor; wherein,
[0316] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the parameter-expandable multi-modal task continuous learning method described in any one of the above embodiments.
[0317] The present invention can be applied to visual question answering: generating accurate answers according to a given image and a natural language question; generating diagnostic information through visual question answering for medical questions based on images (such as X-ray films, CT scans, etc.); analyzing the road conditions ahead by a visual question answering system in the field of autonomous driving. It can also be used for image caption generation: generating a natural language text to describe the content of the input image; automatically generating text descriptions of pictures or videos on social media platforms; generating picture descriptions for visually impaired persons in the field of assistive technology. It can also be used for cross-modal retrieval: given a picture, returning the relevant text description; given a piece of text, returning the relevant image. On e-commerce platforms, automatically retrieving similar products or similar descriptive texts according to the pictures uploaded by users. Retrieving relevant medical literature or case descriptions according to pathological images. It can also be used for sentiment analysis and emotion recognition: analyzing the given text or image content and identifying the emotions or moods therein; analyzing the emotional tendency according to the image and text content posted by users, and applying it to fields such as brand monitoring and marketing; intelligent customer service, analyzing the emotions of the conversation between users and customer service, and adjusting coping strategies according to the emotional changes. It can also be used for visual language recommendation systems: combining multiple data sources (such as images, texts, and user behaviors, etc.) to generate more personalized recommendations; e-commerce recommendation systems, recommending similar products according to the product images and descriptions browsed by users; social media recommendations, recommending similar image or text content according to the interaction behaviors of users with the content of pictures or posts.
[0318] The present invention receives system input data, pre-processes and encodes the original visual data and the original text data respectively, and generates a pre-processing feature matrix; calculates position information and semantic information according to the pre-processing feature matrix to obtain an adaptive feature matrix; performs attention calculation and parameter optimization according to the adaptive feature matrix and the initial parameter matrix to obtain an optimized parameter matrix; performs projection processing and dynamic alignment on the optimized parameter matrix to obtain an aligned feature matrix; performs knowledge extraction and importance evaluation according to the aligned feature matrix to obtain a distillation feature matrix; performs error detection and compensation processing on the distillation feature matrix to obtain an error correction feature matrix; performs task feature extraction and decision processing according to the error correction feature matrix to obtain a prediction result matrix. The present invention realizes continuous optimization and performance improvement of visual language tasks through systematic feature processing, parameter optimization and knowledge transfer mechanisms. First, high-quality representation of input data is ensured through bimodal feature extraction and encoding; then, the expressive power of features is enhanced through position encoding and semantic encoding; then, effective interaction between features is realized through multi-layer parameter optimization and attention calculation; then, the fusion quality of features and the generalization ability of the model are improved through projection alignment and knowledge distillation; finally, the accuracy and reliability of model prediction are ensured through error detection and task adaptation. The present invention not only solves key problems such as feature heterogeneity, semantic misalignment, and knowledge transfer in visual language tasks, but also achieves continuous improvement in model performance through the organic coordination of various modules. In particular, the present invention enables the model to flexibly adapt to the requirements of tasks of different scales and types through the scalable design of parameters, while maintaining the stability and controllability of the optimization process, providing a reliable technical foundation for the further development of visual language tasks.
[0319] The preferred embodiments of the present invention are described in detail above; however, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A parameter-expandable multi-modal task continual learning method, characterized in that It includes the following steps: S1. Receive the system input data and perform preprocessing encoding to generate a preprocessing feature matrix of the input data; where the system input data includes original visual data and original text data; S2. Based on the preprocessing feature matrix of the input data, calculate the feature position information and semantic information respectively to obtain an adaptive feature matrix of the input data; S3. Perform attention calculation and parameter optimization according to the adaptive feature matrix of the input data and the initial parameter matrix pre-stored in the system to obtain an optimized parameter matrix of the input data; S4. Perform projection processing and dynamic alignment on the optimized parameter matrix to obtain an aligned feature matrix of the input data; S5. Perform knowledge extraction and importance evaluation based on the aligned feature matrix to obtain a distilled feature matrix of the input data; S6. Perform error detection and compensation processing on the distilled feature matrix to obtain an error-corrected feature matrix of the input data; S7. Perform task feature extraction and decision-making processing on the error-corrected feature matrix to obtain a prediction result matrix of the input data.
2. The method for continuously learning parameter-expandable multi-modal tasks according to claim 1, wherein Step S1 is further as follows: S11. Receive the original visual data and divide it into primitive blocks; perform feature extraction on each primitive block to generate an initial visual feature sequence; retrieve visual feature items matching the initial visual feature sequence from the pre-stored visual feature dictionary, combine them into a visual feature matrix and input it into a pre-trained encoding model to generate an encoded visual feature matrix; S12. Receive the original text data and perform word segmentation processing to generate an initial token sequence; retrieve text feature items matching the initial token sequence from the pre-stored text feature dictionary, combine them into a text feature matrix and input it into a pre-trained encoding model to generate an encoded text feature matrix; S13. Concatenate the encoded visual feature matrix and the encoded text feature matrix to generate an initial feature matrix; Calculate the mean and standard deviation of each feature in the initial feature matrix and perform normalization processing to obtain a preprocessing feature matrix.
3. The parameter-expandable multi-modal task continuous learning method according to claim 2, characterized in that Step S2 is further as follows: S21. Receive the visual features in the preprocessing feature matrix, calculate sine and cosine position encodings to generate a visual position encoding matrix; retrieve the corresponding semantic information using the visual semantic dictionary according to the visual features to generate a visual semantic encoding matrix; perform weighted superposition of the visual features with the visual position encoding matrix and the visual semantic encoding matrix to obtain an adaptive visual feature matrix of the input data; S22. Receive the text features in the preprocessing feature matrix, calculate sine position encoding and cosine position encoding to generate a text position encoding matrix; retrieve the corresponding semantic information using the text semantic dictionary according to the text features to generate a text semantic encoding matrix; perform weighted superposition of the text features with the text position encoding matrix and the text semantic encoding matrix to obtain an adaptive text feature matrix of the input data; S23. Perform feature concatenation and feature normalization processing on the adaptive visual feature matrix and the adaptive text feature matrix to obtain an adaptive feature matrix of the input data.
4. The parameter-expandable multi-modal task continuous learning method according to claim 3, wherein Step S3 is further as follows: S31. Receive the initial parameter matrix pre-stored and divide it into a predetermined number of sub-matrices, perform independent initialization operations to generate a multi-layer parameter matrix; S32. Generate a query matrix, a key matrix, and a value matrix based on the multi-layer parameter matrix, and perform matrix multiplication operations in combination with the adaptive feature matrix to generate a group of feature transformation matrices; Perform attention calculation on the group of feature transformation matrices to obtain an attention feature matrix; S33. Calculate the gradient value of the attention feature matrix, generate a gradient matrix and input it into a preset optimizer to generate a parameter update matrix; Perform parameter update operations in combination with the multi-layer parameter matrix to obtain an optimized parameter matrix.
5. The method for continuously learning parameter-expandable multi-modal tasks according to claim 4, characterized in that Step S4 is further as follows: S41. Divide the optimized parameter matrix into visual and text parts, perform matrix multiplication operations with the pre-stored projection matrix respectively, generate visual projection features and text projection features and splice them in the feature dimension to obtain a projection feature matrix; S42. Calculate the similarity scores between the features in the projection feature matrix to generate a similarity matrix; Perform matrix multiplication operations in combination with the pre-stored alignment parameter matrix to generate an alignment weight matrix; Perform weighted summation of the alignment weight matrix and the projection feature matrix to obtain a primary alignment feature matrix; S43. Calculate the statistical distribution information of the primary alignment feature matrix to generate a distribution statistics matrix; Perform normalization processing on the primary alignment feature matrix according to the distribution statistics matrix and perform residual connection operations to obtain the final alignment feature matrix.
6. The parameter-expandable multi-modal task continuous learning method according to claim 5, characterized in that Step S5 is further as follows: S51. Based on the pre-stored historical parameter matrix, input the alignment feature matrix into the pre-configured teacher and student models, respectively obtain the teacher and student output matrices and scale them using the temperature coefficient to generate a scaled feature matrix; Calculate the cross-entropy of the outputs of the teacher and student models in the scaled feature matrix to obtain a knowledge feature matrix; S52. Calculate the influence degree of each parameter in the knowledge feature matrix on the model output, generate a parameter sensitivity matrix and normalize it to obtain a parameter importance matrix; Filter and update the knowledge feature matrix based on the parameter importance matrix to obtain a primary distillation feature matrix; S53. Sort the parameters in the primary distillation feature matrix according to their importance, generate a sorted parameter matrix and perform threshold filtering to obtain a filtered parameter matrix; perform parameter merging in combination with the pre-stored historical parameter matrix to update and obtain the final distillation feature matrix.
7. The parameter-expandable multi-modal task continuous learning method according to claim 6, characterized in that, Step S6 is further as follows: S61. Extract label information based on the pre-stored validation data matrix, calculate the prediction error and parameter temporal variation of the distillation feature matrix to obtain an error metric matrix; S62. Calculate a compensation coefficient according to the error metric matrix to generate a compensation coefficient matrix; Perform matrix multiplication operations in combination with the distillation feature matrix, generate a compensation feature matrix and normalize it to obtain a normalized compensation matrix; Perform residual connection and non-linear transformation on the normalized compensation matrix and the distillation feature matrix to generate an activation feature matrix; Perform a validation operation on the activation feature matrix and the pre-stored validation data matrix to obtain an error correction feature matrix.
8. The parameter-expandable multi-modal task continuous learning method according to claim 7, characterized in that Step S7 is further as follows: S71. Group the error correction feature matrix according to the task type to generate a grouped feature matrix; apply task embedding vectors to each group of features to generate an embedded feature matrix; perform feature fusion operations in combination with the pre-stored task data matrix to obtain a task feature matrix; S72. Input the task feature matrix into the pre-stored decision parameter matrix for linear transformation and numerical standardization to generate a standardized decision matrix; Calculate the loss with the pre-stored task label matrix to obtain a primary prediction result matrix; S73. Calculate the confidence of the primary prediction result matrix, generate a confidence matrix and apply threshold filtering to obtain a filtered result matrix; Combine and optimize the filtered result matrix with the primary prediction result matrix, and update to obtain the final prediction result matrix.
9. The parameter-expandable multi-modal task continuous learning method according to claim 8, characterized in that Step S11 is further as follows: S111. Receive the original visual data, perform bilateral filtering operation and adaptive contrast enhancement to obtain enhanced image data and convert it into image tensor data; S112. Use the sliding window algorithm to divide the image tensor data into adjacent overlapping image blocks, and perform boundary compensation and bilinear interpolation resampling to obtain a unified primitive block sequence; S113. Perform multi-scale pyramid transformation and local descriptor extraction on the unified primitive block sequence, generate a local feature sequence and perform feature aggregation operation to obtain an initial visual feature sequence; S114. Perform local sensitive hashing encoding, nearest neighbor search and similarity weight calculation on the initial visual feature sequence to obtain a weighted feature sequence; S115. Rearrange the weighted feature sequence into a matrix form and perform sparsification processing to obtain a visual feature matrix; S116. Input the visual feature matrix into the pre-trained encoding model and perform residual connection operation and layer normalization processing to obtain an encoded visual feature matrix.
10. A parameter-expandable multi-modal task continual learning device, characterized in that, Including: At least one processor; And, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the parameter-expandable multi-modal task continuous learning method according to any one of claims 1 to 9.
Citation Information
Cited By
Multi-modal learning and modal-level adaptive differential privacy clipping method used by multi-modal learning
CN121051796A
Multimodal learning and its use of modal level adaptive differential privacy pruning method
CN121051796B
Closed data space model test adaptive method based on two-stage feature whitening
CN121884036A