An Optimization Method and System for Encoder Architecture Incorporating Multimodal Features
By designing a trainable encoder for each modal, using a hybrid expert model and grouping attention mechanism, dynamically adjusting the number of grids, and using contrast learning technology to align the relationship between modal and language modality, the problem of incomplete feature extraction in multimodal data processing is solved, and more efficient feature fusion and accuracy are achieved.
Patent Information
- Application Number
- CN202510109645.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-23
AI Technical Summary
When processing multimodal data, existing encoders find it difficult to make full use of complementary information between each mode, and lack the ability to flexibly adapt to different mode characteristics, resulting in insufficient feature extraction.
The trainable encoder is designed for each modal, using a hybrid expert model structure and grouping attention mechanism, dynamically adjusting the number of grids, using contrast learning technology to align the relationship between the modal and the language modal, and computed cosine similarity through grouping attention generation in the inference stage.
The effect and performance of multimodal feature fusion is improved, the flexibility and adaptability of the model is enhanced, the accuracy and completeness of feature extraction is improved, and the ability of cross-modal retrieval and multimodal understanding is enhanced.
Smart Images

Figure CN120088346B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an optimization method and system for an encoder architecture that fuses multi-modal features. Background Art
[0002] With the rapid development of information technology, multi-modal data (such as videos, images, audio, depth information, and infrared rays, etc.) has been widely used in various fields. In order to effectively process this multi-modal data and extract valuable information from it, researchers have continuously explored and improved multi-modal feature fusion methods.
[0003] Specifically, when traditional encoders process multi-modal data, some are difficult to fully utilize the complementary information between modalities, resulting in incomplete feature extraction. At the same time, some existing methods use a fixed network structure when processing different modal data, lacking the flexible adaptability to the characteristics of different modalities. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an optimization method and system for an encoder architecture that fuses multi-modal features, which can further improve the effect and performance of multi-modal feature fusion.
[0005] To solve the above technical problem, the technical solution of the present invention is as follows:
[0006] In a first aspect, an optimization method for an encoder architecture that fuses multi-modal features, the method includes:
[0007] Design a corresponding trainable encoder for each modality, where the modalities include videos, images, audio, depth information, and infrared rays, to achieve a unified representation of multi-modal data in a shared embedding space;
[0008] Adopt a mixture of experts model structure as the core of each encoder, with multiple expert networks built-in, and dynamically select the corresponding expert for processing according to the characteristics of the input data;
[0009] Use contrastive learning technology to align the relationship between any modality and the language modality;
[0010] During the image encoding process, through a grouped attention mechanism, multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, each sub-image corresponds to an independent text description, and the attention of patch tokens is calculated within the same group;
[0011] During the training process, dynamically adjust the number of grids in the grid, and allow the sizes of the sub-images within the same synthesized image to be not exactly the same, so as to improve the model's processing ability for different granularity patch tokens;
[0012] In the inference stage, the subgraphs corresponding to multiple task requests are synthesized into a large image and input into the model. Through the grouped attention mechanism, CLS tokens are generated for each group of subgraphs, and the cosine similarity is calculated with the vectors output by the text encoder. The losses of all groups are averaged as the final result.
[0013] Furthermore, contrastive learning techniques are used to align the relationships between any modality and the language modality, including:
[0014] Collect a dataset containing multimodal data, including images, audio, and corresponding texts, sentences;
[0015] For the image modality, a convolutional neural network is used to extract image features; for the audio modality, an audio processing model is used to extract audio features; for the language modality, a word embedding model is used to extract text features;
[0016] Define positive sample pairs and negative sample pairs. Positive sample pairs refer to the pairings between different modalities from the same data source, and negative sample pairs are the modality pairings from different data sources;
[0017] Construct a contrastive loss function, and use the dataset containing positive and negative sample pairs to train a multimodal alignment model. During the training process, the parameters of the model are updated by optimizing the contrastive loss function;
[0018] Use the validation set to evaluate the performance of the model to obtain the evaluation results, and optimize the model according to the evaluation results.
[0019] Furthermore, in the image encoding process, through the grouped attention mechanism, multiple small-resolution subgraphs are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, each subgraph corresponds to an independent text description, and the attention of patch tokens is calculated within the same group, including:
[0020] Generate a set of small-resolution subgraphs so that each subgraph has a corresponding text description; adjust all subgraphs to the same size;
[0021] Use a convolutional neural network to extract features from each subgraph to generate a series of feature vectors;
[0022] Divide the feature map of each subgraph into multiple small blocks, and each small block is converted into a patch token;
[0023] Use a natural language processing model to encode the text description corresponding to each subgraph to generate text feature vectors;
[0024] Group the patch tokens of all sub - graphs according to their corresponding text descriptions, where each group contains all the patch tokens of the sub - graphs from the same text description;
[0025] Within each group, use the self - attention mechanism to calculate the attention weights between the patch tokens.
[0026] Furthermore, based on the evaluation results, during the training process, dynamically adjust the number of grid cells and allow the sizes of the sub - graphs within the same synthetic image to be not exactly the same, so as to improve the model's processing ability for patch tokens of different granularities, including:
[0027] Initialize a model containing an image encoder and a decoder. The encoder is used to generate patch tokens, and the decoder is used to reconstruct the image from the patch tokens;
[0028] Define a strategy to dynamically adjust the number of grid cells, allow the sizes of the sub - graphs within each grid cell to be not exactly the same, set a size range, and randomly select and adjust the size of each sub - graph during each training;
[0029] According to the current number of grid cells and the size of the sub - graphs, dynamically divide each sub - graph to generate patch tokens of the corresponding size;
[0030] Use the image encoder to encode each patch token and extract features;
[0031] Design a loss function that reflects the changes in the dynamic grid cells and the sizes of the sub - graphs. During the training process, according to the dynamically adjusted number of grid cells and the size of the sub - graphs, correspondingly adjust the learning rate and batch size parameters.
[0032] Furthermore, according to the current number of grid cells and the size of the sub - graphs, dynamically divide each sub - graph to generate patch tokens of the corresponding size, including:
[0033] According to the task requirements or user input, determine the current number of grid cells and the size of the sub - graphs in each grid cell; set the initial population size, crossover rate, mutation rate, and number of iterations of the genetic algorithm;
[0034] Use binary encoding to represent the division method of each sub - graph, where each bit represents whether a region of the sub - graph is divided into a patch token; according to the encoding scheme, randomly generate the initial population, and each individual represents a possible division method of the sub - graph;
[0035] Define a fitness function to evaluate the pros and cons of each division method; calculate the fitness value of each individual in the initial population;
[0036] According to the fitness value, select the corresponding individuals in the population to enter the next generation; randomly pair the selected individuals and perform crossover operations according to the set crossover rate to generate new individuals; perform mutation operations on the newly generated individuals according to the set mutation rate, and repeat the fitness calculation, selection, crossover, and mutation operations until the set number of iterations is reached. During the iteration process, record the corresponding individuals as the final solution;
[0037] Decode the final solution obtained by the genetic algorithm into a specific subgraph partitioning method;
[0038] According to the decoded partitioning method, partition each subgraph to generate patch tokens of the corresponding size.
[0039] Furthermore, in designing the loss function that reflects the dynamic grid and subgraph size changes, the calculation formula of the loss function L is:
[0040]
[0041] where, x i is the i-th pixel point in the original image; is the i-th pixel point after reconstruction, N is the total number of pixels in the image; G j represents the size of the j-th dynamically adjusted grid, G ef is the reference grid size, M is the number of grids; S k is the size of the k-th subgraph, S i is the ideal subgraph size, P is the number of subgraphs; i, j, and k represent index values.
[0042] Furthermore, the calculation formula of the fitness function is:
[0043]
[0044] where, Q is the number of subgraphs, M b is the average absolute error of feature extraction corresponding to each subgraph b; A b represents the area of each subgraph, T b represents the calculation time corresponding to the subgraph; G is the number of all images, S c is the structural similarity index of image c; R is the number of patches adjusted during the training process, Δe p is the change amplitude of the patch size; q is the index value; w1, w2, and w3 are weight coefficients.
[0045] In the second aspect, an encoder architecture optimization system that fuses multi-modal features includes:
[0046] A design module for designing a corresponding trainable encoder for each modality, where the modalities include video, image, audio, depth information, and infrared, to achieve a unified representation of multi-modal data in a shared embedding space; adopting a mixture-of-experts model structure as the core of each encoder, with multiple expert networks built-in, and dynamically selecting the corresponding expert for processing according to the characteristics of the input data; using contrastive learning technology to align the relationship between any modality and the language modality;
[0047] A processing module for synthesizing multiple small-resolution sub-images into a high-resolution image of a fixed size in the form of an N-grid during image encoding, arranging them in the form of an N-grid, each sub-image corresponding to an independent text description, and calculating the attention of patch tokens within the same group;
[0048] A training module for dynamically adjusting the number of grids during the training process and allowing the sizes of sub-images within the same synthesized image to be not exactly the same, to improve the model's processing ability for different granularity patch tokens;
[0049] A prediction module for the inference stage, synthesizing sub-images corresponding to multiple task requests into a large image and inputting it into the model, generating CLS tokens for each group of sub-images through the grouped attention mechanism, calculating the cosine similarity with the vector output by the text encoder, and taking the average of the losses of all groups as the final result.
[0050] In a third aspect, a computing device includes:
[0051] One or more processors;
[0052] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described above.
[0053] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the method described above.
[0054] The above solution of the present invention has at least the following beneficial effects:
[0055] Designing a corresponding trainable encoder for each modality (video, image, audio, depth information, and infrared) can achieve a unified representation of multi-modal data in a shared embedding space. This design enables data of different modalities to be fused and processed within the same framework, thereby more comprehensively capturing the information in the data and improving the accuracy and integrity of feature extraction.
[0056] Adopt a mixture of experts model structure as the core of the encoder, with multiple expert networks built-in, which can dynamically select the corresponding expert for processing according to the characteristics of the input data. This design enhances the flexibility and adaptability of the model, enabling the model to efficiently process different modalities of data, and improving the processing efficiency and accuracy.
[0057] Use contrastive learning technology to align the relationship between any modality and the language modality, which helps the model better understand the relevance and complementarity between different modalities of data. This alignment method can improve the performance of the model in tasks such as cross-modal retrieval and multi-modal understanding.
[0058] Introduce a grouped attention mechanism during image encoding, combine multiple small-resolution sub-images into a high-resolution image of a fixed size, and arrange them in the form of an N-grid. This design enables the model to capture local features in the image more finely while maintaining the integrity of global information. Each sub-image corresponds to an independent text description, further enhancing the relevance between the image and the text.
[0059] Dynamically adjust the number of grids during training and allow the sizes of sub-images within the same synthesized image to be not exactly the same. This design enables the model to adapt to different granularity patch tokens, improving the generalization ability and robustness of the model.
[0060] During the inference stage, synthesize the sub-images corresponding to multiple task requests into a large image and input it into the model. Generate CLS tokens for each group of sub-images through the grouped attention mechanism, and calculate the cosine similarity with the vector output by the text encoder. This method can more accurately measure the similarity between the image and the text, thereby providing more accurate inference results. Take the average of the losses of all groups as the final result, further improving the stability and reliability of the model. Brief Description of the Drawings
[0061] Figure 1 It is a schematic flowchart of a method for optimizing the encoder architecture that fuses multi-modal features provided by an embodiment of the present invention.
[0062] Figure 2 It is a schematic diagram of a system for optimizing the encoder architecture that fuses multi-modal features provided by an embodiment of the present invention.
[0063] Figure 3 It is a schematic drawing of a method for optimizing the encoder architecture that fuses multi-modal features provided by an embodiment of the present invention, which generalizes the N-grid to videos.
[0064] Figure 4 It is a schematic diagram of an expert model based on the Transformer architecture for a method for optimizing the encoder architecture that fuses multi-modal features provided by an embodiment of the present invention. Detailed implementation manners
[0065] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0066] As Figure 1 、 Figure 3 and Figure 4 shown, an embodiment of the present invention proposes an optimization method for an encoder architecture that fuses multi-modal features. The method includes the following steps:
[0067] Design corresponding trainable encoders for each modality, where the modalities include video, image, audio, depth information, and infrared rays, to achieve unified representation of multi-modal data in a shared embedding space;
[0068] Adopt a mixture-of-experts model structure as the core of each encoder, with multiple expert networks built in, and dynamically select the corresponding expert for processing according to the characteristics of the input data;
[0069] Use contrastive learning technology to align the relationship between any modality and the language modality;
[0070] During the image encoding process, through a grouped attention mechanism, multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, each sub-image corresponds to an independent text description, and the attention of patch tokens is calculated within the same group;
[0071] During the training process, dynamically adjust the number of grids in the grid, and allow the sizes of the sub-images within the same synthesized image to be not exactly the same, so as to improve the model's processing ability for patch tokens of different granularities;
[0072] In the inference stage, the sub-images corresponding to multiple task requests are synthesized into a large image and input into the model. Through the grouped attention mechanism, CLS tokens are generated for each group of sub-images, the cosine similarity is calculated with the vector output by the text encoder, and the average of the losses of all groups is taken as the final result.
[0073] In the embodiments of the present invention, trainable encoders are designed for each modality (video, image, audio, depth information, and infrared), enabling unified representation of multimodal data in a shared embedding space. This design allows data of different modalities to be fused and processed within the same framework, thereby capturing information in the data more comprehensively and improving the accuracy and integrity of feature extraction. The mixture of experts model structure is adopted as the core of the encoder, with multiple expert networks built-in, which can dynamically select the corresponding expert for processing according to the characteristics of the input data. This design enhances the flexibility and adaptability of the model, enabling the model to efficiently process data of different modalities and improving the processing efficiency and accuracy. The contrastive learning technique is used to align the relationship between any modality and the language modality, which helps the model better understand the correlation and complementarity between different modality data. This alignment method can improve the performance of the model in tasks such as cross-modal retrieval and multimodal understanding. A grouped attention mechanism is introduced during image encoding, where multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size and arranged in the form of an N-grid. This design enables the model to capture local features in the image more precisely while maintaining the integrity of global information. Each sub-image corresponds to an independent text description, further enhancing the correlation between the image and the text. During training, the number of grids is dynamically adjusted, and the sizes of the sub-images within the same synthesized image are allowed to be not exactly the same. This design enables the model to adapt to patch tokens of different granularities, improving the generalization ability and robustness of the model. In the inference stage, the sub-images corresponding to multiple task requests are synthesized into a large image and input into the model. The grouped attention mechanism generates CLS tokens for each group of sub-images and calculates the cosine similarity with the vector output by the text encoder. This method can more accurately measure the similarity between the image and the text, thereby providing more accurate inference results. The average of the losses for all groups is taken as the final result, further improving the stability and reliability of the model.
[0074] In the embodiments of the present invention, the contrastive learning technique is used to align the relationship between any modality and the language modality, including:
[0075] Collect a dataset containing multimodal data, including images, audio, and corresponding texts, sentences;
[0076] For the image modality, use a convolutional neural network to extract image features; for the audio modality, use an audio processing model to extract audio features; for the language modality, use a word embedding model to extract text features;
[0077] Define positive sample pairs and negative sample pairs. Positive sample pairs refer to the pairings between different modalities from the same data source, while negative sample pairs are the modality pairings from different data sources. Specifically, positive sample pairs refer to the pairings between different modality data from the same data source. For example, a picture and its related text description, or an audio clip and its corresponding text record, can both form positive sample pairs. These positive sample pairs reflect the natural correlation between different modalities and are important bases for the model to learn the corresponding relationships between modalities. Negative sample pairs, on the other hand, refer to the modality data pairings from different data sources. This means that there is no direct correlation between these modality data in reality. The introduction of negative sample pairs is to help the model distinguish irrelevant modality data, thereby improving the accuracy of the model in distinguishing the relationships between different modalities.
[0078] Construct a contrastive loss function and train a multi-modal alignment model using a dataset containing positive and negative sample pairs. During the training process, update the model's parameters by optimizing the contrastive loss function. Specifically, the design purpose of the contrastive loss function is to enable the model to maximize the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. Forms of the contrastive loss function include losses based on cosine similarity, triplet loss, etc. These loss functions all aim to widen the distance between positive and negative sample pairs in the feature space. Prepare a dataset containing positive and negative sample pairs for training the multi-modal alignment model. Each sample in the dataset should be clearly labeled as a positive or negative sample pair. During the training process, the model receives paired modality data as input and calculates the loss value according to the contrastive loss function. Through the backpropagation algorithm, the model adjusts its parameters to minimize the loss function, thereby learning the effective representations and alignment methods between different modality data.
[0079] Evaluate the performance of the model using a validation set to obtain the evaluation results and optimize the model according to the evaluation results. Specifically, prepare an independent validation set for evaluating the performance of the model. The validation set contains data similar to but disjoint from the training set to ensure the fairness of the evaluation. Use the validation set to evaluate the trained multi-modal alignment model. Evaluation metrics may include accuracy, recall, F1 score, etc., depending on the requirements of the task. According to the evaluation results on the validation set, optimize the model. Optimization involves adjusting the model architecture, increasing or decreasing the number of network layers, modifying hyperparameters such as the learning rate, and can also try different forms of the contrastive loss function or adjust the hyperparameters in the loss function to optimize the model performance. After one or more optimizations, re-evaluate the performance of the model on the validation set. If the performance improves, save the current model state. If the performance deteriorates or does not meet expectations, continue to adjust the model configuration and retrain.
[0080] In the embodiments of the present invention, through contrastive learning techniques, the model can learn the internal correlations between different modality data, making the relationship between any modality and the language modality closer. This enhanced correlation helps improve the performance of the model in tasks such as cross-modal retrieval and understanding. During the contrastive learning process, the model needs to distinguish between positive and negative sample pairs, which requires the model to extract more discriminative and representative features. Therefore, through contrastive learning techniques, the feature extraction ability of each modality is improved, providing more accurate feature inputs for subsequent tasks. Since contrastive learning uses a large number of positive and negative sample pairs for training, the model can be exposed to more data variability and modality combinations, thereby improving the generalization ability of the model. This ability makes the model more robust when dealing with new data or unseen modality combinations. By constructing an effective contrastive loss function and using this function to update the model's parameters, the training process of the model can be made more efficient and targeted. This optimization not only accelerates the convergence speed of the model but also improves the fitting degree of the model on the training data. Evaluating the performance of the model using a validation set can obtain objective evaluation results. These results not only reflect the performance of the model on the current task but also provide clear guidance for further optimizing the model. By continuously adjusting and optimizing the model parameters, the performance of the model can be further improved.
[0081] In the embodiments of the present invention, during the image encoding process, through a grouped attention mechanism, multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, and each sub-image corresponds to an independent text description, and the attention of patch tokens is calculated within the same group, including:
[0082] Generate a set of small-resolution sub-images so that each sub-image has a corresponding text description; adjust all sub-images to the same size;
[0083] Use a convolutional neural network to extract features from each sub-image, generating a series of feature vectors;
[0084] Divide the feature map of each sub-image into multiple small blocks, and each small block is converted into a patch token;
[0085] Use a natural language processing model to encode the text description corresponding to each sub-image, generating text feature vectors;
[0086] Group all the patch tokens of the sub-images according to their corresponding text descriptions, and each group contains all the patch tokens of the sub-images from the same text description;
[0087] Within each group, use the self-attention mechanism to calculate the attention weights between patch tokens.
[0088] In the embodiments of the present invention, by decomposing a high-resolution image into multiple small-resolution sub-images, the model can capture the local features of the image more meticulously. This way of grouped processing enables the model to perform independent feature extraction for each sub-image, thereby improving the fineness and accuracy of feature extraction. Each sub-image has a corresponding text description, and this clear correspondence helps the model better understand the image content and closely associate the image features with the text description. The enhancement of this relevance helps improve the performance of the model in tasks such as cross-modal retrieval and image understanding. By adjusting the sub-images to the same size, the model can process images of different sizes and resolutions, improving the flexibility of the model. At the same time, the grouped attention mechanism allows the model to dynamically allocate attention resources when processing different sub-images, further enhancing the adaptability of the model. Using a convolutional neural network to extract features from each sub-image to generate a series of feature vectors, this method can greatly reduce the computational complexity and improve the computational efficiency compared to extracting features from the entire high-resolution image. At the same time, the grouped attention mechanism also limits the scope of attention calculation, further optimizing the computational process. By dividing the feature map of each sub-image into multiple small blocks and converting them into patch tokens, the model can understand the image content at a finer-grained level. This processing method helps the model better handle local changes and occlusions in the image, improving the robustness of the model. Encoding the text description corresponding to each sub-image through a natural language processing model to generate text feature vectors and fusing them with the image features. The fusion of this multi-modal information helps the model more comprehensively understand the task requirements and improve the processing ability in complex scenarios.
[0089] In the embodiments of the present invention, based on the evaluation results, the number of grid cells is dynamically adjusted during the training process, and the sizes of the sub-images within the same synthetic image are allowed to be not exactly the same to improve the model's processing ability for different granularity patch tokens, including:
[0090] Initializing a model including an image encoder and a decoder, where the encoder is used to generate patch tokens and the decoder is used to reconstruct the image from the patch tokens;
[0091] Defining a strategy to dynamically adjust the number of grid cells, allowing the sizes of the sub-images within each grid cell to be not exactly the same, setting a size range, and randomly selecting to adjust the size of each sub-image during each training;
[0092] Dynamically dividing each sub-image according to the current number of grid cells and the size of the sub-image to generate patch tokens of corresponding sizes;
[0093] Encode each patch token using an image encoder to extract features;
[0094] Design a loss function that reflects the changes in the dynamic grid and sub - figure sizes. During training, according to the dynamically adjusted number of grids and sub - figure sizes, adjust the learning rate and batch size parameters accordingly.
[0095] In the embodiments of the present invention, by dynamically adjusting the number of grids and allowing the sub - figure sizes to be not exactly the same, the model can encounter more diverse image partitioning methods during training. This flexibility enables the model to better adapt to patch tokens of different granularities, improving the model's processing ability for images of different sizes and resolutions. Due to the dynamic change of the sub - figure size, the model needs to learn to extract effective features from images of different sizes. This training method can force the model to pay more attention to the local details and global structure of the image, thereby enhancing the accuracy and comprehensiveness of feature extraction. In practical applications, the input images often have different sizes and resolutions. By dynamically adjusting the number of grids and sub - figure sizes during training, the model can better handle this change, reduce the sensitivity to the input size, and thus improve the robustness of the model. Dynamically adjusting the learning rate and batch size parameters can reasonably allocate computing resources according to the current number of grids and sub - figure sizes. During training, dynamically adjusting these parameters according to the complexity of the task can effectively improve the training efficiency and avoid resource waste. Due to the dynamic change of the sub - figure size, the model needs to understand and process image information at multiple scales. This multi - scale feature learning helps the model capture richer image information, thereby enhancing its performance in various tasks. By introducing a training strategy with dynamic grid and sub - figure size changes, the model can encounter more diverse data situations. This diverse training data helps improve the generalization ability of the model, enabling it to maintain good performance when facing unseen data.
[0096] In the embodiments of the present invention, according to the current number of grids and sub - figure sizes, dynamically partition each sub - figure to generate patch tokens of corresponding sizes, including:
[0097] Determine the current number of grids and the size of each sub - figure in each grid according to the task requirements or user input; set the initial population size, crossover rate, mutation rate, and number of iterations of the genetic algorithm;
[0098] Use binary encoding to represent the partitioning method of each sub - figure, where each bit represents whether a region of the sub - figure is divided into a patch token; according to the encoding scheme, randomly generate the initial population, and each individual represents a possible sub - figure partitioning method;
[0099] Define a fitness function to evaluate the advantages and disadvantages of each partitioning method; calculate the fitness value for each individual in the initial population;
[0100] According to the fitness value, select the corresponding individuals in the population to enter the next generation; randomly pair the selected individuals and perform crossover operations according to the set crossover rate to generate new individuals; perform mutation operations on the newly generated individuals according to the set mutation rate, and repeat the fitness calculation, selection, crossover, and mutation operations until the set number of iterations is reached. During the iteration process, record the corresponding individuals as the final solutions;
[0101] Decode the final solution obtained by the genetic algorithm into a specific subgraph partitioning method;
[0102] According to the decoded partitioning method, partition each subgraph to generate patch tokens of the corresponding size.
[0103] In the embodiments of the present invention, the genetic algorithm is a global optimization search algorithm that finds the optimal solution to a problem by simulating natural selection and genetic mechanisms. In the embodiments of the present invention, applying the genetic algorithm to dynamically partition each subgraph can more effectively explore various possible partitioning methods, find the global optimal or approximate optimal subgraph partitioning scheme, thereby improving the generation quality of patch tokens. Since the genetic algorithm does not depend on the specific nature of the problem, it can flexibly handle various complex subgraph partitioning problems. At the same time, by dynamically adjusting the number of grid cells and the size of subgraphs and combining the optimization ability of the genetic algorithm, the model can adaptively generate the most suitable patch tokens according to different situations, enhancing the flexibility and self - adaptability of the model. Traditional subgraph partitioning methods may require more manual settings and adjustments, while in the embodiments of the present invention, the genetic algorithm automatically performs subgraph partitioning, reducing human intervention and improving the objectivity and accuracy of partitioning. Although the genetic algorithm may require a certain number of iterations to reach the optimal solution, due to its parallel search characteristics, it can explore better solutions in a relatively short time. In addition, once a suitable partitioning method is found, patch tokens of the corresponding size can be generated quickly, thereby improving the overall computational efficiency. Optimizing subgraph partitioning through the genetic algorithm can generate more reasonable and effective patch tokens, which can more accurately reflect the local features and global structure of the image. This will help improve the performance of subsequent image encoding and decoding, and thus enhance the performance of the entire model.
[0104] In the embodiments of the present invention, in designing the loss function reflecting the changes in dynamic grid cells and subgraph sizes, the calculation formula of the loss function L is:
[0105]
[0106] where x i is the i-th pixel point in the original image; is the i-th pixel point after reconstruction, N is the total number of pixels in the image; G j represents the size of the j-th dynamically adjusted grid, G ef is the reference grid size, M is the number of grids; S k is the size of the k-th sub-image, S i is the ideal sub-image size, P is the number of sub-images; i, j, and k represent index values.
[0107] Furthermore, the calculation formula of the fitness function is:
[0108]
[0109] where Q is the number of sub-images, M b is the average absolute error of feature extraction corresponding to each sub-image b; A b represents the area of each sub-image, T b represents the calculation time corresponding to this sub-image; G is the number of all images, S c is the structural similarity index of image c; R is the number of patches adjusted during training, Δe p is the change range of the patch size; q is the index value; w1, w2, and w3 are weight coefficients.
[0110] In the embodiments of the present invention, the first term in the loss function focuses on the pixel-level difference between the original image and the reconstructed image. By minimizing this difference, the accuracy of image reconstruction can be significantly improved, making the reconstructed image closer to the original image, thereby maintaining the quality and details of the image. The second and third terms of the loss function respectively consider the differences between the grid size and the sub-image size and the ideal sizes. This design helps the model to dynamically adjust the sizes of the grid and the sub-images during the training process, while keeping them within a reasonable range to prevent the grid and sub-images from being too large or too small, which may affect the performance and stability of the model. The fitness function comprehensively considers multiple factors, including feature extraction error, sub-image processing time, image structural similarity, and the variation range of the patch size. This multi-objective optimization design helps the model to find a balance point during the training process, enabling the model to achieve better performance in all aspects. By dynamically adjusting the sizes of the grid and the sub-images and combining the optimization of the fitness function and the loss function, the model can better adapt to image inputs of different sizes and resolutions. This flexibility and adaptability endow the model with stronger generalization ability in practical applications. The area and processing time of the sub-images are considered in the fitness function, which helps the model to reasonably allocate computing resources during the training process and improve the training efficiency. At the same time, by optimizing the variation range of the patch size, unnecessary computational overhead can be reduced, further enhancing the performance of the model.
[0111] An encoder architecture optimization system integrating multi-modal features, comprising:
[0112] A design module for designing a corresponding trainable encoder for each modality, where the modalities include video, image, audio, depth information, and infrared rays, to achieve unified representation of multi-modal data in a shared embedding space; adopting a mixture-of-experts model structure as the core of each encoder, with multiple expert networks built-in, and dynamically selecting the corresponding expert for processing according to the characteristics of the input data; using contrastive learning technology to align the relationship between any modality and the language modality;
[0113] A processing module for synthesizing multiple small-resolution sub-images into a high-resolution image of a fixed size in the form of an N-grid during the image encoding process, with each sub-image corresponding to an independent text description, and calculating the attention of the patch tokens within the same group;
[0114] A training module for dynamically adjusting the number of grids during the training process and allowing the sizes of the sub-images within the same synthesized image to be not exactly the same, so as to improve the model's processing ability for patch tokens of different granularities;
[0115] The prediction module is used in the inference stage. It synthesizes the subgraphs corresponding to multiple task requests into a large image and inputs it into the model. Through the grouped attention mechanism, it generates CLS tokens for each group of subgraphs, calculates the cosine similarity with the vectors output by the text encoder, and takes the average of the losses of all groups as the final result.
[0116] In the embodiments of the present invention, aiming at the limitations of the traditional dense encoder architecture in cross-modal information fusion and representation learning, the present invention proposes an optimization method for a multi-modal encoder architecture based on the Mixture of Experts (MoE). This method integrates multi-modal features (including videos, images, audio, depth information, and infrared rays) to achieve refined representation in a unified embedding space, thereby improving the cross-modal information fusion efficiency and the model's representation ability.
[0117] The mixture of experts model is used as the core of the encoder, combined with the sparse activation mechanism and the dynamic expert selection strategy, enabling the model to flexibly select the most suitable expert network for processing according to the characteristics of different modal data. This design not only enhances the flexibility and self-adaptability of the model but also improves the efficiency and accuracy in processing multi-modal data.
[0118] In addition, to overcome the limitations of traditional multi-modal models in processing fixed-size images and videos, the present invention further introduces a text alignment training method based on synthetic images. This method allows multiple small-resolution subgraphs to be synthesized into a fixed-size high-resolution image and arranged in an N-grid format, with each subgraph corresponding to an independent text description. Through the grouped attention mechanism, the model can more finely capture the local features in the image and effectively align them with the text description.
[0119] During the training process, the number of grids is dynamically adjusted, and the sizes of the subgraphs within the same synthetic image are allowed to be not exactly the same. This dynamic training strategy enhances the model's ability to process patch tokens of different granularities and sizes, improving the model's adaptability and generalization ability.
[0120] Based on the above innovations, the present invention constructs a multi-modal mixture of experts model encoder architecture called MultiModalBind. This architecture significantly improves the performance and applicability of the multi-modal model in the following aspects:
[0121] Through the mixture-of-experts architecture and sparse activation mechanism, the efficient fusion and refined representation of multimodal data in a unified embedding space are achieved, enhancing the model's representation ability. The text alignment training method based on synthetic images optimizes the model's processing efficiency for multi-task requests, avoids interference between subtasks through the grouped attention mechanism, and improves the model's performance in multi-task scenarios. The sparse activation mechanism and the design of fixed-size synthetic images reduce the computational overhead and video memory occupancy during model training and inference, improving the utilization rate of computing resources and inference efficiency.
[0122] The present invention adopts a combination of a mixture-of-experts (MoE) encoder architecture and a text alignment training method based on synthetic images to further optimize the architecture design and training strategy of the multimodal model. Specifically, it includes the following steps:
[0123] 1) Design and apply a mixture-of-experts (MoE) encoder architecture, design corresponding trainable encoders for each modality (including video, image, audio, depth information, and infrared), with the MoE architecture as the core. Such a design enables the data of different modalities to be uniformly represented in a shared embedding space, improving the accuracy and integrity of feature extraction. By dynamically selecting the most suitable expert network to process different modality data, the cross-modal information fusion efficiency and representation learning ability are enhanced.
[0124] 2) When processing image data, introduce the grouped attention mechanism. Synthesize multiple small-resolution sub-images into a high-resolution image of a fixed size and arrange them in the form of an N-grid, with each sub-image corresponding to an independent text description. During the calculation process, only the patch tokens within the same group perform attention calculation to avoid information interference between different grid sub-images, enhance the model's ability to capture local features, and at the same time maintain the integrity of global information.
[0125] 3) Implement text alignment training based on synthetic images. During the training process, synthesize multiple small-resolution sub-images into a large image, with each sub-image corresponding to an independent text description. By dynamically adjusting the number of grids and allowing the sizes of sub-images within the same synthetic image to be not exactly the same, the model's processing ability for patch tokens of different granularities is improved. This method effectively enhances the adaptability and generalization ability of the model.
[0126] 4) In the training stage, use contrastive learning technology to align the relationship between any modality and the language modality. Each group of sub-images in the synthetic image generates a CLS token through the grouped attention mechanism, and performs cosine similarity calculation with the output text vector of the text encoder. Finally, the average of the losses of all groups is taken as the final result to measure the similarity between the image and the text and provide more accurate inference results.
[0127] 5) Implement a sparse activation mechanism where each expert network in the MoE architecture is activated only when needed. This mechanism enables only a portion of the experts to be activated during each forward pass, thereby efficiently processing large-scale multimodal datasets without significantly increasing computational resources and reducing unnecessary computational and memory overhead.
[0128] The advantages of the present invention are that by integrating the mixture of experts (MoE) model encoder architecture for multimodal features and the text alignment training method based on synthetic images, it not only improves the cross-modal information fusion efficiency and representation learning ability of the multimodal model, but also optimizes the multitasking processing ability and significantly reduces the computational resource overhead. In addition, this method flexibly supports the processing of different granularity patch tokens, further enhancing the adaptability and generalization ability of the model, and promoting the development of the multimodal field towards a more efficient and intelligent direction.
[0129] The mixture of experts (MoE) model encoder architecture of this embodiment focuses on innovatively improving the structure of the traditional dense encoder. It adopts the MoE (Mixture of Experts) architecture and combines a new text alignment training method based on synthetic images to address the deficiencies of existing multimodal models in cross-modal information fusion, representation learning ability, single-size input, computational resource utilization, and model generalization. Through these two improvements, the encoder architecture of this embodiment not only solves the deficiencies of existing methods but also provides a new idea for the performance improvement and efficient training of multimodal large models. For the convenience of subsequent representation, we use images to represent other modal inputs.
[0130] The above solution of the present invention is specifically applied as follows in the embodiments:
[0131] S0, Scale multiple images by randomly selecting a template from a predefined template list to synthesize a fixed-size input image.
[0132] S1, For the input text and any other modality (image, video, audio), the text is first converted into a sequence in digital form through a tokenizer, while the input of other modalities is converted into a vector form through vector encoding.
[0133] S2, Obtain a high-dimensional vector by embedding the digital sequence;
[0134] S3, Perform complex inference operations on the vectors of these different modalities through their respective encoders to obtain their respective vector representations. The specific process of S3 is as follows:
[0135] S3.1, The vector sequence obtained by converting the text sequence through tokenization and then embedding is:
[0136] T = [t1, t2, …, t n , where T = [t1, t2, …, t n represents a sequence of text inputs, a text vector of length n, and t i is the vector representation of the i-th word in the text sequence. Each represents the vector representation of the text vocabulary in a d t -dimensional space; d t is the dimension of the text vector, which is usually defined by the word embedding layer.
[0137] The feature vector matrix obtained after preprocessing and vector encoding of the input image is:
[0138] I = [i1, i2, …, i m , where I = [i1, i2, …, i m represents the processed image input, a sequence of image feature vectors of length m, and i j is the j-th image feature vector representation in the image sequence. Each represents the vector representation of the image feature in a d i -dimensional space; d i is the dimension of the image vector, which is determined by the encoding method of the image and the design of the network.
[0139] S3.2. Feed the text vector sequence into the multi-head attention mechanism (Multi-Head Attention, MHA) for feature extraction. The calculation process is as follows:
[0140]
[0141] where:
[0142]
[0143] Slice QT, KT, and VT according to the head dimension to obtain h groups, and then calculate the attention scores for each group. Among them, represents the weight matrix used to calculate the text query (query), with a shape of d t × d q , where: d t is the dimension of the text vector (i.e., the dimension of the text embedding); d q is the dimension of the query vector, usually achieved by mapping the text embedding to a lower dimension; represents the weight matrix used to calculate the text key (key), with a shape of d t × d k, where d k is the dimension of the key vector, which is the same as the dimension d q of the query vector; represents the weight matrix used to calculate the text value (value), with a shape of d t ×d v , where: d v is the dimension of the value vector, usually the same as the query vector d q and the key vector d k . represents that the dimensions of the query, key, and value are usually h t times the text embedding dimension d , where h is the number of attention heads (which can also be understood as the number of heads in the multi-head attention mechanism).
[0144]
[0145] Among them, represents the text query vector (Query) of the j-th head, with a dimension of d q , which is obtained by passing the text input through the weight matrix .
[0146] represents the text key vector (Key) of the j-th head, with a dimension of d k , which is obtained by passing the text input through the weight matrix . represents the text value vector (Value) of the j-th head, with a dimension of d v , which is obtained by passing the text input through the weight matrix ; represents the normalization factor, which is added to avoid too large or too small numerical values during the inner product calculation. softmax represents applying the softmax function to normalize the dot product result of each query vector and key vector to obtain the attention weights. represents the attention output of the j-th head, with a dimension of d v , which is the weighted sum of the query and the key, where the weighting coefficient is obtained by softmax normalization
[0147] S3.3. Concatenate the attention results of each head and project them back to the original dimension through a linear layer:
[0148]
[0149] The calculation of the image vector sequence is similar to that of the text.
[0150] S3.4, The image vector sequence is fed into the multi-head attention mechanism (MHA) for feature extraction, and the calculation process is as follows:
[0151]
[0152] Among them:
[0153]
[0154] QI, KI, and VI are sliced by group, where g is the number of groups:
[0155]
[0156] Among them, concat represents the concatenation operation, which concatenates the attention results of each head along the dimension,
[0157] to obtain an output containing h heads.
[0158] represents the projection matrix, with a shape of d t ×d t , and through it, a linear transformation is performed on the concatenated output to restore it to the original text vector dimension d t .
[0159] Z T represents the final text vector obtained through concatenation and linear projection, with a dimension of d t . I represents the image vector sequence, with a shape of m×d i , where m is the number of image features and d i is the dimension of the image vector.
[0160] are the weight matrices of the query, key, and value of the image respectively, and their dimensions are d i ×d q , d i ×d k and d i ×d v .
[0161] Q I 、K I 、V I are the vector representations of the query, key, and value of the image respectively, and their dimensions are m×d q , m×d k and m×d v .
[0162] Split QI_i, KI_i, and VI_i into h groups according to the head dimension, and then calculate the attention scores for each group:
[0163]
[0164] S3.5, Concatenate the attention results of each head within the group first, then concatenate them by group, and project them back to the original dimension through a linear layer:
[0165] Among them, represents the concatenation of the attention results of the i-th group of images; Z I represents the further concatenation of the concatenation results of all groups; represents the linear transformation matrix of the image, which is used to restore the concatenated output to the original image vector dimension d i .
[0166] S3.6, Introduce the MOE layer after the multi-head attention mechanism. There are E expert networks, denoted as Expert_k(·) (k = 1, …, E), and the structure of each expert network is a small feed-forward network. First, use the gating network to determine the weights assigned to each expert network for each input token. The gating network is a linear layer followed by a SoftMax function:
[0167]
[0168] For each Expert_k(·) (k = 1, …, E), each token obtains the output from each expert network according to the assigned weights and sums them up with weights to obtain the output of the MOE layer:
[0169] Among them,
[0170] Expert k (Z I ) represents the output of the k-th expert network processing the image vector Z I after processing, which is usually the output of a feed-forward network. represents the weight coefficient given by the gating network, indicating the importance of each token on the k-th expert network; T
[0171] expert network; I represents the output of the MOE layer obtained by weighted summation.
[0172] Finally, perform residual connection and layer normalization to obtain the vector representation of the text after passing through the encoder:
[0173] I * = LayerNorm(Z I+Y I );
[0174] S4. Extract the vectors representing the overall semantic information of the text and other modalities and perform further processing to use them as the information embeddings of this modality. During the training phase, the cosine distance will be calculated between the two, and the encoders other than the text encoder will be updated through the contrastive loss and the InfoNCE loss to achieve the alignment between different modalities and the text, and then the alignment between each modality. The specific process of S4 is as follows:
[0175] S4.1. Extract the overall semantic information vectors (information embeddings) of each modality. After the embedding and a series of encoding operations in the previous steps, we extract the vector of the first token at the beginning from the output of the encoder to represent the overall semantic information. Denote it as:
[0176] Z text , Z image , Z video …;
[0177] where, Z text , Z image and Z video respectively represent the vectors of the overall semantic information of different modalities (text, image, video, etc.), which are extracted from the outputs of their respective encoders. Usually, the output of the first token (such as the [CLS] token) is taken as the representation of the overall semantics.
[0178] S4.2. Calculate the Cosine distance. After obtaining the overall semantic information vectors of each modality (image, video, audio, etc.) and the overall semantic vector of the text modality, we calculate the Cosine distance between them to measure the semantic similarity. Taking the image modality and the text modality as an example, the Cosine distance calculation formula between them is as follows:
[0179]
[0180] where, d cos (Z text , Z image ) represents the cosine distance between the text vector and the image vector. The cosine distance is an index to measure the angle size between two vectors and is used to evaluate their similarity; Z text and Z image respectively represent the overall semantic information vectors of the text modality and the image modality.
[0181] S4.3, Calculation of Contrastive Loss and InfoNCE Loss and Update of Encoder. Contrastive Loss: The contrastive loss aims to reduce the distance between positive sample pairs (vector pairs of the same semantics in different modalities, such as corresponding text and image vector pairs), and increase the distance between negative sample pairs (vector pairs of different semantics in different modalities). The form of the contrastive loss function (taking image and text modalities as an example) is as follows:
[0182]
[0183] InfoNCE Loss: The InfoNCE loss is based on the idea of Noise Contrastive Estimation. For a given positive sample (such as a text-image positive sample pair) and a set of negative samples (other mismatched text-image pairs), it measures the distinguishability of the positive sample relative to the negative samples. The formula (taking image and text modalities as an example) is as follows:
[0184]
[0185] L loss = α·L InfoNCE +(1 - α)·L contrastive ;
[0186] During the training process, by calculating these losses and performing a linear combination, and then updating the parameters of the encoder except the text encoder based on gradient backpropagation, different modalities and text can be better aligned, and thus alignment between all modalities can be achieved. Among them, L contrastive represents the contrastive loss function, which is used to reduce the distance between positive sample pairs (vector pairs of the same semantics in different modalities) and increase the distance between negative sample pairs (vector pairs of different semantics in different modalities). represents the square of the cosine distance of positive sample pairs; m is a preset threshold used to define the minimum distance that should be maintained between negative sample pairs; y is an indicator variable used to identify whether the sample pair is a positive sample pair (y = 1) or a negative sample pair (y = 0); N represents the number of samples; L InfoNCE represents the InfoNCE loss function, which is based on the idea of noise contrastive estimation and is used to measure the distinguishability of positive samples relative to negative samples; represents the similarity between positive sample pairs, usually calculated using cosine similarity; τ is a temperature parameter used to control the smoothness of the softmax function and affect the training dynamics of the model; α is a weight parameter used to balance the contributions of contrastive loss and InfoNCE loss to the total loss; L lossrepresents the total loss function, which is a weighted sum of the contrastive loss and the InfoNCE loss. N is the number of samples, indicating how many pairs of text and image samples; y represents the label, which is 1 for positive sample pairs (text and image with the same semantics) and 0 for negative sample pairs (text and image with different semantics); d os is the cosine similarity, representing the similarity between two vectors. respectively represent the embedding vectors of text and image. m: margin respectively represents the minimum distance between negative sample pairs; sim represents the similarity metric, usually the cosine similarity; K represents the number of negative samples, indicating the number of selected negative samples.
[0187] Through the text alignment training method based on the Mixture of Experts (MoE) model encoder architecture and synthetic images proposed by the present invention, significant improvements in experimental results have been successfully achieved in the multi-modal field. Compared with the current mainstream model LanguageBind, the model of the present invention has achieved leading performance in multiple tasks, especially in cross-modal information alignment, complex modal representation learning, and resource utilization efficiency, showing stronger advantages. This achievement verifies the technical breakthrough of the present invention in multi-modal model architecture design and training strategy optimization, providing important support for the development of the multi-modal artificial intelligence field.
[0188] An embodiment of the present invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to execute the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0189] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An optimization method for an encoder architecture that fuses multi-modal features, characterized in that The method includes: Design corresponding trainable encoders for each modality, where the modalities include video, image, audio, depth information, and infrared, to achieve unified representation of multi-modal data in a shared embedding space; Adopt a mixture-of-experts model structure as the core of each encoder, with multiple expert networks built-in, and dynamically select the corresponding expert for processing according to the characteristics of the input data; Use contrastive learning techniques to align the relationships between any modality and the language modality; During the image encoding process, through a grouped attention mechanism, multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, each sub-image corresponds to an independent text description, and the attention of patch tokens is calculated within the same group; Dynamically adjust the number of grids during the training process, and allow the sizes of sub-images within the same synthesized image to be not exactly the same, to improve the model's processing ability for patch tokens of different granularities; In the inference stage, synthesize the sub-images corresponding to multiple task requests into a large image and input it into the model. Through the grouped attention mechanism, generate CLS tokens for each group of sub-images, calculate the cosine similarity with the vector output by the text encoder, and take the average of the losses of all groups as the final result.
2. The method for optimizing an encoder architecture that fuses multi-modal features according to claim 1, characterized in that, Using contrastive learning techniques to align the relationships between any modality and the language modality includes: Collect a dataset containing multi-modal data, including images, audio, and corresponding texts, sentences; For the image modality, use a convolutional neural network to extract image features; for the audio modality, use an audio processing model to extract audio features; for the language modality, use a word embedding model to extract text features; Define positive sample pairs and negative sample pairs. Positive sample pairs refer to the pairings between different modalities from the same data source, and negative sample pairs are modality pairings from different data sources; Construct a contrastive loss function, and use a dataset containing positive sample pairs and negative sample pairs to train a multi-modal alignment model. During the training process, update the model's parameters by optimizing the contrastive loss function; Use the validation set to evaluate the performance of the model to obtain the evaluation results, and tune the model according to the evaluation results.
3. The method for optimizing an encoder architecture that fuses multi-modal features according to claim 2, characterized in that During the image encoding process, through a grouped attention mechanism, multiple small-resolution sub-images are synthesized into a high-resolution image of a fixed size, arranged in the form of an N-grid, each sub-image corresponds to an independent text description, and the attention of patch tokens is calculated within the same group, including: Generate a set of small-resolution sub-images so that each sub-image has a corresponding text description; adjust all sub-images to the same size; Use a convolutional neural network to extract features from each sub-image, generating a series of feature vectors; Divide the feature map of each sub-image into multiple small blocks, and each small block is converted into a patch token; Use a natural language processing model to encode the text description corresponding to each sub-image, generating text feature vectors; Group the patch tokens of all sub-images according to their corresponding text descriptions, and each group contains all the patch tokens of the sub-images from the same text description; Within each group, the self-attention mechanism is used to calculate the attention weights between patch tokens.
4. The method for optimizing the encoder architecture by fusing multimodal features according to claim 3, wherein Based on the evaluation results, during the training process, the number of grid cells is dynamically adjusted, and the sizes of sub-images within the same synthetic image are allowed to be not exactly the same, so as to improve the model's processing ability for patch tokens of different granularities, including: Initialize a model that includes an image encoder and a decoder. The encoder is used to generate patch tokens, and the decoder is used to reconstruct the image from the patch tokens. Define a strategy to dynamically adjust the number of grid cells, allowing the sizes of sub-images within each grid cell to be not exactly the same. Set a size range, and randomly select and adjust the size of each sub-image during each training. According to the current number of grid cells and the size of sub-images, dynamically divide each sub-image to generate patch tokens of corresponding sizes. Use the image encoder to encode each patch token and extract features. Design a loss function that reflects the changes in dynamic grid cells and sub-image sizes. During the training process, according to the dynamically adjusted number of grid cells and sub-image sizes, adjust the learning rate and batch size parameters accordingly.
5. The method for optimizing an encoder architecture that fuses multi-modal features according to claim 4, characterized in that According to the current number of grid cells and the size of sub-images, dynamically divide each sub-image to generate patch tokens of corresponding sizes, including: According to the task requirements or user input, determine the current number of grid cells and the size of sub-images within each grid cell; set the initial population size, crossover rate, mutation rate, and number of iterations of the genetic algorithm. Use binary encoding to represent the division method of each sub-image, where each bit represents whether a region of the sub-image is divided into a patch token; according to the encoding scheme, randomly generate the initial population, and each individual represents a possible division method of the sub-image. Define a fitness function to evaluate the pros and cons of each division method; calculate the fitness value of each individual in the initial population. According to the fitness values, select the corresponding individuals in the population to enter the next generation; randomly pair the selected individuals and perform crossover operations according to the set crossover rate to generate new individuals; perform mutation operations on the newly generated individuals according to the set mutation rate, and repeat the fitness calculation, selection, crossover, and mutation operations until the set number of iterations is reached. During the iteration process, record the corresponding individuals as the final solution. Decode the final solution obtained by the genetic algorithm into a specific sub-image division method. According to the decoded division method, divide each sub-image to generate patch tokens of corresponding sizes.
6. The method for optimizing an encoder architecture integrating multi-modal features according to claim 5, wherein In the designed loss function that reflects the changes in dynamic grid cells and sub-image sizes, the calculation formula of the loss function L is: where x i is the i-th pixel in the original image; is the i-th pixel after reconstruction, N is the total number of pixels in the image; G j represents the size of the j-th dynamically adjusted grid, G ef is the reference grid size, M is the number of grids; S k is the size of the k-th sub-image, S i is the ideal sub-image size, P is the number of sub-images; i, j, and k represent index values.
7. The method for optimizing an encoder architecture by fusing multimodal features according to claim 6, wherein The calculation formula of the fitness function is: Among them, Q is the number of sub - graphs, M b is the mean absolute error of feature extraction corresponding to each sub - graph b; A b represents the area of each sub - graph, T b represents the calculation time corresponding to this sub - graph; G is the number of all images, S c is the structural similarity index of image c; R is the number of patches adjusted during the training process, Δe p is the change range of the patch size; q is the index value; w1, w2 and w3 are weight coefficients.
8. An encoder architecture optimization system that integrates multi-modal features, the system implements the method described in any one of claims 1 to 7, characterized in that, Including: A design module for designing corresponding trainable encoders for each modality, where the modalities include video, image, audio, depth information, and infrared rays, so as to achieve the unified representation of multi-modal data in the shared embedding space; adopt the mixture-of-experts model structure as the core of each encoder, with multiple expert networks built-in, and dynamically select the corresponding expert for processing according to the characteristics of the input data. Use contrastive learning techniques to align the relationships between any modality and the language modality. A processing module, which is used to synthesize multiple small-resolution sub-images into a high-resolution image with a fixed size in the form of an N-grid during image encoding, where each sub-image corresponds to an independent text description, and calculate the attention of patch tokens within the same group through a grouped attention mechanism; A training module, which is used to dynamically adjust the number of grids during the training process and allow the sizes of sub-images within the same synthesized image to be not exactly the same, so as to improve the model's processing ability for patch tokens of different granularities; A prediction module, which is used in the inference stage to synthesize sub-images corresponding to multiple task requests into a large image and input it into the model. Through the grouped attention mechanism, generate CLS tokens for each group of sub-images, calculate the cosine similarity with the vector output by the text encoder, and take the average of the losses of all groups as the final result.
9. A computing device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Hybrid expert visual question-answering method and system based on strong visual semantics
CN118070816A
Multi-modal model and method for fusing characters, images and audios
CN118861988A