Multi-modal general-purpose model collaborative reasoning method based on dynamic routing mechanism
The collaborative reasoning method of multimodal communication models with a dynamic routing mechanism solves the problems of idle resources and insufficient coordination in multimodal data processing, realizes efficient and flexible multimodal data processing and optimal utilization of computing resources, and improves the semantic fusion of multimodal data and task execution efficiency.
Patent Information
- Application Number
- CN202511094204.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing technologies have problems when processing multimodal data, such as idle or overloaded resources, lack of coordination between different models, low information utilization, and failure to maximize the use of computing resources, resulting in low efficiency in multimodal data processing.
A multimodal general-purpose model collaborative reasoning method based on a dynamic routing mechanism is adopted. By receiving task instructions and multimodal data, the task type is determined, a preset processing mapping relationship is constructed, and the task-specific model, modal processing flow and fusion feature weights are dynamically adjusted. The execution performance indicators are monitored in real time, resource allocation is optimized, and efficient fusion of multimodal data and maximum utilization of computing resources are achieved.
It significantly improves the convenience of semantic fusion of multimodal data and the utilization of computing resources, improves the reasoning efficiency and accuracy of different task classifications, and enhances the adaptability and robustness of multimodal collaborative reasoning.
Smart Images

Figure CN120597213A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a collaborative reasoning method for a multimodal communication model based on a dynamic routing mechanism. Background Art
[0002] In the process of processing multimodal data, cross-modal feature fusion can be performed through a static model architecture. Specifically, different modal data can be processed through a network structure with fixed parameters. However, when using a static model architecture to perform cross-modal feature fusion, resources are prone to idleness or overload when processing heterogeneous multimodal data, affecting overall efficiency.
[0003] In the existing technology, different modal data can be independently processed by building task-specific models, feature weights can be dynamically adjusted through the attention mechanism, and computing resources can be allocated by introducing a resource monitoring module. However, in the method of independently processing different modal data by building task-specific models, the lack of coordination between different models leads to low information utilization. In the method of dynamically adjusting feature weights through the attention mechanism, the optimization dimension is relatively single. In the method of allocating computing resources by introducing a resource monitoring module, there is a lack of dynamic routing strategy and the resource angle granularity is relatively coarse, resulting in difficulty in the three-dimensional dynamic coordination between task requirements, modal features and resource allocation, and the inability to optimize processing parameters in real time during the model inference process, resulting in low efficiency in processing multimodal data. Summary of the Invention
[0004] In response to the problems in the existing technology, this application provides a multimodal communication model collaborative reasoning method based on a dynamic routing mechanism, which can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization rate of computing resources.
[0005] In order to solve at least one of the above problems, the present application provides the following technical solutions: In a first aspect, the present application provides a multimodal communication model collaborative reasoning method based on a dynamic routing mechanism, comprising: Receive task instructions and multimodal data corresponding to the task instructions, determine the task type corresponding to the task instructions, extract data features corresponding to the multimodal data, and fuse the data features based on the task general model to obtain a general fusion feature. Task types include classification tasks, generation tasks, and retrieval tasks; Receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to a current task instruction based on a task type in the preset processing mapping relationship; The data features and general fusion features are sent to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and fusion feature weights to obtain special fusion features, and processes the special fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0006] Furthermore, the method further includes: mapping the data features to a preset semantic space in the task general model to obtain mapped data features, determining similarities between the mapped data features, and assigning general feature weights to the mapped data features based on the similarities, wherein the similarities are positively correlated with the feature weights; Based on the feature weights, data features are fused through a task-general model to obtain general fused features.
[0007] Furthermore, before receiving the preset processing mapping relationship, the method further includes: Constructing a preset processing mapping relationship library, which includes a combination of task-specific models, modal processing processes, and fusion feature weights corresponding to each task category; In the preset processing mapping relationship, the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction are determined based on the task type, including: Based on the task type, the task-specific model, modal processing flow and fusion feature weight that match the current task type are determined from the preset processing mapping relationship library.
[0008] Furthermore, the method further includes: when the task type corresponding to the task instruction is a classification task, processing the dedicated fusion feature through a classification algorithm to obtain a classification label corresponding to the task instruction; When the task type corresponding to the task instruction is a generation task, generating at least one of a text generation result, an image generation result, and an audio generation result corresponding to the task instruction by processing the dedicated fusion feature through the generation model; When the task type corresponding to the task instruction is a retrieval task, the dedicated fusion features are processed by the retrieval algorithm to obtain the retrieval results corresponding to the task instruction.
[0009] Furthermore, after determining the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship, the method further includes: Determining a complexity level of the task instructions, wherein the complexity level is determined based on the data size of the multimodal data and the task type; Determine the computing resources required to calculate multimodal data in the process of completing task instructions based on multimodal data, and determine the resource allocation strategy corresponding to the current task instructions based on the complexity level and computing resources, where computing resources include memory resources, video memory resources, and central processing unit computing power requirements; Send data features and common fusion features to task-specific models through routing nodes, including: Based on the resource allocation strategy, data features and general fusion features are sent to the corresponding task-specific model through routing nodes.
[0010] Furthermore, after sending the data features and the universal fusion features to the task-specific model through the routing node, the method further includes: Real-time monitoring of the execution performance indicators of each task-specific model, including accuracy, recall, execution time, and computing resource consumption; Based on the execution performance indicators, the processing parameters in the modal processing flow in the preset processing mapping relationship are dynamically adjusted.
[0011] Furthermore, it also includes: real-time monitoring of the computing resource consumption of each task-specific model in the process of processing data features and general fusion features, and real-time monitoring of the task progress status corresponding to the task instructions; When it is monitored that a task-specific model has completed the processing of data features and general fusion features and there are remaining computing resources, and other task-specific models have not completed the processing of data features and general fusion features, the remaining computing resources are allocated to other task-specific models.
[0012] In a second aspect, the present application provides a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism, comprising: A receiving module is used to receive task instructions and multimodal data corresponding to the task instructions, determine the task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on the task general model to obtain a general fusion feature. Task types include classification tasks, generation tasks, and retrieval tasks; A mapping module is used to receive a preset processing mapping relationship, and determine the task-specific model, modal processing flow and fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship; The processing module is used to send the data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and the fusion feature weights to obtain the dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0013] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism are implemented.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism.
[0015] In a fifth aspect, the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism.
[0016] It can be seen from the above technical solution that the present application provides a multimodal general-purpose model collaborative reasoning method based on a dynamic routing mechanism. By innovatively receiving task instructions and multimodal data corresponding to the task instructions, the task type corresponding to the task instruction is determined, and the data features of the multimodal data are extracted. The data features are fused through the task general model to generate a general fusion feature. The task-specific model, modal processing flow and fusion feature weights of the current task type are matched according to the preset processing mapping relationship. The data features and the general fusion features are transmitted to the task-specific model through the routing node, so that the task-specific model processes the data features and the general fusion features according to the modal processing flow and fusion feature weights to obtain special fusion features, and processes the special fusion features through the task execution component to obtain the task result corresponding to the task instruction. The task-specific model and fusion feature weight can be dynamically determined according to the task instruction. Multiple tasks can also be processed simultaneously according to the same set of multimodal data, thereby improving the reasoning efficiency and accuracy corresponding to different task classifications. Computing resources are allocated on demand through the routing node, reducing redundant processing, and improving the efficiency and overall performance of multimodal semantic fusion. This method can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1Schematic diagram of a process flow of a multimodal communication model collaborative reasoning method based on a dynamic routing mechanism in an embodiment of the present application; Figure 2 This is a structural diagram of a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism in an embodiment of the present application; Figure 3 Schematic diagram of the structure of the electronic device in the embodiment of the present application.
[0019] Reference numerals: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION
[0020] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.
[0022] In existing technologies, multimodal data of different modalities have their own semantic representations, making it difficult to effectively integrate them and perform collaborative reasoning. Faced with different task requirements and changes in data analysis, existing multimodal models find it difficult to quickly adjust their structures and parameters, resulting in unstable performance in different scenarios. Due to the lack of an effective task allocation mechanism, multimodal models may waste resources on some unnecessary calculations, affecting reasoning efficiency.
[0023] In order to effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the use of computing resources, and significantly improve the convenience of semantic fusion and the utilization of computing resources, this application provides an embodiment of a multimodal communication model collaborative reasoning method based on a dynamic routing mechanism, see Figure 1 The multimodal communication model collaborative reasoning method based on dynamic routing mechanism specifically includes the following contents: Step S101: Receive a task instruction and multimodal data corresponding to the task instruction, determine the task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fusion feature.
[0024] Among them, task types include classification tasks, generation tasks and retrieval tasks.
[0025] Optionally, this embodiment receives a task instruction and multimodal data corresponding to the task instruction, wherein the multimodal data includes data of multiple modal types, including but not limited to at least two of image data, text data, and audio data.
[0026] At the same time, the task type corresponding to the task instruction is determined according to the task instruction, wherein the task type includes but is not limited to a classification task, a generation task, and a retrieval task.
[0027] In addition, data features are extracted from the multimodal data to obtain data features, wherein the data features include at least two of text features, image features, and audio features.
[0028] Among them, different feature extraction methods can be used according to the different data types of multimodal data, and image features of image data can be extracted through convolutional neural networks. Among them, the convolution layer and pooling layer of the convolutional neural network can learn features such as edges, textures, and shapes in image data to generate image features; Extract text data using a recurrent neural network or transformer model to obtain text features. Recurrent neural networks can capture sequence information in text data and obtain text features based on this sequence information. Transformer models can process text in parallel using a self-attention mechanism to generate semantic feature vectors (i.e., text features). The audio data is extracted through an audio feature extraction algorithm to obtain audio features.
[0029] In addition, the data features are fused through the task-general model to obtain a general fusion feature. The task-general model is a model with strong generalization ability, which can learn the common features between different modal data. The methods of obtaining the general fusion feature include but are not limited to splicing fusion and dimensionality reduction fusion. Splicing fusion is to directly splice the image features, text features and audio feature vectors required for fusion into a long vector to form a general fusion feature. Dimensionality reduction splicing is to perform dimensionality reduction processing on each feature data and then splice them to obtain a general fusion feature.
[0030] Furthermore, the task-general model can also map data features of different modalities to a unified semantic space, calculate the semantic association weights between feature vectors to reflect the correlation between multimodal data, and fuse data features according to weighted summation or attention-based mechanisms to obtain universal fusion features.
[0031] This embodiment achieves enhanced accuracy and consistency in semantic fusion of multimodal data, improves the adaptability and flexibility of the task-general model to different task types, and improves efficiency and resource utilization.
[0032] Step S102: receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship.
[0033] Optionally, this embodiment receives a preset processing mapping relationship, wherein the preset processing mapping relationship can be constructed based on a large amount of historical task data and expert experience. Through the preset processing mapping relationship, it is possible to quickly find a suitable task-specific model, modal processing flow and fusion feature weight for the current task instruction.
[0034] Among them, the preset processing mapping relationship can be constructed by collecting a large amount of historical task data of different task types (classification tasks, generation tasks, retrieval tasks), and the historical task data includes but is not limited to task instructions, multimodal data samples and corresponding model configurations.
[0035] Analyze and mine historical task data, that is, find the association pattern between task type and optimal model configuration from historical task data, build mapping rules based on the association pattern, take task type as input, and the corresponding optimal task-specific model, modal processing flow and fusion feature weight as output to form a preset processing mapping relationship.
[0036] This embodiment achieves enhanced dynamic adaptability to different task types, improves the efficiency and accuracy of multimodal collaborative reasoning, and enhances the optimal utilization of resources and the overall performance of task processing.
[0037] Step S103: Send the data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and the fusion feature weights to obtain the dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0038] Optionally, this embodiment sends the extracted data features and universal fusion features to the task-specific model through a routing node, wherein the functions of the routing node include but are not limited to data transmission and route selection, wherein data transmission is used to enable the data features and universal fusion features to be accurately transmitted from the feature extraction end to the task-specific model, and routing selection is used to dynamically select the optimal transmission path according to the current task instructions and multimodal data, thereby improving the efficiency and reliability of data transmission.
[0039] In addition, after receiving the data, the task-specific model further processes the data features and general fusion features according to the received modal processing process and fusion feature weights to obtain dedicated fusion features, wherein the modal processing process may include but is not limited to feature preprocessing and feature weighting steps.
[0040] Among them, feature preprocessing is used to perform operations such as normalization, missing value filling, and outlier processing on each multimodal data feature to ensure the quality and consistency of data features. Feature weighted fusion is used to fuse data features and general fusion features through weighted summation according to the fusion feature weights to generate special fusion features.
[0041] In addition, the obtained dedicated fusion features are input into the task execution component for processing to obtain the task result corresponding to the task instruction, wherein the task execution component can obtain the task result through different processing methods according to different task types.
[0042] This embodiment achieves enhanced adaptability to dynamic changes in multimodal data according to task instructions, improves the accuracy of collaborative reasoning and the robustness of semantic fusion, and enhances the utilization efficiency of computing resources and the execution efficiency of the overall reasoning process.
[0043] This embodiment realizes the ability to efficiently process multimodal data, determine the task-specific model, modal processing flow and fusion feature weight corresponding to the task type, output the task results through routing node transmission, task-specific model processing and task execution component collaboration, meet the multi-field multimodal data processing needs, and improve performance.
[0044] In some embodiments, data features are fused based on a task-general model to obtain general fused features, including: Map the data features to the preset semantic space in the task general model respectively to obtain the mapped data features, determine the similarity between the mapped data features, and assign general feature weights to the mapped data features based on the similarity, where the similarity is positively correlated with the feature weight; Based on the feature weights, data features are fused through a task-general model to obtain general fused features.
[0045] Optionally, this embodiment can map the data features corresponding to the extracted multimodal data to a preset semantic space in the task-general model to obtain mapped data features. The preset semantic space is a high-dimensional, unified semantic representation space that is used to eliminate semantic differences between data of different modalities and enable features from different modalities to be compared and integrated in a common space.
[0046] The mapping method may vary according to different modal data, including: Image feature mapping: Image features are mapped into a preset semantic space through a pre-trained visual embedding model (such as a visual Transformer). The visual embedding model can be pre-trained on a large-scale image dataset to learn the mapping relationship between image features and semantic concepts, i.e., the mapped image features.
[0047] Text feature mapping: The text features of text data can be mapped to a preset semantic space through word embedding models (such as Word2Vec and GloVe) and semantic role labeling. The word embedding model can capture the semantic similarity and contextual relationships between words, and semantic role labeling can extract the core semantic components in the sentence, that is, the mapped text features.
[0048] Audio feature mapping: Audio features can be mapped to a semantic space using audio semantic analysis techniques (such as audio embedding models). The audio embedding model analyzes audio features such as spectrum and timing and converts them into semantically relevant representations, i.e., mapped audio features.
[0049] In addition, in a preset semantic space, the similarity between the mapped data features is calculated to measure the semantic relevance of multimodal data from different modalities. Similarity calculation methods include but are not limited to cosine similarity and Euclidean distance.
[0050] In addition, based on the calculated similarity, common feature weights are assigned to the mapped data features, where similarity is positively correlated with feature weight, that is, the higher the similarity, the greater the feature weight. The feature weight can be assigned using linear normalization to map the similarity value to the weight interval.
[0051] In addition, according to the assigned feature weights, the data features are fused through the task-general model to obtain general fusion features, where the fusion methods include but are not limited to weighted summation and feature splicing.
[0052] Furthermore, adversarial training can be used to narrow the distribution differences of different modal data in the preset semantic space and enhance the consistency and accuracy of feature mapping.
[0053] This embodiment achieves the goal of effectively integrating the semantic information of data of different modalities by mapping multimodal data features to a unified preset semantic space and performing feature weight assignment and fusion based on similarity, thereby improving the performance of various task types such as classification tasks, generation tasks, and retrieval tasks. At the same time, the optimized computing process and resource utilization methods make the processing of various task types more efficient and economical.
[0054] In some embodiments, before receiving the preset processing mapping relationship, the method further includes: Constructing a preset processing mapping relationship library, which includes a combination of task-specific models, modal processing processes, and fusion feature weights corresponding to each task category; In the preset processing mapping relationship, the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction are determined based on the task type, including: Based on the task type, the task-specific model, modal processing flow and fusion feature weight that match the current task type are determined from the preset processing mapping relationship library.
[0055] Optionally, before receiving the preset processing mapping relationship, this embodiment constructs a preset processing mapping relationship library, wherein the preset processing mapping relationship library brings together combinations of task-specific models, modal processing processes and fusion feature weights corresponding to various task categories.
[0056] Collect a large number of multimodal data samples and their corresponding annotations from various domains, including examples from various task types such as classification, generation, and retrieval. Also collect relevant task execution logs and model performance evaluation data to understand how different task-specific models perform across various task types.
[0057] The collected multimodal data samples and annotation information are classified into different task types, and the features of each multimodal data sample are extracted, including but not limited to the task objectives, input modality type, and output requirements. The descriptions of each task are semantically analyzed through natural language processing technology to extract the core semantic features of the task description. At the same time, the input and output examples of the task are analyzed to extract data modality features and task structure features.
[0058] For each task type, train multiple candidate task-specific models. Task-specific models can be based on different algorithms and architectures, such as deep learning models and traditional machine learning models. During the training process, evaluate the performance of different task-specific models through cross-validation, including but not limited to accuracy, recall, F1 score, and other metrics. Also, record the performance of each task-specific model under different modality processing flows and fusion feature weight combinations.
[0059] Based on the characteristics of the task samples and the evaluation results of the model, a mapping relationship is established between the task type and the task-specific model, the modal processing flow and the fusion feature weight. Machine learning methods such as decision trees, support vector machines, neural networks, etc. can be used to learn the mapping method between the characteristics of the task samples and the optimal task-specific model configuration, and the established mapping relationship is stored in a preset processing mapping relationship library.
[0060] In addition, in the preset processing mapping relationship library, a query is performed based on the determined task type to find a combination of multiple task-specific models, modal processing flows, and fusion feature weights corresponding to the current task type.
[0061] In addition, based on the specific circumstances of the current task instructions, such as the characteristics of the input modal data and real-time requirements, the best task-specific model, modal processing flow and fusion feature weights are selected from the multiple combinations queried to determine the task-specific model, modal processing flow and fusion feature weights corresponding to the current task instructions.
[0062] Furthermore, each combination can be evaluated by weighted scoring, and the combination with the highest comprehensive score can be selected. The distribution of fusion feature weights is based on the degree of influence of each multimodal data on the execution effect of the task instruction.
[0063] This embodiment realizes efficient and accurate processing of task instructions by constructing and presetting a processing mapping relationship library, thereby improving task execution efficiency and accuracy.
[0064] In some embodiments, the task execution component processes the dedicated fusion features to obtain a task result corresponding to the task instruction, including: When the task type corresponding to the task instruction is a classification task, the classification algorithm is used to process the dedicated fusion features to obtain the classification label corresponding to the task instruction; When the task type corresponding to the task instruction is a generation task, generating at least one of a text generation result, an image generation result, and an audio generation result corresponding to the task instruction by processing the dedicated fusion feature through the generation model; When the task type corresponding to the task instruction is a retrieval task, the dedicated fusion features are processed by the retrieval algorithm to obtain the retrieval results corresponding to the task instruction.
[0065] Optionally, in this embodiment, when the task type corresponding to the task instruction is a classification task, the most suitable classification algorithm is selected from the preset processing mapping relationship library according to the task instruction, such as support vector machine (SVM), decision tree, random forest, and deep neural network classifier.
[0066] For example, in image classification tasks, if the dedicated fusion features have high-dimensional and complex distributions, they can be classified by a deep neural network classifier to obtain the corresponding classification labels.
[0067] Furthermore, multi-label classification tasks can be processed using multi-label classification algorithms to obtain corresponding classification labels. Multi-label classification algorithms can simultaneously predict the output of multiple labels. For example, in news classification, a piece of news may cover multiple topics (such as politics, economics, and technology). Multi-label classification algorithms can comprehensively consider the multimodal information in dedicated fusion features and assign multiple appropriate topic labels to each news article.
[0068] Furthermore, after generating the classification results, post-processing techniques (such as label smoothing and threshold adjustment) are used to optimize the results and reduce the impact of overfitting and noise. At the same time, verification mechanisms such as cross-validation and comparative testing can be introduced to ensure the accuracy and reliability of the results.
[0069] In addition, for text generation tasks, text generation results can be obtained through generative models such as sequence-to-sequence (Seq2Seq) models, transformer models, and generative adversarial networks (GANs).
[0070] Taking product description generation as an example, dedicated fusion features are input into the generation model, so that the generation model can automatically generate descriptive text based on the text generation pattern and semantic logic learned from the training data. During the generation process, the attention mechanism can be used to focus on key feature information to improve the relevance and coherence of the generated text.
[0071] In image generation tasks, dedicated fusion features can be processed through models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) to obtain image generation results.
[0072] For example, in the process of obtaining the corresponding image generation results based on the text description, the generative model converts the text semantic information in the dedicated fusion features into image features, and gradually constructs the image content that meets the description. It can also use the adversarial training mechanism to continuously optimize the quality and details of the image generation results, making them more realistic and in line with expectations.
[0073] In addition, for audio generation tasks, dedicated fusion features can be processed through models such as WaveNet and Tacotron to obtain audio generation results.
[0074] Taking the speech synthesis task as an example, in the process of generating natural and fluent speech audio based on text content, the generative model combines the text semantics and speech features in the dedicated fusion features to generate the corresponding audio waveform. It can focus on the details such as the intonation, speaking speed, and emotion of the speech, thereby improving the naturalness and expressiveness of the audio generation results.
[0075] In addition, when the task type corresponding to the task instruction is a retrieval task, a suitable retrieval algorithm can be selected according to the task instruction, such as a retrieval method based on Euclidean distance, cosine similarity, Hamming distance, etc.
[0076] Taking the retrieval type of cross-modal image data-text data as an example, the similarity between the dedicated fusion feature and each data item in the database is determined, the retrieval results are sorted according to the similarity, and the data item that best matches the query intent is returned.
[0077] In multimodal retrieval tasks, data features from different modalities are fused and processed to improve the accuracy and comprehensiveness of retrieval through multimodal information. For example, in video retrieval tasks, a comprehensive search combining the visual features, audio features, and text description features of a video can more accurately locate video clips that meet the query requirements.
[0078] Furthermore, the preliminary search results can be post-processed and optimized, such as result re-ranking, deduplication, and result fusion. At the same time, the user feedback mechanism can be combined to improve the relevance and satisfaction of the search results, that is, the parameters and model configuration of the search algorithm can be dynamically adjusted according to the user's behavioral data such as clicks and dwell time.
[0079] This embodiment achieves efficient processing of dedicated fusion features through the task execution component, improves the accuracy and quality of task results, and enhances intelligence and adaptability.
[0080] In some embodiments, after determining the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship, the method further includes: Determining a complexity level of the task instructions, wherein the complexity level is determined based on the data size of the multimodal data and the task type; Determine the computing resources required to calculate multimodal data in the process of completing task instructions based on multimodal data, and determine the resource allocation strategy corresponding to the current task instructions based on the complexity level and computing resources, where computing resources include memory resources, video memory resources, and central processing unit computing power requirements; Send data features and common fusion features to task-specific models through routing nodes, including: Based on the resource allocation strategy, data features and general fusion features are sent to the corresponding task-specific model through routing nodes.
[0081] Optionally, after determining the task-specific model, modal processing flow and fusion feature weights in the preset processing mapping relationship, this embodiment needs to further determine the complexity level of the task instructions, wherein the determination of the complexity level comprehensively considers two factors: the data scale of the multimodal data and the task type.
[0082] Data scale includes the evaluation of at least two multimodal data items: image data, text data, and audio data, and must be consistent with the input multimodal data. Image data can be evaluated by evaluating the quantity, resolution, color depth, and other aspects of the image data. For example, high-resolution image data is larger and more difficult to process than low-resolution image data.
[0083] Text data can take into account the length of the text, vocabulary richness, sentence complexity, etc. Among them, long texts or text data containing a large number of professional terms and complex sentences are larger in scale and more complicated to process.
[0084] Audio data can be analyzed for duration, sampling rate, number of channels, etc. Among them, audio data with longer duration, higher sampling rate, and multiple channels is large in scale, and the processing difficulty increases accordingly.
[0085] The task types include the evaluation of the complexity levels of classification tasks, generation tasks, and retrieval tasks. Among them, simple classification tasks may only involve a few categories and single-modal data, while complex classification tasks may involve multiple categories, multi-modal data, and have problems such as category imbalance.
[0086] In generation tasks, simple generation tasks may only generate short texts or simple graphics, while complex generation tasks may require generating long, high-quality texts, images or audios that are highly relevant to the input data.
[0087] Simple retrieval tasks may require single-modal retrieval in small-scale datasets, while complex retrieval tasks may require cross-modal retrieval in large-scale datasets, which have high requirements for retrieval accuracy and recall rate.
[0088] The above factors can be comprehensively considered to establish a complexity assessment model to quantitatively evaluate the complexity level of task instructions and divide the complexity level into three levels: low complexity level, medium complexity level, and high complexity level.
[0089] For example, a task containing a small number of low-resolution images and simple classification instructions can be evaluated as a low complexity level, while a task containing a large number of high-resolution images, long text descriptions, and complex generation instructions can be evaluated as a high complexity level.
[0090] In addition, the computing resources required to calculate the multimodal data in the process of completing the task instructions based on the multimodal data are determined. The computing resources include but are not limited to memory resources, video memory resources and central processing unit (CPU) computing power requirements.
[0091] Based on the task type and data scale, evaluate the computing resource requirements of each modal data processing link. Among them, image data processing requires more video memory resources for storing and processing high-resolution image feature maps, and also has certain requirements for GPU computing power; text data processing mainly consumes memory resources and CPU computing power for storing text sequences and performing natural language processing-related calculations; audio data processing has certain requirements for memory resources and CPU computing power, especially for the processing and analysis of long audio.
[0092] In addition, based on the complexity level and computing resource requirements, the resource allocation strategy corresponding to the current task instruction is determined. For task instructions with a high complexity level, more computing resources can be allocated to meet their processing requirements. For task instructions with a low complexity level, relatively fewer resources can be allocated to achieve rational use of resources.
[0093] For example, for high-complexity image generation tasks, more video memory resources and GPU computing power are allocated to the image generation model to ensure the quality and speed of generated images. For low-complexity text classification tasks, an appropriate amount of memory resources and CPU computing power is allocated to the text classification model to avoid wasting resources.
[0094] In addition, data features and common fusion features are sent to task-specific models through routing nodes, including: The routing node determines the transmission priority and transmission path of data features and general fusion features based on the resource allocation strategy. For modal data processing links with high computing resource requirements, the routing node prioritizes ensuring that the required data can be transmitted to the task-specific model that processes the data features in a timely and efficient manner.
[0095] According to the complexity level of the task instructions and the resource allocation strategy, the bandwidth and rate of data transmission are dynamically adjusted. For high-complexity task instructions, more transmission bandwidth is allocated to speed up data transmission and reduce transmission delay, so that the task-specific model can quickly receive data and start processing.
[0096] Furthermore, during data transmission, routing nodes are also responsible for managing and scheduling data to avoid data congestion and loss.
[0097] For example, when multiple task instructions transmit data simultaneously, the routing node reasonably allocates transmission resources according to the resource allocation strategy and priority of each subtask, ensuring that each multimodal data can smoothly reach the task-specific model of the corresponding subtask.
[0098] Furthermore, we developed resource demand prediction algorithms that can predict changes in computing resource requirements at different stages of a task. Based on these predictions, we can adjust resource allocation strategies in advance and dynamically allocate or release computing resources to tasks, improving resource utilization and task execution efficiency. For example, for a long-running video processing task, if we predict that the video encoding stage will require high CPU and video memory resources, we can allocate sufficient resources in advance to ensure the smooth execution of the task.
[0099] This embodiment achieves this by determining the task-specific model, modal processing flow and fusion feature weight based on the task type in a preset processing mapping relationship, further determining the complexity level of the task instruction, and determining the resource allocation strategy accordingly. Then, according to the strategy, the data features and general fusion features are sent to the task-specific model through the routing node, thereby achieving refined resource management and efficient execution of processing, and improving resource utilization efficiency and overall performance.
[0100] In some embodiments, after sending the data features and the universal fusion features to the task-specific model through the routing node, the method further includes: Real-time monitoring of the execution performance indicators of each task-specific model, including accuracy, recall, execution time, and computing resource consumption; Based on the execution performance indicators, the processing parameters in the modal processing flow in the preset processing mapping relationship are dynamically adjusted.
[0101] Optionally, after sending the data features and general fusion features to the task-specific models through the routing nodes, this embodiment monitors the execution performance indicators of each task-specific model in real time, where the execution performance indicators include accuracy, recall rate, execution time and consumption of computing resources.
[0102] Among them, the accuracy rate is used to measure the degree of match between the task results obtained by the task-specific model and the actual results. For example, in the classification task, the accuracy rate represents the proportion of correctly classified samples to the total number of samples. In the retrieval task, the accuracy rate represents the proportion of relevant results retrieved to the total number of retrieved results.
[0103] The recall rate is used to emphasize the coverage of all true positive examples by the task results obtained by the task-specific model. For example, in the classification task, the recall rate indicates the ratio of the number of correctly classified positive samples to the total number of positive samples. In the retrieval task, the recall rate indicates the ratio of the retrieved relevant results to all relevant results.
[0104] The execution time is used to record the time it takes for the task-specific model to receive data features and general fusion features and obtain task results, reflecting the operating efficiency of the task-specific model.
[0105] The consumption of computing resources, including but not limited to memory usage, video memory usage, and the utilization rate of the Central Processing Unit (CPU) and Graphics Processing Unit (GPU), reflects the resource utilization of the task-specific model in the process of executing task instructions.
[0106] Furthermore, the above performance indicator data can be monitored in real time through special monitoring modules, performance monitoring tools and log records, and stored in a performance monitoring database for subsequent analysis and processing.
[0107] In addition, mathematical models or machine learning algorithms can be used to analyze the relationship between accuracy, recall, execution time, and computing resource consumption and various processing parameters in the modal processing flow. For example, the impact of parameters such as the word vector dimension in text feature extraction and the convolution kernel size and step size in image feature extraction on the accuracy and execution time of classification tasks can be analyzed.
[0108] Determine the performance optimization goal based on the requirements of the task instructions and the current resource status. If the task instructions have high real-time requirements, the goal is to reduce the execution time. If the task instructions focus more on the accuracy of the results, the optimization goal is to improve the accuracy or recall rate.
[0109] At the same time, corresponding adjustment strategies can be formulated, for example, based on optimization algorithms such as gradient descent method and genetic algorithm, to determine the adjustment direction and amplitude of processing parameters.
[0110] Based on the adjustment strategy, processing parameters in the modal processing flow are modified in real time during the execution of the task-specific model. For example, in a classification task, if the current convolution kernel size causes the task-specific model to execute too long but has only a limited improvement in accuracy, the convolution kernel size can be appropriately reduced to improve the efficiency of the task-specific model.
[0111] This embodiment realizes the performance optimization and adaptive adjustment of collaborative reasoning of multimodal general-purpose models by real-time monitoring of the execution performance indicators of each task-specific model and dynamically adjusting the processing parameters of the modal processing flow in the preset processing mapping relationship based on the execution performance indicators, so that the task-general model and the task-specific model can maintain an efficient operating state under different task instructions and different multimodal data conditions, thereby improving the accuracy, recall rate and execution efficiency, and optimizing the utilization of computing resources.
[0112] In some embodiments, dynamically adjusting processing parameters in a modal processing flow in a preset processing mapping relationship based on an execution performance indicator includes: Monitor in real time the computing resource consumption of each task-specific model in the process of processing data features and general fusion features, and monitor in real time the task progress status corresponding to the task instructions; When it is monitored that a task-specific model has completed the processing of data features and general fusion features and there are remaining computing resources, and other task-specific models have not completed the processing of data features and general fusion features, the remaining computing resources are allocated to other task-specific models.
[0113] Optionally, this embodiment monitors in real time the memory and video memory size occupied by the task-specific model during the processing process, wherein performance monitoring tools or specialized hardware monitoring software can be used. For example, when processing image data, the video memory occupancy can be monitored to ensure that the video memory does not overflow.
[0114] Among them, during the monitoring process through performance monitoring tools, the CPU and GPU usage occupied by the task-specific model can be recorded in real time to determine the computing load of the task-specific model in different processing stages. During the data transmission bandwidth monitoring process, the network bandwidth occupied by the task-specific model during the data transmission process can be monitored.
[0115] In addition, based on the processing steps and stages of the task-specific model, the completion percentage of the task instructions in the current stage is determined. For example, for a complex task including multiple processing stages, the overall completion degree is determined by monitoring the completion status of each stage.
[0116] Furthermore, the time from the start of processing of the task-specific model to the current moment can be recorded, and the remaining execution time to complete the current task instruction can be calculated based on the estimated total execution time.
[0117] Furthermore, a status tag (eg, "initializing," "processing," "completed," etc.) may be used to identify the current task progress status of the task-specific model.
[0118] Furthermore, the remaining computational resource requirements for each subtask in each task instruction are determined based on the computational resource consumption of the task-specific model and the task progress status. For example, for a task that has completed half of its processing flow and has high computational resource consumption, its computational resource requirements for subsequent steps are estimated.
[0119] Furthermore, the resource allocation priority of each subtask in the task instruction can be determined based on factors such as the urgency, importance and remaining execution time of the task instruction. For example, urgent and important tasks will be given higher priority and the remaining computing resources will be allocated first.
[0120] In addition, if it is detected that a task-specific model has completed processing of data features and universal fusion features and has remaining computing resources, the remaining computing resources will be allocated to other task-specific models that have not yet completed processing. For example, if an image classification task-specific model has completed processing and has remaining GPU resources, the remaining GPU resources can be allocated to another task-specific model that is performing a complex image generation task.
[0121] Furthermore, resource allocation strategies can be adjusted based on the dynamic needs of each subtask. For example, if a task-specific model runs out of computing resources during processing, resource allocation to other subtasks can be temporarily adjusted to meet the computing resource needs of that subtask.
[0122] This embodiment realizes resource optimization management of collaborative reasoning of multimodal communication models by real-time monitoring of the computing resource consumption and task progress status of each task-specific model and allocating the remaining computing resources to other unfinished task-specific models, thereby improving resource utilization efficiency, optimizing the execution efficiency of task instructions, and enhancing stability and reliability.
[0123] In order to effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the use of computing resources, and significantly improve the convenience of semantic fusion and the utilization of computing resources, the present application provides an embodiment of a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism for realizing all or part of the content of the multimodal communication model collaborative reasoning based on a dynamic routing mechanism, see Figure 2 The multimodal communication model collaborative reasoning device based on a dynamic routing mechanism specifically includes the following contents: A receiving module 10 is configured to receive a task instruction and multimodal data corresponding to the task instruction, determine a task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fused feature. Task types include classification tasks, generation tasks, and retrieval tasks. A mapping module 20 is configured to receive a preset processing mapping relationship and determine a task-specific model, a modal processing flow, and a fusion feature weight corresponding to a current task instruction based on the task type in the preset processing mapping relationship; The processing module 30 is used to send the data features and the general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and the general fusion features based on the modal processing flow and the fusion feature weights to obtain the special fusion features, and processes the special fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0124] From the above description, it can be seen that the embodiment of the present application provides a multimodal general-purpose model collaborative reasoning device based on a dynamic routing mechanism, which can determine the task type corresponding to the task instruction by innovatively receiving task instructions and multimodal data corresponding to the task instructions, and extract the data features of the multimodal data, and generate a universal fusion feature by fusing the data features through the task universal model. According to the preset processing mapping relationship, the task-specific model, modal processing flow and fusion feature weight of the current task type are matched, and the data features and universal fusion features are transmitted to the task-specific model through the routing node, so that the task-specific model processes the data features and universal fusion features according to the modal processing flow and fusion feature weight to obtain a dedicated fusion feature, and processes the dedicated fusion feature through the task execution component to obtain the task result corresponding to the task instruction. It can dynamically determine the task-specific model and fusion feature weight according to the task instruction, and can also process multiple tasks simultaneously according to the same set of multimodal data, thereby improving the reasoning efficiency and accuracy corresponding to different task classifications, allocating computing resources on demand through the routing node, reducing redundant processing, and improving the efficiency and overall performance of multimodal semantic fusion. This method can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization of computing resources.
[0125] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize computing resources, and significantly improve the convenience of semantic fusion and the utilization of computing resources, the present application provides an embodiment of an electronic device for implementing all or part of the content of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism, and the electronic device specifically includes the following content: Processor (processor), memory (memory), communication interface (Communications Interface) and bus; wherein, the processor, memory, and communication interface complete mutual communication through the bus; the communication interface is used to realize information transmission between a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism and related equipment such as a core business system, a user terminal, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited to this. In this embodiment, the logic controller can be implemented with reference to the embodiment of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism in the embodiment, and the embodiment of a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism, the contents of which are merged here, and the repeated parts are not repeated.
[0126] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0127] In practical applications, part of the multimodal communication model collaborative reasoning method based on the dynamic routing mechanism can be executed on the electronic device side as described above, or all operations can be completed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed on the client device, the client device may also include a processor.
[0128] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.
[0129] Figure 3 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0130] In one embodiment, the multimodal communication model collaborative reasoning method based on the dynamic routing mechanism can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control: Step S101: Receive a task instruction and multimodal data corresponding to the task instruction, determine the task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fused feature. Task types include classification tasks, generation tasks, and retrieval tasks. Step S102: receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship; Step S103: Send the data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and the fusion feature weights to obtain the dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0131] From the above description, it can be seen that the electronic device provided in the embodiment of the present application innovatively receives task instructions and multimodal data corresponding to the task instructions, determines the task type corresponding to the task instructions, and extracts data features of the multimodal data, fuses data features through a task general model to generate general fusion features, matches the task-specific model, modal processing flow and fusion feature weights of the current task type according to a preset processing mapping relationship, transmits data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features according to the modal processing flow and fusion feature weights to obtain special fusion features, and processes the special fusion features through the task execution component to obtain the task result corresponding to the task instruction, can dynamically determine the task-specific model and fusion feature weight according to the task instruction, and can also process multiple tasks simultaneously according to the same set of multimodal data, thereby improving the reasoning efficiency and accuracy corresponding to different task classifications, allocating computing resources on demand through the routing node, reducing redundant processing, and improving the efficiency and overall performance of multimodal semantic fusion. This method can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization of computing resources.
[0132] In another embodiment, a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism can be configured separately from the central processing unit 9100. For example, a multimodal communication model collaborative reasoning device based on a dynamic routing mechanism can be configured as a chip connected to the central processing unit 9100, and the function of the multimodal communication model collaborative reasoning method based on the dynamic routing mechanism can be realized through the control of the central processing unit.
[0133] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 3 In addition, the electronic device 9600 may also include all components shown in Figure 3 For components not shown, reference may be made to the prior art.
[0134] like Figure 3As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.
[0135] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.
[0136] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.
[0137] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), or SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is capable of storing additional data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs, or processes used by the central processing unit 9100 to execute operations of the electronic device 9600.
[0138] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, images, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0139] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.
[0140] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless local area network modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130, providing audio output via the speaker 9131 and receiving audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.
[0141] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism in the above-mentioned embodiments, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, all steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism in the above-mentioned embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented: Step S101: Receive a task instruction and multimodal data corresponding to the task instruction, determine the task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fused feature. Task types include classification tasks, generation tasks, and retrieval tasks. Step S102: receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship; Step S103: Send the data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and the fusion feature weights to obtain the dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0142] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application innovatively receives task instructions and multimodal data corresponding to the task instructions, determines the task type corresponding to the task instructions, and extracts data features of the multimodal data, fuses data features through a task general model to generate general fusion features, matches the task-specific model, modal processing flow and fusion feature weights of the current task type according to a preset processing mapping relationship, transmits data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features according to the modal processing flow and fusion feature weights to obtain special fusion features, and processes the special fusion features through the task execution component to obtain the task result corresponding to the task instruction, can dynamically determine the task-specific model and fusion feature weight according to the task instruction, and can also process multiple tasks simultaneously according to the same set of multimodal data, thereby improving the reasoning efficiency and accuracy corresponding to different task classifications, allocating computing resources on demand through routing nodes, reducing redundant processing, and improving the efficiency and overall performance of multimodal semantic fusion. This method can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization of computing resources.
[0143] The embodiments of the present application also provide a computer program product capable of implementing all steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instructions are executed by a processor, the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism are implemented. For example, the computer program / instructions implement the following steps: Step S101: Receive a task instruction and multimodal data corresponding to the task instruction, determine the task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fused feature. Task types include classification tasks, generation tasks, and retrieval tasks. Step S102: receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship; Step S103: Send the data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features based on the modal processing flow and the fusion feature weights to obtain the dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task results corresponding to the task instructions.
[0144] From the above description, it can be seen that the computer program product provided in the embodiment of the present application innovatively receives task instructions and multimodal data corresponding to the task instructions, determines the task type corresponding to the task instructions, and extracts data features of the multimodal data, fuses data features through a task general model to generate general fusion features, matches the task-specific model, modal processing flow and fusion feature weights of the current task type according to a preset processing mapping relationship, transmits data features and general fusion features to the task-specific model through the routing node, so that the task-specific model processes the data features and general fusion features according to the modal processing flow and fusion feature weights to obtain special fusion features, and processes the special fusion features through the task execution component to obtain the task result corresponding to the task instruction, can dynamically determine the task-specific model and fusion feature weight according to the task instruction, and can also process multiple tasks simultaneously according to the same set of multimodal data, thereby improving the reasoning efficiency and accuracy corresponding to different task classifications, allocating computing resources on demand through routing nodes, reducing redundant processing, and improving the efficiency and overall performance of multimodal semantic fusion. This method can effectively solve the shortcomings of traditional technologies in multimodal data processing, such as difficulty in semantic fusion, insufficient model flexibility, and failure to maximize the utilization of computing resources. It can significantly improve the convenience of semantic fusion and improve the utilization of computing resources.
[0145] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatuses, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0146] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0149] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A multimodal communication model collaborative reasoning method based on a dynamic routing mechanism, characterized by: The method comprises: receiving a task instruction and multimodal data corresponding to the task instruction, determining a task type corresponding to the task instruction, extracting data features corresponding to the multimodal data, and fusing the data features based on a general task model to obtain a general fused feature, wherein the task type includes a classification task, a generation task, and a retrieval task; Receiving a preset processing mapping relationship, and determining a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship; The data features and the universal fusion features are sent to the task-specific model through the routing node, so that the task-specific model processes the data features and the universal fusion features based on the modal processing flow and the fusion feature weights to obtain dedicated fusion features, and processes the dedicated fusion features through the task execution component to obtain the task result corresponding to the task instruction.
2. The method according to claim 1, characterized in that The fusing of the data features based on the task-general model to obtain general fusion features includes: Mapping the data features to a preset semantic space in the task general model respectively to obtain mapped data features, determining similarities between the mapped data features, and assigning general feature weights to the mapped data features based on the similarities, wherein the similarities are positively correlated with the feature weights; Based on the feature weights, the data features are fused through the task general model to obtain a general fused feature.
3. The method according to claim 1, characterized in that Before receiving the preset processing mapping relationship, the method further includes: Constructing a preset processing mapping relationship library, wherein the preset processing mapping relationship library includes a combination of task-specific models, modal processing processes, and fusion feature weights corresponding to each of the task categories; The determining of the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship includes: Based on the task type, a task-specific model, a modal processing flow, and a fusion feature weight that match the current task type are determined from the preset processing mapping relationship library.
4. The method according to claim 1, wherein The processing of the dedicated fusion feature by the task execution component to obtain a task result corresponding to the task instruction includes: When the task type corresponding to the task instruction is the classification task, processing the dedicated fusion feature through a classification algorithm to obtain a classification label corresponding to the task instruction; When the task type corresponding to the task instruction is the generation task, processing the dedicated fusion feature through a generation model to generate at least one of a text generation result, an image generation result, and an audio generation result corresponding to the task instruction; When the task type corresponding to the task instruction is the retrieval task, the dedicated fusion feature is processed by a retrieval algorithm to obtain a retrieval result corresponding to the task instruction.
5. The method according to claim 1, wherein After determining the task-specific model, modal processing flow, and fusion feature weight corresponding to the current task instruction based on the task type in the preset processing mapping relationship, the method further includes: determining a complexity level of the task instruction, wherein the complexity level is determined based on a data size of the multimodal data and a type of the task; Determining computing resources required to compute the multimodal data in a process of completing the task instruction based on the multimodal data, and determining a resource allocation strategy corresponding to the current task instruction based on the complexity level and the computing resources, wherein the computing resources include memory resources, video memory resources, and central processing unit computing power requirements; Sending the data features and the universal fusion features to the task-specific model through a routing node includes: The data features and the universal fusion features are sent to the corresponding task-specific model through a routing node based on the resource allocation strategy.
6. The method according to claim 1, characterized in that After sending the data features and the universal fusion features to the task-specific model through the routing node, the method further includes: Real-time monitoring of the execution performance indicators of each task-specific model, wherein the execution performance indicators include accuracy, recall rate, execution time, and consumption of computing resources; Based on the execution performance indicator, the processing parameters in the modal processing flow in the preset processing mapping relationship are dynamically adjusted.
7. The method according to claim 6, characterized in that The dynamically adjusting the processing parameters in the modal processing flow in the preset processing mapping relationship based on the execution performance indicator includes: monitoring in real time the computing resource consumption of each task-specific model in the process of processing the data features and the universal fusion features, and monitoring in real time the task progress status corresponding to the task instruction; When it is monitored that the task-specific model has completed processing the data features and the universal fusion features and there are remaining computing resources, and other task-specific models have not completed processing the data features and the universal fusion features, the remaining computing resources are allocated to other task-specific models.
8. A multimodal communication model collaborative reasoning device based on a dynamic routing mechanism, characterized in that: The device comprises: a receiving module, configured to receive a task instruction and multimodal data corresponding to the task instruction, determine a task type corresponding to the task instruction, extract data features corresponding to the multimodal data, and fuse the data features based on a general task model to obtain a general fused feature, wherein the task types include classification tasks, generation tasks, and retrieval tasks; A mapping module is configured to receive a preset processing mapping relationship, and determine, in the preset processing mapping relationship, a task-specific model, a modal processing flow, and a fusion feature weight corresponding to the current task instruction based on the task type; A processing module is used to send the data features and the general fusion features to the task-specific model through a routing node, so that the task-specific model processes the data features and the general fusion features based on the modal processing flow and the fusion feature weights to obtain a dedicated fusion feature, and processes the dedicated fusion feature through a task execution component to obtain a task result corresponding to the task instruction.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal communication model collaborative reasoning method based on a dynamic routing mechanism described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video classification method, device and equipment based on multi-modal representation, and storage medium
CN113762322A
Information detection method based on modal dynamic feature fusion and cross-modal relation extraction
CN115563573A
Industrial visual inspection system based on multi-modal fusion
CN119444686A
Data decision-making method and system based on multi-modal large model analysis
CN119494079A
Power multi-modal data hierarchical routing feature processing and fusion method and system
CN119513817A
Cited By
Semantic understanding method and engine based on multi-model collaborative reasoning and self-supervised learning
CN121659960A
Semantic understanding method and engine based on multi-model collaborative reasoning and self-supervised learning
CN121659960B
Complex task scene-oriented multi-modal context reasoning method and device
CN121684063A