Dynamic token allocation method and device based on multi-modal fusion and storage medium

By predicting the monomer contribution degree of multimodal data and assigning weight coefficients, and dynamically allocating tokens, the problem of unreasonable allocation of multimodal data is solved, and the performance and resource utilization efficiency of the big model are improved.

CN120448143AActive Publication Date: 2025-08-08BEIJING CENTURY TAL EDUCATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510948582.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-08
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

In the prior art, the token allocation method of multimodal data is carried out in a unified proportion, resulting in waste of resources and insufficient information, which affects the problem-solving ability of the big model.

Method used

By predicting the monomer contribution degree of multimodal data, the weight coefficient is assigned to the multimodal data based on the monomial contribution degree, and the token dynamic allocation is performed.

Benefits of technology

Rationally allocate tokens, reduce invalid calculations, ensure sufficient information, and improve the problem-solving ability and resource utilization efficiency of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448143A_ABST
    Figure CN120448143A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic token allocation method and device for multi-modal fusion and a storage medium, relates to the technical field of data fusion, and mainly aims to improve the token allocation effect on multi-modal data, so that a subsequent large model pays more attention to key information in the multi-modal data when solving problems, computing resources are saved, and the efficiency is improved. And the problem solving capability of the large model based on the multi-modal data is improved. Comprising the steps that to-be-fused multi-modal data needed for processing a target problem is acquired, and in a scene of processing the target problem, the multi-modal data has a complementary relationship; predicting a single contribution degree of the multi-modal data, and distributing a weight coefficient for the multi-modal data based on the single contribution degree; and on the basis of the weight coefficient, dynamic distribution of the token is carried out on the multi-modal data. The method is mainly suitable for a token distribution scene of multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion technology, and in particular to a multimodal fusion dynamic token allocation method, device and storage medium. Background Art

[0002] Currently, large model technology is experiencing explosive growth. Large language models such as GPT-4, Claude, Qwen, and LLaMA, trained with hundreds to trillions of parameters and massive amounts of data, have pushed the boundaries of natural language understanding, logical reasoning, and code generation. They are widely used in fields such as intelligent customer service, educational tutoring, scientific research literature analysis, and legal document generation. Multimodal large models further integrate multimodal inputs such as vision, language, and audio, achieving a unified semantic representation across multiple modalities through cross-modal alignment techniques, significantly improving cognitive accuracy and adaptability for complex tasks. Representative models such as GPT-4V, InternVL, Qwen VL, Gemini, and LLaVA, integrate multimodal data such as vision, language, and audio, and achieve cross-modal joint modeling based on the Transformer architecture. Through large-scale pre-training and instruction fine-tuning, these models significantly improve cross-modal semantic understanding, generation, and reasoning capabilities. They have been widely used in fields such as image captioning, visual question answering, cross-modal retrieval, and intelligent education. Based on this, in order to improve the problem processing accuracy of large models, it is necessary to assign tokens to multimodal data.

[0003] Currently, tokens are typically assigned to all multimodal data using a uniform, artificially determined ratio. However, different modal data contain different information, and this uniform allocation can lead to suboptimal token allocation. This can result in excessive allocations of tokens to data that doesn't need them, leading to wasted resources and the introduction of additional noise. Furthermore, a small number of tokens can be assigned to data that needs more, resulting in insufficient information and hindering the problem-solving capabilities of subsequent large-scale models. Summary of the Invention

[0004] The present invention provides a dynamic token allocation method, device and storage medium for multimodal fusion, which mainly aims to improve the effect of token allocation for multimodal data, save computing resources, and thus enhance the ability of large models to solve problems based on multimodal data.

[0005] According to a first aspect of the present invention, a multimodal fusion dynamic token allocation method is provided, comprising: Acquire multimodal data to be fused required for processing a target problem, wherein there is a complementary relationship between the multimodal data in a scenario of processing the target problem; Predicting individual contributions of the multimodal data, and assigning weight coefficients to the multimodal data based on the individual contributions; Based on the weight coefficients, tokens are dynamically allocated for the multimodal data.

[0006] Optionally, the multimodal data includes text modal data; The method for predicting the monomer contribution of the text modal data includes: Inputting the text modal data multiple times into the text problem processing model to process the target problem, thereby obtaining multiple text processing results; Performing similarity matching on each text processing result and a text processing reference result corresponding to the target question, and determining the text processing accuracy of each text processing result based on the text similarity matching results; Determining the number of precise text processing results of the text processing results having a text processing accuracy greater than a first preset threshold, and taking the ratio of the number of precise text processing results to the total number of text processing results as the precise result hit rate of the text modality data; Based on the precise result hit rate, the individual contribution of the text modality data is determined.

[0007] Optionally, before performing similarity matching on each text processing result and a text processing reference result corresponding to the target question, the method further includes: The result of the target problem is predicted through a multimodal large model to obtain the text processing reference result.

[0008] Optionally, the multimodal data further includes non-text modal data; The method for predicting the individual contribution of the non-text modal data includes: Using the conversion evaluation model to judge multiple times whether the non-text modal data can be generated based on the text modal data, and determining the generation ratio based on the multiple judgment results; Determining a hit rate of accurate processing results when processing the target question based solely on the non-text modal data, and determining an importance score of the non-text modal data based on the hit rate of accurate processing results; Based on the generable determination ratio and the importance score, the individual contribution of the non-text modality data is determined.

[0009] Optionally, the conversion evaluation model is used to determine multiple times whether the non-text modal data can be generated based on the text modal data, including: Extracting text feature vectors from the text modal data and non-text feature vectors from the non-text modal data using the conversion evaluation macromodel; Calculate the cosine similarity between the text feature vector and the non-text feature vector, and determine whether the cosine similarity is greater than a preset similarity threshold; if so, determine that the non-text modal data can be generated based on the text modal data; otherwise, determine that the non-text modal data cannot be generated based on the text modal data.

[0010] Optionally, determining a hit rate of accurate processing results when processing the target question based solely on the non-text modal data includes: Inputting the non-text modal data multiple times into the non-text question processing large model to process the target question, thereby obtaining multiple non-text processing results; Performing similarity matching on each non-text processing result and a non-text processing reference result corresponding to the target question, and determining the non-text processing accuracy of each non-text processing result based on the non-text similarity matching results; The number of accurate non-text processing results of the non-text processing results whose non-text processing accuracy is greater than a second preset threshold is determined, and the ratio of the number of accurate non-text processing results to the total number of non-text processing results is used as the processing accuracy result hit rate.

[0011] Optionally, determining the individual contribution of the non-text modality data based on the generable determination ratio and the importance score includes: Get the proportion of the generated judgment The corresponding first adjustable parameter and the importance score The corresponding second adjustable parameter , based on the determination ratio that can be generated , the first adjustable parameter , the importance score The second adjustable parameter , determine the monomer contribution of the non-text modal data ,in, .

[0012] Optionally, based on the individual contribution, weight coefficients are assigned to the multimodal data respectively, including: If the multimodal data includes image data, determining the image complexity of the image data; Determining a weight coefficient of the image data based on the image complexity, the monomer contribution corresponding to the image data, and a first preset normalization coefficient; A weight coefficient of the residual modal data is determined based on a single contribution of the residual modal data and a second preset normalization coefficient, wherein the multimodal data after removing the image data is used as the residual modal data.

[0013] Optionally, determining the image complexity of the image data includes: Determining an image feature vector corresponding to the image data, and determining a feature dimension of the image feature vector, a dimension value of each feature dimension, and a dimension weight of each feature dimension; The image complexity of the image data is determined based on the feature dimension, the dimension value, and the dimension weight.

[0014] Optionally, dynamically allocating tokens for the multimodal data based on the weight coefficients includes: Divide the weight coefficient corresponding to each multimodal data by the sum of the weight coefficients of the multimodal data to obtain the proportion of the to-be-allocated tokens of the multimodal data in the overall allocated tokens, and use the ratio of each proportion as the token allocation ratio between the multimodal data; Based on the token allocation ratio, tokens are dynamically allocated for the multimodal data.

[0015] Optionally, dynamically allocating tokens for the multimodal data based on the weight coefficients includes: Determine the problem complexity of the target problem and the total token limit ; Based on the multimodal data Weight coefficient of modal data , the complexity of the problem 、Total limit of tokens , the total number of types of multimodal data , dynamically determine the Token quota for modal data ,in, , Problem complexity attenuation factor.

[0016] According to a second aspect of the present invention, a multimodal fusion dynamic token allocation device is provided, comprising: an acquisition unit, configured to acquire multimodal data to be fused required for processing a target problem, wherein, in a scenario of processing the target problem, the multimodal data have a complementary relationship; a prediction unit, configured to predict individual contributions of the multimodal data and assign weight coefficients to the multimodal data based on the individual contributions; A dynamic allocation unit is used to dynamically allocate tokens to the multimodal data based on the weight coefficients.

[0017] According to a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-mentioned dynamic token allocation method for multimodal fusion.

[0018] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned dynamic token allocation method for multimodal fusion when executing the program.

[0019] According to the present invention, a dynamic token allocation method, device, and storage medium for multimodal fusion are provided. Compared with the current method of allocating tokens to all multimodal data based on a manually determined uniform ratio, the present invention predicts the individual contributions of multimodal data and assigns weight coefficients to the multimodal data based on the individual contributions. Finally, based on the weight coefficients, tokens are dynamically allocated to the multimodal data. As a result, the present invention dynamically allocates tokens to the multimodal data based on its contribution to problem solving. It can reasonably allocate tokens to the multimodal data based on the value of the information in the data, avoiding excessive processing of low-contribution modalities under fixed allocation methods and reducing invalid calculations. At the same time, by allocating more tokens to modalities with high contribution, it can ensure the sufficiency of the information required for problem solving and enable subsequent large models to pay more attention to key information in the multimodal data when solving problems, thereby improving the problem-solving ability of subsequent large models. At the same time, the present invention dynamically allocates tokens based on the individual contributions of multimodal data and can flexibly adjust the token allocation ratio according to task requirements, thereby significantly improving model performance, resource utilization efficiency, and application adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A flow chart of a dynamic token allocation method for multimodal fusion provided by an embodiment of the present invention is shown; Figure 2 A flowchart of another multimodal fusion dynamic token allocation method provided by an embodiment of the present invention is shown; Figure 3 A schematic diagram of the structure of a multimodal fusion dynamic token allocation device provided by an embodiment of the present invention is shown; Figure 4 A schematic diagram of the structure of another multimodal fusion dynamic token allocation device provided by an embodiment of the present invention is shown; Figure 5 A schematic diagram of the physical structure of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0021] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0022] Currently, the method of allocating tokens to all multimodal data based on an artificially determined uniform ratio will result in more tokens being allocated to data that does not need too many tokens, resulting in resource waste and the introduction of additional noise. It will also result in a small number of tokens being allocated to data that needs more tokens, resulting in insufficient information, which in turn affects the problem-solving ability of subsequent large models.

[0023] In order to solve the above problems, the embodiment of the present invention provides a dynamic token allocation method for multimodal fusion, such as Figure 1 As shown, the method includes: 101. Obtain multimodal data to be fused required for processing a target problem, wherein, in a scenario of processing the target problem, there is a complementary relationship between the multimodal data.

[0024] Among them, the target problem can be a problem in any scenario such as intelligent customer service, educational guidance, literature analysis, etc., such as answering questions in test papers, question-and-answer questions in intelligent customer service scenarios, visual question-and-answer questions, etc.; multimodal data includes text modal data and non-text modal data, and non-text modal data includes image data, video data, audio data, voice data, etc.; data of different modalities complement and synergistically enhance each other by providing information from different angles or different types, so as to more comprehensively and accurately describe or solve the target problem.

[0025] In an embodiment of the present invention, if a question on a test paper is to be answered, and the question includes both text and image information, the multimodal data, such as the text and image information, is first obtained. Tokens are then dynamically assigned to the text and image information based on their contribution to the answer to the question. Finally, the question is answered based on the token assignment results. A token is the smallest processing unit of data, used to decompose continuous, complex data into structured, actionable segments.

[0026] 102. Predict individual contributions of multimodal data, and assign weight coefficients to the multimodal data based on the individual contributions.

[0027] Among them, the single contribution refers to the contribution of each modal data to solving the target problem when multimodal data is used to collaboratively solve the target problem.

[0028] For an embodiment of the present invention, if the multimodal data includes text modal data and non-text modal data, the contribution of each modal data to solving the target problem when the text modal data and non-text modal data work together to solve the target problem is determined separately. Then, based on the contribution, weight coefficients are assigned to the text modal data and the non-text modal data. That is, the greater the contribution, the more important the modal data is to solving the target problem, and a larger weight should be assigned to it. The smaller the contribution, the less important the modal data is to solving the target problem, and a smaller weight should be assigned to it. In one embodiment of the present invention, a method for assigning weight coefficients to multimodal data includes: if the multimodal data includes image data, determining the image complexity of the image data; determining the weight coefficient of the image data based on the image complexity, the individual contribution corresponding to the image data, and a first preset normalization coefficient; and determining the weight coefficient of the remaining modal data based on the individual contribution of the remaining modal data and a second preset normalization coefficient, wherein the multimodal data after removing the image data is used as the remaining modal data. Among them, the method for determining the image complexity of image data includes: determining the image feature vector corresponding to the image data, and determining the feature dimension of the image feature vector, the dimension value of each feature dimension, and the dimension weight of each feature dimension; based on the feature dimension, the dimension value, and the dimension weight, determining the image complexity of the image data.

[0029] Specifically, in the case where multimodal data includes text modal data and non-text modal data, if the non-text modal data is image data, then in order to determine the distribution weight coefficients corresponding to the text modal data and the image data respectively, it is necessary to first determine the image feature vector of the image data, such as using a pre-built image feature extraction model to extract the image feature vector in the image data, and then determine the length of the image feature vector, that is, the total number of feature dimensions , and directly obtain each feature dimension according to the generation method of the image feature vector Dimension value of If the image feature vector in the embodiment of the present invention is extracted based on the model, the dimension value of the image feature vector is determined by the model architecture, that is, the dimension value can be directly output from the image feature extraction process of the model. At the same time, according to the importance of each feature dimension to the image complexity calculation, Assigning dimension weights , and then determine the image complexity of the image data according to the following formula :

[0030] Furthermore, the weight coefficient of the image data is determined according to the following formula: :

[0031] in, is the individual contribution of image data, The first preset normalization coefficient is set according to actual needs. The embodiment of the present invention determines its weight coefficient by comprehensively analyzing the complexity of image data, monomer contribution and other information. It can assign weights to the information according to the value of the information contained in the image, ensuring the accuracy of weight allocation, and thus ensuring the rationality and accuracy of token allocation. At the same time, the weight coefficient of any remaining modal data, such as text modal data, is determined according to the following formula: :

[0032] in, is the monomer contribution of any remaining modal data, is a second preset normalization coefficient set according to actual needs. This embodiment of the present invention assigns weights based on the contribution of multimodal data to problem solving, and then allocates tokens based on the weights. This allows for reasonable allocation of tokens to multimodal data based on the value of the information contained in the data, avoiding over-processing of low-contribution modalities under a fixed allocation method, reducing ineffective computations, and thus improving problem-solving efficiency.

[0033] 103. Based on the weight coefficient, tokens are dynamically allocated for multimodal data.

[0034] For an embodiment of the present invention, after determining the weight coefficient of the multimodal data, it is necessary to dynamically allocate tokens for the multimodal data based on the weight coefficient. Based on this, step 103 includes: dividing the weight coefficient corresponding to the multimodal data individually by the sum of the weight coefficients of the multimodal data, and obtaining the proportion of the to-be-allocated tokens of the multimodal data in the overall allocated tokens, and taking the ratio of each proportion as the token allocation ratio between the multimodal data; based on the token allocation ratio, dynamically allocate tokens for the multimodal data respectively.

[0035] Specifically, the proportion of tokens to be allocated for multimodal data in the total allocated tokens is calculated according to the following formula: :

[0036] in, is the weight coefficient of the i-th modal data, is the total number of modalities of multimodal data. Finally, the ratio of the proportion of each modal data is used as the token allocation ratio between multimodal data. For example, in multimodal data, the token of text modal data can be set according to actual needs. The token allocation ratio and the token allocation number of text modal data can be used to determine the token allocation number of other modal data. The embodiment of the present invention dynamically allocates tokens to multimodal data according to the degree of contribution to problem solving, and can reasonably allocate tokens to multimodal data, avoiding excessive processing of low-contribution modalities under a fixed allocation method, reducing invalid calculations, and at the same time, by allocating more tokens to modalities with high contribution, it can ensure the sufficiency of the amount of information required for problem solving, thereby improving the problem-solving ability of subsequent large models. At the same time, the present invention dynamically allocates tokens according to the individual contribution of multimodal data, and can flexibly adjust the token allocation ratio according to task requirements, thereby significantly improving model performance, resource utilization efficiency and application adaptability.

[0037] According to a dynamic token allocation method for multimodal fusion provided by the present invention, compared with the current method of allocating tokens to all multimodal data based on a uniform ratio determined artificially, the present invention predicts the individual contribution of multimodal data and assigns weight coefficients to the multimodal data based on the individual contribution. Finally, based on the weight coefficients, tokens are dynamically allocated to the multimodal data. As a result, the present invention dynamically allocates tokens to the multimodal data according to its contribution to problem solving, can reasonably allocate tokens to the multimodal data, avoids excessive processing of low-contribution modalities under a fixed allocation method, and reduces invalid calculations. At the same time, by allocating more tokens to modalities with high contribution, it can ensure the sufficiency of the amount of information required for problem solving, and enable subsequent large models to pay more attention to key information in the multimodal data when solving problems, thereby improving the problem-solving ability of subsequent large models. At the same time, the present invention dynamically allocates tokens based on the individual contribution of multimodal data, and can flexibly adjust the token allocation ratio according to task requirements, thereby significantly improving model performance, resource utilization efficiency, and application adaptability.

[0038] Furthermore, in order to better illustrate the above process of dynamically allocating tokens, as a refinement and extension of the above embodiment, the embodiment of the present invention provides another multimodal fusion dynamic token allocation method, such as Figure 2 As shown, the method includes: 201. Obtain multimodal data to be fused required for processing a target problem, wherein, in the scenario of processing the target problem, there is a complementary relationship between the multimodal data, and the multimodal data includes text modal data and non-text modal data.

[0039] 202. Predict individual contribution of text modal data.

[0040] For the embodiment of the present invention, in order to assign tokens to text modal data, it is first necessary to determine the individual contribution of the text modal data. Based on this, step 202 specifically includes: inputting the text modal data into the text problem processing large model multiple times to process the target problem, and obtaining multiple text processing results; performing similarity matching on each text processing result with the text processing reference result corresponding to the target problem, and determining the text processing accuracy of each text processing result based on the text similarity matching result; determining the number of precise text processing results of text processing results whose text processing accuracy is greater than a first preset threshold, and taking the ratio of the number of precise text processing results to the total number of text processing results as the precise result hit rate of the text modal data; determining the individual contribution of the text modal data based on the precise result hit rate. Among them, the method for determining the text processing reference result corresponding to the target problem includes: performing result prediction on the target problem through a multimodal large model to obtain the text processing reference result.

[0041] Among them, the first preset threshold is set according to actual needs; the text processing large model can be a series of large models such as GPT and Qwen. Specifically, the text problem processing large model is used to require the large model to solve the target problem under the premise of providing only text modal data. For example, if the multimodal data is the descriptive text data and image data of a test question in the test paper, the text problem processing large model is used to require the large model to answer multiple test questions under the premise of providing only descriptive text data, and to determine whether the answer result is the same as the text processing reference result, and then determine the accurate result hit rate according to the following formula :

[0042] in, is the result of text processing, For text processing reference results, For the number of accurate text processing results, is the total number of text processing results. Further, the monomer contribution of text modal data is determined according to the following formula :

[0043] in, These are adjustable parameters that can be set based on actual needs.

[0044] The embodiment of the present invention uses a large model to solve the problem multiple times, and calculates the accuracy rate through a large amount of experimental data, which can more accurately reflect the contribution of text content to problem solving.

[0045] 203. Predicting individual contributions of non-text modal data.

[0046] For the embodiment of the present invention, in order to assign tokens to non-text modal data, it is first necessary to determine the individual contribution of the non-text modal data. Based on this, step 203 specifically includes: using the conversion evaluation large model to judge multiple times whether the non-text modal data can be generated based on the text modal data, and determining the proportion of judgments that can be generated based on the results of multiple judgments; determining the hit rate of accurate processing results when the target problem is processed based on the non-text modal data alone, and determining the importance score of the non-text modal data based on the accurate processing result hit rate; determining the individual contribution of the non-text modal data based on the proportion of judgments that can be generated and the importance score.

[0047] Among them, the method of using a conversion evaluation large model to repeatedly determine whether non-text modal data can be generated based on text modal data includes: using the conversion evaluation large model to respectively extract text feature vectors in the text modal data and non-text feature vectors in the non-text modal data; calculating the cosine similarity between the text feature vector and the non-text feature vector, and determining whether the cosine similarity is greater than a preset similarity threshold; if so, determining that the non-text modal data can be generated based on the text modal data; otherwise, determining that the non-text modal data cannot be generated based on the text modal data.

[0048] Among them, the preset similarity threshold is set according to actual needs. Specifically, if it is determined that the text modal data and the non-text modal data are not similar based on the text feature vector and the non-text feature vector, then it is determined that the non-text modal data cannot be generated based on the text modal data; if it is determined that the similarity between the text modal data and the non-text modal data is high, then it is determined that the non-text modal data can be generated based on the text modal data. Further, the determination ratio can be generated according to the following formula :

[0049] in, To determine the number of times non-text modal data can be generated based on text modal data in multiple judgment results, The total number of judgments for the conversion evaluation model.

[0050] Among them, the method for determining the hit rate of accurate processing results when processing the target problem based solely on non-text modal data includes: inputting the non-text modal data into the non-text problem processing large model multiple times to process the target problem, and obtaining multiple non-text processing results; performing similarity matching on each non-text processing result and the non-text processing reference result corresponding to the target problem, and determining the non-text processing accuracy of each non-text processing result based on the non-text similarity matching result; determining the number of precise non-text processing results of the non-text processing results whose non-text processing accuracy is greater than a second preset threshold, and taking the ratio of the number of precise non-text processing results to the total number of non-text processing results as the hit rate of accurate processing results.

[0051] Among them, the second preset threshold is set according to actual needs, and the non-text processing reference result can be obtained by predicting the result of the target problem based on the multimodal large model. Specifically, the large model for processing non-text problems is required to solve the target problem under the premise of only providing non-text modal data. For example, if the multimodal data is the descriptive text data and image data of a test question in the test paper, the large model for processing non-text problems is required to answer multiple test questions under the premise of only providing image data, and determine whether the answer result is the same as the non-text processing reference result, and then determine the accuracy of the processing result hit rate according to the following formula :

[0052] in, For non-text processing results, For non-text processing reference results, To accurately process the number of non-text results, is the total number of non-text processing results. Further, the importance score of non-text modal data is determined according to the following formula :

[0053] in, It is an adjustable parameter set according to actual needs. Further, after determining the proportion of generative judgments and the importance score, it is necessary to determine the individual contribution of non-text modal data based on the proportion of generative judgments and the importance score. Based on this, the method includes: obtaining the proportion of generative judgments The corresponding first adjustable parameter and the importance score The corresponding second adjustable parameter , based on the determination ratio that can be generated , the first adjustable parameter , the importance score The second adjustable parameter , determine the monomer contribution of the non-text modal data ,in, .

[0054] Among them, the first adjustable parameter , the second adjustable parameter They are constant values set according to actual needs.

[0055] The embodiment of the present invention determines the contribution of non-text modal data by comprehensively considering the ability of text modal data to generate non-text modal data and the importance score of non-text modal data when solving the target problem alone. By increasing the analysis dimension, the comprehensiveness of the problem analysis is ensured, thereby improving the accuracy of determining the contribution of individual non-text modal data.

[0056] 204. Based on the individual contribution, weight coefficients are assigned to the multimodal data respectively.

[0057] Specifically, if the contribution of a single entity is greater, the weight coefficient of the corresponding modal data is greater; if the contribution of a single entity is smaller, the weight coefficient of the corresponding modal data is smaller.

[0058] 205. Based on the weight coefficient, tokens are dynamically allocated for multimodal data.

[0059] For the embodiment of the present invention, after determining the weight coefficients corresponding to the multimodal data, it is necessary to dynamically allocate tokens for the multimodal data according to the weight coefficients. Based on this, step 205 specifically includes: determining the problem complexity of the target problem and the total token limit Based on the multimodal data Weight coefficient of modal data , the complexity of the problem 、Total limit of tokens , the total number of types of multimodal data , dynamically determine the Token quota for modal data ,in, , Problem complexity attenuation factor.

[0060] Specifically, first determine the text features, mathematical features, and logical features of the target problem. Text features include question stem length, number of keywords, density of professional terms, and nesting levels of conditional statements; mathematical features include the number of variables involved, the number of equations or inequalities, and the relationship between unknowns and equations; logical features include the number of logical branches, the number of implicit conditions, and the dependency relationship of multi-step reasoning. Then, assign corresponding weights to different features, and perform feature weighted scoring based on the weights. The complexity of the problem is determined based on the scoring results. The total token limit is determined based on the pre-set value according to actual needs. The token allocation in the embodiment of the present invention not only takes into account the modality weight but also performs dynamic adjustments based on the problem complexity and attenuation factor, ensuring that token resources are efficiently utilized and avoiding excessive consumption on simple problems or low-weight modalities.

[0061] According to another dynamic token allocation method for multimodal fusion provided by the present invention, compared with the current method of allocating tokens to all multimodal data based on a uniform ratio determined artificially, the present invention predicts the individual contribution of multimodal data and assigns weight coefficients to the multimodal data based on the individual contribution. Finally, based on the weight coefficients, tokens are dynamically allocated to the multimodal data. As a result, the present invention dynamically allocates tokens to the multimodal data according to its contribution to problem solving, and can reasonably allocate tokens to the multimodal data, avoiding excessive processing of low-contribution modes under a fixed allocation method and reducing invalid calculations. At the same time, by allocating more tokens to modes with high contribution, the sufficiency of the information required for problem solving can be ensured, and subsequent large models can pay more attention to key information in the multimodal data when solving problems, thereby improving the problem-solving ability of subsequent large models. At the same time, the present invention dynamically allocates tokens based on the individual contribution of multimodal data, and can flexibly adjust the token allocation ratio according to task requirements, thereby significantly improving model performance, resource utilization efficiency and application adaptability.

[0062] Further, as Figure 1 The specific implementation of the present invention provides a multi-modal fusion dynamic token allocation device, such as Figure 3 As shown, the device includes: an acquisition unit 31, a prediction unit 32, and a dynamic allocation unit 33.

[0063] The acquisition unit 31 may be configured to acquire multimodal data to be fused that is required for processing a target problem, wherein, in a scenario of processing the target problem, there is a complementary relationship between the multimodal data.

[0064] The prediction unit 32 may be configured to predict individual contributions of the multimodal data and assign weight coefficients to the multimodal data based on the individual contributions.

[0065] The dynamic allocation unit 33 may be configured to dynamically allocate tokens to the multimodal data based on the weight coefficients.

[0066] In a specific application scenario, the multimodal data includes text modal data; in order to predict the individual contribution of text modal data, such as Figure 4 As shown, the prediction unit 32 includes a processing module 321 , a matching module 322 , and a first determination module 323 .

[0067] The processing module 321 can be used to input the text modal data multiple times into the text problem processing model to process the target problem and obtain multiple text processing results.

[0068] The matching module 322 may be configured to perform similarity matching on each text processing result and the text processing reference result corresponding to the target question, and determine the text processing accuracy of each text processing result based on the text similarity matching results.

[0069] The first determination module 323 can be used to determine the number of precise text processing results of text processing results whose text processing accuracy is greater than a first preset threshold, and use the ratio of the number of precise text processing results to the total number of text processing results as the precise result hit rate of the text modal data.

[0070] The first determination module 323 may also be used to determine the individual contribution of the text modality data based on the precise result hit rate.

[0071] In a specific application scenario, in order to determine the text processing reference result, the first determination module 323 can also be used to predict the result of the target problem through a multimodal large model to obtain the text processing reference result.

[0072] In a specific application scenario, the multimodal data also includes non-text modal data; in order to predict the individual contribution of the non-text modal data, the prediction unit 32 further includes a judgment module 324.

[0073] The judgment module 324 can be used to use the conversion evaluation model to judge multiple times whether the non-text modal data can be generated based on the text modal data, and determine the generation determination ratio based on the multiple judgment results.

[0074] The first determination module 323 can also be used to determine the hit rate of accurate processing results when processing the target problem based solely on the non-text modal data, and determine the importance score of the non-text modal data based on the hit rate of accurate processing results.

[0075] The first determination module 323 may also be configured to determine the individual contribution of the non-text modality data based on the generable determination ratio and the importance score.

[0076] In a specific application scenario, in order to determine whether non-text modal data can be generated based on text modal data, the judgment module 324 can be specifically used to use the conversion evaluation model to respectively extract the text feature vectors in the text modal data and the non-text feature vectors in the non-text modal data; calculate the cosine similarity between the text feature vector and the non-text feature vector, and determine whether the cosine similarity is greater than a preset similarity threshold; if so, determine that the non-text modal data can be generated based on the text modal data; otherwise, determine that the non-text modal data cannot be generated based on the text modal data.

[0077] In a specific application scenario, in order to determine the hit rate of accurate processing results when processing the target problem based solely on non-text modal data, the first determination module 323 can be specifically used to input the non-text modal data multiple times into the non-text problem processing large model to process the target problem, and obtain multiple non-text processing results; perform similarity matching on each non-text processing result with the non-text processing reference result corresponding to the target problem, and determine the non-text processing accuracy of each non-text processing result based on the non-text similarity matching result; determine the number of precise non-text processing results of the non-text processing results whose non-text processing accuracy is greater than a second preset threshold, and use the ratio of the number of precise non-text processing results to the total number of non-text processing results as the hit rate of accurate processing results.

[0078] In a specific application scenario, in order to determine the monomer contribution of non-text modal data, the first determination module 323 can be used to obtain the generative determination ratio. The corresponding first adjustable parameter and the importance score The corresponding second adjustable parameter , based on the determination ratio that can be generated , the first adjustable parameter , the importance score The second adjustable parameter , determine the monomer contribution of the non-text modal data ,in, .

[0079] In a specific application scenario, in order to assign a weight coefficient to multimodal data, the first determination module 323 can also be used to determine the image complexity of the image data if the multimodal data includes image data; determine the weight coefficient of the image data based on the image complexity, the individual contribution corresponding to the image data, and a first preset normalization coefficient; determine the weight coefficient of the remaining modal data based on the individual contribution of the remaining modal data and a second preset normalization coefficient, wherein the multimodal data after removing the image data is used as the remaining modal data.

[0080] In a specific application scenario, in order to determine the image complexity of image data, the first determination module 323 can be specifically used to determine the image feature vector corresponding to the image data, and determine the feature dimension of the image feature vector, the dimension value of each feature dimension, and the dimension weight of each feature dimension; based on the feature dimension, the dimension value, and the dimension weight, the image complexity of the image data is determined.

[0081] In a specific application scenario, in order to dynamically allocate tokens for multimodal data, the dynamic allocation unit 33 includes a division module 331 and a dynamic allocation module 332 .

[0082] The division module 331 can be used to divide the weight coefficient corresponding to each multimodal data by the sum of the weight coefficients of the multimodal data, and obtain the proportion of the to-be-allocated token of the multimodal data in the overall allocated token, and use the ratio of each proportion as the token allocation ratio between the multimodal data.

[0083] The dynamic allocation module 332 may be configured to dynamically allocate tokens to the multimodal data based on the token allocation ratio.

[0084] In a specific application scenario, in order to dynamically allocate tokens for multimodal data, the dynamic allocation unit 33 further includes a second determination module 333 .

[0085] The second determination module 333 can be used to determine the problem complexity of the target problem. and the total token limit Based on the multimodal data Weight coefficient of modal data , the complexity of the problem 、Total limit of tokens , the total number of types of multimodal data , dynamically determine the Token quota for modal data ,in, , Problem complexity attenuation factor.

[0086] It should be noted that for other corresponding descriptions of the functional modules involved in the multimodal fusion dynamic token allocation device provided in the embodiment of the present invention, please refer to Figure 1 The corresponding description of the method shown will not be repeated here.

[0087] Based on the above Figure 1 The method shown, accordingly, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the following steps when executed by a processor: obtaining multimodal data to be fused required for processing the target problem, wherein, in the scenario of processing the target problem, there is a complementary relationship between the multimodal data; predicting the individual contribution of the multimodal data, and based on the individual contribution, assigning weight coefficients to the multimodal data respectively; based on the weight coefficients, dynamically assigning tokens to the multimodal data respectively.

[0088] Based on the above Figure 1 The method shown and Figure 3 The embodiment of the device shown in the figure, the embodiment of the present invention also provides a physical structure diagram of a computer device, such as Figure 5 As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor, wherein the memory 42 and the processor 41 are both arranged on a bus 43, and the processor 41 implements the following steps when executing the program: obtaining multimodal data to be fused required for processing the target problem, wherein, in the scenario of processing the target problem, there is a complementary relationship between the multimodal data; predicting the individual contribution of the multimodal data, and assigning weight coefficients to the multimodal data based on the individual contribution; and dynamically assigning tokens to the multimodal data based on the weight coefficients.

[0089] Through the technical solution of the present invention, the present invention predicts the individual contribution of multimodal data, and based on the individual contribution, assigns weight coefficients to the multimodal data respectively, and finally dynamically assigns tokens to the multimodal data based on the weight coefficients. Therefore, the present invention dynamically assigns tokens to the multimodal data according to the contribution degree of the multimodal data to problem solving, and can reasonably assign tokens to the multimodal data, avoiding excessive processing of low-contribution modes under a fixed allocation method, reducing invalid calculations, and at the same time, by allocating more tokens to modes with high contribution, it can ensure the sufficiency of the amount of information required for problem solving, and enable subsequent large models to pay more attention to the key information in the multimodal data when solving problems, thereby improving the problem-solving ability of subsequent large models. At the same time, the present invention dynamically assigns tokens according to the individual contribution of multimodal data, and can flexibly adjust the token allocation ratio according to task requirements, thereby significantly improving model performance, resource utilization efficiency and application adaptability.

[0090] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, centralized on a single computing device, or distributed across a network of multiple computing devices. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0091] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A dynamic token allocation method for multimodal fusion, characterized in that: include: Acquire multimodal data to be fused required for processing a target problem, wherein there is a complementary relationship between the multimodal data in a scenario of processing the target problem; Predicting individual contributions of the multimodal data, and assigning weight coefficients to the multimodal data based on the individual contributions; Based on the weight coefficients, tokens are dynamically allocated for the multimodal data.

2. The method according to claim 1, characterized in that The multimodal data includes text modal data; The method for predicting the monomer contribution of the text modal data includes: Inputting the text modal data multiple times into the text problem processing model to process the target problem, thereby obtaining multiple text processing results; Performing similarity matching on each text processing result and a text processing reference result corresponding to the target question, and determining the text processing accuracy of each text processing result based on the text similarity matching results; Determining the number of precise text processing results of the text processing results having a text processing accuracy greater than a first preset threshold, and taking the ratio of the number of precise text processing results to the total number of text processing results as the precise result hit rate of the text modality data; Based on the precise result hit rate, the individual contribution of the text modality data is determined.

3. The method according to claim 2, characterized in that Before performing similarity matching on each text processing result and the text processing reference result corresponding to the target question, the method further includes: The result of the target problem is predicted through a multimodal large model to obtain the text processing reference result.

4. The method according to claim 2, characterized in that The multimodal data also includes non-text modal data; The method for predicting the individual contribution of the non-text modal data includes: Using the conversion evaluation model to judge multiple times whether the non-text modal data can be generated based on the text modal data, and determining the generation ratio based on the multiple judgment results; Determining a hit rate of accurate processing results when processing the target question based solely on the non-text modal data, and determining an importance score of the non-text modal data based on the hit rate of accurate processing results; Based on the generable determination ratio and the importance score, the individual contribution of the non-text modality data is determined.

5. The method according to claim 4, characterized in that The conversion evaluation model is used to determine multiple times whether the non-text modal data can be generated based on the text modal data, including: Extracting text feature vectors from the text modal data and non-text feature vectors from the non-text modal data using the conversion evaluation macromodel; Calculate the cosine similarity between the text feature vector and the non-text feature vector, and determine whether the cosine similarity is greater than a preset similarity threshold; if so, determine that the non-text modal data can be generated based on the text modal data; otherwise, determine that the non-text modal data cannot be generated based on the text modal data.

6. The method according to claim 4, characterized in that Determining a hit rate of accurate processing results when processing the target question based solely on the non-text modal data includes: Inputting the non-text modal data multiple times into the non-text question processing large model to process the target question, thereby obtaining multiple non-text processing results; Performing similarity matching on each non-text processing result and a non-text processing reference result corresponding to the target question, and determining the non-text processing accuracy of each non-text processing result based on the non-text similarity matching results; The number of accurate non-text processing results of the non-text processing results whose non-text processing accuracy is greater than a second preset threshold is determined, and the ratio of the number of accurate non-text processing results to the total number of non-text processing results is used as the processing accuracy result hit rate.

7. The method according to claim 4, characterized in that Determining the individual contribution of the non-text modality data based on the generable determination ratio and the importance score includes: Get the proportion of the generated judgment The corresponding first adjustable parameter and the importance score The corresponding second adjustable parameter , based on the determination ratio that can be generated , the first adjustable parameter , the importance score The second adjustable parameter , determine the monomer contribution of the non-text modal data ,in, .

8. The method according to claim 1, characterized in that Based on the individual contribution, weight coefficients are assigned to the multimodal data respectively, including: If the multimodal data includes image data, determining the image complexity of the image data; Determining a weight coefficient of the image data based on the image complexity, the monomer contribution corresponding to the image data, and a first preset normalization coefficient; A weight coefficient of the residual modal data is determined based on a single contribution of the residual modal data and a second preset normalization coefficient, wherein the multimodal data after removing the image data is used as the residual modal data.

9. The method according to claim 8, characterized in that Determining the image complexity of the image data includes: Determining an image feature vector corresponding to the image data, and determining a feature dimension of the image feature vector, a dimension value of each feature dimension, and a dimension weight of each feature dimension; The image complexity of the image data is determined based on the feature dimension, the dimension value, and the dimension weight.

10. The method according to claim 1, characterized in that Based on the weight coefficients, dynamically allocating tokens for the multimodal data includes: Divide the weight coefficient corresponding to each multimodal data by the sum of the weight coefficients of the multimodal data to obtain the proportion of the to-be-allocated tokens of the multimodal data in the overall allocated tokens, and use the ratio of each proportion as the token allocation ratio between the multimodal data; Based on the token allocation ratio, tokens are dynamically allocated for the multimodal data.

11. The method according to claim 1, wherein Based on the weight coefficients, dynamically allocating tokens for the multimodal data includes: Determine the problem complexity of the target problem and the total token limit ; Based on the multimodal data Weight coefficient of modal data , the complexity of the problem 、Total limit of tokens , the total number of types of multimodal data , dynamically determine the Token quota for modal data ,in, , Problem complexity The attenuation factor.

12. A multimodal fusion dynamic token allocation device, characterized in that: include: an acquisition unit, configured to acquire multimodal data to be fused required for processing a target problem, wherein, in a scenario of processing the target problem, the multimodal data have a complementary relationship; a prediction unit, configured to predict individual contributions of the multimodal data and assign weight coefficients to the multimodal data based on the individual contributions; A dynamic allocation unit is used to dynamically allocate tokens to the multimodal data based on the weight coefficients.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Information input system

    CN118210383A

  • Multi-modal data fusion control method and device, equipment and medium

    CN118734250A

  • Multi-modal intention recognition method, system and device under uncertain mode and medium

    CN119646593A

  • Cross-modal question and answer processing method and device and storage medium

    CN119719435A

  • Multi-Modal Fusion Techniques Considering Inter-Modality Correlations and Computer Model Uncertainty

    US20220309295A1