Multi-modal model reasoning method and device and multi-modal model training method and device
By generating and fusing modal encodings based on benchmark data in the multimodal model and replacing the preset placeholder encodings with the multimodal encoder, the problem of inaccurate modal alignment is solved and the inference accuracy of the multimodal model is improved.
Patent Information
- Application Number
- CN202510981436.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
AI Technical Summary
When performing reasoning based on multimodal models, the modal alignment is not accurate enough, resulting in inaccurate reasoning results.
By obtaining any one of the input data of at least two modalities as the benchmark data, determining the benchmark code, and encoding the input data of other modalities as the data to be aligned, a modal code is generated, including a preset start code, a preset end code and a preset placeholder code. After the fusion code is input into the multimodal model, the multimodal encoder is used to replace the preset placeholder code to improve the accuracy of the code alignment.
The encoding alignment accuracy of the input data of each modality of the multimodal model is improved, thereby improving the inference accuracy of the multimodal model.
Smart Images

Figure CN120633867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal model reasoning method, training method and device. Background Art
[0002] When performing inference based on a multimodal model, the model typically receives input data in different modalities, such as text and images, and generates inference output based on the inference request. Inference based on a multimodal model requires encoding the input data in different modalities using different encoders. The encodings corresponding to the input data in different modalities are then aligned, and the multimodal model performs inference based on the aligned encodings of the different modalities. However, this modal alignment is often imprecise and incomplete, resulting in inaccurate inference results from the multimodal model. Summary of the Invention
[0003] The present invention provides a multimodal model inference method, training method and device, which improve the accuracy of the multimodal model's inference of input data of each modality.
[0004] According to one aspect of the present invention, a multimodal model reasoning method is provided, comprising:
[0005] Acquire input data of at least two modalities, and use any one of the at least two modalities of input data as reference data, and determine a reference code corresponding to the reference data;
[0006] Using other modal input data except the reference data among at least two modal input data as data to be aligned, determining a first number of codes generated by encoding each type of the data to be aligned, and determining a modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code;
[0007] Each modal code is fused with the reference code to generate a fused code, and the fused code is input into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
[0008] According to another aspect of the present invention, a multimodal model training method is provided, comprising:
[0009] Acquire at least two multimodal sample data groups; wherein each of the multimodal sample data groups contains sample data of at least two modalities;
[0010] For each of the multimodal sample data groups, taking sample data of any modality in the multimodal sample data group as reference data, determining a reference code corresponding to the reference data;
[0011] Using the sample data other than the reference data in the multimodal sample data group as data to be processed, determining the number of second codes generated by encoding the data to be processed in each modality, and determining a sample code corresponding to the data to be processed based on the second number of codes; wherein the sample code includes a preset start code, a preset end code, and the second number of preset placeholder codes located between the preset start code and the preset end code;
[0012] fusing the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code;
[0013] A preset large language model is trained based on at least two data group codes to generate a multimodal model; wherein, in the process of training the preset large language model based on at least two data group codes, the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model replaces the preset placeholder code between the corresponding preset start code and the preset end code in the data group code.
[0014] According to another aspect of the present invention, there is provided a multimodal model reasoning apparatus, comprising:
[0015] a reference code determination module, configured to obtain input data of at least two modalities, and use any one of the input data of the at least two modalities as reference data to determine a reference code corresponding to the reference data;
[0016] a modal code determination module, configured to use other modal input data, excluding the reference data, from among at least two modal input data as data to be aligned, determine a first number of codes generated by encoding each type of the data to be aligned, and determine a modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code;
[0017] An inference result generation module is used to fuse each of the modal codes with the reference code to generate a fused code, and input the fused code into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
[0018] According to another aspect of the present invention, a multimodal model training device is provided, comprising:
[0019] A multimodal sample data group acquisition module, configured to acquire at least two multimodal sample data groups; wherein each of the multimodal sample data groups contains sample data of at least two modalities;
[0020] a reference code determination module, configured to determine, for each of the multimodal sample data groups, a reference code corresponding to the reference data, using sample data of any modality in the multimodal sample data group as reference data;
[0021] a sample code determination module, configured to treat the sample data other than the reference data in the multimodal sample data group as to-be-processed data, determine the number of second codes generated by encoding the to-be-processed data of each modality, and determine a sample code corresponding to the to-be-processed data based on the second number of codes; wherein the sample code includes a preset start code, a preset end code, and the second number of preset placeholder codes located between the preset start code and the preset end code;
[0022] a data group code generation module, configured to fuse the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code;
[0023] A multimodal model generation module is configured to train a preset large language model based on at least two data group codes to generate a multimodal model; wherein, during the training of the preset large language model based on at least two data group codes, the preset placeholder code between the preset start code and the preset end code corresponding to the data group code is replaced by the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model.
[0024] According to another aspect of the present invention, an electronic device is provided, comprising:
[0025] at least one processor; and
[0026] a memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multimodal model inference method or the multimodal model training method described in any embodiment of the present invention.
[0028] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal model inference method or the multimodal model training method described in any embodiment of the present invention when executed.
[0029] The multimodal model inference scheme of an embodiment of the present invention obtains input data of at least two modalities, and uses any modal input data of the at least two modal input data as reference data to determine the reference code corresponding to the reference data; uses the other modal input data of the at least two modal input data except the reference data as data to be aligned, determines the first number of codes generated by encoding each type of data to be aligned, and determines the modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code; each modal code is fused with the reference code to generate a fused code, and the fused code is input into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code corresponding to the preset start code and the preset end code in the fused code. The technical solution provided by the embodiment of the present invention adds a preset start code, a preset placeholder code and a preset end code corresponding to each type of data to be aligned in the fusion code, so that in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model only replaces the preset placeholder code between the corresponding preset start code and the preset end code in the fusion code, thereby effectively avoiding the influence of the preset placeholder code generated by model reasoning at other positions in the fusion code, improving the accuracy of the coding alignment of the input data of each modality of the multimodal model, and thereby improving the accuracy of the multimodal model reasoning.
[0030] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 A flowchart of a multimodal model reasoning method provided by an embodiment of the present invention;
[0033] Figure 2 A flowchart of a multimodal model training method provided by an embodiment of the present invention;
[0034] Figure 3 A schematic diagram of the structure of a multimodal model inference device provided by an embodiment of the present invention;
[0035] Figure 4 A schematic diagram of the structure of a multimodal model training device provided by an embodiment of the present invention;
[0036] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0038] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0039] Example 1
[0040] Figure 1 A flowchart of a multimodal model reasoning method is provided for the first embodiment of the present invention. This embodiment is applicable to situations where reasoning is performed based on a multimodal model. The method can be performed by a multimodal model reasoning device. The multimodal model reasoning device can be implemented in the form of hardware and / or software. The multimodal model reasoning device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0041] S110 , obtaining input data of at least two modalities, and using any one of the input data of the at least two modalities as reference data, determining a reference code corresponding to the reference data.
[0042] In an embodiment of the present invention, in response to a multimodal model reasoning event being triggered, input data of at least two modalities is obtained. Exemplarily, when a multimodal model reasoning request is detected, it is determined that the multimodal model reasoning event is triggered, wherein the multimodal model reasoning request can be understood as a request for reasoning through a multimodal model. For example, the multimodal model can be a model for reasoning the type of content that the user is interested in (such as the type of pictures that the user is interested in, the type of music that the user is interested in, the type of short videos that the user is interested in, etc.). It should be noted that the embodiment of the present invention does not limit the application scenario of the multimodal model. Optionally, the input data of at least two modalities may include input data of text modality (i.e., text data), input data of audio modality (i.e., audio data), input data of video modality (i.e., video data), and input data of picture modality (i.e., picture data), etc. It should be noted that the embodiment of the present invention does not limit the modal type and number of modalities of the input data.
[0043] In an embodiment of the present invention, input data of any one of at least two modal input data is used as reference data to determine a reference code corresponding to the reference data. Exemplarily, the input data of at least two modalities includes input data of text modality, input data of audio modality, and input data of image modality. In this case, the input data of text modality can be used as the reference data, the input data of audio modality can be used as the reference data, or the input data of image modality can be used as the reference data. The reference data is input into a corresponding encoder, and the encoder obtains a reference code corresponding to the reference data. Optionally, the reference data is text data. Since text data has a relatively small data volume, a low compression ratio, a fast encoding speed, and high multimodal model inference efficiency, data of other modalities requires highly compressed encoding. For example, for 44.1KHz audio, there are 44,100 sampling points per second. The audio information of every 441 sampling points is compressed to generate an audio code, and one second of audio will also generate 100 codes. Similarly, image data and video data also need to be highly compressed, resulting in a slow code generation speed. Therefore, using text data as benchmark data can effectively improve the efficiency of encoding generation, thereby further improving the efficiency of reasoning based on multimodal models.
[0044] S120. Take the other modal input data except the reference data in the input data of at least two modalities as the data to be aligned, determine the first code quantity generated by encoding each type of the data to be aligned respectively, and determine the modal code corresponding to the data to be aligned based on the first code quantity; wherein the modal code includes a preset start code, a preset end code and the first code quantity of preset placeholder codes located between the preset start code and the preset end code.
[0045] In an embodiment of the present invention, input data of other modalities other than the reference data in at least two modalities are used as data to be aligned. Exemplarily, the input data of at least two modalities include input data of text modality, input data of audio modality, and input data of image modality. If the input data of text modality is used as the reference data, the input data of audio modality and the input data of image modality are both data to be aligned. For each type of data to be aligned, the first number of codes generated by encoding the data to be aligned is determined based on the size of the data to be aligned. Exemplarily, the data to be aligned is a 128*128 image, and one code is generated for every 16*16 pixels. In the absence of repeated encoding, the first number of codes generated by encoding the data to be aligned is 64. In another exemplary embodiment, the data to be aligned is 44.1KHz audio, where the audio has 44100 sampling points per second. If the audio information of every 441 sampling points is compressed to generate an audio code, in the absence of repeated encoding, the first number of codes generated by encoding one second of audio is 100.
[0046] In an embodiment of the present invention, the modal code corresponding to the data to be aligned is determined based on the first code quantity M; wherein the modal code includes a preset start code, a preset end code, and a first code quantity of preset placeholder codes located between the preset start code and the preset end code. It can be understood that for each type of data to be aligned, a preset start code is added to the starting position of the data to be aligned, a preset end code is added to the end position, and according to the first code quantity generated by encoding each type of data to be aligned, the data to be aligned is converted into the first code quantity M preset placeholder codes, thereby generating a modal code corresponding to the data to be aligned. For example, if the data to be aligned is a picture, the preset placeholder code can be<img_1> 、<img_feat> It should be noted that, for different types of data to be aligned, the corresponding preset start code, preset end code and preset placeholder code in the modal code are different.
[0047] S130. Fuse each of the modal codes with the reference code to generate a fused code, and input the fused code into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
[0048] In an embodiment of the present invention, based on a preset code fusion strategy, the modal code corresponding to each data to be aligned is fused with the reference code to generate a fused code, which is then input into a pre-trained multimodal model so that the multimodal model analyzes the fused code and obtains an inference result corresponding to the fused code output by the multimodal model. In the process of obtaining the inference result based on the multimodal model, the multimodal encoder in the multimodal model continuously generates actual modal codes corresponding to the data to be aligned. Since the multimodal model is an autoregressive model, its characteristics determine that in the subsequent model inference step, all previous input codes and model inference generated codes are input as input codes to perform model inference. Therefore, in an embodiment of the present invention, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder replaces the preset placeholder code between the corresponding preset start code and the preset end code in the fused code to avoid the influence of the preset placeholder code generated by model inference at other positions in the fused code, thereby improving the accuracy of the replacement of the preset placeholder code. Exemplarily, the data to be aligned includes image data and audio data, then the picture modal code (i.e., the actual modal code) corresponding to the picture data generated by the picture encoder in the multimodal model and the audio modal code (i.e., the actual modal code) corresponding to the audio data generated by the audio encoder in the multimodal model are obtained, and then the preset picture placeholder code between the preset picture start code and the preset picture end code corresponding to the picture data in the fusion code is replaced based on the picture modal code, and the preset audio placeholder code between the preset audio start code and the preset audio end code corresponding to the audio data in the fusion code is replaced based on the audio modal code.
[0049] The multimodal model inference method of an embodiment of the present invention obtains input data of at least two modalities, and uses any modal input data of the at least two modal input data as reference data to determine the reference code corresponding to the reference data; uses the other modal input data of the at least two modal input data except the reference data as data to be aligned, determines the first number of codes generated by encoding each type of data to be aligned, and determines the modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code; each modal code is fused with the reference code to generate a fused code, and the fused code is input into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code corresponding to the preset start code and the preset end code in the fused code. The technical solution provided by the embodiment of the present invention adds a preset start code, a preset placeholder code and a preset end code corresponding to each type of data to be aligned in the fusion code, so that in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model only replaces the preset placeholder code between the corresponding preset start code and the preset end code in the fusion code, thereby effectively avoiding the influence of the preset placeholder code generated by model reasoning at other positions in the fusion code, improving the accuracy of the coding alignment of the input data of each modality of the multimodal model, and thereby improving the accuracy of the multimodal model reasoning.
[0050] In some embodiments, the fusion code is input into a pre-trained multimodal model to obtain an inference result corresponding to the fusion code output by the multimodal model, including: inputting the fusion code into a pre-trained multimodal model to determine a plurality of first core intermediate results output by the multimodal model; feeding back the plurality of first core intermediate results to the input end of the multimodal model to determine a plurality of second core intermediate results output by the multimodal model; and obtaining an inference result corresponding to the fusion code output by the multimodal model based on the plurality of second core intermediate results, or at least one first core intermediate result and at least one second core intermediate result.
[0051] In an embodiment of the present invention, after the fusion code is input into a pre-trained multimodal model, a plurality of first core intermediate results Logits output by the multimodal model are determined, wherein the first core intermediate result Logits is the original value output by the last layer during the inference process of the multimodal model. It can be understood that the first core intermediate result vector constituted by the first core intermediate result is an unnormalized real number vector, and each value in the first core intermediate result vector corresponds to a category. The value range of the plurality of first core intermediate results Logits is any real number, which can be a positive number, a negative number, or even a very large or very small value, and is used to represent the "tendency" of the multimodal model to a certain category when performing inference, wherein the larger the value of the first core intermediate result Logits, the stronger the "tendency" of the multimodal model to the category.
[0052] In an embodiment of the present invention, multiple first core intermediate results are fed back to the input end of the multimodal model, so that the multimodal model performs another inference based on the multiple first core intermediate results and the fusion code to determine the multiple second core intermediate results output by the multimodal model. Then, the inference result corresponding to the fusion code output by the multimodal model is obtained based on the multiple second core intermediate results. That is, the multimodal model performs subsequent calculations, outputs and / or normalization processing based on the multiple second core intermediate results to obtain the corresponding inference result. Optionally, at least one first core intermediate result and at least one second core intermediate result are used to obtain the inference result corresponding to the fusion code output by the multimodal model. That is, the multimodal model performs subsequent calculations, outputs and / or normalization processing based on at least one first core intermediate result and at least one second core intermediate result to obtain the corresponding inference result. The advantage of this setting is that it can effectively improve the accuracy of multimodal model inference and avoid re-inference of all input data from scratch. Only the first core central result output in the second half needs to be re-inferred to obtain the inference result, avoiding waste of computing resources and improving inference efficiency.
[0053] Optionally, before feeding back the multiple first core intermediate results to the input end of the multimodal model, the method further includes: aligning the multiple first core intermediate results with a reference code; feeding back the multiple first core intermediate results to the input end of the multimodal model, and determining the multiple second core intermediate results output by the multimodal model, including: feeding back the aligned reference code and the multiple first core intermediate results to the input end of the multimodal model, and determining the multiple second core intermediate results output by the multimodal model. Exemplarily, the multiple first core intermediate results are aligned with the reference code, and the aligned reference code and the multiple first core intermediate results are fed back to the input end of the multimodal model, so that the multimodal model is again inferred based on the aligned reference code and the multiple first core intermediate results to obtain the multiple second core intermediate results output by the multimodal model. The advantage of such a setting is that it can effectively solve the problem of incorrect alignment and inaccurate inference caused by gradually using the generated actual modal code to replace the preset placeholder code during the first inference process of the multimodal model, thereby further improving the accuracy of multimodal model inference.
[0054] Optionally, before feeding back multiple first core intermediate results to the input end of the multimodal model, it also includes: determining the target code in the multiple first core intermediate results, and deleting the target code from the multiple first core intermediate results; wherein the target code includes the start and end codes and / or placeholder codes generated by the multimodal model in the process of determining the first core intermediate results; the start and end codes include the preset start code and the preset end code; feeding back the multiple first core intermediate results to the input end of the multimodal model, and determining the multiple second core intermediate results output by the multimodal model, including: feeding back the first core intermediate result after deleting the target code to the input end of the multimodal model, and determining the multiple second core intermediate results output by the multimodal model.
[0055] In an embodiment of the present invention, in the process of determining the first core intermediate result by the multimodal model, the multimodal encoder in the multimodal model will continuously generate actual modal codes, wherein the actual modal codes may incorrectly generate start and end codes and / or placeholder codes. At the same time, the multimodal model may also incorrectly generate start and end codes and / or placeholder codes in the reasoning process of determining the first core intermediate result. The start and end codes include a preset start code and a preset end code, and the start and end codes and / or placeholder codes are used as target codes. It is determined whether the target code exists in the multiple first core intermediate results. If so, the target code is deleted from the multiple first core intermediate results, and the multiple first core intermediate results after deleting the target code are fed back to the input end of the multimodal model, so that the multimodal model can perform reasoning again based on the multiple first core intermediate results after deleting the target code, and obtain multiple second core intermediate results output by the multimodal model. The advantage of this setting is that only meaningful codes are retained in the input data of the multimodal model, ensuring that the codes of each modality in the input data of the second inference can be accurately aligned, thereby avoiding the situation where the multimodal model reasoning is not accurate due to the failure of modal alignment, and improving the reasoning accuracy of the multimodal model.
[0056] Optionally, deleting the target code from multiple first core intermediate results includes: determining a portion of the first core intermediate results from multiple first core intermediate results, and deleting the target code from the portion of the first core intermediate results; wherein the portion of the first core intermediate results includes the target code; feeding back the first core intermediate result after deleting the target code to the input end of the multimodal model, and determining the second core intermediate result output by the multimodal model includes: feeding back the portion of the first core intermediate result after deleting the target code to the input end of the multimodal model, so that the multimodal model outputs subsequent multiple second core intermediate results based on the portion of the first core intermediate result after deleting the target code.
[0057] In an embodiment of the present invention, not every first core intermediate result in a plurality of first core intermediate results necessarily contains a target code. Therefore, a portion of the first core intermediate results containing the target code is determined from the plurality of first core intermediate results, and the target code is deleted from the portion of the first core intermediate results. The portion of the first core intermediate results after the target code is deleted is then fed back to the input of the multimodal model, causing the multimodal model to re-infer based on the portion of the first core intermediate results after the target code is deleted, obtaining the multimodal model output corresponding to the portion of the first core intermediate results and the subsequent plurality of second core intermediate results. For example, the plurality of first core intermediate results contain 26 codes, representing the meaning of "The user inputs an image, and the content of the image is the addition operation 1+1=?" If the target code appears in the 13th core intermediate result among the plurality of first core intermediate results, the target code in the portion of the first core intermediate results is deleted and fed back to the input of the multimodal model, causing the multimodal model to re-infer the 13th and subsequent first core intermediate results based on the portion of the first core intermediate results after the target code is deleted, and using the re-inferred core intermediate results as the second core intermediate results. The advantage of this setting is that it can further save computing resources and improve the reasoning efficiency of the multimodal model while ensuring that the actual modal code generated during the second reasoning process can avoid erroneous replacement of the preset placeholder code.
[0058] In some embodiments, before feeding back multiple first core intermediate results to the input end of the multimodal model, it also includes: determining the actual modal coding generated by the multimodal encoder in the multimodal model during the generation of the multiple first core intermediate results, and replacing the preset placeholder coding in the fusion coding based on the actual modal coding; feeding back the multiple first core intermediate results to the input end of the multimodal model, and determining multiple second core intermediate results output by the multimodal model, including: feeding back the multiple first core intermediate results to the input end of the multimodal model so that the multimodal model determines the second core intermediate results output by the multimodal model based on the multiple first core intermediate results and the fusion coding after the preset placeholder coding is replaced by the actual modal coding.
[0059] In an embodiment of the present invention, the actual modal coding generated by the multimodal encoder in the multimodal model during the generation of multiple first core intermediate results is determined, wherein the actual modal coding may include modal coding generated by multimodal encoders such as the picture encoder, audio encoder and video encoder in the multimodal model, and the preset placeholder coding in the fusion coding is replaced based on the actual modal coding, and the fusion coding after the preset placeholder coding is replaced by the actual modal coding and the multiple first core intermediate results are fed back to the input end of the multimodal model, so that the multimodal model can perform inference again based on the multiple first core intermediate results and the fusion coding after the preset placeholder coding is replaced by the actual modal coding to determine the second core intermediate result output by the multimodal model.
[0060] Figure 2 This is a flowchart of a multimodal model training method provided by an embodiment of the present invention. This embodiment is applicable to the case of multimodal model training. The method can be executed by a multimodal model training device. The multimodal model training device can be implemented in the form of hardware and / or software. The multimodal model training device can be configured in an electronic device. Figure 2 As shown, the method includes:
[0061] S210: Acquire at least two multimodal sample data groups; wherein each of the multimodal sample data groups includes sample data of at least two modalities.
[0062] In an embodiment of the present invention, there are at least two multimodal sample data groups, wherein each multimodal sample data group contains sample data of at least two modalities. For example, a multimodal sample data group may contain sample data of different modalities, such as sample data of text modality, sample data of audio modality, sample data of video modality, and sample data of image modality. It should be noted that the modal types of the sample data contained in each multimodal sample data group may be the same or different, and the embodiment of the present invention does not limit the modal types of the sample data contained in each multimodal sample data group.
[0063] S220 : For each of the multimodal sample data groups, use sample data of any modality in the multimodal sample data group as reference data, and determine a reference code corresponding to the reference data.
[0064] In an embodiment of the present invention, for each multimodal sample data group, the sample data of any modality in the multimodal sample data group is used as reference data to determine the reference code corresponding to the reference data. Exemplarily, the multimodal sample data group includes sample data of text modality, sample data of audio modality, sample data of picture modality and sample data of video modality, then the sample data of text modality can be used as reference data, the sample data of audio modality can be used as reference data, the sample data of picture modality can be used as reference data, or the sample data of video modality can be used as reference data. The reference data is input into the encoder, and the reference code corresponding to the reference data is obtained by the encoder. It should be noted that for different multimodal sample data groups, the sample data of different modalities in the multimodal sample data group can be used as reference data, that is, the modality types corresponding to the reference data in each multimodal sample data group can be the same or different.
[0065] S230. Take the other sample data in the multimodal sample data group except the reference data as the data to be processed, determine the second code quantity generated by encoding the data to be processed of each modality respectively, and determine the sample code corresponding to the data to be processed based on the second code quantity; wherein the sample code includes a preset start code, a preset end code and the second code quantity of preset placeholder codes located between the preset start code and the preset end code.
[0066] In an embodiment of the present invention, for each multimodal sample data set, sample data in other modalities in the multimodal sample data set, excluding reference data, is used as data to be processed. For example, if a multimodal sample data set includes sample data in a text modality, sample data in an audio modality, and sample data in an image modality, and if the sample data in the text modality is used as reference data, then both the sample data in the audio modality and the sample data in the image modality are used as data to be processed. For each type of data to be processed, the number of second codes generated by encoding the data to be processed is determined based on the size of the data to be processed. Exemplarily, the data to be processed is a 128*128 picture, and one code is generated for every 16*16 pixels. Without repeated encoding, the number of second codes generated by encoding the data to be processed is 64; another example is that the data to be processed is 44.1KHz audio, where the audio has 44100 sampling points per second. If the audio information of every 441 sampling points is compressed to generate an audio code, without repeated encoding, the number of first codes generated by encoding one second of audio is 100.
[0067] In an embodiment of the present invention, a sample code corresponding to the data to be processed is determined based on a second code quantity N; wherein the sample code includes a preset start code, a preset end code, and a second code quantity of preset placeholder codes located between the preset start code and the preset end code. It is understandable that for each type of data to be processed, a preset start code is added to the start position of the data to be processed, a preset end code is added to the end position, and according to the second code quantity generated by encoding each type of data to be processed, the data to be processed is converted into a second code quantity N of preset placeholder codes, thereby generating a sample code corresponding to the data to be processed. For example, if the data to be processed is a picture, the preset placeholder code can be<img_1> 、<img_feat> It should be noted that for different types of data to be processed, the preset start code, preset end code and preset placeholder code in the corresponding sample code are different.
[0068] S240: Fusing the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code.
[0069] In an embodiment of the present invention, for each multimodal sample data group, the sample codes corresponding to each type of to-be-processed data involved in the multimodal sample data group are fused with the reference code based on a preset code fusion strategy to generate a data group code. It will be appreciated that at least two data group codes can be generated in this manner, where the number of data group codes is the same as the number of multimodal sample data groups.
[0070] S250. Train a preset large language model based on at least two data group codes to generate a multimodal model; wherein, in the process of training the preset large language model based on at least two data group codes, replace the preset placeholder code between the preset start code and the preset end code corresponding to the data group code with the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model.
[0071] In an embodiment of the present invention, at least two data group codes are input into a preset large language model, and the preset large language model is trained based on the at least two data group codes until a preset convergence condition is reached, thereby generating a multimodal model. During the training of the preset large language model based on the at least two data group codes, the multimodal encoder in the large language model continuously generates actual sample codes corresponding to the data to be processed. Since the large language model is an autoregressive model, its characteristics dictate that in subsequent model training steps, all previous input codes and model inference generation codes are input as input codes to perform model training. Therefore, in an embodiment of the present invention, the actual sample codes corresponding to the data to be processed, generated by the multimodal encoder in the large language model, replace the preset placeholder codes between the corresponding preset start code and the preset end code in the data group codes. This avoids the influence of preset placeholder codes generated by model training at other locations in the data group codes, thereby improving the accuracy of the replacement of the preset placeholder codes. Exemplarily, if the data to be processed includes image data and audio data, the image modal code (i.e., actual sample code) corresponding to the image data generated by the image encoder in the large language model and the audio modal code (i.e., actual sample code) corresponding to the audio data generated by the audio encoder in the multimodal model are obtained, and then the preset image placeholder code between the preset image start code and the preset image end code corresponding to the image data in the data group code is replaced based on the image modal code, and the preset audio placeholder code between the preset audio start code and the preset audio end code corresponding to the audio data in the data group code is replaced based on the audio modal code.
[0072] The multimodal model training method of the embodiment of the present invention obtains at least two multimodal sample data groups; wherein each of the multimodal sample data groups contains sample data of at least two modalities; for each of the multimodal sample data groups, the sample data of any modality in the multimodal sample data group is used as reference data, and the reference code corresponding to the reference data is determined; the other sample data in the multimodal sample data group except the reference data is used as the data to be processed, and the second code quantity generated by encoding the data to be processed of each modality is determined respectively, and the sample code corresponding to the data to be processed is determined based on the second code quantity; wherein the sample code includes a preset starting code, a preset end code and the second code number of preset placeholder codes located between the preset start code and the preset end code; respectively fusing the sample code of each modality involved in each of the multimodal sample data groups with the reference code to generate a data group code; training a preset large language model based on at least two data group codes to generate a multimodal model; wherein, in the process of training the preset large language model based on at least two data group codes, the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the data group code. The technical solution provided by the embodiment of the present invention adds a preset start code and a preset end code corresponding to each type of data to be processed to the data group code. Therefore, during the training of the large language model, the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model only replaces the preset placeholder code between the corresponding preset start code and the preset end code in the data group code. This can effectively avoid the influence of the preset placeholder codes generated by model training at other positions in the data group code, improve the accuracy of the code alignment of the sample data of each modality during the multimodal model training process, and thus help to improve the accuracy of subsequent reasoning based on the multimodal model.
[0073] Figure 3 A schematic structural diagram of a multimodal model inference device provided by an embodiment of the present invention.
[0074] like Figure 3 As shown, the device includes:
[0075] A reference code determination module 310 is configured to obtain input data of at least two modalities, and use any one of the input data of the at least two modalities as reference data to determine a reference code corresponding to the reference data;
[0076] The modal code determination module 320 is configured to use the other modal input data, excluding the reference data, from the at least two modal input data as the data to be aligned, determine the number of first codes generated by encoding each type of the data to be aligned, and determine the modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code;
[0077] The inference result generation module 330 is used to fuse each of the modal codes with the reference code to generate a fused code, and input the fused code into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
[0078] Optional, inference result generation module, including:
[0079] a first core intermediate result determining unit, configured to input the fusion code into a pre-trained multimodal model and determine a plurality of first core intermediate results output by the multimodal model;
[0080] a second core intermediate result determining unit, configured to feed back the plurality of first core intermediate results to an input end of the multimodal model, and determine a plurality of second core intermediate results output by the multimodal model;
[0081] An inference result acquisition unit is used to obtain an inference result corresponding to the fusion encoding output by the multimodal model based on the multiple second core intermediate results, or at least one first core intermediate result and at least one second core intermediate result.
[0082] Optionally, also include:
[0083] a code alignment unit, configured to align the plurality of first core intermediate results with a reference code before feeding the plurality of first core intermediate results back to an input end of the multimodal model;
[0084] The second core intermediate result determination unit is configured to:
[0085] The aligned reference codes and the plurality of first core intermediate results are fed back to the input end of the multimodal model to determine a plurality of second core intermediate results output by the multimodal model.
[0086] Optionally, also include:
[0087] a target code deletion unit, configured to determine target codes in the plurality of first core intermediate results and delete the target codes from the plurality of first core intermediate results before feeding the plurality of first core intermediate results back to the input end of the multimodal model; wherein the target codes include start and end codes and / or placeholder codes generated by the multimodal model in the process of determining the first core intermediate results; and the start and end codes include the preset start code and the preset end code;
[0088] The second core intermediate result determination unit includes:
[0089] The second core intermediate result determination subunit is used to feed back the first core intermediate result after deleting the target code to the input end of the multimodal model, and determine multiple second core intermediate results output by the multimodal model.
[0090] Optional, target code removal unit, used to:
[0091] determining a portion of first core intermediate results from the plurality of first core intermediate results, and deleting the target code from the portion of first core intermediate results; wherein the portion of first core intermediate results includes the target code;
[0092] The second core intermediate result determination subunit is configured to:
[0093] The portion of the first core intermediate result after deleting the target code is fed back to the input end of the multimodal model, so that the multimodal model outputs subsequent multiple second core intermediate results based on the portion of the first core intermediate result after deleting the target code.
[0094] Optionally, also include:
[0095] a preset placeholder code replacing unit, configured to determine, before feeding back the plurality of first core intermediate results to the input end of the multimodal model, the actual modal code generated by the multimodal encoder in the multimodal model during the generation of the plurality of first core intermediate results, and replace the preset placeholder code in the fusion code based on the actual modal code;
[0096] The second core intermediate result determination unit is configured to:
[0097] The multiple first core intermediate results are fed back to the input end of the multimodal model so that the multimodal model determines the second core intermediate result output by the multimodal model based on the multiple first core intermediate results and the fusion code after the preset placeholder code is replaced by the actual modality code.
[0098] Optionally, the benchmark data is text data.
[0099] The multimodal model inference device provided in the embodiment of the present invention can execute the multimodal model inference method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0100] Figure 4 A schematic structural diagram of a multimodal model inference device provided by an embodiment of the present invention.
[0101] like Figure 4 As shown, the device includes:
[0102] The multimodal sample data group acquisition module 410 is configured to acquire at least two multimodal sample data groups, wherein each of the multimodal sample data groups includes sample data of at least two modalities;
[0103] A reference code determination module 420 is configured to determine, for each of the multimodal sample data groups, a reference code corresponding to the reference data using sample data of any modality in the multimodal sample data group as reference data;
[0104] The sample code determination module 430 is configured to use the sample data other than the reference data in the multimodal sample data group as data to be processed, determine the number of second codes generated by encoding the data to be processed in each modality, and determine a sample code corresponding to the data to be processed based on the second number of codes; wherein the sample code includes a preset start code, a preset end code, and the second number of preset placeholder codes located between the preset start code and the preset end code;
[0105] a data group code generating module 440 for fusing the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code;
[0106] A multimodal model generation module 450 is configured to train a preset large language model based on at least two data group codes to generate a multimodal model. During the training of the preset large language model based on at least two data group codes, the preset placeholder code between the preset start code and the preset end code corresponding to the data group code is replaced by the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model.
[0107] The multimodal model training device provided in the embodiment of the present invention can execute the multimodal model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0108] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0109] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0110] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0111] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the multimodal model inference method or the multimodal model training method.
[0112] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.
[0113] In some embodiments, the multimodal model reasoning method or the multimodal model training method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the multimodal model reasoning method or the multimodal model training method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the multimodal model reasoning method or the multimodal model training method in any other appropriate manner (for example, by means of firmware).
[0114] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0118] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0119] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0120] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0121] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A multimodal model reasoning method, characterized in that: include: Acquire input data of at least two modalities, and use any one of the at least two modalities of input data as reference data, and determine a reference code corresponding to the reference data; Using other modal input data except the reference data among at least two modal input data as data to be aligned, determining a first number of codes generated by encoding each type of the data to be aligned, and determining a modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code; Each modal code is fused with the reference code to generate a fused code, and the fused code is input into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
2. The method according to claim 1, characterized in that Inputting the fusion code into a pre-trained multimodal model and obtaining an inference result corresponding to the fusion code output by the multimodal model includes: Inputting the fusion code into a pre-trained multimodal model to determine a plurality of first core intermediate results output by the multimodal model; Feeding back the plurality of first core intermediate results to an input end of the multimodal model, and determining a plurality of second core intermediate results output by the multimodal model; According to the multiple second core intermediate results, or at least one first core intermediate result and at least one second core intermediate result, an inference result corresponding to the fusion encoding output by the multimodal model is obtained.
3. The method according to claim 2, characterized in that Before feeding back the plurality of first core intermediate results to the input end of the multimodal model, the method further includes: aligning a plurality of the first core intermediate results with a reference encoding; Feeding back the plurality of first core intermediate results to the input end of the multimodal model, and determining a plurality of second core intermediate results output by the multimodal model, comprises: The aligned reference codes and the plurality of first core intermediate results are fed back to the input end of the multimodal model to determine a plurality of second core intermediate results output by the multimodal model.
4. The method according to claim 2, characterized in that Before feeding back the plurality of first core intermediate results to the input end of the multimodal model, the method further includes: Determine target codes in a plurality of the first core intermediate results, and delete the target codes from the plurality of the first core intermediate results; wherein the target codes include start and end codes and / or placeholder codes generated by the multimodal model in the process of determining the first core intermediate results; the start and end codes include the preset start code and the preset end code; Feeding back the plurality of first core intermediate results to the input end of the multimodal model, and determining a plurality of second core intermediate results output by the multimodal model, comprises: The first core intermediate result after deleting the target code is fed back to the input end of the multimodal model to determine multiple second core intermediate results output by the multimodal model.
5. The method according to claim 4, characterized in that Deleting the target encoding from the plurality of first core intermediate results comprises: determining a portion of first core intermediate results from the plurality of first core intermediate results, and deleting the target code from the portion of first core intermediate results; wherein the portion of first core intermediate results includes the target code; Feeding back the first core intermediate result after deleting the target code to the input end of the multimodal model, and determining a second core intermediate result output by the multimodal model, including: The portion of the first core intermediate result after deleting the target code is fed back to the input end of the multimodal model, so that the multimodal model outputs subsequent multiple second core intermediate results based on the portion of the first core intermediate result after deleting the target code.
6. The method according to any one of claims 2 to 5, characterized in that: Before feeding back the plurality of first core intermediate results to the input end of the multimodal model, the method further includes: Determining the actual modal code generated by the multimodal encoder in the multimodal model during the generation of the plurality of first core intermediate results, and replacing the preset placeholder code in the fusion code based on the actual modal code; Feeding back the plurality of first core intermediate results to the input end of the multimodal model, and determining a plurality of second core intermediate results output by the multimodal model, comprises: The multiple first core intermediate results are fed back to the input end of the multimodal model so that the multimodal model determines the second core intermediate result output by the multimodal model based on the multiple first core intermediate results and the fusion code after the preset placeholder code is replaced by the actual modality code.
7. The method according to any one of claims 1 to 5, characterized in that: The reference data is text data.
8. A multimodal model training method, characterized in that: include: Acquire at least two multimodal sample data groups; wherein each of the multimodal sample data groups contains sample data of at least two modalities; For each of the multimodal sample data groups, taking sample data of any modality in the multimodal sample data group as reference data, determining a reference code corresponding to the reference data; Using the sample data other than the reference data in the multimodal sample data group as data to be processed, determining the number of second codes generated by encoding the data to be processed in each modality, and determining a sample code corresponding to the data to be processed based on the second number of codes; wherein the sample code includes a preset start code, a preset end code, and the second number of preset placeholder codes located between the preset start code and the preset end code; fusing the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code; A preset large language model is trained based on at least two data group codes to generate a multimodal model; wherein, in the process of training the preset large language model based on at least two data group codes, the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model replaces the preset placeholder code between the corresponding preset start code and the preset end code in the data group code.
9. A multimodal model inference device, characterized in that: include: a reference code determination module, configured to obtain input data of at least two modalities, and use any one of the input data of the at least two modalities as reference data to determine a reference code corresponding to the reference data; a modal code determination module, configured to use other modal input data, excluding the reference data, from among at least two modal input data as data to be aligned, determine a first number of codes generated by encoding each type of the data to be aligned, and determine a modal code corresponding to the data to be aligned based on the first number of codes; wherein the modal code includes a preset start code, a preset end code, and the first number of preset placeholder codes located between the preset start code and the preset end code; An inference result generation module is used to fuse each of the modal codes with the reference code to generate a fused code, and input the fused code into a pre-trained multimodal model to obtain an inference result corresponding to the fused code output by the multimodal model; wherein, in the process of obtaining the inference result based on the multimodal model, the actual modal code corresponding to the data to be aligned generated by the multimodal encoder in the multimodal model replaces the preset placeholder code between the preset start code and the preset end code corresponding to the fused code.
10. A multimodal model training device, characterized in that: include: A multimodal sample data group acquisition module, configured to acquire at least two multimodal sample data groups; wherein each of the multimodal sample data groups contains sample data of at least two modalities; a reference code determination module, configured to determine, for each of the multimodal sample data groups, a reference code corresponding to the reference data, using sample data of any modality in the multimodal sample data group as reference data; a sample code determination module, configured to treat the sample data other than the reference data in the multimodal sample data group as to-be-processed data, determine the number of second codes generated by encoding the to-be-processed data of each modality, and determine a sample code corresponding to the to-be-processed data based on the second number of codes; wherein the sample code includes a preset start code, a preset end code, and the second number of preset placeholder codes located between the preset start code and the preset end code; a data group code generation module, configured to fuse the sample code of each modality involved in each multimodal sample data group with the reference code to generate a data group code; A multimodal model generation module is configured to train a preset large language model based on at least two data group codes to generate a multimodal model; wherein, during the training of the preset large language model based on at least two data group codes, the preset placeholder code between the preset start code and the preset end code corresponding to the data group code is replaced by the actual sample code corresponding to the data to be processed generated by the multimodal encoder in the large language model.