Model training method, model reasoning method, electronic device and storage medium

By adopting a training method of mixed unidirectional and bidirectional parallel decoding in a unified multimodal large model, the problem of low model training efficiency in the existing technology is solved, and more efficient training and inference effects are achieved.

CN118586525BActive Publication Date: 2025-05-13SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410912446.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-05-13
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

In the prior art, the unified multimodal large model leads to very low model training through one-way autoregressive inference training method that decodes multimodal word sequences one by one.

Method used

By determining the first word segmentation to be trained with one-way decoding and the second word segmentation to be trained with two-way decoding based on the respective word segmentation of the visual mode and the language mode, the initial unified multimodal model is trained to mix unidirectional and bidirectional parallel decoding until the preset training stop condition is met.

Benefits of technology

The training efficiency of each modal word element segment is improved, and the word element generation effect of the predicted word element obtained after decoding training of each modal word element segment is improved. This greatly improves the inference efficiency of the unified multimodal big model while maintaining the model's prediction effect on different modal word elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118586525B_ABST
    Figure CN118586525B_ABST
Patent Text Reader

Abstract

The present invention provides a model training method, a model reasoning method, an electronic device and a storage medium, wherein the model training method comprises: determining a first word segmentation to be trained for unidirectional decoding and a second word segmentation to be trained for bidirectional decoding based on word segmentation of each visual modality and language modality; training an initial unified multimodal large model for mixed unidirectional and bidirectional parallel decoding based on the first word segmentation and the second word segmentation and the modality identifiers carried by each; until determining that the training result meets the preset training stop condition, the corresponding target unified multimodal large model. The present invention not only improves the training efficiency of each modal word segmentation, but also improves the word generation effect of the predicted word obtained after decoding training of each modal word segmentation, thereby maintaining the model prediction effect on different modal word segments while also improving the reasoning efficiency of the unified multimodal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method, a model reasoning method, an electronic device and a storage medium. Background Art

[0002] In order to perform word-gram inference on multimodal information such as video, image and language (such as text), a unified multimodal large model came into being. This unified multimodal large model can mix word-gram sequences of different modalities and use special words (such as <language>, <vision>, etc.) to distinguish word-gram sequences of different modalities; and in order to ensure that the model inference stage is more efficient and accurate, this unified multimodal large model can be trained to enable it to have the function of inferring and decoding mixed word-gram sequences corresponding to multimodal information; therefore, how to train a multimodal large model has become a key issue that needs to be solved urgently.

[0003] In the related art, multimodal information is usually mixed and input into a unified multimodal large model in the form of word units, and then the unidirectional autoregressive decoding strategy of the large language model is used to predict the next word unit of the mixed multimodal word unit sequence one by one. However, since the word unit sequence of the visual modality (such as multiple video frame images of a video) usually contains a large number of word units, the unidirectional autoregressive reasoning of decoding multimodal word units one by one will result in very low word unit prediction efficiency, thereby resulting in very low model training efficiency. Summary of the invention

[0004] The present invention provides a model training method, a model reasoning method, an electronic device and a storage medium, which are used to solve the defect of low model training efficiency caused by the unidirectional autoregressive reasoning training method of decoding multimodal word-gram sequences word by word in the prior art. Not only the training efficiency of each modal word-gram segment is improved, but also the word-gram generation effect of the predicted word-gram obtained after decoding training of each modal word-gram segment is improved, so as to maintain the model prediction effect on different modal word-grams while improving the reasoning efficiency of the unified multimodal large model.

[0005] The present invention provides a model training method, comprising the following steps.

[0006] Based on the word segmentation of each visual modality and language modality, the first word segmentation to be trained for unidirectional decoding and the second word segmentation to be trained for bidirectional decoding are determined; based on the first word segmentation and the second word segmentation and the modality identifiers they carry, the initial unified multimodal large model is trained for mixed unidirectional and bidirectional parallel decoding; until it is determined that the training result meets the preset training stop condition and the corresponding target unified multimodal large model is obtained.

[0007] According to a model training method provided by the present invention, the initial unified multimodal large model is trained with mixed unidirectional and bidirectional parallel decoding based on the first word segmentation and the second word segmentation and the modal identifiers carried by each, including: based on the preset attention mask corresponding to the first modality of the first word segmentation and the modal identifier representing the first modality, the unified multimodal large model is trained with unidirectional autoregressive decoding and unidirectional Jacobi decoding; based on the preset random mask corresponding to the second modality of the second word segmentation and the modal identifier representing the second modality, the unified multimodal large model is trained with bidirectional random demasking.

[0008] According to a model training method provided by the present invention, the process of determining the preset attention mask and the preset random mask includes: when the first modality is the language modality, determining the causal mask of the first word segment based on a preset self-attention mechanism, and determining the causal mask as the preset attention mask; when the second modality is the visual modality, performing mask learning on the full mask segment corresponding to the second word segment based on a preset random mask learning strategy to obtain the preset random mask.

[0009] According to a model training method provided by the present invention, the target unified multimodal large model corresponding to the training result satisfying the preset training stop condition is determined, including: determining the autoregressive loss function corresponding to the unidirectional autoregressive decoding training, the Jacobi decoding loss function corresponding to the unidirectional Jacobi decoding training, and the demasking loss function corresponding to the bidirectional random demasking training; based on the matching relationship between the autoregressive loss function, the preset autoregressive decoding target, the Jacobi decoding loss function, the preset Jacobi decoding target, the demasking loss function and the preset random demasking target, and the training results, respectively, determining the target unified multimodal large model.

[0010] According to a model training method provided by the present invention, the process of determining the word segmentation of each of the different modalities includes: inputting the multimodal original information into the word segmenter of the corresponding modality for word conversion processing to obtain the word segmentation of each of the different modalities; and there is an association relationship between adjacent modal original information in the multimodal original information.

[0011] The present invention provides a model reasoning method, comprising the following steps.

[0012] Determine the target word segments of different target modalities and the target unified multimodal large model trained by the aforementioned model training method; based on the target unified multimodal large model, perform parallel decoding and reasoning on each of the target word segments and each of the target modalities, and determine the model reasoning results that match the preset reasoning requirements based on the decoding and reasoning results.

[0013] According to a model inference method provided by the present invention, the target unified multimodal large model is used to perform parallel decoding inference on each target word segment, including: when the different target modalities include language modality and visual modality, the target unified multimodal large model is used to perform Jacobi decoding prediction of each target word segment and each target modality to obtain multiple predicted word units; based on the target unified multimodal large model, the multiple predicted word units and each target word segment and each target modality, a preset initial full mask sequence is subjected to demasked decoding prediction of bidirectional multi-step reasoning.

[0014] According to a model reasoning method provided by the present invention, the method also includes: when the demasked decoding prediction result does not meet the preset reasoning end condition, performing unidirectional multi-step reasoning Jacobi decoding prediction on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment and each target modality; or, based on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment, each target modality and the target unified multimodal large model, performing bidirectional multi-step reasoning demasked decoding prediction on the initial fully masked sequence.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements any one of the above-described model training methods or any one of the above-described model inference methods.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described model training methods or any of the above-described model inference methods.

[0017] The present invention provides a model training method, a model reasoning method, an electronic device and a storage medium, wherein the model training method performs mixed unidirectional and bidirectional parallel decoding training on an initial unified multimodal large model based on the first word segment to be trained for unidirectional decoding and the second word segment to be trained for bidirectional decoding in the word segment of each of the visual modality and the language modality and the modality identifiers carried by each; until it is determined that the training result meets the preset training stop condition and the corresponding target unified multimodal large model. In this way, the unified multimodal large model is trained by hybrid unidirectional and bidirectional parallel decoding through multimodal word segmentation, ensuring that the unified multimodal large model has unidirectional decoding reasoning ability and bidirectional demasking reasoning ability after training, avoiding the defect of low word prediction efficiency caused by only unidirectional autoregressive decoding training of the existing unified multimodal large model. It not only improves the training efficiency of each modal word segmentation, but also improves the word generation effect of the predicted word obtained after decoding training of each modal word segmentation. Therefore, while maintaining the model's prediction effect on different modal word units, the reasoning efficiency of the unified multimodal large model is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is one of the flow charts of the model training method provided by the present invention.

[0020] Figure 2 This is the second flow chart of the model training method provided by the present invention.

[0021] Figure 3 It is one of the flow charts of the model reasoning method provided by the present invention.

[0022] Figure 4 This is the second flow chart of the model reasoning method provided by the present invention.

[0023] Figure 5 It is a structural schematic diagram of the model training device provided by the present invention.

[0024] Figure 6 It is a structural schematic diagram of the model reasoning device provided by the present invention.

[0025] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] In the embodiments of the present invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. In the textual description of the present invention, the character " / " generally indicates that the previous and next associated objects are in an "or" relationship. In addition, it should be noted that the serial numbers themselves, such as "first", "second", etc., used to distinguish the objects described in the present invention are only used to distinguish the objects described, and do not have any order or technical meaning.

[0028] With the development of large language models, their model structure and computational optimization methods have gradually matured; this is because of the great success of large language models, so artificial intelligence (AI) models in other fields (such as vision) have gradually applied some of the concepts used by large language models, such as using the token representation of large language models. In large language models, a token can be part of a word (such as "building" can be composed of two tokens "build" and "ing"); similarly, for visual modalities, a token can represent a certain combination of several pixels (such as the operation of converting text or pixels into tokens is generated by the tokenizer).

[0029] In order to perform word-gram inference on multimodal information such as video, image, and language (such as text), a unified multimodal large model of language model structure came into being. Tokens of different modalities (such as language and vision) can be mixed and input into the unified multimodal large model, and then the autoregressive decoding used by the large language model is performed to infer the next token. When all tokens are generated, they are converted into the original media modality (such as text or image, etc.) through detokenization. This unified multimodal large model can mix word-gram sequences of different modalities and use special words (such as <language>, <vision>, etc.) to distinguish word-gram sequences of different modalities; and in order to ensure that the model inference stage is more efficient and accurate, this unified multimodal large model can be trained to enable it to have the function of inferring and decoding mixed word-gram sequences corresponding to multimodal information; therefore, how to train a multimodal large model has become a key issue that needs to be solved urgently.

[0030] In the related art, multimodal information is usually mixed and input into a unified multimodal large model in the form of word-grams, and then the unidirectional autoregressive decoding strategy of the large language model is used to predict the next word-gram of the mixed multimodal word-gram sequence one by one, such as performing unidirectional autoregressive decoding training in a teacher-forcing manner of the large language model. However, since word-gram sequences of visual modalities (such as multiple video frame images of a video) usually contain a large number of word-grams, unidirectional autoregressive reasoning for decoding multimodal word-grams one by one will result in very low word-gram prediction efficiency, thereby resulting in very low model training efficiency.

[0031] Combine the following Figure 1-Figure 7The model training method, model reasoning method, electronic device and storage medium of the present invention are described, wherein the execution subject of the model training method can be an electronic device or storage medium deployed with a unified multimodal large model, the unified multimodal large model can also be called a multimodal model, and one of the English full names of the multimodal model can be Mixed-Modal Early-Fusion Foundation Models; the unified multimodal large model (that is, the multimodal model) can understand and generate any arbitrary sequence of images and texts. In addition, the electronic device can be a personal computer (PC), a portable device, a laptop, a smart phone, a tablet computer, a portable wearable device and other devices; the server can refer to a server, or a server cluster composed of multiple servers, a cloud server, etc. The present invention does not specifically limit the specific form of the electronic device or server. Furthermore, the execution subject of the model training method can also be applied to a model training device set in an electronic device or a server, and the model training device can be implemented by software, hardware or a combination of the two. The following describes the model training method by taking the execution subject of the model training method as an electronic device deployed with a unified multimodal large model as an example.

[0032] In order to facilitate understanding of the model training method provided by the embodiment of the present invention, the model training method provided by the present invention will be described in detail below through the following exemplary embodiments. It is understandable that the following exemplary embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0033] Reference Figure 1 , which is one of the flow charts of the model training method provided by the present invention, such as Figure 1 As shown, the model training analysis method includes the following steps 110 and 120.

[0034] Step 110 : Determine a first word segment to be trained for unidirectional decoding and a second word segment to be trained for bidirectional decoding based on the word segmentations of the visual modality and the language modality.

[0035] Among them, each word-gram segment includes a modal identifier that identifies the modality of the corresponding word-gram segment. The modal identifier can be located in the first word segment of the corresponding word-gram segment in the form of a special word-gram, such as the modal identifier can be <language>, <vision>, etc.

[0036] Word-gram segmentation of the visual modality may include but is not limited to image word-gram segmentation and video word-gram segmentation, etc., and word-gram segmentation of the language modality may include but is not limited to text word-gram segmentation and speech text word-gram segmentation.

[0037] The number of the first word-unit segments may be one or more, and the number of the second word-unit segments may be one or more, which is not specifically limited in the present invention.

[0038] At least one first word-gram segment is a word-gram segment of a language modality. For example, each first word-gram segment can be a text word-gram segment or a speech-text word-gram segment.

[0039] At least one second word-gram segmentation is a word-gram segmentation of a visual modality. For example, each second word-gram segmentation can be an image word-gram segmentation and a video word-gram segmentation.

[0040] Specifically, the unified multimodal large model built into the electronic device first segments the word segments of different input modalities according to the modality identifier carried in each word segment, and takes the word segments of other visual modalities such as image word segmentation and video word segmentation in all word segmentations as the first word segmentation, and each first word segmentation is used for one-way decoding training; at the same time, the word segmentations of other language modalities such as text word segmentation and audio word segmentation in all word segmentations are taken as the second word segmentation.

[0041] Step 120: Based on the first word-unit segmentation and the second word-unit segmentation and the modal identifiers carried by each, the initial unified multimodal large model is trained with mixed unidirectional and bidirectional parallel decoding; until it is determined that the training result meets the preset training stop condition and the corresponding target unified multimodal large model is obtained.

[0042] Specifically, when all first word segments and all second word segments are mixed and input into the initial unified multimodal large model, the unified multimodal large model can identify the modality and training method of the corresponding word segment based on each modal identifier. For example, based on the modal identifier of <language> carried by each first word segment, it is determined that the training method of the corresponding first word segment is unidirectional decoding training; for another example, based on the modal identifier of <vision> carried by each second word segment, it is determined that the training method of the corresponding second word segment is bidirectional decoding training.

[0043] In addition, in addition to indicating the modality to which the corresponding first word segmentation or the corresponding second word segmentation belongs, each modal identifier can also carry the target word information required for the corresponding first word segmentation or the corresponding second word segmentation after the end of this training. Each target word information can be one of the target number of word segments obtained by the text word segmentation after the unidirectional decoding training, the target image size and the target number of word segments contained in the target image obtained by the image word segmentation after the bidirectional decoding training, and the target image size, target number of images and the target number of word segments contained in each target image obtained by the video word segmentation after the bidirectional decoding training.

[0044] For example, when the target word information is When , it can be shown that after this one-way decoding training, the prediction tokens; when the target word unit information is 384×384, it can indicate that after this bidirectional decoding training, an image with a size of 384×384 and a number of pixels of 384×384 needs to be predicted; when the target word unit information is 16×384×384, it can indicate that after this bidirectional training, 16 video frame images with a size of 384×384 need to be predicted and each video frame image contains 384×384 pixels.

[0045] At this point, based on the modal identifiers carried by all first word segments and all second word segments, the unified multimodal large model is trained with mixed unidirectional and bidirectional parallel decoding until the cumulative number of iterations reaches the preset number of iterations, or the loss values ​​of the unidirectional training loss function and the bidirectional training loss function respectively meet the corresponding loss thresholds, then the training is stopped, and the target unified multimodal large model corresponding to the time when the iteration stops is obtained.

[0046] The model training method provided by the embodiment of the present invention performs mixed unidirectional and bidirectional parallel decoding training on the initial unified multimodal large model based on the first word segment to be trained for unidirectional decoding and the second word segment to be trained for bidirectional decoding in the word segment of each visual modality and language modality and the modality identifiers carried by each; until the target unified multimodal large model corresponding to the training result is determined to meet the preset training stop condition. In this way, the training method of mixed unidirectional and bidirectional parallel decoding of the unified multimodal large model by multimodal word segmentation ensures that the unified multimodal large model has unidirectional decoding reasoning ability and bidirectional demasking reasoning ability after training, avoiding the defect of low word prediction efficiency caused by only unidirectional autoregressive decoding training of the existing unified multimodal large model, not only improving the training efficiency of each modal word segment, but also improving the word generation effect of the predicted word obtained after decoding training of each modal word segment, thereby maintaining the model prediction effect on different modal word segments while also improving the reasoning efficiency of the unified multimodal large model.

[0047] Based on the above Figure 1 In the model training method shown, in an exemplary embodiment, the process of determining the word-gram segments of different modalities in step 110 includes the following steps.

[0048] The multimodal original information is input into the word segmenter of the corresponding modality for word unit conversion processing to obtain word unit segmentations of different modalities; there is a correlation relationship between adjacent modal original information in the multimodal original information.

[0049] Specifically, refer to Figure 2 The second flowchart of the model training method shown in FIG. Figure 2 As shown, the multimodal original information includes visual modal information and language modal information, and the visual modal information includes image information containing two kittens and video information containing multiple video frame images, and the language modal information includes first text information containing "What's the content of this picture? Please help generate a related video", and second text information containing "There are two kittens in this picture. The video of two kittens playing with each other is as follows."

[0050] At this time, each language modality original information in the multimodal original information is respectively input into the language tokenizer of the corresponding language modality for token conversion processing, such as tokenization processing, to obtain token segmentation of multiple language modalities; at the same time, each visual modality original information in the multimodal original information is first compressed by a compression module, such as by a vector quantized variational autoencoder (VQ-VAE) encoder to perform pixel compression processing, and then input into the visual tokenizer corresponding to the visual modality for token conversion processing, such as tokenization processing and conversion into tokens; thereby obtaining token segmentation of different modalities.

[0051] It should be noted that, considering that the original information of the speech modal specifically presents the specified text content in the form of language, the original information of the speech modal can be first subjected to text content extraction, and then the extracted text content can be subjected to word-gram conversion processing on the text content contained therein, such as converting it into tokens after word segmentation (tokenization); considering that the word-gram segmentation obtained after the original information of the speech modal is processed is essentially the word-gram conversion of the text content contained therein, it can be determined that the word-gram segmentation of the language modal includes text word-gram segmentation and speech-text word-gram segmentation.

[0052] Based on the above Figure 1 In the model training method shown in FIG. 1 , in an exemplary embodiment, the specific process of step 120 includes:

[0053] Based on the preset attention mask corresponding to the first modality of the first word segment and the modal identifier representing the first modality, the unified multimodal large model is subjected to unidirectional autoregressive decoding training and unidirectional Jacobi decoding training; based on the preset random mask corresponding to the second modality of the second word segment and the modal identifier representing the second modality, the unified multimodal large model is subjected to bidirectional random demasking training.

[0054] For details, please refer to Figure 2 The model training method shown is Figure 2As shown, the word segments of different modalities include the language modality ~ this N 1 tokens, visual modality ~ this N 2 tokens, language modality ~ this N 3 tokens, and visual modality ~ this N 4 tokens, the above N 1 tokens and N 3 tokens and their respective first modalities, and N 2 tokens and N 4 The tokens and their respective second modalities are mixed as unified modal words and then enter the unified multimodal model for mixed unidirectional and bidirectional parallel decoding training, that is, the unified multimodal model is used to train the words containing +1 tokens preset attention mask for one-way autoregressive decoding training for next token prediction, N 1 tokens for one-way Jacobi decoding training, The preset random mask of tokens is used for bidirectional random demasking training, and the The preset attention mask of tokens is used to perform one-way autoregressive decoding training for the next token prediction, and the language modality Tokens are trained for one-way Jacobi decoding, and Bidirectional random demasking training is performed with the preset random mask of .

[0055] It should be noted that if Figure 2 As shown, the language modality After one-way autoregressive decoding training and one-way Jacobi decoding training, the language modality can be obtained. The prediction probability of the target prediction word , visual modality N 2 After bidirectional random masking training, the visual modality can be obtained. N2 The prediction probability of the target prediction word , language modality N 3 After one-way autoregressive decoding training and one-way Jacobi decoding training, the language modality can be obtained. N 3 The prediction probability of the target prediction word , and visual modality N 4 After bidirectional random masking training, the visual modality can be obtained. N 4 The prediction probability of the target prediction word In this way, the Mask TokenModel (MTM) in the field of visual models and the Jacobi parallel decoding method for accelerating autoregressive reasoning in the field of language models are combined, and both the masked reasoning and the parallel decoding reasoning can improve the computing power utilization by reducing the number of model reasoning times, which can ensure that the model trained to convergence has a large proportion of acceleration effect in both reasoning language-like word segmentation and reasoning visual-like word segmentation, thereby reducing bandwidth requirements and solving the problem that the existing model autoregressive reasoning is usually limited by bandwidth.

[0056] It should be noted that, for the word-gram prediction results output by each first word-gram segment after unidirectional decoding training of the unified multimodal large model or the word-gram prediction results output by each second word-gram segment after bidirectional decoding training of the unified multimodal large model, the word-gram prediction results can be first converted into latent space variables, and then the latent space variables are matrix multiplied with the dictionary size containing different word-gram prediction probabilities, thereby obtaining the probability of each target predicted word-gram in the word-gram prediction result.

[0057] In addition, for all first word-unit segments and all second word-unit segments, the following can be determined based on their respective modal identifiers and the number of segmented words they contain: Figure 2 The mixed modal attention mask shown is used to facilitate determining a preset attention mask corresponding to each first word-unit segment and a preset random mask corresponding to each second word-unit segment from the mixed modal attention mask.

[0058] Therefore, in an exemplary embodiment, for each first word-unit segment and each second word-unit segment, the preset attention mask corresponding to one of the first word-unit segments and the preset random mask corresponding to one of the second word-unit segments, the specific determination process includes the following steps.

[0059] When the first modality is a language modality, the causal mask of the first word segment is determined based on a preset self-attention mechanism, and the causal mask is determined as a preset attention mask; when the second modality is a visual modality, mask learning is performed on the full-mask segment corresponding to the second word segment based on a preset random mask learning strategy to obtain a preset random mask.

[0060] For details, please refer to Figure 2 , will be targeted at language modal N 1 tokens, for visual modality N 2 tokens, for language modality N 3 tokens and for the visual modality N 4 When tokens are input into the unified multimodal model, the corresponding mask can be determined for each modal token segment and then input into the multimodal model. That is, for the first word segment of the language modality, the causal mask of the first word segment can be determined by calculating the attention correlation between tokens using the pre-set self-attention mechanism. For example, this causal mask can be specifically Figure 2 The upper left corner of the mixed modality attention mask is 5×5 (at this time N 1 +1=5) in the upper triangular attention mask in the matrix, or Figure 2 The upper left corner of the mixed modality attention mask is 15×15 (at this time N 1 +1=5, N 2 +1=5, N 3 +1=5) upper triangular attention mask in the matrix, Figure 2 middle N 4 +1=10.

[0061] For the second word-unit segmentation of the visual modality, a preset random mask learning strategy can be used to perform mask learning on the full-mask segmentation corresponding to the second word-unit segmentation, so as to learn masks of different ratios and obtain a preset random mask corresponding to the second word-unit segmentation.

[0062] Exemplarily, the mask learning process can be determined using formula (1).

[0063] (1).

[0064] In formula (1), Indicates the fully masked segment corresponding to the second word segment M Middle mask elements, , Indicates the fully masked segment corresponding to the second word segment M The total number of mask elements in , or the total number of tokens contained in the second token segment, such as Specifically, it can be Figure 2 In N 1 or N 2 ; Indicated in domain, and satisfies , ;For example, or ; Represents a random variable and All of them are uniformly distributed in the range of (0,1); when When , it indicates that it is a non-masked position and has a masked value; when , it indicates that it is a mask position and corresponds to a predicted token.

[0065] At this point, the preset random mask shown in formula (2) can be determined by formula (1) .

[0066] (2).

[0067] In formula (2), Indicates Mask elements It is a mask value or a predicted token, such as a preset random mask. Specifically, it can be , b Represents the mask value and is usually the dictionary size of the VQ-VAE codebook; this is because all latent space tokens have values ​​between 0 and the dictionary size of 1, so the mask value is set to a value outside its range, such as the dictionary size.

[0068] In this way, by combining the unidirectional Jacobi trajectory objective and the bidirectional demasking objective in the model training stage, the reasoning efficiency of the unified multimodal large model can be greatly improved while maintaining the model's prediction effect on tokens of different modalities.

[0069] Based on the above Figure 1In the model training method shown in FIG. 1 , in an exemplary embodiment, the specific implementation process of step 120 may include the following steps.

[0070] Determine the autoregressive loss function corresponding to the unidirectional autoregressive decoding training, the Jacobi decoding loss function corresponding to the unidirectional Jacobi decoding training, and the demasking loss function corresponding to the bidirectional random demasking training; based on the matching relationship between the autoregressive loss function, the preset autoregressive decoding target, the Jacobi decoding loss function, the preset Jacobi decoding target, the demasking loss function and the preset random demasking target and the training results, respectively, determine the target unified multimodal large model.

[0071] Specifically, for a unified multimodal large model trained with mixed unidirectional and bidirectional parallel decoding, when its head module or fully connected layer outputs logits, the autoregressive loss function corresponding to the unidirectional autoregressive decoding training can be determined by referring to formula (3): , refer to formula (4) to determine the Jacobi decoding loss function corresponding to the one-way Jacobi decoding training , and refer to formula (5) to determine the demasking loss function corresponding to the bidirectional random demasking training .

[0072] (3).

[0073] (4).

[0074] (5).

[0075] In formulas (3) to (5), Represents the model parameters of the unified multimodal large model and includes the model parameters to be optimized for the one-way auto-regressive (AR) model and the MTM model; Represents the output probability of the AR model or the output probability of the MTM model, Before use Tokens prediction Tokens, represents the tokens decoded and inferred in the jth step in the kth Jacobi trajectory (including correctly ordered tokens and random tokens); n represents the number of tokens decoded and inferred in each step of a Jacobi trajectory, that is, the number of tokens in a Jacobi trajectory; m represents the total number of decoding and inference steps required for each Jacobi trajectory. represents all correctly ordered tokens predicted by the last step of reasoning in the k-th Jacobi trajectory, Indicates that the content in brackets will not participate in the model parameter gradient calculation through back propagation. Represents an unmasked tokens segment.

[0076] In addition, according to the autoregressive loss function And Jacobi decoding loss function , we can also get the total loss function of the first word segment of the language modality , as shown in formula (6).

[0077] (6).

[0078] In formula (6), Indicates the Jacobi decoding ratio.

[0079] It should be noted that regarding the autoregressive loss function , Jacobi decoding loss function And the demasking loss function , here Used for training one-way language tokens segments, It is used for training bidirectional visual tokens segments; for details, please refer to Table 1.

[0080] Table 1

[0081]

[0082] In Table 1, when the first word segment of the language mode is Y and it contains When there are N tokens, these N tokens can be divided into multiple sub-word segments of length n, where the first sub-word segment includes For Jacobi decoding inference, we can randomly select n random tokens as the initial state of random tokens corresponding to the first sub-word segmentation, that is, , after performing Jacobi decoding inference on the n random tokens in step 1, we get , after performing 2-step Jacobi decoding inference on the n random tokens, we get ; And so on, until the m-step Jacobi decoding inference is performed on the n random tokens to obtain . At this time, you can , ,……, As the Jacobi trajectory of the first sub-word segment mentioned above; traverse the other sub-word segments in N tokens according to the above steps; until the Jacobi trajectory of all sub-word segments in N tokens is obtained.

[0083] In combination with the above formulas (3) to (6) and the exemplary description of Table 1, it can be understood that after performing a training of mixed unidirectional and bidirectional parallel decoding or a preset number of trainings of mixed unidirectional and bidirectional parallel decoding, the training results can be compared with the loss values ​​of the loss functions corresponding to different training targets, that is, the function value of the autoregressive loss function can be compared with the preset autoregressive decoding target to determine whether the unidirectional autoregressive decoding training has converged, and the function value of the Jacobi decoding loss function can be compared with the preset Jacobi decoding target to determine whether the model has learned the Jacobi trajectory and whether the Jacobi decoding speed of the model has been accelerated; and the function value of the demasking loss function can be compared with the preset random demasking target to determine whether the bidirectional demasking training has converged; if the training results meet the above training targets, the training is stopped; otherwise, the training of mixed unidirectional and bidirectional parallel decoding is continued. Until the target unified multimodal large model corresponding to the preset training stop condition is obtained. In this way, by introducing the Jacobi decoding loss function term based on the autoregressive objective, it is ensured that the model can learn the Jacobi trajectory, thereby accelerating the Jacobi decoding speed of the model during inference (increasing the tokens acceptance rate).

[0084] Below, the model reasoning method provided by the present invention is described, wherein the execution subject of the model reasoning method is an electronic device or server deployed with a target unified multimodal large model trained according to the aforementioned model training method. In addition, the execution subject of the model reasoning method can also be applied to a model reasoning device provided in an electronic device or a server, and the model reasoning device can be implemented by software, hardware, or a combination of both. Below, the model reasoning method is described by taking the execution subject of the model reasoning method as an electronic device deployed with a target unified multimodal large model as an example.

[0085] Reference Figure 3 , which is one of the flow charts of the model reasoning method provided by the present invention, such as Figure 3 As shown, the model reasoning method includes the following steps 310 and 320.

[0086] Step 310: determine the target word segments of different target modalities and the target unified multimodal large model trained by the aforementioned model training method.

[0087] Step 320: Based on the target unified multimodal large model, parallel decoding and reasoning are performed on each target word segment and each target modality, and based on the decoding and reasoning results, a model reasoning result that matches the preset reasoning requirements is determined.

[0088] Among them, each target modality can be other modalities such as visual modality or language modality; the target word segments of adjacent target modalities can be associated with each other, and adjacent target modalities can be the same modality or different modalities.

[0089] Exemplarily, when the multiple target modalities are specifically language modality and visual modality, the decoding reasoning result of the language modality can summarize the decoding reasoning result of the visual modality in text form. For example, the decoding reasoning result of the language modality is specifically "There are two kittens in this picture, and the video of the two kittens playing with each other is as follows:", and the decoding reasoning result of the visual modality is specifically multiple video frame images of two kittens playing with each other, that is, a video of two kittens playing with each other.

[0090] Specifically, when the number of target modalities is at least 2, the target word-unit segmentation corresponding to the visual modality original information and the language modality original information can be determined. For example, the language modality original information can be processed by the corresponding language word segmenter to obtain the corresponding target word-unit segmentation, and the visual modality original information can be first compressed by the compression module and then processed by the corresponding visual word segmenter to obtain the corresponding target word-unit segmentation.

[0091] At this time, the target unified multimodal large model can be used to perform different parallel decoding reasoning on target word segments of different modalities. The multiple predicted tokens obtained each time will be merged to the right of the multiple target word segments that currently carry the corresponding target modalities as the new input of the target unified multimodal large model for decoding reasoning prediction again; until the decoding reasoning output result meets the preset reasoning requirements, the multiple word units that meet the preset reasoning requirements are reversed and converted, and then the reversed and converted results are converted to the original space where the corresponding target modality is located, so as to obtain the model reasoning result that matches the preset reasoning requirements, that is, the required multimodal input result is obtained, such as Figure 4 "There are two kittens in this picture. The following is a video of the two kittens playing with each other," and multiple video frame pictures of kittens playing with each other.

[0092] It should be noted that, for the de-word segmentation processing, the de-word segmentation device (de-tokenize) corresponding to the corresponding target modality can be used to perform the de-word segmentation conversion processing. For example, when the corresponding target modality is the visual modality, the visual de-word segmentation device is used to perform the de-word segmentation conversion processing; similarly, when the corresponding target modality is the language modality, the language de-word segmentation device is used to perform the de-word segmentation conversion processing.

[0093] The model inference method provided in the embodiment of the present invention improves the model inference efficiency and model prediction effect by using a target unified multimodal large model trained to convergence to infer language word segmentation and visual word segmentation.

[0094] Based on the above Figure 3 The model training method shown, in an exemplary embodiment, in step 320, parallel decoding reasoning is performed on each target word segment based on the target unified multimodal large model, including the following steps.

[0095] In the case of different target modalities including language modality and visual modality, the target unified multimodal large model is used to perform unidirectional multi-step reasoning Jacobi decoding prediction on each target word segment and each target modality to obtain multiple predicted word units; based on the target unified multimodal large model, multiple predicted word units and each target word segment and each target modality, a bidirectional multi-step reasoning demasked decoding prediction is performed on the preset initial full-mask sequence.

[0096] Among them, there is at least one target word segment for the language modality and at least one target word segment for the visual modality; and the target word segments of adjacent target modalities may be associated with each other, and the adjacent target modalities may be the same modality or different modalities.

[0097] Specifically, refer to Figure 4 The second flowchart of the model reasoning method shown in FIG. Figure 4 As shown in , for the original language modality information "What does this picture contain? Please help generate a related video", the target word segmentation of the corresponding generated language modality (such as <language>) can be this N 1 word units; and Figure 4 For an image with two kittens in , the target word segmentation corresponding to the generated visual modality (such as <vision>) can be this N 2 A word.

[0098] The above N 1 Term and N 2 The word units and their respective target modalities are mixed as unified modal word units and input into the target unified multimodal large model for Jacobi decoding prediction of unidirectional multi-step reasoning, that is, <language> <Visual> The number of tokens required for Jacobi decoding prediction N 3It can be pre-marked in the language mode, and because the language mode exists in a special token, it also participates in Jacobi decoding prediction, so it is necessary to repeat multiple steps of Jacobi decoding prediction and output N 3 +1 prediction token, which is the output this N 3 predicted word units and the special token <language>, which N 3 The predicted word unit can be processed by the language reverse segmenter for multi-modal output, that is, the output is "There are two kittens in this picture, and the video of two kittens playing with each other is as follows:".

[0099] At the same time, the N 3 The predicted word and its target modality are placed in the N 1 Term and N 2 The right side of each word and its target modality are used as new input and then input into the target unified multimodal model again, that is, <language> <Visual> <Language> As the new input of the target unified multimodal large model, the target unified multimodal large model performs bidirectional multi-step reasoning on the preset initial full-mask sequence to predict the number of tokens required for demasking decoding prediction N 4 It can be pre-marked in the visual modality, and because the visual modality exists in a special token, it also participates in the demasking decoding prediction. Therefore, it is necessary to repeat the demasking decoding prediction multiple times and output this N 4 predicted word units and the special token <Visual>; at this time N 4 The predicted words can be processed by the visual de-segmenter to obtain a video that meets the requirements, that is, a video of two kittens playing with each other.

[0100] Exemplarily, for the word-meta reasoning decoding of the target word-meta segment of the language modality (such as <language>): a segment of tokens is randomly selected from the text dictionary of the target unified multimodal large model, recorded as tokens segment A and input into the target unified multimodal large model to generate a segment of next tokens prediction probabilities corresponding to the segment of tokens. Afterwards, the tokens with the highest probability are selected from the segment of next tokens prediction probabilities as new inputs and input into the target unified multimodal large model; the above steps are repeated until the target unified multimodal large model outputs the tokens with the highest prediction probability that are consistent with the input tokens segment A, then the reasoning of the current target word-meta segment is accepted, and then the reasoning of the next segment of the target word-meta segment is performed. Here, each step of reasoning in each segment will accept at least one token. Considering that the target unified multimodal large model has been trained with the Jacobi decoding target, it can be largely ensured that multiple tokens are accepted each time.

[0101] For the target word segmentation in the visual modality (e.g., <vision>), the MTM model is based on The preset initial full mask sequence is used as the initial state to perform multi-step unmasking token prediction. , , Indicates the number of unmasking inferences of the MTM model, represents the average calculation, Indicated in A monotonically increasing function in the domain that satisfies , ; Indicates the start of demasking decoding. Indicates the end of demasking decoding. Represents the total number of demasking decoding steps. The number of demasking prediction tokens in each step of the MTM model is , for The index value of ;pass It is guaranteed that at least one token is decoded; by It can be ensured that the total number of decoded tokens is If a token that has already been predicted is encountered during the mask-free decoding prediction process, the position is skipped.

[0102] It is understandable that if the multiple predicted word units obtained by the demasked decoding prediction do not meet the preset reasoning end conditions, such as the number of predicted word units does not reach the word unit length threshold preset by the user, the preset end symbol does not appear, or the cumulative number of reasoning times does not reach the preset number threshold, etc., the new target word unit segment of the new target modality can be obtained again to perform different parallel decoding reasoning again. Based on this, the model reasoning method provided by the present invention can also include the following steps.

[0103] When the demasked decoding prediction result does not meet the preset reasoning end condition, based on the target unified multimodal large model, the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment and each target modality are subjected to unidirectional multi-step reasoning Jacobi decoding prediction; or, based on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment, each target modality and the target unified multimodal large model, the initial fully masked sequence is subjected to bidirectional multi-step reasoning demasked decoding prediction.

[0104] For example, continue to refer to Figure 4 , if the mask decoding prediction this N 4 When the predicted word unit and the special token <vision> do not meet the preset reasoning end condition, the reasoning prediction can continue, such as using the target unified multimodal large model to place the unmasked decoding prediction result in <language> <Visual> <Language> The new input B formed on the right side of continues to perform Jacobi decoding prediction of unidirectional multi-step reasoning. The new input B here can be specifically <language> <Visual> <Language> <Visual> Or, using the target unified multimodal large model, under the prompt of the above new input B, continue to perform bidirectional multi-step reasoning on the initial full mask sequence to perform mask-free decoding prediction until the preset reasoning end condition is met.

[0105] The model training device provided by the present invention is described below. The model training device described below and the model training method described above can be referenced to each other.

[0106] Reference Figure 5 , which is a schematic diagram of the structure of the model training device provided by the present invention, such as Figure 5 As shown, the model training device 500 includes: a word unit determination unit 510 and a model training unit 520.

[0107] The word-gram determining unit 510 is used to determine a first word-gram segment to be trained for unidirectional decoding and a second word-gram segment to be trained for bidirectional decoding based on the word-gram segmentations of the visual modality and the language modality.

[0108] The model training unit 520 is used to perform mixed unidirectional and bidirectional parallel decoding training on the initial unified multimodal large model based on the first word segmentation and the second word segmentation and the modal identifiers carried by each; until it is determined that the training result meets the preset training stop condition and the corresponding target unified multimodal large model.

[0109] Optionally, the word unit determination unit 510 is specifically used to input the multimodal original information into the word segmenter of the corresponding modality for word unit conversion processing to obtain word unit segmentations of different modalities; adjacent modal original information in the multimodal original information has an association relationship.

[0110] Optionally, the model training unit 520 is specifically used to perform unidirectional autoregressive decoding training and unidirectional Jacobi decoding training on the unified multimodal large model based on the preset attention mask corresponding to the first modality of the first word segment and the modality identifier representing the first modality; and to perform bidirectional random demasking training on the unified multimodal large model based on the preset random mask corresponding to the second modality of the second word segment and the modality identifier representing the second modality.

[0111] Optionally, the model training unit 520 is specifically used to determine the causal mask of the first word segment based on a preset self-attention mechanism when the first modality is a language modality, and determine the causal mask as a preset attention mask; when the second modality is a visual modality, perform mask learning on the full mask segment corresponding to the second word segment based on a preset random mask learning strategy to obtain a preset random mask.

[0112] Optionally, the model training unit 520 is specifically used to determine the autoregressive loss function corresponding to the unidirectional autoregressive decoding training, the Jacobi decoding loss function corresponding to the unidirectional Jacobi decoding training, and the demasking loss function corresponding to the bidirectional random demasking training; based on the matching relationship between the autoregressive loss function, the preset autoregressive decoding target, the Jacobi decoding loss function, the preset Jacobi decoding target, the demasking loss function and the preset random demasking target and the training results, respectively, the target unified multimodal large model is determined.

[0113] The model reasoning device provided by the present invention is described below. The model reasoning device described below and the model reasoning method described above can be referenced to each other.

[0114] Reference Figure 6 , which is a schematic diagram of the structure of the model reasoning device provided by the present invention, such as Figure 6As shown, the model reasoning device 600 includes: an information determination unit 610 and a model reasoning unit 620.

[0115] The information determination unit 610 is used to determine the target word segmentation of each target modality and the target unified multimodal large model trained by the aforementioned model training method.

[0116] The model reasoning unit 620 is used to perform parallel decoding reasoning on each target word segment and each target modality based on the target unified multimodal large model, and determine the model reasoning result that matches the preset reasoning requirement based on the decoding reasoning result.

[0117] Optionally, the model inference unit 620 is specifically used to perform unidirectional multi-step reasoning Jacobi decoding prediction on each target word segment and each target modality using the target unified multimodal large model in different target modalities including language modality and visual modality to obtain multiple predicted word elements; based on the target unified multimodal large model, multiple predicted word elements and each target word segment and each target modality, perform bidirectional multi-step reasoning de-masked decoding prediction on a preset initial full-mask sequence.

[0118] Optionally, the model inference unit 620 is specifically used to perform Jacobi decoding prediction with one-way multi-step reasoning on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment and each target modality based on the target unified multimodal large model when the demasked decoding prediction result does not meet the preset reasoning end condition; or, perform demasked decoding prediction with two-way multi-step reasoning on the initial full-masked sequence based on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment, each target modality and the target unified multimodal large model.

[0119] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the model training method, which includes: based on the word-gram segmentation of each visual modality and the language modality, determining the first word-gram segmentation to be trained for unidirectional decoding and the second word-gram segmentation to be trained for bidirectional decoding; based on the first word-gram segmentation and the second word-gram segmentation and the modality identifiers carried by each, training the initial unified multimodal large model for mixed unidirectional and bidirectional parallel decoding; until determining that the training result meets the preset training stop condition, the corresponding target unified multimodal large model.

[0120] Alternatively, a model reasoning method is executed, which includes: determining the target word segments of different target modalities and the target unified multimodal large model trained by the aforementioned model training method; based on the target unified multimodal large model, performing parallel decoding and reasoning on each target word segment and each target modality, and determining a model reasoning result that matches the preset reasoning requirement based on the decoding and reasoning result.

[0121] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0122] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model training method provided by the above-mentioned methods, which includes: based on the word segmentation of each visual modality and the language modality, determining the first word segmentation to be trained for unidirectional decoding and the second word segmentation to be trained for bidirectional decoding; based on the first word segmentation and the second word segmentation and the modality identifiers carried by each, training the initial unified multimodal large model for mixed unidirectional and bidirectional parallel decoding; until it is determined that the training result meets the preset training stop condition and the corresponding target unified multimodal large model.

[0123] Alternatively, a model reasoning method is executed, which includes: determining the target word segments of different target modalities and the target unified multimodal large model trained by the aforementioned model training method; based on the target unified multimodal large model, performing parallel decoding and reasoning on each target word segment and each target modality, and determining a model reasoning result that matches the preset reasoning requirement based on the decoding and reasoning result.

[0124] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the model training method provided by the above-mentioned methods, the method comprising: determining a first word segment to be trained for unidirectional decoding and a second word segment to be trained for bidirectional decoding based on the word segmentations of the visual modality and the language modality respectively; training an initial unified multimodal large model for mixed unidirectional and bidirectional parallel decoding based on the first word segmentation and the second word segmentation and the modality identifiers carried by each; until determining that the training result satisfies a preset training stop condition corresponding to the target unified multimodal large model.

[0125] Alternatively, a model reasoning method is executed, which includes: determining the target word segments of different target modalities and the target unified multimodal large model trained by the aforementioned model training method; based on the target unified multimodal large model, performing parallel decoding and reasoning on each target word segment and each target modality, and determining a model reasoning result that matches the preset reasoning requirement based on the decoding and reasoning result.

[0126] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model training method, characterized in that: include: Determining a first word segment to be trained for unidirectional decoding and a second word segment to be trained for bidirectional decoding based on the word segmentations of the visual modality and the language modality; Based on the first word-unit segmentation and the second word-unit segmentation and the modal identifiers carried by each, the initial unified multimodal large model is trained by hybrid unidirectional and bidirectional parallel decoding; until it is determined that the training result satisfies the preset training stop condition and the corresponding target unified multimodal large model is obtained; The training of hybrid unidirectional and bidirectional parallel decoding on the initial unified multimodal large model based on the first word-unit segmentation and the second word-unit segmentation and the modal identifiers carried by each of them includes: Based on the preset attention mask of the first modality corresponding to the first word segment and the modality identifier representing the first modality, performing unidirectional autoregressive decoding training and unidirectional Jacobi decoding training on the unified multimodal large model; Based on the preset random mask of the second modality corresponding to the second word segment and the modality identifier representing the second modality, bidirectional random demasking training is performed on the unified multimodal large model.

2. The model training method according to claim 1, characterized in that: The process of determining the preset attention mask and the preset random mask includes: In the case where the first modality is the language modality, determining a causal mask of the first word segment based on a preset self-attention mechanism, and determining the causal mask as the preset attention mask; When the second modality is the visual modality, mask learning is performed on the full-mask segment corresponding to the second word segment based on a preset random mask learning strategy to obtain the preset random mask.

3. The model training method according to claim 1 or 2, characterized in that: The target unified multimodal large model corresponding to when the training result satisfies the preset training stop condition includes: Determine an autoregressive loss function corresponding to the unidirectional autoregressive decoding training, a Jacobi decoding loss function corresponding to the unidirectional Jacobi decoding training, and a demasking loss function corresponding to the bidirectional random demasking training; Based on the matching relationship between the autoregressive loss function, the preset autoregressive decoding target, the Jacobi decoding loss function, the preset Jacobi decoding target, the demasking loss function and the preset random demasking target, respectively, and the training results, the target unified multimodal large model is determined.

4. The model training method according to claim 1 or 2, characterized in that: The process of determining the word-unit segmentation of each of the different modalities includes: The multimodal original information is input into the word segmenter of the corresponding modality for word unit conversion processing to obtain word unit segmentations of the different modalities; and adjacent modal original information in the multimodal original information has an association relationship.

5. A model reasoning method, characterized in that: include: Determine the target word segmentation of each of the different target modalities and the target unified multimodal large model trained by the model training method of any one of claims 1 to 4 above; Based on the target unified multimodal large model, parallel decoding and reasoning are performed on each of the target word segments and each of the target modalities, and based on the decoding and reasoning results, a model reasoning result that matches the preset reasoning requirements is determined.

6. The model reasoning method according to claim 5, characterized in that: The performing parallel decoding reasoning on each of the target word-unit segments based on the target unified multimodal large model includes: In the case where the different target modalities include language modality and visual modality, using the target unified multimodal large model to perform Jacobi decoding prediction of each target word unit segment and each target modality of the target modality to obtain multiple predicted word units; Based on the target unified multimodal large model, the multiple predicted word units and each of the target word unit segments and each of the target modalities, a de-masked decoding prediction of a preset initial full-mask sequence is performed by bidirectional multi-step reasoning.

7. The model reasoning method according to claim 6, characterized in that: The method further comprises: When the demasked decoding prediction result does not meet the preset reasoning end condition, based on the target unified multimodal large model, the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment and each target modality are subjected to unidirectional multi-step reasoning Jacobi decoding prediction; or, based on the demasked decoding prediction result, the visual modality corresponding to the demasked decoding prediction result, each target word segment, each target modality and the target unified multimodal large model, the initial fully masked sequence is subjected to bidirectional multi-step reasoning demasked decoding prediction.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the model training method as described in any one of claims 1 to 4, or the model reasoning method as described in any one of claims 5 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the model training method as described in any one of claims 1 to 4, or the model reasoning method as described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Cross-modal sponsored search method and system based on fine-grained alignment of VLP input end

    CN117609597A

  • Motion recognition multi-modal large model construction method fusing text and video space-time signals

    CN117612263A