Training method, device and electronic equipment for multimodal large model
By splicing and modality-labeling the training data of large multimodal models, combined with a unidirectional attention mechanism and a variable-dimensionality strategy, the problem of limited applicability of large multimodal models is solved, achieving wider task applicability and higher training accuracy.
Patent Information
- Application Number
- CN202411366471.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-09-27
AI Technical Summary
When existing large multimodal models handle multimodal understanding and generation tasks, network cascades are prone to cascade errors, and their applicable scenarios are limited, making them unsuitable for a variety of tasks.
By obtaining a sequence of sample data blocks in the training data, performing splicing processing and modality labeling, and combining the unidirectional attention mechanism and dimensionality change strategy, the multimodal large model is trained to expand its applicable scenarios.
It enables large multimodal models to be applicable to a variety of multimodal understanding and generation tasks, avoids cascading errors, and improves training accuracy and efficiency.
Smart Images

Figure CN119476386B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as deep learning, natural language processing, computer vision, speech technology, and large models, and in particular to a training method, device, and electronic equipment for a multimodal large model. Background Art
[0002] When processing multimodal understanding tasks, the current large multimodal model needs to connect the understanding network of the corresponding modality before the large multimodal model; when processing multimodal generation tasks, the generation network of the corresponding modality needs to be connected after the large multimodal model.
[0003] In the above scheme, the cascade between multiple networks is prone to cascade errors; and when the connected networks are fixed, it can only be applied to specific multimodal understanding tasks or multimodal generation tasks, and the applicable scenarios are limited. Summary of the Invention
[0004] The present disclosure provides a method, device, and electronic device for training a multimodal large model.
[0005] According to one aspect of the present disclosure, a method for training a large multimodal model is provided, the method comprising: obtaining training data and an initial large multimodal model; the training data comprising: a sequence of sample data blocks corresponding to sample data under at least two modalities; splicing the sequence of sample data blocks under the at least two modalities to obtain a sample splicing sequence; and training the large multimodal model in combination with the sample splicing sequence.
[0006] According to another aspect of the present disclosure, a method for processing a multimodal task is provided, the method comprising: obtaining a multimodal task; the multimodal task comprising data in at least one modality; data in a text modality in the at least one modality being used to indicate the need to generate data in a target modality; obtaining a multimodal large model; the multimodal large model being determined based on the training method for the multimodal large model as described above; and determining the generated data in the target modality in combination with the data in the at least one modality and the multimodal large model.
[0007] According to another aspect of the present disclosure, a training device for a multimodal large model is provided, the device comprising: an acquisition module for acquiring training data and an initial multimodal large model; the training data comprising: a sequence of sample data blocks corresponding to sample data under at least two modalities; a splicing processing module for splicing the sequence of sample data blocks under the at least two modalities to obtain a sample splicing sequence; and a training module for training the multimodal large model in combination with the sample splicing sequence.
[0008] According to another aspect of the present disclosure, a device for processing a multimodal task is provided, the device comprising: a first acquisition module for acquiring a multimodal task; the multimodal task comprises data in at least one modality; data in a text modality in the at least one modality is used to indicate that data in a target modality needs to be generated; a second acquisition module for acquiring a multimodal large model; the multimodal large model is determined based on the training method for the multimodal large model as described above; and a determination module for determining the generated data in the target modality in combination with the data in the at least one modality and the multimodal large model.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the multimodal large model proposed above in the present disclosure; or, to execute the processing method of the multimodal task proposed above in the present disclosure.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the training method of the multimodal large model proposed above in the present disclosure; or, to execute the processing method of the multimodal task proposed above in the present disclosure.
[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method for training a multimodal large model proposed above in the present disclosure; or, implements the steps of the method for processing a multimodal task proposed above in the present disclosure.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0014] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0015] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0016] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure;
[0020] Figure 7 It is a block diagram of an electronic device used to implement the training method of a multimodal large model or the processing method of a multimodal task of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] When processing multimodal understanding tasks, the current large multimodal model needs to connect the understanding network of the corresponding modality before the large multimodal model; when processing multimodal generation tasks, the generation network of the corresponding modality needs to be connected after the large multimodal model.
[0023] In the above scheme, the cascade between multiple networks is prone to cascade errors; and when the connected networks are fixed, it can only be applied to specific multimodal understanding tasks or multimodal generation tasks, and the applicable scenarios are limited.
[0024] To address the above issues, the present disclosure proposes a training method, device, and electronic device for a multimodal large model.
[0025] Figure 1 This is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the multimodal large model training method of the present embodiment can be applied to a multimodal large model training device. This device can be configured in an electronic device so that the electronic device can perform the multimodal large model training function. The following embodiments are described using an electronic device as an example.
[0026] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.
[0027] The training device for the multimodal large model may also be software in an electronic device, such as training software for the multimodal large model, etc. In the following embodiments, the training device for the multimodal large model is described as an electronic device.
[0028] like Figure 1 As shown, the training method of the multimodal large model may include the following steps:
[0029] Step 101: Acquire training data and an initial multimodal large model; the training data includes: a sequence of sample data blocks corresponding to sample data under at least two modalities.
[0030] In the disclosed embodiments, the modalities in the training data may include at least two of the following: text, image, audio, and video. The number of sample data block sequences for each modality in the training data may be one or more. The sample data block sequences may be obtained by segmenting the sample data.
[0031] The number of training data items can be at least one. For example, one training data item can include M text items, N images, K videos, L audio files, etc. The value of M can be an integer greater than or equal to 0; the value of N can be an integer greater than or equal to 0; the value of K can be an integer greater than or equal to 0; and the value of L can be an integer greater than or equal to 0. The values of M, N, K, and L cannot all be 0.
[0032] M texts, N images, K videos, and L audio files can be combined in any order. Data from each modality can be interspersed. For example, a text file, an image, a video, a text file, and an audio file can be combined.
[0033] Among them, the modalities in the training data are at least two of text, images, audio and video, and the number is one or more, so that the multimodal large model can be combined with any number of sample data block sequences under any modality for training processing, so that it can fully learn the features in the data under each modality, so that the trained multimodal large model can be applicable to a variety of multimodal understanding tasks and multimodal generation tasks, expanding the applicable scenarios of the multimodal large model.
[0034] Step 102: Splice sample data block sequences under at least two modalities to obtain a sample splicing sequence.
[0035] In an embodiment of the present disclosure, in one example, an electronic device may splice sample data block sequences under at least two modalities according to the arrangement order of the sample data block sequences under at least two modalities in the training data to obtain a sample splicing sequence for training a multimodal large model.
[0036] In another example, in order to expand the amount of training data so that the multimodal large model can learn the interaction characteristics between data under multiple modalities, the electronic device can shuffle the arrangement order of the sample data block sequences under at least two modalities in the training data; splice the sample data block sequences under at least two modalities according to the shuffled order to obtain a shuffled sample splicing sequence for training the multimodal large model.
[0037] In the embodiment of the present disclosure, in order to ensure that the multimodal large model can distinguish the sample data block sequences under each modality in the sample splicing sequence, thereby being able to extract features in a targeted manner and improve the accuracy of feature extraction, the electronic device can perform head and tail modality marking processing on the sample data block sequence to obtain a processed sample data block sequence; and splice the processed sample data block sequences to obtain a sample splicing sequence.
[0038] Performing head and tail modality tagging on the sample data chunk sequence refers to adding a modality tag before the first sample data chunk in the sample data chunk sequence, and adding a modality tag after the last sample data chunk in the sample data chunk sequence. Sample data chunk sequences in different modalities may be tagged using different modality tags.
[0039] Step 103: Combine the sample splicing sequence to train the multimodal large model.
[0040] In the disclosed embodiments, the electronic device may further perform the following processes: obtaining a sequence of sample data blocks corresponding to sample data in a single modality; and training a multimodal large model by combining the sequence of sample data blocks in the single modality. The single modality may be any one of text, image, video, and audio.
[0041] The training method of the multimodal large model of the embodiment of the present disclosure obtains training data and an initial multimodal large model; the training data includes: a sequence of sample data blocks corresponding to sample data under at least two modalities; the sample data block sequences under at least two modalities are spliced to obtain a sample splicing sequence; the multimodal large model is trained in combination with the sample data block sequences under at least two modalities; wherein, the multimodal large model is trained in combination with the sample data block sequences under at least two modalities, so that the trained multimodal large model can be applicable to a variety of multimodal understanding tasks and multimodal generation tasks, thereby expanding the applicable scenarios of the multimodal large model; and only using the multimodal large model can avoid cascade errors caused by multiple network cascades.
[0042] In order to further improve the accuracy of obtaining the sample data block sequence under at least two modalities in the training data, the electronic device can obtain the sample data block sequence by segmentation, horizontal arrangement or mapping of the sample data under at least two modalities. Figure 2 As shown, Figure 2 is a schematic diagram according to a second embodiment of the present disclosure, Figure 2 The illustrated embodiment may include the following steps:
[0043] Step 201: Obtain an initial multimodal large model.
[0044] Step 202: Obtain sample data under at least two modalities.
[0045] Step 203 : Segment and horizontally arrange the sample data in the non-text mode among the at least two modes to obtain a sequence of sample data blocks in the non-text mode.
[0046] In the embodiment of the present disclosure, it is assumed that the at least two modalities include an image modality, a video modality, and an audio modality. For the image modality, the electronic device may determine a sample data block sequence in the image modality by, for example, performing two-dimensional segmentation processing on an image in the image modality to obtain a plurality of image blocks; horizontally arranging the plurality of image blocks from left to right and from top to bottom according to the image dimensions to obtain an image block sequence; and determining the image block sequence as a sample data block sequence in the image modality.
[0047] For the video modality, the process of the electronic device determining the sample data block sequence in the video modality can, for example, be: performing three-dimensional segmentation processing on the video in the video modality to obtain multiple video blocks (wherein the video blocks are obtained by combining image blocks at the same position in multiple consecutive frames of images); horizontally arranging the multiple video blocks from front to back in the time dimension and from left to right and from top to bottom in the image dimension to obtain a video block sequence; and determining the video block sequence as a sample data block sequence in the video modality.
[0048] For the audio modality, when the audio in the audio modality is audio in the time domain, in one example, the process of the electronic device determining the sample data block sequence in the audio modality can be, for example, dividing the audio according to time periods to obtain multiple audio segments; arranging the multiple audio segments in chronological order to obtain an audio segment sequence; and determining the audio segment sequence as a sample data block sequence in the audio modality.
[0049] For the audio modality, when the audio in the audio modality is audio in the time domain, in another example, the process of the electronic device determining the sample data block sequence in the audio modality can be, for example, performing frequency domain conversion processing on the audio to obtain a spectrogram; dividing the spectrogram according to frequency bands to obtain multiple audio frequency bands; arranging the multiple audio frequency bands in frequency band order to obtain an audio frequency band sequence; and determining the audio frequency band sequence as a sample data block sequence in the audio modality.
[0050] Among them, the spectrum diagram can reflect the distribution of audio in frequency and has a large amount of information, which enables the multimodal large model to extract more features, thereby enabling the multimodal large model to learn more features and improve the accuracy of the trained multimodal large model.
[0051] Step 204 : When the at least two modalities include a text modality, segment and map the sample data in the text modality in combination with the text vocabulary to obtain a sequence of sample data blocks in the text modality.
[0052] In the disclosed embodiment, the text vocabulary may include words and integer identifiers corresponding to the words. Accordingly, the electronic device may execute step 204 by, for example, performing word segmentation processing on the sample data in the text modality based on the words in the text vocabulary to obtain a word sequence; querying the text vocabulary based on the words in the word sequence to obtain integer identifiers corresponding to the words in the text vocabulary; and determining a sequence consisting of the integer identifiers corresponding to the words in the word sequence as a sample data block sequence in the text modality.
[0053] Among them, combining the various words in the text vocabulary and the corresponding integer identifiers to determine the sequence of sample data blocks under the text modality can reduce the amount of data that needs to be processed by the multimodal large model and improve the processing efficiency of the multimodal large model.
[0054] Step 205 : performing splicing processing on the sample data block sequences under at least two modalities to obtain a sample splicing sequence.
[0055] In the embodiment of the present disclosure, in order to further reduce the data processing volume of the multimodal large model, before step 205, the electronic device may also perform the following process: combining the dimensionality change strategy to perform dimensionality reduction processing on the sample data block sequence in the non-text mode in at least two modes.
[0056] Among them, the dimension change strategy includes at least one of the following: a dimension change strategy based on variational autoencoder, and a dimension change strategy based on matrix transformation.
[0057] The variational autoencoder may include an encoder and a decoder. The encoder may map high-dimensional input data to a low-dimensional latent space, achieving dimensionality reduction. The decoder may map the low-dimensional latent space to high-dimensional output data, achieving dimensionality increase. The electronic device may combine the encoding network in the variational autoencoder to perform dimensionality reduction processing on each sample data block in a sequence of sample data blocks in a non-text modality.
[0058] Among them, the dimension-changing strategy based on matrix transformation, for example, performs deformation processing based on a deformation matrix, that is, reshaping a matrix into a matrix with different dimensions, but keeping the total number of elements unchanged.
[0059] Step 206: Combine the sample splicing sequence to train the multimodal large model.
[0060] It should be noted that the details of steps 205 to 206 can be found in Figure 1 Steps 102 to 103 in the illustrated embodiment will not be described in detail here.
[0061] The training method of the multimodal large model of the embodiment of the present disclosure is as follows: obtaining an initial multimodal large model; obtaining sample data under at least two modalities; segmenting and horizontally arranging the sample data under the non-text modality of the at least two modalities to obtain a sequence of sample data blocks under the non-text modality; when the at least two modalities include a text modality, segmenting and mapping the sample data under the text modality in combination with a text vocabulary to obtain a sequence of sample data blocks under the text modality; splicing the sample data block sequences under the at least two modalities to obtain a sample splicing sequence; and training the multimodal large model in combination with the sample splicing sequence; wherein, by segmenting, horizontally arranging, or mapping the sample data under the at least two modalities to obtain a sample data block sequence, the accuracy of obtaining the sample data block sequence under the at least two modalities in the training data is further improved, thereby further improving the accuracy of the trained multimodal large model.
[0062] In order to further improve the training speed and accuracy of the multimodal large model, the electronic device can input the sample splicing sequence into the multimodal large model to obtain the predicted splicing sequence; and then determine the loss function value to perform training processing. Figure 3 As shown, Figure 3 is a schematic diagram according to a third embodiment of the present disclosure, Figure 3 The illustrated embodiment may include the following steps:
[0063] Step 301: Acquire training data and an initial multimodal large model; the training data includes: a sequence of sample data blocks corresponding to sample data under at least two modalities.
[0064] Step 302: Splice sample data block sequences under at least two modalities to obtain a sample splicing sequence.
[0065] Step 303: Input the sample splicing sequence into the multimodal large model to obtain a predicted splicing sequence.
[0066] In the disclosed embodiment, the multimodal large model adopts a unidirectional attention mechanism. It is assumed that the sample splicing sequence is obtained by splicing a sample data block sequence under the first modality, a sample data block sequence under the second modality, and a sample data block sequence under the third modality. Accordingly, the multimodal large model processes the sample splicing sequence to obtain a predicted splicing sequence in the following manner: for the last sample data block in the sample data block sequence under the first modality, based on the last sample data block under the first modality, predict the first predicted data block under the second modality; for the last sample data block in the sample data block sequence under the second modality, based on the last sample data block under the second modality, predict the first predicted data block under the third modality.
[0067] For the i-th sample data block in the sample data block sequence under the first mode, if the i-th sample data block is not the last sample data block, the i-th sample data block may be combined with the i-th sample data block to predict the i+1-th prediction data block in the prediction data block sequence under the first mode. For the i-th sample data block in the sample data block sequence under the second mode, if the i-th sample data block is not the last sample data block, the i-th sample data block may be combined with the i-th sample data block to predict the i+1-th prediction data block in the prediction data block sequence under the second mode. For the i-th sample data block in the sample data block sequence under the third mode, if the i-th sample data block is not the last sample data block, the i-th sample data block may be combined with the i-th sample data block to predict the i+1-th prediction data block in the prediction data block sequence under the third mode.
[0068] Among them, the adoption of the unidirectional attention mechanism enables the multimodal large model to predict the first predicted data block in the predicted data block sequence under the next modality based on the last sample data block in the sample data block sequence under the previous modality in the sample splicing sequence, so that the multimodal large model can learn the interactive features between data under multiple modalities and realize cross-modal feature transfer.
[0069] Step 304: Determine the loss function value of the multimodal large model based on the sample splicing sequence and the predicted splicing sequence.
[0070] In the disclosed embodiment, the sample data block sequences in the sample splicing sequence correspond one-to-one to the prediction data block sequences in the prediction splicing sequence. Accordingly, the electronic device may perform step 304 by, for example, determining a sub-loss function value for each sample data block sequence in the sample splicing sequence based on the sample data block sequence and the prediction data block sequence corresponding to the sample data block sequence in the prediction splicing sequence; and determining a total loss function value based on at least two sub-loss function values.
[0071] The sample data blocks in the sample data block sequence have a one-to-one correspondence with the prediction data blocks in the prediction data block sequence corresponding to the sample data block sequence. Accordingly, the electronic device may determine the value of the sub-loss function by, for example, determining, for each sample data block in the sample data block sequence, a difference value between the sample data block and the prediction data block corresponding to the sample data block; and determining the value of the sub-loss function based on the difference value corresponding to each sample data block in the sample data block sequence.
[0072] Among them, the determination of the difference value between the sample data block and the prediction data block, and the determination of the sub-loss function value between the sample data block sequence and the prediction data block sequence can improve the accuracy of the determined total loss function value and further improve the training accuracy of the multimodal large model.
[0073] Step 305: Adjust the parameters of the multimodal large model according to the loss function value to achieve training.
[0074] It should be noted that the details of steps 301 to 302 can be found in Figure 1 Steps 101 to 102 in the illustrated embodiment will not be described in detail here.
[0075] The training method of the multimodal large model of the embodiment of the present disclosure is through obtaining training data and an initial multimodal large model; the training data includes: a sample data block sequence corresponding to sample data under at least two modalities; the sample data block sequence under at least two modalities is spliced to obtain a sample spliced sequence; the sample spliced sequence is input into the multimodal large model to obtain a predicted spliced sequence; the loss function value of the multimodal large model is determined according to the sample spliced sequence and the predicted spliced sequence; the parameters of the multimodal large model are adjusted according to the loss function value to achieve training; wherein, the loss function value is determined in combination with the sample spliced sequence and the predicted spliced sequence for training processing, which can further improve the training speed and training accuracy of the multimodal large model.
[0076] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure. It should be noted that the multimodal task processing method of the present embodiment can be applied to a multimodal task processing device. This device can be configured in an electronic device to enable the electronic device to perform multimodal task processing functions. The following embodiments are described using an electronic device as an example.
[0077] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.
[0078] The multimodal task processing device may also be software in an electronic device, such as multimodal task processing software, etc. In the following embodiments, the multimodal task processing device is described as an electronic device.
[0079] like Figure 4 As shown, the processing method of the multimodal task may include the following steps:
[0080] Step 401: Acquire a multimodal task; the multimodal task includes data in at least one modality; data in a text modality in the at least one modality is used to indicate that data in a target modality needs to be generated.
[0081] In the disclosed embodiments, the data in the Chinese modality can be used to describe a multimodal task, or to describe the target modality to which the data generated by the multimodal task belongs. For example, the data in the Chinese modality can include the name of the multimodal task, the target modality, etc.
[0082] Step 402: Obtain a multimodal large model; the multimodal large model is based on Figures 1 to 3The training method of the multimodal large model in any of the embodiments is determined.
[0083] Step 403 : Determine generated data in the target modality by combining data in at least one modality and the multimodal large model.
[0084] In an embodiment of the present disclosure, the process of the electronic device executing step 403 may, for example, be to segment and map the data in the text mode in combination with the text vocabulary to obtain a data block sequence in the text mode; when at least one mode includes a non-text mode, segment and horizontally arrange the data in the non-text mode to obtain a data block sequence in the non-text mode; splice the data block sequence in at least one mode to obtain a spliced sequence; input the spliced sequence into a multimodal large model to obtain a predicted sequence; and determine the generated data in the target mode based on the predicted data block sequence in the target mode in the predicted sequence.
[0085] In which, when at least one modality includes an audio modality, the data block sequence under the audio modality can be determined in combination with the audio in the time domain; or, the audio in the time domain can be transformed into the frequency domain to obtain the audio in the frequency domain, and then the data block sequence under the audio modality can be determined in combination with the audio in the frequency domain.
[0086] Among them, the segmentation processing, horizontal arrangement processing or mapping processing of the data under at least one modality can be carried out in units of data blocks in the data, so as to obtain the accuracy of the acquired data block sequence, and enable the multimodal large model to be trained and processed in units of data blocks, thereby further improving the accuracy of the trained multimodal large model.
[0087] In an embodiment of the present disclosure, to further reduce the amount of data processing required by the multimodal large model, before splicing the data block sequence under at least one modality to obtain the spliced sequence, the electronic device may further perform the following process: when at least one modality includes a non-text modality, the data block sequence under the non-text modality may be subjected to dimensionality reduction processing in combination with a dimensionality change strategy. Correspondingly, after inputting the spliced sequence into the multimodal large model to obtain a predicted sequence, the electronic device may further perform the following process: when the target modality includes a target non-text modality, the data block sequence under the target non-text modality in the predicted sequence may be subjected to dimensionality increase processing in combination with a dimensionality change strategy.
[0088] Among them, the dimension change strategy may include at least one of the following: a dimension change strategy based on a variational autoencoder, and a dimension change strategy based on matrix transformation.
[0089] The method for processing a multimodal task according to an embodiment of the present disclosure comprises the following steps: obtaining a multimodal task; the multimodal task comprises data in at least one modality; data in a text modality of at least one modality is used to indicate that data in a target modality needs to be generated; obtaining a multimodal large model; and the multimodal large model is based on the following example. Figures 1 to 3 The training method of the multimodal large model in any of the embodiments is determined; combining the data under at least one modality and the multimodal large model, the generated data under the target modality is determined; wherein the multimodal large model is trained based on a sequence of sample data blocks corresponding to sample data under at least two modalities, and can be applied to processing a variety of multimodal understanding tasks and multimodal generation tasks, thereby improving the processing efficiency of multimodal tasks.
[0090] In order to implement the above embodiment, the present disclosure also provides a training device for a multimodal large model. Figure 5 As shown, Figure 5 Schematic diagram of the fifth embodiment of the present disclosure. The multimodal large model training device 50 may include: an acquisition module 501 , a splicing processing module 502 , and a training module 503 .
[0091] Among them, the acquisition module 501 is used to obtain training data and an initial multimodal large model; the training data includes: a sequence of sample data blocks corresponding to sample data under at least two modalities; the splicing processing module 502 is used to splice the sample data block sequences under the at least two modalities to obtain a sample splicing sequence; the training module 503 is used to train the multimodal large model in combination with the sample splicing sequence.
[0092] As a possible implementation of an embodiment of the present disclosure, the acquisition module 501 includes a first acquisition unit, a first processing unit, and a second processing unit; the first acquisition unit is used to acquire sample data under the at least two modalities; the first processing unit is used to segment and horizontally arrange the sample data under the non-text modality among the at least two modalities to obtain a sequence of sample data blocks under the non-text modality; and the second processing unit is used to segment and map the sample data under the text modality in combination with a text vocabulary when the at least two modalities include a text modality to obtain a sequence of sample data blocks under the text modality.
[0093] As a possible implementation manner of an embodiment of the present disclosure, the at least two modalities include an audio modality; the first processing unit is specifically configured to perform frequency domain transform processing on the audio data in the audio modality to obtain frequency domain data; and perform frequency band segmentation processing and horizontal arrangement processing on the frequency domain data to obtain a sequence of sample data blocks in the audio modality.
[0094] As a possible implementation of an embodiment of the present disclosure, the at least two modalities include a video modality and an image modality; the first processing unit is specifically configured to, with respect to video data in the video modality, perform three-dimensional segmentation processing on the video data and perform horizontal arrangement processing according to a time dimension and an image dimension to obtain a sequence of sample data blocks in the video modality; and, with respect to image data in the image modality, perform two-dimensional segmentation processing on the image data and perform horizontal arrangement processing according to an image dimension to obtain a sequence of sample data blocks in the image modality.
[0095] As a possible implementation method of an embodiment of the present disclosure, the text vocabulary includes words and integer identifiers corresponding to the words; the second processing unit is specifically used to perform word segmentation processing on the sample data under the text modality in combination with the words in the text vocabulary to obtain a word sequence; query the text vocabulary according to the words in the word sequence to obtain the integer identifiers corresponding to the words in the text vocabulary; and determine the sequence composed of the integer identifiers corresponding to each word in the word sequence as the sample data block sequence under the text modality.
[0096] As a possible implementation of the embodiment of the present disclosure, the splicing processing module 502 is specifically configured to: perform head and tail modality marking processing on each sample data chunk sequence to obtain a processed sample data chunk sequence; and perform splicing processing on each of the processed sample data chunk sequences to obtain a sample splicing sequence.
[0097] As a possible implementation of the embodiment of the present disclosure, the device further includes: a scrambling processing module, which is used to scramble the splicing order of each sample data block sequence in the sample splicing sequence to obtain a scrambled sample splicing sequence; the training module 503 is also used to train the multimodal large model in combination with the scrambled sample splicing sequence.
[0098] As a possible implementation of the embodiment of the present disclosure, the apparatus further includes: a dimension change processing module configured to perform dimension reduction processing on a sequence of sample data blocks in a non-text mode among the at least two modes in combination with a dimension change strategy.
[0099] As a possible implementation of an embodiment of the present disclosure, the dimension change strategy includes at least one of the following: a dimension change strategy based on a variational autoencoder, and a dimension change strategy based on matrix transformation.
[0100] As a possible implementation method of an embodiment of the present disclosure, the training module 503 includes: a second acquisition unit, a determination unit and an adjustment processing unit; the second acquisition unit is used to input the sample splicing sequence into the multimodal large model to obtain a predicted splicing sequence; the determination unit is used to determine the loss function value of the multimodal large model based on the sample splicing sequence and the predicted splicing sequence; the adjustment processing unit is used to perform parameter adjustment processing on the multimodal large model according to the loss function value to achieve training.
[0101] As a possible implementation of an embodiment of the present disclosure, the multimodal large model adopts a unidirectional attention mechanism; the sample splicing sequence is obtained by splicing a sample data block sequence under the first modality, a sample data block sequence under the second modality, and a sample data block sequence under the third modality; the multimodal large model processes the sample splicing sequence to obtain the predicted splicing sequence in the following manner: for the last sample data block in the sample data block sequence under the first modality, based on the last sample data block under the first modality, predicting the first predicted data block under the second modality; for the last sample data block in the sample data block sequence under the second modality, based on the last sample data block under the second modality, predicting the first predicted data block under the third modality.
[0102] As a possible implementation manner of an embodiment of the present disclosure, a sample data block sequence in the sample splicing sequence corresponds one-to-one to a prediction data block sequence in the prediction splicing sequence; the determining unit is specifically configured to determine, for each sample data block sequence in the sample splicing sequence, a sub-loss function value based on the sample data block sequence and a prediction data block sequence corresponding to the sample data block sequence in the prediction splicing sequence; and determine a total loss function value based on at least two of the sub-loss function values.
[0103] As a possible implementation manner of the embodiment of the present disclosure, the sample data blocks in the sample data block sequence have a one-to-one correspondence with the prediction data blocks in the prediction data block sequence corresponding to the sample data block sequence; the determining unit is further configured to determine, for each sample data block in the sample data block sequence, a difference value between the sample data block and the prediction data block corresponding to the sample data block; and determine the value of the sub-loss function based on the difference value corresponding to each sample data block in the sample data block sequence.
[0104] As a possible implementation of an embodiment of the present disclosure, the modalities in the training data include at least two of the following: text, image, audio, and video; and the number of sample data block sequences under each modality in the training data is one or more.
[0105] The training device for a multimodal large model of an embodiment of the present disclosure obtains training data and an initial multimodal large model; the training data includes: a sequence of sample data blocks corresponding to sample data under at least two modalities; the sample data block sequences under at least two modalities are spliced to obtain a sample splicing sequence; and the multimodal large model is trained in combination with the sample data block sequences under at least two modalities; wherein, the multimodal large model is trained in combination with the sample data block sequences under at least two modalities, so that the trained multimodal large model can be applicable to a variety of multimodal understanding tasks and multimodal generation tasks, thereby expanding the applicable scenarios of the multimodal large model; and by using only the multimodal large model, cascade errors caused by multiple network cascades can be avoided.
[0106] In order to implement the above embodiment, the present disclosure also provides a multi-modal task processing device. Figure 6 As shown, Figure 6 6 is a schematic diagram of a sixth embodiment of the present disclosure. The multimodal task processing device 60 may include: a first acquisition module 601 , a second acquisition module 602 , and a determination module 603 .
[0107] Among them, the first acquisition module 601 is used to acquire a multimodal task; the multimodal task includes data under at least one modality; the data under the text modality in the at least one modality is used to indicate that data under the target modality needs to be generated; the second acquisition module 602 is used to acquire a multimodal large model; the multimodal large model is based on Figures 1 to 3 The training method of the multimodal large model in any embodiment is determined; a determination module 603 is used to combine the data under the at least one modality and the multimodal large model to determine the generated data under the target modality.
[0108] As a possible implementation method of an embodiment of the present disclosure, the determination module 603 is specifically used to segment and map the data under the text modality in combination with the text vocabulary to obtain a data block sequence under the text modality; when the at least one modality includes a non-text modality, segment and horizontally arrange the data under the non-text modality to obtain a data block sequence under the non-text modality; splice the data block sequence under the at least one modality to obtain a spliced sequence; input the spliced sequence into the multimodal large model to obtain a predicted sequence; and determine the generated data under the target modality based on the predicted data block sequence under the target modality in the predicted sequence.
[0109] As a possible implementation of an embodiment of the present disclosure, the device also includes: a dimension change processing module, which is used to, when the at least one modality includes a non-text modality, combine the dimension change strategy to perform dimensionality reduction processing on the data block sequence under the non-text modality; the dimension change processing module is also used to, when the target modality includes a target non-text modality, combine the dimension change strategy to perform dimensionality increase processing on the data block sequence under the target non-text modality in the prediction sequence.
[0110] The processing device of the multimodal task of the embodiment of the present disclosure obtains a multimodal task; the multimodal task includes data in at least one modality; data in a text modality in at least one modality is used to indicate that data in a target modality needs to be generated; obtains a multimodal large model; the multimodal large model is based on Figures 1 to 3 The training method of the multimodal large model in any of the embodiments is determined; combining the data under at least one modality and the multimodal large model, the generated data under the target modality is determined; wherein the multimodal large model is trained based on a sequence of sample data blocks corresponding to sample data under at least two modalities, and can be applied to processing a variety of multimodal understanding tasks and multimodal generation tasks, thereby improving the processing efficiency of multimodal tasks.
[0111] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, comply with relevant laws and regulations, and do not violate public order and good morals.
[0112] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0113] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0114] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0115] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0116] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the training method of a multimodal large model or the processing method of a multimodal task. For example, in some embodiments, the training method of a multimodal large model or the processing method of a multimodal task can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method of the multimodal large model or the processing method of the multimodal task described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute a training method for a multimodal large model or a processing method for a multimodal task in any other appropriate manner (e.g., by means of firmware).
[0117] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0118] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0119] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0121] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0122] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0123] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0124] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a large multimodal model, the method comprising: Obtain training data and an initial multimodal large model; The training data includes: a sequence of sample data blocks corresponding to sample data in at least two modalities; the modalities include at least two of the following: text, image, audio, and video; performing splicing processing on the sample data block sequences under the at least two modalities to obtain a sample splicing sequence; The multimodal large model is trained in combination with the sample splicing sequence; the sample splicing sequence is used to input the multimodal large model to obtain a predicted splicing sequence; The multimodal large model adopts a unidirectional attention mechanism; the sample splicing sequence is obtained by splicing the sample data block sequence under the first modality, the sample data block sequence under the second modality, and the sample data block sequence under the third modality; The multimodal large model processes the sample splicing sequence to obtain the predicted splicing sequence, including: Predicting a first predicted data block under the second modality based on a last sample data block in the sequence of sample data blocks under the first modality; The first predicted data block under the third mode is predicted based on the last sample data block in the sequence of sample data blocks under the second mode.
2. The method according to claim 1, wherein Get training data, including: Acquiring sample data under the at least two modalities; Performing segmentation and horizontal arrangement on the sample data in the non-text mode of the at least two modes to obtain a sequence of sample data blocks in the non-text mode; In a case where the at least two modalities include a text modality, the sample data in the text modality is segmented and mapped in combination with a text vocabulary to obtain a sequence of sample data blocks in the text modality.
3. The method according to claim 2, wherein: The at least two modalities include an audio modality; and the segmentation and horizontal arrangement of the sample data in the non-text modality of the at least two modalities to obtain a sequence of sample data blocks in the non-text modality includes: For the audio data in the audio mode, performing frequency domain transformation processing on the audio data to obtain frequency domain data; The frequency domain data is subjected to frequency band segmentation processing and horizontal arrangement processing to obtain a sample data block sequence under the audio mode.
4. The method according to claim 2, wherein: The at least two modalities include a video modality and an image modality; and the segmentation and horizontal arrangement of the sample data in the non-text modality of the at least two modalities to obtain a sequence of sample data blocks in the non-text modality includes: For the video data in the video modality, performing three-dimensional segmentation processing on the video data and horizontal arrangement processing according to the time dimension and the image dimension to obtain a sequence of sample data blocks in the video modality; For the image data in the image modality, two-dimensional segmentation processing is performed on the image data and horizontal arrangement processing is performed according to the image dimension to obtain a sample data block sequence in the image modality.
5. The method according to claim 2, wherein: The text vocabulary includes words and integer identifiers corresponding to the words; the text vocabulary is combined to segment and map the sample data in the text modality to obtain a sequence of sample data blocks in the text modality, including: Performing word segmentation processing on the sample data in the text modality in combination with the words in the text vocabulary to obtain a word sequence; Searching the text vocabulary according to the words in the word sequence to obtain an integer identifier corresponding to the word in the text vocabulary; A sequence consisting of integer identifiers corresponding to each word in the word sequence is determined as a sample data block sequence in the text modality.
6. The method according to claim 1, wherein The step of splicing the sample data block sequences under the at least two modalities to obtain a sample splicing sequence includes: For each sample data block sequence, performing head and tail modality marking processing on the sample data block sequence to obtain a processed sample data block sequence; The processed sample data block sequences are spliced together to obtain a sample splicing sequence.
7. The method according to claim 1 or 6, wherein: The method further comprises: Performing a shuffling process on the splicing order of each sample data block sequence in the sample splicing sequence to obtain a shuffled sample splicing sequence; The multimodal large model is trained in combination with the shuffled sample splicing sequence.
8. The method according to claim 1, wherein Before splicing the sample data block sequences under the at least two modalities to obtain a sample splicing sequence, the method further includes: performing dimensionality reduction processing on the sample data block sequence under the non-text modality of the at least two modalities in combination with a dimensionality change strategy.
9. The method according to claim 8, wherein The dimension change strategy includes at least one of the following: a dimension change strategy based on a variational autoencoder, and a dimension change strategy based on matrix transformation.
10. The method according to claim 1, wherein The step of training the multimodal large model by combining the sample splicing sequence includes: Inputting the sample splicing sequence into the multimodal large model to obtain a predicted splicing sequence; Determining a loss function value of the multimodal large model according to the sample splicing sequence and the predicted splicing sequence; The parameters of the multimodal large model are adjusted according to the loss function value to achieve training.
11. The method according to claim 10, wherein: The sample data block sequence in the sample splicing sequence corresponds one-to-one to the prediction data block sequence in the prediction splicing sequence; and determining the loss function value of the multimodal large model based on the sample splicing sequence and the prediction splicing sequence includes: For each sample data block sequence in the sample splicing sequence, determining a sub-loss function value according to the sample data block sequence and a prediction data block sequence corresponding to the sample data block sequence in the prediction splicing sequence; Determine a total loss function value based on at least two of the sub-loss function values.
12. The method according to claim 11, wherein The sample data blocks in the sample data block sequence correspond one-to-one with the prediction data blocks in the prediction data block sequence corresponding to the sample data block sequence; and determining a sub-loss function value based on the sample data block sequence and the prediction data block sequence corresponding to the sample data block sequence in the prediction concatenated sequence, comprising: determining, for each specimen data chunk in the sequence of specimen data chunks, a difference value between the specimen data chunk and a prediction data chunk corresponding to the specimen data chunk; The sub-loss function value is determined according to the difference value corresponding to each sample data block in the sample data block sequence.
13. The method according to claim 1, wherein The number of sample data block sequences under each modality in the training data is one or more.
14. A method for processing a multimodal task, the method comprising: Acquire multimodal tasks; The multimodal task includes data under at least one modality; The data in the text mode in the at least one mode is used to indicate that data in the target mode needs to be generated; Obtaining a large multimodal model; the large multimodal model is determined based on the training method of the large multimodal model according to any one of claims 1 to 13; The generated data under the target modality is determined by combining the data under the at least one modality and the multimodal large model.
15. The method according to claim 14, wherein The combining the data under the at least one modality and the multimodal large model to determine the generated data under the target modality includes: Segmenting and mapping the data in the text mode in combination with the text vocabulary to obtain a sequence of data blocks in the text mode; In a case where the at least one modality includes a non-text modality, segmenting and horizontally arranging the data in the non-text modality to obtain a data block sequence in the non-text modality; performing splicing processing on the data block sequence under the at least one modality to obtain a spliced sequence; Inputting the spliced sequence into the multimodal large model to obtain a predicted sequence; Determine generated data under the target modality based on a prediction data block sequence under the target modality in the prediction sequence.
16. The method according to claim 15, wherein Before splicing the data block sequence under the at least one modality to obtain the spliced sequence, the method further includes: when the at least one modality includes a non-text modality, performing dimensionality reduction processing on the data block sequence under the non-text modality in combination with a dimensionality change strategy; Correspondingly, after inputting the spliced sequence into the multimodal large model and obtaining the predicted sequence, the method also includes: when the target modality includes a target non-text modality, combining the dimensionality change strategy to perform dimensionality increase processing on the data block sequence under the target non-text modality in the predicted sequence.
17. A multimodal large model training device, comprising: The acquisition module is used to obtain training data and the initial multimodal large model; The training data includes: a sequence of sample data blocks corresponding to sample data in at least two modalities; the modalities include at least two of the following: text, image, audio, and video; a splicing processing module, configured to splice the sample data block sequences under the at least two modalities to obtain a sample splicing sequence; A training module, configured to train the multimodal large model in combination with the sample splicing sequence; the sample splicing sequence is used to input the multimodal large model to obtain a predicted splicing sequence; The multimodal large model adopts a unidirectional attention mechanism; the sample splicing sequence is obtained by splicing the sample data block sequence under the first modality, the sample data block sequence under the second modality, and the sample data block sequence under the third modality; The multimodal large model processes the sample splicing sequence to obtain the predicted splicing sequence, including: Predicting a first predicted data block under the second modality based on a last sample data block in the sequence of sample data blocks under the first modality; The first predicted data block under the third mode is predicted based on the last sample data block in the sequence of sample data blocks under the second mode.
18. The device according to claim 17, wherein The acquisition module includes a first acquisition unit, a first processing unit and a second processing unit; The first acquisition unit is configured to acquire sample data under the at least two modalities; The first processing unit is configured to perform segmentation and horizontal arrangement on the sample data in the non-text mode of the at least two modes to obtain a sequence of sample data blocks in the non-text mode; The second processing unit is configured to, when the at least two modalities include a text modality, segment and map the sample data in the text modality in combination with a text vocabulary to obtain a sequence of sample data blocks in the text modality.
19. The device according to claim 18, wherein The at least two modes include an audio mode; the first processing unit is specifically configured to: For the audio data in the audio mode, performing frequency domain transformation processing on the audio data to obtain frequency domain data; The frequency domain data is subjected to frequency band segmentation processing and horizontal arrangement processing to obtain a sample data block sequence under the audio mode.
20. The apparatus according to claim 18, wherein The at least two modalities include a video modality and an image modality; the first processing unit is specifically configured to: For the video data in the video modality, performing three-dimensional segmentation processing on the video data and horizontal arrangement processing according to the time dimension and the image dimension to obtain a sequence of sample data blocks in the video modality; For the image data in the image modality, two-dimensional segmentation processing is performed on the image data and horizontal arrangement processing is performed according to the image dimension to obtain a sample data block sequence in the image modality.
21. The apparatus according to claim 18, wherein The text vocabulary includes words and integer identifiers corresponding to the words; the second processing unit is specifically configured to: Performing word segmentation processing on the sample data in the text modality in combination with the words in the text vocabulary to obtain a word sequence; Searching the text vocabulary according to the words in the word sequence to obtain an integer identifier corresponding to the word in the text vocabulary; A sequence consisting of integer identifiers corresponding to each word in the word sequence is determined as a sample data block sequence in the text modality.
22. The apparatus according to claim 17, wherein The splicing processing module is specifically used to: For each sample data block sequence, performing head and tail modality marking processing on the sample data block sequence to obtain a processed sample data block sequence; The processed sample data block sequences are spliced together to obtain a sample splicing sequence.
23. The device according to claim 17 or 22, wherein The device further comprises: a scrambling processing module, configured to scramble the splicing order of each sample data block sequence in the sample splicing sequence to obtain a scrambled sample splicing sequence; The training module is also used to train the multimodal large model in combination with the shuffled sample splicing sequence.
24. The apparatus according to claim 17, wherein The device further includes: a dimension change processing module, configured to perform dimension reduction processing on a sequence of sample data blocks in a non-text mode among the at least two modes in combination with a dimension change strategy.
25. The apparatus according to claim 24, wherein The dimension change strategy includes at least one of the following: a dimension change strategy based on a variational autoencoder, and a dimension change strategy based on matrix transformation.
26. The apparatus according to claim 17, wherein The training module includes: a second acquisition unit, a determination unit and an adjustment processing unit; The second acquisition unit is configured to input the sample splicing sequence into the multimodal large model to obtain a predicted splicing sequence; The determining unit is configured to determine a loss function value of the multimodal large model based on the sample splicing sequence and the predicted splicing sequence; The adjustment processing unit is used to perform parameter adjustment processing on the multimodal large model according to the loss function value to achieve training.
27. The device according to claim 26, wherein The sample data block sequence in the sample splicing sequence corresponds one-to-one to the prediction data block sequence in the prediction splicing sequence; the determining unit is specifically configured to: For each sample data block sequence in the sample splicing sequence, determining a sub-loss function value according to the sample data block sequence and a prediction data block sequence corresponding to the sample data block sequence in the prediction splicing sequence; Determine a total loss function value based on at least two of the sub-loss function values.
28. The apparatus according to claim 27, wherein The sample data blocks in the sample data block sequence correspond one-to-one to the prediction data blocks in the prediction data block sequence corresponding to the sample data block sequence; the determining unit is further configured to: determining, for each specimen data chunk in the sequence of specimen data chunks, a difference value between the specimen data chunk and a prediction data chunk corresponding to the specimen data chunk; The sub-loss function value is determined according to the difference value corresponding to each sample data block in the sample data block sequence.
29. The apparatus according to claim 17, wherein The number of sample data block sequences under each modality in the training data is one or more.
30. A multimodal task processing device, comprising: A first acquisition module is used to acquire multimodal tasks; The multimodal task includes data under at least one modality; The data in the text mode in the at least one mode is used to indicate that data in the target mode needs to be generated; A second acquisition module is configured to acquire a large multimodal model; the large multimodal model is determined based on the training method for the large multimodal model according to any one of claims 1 to 13; A determination module is used to determine the generated data under the target modality by combining the data under the at least one modality and the multimodal large model.
31. The apparatus according to claim 30, wherein The determining module is specifically configured to: Segmenting and mapping the data in the text mode in combination with the text vocabulary to obtain a sequence of data blocks in the text mode; In a case where the at least one modality includes a non-text modality, segmenting and horizontally arranging the data in the non-text modality to obtain a data block sequence in the non-text modality; performing splicing processing on the data block sequence under the at least one modality to obtain a spliced sequence; Inputting the spliced sequence into the multimodal large model to obtain a predicted sequence; Determine generated data under the target modality based on a prediction data block sequence under the target modality in the prediction sequence.
32. The apparatus according to claim 31, wherein The device further includes: a dimension change processing module for performing dimension reduction processing on a data block sequence in the non-text mode in combination with a dimension change strategy when the at least one mode includes a non-text mode; The dimension change processing module is further configured to perform dimension increase processing on the data block sequence under the target non-text modality in the prediction sequence in combination with the dimension change strategy when the target modality includes a target non-text modality.
33. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13; Alternatively, the method according to any one of claims 14 to 16 is performed.
34. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13; or to execute the method according to any one of claims 14 to 16.
35. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 13; or implements the method according to any one of claims 14 to 16.
Citation Information
Patent Citations
Pre-training method, device and equipment of image-text understanding model and storage medium
CN116796287A
Multi-modal pre-training model training method and device and multi-modal data processing method and device
CN116861995A