Multi-modal large model training method and device and electronic device

The training method for a multimodal large-scale model addresses the challenge of separate modality training by using collaborative training of codec networks with a shared word list, resulting in improved efficiency and accuracy.

JP2025081719APending Publication Date: 2025-05-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025031574
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-27
Filing Date
2025-02-28
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current multimodal large-scale models, such as video generation models, require separate training of encoding networks for different modalities (video, text, audio), increasing the difficulty and cost of model training.

Method used

A training method for a multimodal large-scale model that involves obtaining initial training data and an initial model with a backbone network and codec networks for multiple non-text modalities, performing collaborative training on the codec networks and a shared word list, and then training the backbone network using multimodal sample reference data and sample generation data.

Benefits of technology

This approach reduces the complexity and cost of training by allowing the codec networks to share a common word list, improving training efficiency and accuracy for multimodal large-scale models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081719000001_ABST
    Figure 2025081719000001_ABST
Patent Text Reader

Abstract

To reduce the difficulty of model training and the cost of model training.SOLUTION: A method includes the steps of: acquiring first training data and second training data; acquiring an initial multi-modal large model; causing a backbone network and multiple codec networks corresponding to multiple non-text modalities included in the multi-modal large model to perform coding and decoding processing on the basis of the same multi-modal word list; performing joint training processing on the multiple codec networks and the multi-modal word list on the basis of the data in the multiple non-text modalities in the first training data; performing training processing on the backbone network on the basis of multi-modal sample reference data and sample generation data under the target task in the second training data; and causing the multiple codec networks to perform the coding and decoding processing on the basis of the same multi-modal word list.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as deep learning, natural language processing, computer vision, speech technology, large-scale models, etc. Specifically, it relates to a training method, apparatus, and electronic device for a multimodal large-scale model.

Background Art

[0002] Current multimodal large-scale models such as video generation models include an encoding network, a backbone network, and a decoding network. Here, the encoding network includes three types: an encoding network for the video modality, an encoding network for the text modality, and an encoding network for the audio modality.

[0003] In the above video generation model, the encoding networks of different modalities are obtained by using word lists of different modalities and training separately using data of different modalities. In the training process of the video generation model, it is necessary to separately train the word lists of different modalities, which increases the difficulty and cost of model training.

Summary of the Invention

[0004] The present disclosure provides a training method, apparatus, and electronic device for a multimodal large-scale model.

[0005] According to one aspect of the present disclosure, a method for training a multimodal large-scale model is provided. The method includes: obtaining first training data and second training data, where the first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data for a target task; obtaining an initial multimodal large-scale model, where the multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities, and the plurality of codec networks perform encoding and decoding processes based on the same multimodal word list; performing collaborative training processing on the plurality of codec networks and the multimodal word list based on the data in the plurality of non-text modalities; and when the training of the plurality of codec networks and the multimodal word list is completed, training the backbone network based on the multimodal sample reference data and sample generation data for the target task.

[0006] According to another aspect of the present disclosure, a method for processing a target task is provided. The method includes: obtaining a target task including data in at least two modalities; obtaining a multimodal large-scale model obtained based on the above method for training a multimodal large-scale model; and inputting the data in the at least two modalities into the multimodal large-scale model to obtain generated data output from the multimodal large-scale model.

[0007] According to another aspect of the present disclosure, there is provided a training apparatus for a multimodal large-scale model, the apparatus comprising: a first acquisition module for acquiring first training data and second training data, wherein the first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data for a target task; a second acquisition module for acquiring an initial multimodal large-scale model, wherein the multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities, and the plurality of codec networks perform encoding and decoding processes based on the same multimodal word list; a first training processing module for performing cooperative training processing on the plurality of codec networks and the multimodal word list based on data in a plurality of non-text modalities; and a second training processing module for training the backbone network based on the multimodal sample reference data and sample generation data for the target task when the training of the plurality of codec networks and the multimodal word list is completed.

[0008] According to another aspect of the present disclosure, there is provided a processing apparatus for a target task, the apparatus comprising: a first acquisition module for acquiring a target task including data in at least two modalities; a second acquisition module for acquiring a multimodal large-scale model obtained based on the above-described training method of the multimodal large-scale model; and a third acquisition module for inputting the data in the at least two modalities into the multimodal large-scale model and acquiring generated data output from the multimodal large-scale model.

[0009] According to another aspect of the present disclosure, an electronic device is provided, including at least one processor and a memory communicatively connected to the at least one processor. Here, instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor is caused to execute the above-mentioned multi-modal large-scale model training method of the present disclosure or the above-mentioned target task processing method of the present disclosure.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions cause a computer to execute the above-mentioned multi-modal large-scale model training method of the present disclosure or the above-mentioned target task processing method of the present disclosure.

[0011] According to another aspect of the present disclosure, a computer program is provided, and when the computer program is executed by a processor, the above-mentioned multi-modal large-scale model training method of the present disclosure or the above-mentioned target task processing method of the present disclosure is realized.

[0012] It should be noted that the content described in this part is not intended to identify essential or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be more easily understood from the following description.

Brief Description of the Drawings

[0013] The drawings are for a better understanding of the solution and do not limit the present disclosure.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0014] Hereinafter, exemplary embodiments of the present disclosure will be described in combination with the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, but they should be considered merely as examples. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the described embodiments without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the following description omits the description of well-known functions and structures.

[0015] Current multimodal large models such as video generation models include an encoding network, a backbone network, and a decoding network. Here, the encoding network includes three types: an encoding network for the video modality, an encoding network for the text modality, and an encoding network for the audio modality.

[0016] In the above video generation model, the encoding networks of different modalities are obtained by using word lists of different modalities and training them individually using data of different modalities. In the training process of the video generation model, it is necessary to separately train the word lists of different modalities, which increases the difficulty and cost of model training.

[0017] In response to the above problems, the present disclosure proposes a training method, apparatus, and electronic device for a multimodal large-scale model.

[0018] FIG. 1 is a schematic diagram of a first embodiment according to the present disclosure. The training method for a multimodal large-scale model according to the embodiments of the present disclosure is applicable to a training apparatus for a multimodal large-scale model, and this apparatus may be configured in an electronic device so that the electronic device can execute the training function of the multimodal large-scale model. In the following embodiments, the example where the execution subject is an electronic device will be described.

[0019] Here, the electronic device may be any device with computing functions such as a personal computer (abbreviated as PC), a mobile terminal, a server, etc., and the mobile terminal may be, for example, an in-vehicle device, a mobile phone, a tablet, a personal digital assistant, a wearable device, etc., and may be a hardware device with various operating systems, touch screens, and / or displays.

[0020] Here, the training apparatus for a multimodal large-scale model may be software in an electronic device, such as training software for a multimodal large-scale model. In the following embodiments, the example where the training apparatus for a multimodal large-scale model is an electronic device will be described.

[0021] As shown in FIG. 1, the training method for this multimodal large-scale model may include the following steps.

[0022] In step 101, acquire first training data and second training data, where the first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data for a target task.

[0023] In an embodiment of the present disclosure, the non-text modality includes at least one of an audio modality, a silent video modality, and an image modality. Accordingly, the data in the non-text modality is, for example, audio, silent video, an image, etc. Here, for audio-visual, the data can be divided into data in two non-text modalities. For example, the audio-visual may be divided into data in the audio modality and data in the video modality.

[0024] Here, with various non-text modality settings, the multi-modal large-scale model can flexibly process data in different text modalities, the application scenarios of the multi-modal large-scale model are extended, and the processing efficiency of the multi-modal large-scale model is improved.

[0025] In an embodiment of the present disclosure, the target task may include at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multi-modal understanding task.

[0026] Here, the multi-modal sample reference data in multiple target tasks can include, for example, sample reference data in at least one modality of a text modality, an image modality, a video modality, and an audio modality. Here, the multi-modal understanding task is a task of understanding and processing the input multi-modal data and outputting the understood text. Here, during the understanding process, the input multi-modal data can include sample reference data in at least one modality of a text modality, an image modality, a video modality, and an audio modality.

[0027] Here, by setting the multi-modal sample reference data and sample generation data for various target tasks, the multi-modal large-scale model is trained based on the training data to obtain a model applicable to different target tasks, the application scenario of the multi-modal large-scale model is expanded, and the processing efficiency of the multi-modal large-scale model is further improved.

[0028] In step 102, an initial multi-modal large-scale model is obtained. The multi-modal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multi-modal word list.

[0029] In the embodiments of the present disclosure, there may be a plurality of codec networks. For example, it may include a one-dimensional codec network, a two-dimensional codec network, and a three-dimensional codec network. Here, taking the non-text modality including an audio modality, a silent video modality, and an image modality as an example, the codec network corresponding to the audio modality may be a one-dimensional codec network, the codec network corresponding to the image modality may be a two-dimensional codec network, and the codec network corresponding to the video modality may be a three-dimensional codec network.

[0030] Here, the plurality of codec networks use the same multi-modal word list. Here, the multi-modal word list may include a plurality of integer identifiers and vectors corresponding to each integer identifier. Here, one vector in the multi-modal word list can represent an image, audio, and video simultaneously. That is, one vector in the multi-modal word list may be a vector of an image block in the image, or a vector of an audio segment in the audio, or a vector of a video block in the video.

[0031] Here, by using different codec network processes for data in different non-text modalities, as many features as possible in the data of the non-text modality can be extracted, and the accuracy of the extracted features can be improved.

[0032] In an embodiment of the present disclosure, taking a two-dimensional codec network as an example, processing an image in the image modality based on the encoding network in the two-dimensional codec network includes inputting the image into the encoding network in the two-dimensional codec network to obtain two-dimensional image features, performing a one-dimensional transformation process on the two-dimensional image features to obtain a transformed one-dimensional feature vector, and performing a mapping process on the features in the one-dimensional feature vector based on a multimodal word list to obtain an integer sequence corresponding to the image.

[0033] Here, taking a one-dimensional codec network as an example, processing audio in the audio modality based on the encoding network in the one-dimensional codec network includes inputting the audio into the encoding network in the one-dimensional codec network to obtain one-dimensional audio features, where the one-dimensional audio features are one-dimensional feature vectors, and performing a mapping process on the features in the one-dimensional feature vector based on a multimodal word list to obtain an integer sequence corresponding to the audio.

[0034] Here, taking a three-dimensional codec network as an example, processing a video in the video modality based on the encoding network in the three-dimensional codec network includes inputting the video into the encoding network in the three-dimensional codec network to obtain three-dimensional video features, performing a one-dimensional transformation process on the three-dimensional video features to obtain a transformed one-dimensional feature vector, and performing a mapping process on the features in the one-dimensional feature vector based on a multimodal word list to obtain an integer sequence corresponding to the video.

[0035] Here, the encoding network in each dimensional codec network can obtain one-dimensional feature vectors for data in multiple modalities by dimensionally transforming the features obtained by encoding the data, can realize feature mapping processing with the same multimodal word list, can realize the interaction between multiple modality features, and can further improve the accuracy of feature extraction of the encoding network in each dimensional codec network.

[0036] In step 103, based on the data in multiple non-text modalities, cooperative training processing is performed on multiple codec networks and a multimodal word list.

[0037] In the embodiments of the present disclosure, the electronic device determines the total loss function value of multiple codec networks based on the data in multiple non-text modalities and multiple codec networks, and further adjusts the parameters of the multiple codec networks and the multimodal word list simultaneously based on the total loss function value to realize cooperative training.

[0038] In step 104, when the training of multiple codec networks and the multimodal word list is completed, the backbone network is trained based on the multimodal sample reference data and sample generation data in the target task.

[0039] In the embodiments of the present disclosure, the backbone network may be, for example, an autoregressive model. The autoregressive model is a time series analysis model that predicts the future value of a variable by adding a random error term to the linear combination of the past values of the variable.

[0040] Here, the electronic device can determine an integer sequence or a combination of integer sequences corresponding to the predicted generated data based on the multimodal sample reference data, a plurality of encoding networks corresponding to a plurality of non-text modalities, and a backbone network, determine the modality to which the sample generated data belongs, and determine an integer sequence or a combination of integer sequences corresponding to the sample generated data based on the encoding network in the codec network corresponding to the sample generated data and the modality to which it belongs. Based on the integer sequence or combination of integer sequences corresponding to the predicted generated data and the integer sequence or combination of integer sequences corresponding to the sample generated data, a loss function value is determined, and further, the parameters of the backbone network are adjusted to realize training.

[0041] In the method for training a multimodal large-scale model according to an embodiment of the present disclosure, first training data and second training data are acquired. The first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data in a target task. An initial multimodal large-scale model is acquired. The multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multimodal word list. Based on the data in the plurality of non-text modalities, a collaborative training process is performed on the plurality of codec networks and the multimodal word list. When the training of the plurality of codec networks and the multimodal word list is completed, the backbone network is trained based on the multimodal sample reference data and the sample generation data in the target task. Here, the plurality of codec networks perform encoding and decoding processes, and a collaborative training process based on the same multimodal word list, and it is possible to avoid separately training word lists of different modalities in the training process, and reduce the difficulty and cost of model training.

[0042] Here, in order to further improve the training speed of the plurality of codec networks corresponding to the plurality of non-text modalities, based on the data in the plurality of non-text modalities, the loss function values of the plurality of codec networks are respectively determined, and further the plurality of loss function values are summed to adjust the plurality of codec networks and the multimodal word list. As shown in FIG. 2, FIG. 2 is a schematic diagram of a second embodiment according to the present disclosure. The embodiment shown in FIG. 2 can include the following steps.

[0043] In step 201, the first training data and the second training data are obtained. The first training data includes data in a plurality of non-text modalities, and the second training data includes multi-modal sample reference data and sample generation data in the target task.

[0044] In step 202, an initial multi-modal large-scale model is obtained. The multi-modal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multi-modal word list.

[0045] In step 203, based on the data in the plurality of non-text modalities and the plurality of codec networks corresponding to the plurality of non-text modalities, the loss function values of the plurality of codec networks corresponding to the plurality of non-text modalities are determined.

[0046] In the embodiments of the present disclosure, in the process of step 203, for example, the electronic device determines prediction data corresponding to a plurality of data based on the data in the plurality of non-text modalities and the plurality of codec networks corresponding to the plurality of non-text modalities. For each non-text modality, based on the data in the non-text modality, the prediction data corresponding to the data, and the discrimination network in the codec network corresponding to the non-text modality, a true / false discrimination result is determined. Based on at least one of the true / false discrimination result, the difference between the data and the prediction data, and the characteristic difference between the data and the prediction data, the loss function value of the codec network corresponding to the non-text modality may be determined.

[0047] In an embodiment of the present disclosure, when there are no at least two candidate data with a correlation among data in a plurality of non-text modalities, the process by which an electronic device determines prediction data corresponding to the plurality of data may be, for example, to input, in order, data in each non-text modality into a codec network corresponding to the non-text modality to obtain prediction data corresponding to the data in the non-text modality.

[0048] Here, specifically, for each non-text modality, the codec network corresponding to this non-text modality may include an encoding network and a decoding network. The electronic device may determine an integer sequence corresponding to this data based on the data in this non-text modality and the encoding network, and may determine prediction data corresponding to this data based on the integer sequence and the decoding network.

[0049] Here, specifically, the electronic device inputs the data in this non-text modality into the encoding network to obtain the output data features. When the data features are multi-dimensional feature vectors, the data features are processed by one-dimensional transformation to obtain one-dimensional feature vectors, and the features in the one-dimensional feature vectors are mapped based on the multi-modal word list to obtain an integer sequence corresponding to this data.

[0050] Here, specifically, the electronic device determines prediction data features corresponding to this data based on the multi-modal word list and the integer sequence corresponding to this data, and further inputs the prediction data features corresponding to this data into the decoding network to obtain prediction data corresponding to this data output from the decoding network.

[0051] Here, at least two combinations of candidate data related to each other can include audio - video. Here, by setting the audio - video and the image with audio, the multi - modal large - scale model can extract the interaction features between at least two modalities and improve the accuracy of the extracted features.

[0052] Here, if there are no at least two candidate data related to each other among the data in multiple non - text modalities, there is no relationship between the data in multiple non - text modalities. The data in multiple non - text modalities can be processed individually to obtain prediction data, whereby the data in multiple non - text modalities can be processed in parallel, and the data processing efficiency can be improved.

[0053] In the embodiments of the present disclosure, when there are at least two candidate data related to each other among the data in multiple non - text modalities, for these at least two candidate data, the electronic device inputs the at least two candidate data into the encoding network in the codec network corresponding to the modalities to which they respectively belong, obtains the data features corresponding to the at least two candidate data respectively, performs one - dimensional transformation processing on the at least two data features and addition processing bit by bit to obtain a processed one - dimensional feature vector, performs mapping processing on the features in the processed one - dimensional feature vector based on the multi - modal word list to obtain a processed integer sequence, and determines the prediction data corresponding to the at least two candidate data based on the processed integer sequence and the decoding network in the codec network corresponding to the modalities to which the at least two candidate data belong.

[0054] Specifically, the electronic device can determine prediction data features based on the multimodal word list and the processed integer sequence, and input the prediction data features into the decoding network in the codec network corresponding to the modality to which at least two candidate data belong respectively, so as to obtain prediction data corresponding to at least two candidate data.

[0055] Here, after adding the data features corresponding to at least two related candidate data bit by bit, and performing feature mapping processing based on the multimodal word list, the interaction of features between at least two candidate data can be realized, fragmentation between at least two candidate data can be avoided, and the training efficiency of the codec network can be further improved.

[0056] Here, the plurality of codec networks corresponding to a plurality of non-text modalities can further include a discrimination network. The discrimination network discriminates the authenticity of the prediction data corresponding to this data based on the data in the non-text modality. Here, the authenticity discrimination result is that the prediction data is true, or the prediction data is false. Based on the authenticity discrimination result, the difference in authenticity discrimination between the data and the prediction data can be determined.

[0057] Here, the difference between the data and the prediction data can be determined based on the similarity between the data and the prediction data. Here, the feature difference between the data and the prediction data may be at least one of the difference between the target information in the data and the target information in the prediction data, the difference between the description content of the data and the description content of the prediction data, the difference between the Fourier transform result of the data and the Fourier transform result of the prediction data, etc., and can be set according to actual needs.

[0058] Here, based on the differences in multiple aspects between the data in the non-text modality and the predicted data, by determining the loss function value of the codec network corresponding to the non-text modality, the accuracy of the determined loss function value can be further improved, and the training accuracy of the multiple codec networks can be further improved.

[0059] In step 204, based on the loss function values of the multiple codec networks corresponding to the multiple non-text modalities, the multiple codec networks and the multimodal word list are adjusted to realize collaborative training.

[0060] In the embodiment of the present disclosure, the process in which the electronic device executes step 204 may be, for example, adding the loss function values of the multiple codec networks corresponding to the multiple non-text modalities to obtain a total loss function value, and based on the total loss function value, adjusting the parameters of the multiple codec networks and adjusting the vector corresponding to the integer identifier in the multimodal word list, so as to realize collaborative training.

[0061] Here, based on the total loss function value, by adjusting the parameters of the multiple codec networks and adjusting the vector corresponding to the integer identifier in the multimodal word list, the vector corresponding to the integer identifier in the multimodal word list can reflect the features in the multiple non-text modalities, and the accuracy of the codec network obtained by training can be further improved.

[0062] In step 205, when the training of the multiple codec networks and the multimodal word list is completed, the backbone network is trained based on the multimodal sample reference data and the sample generation data in the target task.

[0063] Here, in steps 201 to 202, and the details of step 205, reference can be made to steps 101 to 102 and step 104 in the embodiment shown in FIG. 1, and detailed descriptions are omitted here.

[0064] In the training method of the multimodal large-scale model according to the embodiment of the present disclosure, first training data and second training data are obtained. The first training data includes data in a plurality of non-text modalities. The second training data includes multimodal sample reference data and sample generation data in a target task. An initial multimodal large-scale model is obtained. The multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multimodal word list. Based on the data in the plurality of non-text modalities and the plurality of codec networks corresponding to the plurality of non-text modalities, loss function values of the plurality of codec networks corresponding to the plurality of non-text modalities are determined. Based on the loss function values of the plurality of codec networks corresponding to the plurality of non-text modalities, the plurality of codec networks and the multimodal word list are adjusted to realize collaborative training. When the training of the plurality of codec networks and the multimodal word list is completed, the backbone network is trained based on the multimodal sample reference data and sample generation data in the target task. Here, the loss function values of the plurality of codec networks are respectively determined, and further the plurality of loss function values are summed up, and by adjusting the plurality of codec networks and the multimodal word list, the training speed and training accuracy of the plurality of codec networks corresponding to the plurality of non-text modalities can be further improved.

[0065] Here, in order to further improve the training accuracy of the backbone network, the electronic device can determine prediction-generated data based on multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network, and then determine the loss function value of the backbone network to perform an adjustment process. As shown in FIG. 3, FIG. 3 is a schematic diagram of a third embodiment according to the present disclosure. The embodiment shown in FIG. 3 can include the following steps.

[0066] In step 301, first training data and second training data are obtained. The first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample-generated data in a target task.

[0067] In step 302, an initial multimodal large-scale model is obtained. The multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multimodal word list.

[0068] In step 303, a cooperative training process is performed on the plurality of codec networks and the multimodal word list based on the data in the plurality of non-text modalities.

[0069] In step 304, prediction-generated data is determined based on multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network.

[0070] In an embodiment of the present disclosure, the multimodal sample reference data may include data in at least two modalities among an audio modality, a silent video modality, an image modality, and a text modality, and the sample generation data may include data in at least one modality among an audio modality, a silent video modality, an image modality, and a text modality.

[0071] Here, through the flexible setting of data in various modalities within the multimodal sample reference data and the sample generation data, the multimodal large-scale model obtained through training can be applied to various target tasks, and the application scenarios of the multimodal large-scale model are expanded.

[0072] In an embodiment of the present disclosure, the process of the electronic device executing step 304 may be, for example, determining a multimodal integer sequence combination based on the multimodal sample reference data and the encoding network in a plurality of codec networks corresponding to a plurality of non-text modalities. In the multimodal integer sequence combination, integer sequences in different modalities are distinguished by modality markers, inputting the multimodal integer sequence combination into a backbone network to obtain a predicted integer sequence or a predicted integer sequence combination output from the backbone network, and inputting the predicted integer sequence in the predicted integer sequence or the predicted integer sequence combination into the decoding network in the codec network corresponding to the modality to which it belongs to obtain predicted generation data.

[0073] Here, the multimodal sample reference data may include sample reference data in a plurality of candidate non-text modalities and sample text data in the text modality. Accordingly, the process of determining the electronic device multimodal integer sequence combination may be, for example, for each candidate non-text modality, based on the sample reference data in the candidate non-text modality and the encoding network in the codec network corresponding to the candidate non-text modality, determining the integer sequence in the candidate non-text modality, and according to the modality marker, splicing the integer sequences in the plurality of candidate non-text modalities and the integer sequence corresponding to the sample text data to obtain the multimodal integer sequence combination.

[0074] Here, still, the electronic device can set a text word list for the text, and the text word list includes integer identifiers and corresponding words. The electronic device performs word segmentation processing on the sample text data to obtain a plurality of words, searches the text word list based on the plurality of words to obtain the integer identifiers corresponding to the plurality of words, and then combines the integer identifiers to obtain the integer sequence corresponding to the sample text data.

[0075] Here, when the target task is an image generation task, a video generation task, an audio generation task, or a text generation task, the predicted generation data output from the multimodal large-scale model is data in a single modality. Accordingly, the backbone network can output a predicted integer sequence. When the target task is a multi-output task, for example, a graphic generation task, an audio-video generation task, etc., the predicted generation data output from the multimodal large-scale model is data in a multimodality. Accordingly, the backbone network can output a combination of predicted integer sequences.

[0076] Here, for the predicted integer sequence output from the backbone network, a modality marker can be set. For each predicted integer sequence in the combination of predicted integer sequences output from the backbone network, a modality marker can be set. Based on the modality marker, a decoding network corresponding to a non-text modality for decoding the predicted integer sequence can be determined, and the target decoding process can be performed.

[0077] Here, by the modality marker for the integer sequence in a plurality of candidate non-text modalities and the word sequence corresponding to the sample text data, the backbone network can distinguish integer sequences of different modalities and perform learning processing, and can further improve the training speed and training accuracy of the backbone network.

[0078] In an embodiment of the present disclosure, when there are at least two pieces of sample reference data in candidate non-text modalities that are related in the multi-modal sample reference data, for example, the multi-modal sample reference data includes audio-video data, the electronic device determines the data features of the at least two pieces of sample reference data in the candidate non-text modalities based on the encoding network in the codec network in multiple modalities. When the data features are multi-dimensional feature vectors, the data features are processed by one-dimensional transformation to obtain one-dimensional feature vectors, the at least two one-dimensional feature vectors are added bit by bit to obtain a processed one-dimensional feature vector, the features in the processed one-dimensional feature vector are mapped based on the multi-modal word list to obtain an integer sequence, and the integer sequence can be subjected to combined modality marker processing. Here, the combined modality marker is, for example, an audio-video marker, etc.

[0079] In an embodiment of the present disclosure, when there is sample generation data in at least two candidate non-text modalities that are related among the sample generation data, when the backbone network outputs a predicted integer sequence combination, for at least two candidate non-text modalities that are related, an integer sequence of one combined modality may be generated, and this integer sequence is input into the decoding network in the codec network corresponding to at least two candidate non-text modalities, and predicted generation data in at least two candidate non-text modalities may be obtained.

[0080] Here, through the generation process of the integer sequence of the combined modality, the multi-modal large-scale model can realize the synchronous generation process of the predicted generation data in at least two candidate non-text modalities.

[0081] In step 305, based on the sample generation data and the predicted generation data, the loss function value of the backbone network is determined.

[0082] In an embodiment of the present disclosure, the electronic device can determine the loss function value of the backbone network based on the sample generation data, the predicted generation data, and the loss function of the multi-modal large-scale model.

[0083] In step 306, based on the loss function value, the parameters of the backbone network are adjusted to realize training.

[0084] Here, for the details of steps 301 to 303, reference can be made to steps 101 to 103 in the embodiment shown in FIG. 1, and detailed descriptions are omitted here.

[0085] In the method for training a multimodal large-scale model according to an embodiment of the present disclosure, first training data and second training data are obtained. The first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data for a target task. An initial multimodal large-scale model is obtained. The multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multimodal word list. Based on the data in the plurality of non-text modalities, a cooperative training process is performed on the plurality of codec networks and the multimodal word list. Based on the multimodal sample reference data, the plurality of codec networks corresponding to the plurality of non-text modalities, and the backbone network, predicted generation data is determined. Based on the sample generation data and the predicted generation data, a loss function value of the backbone network is determined. Based on the loss function value, the parameters of the backbone network are adjusted to realize training. Here, based on the multimodal sample reference data, the plurality of codec networks corresponding to the plurality of non-text modalities, and the backbone network, the predicted generation data is determined, and further, the loss function value of the backbone network is determined and adjustment processing is performed, so that the training speed and training accuracy of the backbone network can be improved.

[0086] The following examples will be used for explanation. As shown in FIG. 4, it is a schematic diagram of the training of a plurality of codec networks in a multimodal large-scale model. In FIG. 4, the following steps can be included. In step 401, the input is fed into the encoding network in the audio 1D-CNN (the codec network corresponding to the audio modality) to obtain the output audio features, the video is fed into the encoding network in the 3D-CNN (the codec network corresponding to the video modality) to obtain the output video features, and the image is fed into the encoding network in the 2D-CNN (the codec network corresponding to the image modality) to obtain the output image features. In step 402, the image features and the video features are each dimensionally transformed into one dimension to obtain a one-dimensional feature vector corresponding to the image and a one-dimensional feature vector corresponding to the video. When the audio is the audio corresponding to the video, that is, when the audio and the video are combined to obtain an audiovisual, the audio features (one-dimensional feature vector) corresponding to the audio and the one-dimensional feature vector corresponding to the video are subjected to a sum (addition bit by bit) process to obtain the processed one-dimensional feature vector. In step 403, the one-dimensional feature vector corresponding to the image and the processed one-dimensional feature vector are each mapped based on a codebook (a multimodal word list) to obtain an integer sequence corresponding to the image and a processed integer sequence, and this processed integer sequence can be used as an integer sequence corresponding to the audio and an integer sequence corresponding to the video, respectively. In step 404, based on the integer sequence corresponding to the audio, the integer sequence corresponding to the image, the integer sequence corresponding to the video, the decoding network of the 1D-CNN, the decoding network of the 2D-CNN, and the decoding network of the 3D-CNN, a predicted audio, a predicted image, and a predicted video for the collaborative training process of the 1D-CNN, 2D-CNN, and 3D-CNN are determined and obtained.

[0087] FIG. 5 is a schematic diagram of a fourth embodiment according to the present disclosure. Note that the method for processing a target task according to an embodiment of the present disclosure is applicable to a processing apparatus for a target task, and this apparatus may be configured in an electronic device so that the electronic device can execute the processing function of the target task. In the following embodiments, it will be described by taking the execution subject as an electronic device as an example.

[0088] Here, the electronic device may be any device having a computing function such as a personal computer (abbreviated as PC), a mobile terminal, a server, etc. The mobile terminal may be, for example, a vehicle-mounted device, a mobile phone, a tablet, a personal digital assistant, a wearable device, etc., and may be a hardware device equipped with various operating systems, a touch screen, and / or a display.

[0089] Here, the processing apparatus for the target task may be software in the electronic device, for example, the processing software for the target task. In the following embodiments, it will be described by taking the processing apparatus for the target task as an electronic device as an example.

[0090] As shown in FIG. 5, the method for processing this target task may include the following steps.

[0091] In step 501, a target task including data in at least two modalities is acquired.

[0092] In an embodiment of the present disclosure, the target task may include at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multimodality understanding task. Here, the data in at least two modalities included in the target task is the data necessary for executing the target task.

[0093] Here, taking the case where the target task is an image generation task as an example, assume that the image generation task performs image generation processing based on an image and text. Accordingly, the image generation task may include data in the image modality and data in the text modality. Taking the case where the target task is a video generation task as an example, assume that the video generation task performs video generation processing based on an image and audio. Accordingly, the video generation task may include data in the image modality and data in the audio modality.

[0094] Here, through the setting of multiple target tasks, the multimodal large-scale model can process data in multiple target tasks, and the task application scenario of the multimodal large-scale model is extended.

[0095] In step 502, obtain a multimodal large-scale model obtained based on the training method of the multimodal large-scale model in any of the embodiments shown in FIGS. 1 to 3.

[0096] In an embodiment of the present disclosure, the multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities, and the plurality of codec networks perform encoding and decoding processing based on the same multimodal word list.

[0097] Here, there may be a plurality of codec networks. For example, it may include a one-dimensional codec network, a two-dimensional codec network, a three-dimensional codec network, etc. The one-dimensional codec network is a codec network corresponding to the audio modality, the two-dimensional codec network is a codec network corresponding to the image modality, and the three-dimensional codec network is a codec network corresponding to the video modality.

[0098] In step 503, input data in at least two modalities into a multimodal large-scale model, and obtain the generated data output from the multimodal large-scale model.

[0099] In an embodiment of the present disclosure, taking as an example that at least two modalities are respectively an image modality, an audio modality, a video modality, and a text modality, the process of the electronic device executing step 503 is, for example, inputting an image in the image modality into a two-dimensional codec network to determine an integer sequence corresponding to the image, inputting audio in the audio modality into a one-dimensional codec network to determine an integer sequence corresponding to the audio, inputting a video in the video modality into a three-dimensional codec network to determine an integer sequence corresponding to the video, performing word segmentation processing on the text in the text modality, and determining and obtaining an integer sequence corresponding to the text based on integer identifiers corresponding to a plurality of words in the text word list. Perform modality marking and splicing processing on the integer sequence corresponding to the image, the integer sequence corresponding to the audio, the integer sequence corresponding to the video, and the integer sequence corresponding to the text to obtain a combined processed integer sequence. Input the combined processed integer sequence into a backbone network to obtain a predicted integer sequence or a predicted combined integer sequence output from the backbone network. Based on the predicted integer sequence or the predicted combined integer sequence and the decoding network in a plurality of codec networks, it may also be possible to determine the generated data.

[0100] In an embodiment of the present disclosure, there may be a correlation between data in at least two modalities in a target task. Assume that the target task includes data in an audio modality, data in a video modality, and data in an image modality. Here, there is a correlation between the data in the audio modality and the data in the video modality. The electronic device inputs the data in the audio modality into an encoding network in a codec network corresponding to the audio modality to obtain the output audio data features. Here, the audio data features are one-dimensional feature vectors. The data in the video modality is input into an encoding network in a codec network corresponding to the video modality to obtain the output video data features. The video data features are processed by one-dimensional transformation to obtain one-dimensional feature vectors. The two one-dimensional feature vectors are added bit by bit, and then feature mapping processing is performed based on a multimodal word list to obtain an integer sequence. For this integer sequence, combined modality marker processing, that is, audio-video marker processing, can be performed.

[0101] In an embodiment of the present disclosure, when the target task is to generate data in at least two non-text modalities with a correlation, when the backbone network outputs a predicted integer sequence combination, for at least two candidate non-text modalities with a correlation, an integer sequence of one combined modality can be generated. This integer sequence is respectively input into a decoding network in a codec network corresponding to at least two candidate non-text modalities, and predicted generated data in at least two candidate non-text modalities may be obtained. Thereby, synchronous generation processing of predicted generated data in at least two non-text modalities with a correlation is realized.

[0102] In the method for processing a target task according to an embodiment of the present disclosure, a target task including data in at least two modalities is obtained, a multimodal large-scale model obtained based on the training method of the multimodal large-scale model in any one of the embodiments shown in FIGS. 1 to 3 is obtained, data in at least two modalities is input into the multimodal large-scale model, and generated data output from the multimodal large-scale model is obtained. Here, a plurality of codec networks in the multimodal large-scale model perform encoding and decoding processes based on the same multimodal word list, thereby realizing unified modeling of data in a plurality of modalities, and improving the accuracy of the obtained generated data determined.

[0103] To implement the above embodiment, the present disclosure further provides a training apparatus for a multimodal large-scale model. As shown in FIG. 6, FIG. 6 is a schematic diagram of a fifth embodiment according to the present disclosure. The training apparatus 60 for the multimodal large-scale model may include a first acquisition module 601, a second acquisition module 602, a first training processing module 603, and a second training processing module 604.

[0104] Here, the first acquisition module 601 acquires the first training data and the second training data. The first training data includes data in a plurality of non-text modalities. The second training data includes multi-modal sample reference data and sample generation data for the target task. The second acquisition module 602 acquires an initial multi-modal large-scale model. The multi-modal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multi-modal word list. The first training processing module 603 performs cooperative training processing on the plurality of codec networks and the multi-modal word list based on the data in the plurality of non-text modalities. When the training of the plurality of codec networks and the multi-modal word list is completed, the second training processing module 604 performs training processing on the backbone network based on the multi-modal sample reference data and sample generation data for the target task.

[0105] As a possible implementation form of an embodiment of the present disclosure, the first training processing module 603 includes a first determination unit and a first adjustment processing unit. The first determination unit determines a loss function value of a plurality of codec networks corresponding to a plurality of non-text modalities based on the data in the plurality of non-text modalities and the plurality of codec networks corresponding to the plurality of non-text modalities. The first adjustment processing unit adjusts the plurality of codec networks and the multi-modal word list based on the loss function value of the plurality of codec networks corresponding to the plurality of non-text modalities to realize cooperative training.

[0106] As a possible implementation form of an embodiment of the present disclosure, the first determination unit includes a first determination subunit, a second determination subunit, and a third determination subunit. The first determination subunit determines prediction data corresponding to a plurality of pieces of the data based on the data in a plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities. The second determination subunit determines a true / false discrimination result for each non-text modality based on the data in the non-text modality, the prediction data corresponding to the data, and a discrimination network in the codec network corresponding to the non-text modality. The third determination subunit determines a loss function value of the codec network corresponding to the non-text modality based on at least one of the true / false discrimination result, the difference between the data and the prediction data, and the characteristic difference between the data and the prediction data.

[0107] As a possible implementation form of an embodiment of the present disclosure, specifically, the first determination subunit determines whether there are at least two candidate data with a relevant relationship among the data in a plurality of non-text modalities. The at least two candidate data belong to different modalities. When there are no at least two candidate data among the data in a plurality of non-text modalities, for each non-text modality in turn, the data in the non-text modality is input into the codec network corresponding to the non-text modality to obtain prediction data corresponding to the data in the non-text modality.

[0108] As a possible implementation form of an embodiment of the present disclosure, when at least two of the candidate data exist in data in a plurality of non-text modalities, the first determination subunit further inputs at least two of the candidate data into an encoding network in a codec network corresponding to the modality to which each belongs, obtains data features corresponding to at least two of the candidate data respectively, performs one-dimensional transformation processing on at least two of the data features and addition processing for each bit to obtain a processed one-dimensional feature vector, performs mapping processing on features in the processed one-dimensional feature vector based on the multimodal word list to obtain a processed integer sequence, and determines prediction data corresponding to at least two of the candidate data based on the processed integer sequence and a decoding network in a codec network corresponding to the modality to which at least two of the candidate data belong.

[0109] As a possible implementation form of an embodiment of the present disclosure, specifically, the first adjustment processing unit adds loss function values of a plurality of codec networks corresponding to a plurality of non-text modalities to obtain a total loss function value, adjusts parameters of the plurality of codec networks based on the total loss function value, and adjusts a vector corresponding to an integer identifier in the multimodal word list, thereby realizing cooperative training.

[0110] As a possible implementation form of an embodiment of the present disclosure, the non-text modality includes at least one of an audio modality, a silent video modality, and an image modality. The codec network corresponding to the audio modality is a one-dimensional codec network, the codec network corresponding to the silent video modality is a three-dimensional codec network, and the codec network corresponding to the image modality is a two-dimensional codec network.

[0111] As a possible implementation form of an embodiment of the present disclosure, processing an image in an image modality based on an encoding network in the two-dimensional codec network includes inputting the image into the encoding network in the two-dimensional codec network to obtain two-dimensional image features, performing a one-dimensional transformation process on the two-dimensional image features to obtain a transformed one-dimensional feature vector, and performing a mapping process on the features in the one-dimensional feature vector based on the multimodal word list to obtain an integer sequence corresponding to the image.

[0112] As a possible implementation form of an embodiment of the present disclosure, a combination of at least two candidate data with a relevant relationship includes audio and video.

[0113] As a possible implementation form of an embodiment of the present disclosure, the second training processing module 604 includes a second determination unit, a third determination unit, and a second adjustment processing unit. The second determination unit determines prediction generation data based on the multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network. The third determination unit determines a loss function value of the backbone network based on the sample generation data and the prediction generation data. The second adjustment processing unit adjusts the parameters of the backbone network based on the loss function value to realize training.

[0114] As a possible implementation form of an embodiment of the present disclosure, specifically, the second determination unit determines a multimodal integer sequence combination based on the multimodal sample reference data and an encoding network in a plurality of codec networks corresponding to a plurality of non-text modalities. In the multimodal integer sequence combination, integer sequences in different modalities are distinguished by modality markers. The multimodal integer sequence combination is input into the backbone network to obtain a predicted integer sequence or a predicted integer sequence combination output from the backbone network. The predicted integer sequence in the predicted integer sequence or the predicted integer sequence combination is input into a decoding network in a codec network corresponding to the modality to which it belongs to obtain the predicted generated data.

[0115] As a possible implementation form of an embodiment of the present disclosure, the multimodal sample reference data includes sample reference data in a plurality of candidate non-text modalities and sample text data in the text modality. The second determination unit further determines, for each candidate non-text modality, an integer sequence in the candidate non-text modality based on the sample reference data in the candidate non-text modality and an encoding network in a codec network corresponding to the candidate non-text modality, and splices the integer sequences in the plurality of candidate non-text modalities and the integer sequence corresponding to the sample text data according to modality markers to obtain the multimodal integer sequence combination.

[0116] As a possible implementation form of an embodiment of the present disclosure, the multimodal sample reference data includes data in at least two modalities among audio modality, silent video modality, image modality, and text modality, and the sample generated data includes data in at least one modality among audio modality, silent video modality, image modality, and text modality.

[0117] As a possible implementation form of the embodiments of the present disclosure, the target task includes at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multimodality understanding task.

[0118] In the training device of the multimodal large-scale model according to the embodiments of the present disclosure, first training data and second training data are obtained. The first training data includes data in a plurality of non-text modalities, and the second training data includes multimodal sample reference data and sample generation data for the target task. An initial multimodal large-scale model is obtained. The multimodal large-scale model includes a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities. The plurality of codec networks perform encoding and decoding processes based on the same multimodal word list. Based on the data in the plurality of non-text modalities, a cooperative training process is performed on the plurality of codec networks and the multimodal word list. When the training of the plurality of codec networks and the multimodal word list is completed, the backbone network is trained based on the multimodal sample reference data and sample generation data for the target task. Here, by performing encoding and decoding processes and a cooperative training process based on the same multimodal word list by the plurality of codec networks, it is possible to avoid separately training the word lists of different modalities during the training process, and reduce the difficulty and cost of model training.

[0119] To implement the above embodiments, the present disclosure further provides a processing device for the target task. As shown in FIG. 7, FIG. 7 is a schematic diagram of the sixth embodiment according to the present disclosure. The processing device 70 for the target task may include a first acquisition module 701, a second acquisition module 702, and a third acquisition module 703.

[0120] Here, the first acquisition module 701 acquires a target task including data in at least two modalities, the second acquisition module 702 acquires a multimodal large-scale model obtained based on the training method of the multimodal large-scale model in any one of the embodiments shown in FIGS. 1 to 3, and the third acquisition module 703 inputs the data in the at least two modalities into the multimodal large-scale model to acquire generated data output from the multimodal large-scale model.

[0121] As a possible implementation form of the embodiment of the present disclosure, the target task includes at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multimodal understanding task.

[0122] In the target task processing device according to the embodiment of the present disclosure, a target task including data in at least two modalities is acquired, a multimodal large-scale model obtained based on the training method of the multimodal large-scale model in any one of the embodiments shown in FIGS. 1 to 3 is acquired, the data in at least two modalities is input into the multimodal large-scale model to acquire generated data output from the multimodal large-scale model. Here, a plurality of codec networks in the multimodal large-scale model perform encoding and decoding processing based on the same multimodal word list, thereby realizing unified modeling of data in a plurality of modalities, and improving the accuracy of the determined generated data.

[0123] It should be noted that in the technical solution of the present disclosure, the acquisition, storage, application, etc. of relevant user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0124] According to the embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium. According to an embodiment of the present disclosure, the present disclosure provides a computer program, and when the computer program is executed by a processor, the training method of the multimodal large-scale model or the processing method of the target task proposed by the present disclosure is realized.

[0125] FIG. 8 is a schematic block diagram of an exemplary electronic device 800 for executing an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices such as personal digital processing, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure required.

[0126] As shown in FIG. 8, the electronic device 800 includes a computing unit 801 that executes various appropriate operations and processes according to a computer program / instructions stored in a read-only memory (ROM) 802 or a computer program / instructions loaded from a storage unit 806 into a random access memory (RAM) 803. Various programs and data necessary for the operation of the electronic device 800 may also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0127] The multiple components of the electronic device 800 are connected to the I / O interface 805 and include an input unit 806 such as a keyboard and a mouse, an output unit 807 such as various types of displays and speakers, a storage unit 808 such as a magnetic disk and an optical disk, and a communication unit 809 such as a network card, a modem, and a wireless communication transceiver. The communication unit 809 enables the electronic device 800 to improve information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0128] The computing unit 801 may be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units for various machine learning models algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes each of the methods and processes described above, for example, the training method of the multimodal large-scale model or the processing method of the target task. For example, in some embodiments, the training method of the multimodal large-scale model or the processing method of the target task can be realized as a computer software program tangibly included in a machine-readable medium such as the storage unit 806, and part or all of the computer program / instructions may be loaded and / or installed into the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program / instructions are instruction-loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the multimodal large-scale model or the processing method of the target task described above may be executed. Alternatively, in other embodiments, the computing unit 801 may be arranged in any other suitable manner (e.g., via firmware) to execute the training method of the multimodal large-scale model or the processing method of the target task.

[0129] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented by one or more computer programs / instructions, which can be executed and / or interpreted in a programmable system including at least one programmable processor, where the programmable processor can be a special-purpose or general-purpose programmable processor, and which receives data and instructions from a storage system, at least one input device, and at least one output device, and can transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] The program code for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus so that, when executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes, or is capable of storing, a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0132] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with a user, for example, feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein can be implemented on a computing system that includes back-end components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, where the user can interact with embodiments of the systems and techniques described herein via the graphical user interface or the web browser), or a computing system that includes any combination of such back-end components, middleware components, and front-end components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0134] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. A client-server relationship is created by computer programs / instructions that are executed on corresponding computers and have a client-server relationship with each other. The server can be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0135] It should be understood that the steps can be rearranged, added, or deleted using the various forms of flow shown above. For example, each step described in this disclosure can be executed in parallel, sequentially, or in a different order, but this specification is not limited herein as long as the technical solutions disclosed in this disclosure can achieve the desired results.

[0136] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall all be included within the protection scope of the present disclosure.

Claims

1. A method for training a multimodal large-scale model, comprising: obtaining first training data and second training data, the first training data including data in a plurality of non-text modalities, and the second training data including multimodal sample reference data and sample generation data for a target task; obtaining an initial multimodal large-scale model, the multimodal large-scale model including a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities, the plurality of codec networks performing encoding and decoding processes based on a same multimodal word list; performing a joint training process for the codec networks and the multimodal word list based on data in a plurality of non-text modalities; training the backbone network based on multimodal sample reference data and sample generation data for the target task when training of the codec networks and the multimodal word list is completed; A method for training large-scale multimodal models, including:

2. performing a joint training process on the plurality of codec networks and the multimodal word list based on data in the plurality of non-text modalities, determining loss function values ​​of the multiple codec networks corresponding to the multiple non-text modalities based on the data in the multiple non-text modalities and the multiple codec networks corresponding to the multiple non-text modalities; According to loss function values ​​of a plurality of codec networks corresponding to a plurality of non-text modalities, adjusting the plurality of codec networks and the multimodal word list to realize joint training; The method for training a multimodal large scale model according to claim 1 , comprising:

3. The step of determining loss function values ​​of a plurality of codec networks corresponding to a plurality of non-text modalities based on the data in the plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities includes: determining predicted data corresponding to a plurality of said data based on data in a plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities; determining, for each non-text modality, a true / false discrimination result based on the data in the non-text modality, the predicted data corresponding to the data, and a discrimination network in the codec network corresponding to the non-text modality; determining a loss function value of a codec network corresponding to the non-text modality based on at least one of the authenticity determination result, the difference between the data and the predicted data, and a characteristic difference between the data and the predicted data; The method for training a multimodal large scale model according to claim 2 , comprising:

4. determining predicted data corresponding to the plurality of data based on data in the plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities, determining whether there are at least two candidate data items related to each other among the data items in a plurality of non-text modalities, the at least two candidate data items belonging to different modalities; if there are not at least two candidate data among the data in a plurality of non-text modalities, inputting the data in the non-text modality into a codec network corresponding to the non-text modality in turn for each non-text modality to obtain predicted data corresponding to the data in the non-text modality; The method for training a multimodal large scale model according to claim 3, comprising:

5. determining predicted data corresponding to the plurality of data based on data in the plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities, When at least two pieces of candidate data exist among data in a plurality of non-text modalities, inputting the at least two pieces of candidate data into an encoding network in a codec network corresponding to each of the modalities to which the candidate data belongs, and acquiring data features corresponding to the at least two pieces of candidate data; a step of one-dimensionally transforming and bit-wise adding at least two of the data features to obtain a processed one-dimensional feature vector; mapping features in the processed one-dimensional feature vector based on the multimodal word list to obtain a processed integer sequence; determining prediction data corresponding to at least two of the candidate data based on the processed integer sequence and a decoding network in a codec network corresponding to a modality to which the at least two candidate data belong; The method for training a multimodal large scale model according to claim 4, further comprising:

6. According to loss function values ​​of the codec networks corresponding to the non-text modalities, adjusting the codec networks and the multimodal word list to realize joint training is performed, summing loss function values ​​of multiple codec networks corresponding to multiple non-text modalities to obtain a total loss function value; adjusting parameters of the plurality of codec networks based on the total loss function value, and adjusting vectors corresponding to integer identifiers in the multimodal word list to achieve joint training; The method for training a multimodal large scale model according to claim 2 , comprising:

7. the non-text modality includes at least one of an audio modality, a silent video modality, and an image modality; The method for training a multimodal large scale model according to claim 1 , wherein the codec network corresponding to the audio modality is a one-dimensional codec network, the codec network corresponding to the silent video modality is a three-dimensional codec network, and the codec network corresponding to the image modality is a two-dimensional codec network.

8. Processing an image in an image modality based on an encoding network in the two-dimensional codec network includes: inputting the image into an encoding network in the two-dimensional codec network to obtain two-dimensional image features; transforming the two-dimensional image features one-dimensionally to obtain a transformed one-dimensional feature vector; mapping features in the one-dimensional feature vector based on the multimodal word list to obtain a sequence of integers corresponding to the image; The method for training a multimodal large scale model according to claim 7, comprising:

9. The method for training a multimodal large-scale model according to claim 4 , wherein the combination of at least two candidate data in a related relationship includes audio-video.

10. training the backbone network based on multimodal sample reference data and sample generation data of the target task, determining predicted generation data based on the multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network; determining a loss function value for the backbone network based on the sample generation data and the predicted generation data; adjusting parameters of the backbone network based on the loss function value to achieve training; The method for training a multimodal large scale model according to claim 1 , comprising:

11. determining predicted generation data based on the multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network, determining a multimodal integer sequence combination based on the multimodal sample reference data and a coding network in a plurality of codec networks corresponding to a plurality of non-text modalities, in which integer sequences in different modalities in the multimodal integer sequence combination are distinguished by a modality marker; inputting the multi-modal integer sequence combination into the backbone network to obtain a predicted integer sequence or a predicted integer sequence combination output from the backbone network; inputting the predicted integer sequence or the predicted integer sequence in the predicted integer sequence combination into a decoding network in a codec network corresponding to the modality to which it belongs, to obtain the predicted generated data; The method for training a multimodal large scale model according to claim 10, comprising:

12. the multimodal sample reference data includes sample reference data in a plurality of candidate non-text modalities and sample text data in a text modality; determining a multimodal integer sequence combination based on the multimodal sample reference data and a coding network in a plurality of codec networks corresponding to a plurality of non-text modalities, comprising: for each candidate non-text modality, determining a sequence of integers in said candidate non-text modality based on sample reference data in said candidate non-text modality and on an encoding network in a codec network corresponding to said candidate non-text modality; splicing integer sequences in a plurality of candidate non-text modalities and integer sequences corresponding to the sample text data according to modality markers to obtain the multi-modal integer sequence combination; The method for training a multimodal large scale model according to claim 11, comprising:

13. the multimodal sample reference data includes data in at least two of the following modalities: an audio modality, a silent video modality, an image modality, and a text modality; 2. The method of claim 1, wherein the sample generated data includes data in at least one of the following modalities: audio, silent video, image, and text.

14. The method of claim 1 , wherein the target task includes at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multi-modality understanding task.

15. A method for processing a target task, comprising the steps of: acquiring a target task including data in at least two modalities; Obtaining a multimodal large-scale model obtained according to the method for training a multimodal large-scale model according to claim 1; inputting data from the at least two modalities into the multimodal large-scale model to obtain generated data output from the multimodal large-scale model; and a method for processing the target task, including:

16. The method of claim 15 , wherein the target task includes at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multi-modality understanding task.

17. A multimodal large scale model training apparatus, comprising: a first acquisition module for acquiring first training data and second training data, the first training data including data in a plurality of non-text modalities, and the second training data including multimodal sample reference data and sample generation data for a target task; a second acquisition module for acquiring an initial multimodal large-scale model, the multimodal large-scale model including a backbone network and a plurality of codec networks corresponding to a plurality of non-text modalities, the plurality of codec networks performing encoding and decoding processes based on a same multimodal word list; a first training module for performing a joint training process on the plurality of codec networks and the multimodal word list based on data in a plurality of non-text modalities; a second training module for training the backbone network based on multimodal sample reference data and sample generation data in the target task when training of the plurality of codec networks and the multimodal word list is completed; A training device for multimodal large-scale models, including:

18. the first training processing module includes a first determination unit and a first adjustment processing unit; The first determination unit determines loss function values ​​of a plurality of codec networks corresponding to the plurality of non-text modalities based on data in the plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities; The apparatus for training a multimodal large-scale model according to claim 17, wherein the first training unit is configured to train a plurality of codec networks and the multimodal word list based on loss function values ​​of the codec networks corresponding to a plurality of non-text modalities to realize joint training.

19. the first determining unit includes a first determining subunit, a second determining subunit and a third determining subunit; The first determination subunit determines predicted data corresponding to the plurality of data based on data in a plurality of non-text modalities and a plurality of codec networks corresponding to the plurality of non-text modalities; The second determination subunit determines, for each non-text modality, a true / false determination result based on the data in the non-text modality, the prediction data corresponding to the data, and a discrimination network in the codec network corresponding to the non-text modality; The multimodal large scale model training apparatus of claim 18, wherein the third determination subunit determines a loss function value of a codec network corresponding to the non-text modality based on at least one of the truth / falseness determination result, a difference between the data and the predicted data, and a feature difference between the data and the predicted data.

20. The first determination subunit: determining whether there are at least two candidate data items related to each other among the data items in a plurality of non-text modalities, the at least two candidate data items belonging to different modalities; 20. The apparatus for training a multimodal large scale model according to claim 19, wherein, if there are not at least two candidate data among the data in a plurality of non-text modalities, for each non-text modality in turn, the data in the non-text modality is input to a codec network corresponding to the non-text modality to obtain predicted data corresponding to the data in the non-text modality.

21. The first determining subunit further comprises: When at least two pieces of candidate data exist among the data in a plurality of non-text modalities, inputting the at least two pieces of candidate data into an encoding network in a codec network corresponding to each of the modalities to which the candidate data belongs, and obtaining data features corresponding to the at least two pieces of candidate data, respectively; a one-dimensional transformation process is performed on at least two of the data features and a bit-by-bit addition process is performed to obtain a one-dimensional feature vector after the transformation process; mapping features in the processed one-dimensional feature vector based on the multimodal word list to obtain a processed integer sequence; The apparatus for training a multimodal large scale model according to claim 20, further comprising: determining predicted data corresponding to at least two of the candidate data based on the processed integer sequence and a decoding network in a codec network corresponding to the modality to which the at least two of the candidate data belong.

22. The first adjustment processing unit, Adding the loss function values ​​of the multiple codec networks corresponding to the multiple non-text modalities to obtain a total loss function value; 20. The apparatus for training a multimodal large scale model of claim 18, further comprising: adjusting parameters of a plurality of said codec networks based on said total loss function value; and adjusting vectors corresponding to integer identifiers in said multimodal word list, thereby achieving joint training.

23. the non-text modality includes at least one of an audio modality, a silent video modality, and an image modality; 18. The apparatus for training a multimodal large scale model according to claim 17, wherein the codec network corresponding to the audio modality is a one-dimensional codec network, the codec network corresponding to the silent video modality is a three-dimensional codec network, and the codec network corresponding to the image modality is a two-dimensional codec network.

24. Processing an image in an image modality based on an encoding network in the two-dimensional codec network includes: inputting the image into an encoding network in the two-dimensional codec network to obtain two-dimensional image features; transforming the two-dimensional image features one-dimensionally to obtain a transformed one-dimensional feature vector; mapping features in the one-dimensional feature vector based on the multimodal word list to obtain a sequence of integers corresponding to the image; 24. The multimodal large scale model training apparatus of claim 23, comprising:

25. 22. The apparatus for training a multimodal large scale model according to claim 20 or 21, wherein the combination of at least two candidate data in a related relationship comprises audio-video.

26. the second training processing module includes a second determining unit, a third determining unit and a second adjusting processing unit; The second determining unit determines predicted generation data based on the multimodal sample reference data, a plurality of codec networks corresponding to a plurality of non-text modalities, and the backbone network; The third determination unit determines a loss function value of the backbone network based on the sample generation data and the predicted generation data; The multi-modal large-scale model training apparatus according to claim 17 , wherein the second tuning processing unit adjusts parameters of the backbone network based on the loss function value to achieve training.

27. The second determination unit: determining a multimodal integer sequence combination based on the multimodal sample reference data and a coding network in a plurality of codec networks corresponding to a plurality of non-text modalities, in which integer sequences in different modalities are distinguished by modality markers in the multimodal integer sequence combination; inputting the multi-modal integer sequence combination into the backbone network to obtain a predicted integer sequence or a predicted integer sequence combination output from the backbone network; 27. The apparatus for training a multimodal large scale model according to claim 26, further comprising: inputting the predicted integer sequence or the predicted integer sequence in the predicted integer sequence combination into a decoding network in a codec network corresponding to the modality to which it belongs to obtain the predicted generation data.

28. the multimodal sample reference data includes sample reference data in a plurality of candidate non-text modalities and sample text data in a text modality; The second determination unit further comprises: for each candidate non-text modality, determining a sequence of integers in the candidate non-text modality based on sample reference data in the candidate non-text modality and an encoding network in a codec network corresponding to the candidate non-text modality; 28. The apparatus for training a multimodal large scale model according to claim 27, further comprising: splicing integer sequences in a plurality of candidate non-text modalities and integer sequences corresponding to the sample text data according to modality markers to obtain the multimodal integer sequence combinations.

29. the multimodal sample reference data includes data in at least two of the following modalities: an audio modality, a silent video modality, an image modality, and a text modality; 27. An apparatus for training a multimodal large scale model according to claim 17 or 26, wherein the sample generated data includes data in at least one of the following modalities: audio, silent video, image and text modalities.

30. 20. The apparatus for training a multimodal large scale model of claim 17, wherein the target task comprises at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multi-modality understanding task.

31. A processing device for a target task, a first acquisition module for acquiring a target task including data in at least two modalities; a second acquisition module for acquiring a multimodal large-scale model obtained according to the method for training a multimodal large-scale model according to any one of claims 1 to 14; a third acquisition module for inputting data in the at least two modalities into the multimodal large-scale model and acquiring generated data output from the multimodal large-scale model; a target task processing device including:

32. 32. The target task processing apparatus of claim 31, wherein the target task includes at least one of an image generation task, a video generation task, an audio generation task, a text generation task, and a multi-modality understanding task.

33. 1. An electronic device comprising: At least one processor; a memory communicatively connected to the at least one processor; An electronic device, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform a method according to any one of claims 1 to 14 or any one of claims 15 to 16.

34. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: A non-transitory computer readable storage medium, the computer instructions causing a computer to perform the method of any of claims 1 to 14 or any of claims 15 to 16.

35. A computer program comprising: A computer program product, which when executed by a processor, results in the method according to any one of claims 1 to 14 or the method according to any one of claims 15 to 16.

Citation Information

Patent Citations

  • Defect detection system, method and storage medium for display device

    JP2024022577A

  • Enhanced user experience through bi-directional audio and visual signal generation

    US20220343543A1