Multi-modal interaction-oriented large language model joint training method and system
By disassembling the multimodal data set from the original video and training the large language model in stages, the problem of inaccurate understanding of modal information in virtual human technology is solved, and more natural human-computer interaction is achieved.
Patent Information
- Application Number
- CN202510395320.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing virtual human technology cannot effectively understand and respond to users' multiple modal information, resulting in inaccurate understanding of interactive intentions and unnatural responses.
By disassembling the multimodal data set from the original video containing subtitles and sounds, and training the large language model in two stages, first trained with single-modal data, and then trained with multimodal data, the cross-modal feature fusion is performed using a model composed of Transformer encoder, CNN encoder, ViT encoder, Transformer decoder and WaveNet decoder.
The large language model has improved its understanding and generation ability of multiple modal information for text, speech and images, and is suitable for multimodal interaction scenarios to achieve more natural human-computer interaction.
Smart Images

Figure CN120509476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model training technology, and in particular to a large language model joint training method and system for multimodal interaction. Background Art
[0002] As large language models become more commonplace, people are demanding higher levels of intelligence from them. For example, when interacting with virtual humans, users expect the machine to be as intelligent as a real person, correctly understanding their intent and providing appropriate feedback.
[0003] However, current virtual human technology can only understand user intentions based on text or text generated by voice recognition, and cannot provide actions, expressions, etc. that are highly matched with the response content in actual interaction scenarios. For example, when a virtual human speaks, it always uses fixed gestures and expressions, which leads to inaccurate understanding of the interaction intention and inappropriate and unnatural responses.
[0004] Therefore, how to enable large language models to understand users' multi-modal information such as voice, text, and human images, and to output multi-modal information such as voice, text, and human images, so as to achieve true multimodal interaction, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] To solve the above-mentioned technical problems, the present invention provides a large language model joint training method and system for multimodal interaction, which can effectively improve the large language model's ability to understand and generate multimodal information such as text, voice and images. It is suitable for scenarios facing multimodal interaction and is closer to natural and actual human-to-human interaction.
[0006] The technical solution adopted in the present invention is as follows: A method for jointly training a large language model for multimodal interaction includes the following steps: obtaining an original video containing subtitles and sound, and decomposing the original video into multiple data groups, wherein each data group includes mutually corresponding text data, voice data, and image data; extracting data of each modality from the multiple data groups separately to obtain first training data consisting of multiple text data, second training data consisting of multiple voice data, and third training data consisting of multiple image data; in a first training stage, training the large language model using the first training data, the second training data, and the third training data, respectively; and in a second training stage, using the multiple data groups as fourth training data, and training the large language model using the fourth training data.
[0007] The large language model includes a first encoder, a second encoder, a third encoder, a cross-attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder. The first encoder, the second encoder and the third encoder are respectively connected to the cross-attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross-attention layer.
[0008] The first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
[0009] The first encoder, the second encoder, and the third encoder are respectively a Transformer encoder, a CNN, and a ViT (Vision Transformer) encoder, and the first decoder, the second decoder, and the third decoder are respectively a Transformer decoder, a WaveNet, and a diffusion model.
[0010] Each text data in the first training data is used as a training sample data, and at least one token in the text data is masked; each voice data in the second training data is used as a training sample data, and at least one phone in the voice data is masked; each image data in the third training data is used as a training sample data, and at least one patch in the image data is masked.
[0011] Each data group in the fourth training data serves as a training sample data, and one or two random modes in the data group are masked.
[0012] A large language model joint training system for multimodal interaction includes: an acquisition module, the acquisition module is used to acquire an original video containing subtitles and sound, and disassemble the original video into multiple data groups, wherein each data group includes mutually corresponding text data, voice data and image data; an extraction module, the extraction module is used to extract the data of each modality from the multiple data groups separately, and obtain first training data consisting of multiple text data, second training data consisting of multiple voice data and third training data consisting of multiple image data; the training module is used to train the large language model with the first training data, the second training data and the third training data in a first training phase, and train the large language model with the fourth training data in a second training phase using the multiple data groups as fourth training data.
[0013] The large language model includes a first encoder, a second encoder, a third encoder, a cross-attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder. The first encoder, the second encoder and the third encoder are respectively connected to the cross-attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross-attention layer.
[0014] The first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
[0015] The first encoder, the second encoder, and the third encoder are respectively a Transformer encoder, a CNN, and a ViT encoder, and the first decoder, the second decoder, and the third decoder are respectively a Transformer decoder, a WaveNet, and a diffusion model.
[0016] Each text data in the first training data is used as a training sample data, and at least one token in the text data is masked; each voice data in the second training data is used as a training sample data, and at least one phone in the voice data is masked; each image data in the third training data is used as a training sample data, and at least one patch in the image data is masked.
[0017] Each data group in the fourth training data serves as a training sample data, and one or two random modes in the data group are masked.
[0018] Beneficial effects of the present invention: The present invention extracts a data group containing multimodal data from the original video containing subtitles and sound, and uses this data group to train a large language model in two stages. In the first stage, each single modal data is trained separately, and in the second stage, multimodal data is trained. This can effectively improve the large language model's ability to understand and generate multiple modal information such as text, voice, and images. It is suitable for scenarios facing multimodal interaction and is closer to natural and actual human-to-human interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flowchart of a method for joint training of a large language model for multimodal interaction according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a large language model according to an embodiment of the present invention; Figure 3 Schematic diagram of a large language model joint training system for multimodal interaction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] like Figure 1 As shown, the large language model joint training method for multimodal interaction according to an embodiment of the present invention includes the following steps: S1, obtaining an original video containing subtitles and sound, and decomposing the original video into multiple data groups, wherein each data group includes corresponding text data, voice data and image data.
[0022] In one embodiment of the present invention, the original video may be an interview program video, a news interview video, etc. containing subtitles and dialogue audio, wherein the subtitles are text corresponding to the dialogue content.
[0023] In one embodiment of the present invention, the sound and video can be segmented based on the switching time points of each subtitle, resulting in a subtitle, a voice segment, and a video segment within each time period. The text content of a subtitle within a time period is treated as text data, the voice segment within the time period is treated as the corresponding voice data, and an image frame within the video segment within the time period is treated as the corresponding image data. The image frame can be an image frame at the middle time point of the video segment, or an image frame randomly selected from the video segment. Since the original video is an interview program video, a news interview video, etc., the image data in each data group are all character images, which contain information such as character movements and / or expressions.
[0024] The number of data sets obtained in this step is used to train a large language model. The larger the number, the better the subsequent training effect. However, it should also be considered in combination with the computing power of the processor running the model.
[0025] S2, extracting data of each modality from the multiple data groups separately to obtain first training data consisting of multiple text data, second training data consisting of multiple voice data, and third training data consisting of multiple image data.
[0026] S3, in a first training stage, respectively training the large language model with the first training data, training the large language model with the second training data, and training the large language model with the third training data.
[0027] S4, in the second training stage, using the plurality of data groups as fourth training data, and training the large language model with the fourth training data.
[0028] The large language model of the embodiment of the present invention can be applied to the virtual human displayed on the display screen, and is suitable for virtual human interaction scenarios such as games, intelligent customer service, and questionnaires. Figure 2 As shown, the large language model includes a first encoder, a second encoder, a third encoder, a cross-attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder. The first encoder, the second encoder and the third encoder are respectively connected to the cross-attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross-attention layer.
[0029] Among them, the first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
[0030] In the second training stage, that is, when the large language model inputs data from multiple modalities at the same time, after the first to third encoders encode the data of the corresponding modalities respectively, the cross-attention layer can calculate the attention weights between the features of different modalities to achieve cross-modal feature fusion; the feature decomposition layer can extract relevant features from the fused features according to the query vectors of each modality through attention decomposition, and output them to the decoders corresponding to each modality for decoding.
[0031] In the first training stage, that is, when the large language model only inputs data of a single modality at the same time, the queries, keys, and values of the cross-attention are unified into features of the same modality, that is, they are degenerated into self-attention layers, thereby realizing the processing of single modality features; correspondingly, the feature decomposition layer can directly transmit the features to the corresponding decoder for decoding.
[0032] In a specific embodiment of the present invention, the first encoder is a Transformer encoder, the second encoder is a CNN, and the third encoder is a ViT encoder; the first decoder is a Transformer decoder, the second decoder is a WaveNet, and the third decoder is a diffusion model.
[0033] In one embodiment of the present invention, each character data in the first training data is used as a training sample data, and at least one token in the character data is masked. A token is a basic unit of character data processed by the model, and this unit can be a word, a subword, a character, etc.
[0034] Each speech data in the second training data is used as a training sample data, and at least one phone in the speech data is masked, where phone refers to a phoneme.
[0035] Each image data in the third training data is used as a training sample data, and at least one patch in the image data is masked, where a patch refers to an image block.
[0036] Through the training in the first training phase of step S3, the ability of the large language model to understand and generate each modal data can be improved.
[0037] In one embodiment of the present invention, each data set in the fourth training data serves as training sample data, and one or two modalities in the data set are randomly masked. Masking of text data, speech data, and image data may respectively mean masking all tokens of text, setting the amplitude or intensity of the speech waveform data or spectrogram to zero, and setting all pixel values of the image to zero.
[0038] The second training phase is performed after the first training phase. Through the training in the second training phase of step S4, the large language model's ability to understand the correspondence between various modalities and its cross-modal generation capability can be improved. Moreover, based on the training foundation of the first training phase, the large language model already has the ability to understand and generate single modalities, and the training effect of cross-modal capabilities is better.
[0039] According to an embodiment of the present invention, a large language model joint training method for multimodal interaction is implemented by extracting a data group containing multimodal data from an original video containing subtitles and sound, and using this data group to train the large language model in two stages. In the first stage, each unimodal data is trained separately, and in the second stage, multimodal data is trained. This method can effectively improve the large language model's ability to understand and generate multiple modal information such as text, voice, and images, and is suitable for scenarios oriented towards multimodal interaction, closer to natural and actual human-to-human interaction.
[0040] Corresponding to the large language model joint training method for multimodal interaction in the above embodiment, the present invention also proposes a large language model joint training system for multimodal interaction.
[0041] like Figure 3 As shown, the large language model joint training system for multimodal interaction of an embodiment of the present invention includes: an acquisition module 10, an extraction module 20 and a training module 30. The acquisition module 10 is used to acquire the original video containing subtitles and sound, and disassemble the original video into multiple data groups, wherein each data group includes mutually corresponding text data, voice data and image data; the extraction module 20 is used to extract the data of each modality from the multiple data groups separately, and obtain first training data composed of multiple text data, second training data composed of multiple voice data and third training data composed of multiple image data; the training module 30 is used to train the large language model with the first training data, the second training data and the third training data in the first training stage, and train the large language model with the fourth training data using the multiple data groups as the fourth training data in the second training stage.
[0042] In one embodiment of the present invention, the original video may be an interview program video, a news interview video, etc. containing subtitles and dialogue audio, wherein the subtitles are text corresponding to the dialogue content.
[0043] In one embodiment of the present invention, the acquisition module 10 can segment the sound and video based on the switching time points of each subtitle, and obtain a subtitle, a voice segment, and a video segment within each time period. The text content of a subtitle within a time period is used as text data, the voice segment within the time period is used as the corresponding voice data, and an image frame in the video segment within the time period is used as the corresponding image data. The image frame can be an image frame at the middle time point of the video segment, or an image frame randomly selected from the video segment. Since the original video is an interview program video, a news interview video, etc., the image data in each data group are character images, which contain information such as character movements and / or expressions.
[0044] The number of data groups acquired by the acquisition module 10 is used for training a large language model. The larger the number, the better the subsequent training effect. However, it should also be considered in combination with the computing power of the processor running the model.
[0045] The large language model of the embodiment of the present invention can be applied to the virtual human displayed on the display screen, and is suitable for virtual human interaction scenarios such as games, intelligent customer service, and questionnaires. Figure 2 As shown, the large language model includes a first encoder, a second encoder, a third encoder, a cross-attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder. The first encoder, the second encoder and the third encoder are respectively connected to the cross-attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross-attention layer.
[0046] Among them, the first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
[0047] In the second training stage, that is, when the large language model inputs data from multiple modalities at the same time, after the first to third encoders encode the data of the corresponding modalities respectively, the cross-attention layer can calculate the attention weights between the features of different modalities to achieve cross-modal feature fusion; the feature decomposition layer can extract relevant features from the fused features according to the query vectors of each modality through attention decomposition, and output them to the decoders corresponding to each modality for decoding.
[0048] In the first training stage, that is, when the large language model only inputs data of a single modality at the same time, the queries, keys, and values of the cross-attention are unified into features of the same modality, that is, they are degenerated into self-attention layers, thereby realizing the processing of single modality features; correspondingly, the feature decomposition layer can directly transmit the features to the corresponding decoder for decoding.
[0049] In a specific embodiment of the present invention, the first encoder is a Transformer encoder, the second encoder is a CNN, and the third encoder is a ViT encoder; the first decoder is a Transformer decoder, the second decoder is a WaveNet, and the third decoder is a diffusion model.
[0050] In one embodiment of the present invention, each character data in the first training data is used as a training sample data, and at least one token in the character data is masked. A token is a basic unit of character data processed by the model, and this unit can be a word, a subword, a character, etc.
[0051] Each speech data in the second training data is used as a training sample data, and at least one phone in the speech data is masked, where phone refers to a phoneme.
[0052] Each image data in the third training data is used as a training sample data, and at least one patch in the image data is masked, where a patch refers to an image block.
[0053] Through the training in the first training stage, the large language model's ability to understand and generate each modal data can be improved.
[0054] In one embodiment of the present invention, each data set in the fourth training data serves as training sample data, and one or two modalities in the data set are randomly masked. Masking of text data, speech data, and image data may respectively mean masking all tokens of text, setting the amplitude or intensity of the speech waveform data or spectrogram to zero, and setting all pixel values of the image to zero.
[0055] The second training phase is conducted after the first training phase. Through the training in the second training phase, the large language model's ability to understand the correspondence between various modalities can be improved, and its cross-modal generation capability can be improved. Moreover, with the training foundation of the first training phase, the large language model already has the ability to understand and generate single modalities, and the training effect of cross-modal capabilities is better.
[0056] According to an embodiment of the present invention, a large language model joint training system for multimodal interaction is developed by extracting a data group containing multimodal data from an original video containing subtitles and sound, and using this data group to train the large language model in two stages. In the first stage, each single modal data is trained separately, and in the second stage, multimodal data is trained. This can effectively improve the large language model's ability to understand and generate multiple modal information such as text, voice, and images, and is suitable for scenarios oriented towards multimodal interaction, closer to natural and actual human-to-human interaction.
[0057] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. "Multiple" means two or more, unless otherwise specifically defined.
[0058] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0059] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.
[0060] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0061] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0062] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0063] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0064] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0065] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0066] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A large language model joint training method for multimodal interaction, characterized by: The following steps are involved: Acquire an original video containing subtitles and sound, and decompose the original video into multiple data groups, wherein each data group includes corresponding text data, voice data, and image data; Extracting data of each modality from the plurality of data groups separately to obtain first training data consisting of a plurality of text data, second training data consisting of a plurality of voice data, and third training data consisting of a plurality of image data; In a first training phase, the large language model is trained using the first training data, the large language model is trained using the second training data, and the large language model is trained using the third training data, respectively; In the second training phase, the plurality of data groups are used as fourth training data, and the large language model is trained using the fourth training data.
2. The large language model joint training method for multimodal interaction according to claim 1 is characterized in that: The large language model includes a first encoder, a second encoder, a third encoder, a cross attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder, wherein the first encoder, the second encoder and the third encoder are respectively connected to the cross attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross attention layer. The first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
3. The large language model joint training method for multimodal interaction according to claim 2 is characterized in that: The first encoder, the second encoder, and the third encoder are respectively a Transformer encoder, a CNN, and a ViT encoder, and the first decoder, the second decoder, and the third decoder are respectively a Transformer decoder, a WaveNet, and a diffusion model.
4. The method for joint training of a large language model for multimodal interaction according to any one of claims 1 to 3, characterized in that: Each character data in the first training data is used as a training sample data, and at least one token in the character data is masked; Each voice data in the second training data is used as a training sample data, and at least one phone in the voice data is masked; Each image data in the third training data serves as a training sample data, and at least one patch in the image data is masked.
5. The method for joint training of a large language model for multimodal interaction according to any one of claims 1 to 3, characterized in that: Each data group in the fourth training data serves as a training sample data, and one or two random modes in the data group are masked.
6. A large language model joint training system for multimodal interaction, characterized by: include: An acquisition module, configured to acquire an original video containing subtitles and audio, and decompose the original video into a plurality of data groups, wherein each data group includes corresponding text data, voice data, and image data; an extraction module, configured to extract data of each modality from the plurality of data groups separately to obtain first training data consisting of a plurality of text data, second training data consisting of a plurality of voice data, and third training data consisting of a plurality of image data; A training module is used to train the large language model using the first training data, the second training data, and the third training data in a first training phase, and to train the large language model using the fourth training data using the plurality of data groups as fourth training data in a second training phase.
7. The large language model joint training system for multimodal interaction according to claim 6, characterized in that: The large language model includes a first encoder, a second encoder, a third encoder, a cross attention layer, a feature decomposition layer, a first decoder, a second decoder and a third decoder, wherein the first encoder, the second encoder and the third encoder are respectively connected to the cross attention layer, the first decoder, the second decoder and the third decoder are respectively connected to the feature decomposition layer, and the feature decomposition layer is connected to the cross attention layer. The first training data, the second training data and the third training data are encoded by the first encoder, the second encoder and the third encoder respectively, the text data, voice data and image data in the fourth training data are encoded by the first encoder, the second encoder and the third encoder respectively, and the first decoder, the second decoder and the third decoder decode the text features, voice features and image features respectively.
8. The large language model joint training system for multimodal interaction according to claim 7, characterized in that: The first encoder, the second encoder, and the third encoder are respectively a Transformer encoder, a CNN, and a ViT encoder, and the first decoder, the second decoder, and the third decoder are respectively a Transformer decoder, a WaveNet, and a diffusion model.
9. The large language model joint training system for multimodal interaction according to any one of claims 6 to 8, characterized in that: Each character data in the first training data is used as a training sample data, and at least one token in the character data is masked; Each voice data in the second training data is used as a training sample data, and at least one phone in the voice data is masked; Each image data in the third training data serves as a training sample data, and at least one patch in the image data is masked.
10. The large language model joint training system for multimodal interaction according to any one of claims 6 to 8, characterized in that: Each data group in the fourth training data serves as a training sample data, and one or two random modes in the data group are masked.
Citation Information
Patent Citations
Multi-modal Transform semantic segmentation algorithm for coping with RGB-D modal deficiency
CN117671265A
Large and small model cooperative training method and device for multi-modal large language model
CN119514645A
Cross-modal retrieval method and device, electronic equipment and storage medium
CN119597939A