An Interaction Method, System, Device and Medium Based on Multimodal Large Model

By performing classification processing and modal correlation calculation on the real-time input information set, combined with splicing processing and decoding output, the problem of multimodal large models being difficult to achieve cross-modal correlation and complex combinations is solved, and natural multimodal interaction and long context processing is realized.

CN119884691BActive Publication Date: 2025-06-20CHENGDU KOALA URAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510352943.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-20
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing multimodal large-modal model technology is difficult to achieve cross-modal correlation and complex multimodal combinations, and cannot naturally handle complex multimodal interaction scenarios.

Method used

By obtaining the real-time input information set, including video, audio and text information, classifying and performing correlation calculations through preset modal correlation models, obtaining correlation loss information, and then splicing and decoding output of the processed data, realizing correlation and complex multimodal combinations between modes.

Benefits of technology

It realizes complete information interaction based on cross-modality, can naturally handle complex multimodal interaction scenarios, breaks the dependence of traditional multimodal systems on strict modal classification, and enhances the model's processing ability of long contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884691B_ABST
    Figure CN119884691B_ABST
Patent Text Reader

Abstract

The present invention provides an interaction method, system, device and medium based on a multi-modal large model, relating to the technical field of multi-modal large models. The method includes: obtaining a real-time input information set; respectively processing the real-time input information set to obtain processed data, where the processed data includes first processed information, second processed information and third processed information, and the first processed information is obtained by processing real-time video information, the second processed information is obtained by processing real-time audio information, and the third processed information is obtained by processing real-time text information; performing correlation calculation on the processed data through a preset modal correlation model; performing splicing processing on the processed data according to the correlation loss information to obtain a spliced data set; decoding and outputting the spliced data set to obtain interaction response data, and the interaction response data is used to feedback interaction information. This method solves the problem of realizing cross-modal correlation of real-time input data and is convenient to be extended to more complex multi-modal combinations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal large models, and more particularly, to an interaction method, system, device, and medium based on a multimodal large model. Background Art

[0002] In recent years, the field of artificial intelligence has been undergoing a revolutionary transformation led by large language models. From the initial text generation to the current integration of multimodal information such as speech, audio, images, and videos, large models have become a research hotspot in the field of artificial intelligence. The initial large models mainly focused on text interaction, aiming to achieve natural language processing tasks such as machine translation, text summarization, and question-answering systems. With technological progress, researchers have attempted to incorporate multimodal information into large models to enhance their generalization ability in real-world scenarios. Multimodal large models combine language processing capabilities with the ability to generate multiple information modalities, enabling them to not only process text information but also various modal data such as audio and video, thereby achieving richer and more natural interaction methods. However, in existing methods, multimodality often overly relies on strict modal classification, fails to establish cross-modal associations, and does not scale to more complex multimodal combinations, making it impossible to naturally handle complex multimodal interaction scenarios. Summary of the Invention

[0003] The purpose of the present invention is to provide an interaction method, system, device, and medium based on a multimodal large model to solve the problem of establishing cross-modal associations for real-time input data and facilitating expansion to more complex multimodal combinations.

[0004] To achieve the above objective, the technical solutions adopted by the present invention are as follows:

[0005] In a first aspect, the present application provides an interaction method based on a multimodal large model, the method comprising:

[0006] Obtaining a real-time input information set, the real-time input information set including real-time video information, real-time audio information, and real-time text information;

[0007] Processing the real-time input information set respectively to obtain processed data, the processed data including first processed information, second processed information, and third processed information, wherein processing the real-time video information obtains the first processed information, processing the real-time audio information obtains the second processed information, and processing the real-time text information obtains the third processed information;

[0008] Performing association calculation on the processed data through a preset modal association model to obtain association loss information of the first processed information, the second processed information, and the third processed information;

[0009] Perform splicing processing on the processed data according to the associated loss information to obtain a spliced dataset;

[0010] Perform decoding output on the spliced dataset to obtain interactive response data, where the interactive response data is used to feedback interactive information.

[0011] In a second aspect, the present application also provides an interaction system based on a multi-modal large model, and the system includes:

[0012] An acquisition module, configured to acquire a real-time input information set, where the real-time input information set includes real-time video information, real-time audio information, and real-time text information;

[0013] A processing module, configured to respectively process the real-time input information set to obtain processed data, where the processed data includes first processed information, second processed information, and third processed information, and where the real-time video information is processed to obtain the first processed information, the real-time audio information is processed to obtain the second processed information, and the real-time text information is processed to obtain the third processed information;

[0014] An association module, configured to perform association calculation on the processed data through a preset modality association model to obtain the associated loss information of the first processed information, the second processed information, and the third processed information;

[0015] A splicing module, configured to perform splicing processing on the processed data according to the associated loss information to obtain a spliced dataset;

[0016] A decoding output module, configured to perform decoding output on the spliced dataset to obtain interactive response data, where the interactive response data is used to feedback interactive information.

[0017] In a third aspect, the present application also provides an interaction device based on a multi-modal large model, including:

[0018] A memory, configured to store a computer program;

[0019] A processor, configured to implement the steps of the interaction method based on the multi-modal large model when executing the computer program.

[0020] In a fourth aspect, the present application also provides a readable storage medium, where a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned interaction method based on the multi-modal large model are implemented.

[0021] The beneficial effects of the present invention are:

[0022] The present invention realizes complete information interaction on the basis of cross-modal by classifying, associating and calculating, splicing and expanding, and decoding and outputting the real-time input information set in sequence, and can naturally process complex multi-modal interaction scenarios. Among them: in the processing of data, the feature processing of different modalities is considered; in the association calculation, a preset modality association model is introduced to realize the association calculation of the processed information of different modalities, breaking the dependence of traditional multi-modal systems on strict modality classification, solving the association between cross-modalities for real-time input data, and facilitating the extension to more complex multi-modal combinations; the introduction of splicing and expansion realizes the processing ability for long contexts.

[0023] In addition, the association calculation in this method reflects progressive association. It starts from the association of bimodality, such as vision-text, audio-text, and then gradually expands to more complex multi-modal combinations. This association algorithm not only reduces the difficulty of model association calculation, but also helps the model establish a more stable cross-modal understanding ability.

[0024] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification, or be understood by implementing the embodiments of the present invention. Brief Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic flowchart of the interaction method based on a multi-modal large model described in the embodiments of the present invention;

[0027] Figure 2 It is a schematic structural diagram of the interaction system based on a multi-modal large model described in the embodiments of the present invention;

[0028] Figure 3 It is a schematic structural diagram of the interaction device based on a multi-modal large model described in the embodiments of the present invention;

[0029] Markings in the figure:

[0030] 1. Acquisition module; 2. Processing module; 3. Association module; 4. Splicing module; 5. Decoding and output module; 31. First association unit; 32. Second association unit; 33. Third association unit; 800. Interaction device based on a multi-modal large model; 801. Processor; 802. Memory; 803. Multimedia component; 804. I / O interface; 805. Communication component. Detailed implementation mode

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein generally can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0032] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first", "second", etc. are only used for differential description and cannot be construed as indicating or implying relative importance.

[0033] Embodiment 1:

[0034] This embodiment provides an interaction method based on a multimodal large model.

[0035] See Figure 1 , which shows that this method includes steps S1 to S5, specifically:

[0036] S1: Obtain a real-time input information set, which includes real-time video information, real-time audio information, and real-time text information;

[0037] Among them, in the real-time video information, it not only includes video information such as real-time video, but also includes picture information.

[0038] S2: Process the real-time input information set respectively to obtain processed data, which includes first processed information, second processed information, and third processed information. Among them, the real-time video information is processed to obtain the first processed information, the real-time audio information is processed to obtain the second processed information, and the real-time text information is processed to obtain the third processed information;

[0039] In step S2, to clarify the processing process of the processed data, step S2 includes steps S21 to S24, specifically:

[0040] S21: Solve the real-time video information through a video sampling model to obtain a key frame data set, where the video sampling model is:

[0041] ; (1)

[0042] In the above formula (1), represents the key frame data set, (.) represents the video sampling function, represents the real-time video information, which is the original image or video sequence, represents the preset number of sampling frames;

[0043] In step S21, the video sampling function (.) can sample the key frame data and unify the sampling length.

[0044] S22: Process each frame of data in the key frame data set through the video modality processing model to obtain the first processing information, where the video modality processing model is:

[0045] ; (2)

[0046] In the above formula (2), represents the key frame data set, respectively represent different key frame data, represents the model for dynamically partitioning the key frame data for processing, represents the key frame data of the feature vector, represents the key frame data of the first processing information, {.} represents the bridging model;

[0047] S23: Process the real-time audio information through the audio modality processing model to obtain the second processing information, where the audio modality processing model is:

[0048] ; (3)

[0049] In the above formula (3), represents the real-time audio information, respectively represent the audio sequence data of different segments after segmenting the real-time audio information, represents the feature extraction of the current segment of audio sequence data through the Mel filter model, represents the data obtained by convolutional downsampling after feature extraction through the Mel filter model by , {.} represents the encoding model, represents the second processing information of the segment of audio sequence;

[0050] In this method, algorithms such as fixed-length segmentation, sliding window, and feature analysis can be used to segment real-time audio information. Among them, fixed-length segmentation divides real-time audio information by a fixed time length and is suitable for situations where the fixed frequency of audio content is relatively high; sliding window segmentation is suitable for processing audio data containing multiple speakers; feature analysis is applicable to scenarios where audio content needs to be analyzed in detail before segmentation.

[0051] S24: Process the real-time text information through a preset text modality processing model to obtain third processed information.

[0052] In this step, the preset text modality processing model can adopt tokenization to process the text information to obtain third processed information, and the third processed information is text sequence data of different segments. During the tokenization process, an algorithm for byte pair encoding can also be adopted.

[0053] In existing methods, multimodality often overly relies on strict modality classification, fails to achieve cross-modal association, and will not be extended to more complex multimodal combinations. Therefore, this method performs association calculation on the processed data to break the dependence of traditional multimodal systems on strict modality classification.

[0054] S3: Perform association calculation on the processed data through a preset modality association model to obtain association loss information of the first processed information, the second processed information, and the third processed information;

[0055] In step S3, to clarify the specific content of the association calculation, step S3 includes S31 to S33:

[0056] S31: Obtain loss information related to video text association;

[0057] In this step, the specific calculation process of the loss information related to video text association includes steps S311 to S313, specifically:

[0058] S311: Perform prediction probability calculation on the first processed information and the processed information of the preset text to obtain a first prediction information;

[0059] Among them, the calculation formula for the first prediction information is:

[0060] ; (4)

[0061] In the above formula (4), represents the first prediction information, represents the probability of the first processed information when it is the processed information of the preset text, represents the first processed information, represents the processed information of the preset text.

[0062] S312: Calculate the prediction probability of the third processing information and the processing information of the preset video to obtain the second prediction information;

[0063] Among them, the calculation formula of the second prediction information is:

[0064] ; (5)

[0065] In the above formula (5), represents the second prediction information, represents the probability of the third processing information when it is the processing information of the preset video, represents the third processing information, represents the processing information of the preset video.

[0066] S313: Calculate the loss of the first prediction information and the second prediction information to obtain the loss information related to the video text.

[0067] In step S313, the loss calculation is:

[0068] ; (6)

[0069] In the above formula (6), represents the first prediction information, represents the probability of the first processing information when it is the processing information of the preset text, represents the first processing information, represents the processing information of the preset text, represents the second prediction information, represents the probability of the third processing information when it is the processing information of the preset video, represents the third processing information, represents the processing information of the preset video, represents the loss information related to the video text.

[0070] In this method, the loss information related to the video text is used to characterize the relevance between the video content (including images) and the corresponding description text.

[0071] S32: Obtain the loss information related to the audio text;

[0072] In step S32, to clarify the specific calculation process of the loss information related to the audio text, step S32 includes S321 to S323, specifically:

[0073] S321: Calculate the prediction probability of the second processing information and the processing information of the preset text to obtain the third prediction information;

[0074] Among them, the calculation formula of the third prediction information is:

[0075] ; (7)

[0076] In the above formula (7), represents the third prediction information, represents the probability of the second processing information when representing the processing information of the preset text, represents the second processing information, represents the processing information of the preset text.

[0077] S322: Calculate the prediction probability by comparing the third processing information with the processing information of the preset audio to obtain the fourth prediction information;

[0078] Among them, the calculation formula of the fourth prediction information is:

[0079] ; (8)

[0080] In the above formula (8), represents the fourth prediction information, represents the probability of the third processing information when representing the processing information of the preset audio, represents the third processing information, represents the processing information of the preset audio.

[0081] S323: Calculate the loss by comparing the third prediction information and the fourth prediction information to obtain the loss information related to the audio text.

[0082] In step S323, the loss calculation is:

[0083] ; (9)

[0084] In the above formula (9), represents the third prediction information, represents the probability of the second processing information when representing the processing information of the preset text, represents the second processing information, represents the processing information of the preset text, represents the fourth prediction information, represents the probability of the third processing information when representing the processing information of the preset audio, represents the third processing information, represents the processing information of the preset audio, represents the loss information related to the audio text.

[0085] The loss information related to the audio text is used to associate the audio content with the text transcription.

[0086] S33: Obtain the overall association loss information based on the loss information associated with the video text and the loss information associated with the audio text. In this step, the calculation of the overall association loss information is as follows:

[0087] ; (10)

[0088] In the above formula (10), represents the overall association loss information, represents the loss information associated with the video text, represents the loss information associated with the audio text, represents the regularization term.

[0089] The overall association loss information reflects the association during multi-modal combination. Through the feedback of the overall association loss information, a unified framework is established to handle the conversion and association between modalities, enabling the multi-modal model to handle complex multi-modal interaction scenarios more naturally.

[0090] In this method, by comparing the loss information associated with the video text, the loss information associated with the audio text, and the overall association loss information with the preset target loss information one by one, it is judged whether the association degree of the current different modalities meets the standard, so as to realize the natural learning and processing of the complex association between modalities.

[0091] In addition, the association calculation in this method reflects progressive association. It starts from the association of bimodalities, such as vision-text, audio-text, and then gradually expands to more complex multi-modal combinations. This kind of association algorithm not only reduces the difficulty of model association calculation, but also helps the model establish a more stable cross-modal understanding ability.

[0092] In this method, to achieve the processing ability of the multi-modal large model for long context, step S4 splicing processing is introduced to realize the joint optimization of multi-modal context.

[0093] S4: Perform splicing processing on the processed data according to the association loss information to obtain a spliced dataset;

[0094] After comparing the loss information associated with the video text, the loss information associated with the audio text, and the overall association loss information with the preset target loss information one by one, if all meet the requirements of the preset target loss information, splicing processing can be performed at this time. The requirements of the preset target loss information are set according to the association requirements of the large model for different modalities. For example, if the large model needs to strongly associate videos, audios, and texts, it is required that the loss information associated with the video text, the loss information associated with the audio text, and the overall association loss information are all less than the preset target loss information. If there is one item greater than the preset target loss information, parameter adjustment needs to be performed through an optimization algorithm.

[0095] In this method, to clarify the specific process of splicing processing, step S4 includes S41 to S44, specifically:

[0096] S41: Perform splicing processing on the first processing information to obtain first splicing information, where the splicing calculation is:

[0097] ; (11)

[0098] In the above formula (11), represents the first splicing information; represents the first processing information of the th key frame data, where the value range of in

[0099] is from 1 to n;

[0100] ; (12)

[0101] In the above formula (12), represents the second splicing information, represents the second processing information of the th audio sequence segment, where the value range of in

[0102] is from 1 to m;

[0103] ; (13)

[0104] In the above formula (13), represents the third splicing information, represents the content corresponding to the th text sequence data segment, where the value range of in

[0105] S44: Perform overall splicing on the first splicing information, the second splicing information, and the third splicing information to obtain a splicing data set, where the overall splicing calculation is:

[0106] ; (14)

[0107] In the above formula (14), Represents the concatenated dataset, Represents the first concatenation information, Represents the second concatenation information, Represents the third concatenation information, Represents the maximum sequence length allowed in the preset concatenation model.

[0108] The introduction of the concatenated dataset realizes the formation of a unified sequence representation of the features of text, images, videos, and audio through concatenation processing on the basis of multimodal input. In addition, during the concatenation process, the concatenation strategy can be dynamically adjusted according to real-time input information, and the concatenation method and length can be automatically adjusted according to different scenarios to achieve the optimal feature combination, significantly improving the performance and efficiency of the multimodal large model in practical applications.

[0109] S5: Decode and output the concatenated dataset to obtain interactive response data, which is used to feedback interactive information.

[0110] In the decoding output, to achieve dynamic allocation of the importance of different modalities, in the feature fusion model included in the decoding output, the weight of each modality is solved through the attention mechanism. At this time, the following can be achieved: ① Dynamically allocate the proportion in the total sequence length according to the importance of different modalities; ② Prioritize ensuring the integrity of key interactive response data.

[0111] In step S5, the calculation of the decoding output is:

[0112] ; (15)

[0113] In the above formula (15), Represents the interactive response data, Represents the language decoding model, Represents the feature fusion model, Represents the feature projection model, Represents the concatenated dataset.

[0114] In formula (15), the concatenated dataset outputs text in an autoregressive manner with continuous tokens in each model. This method makes full use of the powerful sequence modeling ability of the large model and can naturally understand and generate interactive information for feedback.

[0115] During the decoding output, to achieve a unified representation of the embedding spaces of videos, images, audio, and text, the feature vectors in the embedding space can be aligned. The means of alignment can be based on the gradient descent method, and the aligned feature vectors are used to characterize the feature consistency of different modalities.

[0116] In addition, the text of the decoding output can be converted through a text-to-speech model, and the conversion formula is:

[0117] ; (16)

[0118] In the above formula (16), represents the final interactive response audio information, represents an external conversion model, represents the interactive response data.

[0119] The introduction of the external conversion model can integrate a part of the decoding process into the inference process of the large model and achieve text-to-speech. During this process, the speech generation delay is reduced through joint optimization.

[0120] In this method, to improve the deep association of different modalities and enhance the generalization ability of the model for data processing, the processed data is marked and the state label prediction is maximized, that is: between step S3 and step S4, there are also S34 and S35, specifically:

[0121] S34: Perform state marking on the processed data to obtain multimodal marking information, where the multimodal marking information includes video marking information, audio marking information, and text marking information;

[0122] Among them, the state marking is used to distinguish different information, specifically:

[0123] The video marking information corresponds to , and video marking is performed through a visual encoder. The video marking information includes images and videos;

[0124] The audio marking information corresponds to , and audio marking is performed through an audio encoder;

[0125] The text marking information corresponds to , and text marking is performed through a text encoder.

[0126] S35: Predict the multimodal marking information through a maximum state label prediction model, and the maximum state label prediction model is calculated as:[[]]

[0127] ; (17)

[0128] In the above formula (17), Maximum state label prediction loss information; represents the number of samples of the processed data; represents the true label of the state label, and its value is 1 or 0; represents the probability of the predicted state label, where represents the multimodal marking information, represents the current preset data.

[0129] In the above process, the maximized state marker prediction model realizes the marker feedback of the processed data, maximizing the accuracy of the state marker prediction loss information representation model for different information markers. When the maximized state marker prediction loss information meets the preset prediction loss information, the association of different modalities at a deeper level is further improved, enhancing the generalization ability of the model for data processing.

[0130] Embodiment 2:

[0131] As Figure 2 shown, this embodiment provides an interaction system based on a multimodal large model, and the system includes:

[0132] An acquisition module 1, configured to acquire a real-time input information set, where the real-time input information set includes real-time video information, real-time audio information, and real-time text information;

[0133] A processing module 2, configured to process the real-time input information set respectively to obtain processed data, where the processed data includes first processed information, second processed information, and third processed information. Among them, the real-time video information is processed to obtain the first processed information, the real-time audio information is processed to obtain the second processed information, and the real-time text information is processed to obtain the third processed information;

[0134] An association module 3, configured to perform association calculation on the processed data through a preset modality association model to obtain association loss information of the first processed information, the second processed information, and the third processed information;

[0135] A splicing module 4, configured to perform splicing processing on the processed data according to the association loss information to obtain a spliced data set;

[0136] A decoding and output module 5, configured to perform decoding and output on the spliced data set to obtain interaction response data, where the interaction response data is used to feedback interaction information.

[0137] In an implementation method disclosed in the present invention, the association module 3 includes:

[0138] A first association unit 31, configured to obtain loss information of video-text association;

[0139] A second association unit 32, configured to obtain loss information of audio-text association;

[0140] A third association unit 33, configured to obtain overall association loss information according to the loss information of video-text association and the loss information of audio-text association.

[0141] It should be noted that regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0142] Embodiment 3:

[0143] Corresponding to the above method embodiments, in this embodiment, an interaction device based on a multimodal large model is also provided. An interaction device based on a multimodal large model described below can be correspondingly referred to the interaction method based on a multimodal large model described above.

[0144] Figure 3 is a block diagram of an interaction device 800 based on a multimodal large model shown according to an exemplary embodiment. As Figure 3 shown, the interaction device 800 based on a multimodal large model may include: a processor 801, a memory 802. The interaction device 800 based on a multimodal large model may further include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0145] Among them, the processor 801 is used to control the overall operation of the multimodal large model-based interaction device 800 to complete all or part of the steps in the above-mentioned multimodal large model-based interaction method. The memory 802 is used to store various types of data to support the operation of the multimodal large model-based interaction device 800. These data may include, for example, instructions for any application or method operating on the multimodal large model-based interaction device 800, as well as application-related data, such as messages, pictures, audio, video, etc. received and sent. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 803 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or sent through the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the multimodal large model-based interaction device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G or 4G, or a combination of one or more of them. Therefore, the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0146] In an exemplary embodiment, the multimodal large model-based interaction device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above multimodal large model-based interaction method.

[0147] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When the program instructions are executed by a processor, the steps of the above multimodal large model-based interaction method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 802 including program instructions, and the above program instructions can be executed by the processor 801 of the multimodal large model-based interaction device 800 to complete the above multimodal large model-based interaction method.

[0148] Embodiment 4:

[0149] Corresponding to the above method embodiment, a readable storage medium is also provided in this embodiment. A readable storage medium described below can be correspondingly referred to with a multimodal large model-based interaction method described above.

[0150] A readable storage medium has a computer program stored thereon. When the computer program is executed by a processor, the steps of the multimodal large model-based interaction method of the above method embodiment are implemented.

[0151] The readable storage medium can specifically be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0152] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0153] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. An interactive method based on a multimodal large model, characterized in that: include: Acquire a real-time input information set, wherein the real-time input information set includes real-time video information, real-time audio information, and real-time text information; Processing the real-time input information sets respectively to obtain processed data, the processed data including first processed information, second processed information and third processed information, wherein the real-time video information is processed to obtain the first processed information, the real-time audio information is processed to obtain the second processed information, and the real-time text information is processed to obtain the third processed information; The processing data is correlated and calculated by a preset modal correlation model to obtain correlation loss information of the first processing information, the second processing information and the third processing information; wherein, the correlation loss information includes: Obtain loss information associated with video text; Get the loss information associated with the audio text; Obtaining overall associated loss information according to the loss information associated with the video text and the loss information associated with the audio text; The loss of the loss information associated with the video text is calculated as: ; In the above formula, represents the first prediction information, represents the probability of the first processing information when the processing information of the preset text is the first processing information, represents the first processing information, Indicates the processing information of the preset text. represents the second prediction information, represents the probability that the processing information of the preset video is the third processing information, represents the third processing information, Indicates the processing information of the preset video. Indicates the loss information of the video text association; The loss of the loss information associated with the audio text is calculated as: ; In the above formula, represents the third prediction information, represents the probability of the second processing information when the processing information of the preset text is the second processing information, represents the second processing information, Indicates the processing information of the preset text. represents the fourth prediction information, represents the probability that the processing information of the preset audio is the third processing information, represents the third processing information, Indicates the processing information of preset audio. Indicates the loss information of audio text association; The loss of the overall associated loss information is calculated as: ; In the above formula, Represents the overall associated loss information, Indicates the loss information of video text association, Indicates the loss information of audio text association, represents the regularization term; The processed data are spliced ​​according to the associated loss information to obtain a spliced ​​data set; when the loss information associated with the video text, the loss information associated with the audio text, and the overall associated loss information are compared with the preset target loss information one by one, they all meet the preset target loss information requirements and are spliced; the preset target loss information requirements are set according to the association requirements of different modalities of the large model; The spliced ​​data set is decoded and output to obtain interactive response data, and the interactive response data is used to feed back interactive information.

2. The interactive method based on a multimodal large model according to claim 1, characterized in that: The loss information calculation of the video text association includes: Performing prediction probability calculation on the first processed information and the processed information of the preset text to obtain first prediction information; Performing prediction probability calculation on the third processing information and the processing information of the preset video to obtain second prediction information; Loss calculation is performed on the first prediction information and the second prediction information to obtain loss information associated with the video text.

3. The interactive method based on a multimodal large model according to claim 1, characterized in that: The loss information calculation of the audio text association includes: Performing prediction probability calculation on the second processed information and the processed information of the preset text to obtain third prediction information; Performing prediction probability calculation on the third processing information and the processing information of the preset audio to obtain fourth prediction information; Loss calculation is performed on the third prediction information and the fourth prediction information to obtain loss information associated with the audio text.

4. The interactive method based on a multimodal large model according to claim 1, characterized in that: After performing correlation calculation on the processed data through a preset modal correlation model to obtain correlation loss information of the first processed information, the second processed information and the third processed information, the method further includes: Performing status marking on the processed data to obtain multimodal marking information, wherein the multimodal marking information includes video marking information, audio marking information, and text marking information; The multimodal label information is predicted by a maximizing state label prediction model, and the calculation of the maximizing state label prediction loss information in the maximizing state label prediction model is: ; In the above formula, Maximize the state label prediction loss information; Indicates the number of samples of processed data; The true label of the state mark, whose value is 1 or 0; represents the predicted state label probability, where Represents multimodal labeling information, Indicates the current preset data.

5. The interactive method based on a multimodal large model according to claim 1, characterized in that: In the feature fusion model included in the decoded output, the weight of each modality is solved through the attention mechanism.

6. The interactive method based on a multimodal large model according to claim 1, characterized in that: The decoded output is calculated as: ; In the above formula, Indicates interactive response data. represents the language decoding model, represents the feature fusion model, represents the feature projection model, Represents a concatenated dataset.

7. The interactive method based on a multimodal large model according to claim 1, characterized in that: The real-time input information sets are processed respectively to obtain processed data, wherein the processed data includes first processed information, second processed information and third processed information, including: The real-time video information is solved by a video sampling model to obtain a key frame data set, wherein the video sampling model is: ; In the above formula, represents a key frame dataset, (.) represents the video sampling function, Represents real-time video information, which is the original image or video sequence, Indicates the preset number of sampling frames; Each frame of data in the key frame data set is processed by a video modality processing model to obtain first processing information, wherein the video modality processing model is: ; In the above formula, represents a key frame dataset, Respectively represent different key frame data, Represents key frame data Model for dynamic block processing, Represents key frame data The characteristic vector of Represents key frame data First process information, {.} represents a bridge model; The real-time audio information is processed by an audio modal processing model to obtain second processing information, wherein the audio modal processing model is: ; In the above formula, Represents real-time audio information, They respectively represent the audio sequence data of different segments after segmenting the real-time audio information. Indicates that the feature extraction of the audio sequence data of the current segment is performed through the Mel filter model. Indicates that after feature extraction through the Mel filter model The data to be convolved and downsampled, {.} indicates the encoding model, express second processing information of the audio sequence; The real-time text information is processed by a preset text modality processing model to obtain third processing information.

8. The interactive method based on a multimodal large model according to claim 1, characterized in that: The processed data are spliced ​​according to the associated loss information to obtain a spliced ​​data set, including: The first processed information is spliced ​​to obtain first spliced ​​information, wherein the splicing calculation is: ; In the above formula, Indicates the first splicing information; Indicates The first processing information of key frame data, wherein middle The value range of is 1 to n; Indicates splicing processing; The second processed information is spliced ​​to obtain second spliced ​​information, wherein the splicing calculation is: ; In the above formula, Indicates the second splicing information, express The second processing information of the audio sequence is middle The value range of is 1 to m; Indicates splicing processing; The third processed information is spliced ​​to obtain third spliced ​​information, wherein the splicing calculation is: ; In the above formula, Indicates the third splicing information, Indicates The content corresponding to the text sequence data, where middle The value range of is from 1 to w; Indicates splicing processing; The first splicing information, the second splicing information and the third splicing information are spliced ​​as a whole to obtain a spliced ​​data set, wherein the overall splicing calculation is: ; In the above formula, represents the concatenated dataset, Indicates the first splicing information, Indicates the second splicing information, Indicates the third splicing information, Indicates the maximum sequence length allowed in the preset splicing model.

9. An interactive system based on a multimodal large model, characterized in that: include: An acquisition module, used to acquire a real-time input information set, wherein the real-time input information set includes real-time video information, real-time audio information and real-time text information; A processing module, used to process the real-time input information set respectively to obtain processed data, wherein the processed data includes first processed information, second processed information and third processed information, wherein the real-time video information is processed to obtain the first processed information, the real-time audio information is processed to obtain the second processed information, and the real-time text information is processed to obtain the third processed information; The association module is used to perform association calculation on the processed data through a preset modal association model to obtain the association loss information of the first processing information, the second processing information and the third processing information; wherein, it includes: Obtain loss information associated with video text; Get the loss information associated with the audio text; Obtaining overall associated loss information according to the loss information associated with the video text and the loss information associated with the audio text; The loss of the loss information associated with the video text is calculated as: ; In the above formula, represents the first prediction information, represents the probability of the first processing information when the processing information of the preset text is the first processing information, represents the first processing information, Indicates the processing information of the preset text. represents the second prediction information, represents the probability that the processing information of the preset video is the third processing information, represents the third processing information, Indicates the processing information of the preset video. Indicates the loss information of the video text association; The loss of the loss information associated with the audio text is calculated as: ; In the above formula, represents the third prediction information, represents the probability of the second processing information when the processing information of the preset text is represents the second processing information, Indicates the processing information of the preset text. represents the fourth prediction information, represents the probability that the processing information of the preset audio is the third processing information, represents the third processing information, Indicates the processing information of preset audio. Indicates the loss information of audio text association; The loss of the overall associated loss information is calculated as: ; In the above formula, Represents the overall associated loss information, Indicates the loss information of video text association, Indicates the loss information of audio text association, represents the regularization term; A splicing module is used to splice the processed data according to the associated loss information to obtain a spliced ​​data set; when the loss information associated with the video text, the loss information associated with the audio text, and the overall associated loss information are compared with the preset target loss information one by one, they all meet the preset target loss information requirements and are spliced; the preset target loss information requirements are set according to the association requirements of different modalities of the large model; The decoding output module is used to decode and output the spliced ​​data set to obtain interactive response data, and the interactive response data is used to feed back interactive information.

10. An interactive device based on a multimodal large model, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the interaction method based on a multimodal large model as claimed in any one of claims 1 to 8 when executing the computer program.

11. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the interaction method based on a multimodal large model as claimed in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Video and text cross-modal retrieval method based on relational reasoning network

    CN113239159A

  • Multi-modal entity recognition and relation extraction method based on learnable prompt

    CN117151223A