Methods, apparatuses, devices, and media for processing multi-modal data

CN122497983APending Publication Date: 2026-07-31BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-01-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing machine learning models fail to effectively consider the temporal correspondence between multimodal data when processing multimodal data, resulting in poor performance, especially in terms of the synchronization and integrity of audio and video information.

Method used

By receiving multimodal data and processing requests, the system extracts video and audio tags at predetermined time intervals, interweaves them to generate multimodal features, processes the multimodal data based on text features, and trains and adjusts the encoder, aligner, and language model of the machine learning model to optimize the synchronization of audio and video information.

Benefits of technology

It improves the performance of machine learning models in processing multimodal data, enhances the synchronous understanding and integrity of audio and video information, and improves the accuracy and efficiency of processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497983A_ABST
    Figure CN122497983A_ABST
Patent Text Reader

Abstract

Methods, apparatus, devices, and media for processing multimodal data are provided. In one method, multimodal data and a processing request for the multimodal data are received. The multimodal data includes at least video data and audio data, and the processing request is represented in natural language. Multiple video tags and multiple audio tags are extracted from the video data and the audio data respectively, according to predetermined time intervals. The multiple video tags and multiple audio tags are interleaved according to predetermined time intervals to generate multimodal features of the multimodal data. The multimodal data is processed based on the multimodal features and the text features of the processing request. Using exemplary implementations of this disclosure, audio and video data in multimodal data can be processed simultaneously, thereby enhancing the model's understanding of the synchronization of audio and video information and improving the processing performance of multimodal data.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatuses, devices, and media for processing multi-modal data TECHNICAL FIELD

[0001] Exemplary implementations of the present disclosure generally relate to machine learning models, and in particular, to methods, apparatuses, devices, and computer-readable storage media for processing multi-modal data using machine learning models. BACKGROUND

[0002] Machine learning techniques have been widely used to perform a variety of tasks. Currently, a variety of machine learning models have been proposed for processing multi-modal data. Multi-modal data generally includes video data, audio data, text data, and the like. However, existing solutions do not consider the temporal correspondence between multi-modal data as a whole, and thus the performance of machine learning models is not satisfactory. At this time, it is desirable to utilize machine learning models to process multi-modal data in a more effective manner. SUMMARY

[0003] In a first aspect of the present disclosure, a method for processing multi-modal data is provided. In the method, multi-modal data and a processing request for the multi-modal data are received, the multi-modal data including at least video data and audio data, and the processing request being expressed in natural language. According to a predetermined time interval, a plurality of video tokens of the video data and a plurality of audio tokens of the audio data are extracted. According to the predetermined time interval, the plurality of video tokens and the plurality of audio tokens are interleaved to generate multi-modal features of the multi-modal data. The multi-modal data is processed based on the multi-modal features and text features of the processing request.

[0004] In a second aspect of the present disclosure, an apparatus for processing multi-modal data is provided. The apparatus includes a receiving module configured to receive multi-modal data and a processing request for the multi-modal data, the multi-modal data including at least video data and audio data, and the processing request being expressed in natural language; an extracting module configured to extract, according to a predetermined time interval, a plurality of video tokens of the video data and a plurality of audio tokens of the audio data; an interleaving module configured to interleave, according to the predetermined time interval, the plurality of video tokens and the plurality of audio tokens to generate multi-modal features of the multi-modal data; and a processing module configured to process the multi-modal data based on the multi-modal features and text features of the processing request.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0008] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the present disclosure or to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description, which, taken in conjunction with the drawings, disclose various implementations. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, advantages and aspects of the implementations of the present disclosure will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings. In the drawings like reference numerals designate like elements, wherein:

[0010] FIG. 1 shows a block diagram of an application environment according to one example implementation of the present disclosure;

[0011] FIG. 2 shows a block diagram of a method for processing multi-modal data according to some implementations of the present disclosure;

[0012] FIG. 3 shows a block diagram of a structure of a machine learning model according to some implementations of the present disclosure;

[0013] FIG. 4 shows a block diagram of a training process of a machine learning model according to some implementations of the present disclosure;

[0014] FIG. 5 shows a block diagram of a reinforcement learning process according to some implementations of the present disclosure;

[0015] FIG. 6 shows a flowchart of a method for processing multi-modal data according to some implementations of the present disclosure;

[0016] FIG. 7 shows a block diagram of an apparatus for processing multi-modal data according to some implementations of the present disclosure; and

[0017] FIG. 8 shows a block diagram of a device capable of implementing the implementations of the present disclosure. DETAILED DESCRIPTION

[0018] Implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several implementations of the present disclosure are described, it should be understood that the present disclosure can be embodied in many other forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and implementations described are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0019] In the description of implementations of the present disclosure, the term "includes" and its derivatives mean "including but not limited to". The term "based on" means "based at least in part on". The term "one implementation" or "the implementation" means "at least one implementation". The term "some implementations" means "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions known at present and / or to be developed in the future.

[0020] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0022] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware that executes the operation of the technical solutions of the present disclosure according to the prompt information.

[0023] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending prompt information to the user, for example, can be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0024] It can be understood that the above-mentioned notification and acquisition of user authorization process is only illustrative, and does not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.

[0025] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the execution of a subsequent action performed in response to the event or condition is not necessarily strongly correlated with the time at which the event occurs or the condition is established. For example, in some cases, a subsequent action can be performed immediately upon the occurrence of an event or the establishment of a condition; in other cases, a subsequent action can be performed after a period of time has elapsed since the occurrence of an event or the establishment of a condition.

[0026] Example Environment

[0027] A variety of machine learning models have been proposed for processing multi-modal data. Multi-modal data often includes video data, audio data, textual data, and so on. An application environment according to some implementations of the present disclosure is described with reference to FIG. 1, which illustrates a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, multi-modal data 110 can include video data 112, audio data 114, and so on. It will be appreciated that FIG. 1 merely schematically illustrates various modalities in multi-modal data, and alternatively and / or additionally, multi-modal data can further include textual data, and / or data of other modalities.

[0028] A processing request 120 from a user can be received, which can be expressed in natural language and is for processing multi-modal data 110. For example, the processing request 120 can express: please extract a summary from the multi-modal data, please answer the following question based on the multi-modal data, and so on. A machine learning model 130 can process the multi-modal data 110 and provide a processing result that matches the processing request 120.

[0029] For ease of description, in the following, only a multi-modal large language model is taken as an example to describe the specific process of processing multi-modal data. At present, various multi-modal models have been proposed, for example, a multi-modal model capable of understanding picture content and video content, a multi-modal model capable of understanding audio and video in a video. However, the prior art cannot consider the time correspondence between multi-modal data as a whole, and the performance of the machine learning model is not satisfactory. Specifically, the existing model only focuses on video information, although it can well understand the video information of the video, such as detailed description of the video content, answering questions about the video content, etc., but usually ignores the audio information in the video. The model with audio and video understanding capability usually cannot well align the audio and video information, and the attention to the time synchronization relationship of the audio and video information is missing in the training process. Further, the existing model is difficult to completely capture various information in the audio and video, resulting in information omission, and there may be a lot of repeated content in the processing result. At this time, it is expected to utilize the machine learning model to process multi-modal data in a more effective way.

[0030] Summary of processing multi-modal data

[0031] To at least partially solve the deficiencies in the prior art, according to one exemplary implementation of the present disclosure, a method for processing multi-modal data is proposed. Specifically, multi-modal data and a processing request in natural language representation for the multi-modal data can be received, the multi-modal data at least including video data and audio data. With some implementations of the present disclosure, the video data and the audio data can be aligned, thereby improving the performance of processing multi-modal data.

[0032] A summary according to one exemplary implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 for processing multi-modal data according to some implementations of the present disclosure. As shown in FIG. 2, a plurality of video tokens 214, 216, …, 218 of the video data 112 can be extracted according to a predetermined time interval. Similarly, a plurality of audio tokens 224, 226, …, 228 of the audio data 114 can be extracted according to a predetermined time interval. Specifically, video features 210 can be extracted from the video data 112, and the video features 210 are parsed to determine the plurality of video tokens 214, 216, …, 218. Similarly, audio features 220 can be extracted from the audio data 112, and the audio features 220 are parsed to determine the plurality of audio tokens 224, 226, …, 228.

[0033] Further, the plurality of video tokens 214, 216, …, 218 and the plurality of audio tokens 224, 226, …, 228 can be interleaved according to predetermined time intervals to generate a multimodal feature 230 of the multimodal data. The multimodal data can be processed based on the multimodal feature 230 and a text feature of the processing request. With the exemplary implementation of the present disclosure, the audio data and the video data in the multimodal data can be processed simultaneously, which enhances the model’s understanding of the synchronization of audio-video information and thus improves the data processing performance.

[0034] Detailed process of processing multimodal data

[0035] In the context of the present disclosure, for ease of description, the process of processing multimodal data is described only by taking a query request as an example of a processing request. Here, the multimodal data can be input to a machine learning model, and the query request can instruct the machine learning model to output a summary caption of the multimodal data in a text format. Alternatively and / or additionally, the processing request can perform other tasks, such as detailed description of video content, answering questions about video content, generating a short video including summary content, and the like.

[0036] According to some implementations of the present disclosure, the above-mentioned process can be performed by utilizing a machine learning model, which can include a plurality of encoders, an aligner, a language model, and / or other models. The machine learning model can be trained by utilizing a large amount of training data, for example, the training process can include a pre-training phase, a supervised fine-tuning phase, a reinforcement learning phase, and the like. Specifically, a multimodal audio-video large language model can be constructed, and original video information and audio information can be input to corresponding encoders to extract corresponding key information (e.g., video tokens and audio tokens). The video tokens and the audio tokens can be arranged in a time-sequential interleaved manner to obtain audio-video information with time synchronization, which is input to the large language model. The large language model can be utilized to determine a processing result based on the extracted audio-video information and an input natural language prompt, so as to convert the audio-video information into a corresponding processing result.

[0037] The mechanism of the machine learning model is described with reference to FIG. 3, which shows a block diagram 300 of the structure of a machine learning model according to some implementations of the present disclosure. As shown in FIG. 3, the machine learning model can include a video encoder 310, which can process a video portion in the input multimodal data to determine a corresponding video feature. Further, the machine learning model can include a video aligner 312, which can extract a plurality of video tokens 214, 216, …, 218 from the video feature according to predetermined time intervals.

[0038] Similar to the way of processing the video, the machine learning model can include an audio encoder 320 that can process the audio portion in the inputted multi-modal data to determine corresponding audio features. Further, the machine learning model can include an audio aligner 312 that can extract a plurality of audio tokens 224, 226, …, 228 from the audio features at predetermined time intervals.

[0039] According to some implementations of the present disclosure, the machine learning model can be utilized to extract a plurality of video tokens. Specifically, a video encoder in the machine learning model can be utilized to extract video features of the video data, and a video aligner can be utilized to extract a plurality of video tokens from the video features at predetermined time intervals. As shown in FIG. 3, the video encoder 310 can be inputted with the video data in the multi-modal data to obtain the video features. The video aligner 312 can be inputted with the video features to obtain the plurality of video tokens 214, 216, …, 218.

[0040] According to some implementations of the present disclosure, the machine learning model can be utilized to extract a plurality of audio tokens. Specifically, an audio encoder in the machine learning model can be utilized to extract audio features of the audio data, and an audio aligner can be utilized to extract a plurality of audio tokens from the audio features at predetermined time intervals. As shown in FIG. 3, the audio encoder 320 can be inputted with the audio data in the multi-modal data to obtain the audio features. The audio aligner 322 can be inputted with the audio features to obtain the plurality of audio tokens 224, 226, …, 228.

[0041] According to some implementations of the present disclosure, the machine learning model can include an interleaving combiner 330. The interleaving combiner 330 can interleave the plurality of video tokens and the plurality of audio tokens to generate multi-modal features of the multi-modal data at the predetermined time intervals. According to some implementations of the present disclosure, the predetermined time interval can be set to 1 second (or other data), which can be divided into smaller sub-time intervals. For example, the sub-time interval can be set based on the sampling frequency of the video data, assuming that the video data includes 25 data frames per second, the sub-time interval can be 1 / 25 second. The video tokens and the audio tokens can be extracted at the same time intervals. For example, 25 video tokens can be extracted from 1 second of video data, and 25 audio tokens can be extracted from 1 second of audio data.

[0042] The interleaving combiner 330 can be utilized to process the multiple video tokens and the multiple audio tokens. For example, the audio token 224 can be placed after the video token 214 to generate a token 334 corresponding to a first sub-time interval; the audio token 226 can be placed after the video token 216 to generate a token 336 corresponding to a second sub-time interval; …; the audio token 228 can be placed after the video token 218 to generate a token 338 corresponding to an Nth sub-time interval. Further, the tokens 334, 336, …, 338 can be concatenated to generate a multi-modal feature of the multi-modal data. Here, the multi-modal feature can include video features and audio features that are time-aligned, thereby describing multiple contents of the multi-modal data in a more accurate manner.

[0043] As shown in FIG. 3, the machine learning model can include a text encoder 340 that can extract text features 344, 346, …, and 348 from the processing request, and a language model 350 that can receive the multi-modal feature and the text features, and generate a corresponding processing result. The text features and the multi-modal feature can be input to the language model 350 to output a corresponding processing result by the language model 350. Further, the machine learning model can include a low-rank adaptation (LoRA) model 360 for adjusting the language model 350.

[0044] According to some implementations of the present disclosure, the machine learning model is a pre-trained model, and can perform a corresponding processing on multi-modal data according to a processing request. In an inference process, the machine learning model can receive input multi-modal data, and output a processing result corresponding to the processing request. Assuming that the multi-modal data is a video of a conversation between two persons, and the processing request is: please extract the summary content of the video. At this time, the machine learning model can output the text: the video includes a conversation between two persons, and the conversation content is to discuss the weather…. For another example, assuming that the multi-modal data is a video introducing the history of a city, and the previous processing request is: please generate a short video describing the summary content of the video, with a length not exceeding 1 minute. At this time, the machine learning model can output a corresponding short video.

[0045] According to some implementations of the present disclosure, the machine learning model can be trained based on multiple steps, see FIG. 4 for more details, which illustrates a block diagram 400 of a training process of a machine learning model according to some implementations of the present disclosure. As shown in FIG. 4, in visual training 410, visual samples can be received, which include visual data and visual annotations describing the visual data, the visual data including at least any of image data or video data, and the visual annotations including textual descriptions of the visual data; and the visual samples are used to update the video aligner and the language model in the machine learning model.

[0046] Specifically, a large amount of visual data, including images and videos, can be used for training, so as to obtain a pure visual large model with good performance. Here, the visual samples may, for example, include image data and corresponding annotations. Assuming that the image data includes a natural landscape, the annotations may, for example, indicate that the image includes a blue sky with white clouds, a mountain and a lake, etc. The visual samples may, for example, include video data and corresponding annotations. Assuming that the video data includes a scene of a person talking, the annotations may, for example, indicate that an old man and a child are talking, the old man is on the left and the child is on the right, the background of the picture is a park, etc. It should be understood that the annotations here only relate to the visual content in the video and do not relate to the sound. For example, the LLaVA-Video scheme can be used for training. At this time, other parameters in the machine learning model can be fixed, and the video aligner and the language model in the machine learning model are updated. In this way, the machine learning model can be made to fully master the correlation between the visual content and the textual description, so that the machine learning model can process the video part in the multi-modal data in a more accurate manner.

[0047] In audio training 420, audio samples can be received, which can include audio data and audio annotations describing the audio data, the audio annotations including at least any of language recognition annotations or textual description annotations; and the audio samples are used to update the audio aligner. Specifically, a large amount of audio data can be used for training, which can train the machine learning model to specify a speech recognition task and a task of extracting textual descriptions of the audio. Assuming that the audio data includes a conversation, for the speech recognition task, the corresponding annotations may, for example, include the specific text involved in the conversation, such as the specific content of the conversation between two characters. For the task of extracting textual descriptions of the audio, the corresponding annotations may, for example, include a summary of the content of the conversation.

[0048] According to some implementations of the present disclosure, other parameters in the machine learning model can be fixed, and only the audio aligner in the machine learning model is updated. With some implementations of the present disclosure, the machine learning model can be made to fully master the knowledge about processing the audio task, so that the machine learning model processes the audio part in the multi-modal data in a more accurate manner.

[0049] In the audio-video fine-tuning 430, a multi-modal sample can be received, which can include a video sample, an audio sample, and a multi-modal label describing the multi-modal sample, which can include at least one of the following: a description text label or a question and answer label; and the multi-modal sample is used to update at least one of the following: an audio aligner, a language model, or a low-rank adaptive model in the machine learning model. Specifically, based on the supervised fine-tuning technical solution, the video with audio of the fine label can be used for training, so as to train the machine learning model to perform the audio-video extraction description text (audio-visual caption) task and the audio-video question and answer (audio-visual QA) task. Specifically, other parameters in the machine learning model can be fixed, and only the LoRA model of the audio aligner and the language model is updated.

[0050] It should be understood that the multi-modal sample herein can include multi-modal data, and the multi-modal data includes both audio and video parts. In other words, training is performed at this time using video data with sound, and the label can include labels about both the audio part and the video part. Assuming that the video data includes a person conversation scene, the label at this time can represent that an old man and a child are in conversation, the old man is on the left, the child is on the right, the picture background is a park, and the background sound has bird calls and background music. The old man says: XXXX, and the child says: YYY.

[0051] According to some implementations of the present disclosure, a certain amount of multi-modal data with rich label information can be provided to update the machine learning model, so that the machine learning model fully masters the knowledge about processing the task combining video and audio, so that the machine learning model processes the video part and the audio part in the multi-modal data in a more accurate manner.

[0052] In the reinforcement learning 440, a loss function can be determined based on at least one reinforcement learning indicator associated with the machine learning model; and the machine learning model is updated based on the loss function. Specifically, reinforcement learning fine-tuning can be performed on the machine learning model trained in the above manner.

[0053] Currently, various reinforcement learning technical solutions have been proposed. For example, an existing evaluation model (e.g., a large language model, etc.) can be used to determine the evaluation of the processing result output by the machine learning model. However, this technical solution relies too much on the ability of the evaluation model, and the effect usually has strong randomness and poor credibility, and the final feedback effect is not satisfactory. In addition, currently, multiple evaluation models have been proposed to determine the evaluation for multiple indicators respectively, and then select the sample that is good for all indicators. However, this technical solution has poor sample utilization rate, and the final feedback effect is also cross. At this time, it is expected to find a technical solution that can simultaneously have high credibility and high sample utilization rate of feedback.

[0054] According to some implementations of the present disclosure, reinforcement learning can include multiple indicators, thereby further improving the performance of the machine learning model. According to some implementations of the present disclosure, the reinforcement learning indicators can be determined according to the multi-faceted needs for the machine learning model, for example, the machine learning model can be updated based on completeness, repeatability, and traceability, etc. The multi-faceted evaluation can be used as a reward in reinforcement learning, and then the machine learning model is updated. See Figure 5 for more details, which shows a block diagram 500 of a reinforcement learning process according to some implementations of the present disclosure.

[0055] As shown in Figure 5, the reinforcement learning 440 process can be performed based on multiple reinforcement learning indicators 550. Specifically, the reinforcement learning indicators 550 can include completeness 510, repeatability 520, and traceability 530. The completeness can indicate whether the processing result output by the machine learning model involves all events in the multi-modal data. Here, the completeness can include two aspects: missing rate and illusion rate. Assuming that the multi-modal data includes multiple events, the missing rate can represent the proportion between the number of missing events in the output processing result and the number of multiple events. The higher the missing rate, the worse the performance of the machine learning model. It should be understood that the machine learning model can output illusions, i.e., events imagined by the machine learning model, and the illusion rate can represent the proportion between the number of events belonging to illusions in the output result and the number of multiple events. The higher the illusion rate, the worse the performance of the machine learning model. In this way, the reliability of the processing result of the machine learning model can be evaluated from the aspect of completeness.

[0056] According to some implementations of the present disclosure, the at least one reinforcement indicator can include a completeness indicator. A training sample can be obtained, and an evaluation of completeness of the training sample can be determined. Specifically, the training sample can include multi-modal data and a corresponding textual annotation. In determining the evaluation, the textual annotation can be divided into a plurality of events; a prediction of the textual annotation generated by processing the multi-modal data by the machine learning model can be compared with the plurality of events; and the evaluation can be determined based on a result of the comparison. In this way, the evaluation related to completeness can be determined in a simple and effective manner, and the machine learning model can be updated accordingly towards improving completeness.

[0057] According to some implementations of the present disclosure, the training data can include video data and a corresponding textual annotation, and the textual annotation can be split into a plurality of atomic events using a known machine learning model. Assume the textual description is: “a student opens the backpack, takes out the pencil box, takes out the book, starts to write homework…”. At this time, the plurality of atomic events can include, for example, 1) the student opens the backpack, 2) the student takes out the pencil box, 3) the student takes out the book, 4) the student starts to write homework… The video data in the training data can be input into the machine learning model, and a prediction of the processing result can be determined. Assume the prediction is: “a student opens the backpack, takes out the pencil box”, the prediction includes 2 events, while the ground truth includes 4 events, at this time, the completeness can be expressed as 2 / 4 = 50%, for example. According to some implementations of the present disclosure, the list of atomic events and the prediction generated by the model can be input into an evaluation model, and the evaluation model can be used to predict which events are missing, which events are described incorrectly, and which events are hallucinations of the model. According to some implementations of the present disclosure, the higher the completeness, the higher the reward value can be specified. In this way, the completeness of the output result of the machine learning model can be improved.

[0058] According to some implementations of the present disclosure, the at least one reinforcement indicator can include a repetition indicator. A training sample can be obtained, and an evaluation of repetition of the training sample can be determined. Specifically, the training sample can include multi-modal data, and determining the evaluation can include: dividing a prediction of the textual annotation generated by processing the multi-modal data by the machine learning model into a plurality of phrases; determining a number of repeated words in the words in the plurality of phrases; and determining the evaluation based on the number of repeated words and a total number of words in the prediction of the textual annotation. In this way, the evaluation related to repetition can be determined in a simple and effective manner, and the machine learning model can be updated accordingly towards reducing repetition.

[0059] According to some implementations of the present disclosure, the prediction of the processing result of the machine learning model output can be divided into several phrases according to punctuation marks, and the number of occurrences of each phrase can be counted. Further, the number of repeated words in the words in the plurality of phrases can be determined, where the number of repeated words = the number of words contained in each phrase * (the number of occurrences of the phrase - 1). Further, the redundancy can be determined based on the following formula: redundancy = the number of repeated words / the total number of words. Specifically, assuming that the output result of the model is: "My name is XXX, nice to meet you, …", and the phrase "nice to meet you" occurs 10 times in the entire output result. At this time, the phrase "nice to meet you" includes 4 words, and the phrase occurs 10 times. At this time, the number of repeated words = 4 * (10 - 1) = 36. Assuming that the total number of words in the output result is 100, the redundancy is 36 / 100 = 36%. With some implementations of the present disclosure, it can be specified that the higher the redundancy, the lower the reward value. In this way, the redundancy of the output result of the machine learning model can be reduced.

[0060] According to some implementations of the present disclosure, the at least one reinforcement indicator includes a backtracking indicator. A training sample can be obtained, and an evaluation of the backtracking degree of the training sample can be determined. Specifically, the training sample can include multi-modal data and a script annotation corresponding to a data segment (i.e., a video within a certain time window) in the multi-modal data. Determining the evaluation can include: dividing the script annotation into a plurality of events; comparing the prediction of the script generated by processing the data segment by the machine learning model with the plurality of events; and determining the evaluation based on the result of the comparison. In this way, the evaluation related to the backtracking degree can be determined in a simple and effective manner, and the machine learning model can be updated accordingly in the direction of improving the backtracking degree.

[0061] It should be understood that time backtracking refers to: a specified time window, the model needs to describe all the audio and video information within the time window. The evaluation method is similar to the event completeness indicator, the difference is that the list of atomic events provided at this time is the plurality of events within the time window. Assuming that the training data includes 2 minutes, the machine learning model can be specified to provide the script within the first minute. At this time, each event in the generated script can be compared with the plurality of true value events within the first minute in the training data in order to determine the corresponding backtracking degree. With some implementations of the present disclosure, it can be specified that the higher the backtracking degree, the higher the reward value. In this way, the backtracking degree of the output result of the machine learning model can be improved.

[0062] According to some implementations of the present disclosure, the overall loss function can be determined based on the above aspects. To improve the training efficiency, the training data can include first samples and second samples. Specifically, in determining the at least one loss function corresponding to the at least one reinforcement learning indicator, for a reinforcement indicator in the at least one reinforcement indicator, a first evaluation corresponding to the reinforcement learning indicator can be determined based on the first samples, and a second evaluation corresponding to the reinforcement learning indicator can be determined based on the second samples; a difference between the first evaluation and the second evaluation can be determined; and a loss function corresponding to the reinforcement learning indicator can be determined based on the difference. In this way, the influence of different samples on different reinforcement learning indicators of the machine learning model can be measured in a simple and effective manner, thereby improving the training efficiency of the machine learning model.

[0063] According to some implementations of the present disclosure, for each reinforcement learning indicator, the first samples and the second samples can be processed separately in the manner described above to determine the respective evaluations. For example, for the completeness, the first samples include the first multi-modal data and the first textual annotation, and determining the first evaluation includes: dividing the first textual annotation into a plurality of events; comparing the prediction of the textual generated by processing the first multi-modal data by the machine learning model with the plurality of events; and determining the first evaluation based on the result of the comparison. The second evaluation of the second samples can be determined in a similar manner.

[0064] For example, for the repetitiveness, the first samples include the first multi-modal data, and determining the first evaluation includes: dividing the prediction of the textual generated by processing the first multi-modal data by the machine learning model into a plurality of phrases; determining the number of repeated words in the words in the plurality of phrases; and determining the first evaluation based on the number of repeated words and the total number of words in the prediction of the textual. The second evaluation of the second samples can be determined in a similar manner. For example, for the backtracking, the first samples include the first multi-modal data and the first textual annotation corresponding to a data segment in the first multi-modal data, and determining the first evaluation includes: dividing the first textual annotation into a plurality of events; comparing the prediction of the textual generated by processing the data segment by the machine learning model with the plurality of events; and determining the first evaluation based on the result of the comparison. The second evaluation of the second samples can be determined in a similar manner.

[0065] According to some implementations of the present disclosure, a reinforcement learning sample can be received, the reinforcement learning sample including a first sample and a second sample; based on the first sample and the second sample, at least one loss function corresponding to at least one reinforcement learning indicator can be determined; and based on the at least one loss function, a loss function can be determined. Specifically, the multiple reinforcement learning indicators described above can be introduced into a direct preference optimization (DPO) technical solution. Specifically, in order to comprehensively consider various indicators, various indicators can be normalized, and the original DPO loss function can be rewritten as follows:

[0066] In the above formula, Δ i (x) represents the normalized difference between sample 1 and sample 2 on the i-th (i is less than the number n of indicators, n = 3 in this paper) indicator, k, m are hyperparameters for adjusting the loss function, and σ represents an activation function. In this way, a smooth decision coefficient can be obtained on an n-dimensional hyperplane, so that the weight of each sample in the DPO optimization process can vary smoothly. With some implementations of the present disclosure, two samples with a large gap have a greater impact on the optimization process, while two samples with a small gap have a smaller impact on the optimization process.

[0067] Generally speaking, when evaluating the pros and cons of the output results of the machine learning model, all the content is often first converted into text, and a pure text large language model is used to give feedback, which leads to that the feedback for audio and video is not direct enough and the feedback can be inaccurate. It should be understood that the machine learning model of the present disclosure has been trained using rich audio data and video data and has strong audio and video understanding capability before reinforcement learning. At this time, the model itself can be considered for self-feedback, and the self-feedback model is made online, that is, real-time self-feedback in the model updating process. Alternatively and / or additionally, the self-feedback training can be used to replace the reinforcement learning 440 described above.

[0068] According to some implementations of the present disclosure, in the process of determining the loss function, the machine learning model can be used to determine the loss function. Specifically, an online self-feedback training method for audio and video is proposed. Specifically, the flow of online self-feedback is as follows. First, the machine learning model can be input with multi-modal data, and the machine learning model can be used to generate two pieces of explanatory text online: caption1 and caption2. More audio and video information (e.g., background information, etc.) can be provided to the machine learning model, and the probability values of caption1 and caption2 can be calculated by the machine learning model, respectively, with the positive sample being the one with a high probability value and vice versa. Based on the probability value, the importance weight of the sample pair can also be calculated, that is, Δ i(x) Further, the DPO loss can be calculated and the model can be updated accordingly. According to some implementations of the present disclosure, the above process can be performed iteratively, for example, another multi-modal data can be input to the machine learning model, and the above online self-feedback training process can be repeated.

[0069] It should be appreciated that although the above describes the process of processing multi-modal data with query request as an example of processing request. Here, the multi-modal data can be input to the machine learning model, and the query request can instruct the machine learning model to output summary text of the multi-modal data in text format. Alternatively and / or additionally, the processing request can perform other tasks, for example, detailed description of video content, answering questions about video content, generating short videos including summary content, etc.

[0070] According to some implementations of the present disclosure, the structure and output of the network layers downstream of the machine learning model can be adjusted according to the task performed by the processing request. With some implementations of the present disclosure, the multi-modal data processing capability in the backbone of the machine learning model can be adjusted to adapt to perform different downstream tasks. In this way, the backbone structure of the machine learning model can be reused as much as possible to perform different tasks.

[0071] According to some implementations of the present disclosure, an automated index for evaluating audio-video description is proposed, making it possible to automatically evaluate audio-video description. Reinforcement learning fine-tuning can be performed on an audio-video multi-modal large model, balancing the understanding of audio and video modalities, maintaining the visual understanding ability of the model, and preventing the performance of video modality understanding from declining due to the addition of audio modality. Further, automatic balancing can be performed between multiple reinforcement learning indexes, and the multi-dimensional performance of the model can be considered comprehensively. A self-feedback online reinforcement learning is proposed, which supports the model to check the text generated by itself, thereby realizing self-feedback learning. With exemplary implementations of the present disclosure, audio data and video data in multi-modal data can be processed simultaneously, thereby enhancing the synchronization of the model in understanding audio-video information, and improving the data processing performance.

[0072] Example process

[0073] FIG. 6 illustrates a flowchart of a method 600 for processing multi-modal data according to some implementations of the present disclosure. At block 610, multi-modal data and a processing request for the multi-modal data are received, the multi-modal data including at least video data and audio data, and the processing request being expressed in natural language. At block 620, a plurality of video tokens of the video data and a plurality of audio tokens of the audio data are extracted according to a predetermined time interval. At block 630, the plurality of video tokens and the plurality of audio tokens are interleaved to generate multi-modal features of the multi-modal data according to the predetermined time interval. At block 640, the multi-modal data is processed based on the multi-modal features and a text feature of the processing request.

[0074] According to some implementations of the present disclosure, extracting the plurality of video tokens comprises: extracting, with a video encoder in the machine learning model, video features of the video data; and extracting, with the video aligner, the plurality of video tokens from the video features according to the predetermined time interval.

[0075] According to some implementations of the present disclosure, extracting the plurality of audio tokens comprises: extracting, with an audio encoder in the machine learning model, audio features of the video data; and extracting, with the audio aligner, the plurality of audio tokens from the audio features according to the predetermined time interval.

[0076] According to some implementations of the present disclosure, the machine learning model is obtained by: receiving a vision sample, the vision sample comprising vision data and a vision annotation describing the vision data, the vision data comprising at least one of: image data or video data, and the vision annotation comprising a textual description of the vision data; and updating the video aligner and a language model in the machine learning model with the vision sample.

[0077] According to some implementations of the present disclosure, the machine learning model is obtained by: receiving an audio sample, the audio sample comprising audio data and an audio annotation describing the audio data, the audio annotation comprising at least one of: a language recognition annotation or a textual description annotation; and updating the audio aligner with the audio sample.

[0078] According to some implementations of the present disclosure, the machine learning model is obtained by: receiving a multi-modal sample, the multi-modal sample comprising a video sample, an audio sample, and a multi-modal annotation describing the multi-modal sample, the multi-modal annotation comprising at least one of: a textual description annotation or a question-answer annotation; and updating at least one of: the audio aligner, the language model, or a low-rank adaptive model in the machine learning model with the multi-modal sample.

[0079] According to some implementations of the present disclosure, the machine learning model is obtained by: determining a loss function based on at least one reinforcement learning indicator associated with the machine learning model; and updating the machine learning model based on the loss function.

[0080] According to some implementations of the present disclosure, determining the loss function comprises: receiving a reinforcement learning sample, the reinforcement learning sample comprising a first sample and a second sample; determining at least one loss function corresponding to the at least one reinforcement learning indicator based on the first sample and the second sample; and determining the loss function based on the at least one loss function.

[0081] According to some implementations of the present disclosure, determining the at least one loss function corresponding to the at least one reinforcement learning indicator comprises, for a reinforcement indicator in the at least one reinforcement indicator, determining a first evaluation corresponding to the reinforcement learning indicator based on the first sample, and determining a second evaluation corresponding to the reinforcement learning indicator based on the second sample; determining a difference between the first evaluation and the second evaluation; and determining the loss function corresponding to the reinforcement learning indicator based on the difference.

[0082] According to some implementations of the present disclosure, the at least one reinforcement indicator comprises a completeness indicator, the first sample comprises first multi-modal data and first narrative annotations, and determining the first evaluation comprises: dividing the first narrative annotations into a plurality of events; comparing a prediction of a narrative generated by processing the first multi-modal data by the machine learning model with the plurality of events; and determining the first evaluation based on a result of the comparison.

[0083] According to some implementations of the present disclosure, the at least one reinforcement indicator comprises a repetition indicator, the first sample comprises first multi-modal data, and determining the first evaluation comprises: dividing a prediction of a narrative generated by processing the first multi-modal data by the machine learning model into a plurality of phrases; determining a number of repeated words in the plurality of phrases; and determining the first evaluation based on the number of repeated words and a total number of words in the prediction of the narrative.

[0084] According to some implementations of the present disclosure, the at least one reinforcement indicator comprises a backtracking indicator, the first sample comprises first multi-modal data and first narrative annotations corresponding to a data segment in the first multi-modal data, and determining the first evaluation comprises: dividing the first narrative annotations into a plurality of events; comparing a prediction of a narrative generated by processing the data segment by the machine learning model with the plurality of events; and determining the first evaluation based on a result of the comparison.

[0085] According to some implementations of the present disclosure, determining the loss function comprises: determining the loss function with the machine learning model.

[0086] According to some implementations of the present disclosure, the request is a query request for content of the multi-modal data.

[0087] Example apparatus and devices

[0088] FIG. 7 illustrates a block diagram of an apparatus 700 for processing multi-modal data, according to some implementations of the present disclosure. The apparatus includes a receiving module 710 configured to receive multi-modal data and a processing request for the multi-modal data, the multi-modal data including at least video data and audio data, the processing request being in a natural language; an extracting module 720 configured to extract a plurality of video tokens of the video data and a plurality of audio tokens of the audio data according to a predetermined time interval; an interleaving module 730 configured to interleave the plurality of video tokens and the plurality of audio tokens according to the predetermined time interval to generate a multi-modal feature of the multi-modal data; and a processing module 740 configured to process the multi-modal data based on the multi-modal feature and a text feature of the processing request.

[0089] According to some implementations of the present disclosure, the extracting module is further configured to extract video features of the video data using a video encoder in the machine learning model, and extract the plurality of video tokens from the video features according to the predetermined time interval using the video aligner.

[0090] According to some implementations of the present disclosure, the extracting module is further configured to extract audio features of the video data using an audio encoder in the machine learning model, and extract the plurality of audio tokens from the audio features according to the predetermined time interval using the audio aligner.

[0091] According to some implementations of the present disclosure, the apparatus further includes a training module configured to receive a visual sample, the visual sample including visual data and a visual annotation describing the visual data, the visual data including at least either of image data or video data, and the visual annotation including a textual description of the visual data, and update the video aligner and a language model in the machine learning model using the visual sample.

[0092] According to some implementations of the present disclosure, the training module is further configured to receive an audio sample, the audio sample including audio data and an audio annotation describing the audio data, the audio annotation including at least either of a language recognition annotation or a textual description annotation, and update the audio aligner using the audio sample.

[0093] According to some implementations of the present disclosure, the training module is further configured to receive a multi-modal sample, the multi-modal sample including the visual sample, the audio sample, and a multi-modal annotation describing the multi-modal sample, the multi-modal annotation including at least either of a textual description annotation or a question-answer annotation, and update at least either of the audio aligner, the language model, or a low-rank adaptive model in the machine learning model using the multi-modal sample.

[0094] According to some implementations of the present disclosure, the training module is further configured to: determine the loss function based on at least one reinforcement learning metric associated with the machine learning model; and update the machine learning model based on the loss function.

[0095] According to some implementations of the present disclosure, the training module is further configured to: receive a reinforcement learning sample, the reinforcement learning sample comprising a first sample and a second sample; determine at least one loss function corresponding to at least one reinforcement learning metric based on the first sample and the second sample; and determine the loss function based on the at least one loss function.

[0096] According to some implementations of the present disclosure, the training module is further configured to: determine, for a reinforcement metric in the at least one reinforcement metric, a first evaluation corresponding to the reinforcement learning metric based on the first sample and a second evaluation corresponding to the reinforcement learning metric based on the second sample; determine a difference between the first evaluation and the second evaluation; and determine the loss function corresponding to the reinforcement learning metric based on the difference.

[0097] According to some implementations of the present disclosure, the at least one reinforcement metric comprises a completeness metric, the first sample comprises first multi-modal data and first narrative annotation, and the training module is further configured to: divide the first narrative annotation into a plurality of events; compare a prediction of a narrative generated by processing the first multi-modal data by the machine learning model with the plurality of events; and determine the first evaluation based on a result of the comparison.

[0098] According to some implementations of the present disclosure, the at least one reinforcement metric comprises a repetition metric, the first sample comprises first multi-modal data, and the training module is further configured to: divide a prediction of a narrative generated by processing the first multi-modal data by the machine learning model into a plurality of phrases; determine a number of repeated words of words in the plurality of phrases; and determine the first evaluation based on the number of repeated words and a total number of words of the prediction of the narrative.

[0099] According to some implementations of the present disclosure, the at least one reinforcement metric comprises a backtracking metric, the first sample comprises first multi-modal data and first narrative annotation corresponding to a data segment in the first multi-modal data, and the training module is further configured to: divide the first narrative annotation into a plurality of events; compare a prediction of a narrative generated by processing the data segment by the machine learning model with the plurality of events; and determine the first evaluation based on a result of the comparison.

[0100] According to some implementations of the present disclosure, the training module is further configured to: determine the loss function using the machine learning model.

[0101] According to some implementations of the present disclosure, the processing request is a query request for content of the multi-modal data.

[0102] FIG. 8 illustrates a block diagram of a device 800 that can implement a number of implementations of the present disclosure. It should be understood that the computing device 800 illustrated in FIG. 8 is merely an example and should not be construed as any limitation of the functionality and scope of the implementations described herein. The computing device 800 illustrated in FIG. 8 can be used to implement the methods described above.

[0103] As illustrated in FIG. 8, the computing device 800 is in the form of a general- purpose computing device. Components of the computing device 800 can include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 820. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of the computing device 800.

[0104] The computing device 800 typically includes a plurality of computer storage media. Such media can be volatile and nonvolatile media and removable and non-removable media implemented in any method or technology for storage of information such as program modules, data, and / or data structures. The memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable media and can include machine readable media such as a flash drive, a magnetic disk drive, or any other media that can be used to store information and / or data (e.g., training data for training) and that can be accessed by the computing device 800.

[0105] The computing device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 8, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media such as a floppy disk, a ZIP® disk, a magnetic tape, or an optical disk, for example. In these instances, each drive can be connected to the bus by one or more data media interfaces. The memory 820 can include a computer program product 825 having one or more program modules configured to carry out the various methods or actions of the implementations of the present disclosure.

[0106] The communication units 840 enable communications with other computing devices over a communication medium. Additionally, the functionality of the components of the computing device 800 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with each other through a communication connection. Thus, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0107] The input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 800 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc. through the communication unit 840, as needed, communicate with one or more devices that enable a user to interact with the computing device 800, or any devices (e.g., a network card, a modem, etc.) that enable the computing device 800 to communicate with one or more other computing devices. Such communication can be carried out through an input / output (I / O) interface (not shown).

[0108] According to an example implementation of the present disclosure, a computer readable storage medium is provided, having stored thereon computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided, having stored thereon a computer program, which when executed by a processor implements the method described above.

[0109] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0110] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0111] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0113] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for processing multi-modal data, comprising: receiving multi-modal data and a processing request for the multi-modal data, the multi-modal data comprising at least video data and audio data, the processing request being expressed in natural language; extracting a plurality of video tokens of the video data and a plurality of audio tokens of the audio data according to a predetermined time interval; interleaving the plurality of video tokens and the plurality of audio tokens according to the predetermined time interval to generate multi-modal features of the multi-modal data; and processing the multi-modal data based on the multi-modal features and text features of the processing request. 2.The method of claim 1, wherein the extracting the plurality of video tokens comprises: extracting video features of the video data with a video encoder in a machine learning model; and extracting the plurality of video tokens from the video features according to the predetermined time interval with a video aligner. 3.The method of claim 2, wherein the extracting the plurality of audio tokens comprises: extracting audio features of the video data with an audio encoder in the machine learning model; and extracting the plurality of audio tokens from the audio features according to the predetermined time interval with an audio aligner. 4.The method of claim 3, wherein the machine learning model is obtained by: receiving visual samples comprising visual data and visual annotations describing the visual data, the visual data comprising at least either of image data or video data, and the visual annotations comprising textual descriptions of the visual data; and updating the video aligner and a language model in the machine learning model with the visual samples. 5.The method of claim 4, wherein the machine learning model is obtained by: receiving audio samples comprising audio data and audio annotations describing the audio data, the audio annotations comprising at least either of language recognition annotations or textual description annotations; and updating the audio aligner with the audio samples. 6.The method of claim 5, wherein the machine learning model is obtained by: receiving multi-modal samples comprising video samples, audio samples, and multi-modal annotations describing the multi-modal samples, the multi-modal annotations comprising at least either of textual description annotations or question-answer annotations; and updating at least either of the audio aligner, the language model, or a low-rank adaptive model in the machine learning model with the multi-modal samples. 7.The method of claim 6, wherein the machine learning model is obtained by: determining a loss function based on at least one reinforcement learning indicator associated with the machine learning model; and updating the machine learning model based on the loss function. 8.The method of claim 7, wherein the determining the loss function comprises: receiving reinforcement learning samples comprising a first sample and a second sample; ​ ​ ​ determine at least one loss function corresponding to the at least one reinforcement learning indicator based on the first sample and the second sample; and determine the loss function based on the at least one loss function. 9.The method of claim 8, wherein determining the at least one loss function corresponding to the at least one reinforcement learning indicator comprises: for a reinforcement indicator in the at least one reinforcement indicator, determining a first evaluation corresponding to the reinforcement learning indicator based on the first sample, and determining a second evaluation corresponding to the reinforcement learning indicator based on the second sample; determining a difference between the first evaluation and the second evaluation; and determining a loss function corresponding to the reinforcement learning indicator based on the difference. 10.The method of claim 9, wherein the at least one reinforcement indicator comprises a completeness indicator, the first sample comprises first multi-modal data and first transcription annotation, and determining the first evaluation comprises: dividing the first transcription annotation into a plurality of events; comparing a prediction of transcription generated by processing the first multi-modal data by the machine learning model with the plurality of events; and determining the first evaluation based on a result of the comparison. 11.The method of claim 9, wherein the at least one reinforcement indicator comprises a repetition indicator, the first sample comprises first multi-modal data, and determining the first evaluation comprises: dividing a prediction of transcription generated by processing the first multi-modal data by the machine learning model into a plurality of phrases; determining a number of repeated words of words in the plurality of phrases; and determining the first evaluation based on the number of repeated words and a total number of words of the prediction of transcription. 12.The method of claim 9, wherein the at least one reinforcement indicator comprises a backtracking indicator, the first sample comprises first multi-modal data and first transcription annotation corresponding to a data segment in the first multi-modal data, and determining the first evaluation comprises: dividing the first transcription annotation into a plurality of events; comparing a prediction of transcription generated by processing the data segment by the machine learning model with the plurality of events; and determining the first evaluation based on a result of the comparison. determining the loss function using the machine learning model. 14.The method of claim 1, wherein the processing request is a query request for content of the multi-modal data. 15.An apparatus for processing multi-modal data, comprising: a receiving module configured to receive multi-modal data and a processing request for the multi-modal data, the multi-modal data comprising at least video data and audio data, the processing request being expressed in natural language; 13. The method of claim 1, wherein determining the loss function comprises: an extracting module configured to extract a plurality of video tokens of the video data and a plurality of audio tokens of the audio data according to a predetermined time interval; an interleaving module configured to interleave the plurality of video tokens and the plurality of audio tokens according to the predetermined time interval to generate multi-modal features of the multi-modal data; and a processing module configured to process the multi-modal features to generate a result of the processing request. ​ ​ ​ ​ a processing module configured to process the multi-modal data based on the multi-modal features and text features of the processing request.

16. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-14.

17. A computer readable storage medium having stored thereon a computer program, the computer program when executed by a processor cause the processor to implement the method according to any one of claims 1-14.

18. A computer program product comprising a computer program, wherein the computer program when executed by a processor implements the method according to any one of claims 1-14.