Video processing method and device, equipment, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2024-09-13
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies rely solely on textual descriptions of video clips in video processing, resulting in missing information and insufficient processing accuracy, and failing to effectively combine other modal information for task processing.
By acquiring multimodal information of the target video segment, including video frames, audio, and text, and combining it with reference and supplementary information, the input to the machine learning model is determined, and the machine learning model is used to perform video processing tasks.
It improves the accuracy of video processing, especially in long video processing, and can better meet users' multimodal needs and question-and-answer tasks.
Smart Images

Figure CN122070698A_ABST
Abstract
Description
Method, apparatus, device, storage medium and program product for video processing TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, electronic device, computer-readable storage medium and computer program product for video processing. BACKGROUND
[0002] With the rapid development of computer technology, more and more applications, platforms or systems currently provide processing services for multimedia information, which brings great convenience to the general public. Multimedia information can include various types of content such as video, image, image set, text, audio, etc. The processing service for video can also be referred to as video processing service. The application, platform or system with video processing service can use video processing technology to process video for video editing, video understanding, video compression, etc.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for video processing is provided. The method comprises: obtaining a target video segment of a target video to be processed, the target video comprising a plurality of video segments; determining multi-modal information corresponding to the target video segment, the multi-modal information comprising video frames of the target video segment and further comprising one of: audio corresponding to the target video segment, text in each video frame of the target video segment; determining a model input for a machine learning model based at least on the multi-modal information and reference information associated with the target video segment, the reference information indicating information associated with at least one processed video segment of the plurality of video segments, and / or supplementary information matching the target video; and determining a processing result of a video processing task for the target video segment by providing the model input to the machine learning model and utilizing the machine learning model.
[0005] In a second aspect of the disclosure, an apparatus for video processing is provided. The apparatus includes: a segment obtaining module configured to obtain a target video segment of a target video to be processed, the target video including a plurality of video segments; an information determining module configured to determine multi-modal information corresponding to the target video segment, the multi-modal information including video frames of the target video segment and further including one of: audio corresponding to the target video segment, text in each video frame of the target video segment; an input determining module configured to determine a model input for a machine learning model based at least on the multi-modal information and reference information associated with the target video segment, the reference information indicating information associated with at least one processed video segment of the plurality of video segments, and / or supplementary information matching the target video; and a result determining module configured to determine a processing result for a video processing task of the target video segment by providing the model input to the machine learning model and utilizing the machine learning model.
[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program, which, when executed by a processor, implements the method of the first aspect.
[0008] In a fifth aspect of the disclosure, a computer program product is provided. The product includes a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.
[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the disclosure, nor is it used to limit the scope of the disclosure. Other features of the disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG. 2 shows a schematic diagram of an example architecture of video processing according to some embodiments of the present disclosure;
[0013] FIG. 3 illustrates an example of performing a video processing task using a machine learning model, according to some embodiments of the present disclosure;
[0014] FIG. 4 illustrates a flowchart of a method of video processing, according to some embodiments of the present disclosure;
[0015] FIG. 5 illustrates an exemplary structural block diagram of an apparatus for video processing, according to some embodiments of the present disclosure; and
[0016] FIG. 6 illustrates a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are illustrated in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted in a limited sense as set forth in the embodiments set forth herein. Rather, the embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] In the description of embodiments of the present disclosure, the term "comprising" and similar terms are understood to encompass open-ended inclusion, i.e., "including but not limited to". The term "based on" is understood to mean "based at least in part on". The term "one embodiment" or "the embodiment" is understood to mean "at least one embodiment". The term "some embodiments" is understood to mean "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0019] In this document, unless explicitly stated, performing a step "in response to" A does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0020] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, obtaining, using, storing or deleting) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained by appropriate means, wherein the relevant user can include any type of right subject, such as individuals, enterprises, groups.
[0022] For example, in response to receiving an active request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using information of the relevant user, so that the relevant user can autonomously select whether to provide information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.
[0023] As an optional but non-limiting implementation manner, in response to receiving an active request of a relevant user, the manner of sending a prompt information to the relevant user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0024] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0025] As used herein, the term "model" can learn an association between respective inputs and outputs from training data, such that a corresponding output can be generated for a given input after training is completed. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network", or a "learning network", which terms are used interchangeably herein.
[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.
[0027] Generally, machine learning can include three stages, i.e., a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to obtain consistent inferences from the training data that satisfy an expected objective. Through training, the model can be considered to have learned an association (also referred to as a mapping) from input to output from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to determine whether the model is able to provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the parameter values obtained through training to determine corresponding outputs.
[0028] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example 100, a user 140 can initiate a user request or provide a user input to a video processing system 110, which can instruct the video processing system 110 to perform a video processing task on a target video 120 to determine a processing result 150 corresponding to the target video 120. The user request or the user input can be any suitable type of request or input, e.g., a voice type, a text type, a gesture type, etc. For example, the video processing system 110 can provide an interactive interface for receiving the user request, which includes an input box in the interactive interface. The video processing system 110 can receive the target video 120 and the user input provided by the user 140 via the input box.
[0029] For example, the user 140 can upload the target video 120 to be processed and the user input “briefly describe the target video” to the video processing system 110. Based on the user input, the video processing system 110 can determine that the video processing task on the target video 120 is to determine description information of the target video 120. The video processing system 110 can further perform video processing on the target video 120 to determine the description information of the target video 120.
[0030] The target video 120 can include a plurality of video clips 125 (e.g., which can include video clips 125-1, 125-2, …, 125-N, N being any suitable positive integer, for convenience of description, one or more video clips can be collectively referred to as video clips 125). The plurality of video clips 125 can each have a duration less than a predetermined duration threshold. For example, the plurality of video clips 125 can each have a duration less than 30 seconds. Alternatively or additionally, the plurality of video clips 125 can also have the same duration. For example, the plurality of video clips 125 can each be a 25-second video clip. It can be understood that each video clip 125 includes at least one video frame.
[0031] In some embodiments, the video processing system 110 can perform the video processing task on the plurality of video segments 125 respectively to obtain respective processing results of the plurality of video segments 125. The video processing system 110 can in turn determine the processing result 150 of performing the video processing task on the target video 120 based on the respective processing results of the plurality of video segments 125.
[0032] In some embodiments, the video processing system 110 can utilize the machine learning model 130 to perform the video processing task on the target video 120 to determine the processing result 150 corresponding to the target video 120. The machine learning model 130 can be deployed locally at the video processing system 110 or at other devices / systems (e.g., remote devices). The video processing system 110 can process the target video 120 directly utilizing the machine learning model 130 deployed locally or by invoking the machine learning model 130 deployed at other devices / systems through a communication connection between the video processing system 110 and the other devices / systems.
[0033] The machine learning model 130 can be based on any appropriate model structure, including but not limited to, a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), etc. In some embodiments, the machine learning model 130 can be based on a visual language model (VLM). A visual language model is capable of answering questions based on learning from a large amount of images, videos, and corpus. The visual language model is capable of generating text output based on input images, videos, and text. The machine learning model 130 can also be based on other appropriate models. Although only a single machine learning model 130 is shown in the figure, there can be multiple machine learning models for the video processing system 110 to use. If the machine learning model 130 includes multiple models, the multiple models can include models configured to perform the same or similar tasks or models configured to perform different tasks.
[0034] The video processing system 110 can run on a suitable electronic device. The electronic device herein can include any computing system with computing capability, such as various computing devices / systems, end devices, server devices, etc. The end device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
[0035] The server device can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and basic cloud computing services such as big data and artificial intelligence platform. The server device may, for example, include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0036] It should be understood that the structure and function of the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0037] Traditionally, a video processing system can divide a long video into multiple video clips, and perform a video understanding task on each of the multiple video clips to obtain a respective text description information of each of the multiple video clips, for briefly summarizing the content in the video clip. The video processing system further determines an understanding task on the entire video by summarizing, inducing, and analyzing the respective text description information of each of the multiple video clips, to determine an induction summary of the entire video. That is, the summary of the entire video is determined based on the respective text corresponding to each of the multiple video clips. In this case, the final processing result for the entire video can only be determined based on the intermediate text description information of each video clip, without combining information of other modalities (e.g., audio of the target video, etc.), which can cause information missing problems and affect the accuracy of processing the target video. In addition, based on the text description information of the video clip, it is often impossible to generate other task processing results for the entire video or a single video clip, resulting in task limitations. For example, some tasks may wish to extract a clip related to a certain object from a long video. However, based on only the summary description of each video clip, such a task is difficult to achieve.
[0038] In view of this, according to an embodiment of the present disclosure, an improved scheme for video processing is provided. According to the scheme of the present embodiment, a target video clip of a target video to be processed is obtained, the target video comprising multiple video clips. Multiple modal information corresponding to the target video clip is determined, the multiple modal information comprising video frames of the target video clip, and in addition comprising audio corresponding to the target video clip and / or text in each video frame of the target video clip. A model input for a machine learning model is determined based at least on the multiple modal information and reference information associated with the target video clip. The reference information indicates information associated with at least one processed video clip of the multiple video clips, and / or supplementary information matching the target video. By providing the model input to the machine learning model, a processing result of a video processing task for the target video clip is determined using the machine learning model.
[0039] In this way, the machine learning model can be used to perform a video processing task on the target video clip based on the multiple modal information of the target video or the video clip of the target video and the relevant information of the processed video clip. By simultaneously considering various modal information (video, text, etc.) related to the video and various reference information, it is helpful to improve the accuracy of video processing and improve question and answer tasks based on video. In addition, since the relevant information of the previously processed video is continuously relied on in the processing of a single video clip, it can improve the task accuracy in the case of processing a long video.
[0040] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0041] FIG. 2 illustrates a schematic diagram of an example architecture 200 of video processing, according to some embodiments of the present disclosure. The architecture 200 can be implemented at the video processing system 110 of FIG. 1. For ease of discussion, the architecture 200 will be described with reference to the environment 100 of FIG. 1.
[0042] The video processing system 110 can obtain a target video 120 to be processed. In some embodiments, the target video 120 can be a video provided by a user (e.g., the user 140) to the video processing system 110 in real time, or can be a video pre-stored locally at the video processing system 110 or accessible by the video processing system 110 via other means.
[0043] The target video 120 can include or can be divided into a plurality (e.g., N) of video segments 125, and the present disclosure does not limit the length of the target video and the number of the video segments. The granularity of the target video and the video segments can be set according to actual application. For example, the target video 120 can be a TV episode, and the video segments 125 can be individual episodes in the TV episode. In another example, the target video 120 can be a single continuous video (e.g., a movie or a whole video shot by a user), and the video segments 125 can be divided from the target video 120 according to predetermined time intervals or in other manners.
[0044] In some embodiments, the plurality of video segments 125 has a certain order, e.g., the plurality of video segments 125 can be sequentially ordered according to the respective start times of the plurality of video segments 125. The video processing system 110 can process the plurality of video segments 125 sequentially based on the order of the plurality of video segments, for example.
[0045] In some embodiments, the video processing system 110 can also obtain a user input. The user input can be obtained synchronously with the target video 120 or separately, for example. The video processing system 110 can obtain the user input in any appropriate manner, including but not limited to obtaining a user input in text form via an input box, receiving a user input in voice form via a microphone, and the like. The user input can indicate a video processing task. In some embodiments, the user input can indicate a task processing requirement related to the visual information and the audio information of the target video 120, i.e., the user input can indicate a video processing task of processing the visual information and the audio information of the target video.
[0046] The video processing task can be, for example, a tracking task of at least one object in the target video 120, an identification task of the at least one object, a content extraction task of the target video 120, an editing task of the target video 120, and the like. The content extraction task of the target video 120 can indicate, for example, extracting a specified video from the target video 120, extracting a specified image from the target video 120, extracting a specified audio from the target video 120, extracting a specified text from the target video 120, and the like. The editing task of the target video 120 can include adding special effects, filters, stickers, speed, and the like to the target video.
[0047] If the video processing task is the tracking task of the at least one object, the processing result obtained by performing the video processing task on the target video 120 can at least include coordinates of the at least one object in each video frame of the target video 120. If the video processing task includes the identification task of the at least one object, the processing result obtained by performing the video processing task on the target video 120 can include respective identification results of the at least one object. If the object is a character, the identification result corresponding to each object can indicate information such as identity, appearance, and camp of the object. If the object is an item, the identification result corresponding to each object can indicate information such as purpose and composition of the object.
[0048] If the video processing task is the content extraction task of the target video 120, the processing result obtained by performing the video processing task on the target video 120 can include the content extracted from the target video 120, which can be determined based on the specific content of the video processing task. For example, if the video processing task indicates extracting a specific audio from the target video 120, the processing result can include the start time and end time of the audio in the target video, and the video processing system 110 can subsequently extract the audio from the target video 120 based on the start time and end time. Alternatively or additionally, the processing result can also directly include the audio extracted from the target video 120.
[0049] If the video processing task is the editing task of the target video 120, the processing result obtained by performing the video processing task on the target video 120 can include the edited target video 120. For example, if the video processing task indicates adding special effects to the target video 120, the processing result is the target video after adding the special effects. It can be understood that the video processing task can also be any other appropriate task, and the present disclosure does not limit the specific task and the corresponding processing result.
[0050] In some embodiments, the video processing system 110 can directly perform the video processing task indicated by the user input on the target video 120 to determine the processing result of the video processing task on the target video 120. Illustratively, the video processing system 110 can obtain the multi-modal information of the target video 120. The multi-modal information of the target video 120 at least includes the video frames of the target video 120. The multi-modal information of the target video 120 can further include the audio corresponding to the target video 120 and / or the text in each video frame of the target video 120.
[0051] The text in each video frame can be, for example, the caption of each video frame, the comment (barrage) in each video frame, the text presented in each video frame, etc. In some embodiments, the video processing system 110 can directly obtain the caption of the target video and determine the caption as the text of each video frame. In some embodiments, the video processing system 110 can further extract the text from each video frame using an optical character recognition (OCR) algorithm, etc. In some embodiments, the video processing system 110 can further extract the text from the audio of the target video 120 using an automatic speech recognition (ASR) algorithm, etc.
[0052] The video processing system 110 can determine the processing result 150 of the video processing task on the target video 120 based at least on the multi-modal information of the target video 120, for example, using the machine learning model 130. In some embodiments, to ensure the accuracy of the video processing performed on the target video 120, the video processing system 110 can sequentially process each video segment 125 of the target video 120 to determine the processing result 245 of each video segment 125 (which can include the processing result 245-1 corresponding to the video segment 125-1, the processing result 245-2 corresponding to the video segment 125-2, …, the processing result 245-N corresponding to the video segment 125-N, for convenience of description, one or more processing results corresponding to one or more video segments can be collectively referred to as the processing result 245) in view of the limitation of the model capability. The video processing system 110 can further determine the processing result 150 of the video processing task on the target video 120 based on the processing result 245 of each video segment 125.
[0053] In some embodiments, the video segment to be processed at present (e.g., the video segment 125-2) can be referred to as the target video segment. For example, if the target video 120 includes 10 video segments in total, and the video processing system 110 has performed the video processing task on the first video segment, and is currently performing the video processing task on the second video segment, the second video segment can be regarded as the target video segment. In some embodiments, the video processing system 110 can process at least one video segment at the same time, in which case the target video segment can include the at least one video segment.
[0054] The video processing system 110 can obtain a target video segment, and determine multi-modal information 220 corresponding to the target video segment. The multi-modal information 220 at least includes video frames 222 of the target video segment. In some embodiments, the multi-modal information 220 can further include audio 224 corresponding to the target video segment, and / or text 226 in the video frames of the target video segment. The audio 224 can be, for example, audio of the target video segment, and the text 226 can be, for example, subtitles of the target video segment. Regarding the text 226, in some embodiments, the video processing system 110 can directly obtain subtitles of the target video segment, and determine the subtitles as the text 226. In some embodiments, the video processing system 110 can further extract the text 226 from the video frames 222, the audio 224 of the target video segment, using OCR algorithm, ASR algorithm, and the like.
[0055] The video processing system 110 can determine model input for the machine learning model 130 based on at least the multi-modal information 220, for example. For example, the video processing system 110 can determine the model input for the machine learning model 130 based on the user input and the multi-modal information 220. In some embodiments, the video processing system 110 can further obtain reference information associated with the target video segment. The reference information can include information associated with at least one video segment that has been processed in the plurality of video segments (also referred to as information of the processed video segment) 212, and / or supplementary information 232 matching the target video 120.
[0056] In some embodiments, the video processing system 110 processes the video segment using the machine learning model 130, and the information 212 associated with at least one video segment that has been processed in the plurality of video segments includes a processing result of the at least one video segment by the machine learning model 130. The video processing system 110 can store the processing result of the at least one video segment by the machine learning model 130. For example, the video processing system 110 can store the processing result of the at least one video segment by the machine learning model 130 to a processing result library 210.
[0057] In some embodiments, the information 212 associated with the at least one processed video segment of the plurality of video segments 125 can further include at least part of the multi-modal information of the at least one video segment. Illustratively, for each of the at least one processed video segment, the video processing system 110 can further store at least part of the multi-modal information of the video segment together with the processing result of the video segment. The at least part of the multi-modal information can include, for example, the audio corresponding to the video segment. The video processing system 110 can obtain the information 212 associated with the at least one processed video segment of the plurality of video segments 125 from the processing result library 210, for example, in response to a request to perform a video processing task on a target video segment.
[0058] The supplemental information 232 can be, for example, supplemental information determined from a knowledge base (e.g., the external knowledge base 230) that matches the target video 120, and the information in the knowledge base can be pre-stored in the knowledge base via any suitable manner. The supplemental information 232 can include, for example, description information of at least one object associated with the target video 120. The at least one object can be, for example, at least one character in the target video 120, and the description information of the at least one character can include information such as the character’s name, identity, appearance, affiliation, image, actor’s name, etc. The supplemental information 232 can further include meta information (also referred to as metadata) of the target video 120, which can include information such as the name of the target video, the synopsis, the creator, etc.
[0059] The video processing system 110 can determine the model input for the machine learning model 130 based at least on the multi-modal information 220 and the reference information associated with the target video segment, which can include the information 212 associated with the at least one processed video segment of the plurality of video segments 125, and / or the supplemental information 232 that matches the target video 120. The video processing system 110 can determine the model input based on, for example, the user input, the multi-modal information 220, and the reference information. The video processing system 110 can utilize the machine learning model 130 to perform the video processing task on the target video segment by providing the model input to the machine learning model 130.
[0060] Alternatively or additionally, in some embodiments, the video processing system 110 can further obtain a prompt template, and determine the prompt input for the machine learning model 130 by filling the user input, the multi-modal information 220, and the reference information into the prompt template. The prompt input is the model input for the machine learning model 130. The video processing system 110 can utilize the machine learning model 130 to perform the video processing task on the target video segment by providing the prompt input to the machine learning model 130.
[0061] With respect to the specific way that the machine learning model 130 performs the video processing task on the target video segment, reference is made to FIG. 3, which illustrates an example 300 of performing a video processing task with the machine learning model 130, according to some embodiments of the present disclosure. The machine learning model 130 can include at least an image encoder 310 and an audio encoder 320. It is noted that the machine learning model 130 can include at least one machine learning model, and the image encoder 310, the audio encoder 320 and the machine learning model for determining the processing result can belong to the same machine learning model (i.e., the image encoder 310 and the audio encoder 320 can be part of the machine learning model), or can belong to different machine learning models.
[0062] The machine learning model 130 can encode the first information 302 of image type (e.g., can include at least the video frames 222 of the target video segment) in the model input into a visual feature representation 312 that matches the input space of the machine learning model 130 with the image encoder 310. The machine learning model 130 can encode the second information 304 of audio type (e.g., can include at least the audio 224 corresponding to the target video segment) in the model input into an audio feature representation 322 that matches the input space of the machine learning model 130 with the audio encoder 320. The machine learning model 130 can also determine a text feature representation 332 corresponding to the third information 306 of text type (e.g., can include the text 226 and the reference information) in the model input. For example, the machine learning model 130 can include a text encoder, and the machine learning model 130 can determine the text feature representation 332 corresponding to the third information 306 with the text encoder. The machine learning model 130 can determine the processing result 245 for the video processing task on the target video segment based on the visual feature representation 312, the audio feature representation 322 and the text feature representation 332.
[0063] Referring back to FIG. 2, taking the target video segment as the video segment 125-2 for example, the processing result corresponding to the target video segment is the processing result 245-2. In some embodiments, the video processing system 110 can store the processing result for the target video segment for use in processing of the next video segment of the target video segment. For example, the video processing system 110 can store the processing result 245-2 to the processing result library 210 for use in processing of the next video segment of the video segment 125-2.
[0064] The video processing system 110 can perform the video processing task on the plurality of video segments 125 in a similar manner to obtain a plurality of processing results 245 corresponding to the plurality of video segments 125. The video processing system 110 can determine the processing result 150 of the video processing task for the target video 120 based on the plurality of processing results 245 of the plurality of video segments 125. The video processing system 110 can analyze the plurality of processing results 245 of the plurality of video segments 125, for example, to determine the processing result 150 of the video processing task for the target video 120, and the disclosure does not limit the specific manner of determining the processing result 150 of the video processing task for the target video 120 based on the plurality of processing results 245 of the plurality of video segments 125.
[0065] In some embodiments, the video processing system 110 can determine the processing result 150 of the video processing task for the target video 120 as a reply to the user input. The video processing system 110 can further provide the reply to the user. For example, the video processing system 110 can provide a reply presentation interface, and can provide the reply to the user via the reply presentation interface. For another example, the interactive interface for receiving the user input can be a conversation interface between the user and a digital assistant. The video processing system 110 can provide the reply in the conversation interface. It can be understood that the video processing system 110 can provide the reply in any appropriate manner, and the disclosure does not limit the specific manner of providing the reply.
[0066] The video processing system 110 of the disclosure can perform the task processing in combination with the multi-modal information of the target video, so that the user's query can be accurately satisfied even if the query involves requirements other than the vision of the target video, such as audio-related requirements, and a processing result meeting the user's expectation can be generated.
[0067] In some embodiments, the video processing task for the target video 120 is a video processing task for the plurality of video segments 125 of the target video 120. For a target video segment, if the video processing task indicates tracking of at least one object in the target video segment, the processing result 245 at least includes coordinates of the at least one object in each video frame of the target video segment. It can be understood that if the at least one object includes a plurality of objects, the processing result 245 can include respective identities of the plurality of objects and respective coordinates of the plurality of objects in each video frame of the target video segment. Thus, the video processing system 110 can determine coordinates of the at least one object in each video frame of the target video 120 based on the coordinates of the at least one object in each video frame of each video segment 125. The video processing system 110 can determine the coordinates of the at least one object in each video frame of the target video 120 as a reply to the user input, for example.
[0068] Similarly, if the video processing task indicates to obtain a specific audio in the target video segment, the processing result 245 indicates at least a start time and an end time of the specific audio in the target video segment. The specific audio here can be background music, an intro song, an outro song, speech of at least one speaker, etc. in the target video segment. Based on the plurality of processing results 245 corresponding to the plurality of video segments 125, the video processing system 110 can determine at least one start time and a corresponding at least one end time of the specific audio in the target video 120.
[0069] It can be appreciated that the specific audio in the target video can be continuous or can appear intermittently. If a start time of the specific audio in a video segment A and an end time of the specific audio in a video segment B are adjacent or coincident, and the video segment A and the video segment B are two consecutive video segments, it can be determined that the specific audio does not interrupt in the two video segments, and the video processing system 110 can determine a start time of the specific audio in the two video segments as the start time of the specific audio in the video segment A and an end time of the specific audio in the two video segments as the end time of the specific audio in the video segment B. The video processing system 110 may, for example, directly determine the at least one start time and the corresponding at least one end time of the specific audio in the target video 120 as the reply to the user input. Alternatively or additionally, the video processing system 110 may, for example, extract the specific audio from the audio of the target video 120 based on the at least one start time and the corresponding at least one end time of the specific audio in the target video 120, and determine the specific audio as the reply to the user input.
[0070] If the video processing task indicates to recognize at least one object in the target video segment, the processing result 245 includes at least a description corresponding to each of the at least one object. Illustratively, for each object, if the object is a character, the description corresponding to the object can be a piece of text introducing an identity, an appearance, an affiliation, a gender, etc. of the character. It can be appreciated that if the at least one object includes a plurality of objects, the processing result 245 can include a respective identification and a respective description of each of the plurality of objects. In some embodiments, for different video segments 125, descriptions of different objects can be determined. Alternatively or additionally, in some embodiments, for the same video segment 125, descriptions of a plurality of objects can also be determined. Thus, the video processing system 110 can determine the respective description of the at least one object based on the plurality of video segments 125. The video processing system 110 may, for example, determine the description of the at least one object as the reply to the user input.
[0071] In some embodiments, the video processing task can also be a task involving multiple modalities. For example, the user input indicates a task processing requirement related to visual information and audio information of the target video 120, the video processing system 110 can extract a video from the target video 120 that satisfies the requirement and determine the video as a reply (or also said as a reply video) to the user input. Illustratively, the user input can indicate an audio requirement for the target video (or in other words, indicate to extract a video from the target video 120 that includes a specified audio, the specified audio being an audio that satisfies the audio requirement). Illustratively, the user input can indicate a requirement for a speech of at least one speaker appearing in the target video 120. The requirement can also indicate a scenario in which the at least one speaker speaks (e.g., monologue, dialogue, dialogue under a specific emotion, etc.). For example, the user input can indicate to extract all segments in which two specified speakers have a dialogue from the target video.
[0072] The multi-modal information of the target video 120 includes audio of the target video 120 (or in other words, the multi-modal information 220 of each of the plurality of video segments 125 includes audio 224), which can help the machine learning model 130 to better extract a video from the target video 120 that includes a specified audio. The plurality of processing results 245 corresponding to the plurality of video segments 125 can indicate a start time and an end time of the audio that satisfies the audio requirement in each of the video segments.
[0073] The video processing system 110 can in turn determine a start time and an end time of the audio that satisfies the audio requirement in the target video 120 based on the plurality of processing results 245 corresponding to the plurality of video segments 125. The video processing system 110 can extract a reply video from the target video based on the start time and the end time of the audio that satisfies the audio requirement in the target video 120, the audio corresponding to the reply video being the audio that satisfies the audio requirement, and the start time and the end time of the reply video in the target video 120 being the start time and the end time of the audio that satisfies the audio requirement in the target video 120.
[0074] Illustratively, the user input can indicate a text requirement for the target video (or in other words, indicate to extract a video from the target video 120 that includes a specified text, the specified text being a text that satisfies the text requirement). Illustratively, the user input can indicate a requirement for a subtitle appearing in the target video 120. The multi-modal information of the target video 120 includes text of the target video 120 (or in other words, the multi-modal information 220 of each of the plurality of video segments 125 includes text 226), which can help the machine learning model 130 to better extract a video from the target video 120 that includes a specified text. The plurality of processing results 245 corresponding to the plurality of video segments 125 can indicate a time of appearance of the text that satisfies the text requirement in each of the video segments.
[0075] The video processing system 110 can further determine, based on the plurality of processing results 245 corresponding to the plurality of video clips 125, a plurality of occurrence times of the text satisfying the text requirement in the target video 120. The video processing system 110 can extract, based on the plurality of occurrence times of the text satisfying the text requirement in the target video 120, a reply video from the target video, the text corresponding to the reply video being the text satisfying the text requirement.
[0076] In summary, according to embodiments of the present disclosure, a machine learning model can be utilized to perform a video processing task on a target video clip based on multi-modal information of the target video or the target video clip and related information of processed video clips. By considering various types of modal information (video, text, etc.) related to the video and various types of reference information at the same time, this helps to improve the accuracy of video processing and improve the task accuracy based on the video. In addition, since the related information of the previously processed video is constantly relied on in the processing process of a single video clip, this can improve the task accuracy in the case of processing a long video.
[0077] FIG. 4 illustrates a flowchart of a method 400 of video processing, according to some embodiments of the present disclosure. The method 400 can be implemented at the video processing system 110.
[0078] At block 410, the video processing system 110 obtains a target video clip of a target video to be processed, the target video including a plurality of video clips.
[0079] At block 420, the video processing system 110 determines multi-modal information corresponding to the target video clip, the multi-modal information including video frames of the target video clip and further including one of: audio corresponding to the target video clip, text in each video frame of the target video clip.
[0080] At block 430, the video processing system 110 determines, based at least on the multi-modal information and reference information associated with the target video clip, a model input for the machine learning model, the reference information indicating information associated with at least one processed video clip of the plurality of video clips, and / or supplementary information matching the target video.
[0081] At block 440, the video processing system 110 determines, by providing the model input to the machine learning model, a processing result of the video processing task for the target video clip using the machine learning model.
[0082] In some embodiments, determining, based at least on the multi-modal information and the reference information associated with the target video clip, the model input for the machine learning model includes: obtaining a user input indicating the video processing task; and determining the model input based on the user input, the multi-modal information, and the reference information.
[0083] In some embodiments, the machine learning model comprises at least an image encoder and an audio encoder, the machine learning model is configured to process the model input by: encoding first information of an image type in the model input into a visual feature representation matching an input space of the machine learning model using the image encoder, the first information comprising at least video frames of a target video segment; encoding second information of an audio type in the model input into an audio feature representation matching the input space of the machine learning model using the audio encoder, the second information comprising at least audio corresponding to the target video segment; determining a text feature representation corresponding to third information of a text type in the model input, the third information comprising at least text presented in each video frame and reference information; and determining a processing result of a video processing task for the target video segment based on the visual feature representation, the audio feature representation, and the text feature representation.
[0084] In some embodiments, if the video processing task indicates tracking of at least one object in the target video segment, the processing result comprises at least coordinates of the at least one object in each video frame of the target video segment; if the video processing task indicates obtaining of a specific audio in the target video segment, the processing result indicates at least a start time and an end time of the specific audio in the target video segment; and / or if the video processing task indicates recognition of at least one object in the target video segment, the processing result comprises at least a respective description of the at least one object.
[0085] In some embodiments, the information associated with the at least one processed video segment of the plurality of video segments comprises a processing result of the machine learning model on the at least one video segment, and wherein the supplementary information comprises supplementary information determined from a knowledge base matching the target video.
[0086] In some embodiments, the method 400 further comprises storing the processing result for the target video segment for use in processing of a next video segment of the target video segment.
[0087] In some embodiments, the method 400 further comprises determining a processing result of a video processing task for the target video based on respective processing results of the plurality of video segments in the target video.
[0088] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG. 5 shows an exemplary structural block diagram of an apparatus 500 for video processing according to some embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the video processing system 110. The various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0089] As shown in FIG. 5, the apparatus 500 includes a segment obtaining module 510 configured to obtain a target video segment of a target video to be processed, the target video including a plurality of video segments. The apparatus 500 further includes an information determining module 520 configured to determine multi-modal information corresponding to the target video segment, the multi-modal information including video frames of the target video segment and further including one of: audio corresponding to the target video segment, and text presented in the video frames of the target video segment. The apparatus 500 further includes an input determining module 530 configured to determine, based at least on the multi-modal information and reference information associated with the target video segment, a model input for a machine learning model, the reference information indicating information associated with at least one processed video segment of the plurality of video segments, and / or supplemental information matching the target video. The apparatus 500 further includes a result determining module 540 configured to determine, by providing the model input to the machine learning model, a processing result for a video processing task of the target video segment using the machine learning model.
[0090] In some embodiments, the input determining module 530 is further configured to: obtain a user input indicating the video processing task; and determine the model input based on the user input, the multi-modal information, and the reference information.
[0091] In some embodiments, the machine learning model at least includes an image encoder and an audio encoder, the machine learning model processing the model input via: encoding, using the image encoder, first information of an image type in the model input into a visual feature representation matching an input space of the machine learning model, the first information at least including the video frames of the target video segment; encoding, using the audio encoder, second information of an audio type in the model input into an audio feature representation matching the input space of the machine learning model, the second information at least including the audio corresponding to the target video segment; determining a text feature representation corresponding to third information of a text type in the model input, the third information at least including the text presented in the video frames and the reference information; and determining, based on the visual feature representation, the audio feature representation, and the text feature representation, the processing result for the video processing task of the target video segment.
[0092] In some embodiments, if the video processing task indicates tracking of at least one object in the target video segment, the processing result at least includes coordinates of the at least one object in the video frames of the target video segment; if the video processing task indicates obtaining of a specific audio in the target video segment, the processing result at least indicates a start time and an end time of the specific audio in the target video segment; and / or if the video processing task indicates recognition of at least one object in the target video segment, the processing result at least includes respective descriptions of the at least one object.
[0093] In some embodiments, the information associated with the at least one processed video segment of the plurality of video segments comprises a processing result of the at least one video segment by a machine learning model, and wherein the supplemental information comprises supplemental information determined from a knowledge base that matches the target video.
[0094] In some embodiments, the apparatus 500 further comprises a result storage module configured to store the processing result for the target video segment for use in processing of a next video segment of the target video segment.
[0095] In some embodiments, the apparatus 500 further comprises a result determination module configured to determine the processing result for the video processing task of the target video based on respective processing results of the plurality of video segments in the target video.
[0096] The units and / or modules included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 500 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0097] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices can include, for example, the video processing system 110 in FIG. 1.
[0098] FIG. 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 illustrated in FIG. 6 is merely exemplary and should not be construed as limiting on the functionality and scope of the embodiments described herein. The electronic device 600 illustrated in FIG. 6 can be used to implement the video processing system 110 of FIG. 1 or the apparatus 500 of FIG. 5.
[0099] As shown in FIG. 6, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 can include, but are not limited to, one or more processors or processing units 610, memory 620, storage 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit(s) 610 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 620. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 600.
[0100] Electronic device 600 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 620 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 630 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 600.
[0101] Electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard drive"), and a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM or other optical medium). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0102] Communication unit(s) 640 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines capable of communication through a communication connection. Accordingly, electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.
[0103] The input device 650 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., one or more devices that enable a user to interact with the electronic device 600, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 600 to communicate with one or more other electronic devices, as desired, via the communication unit 640. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0104] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a manufactured product, such that the instructions which run on the computer, the programmable data processing apparatus, and / or other devices implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0107] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0108] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The computer program product can also be an article of manufacture that comprises a computer readable medium. The medium can comprise a hard disk drive, a memory, or a floppy diskette, which can be accessed using a drive unit. Additionally, the medium can comprise a storage device that can store program codes. The storage device can include, but is not limited to, devices needing a platter and a read / write head, optical disk drives such as CD-ROM, DVD, Blu-ray Disc™ drives, memory devices such as flash drives, memory sticks, or any device that stores digital information. Additionally, the medium can include a computer readable file that can be downloaded into a memory or a storage device, such as a hard drive or a memory stick. The computer program product described herein can be implemented as a computer program product comprising a computer readable medium having stored instances of computer program code.
[0109] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Various implementations of the disclosure can be made and utilized in various ways and for various purposes, and some implementations of the disclosure can have additional elements which are incorporated to add additional functionality to the variations of the disclosure. It will be apparent to those of ordinary skill in the art that many modifications and variations can be made to the present disclosure without departing from the spirit or scope of the disclosure. Thus, it is intended that the present disclosure cover the modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.
Claims
1. A method for video processing, comprising: obtaining a target video segment of a target video to be processed, the target video comprising a plurality of video segments; determining multi-modal information corresponding to the target video segment, the multi-modal information comprising video frames of the target video segment and further comprising one of: audio corresponding to the target video segment, text presented in respective video frames of the target video segment; determining a model input for a machine learning model based at least on the multi-modal information and reference information associated with the target video segment, the reference information indicating information associated with at least one processed video segment of the plurality of video segments, and / or supplemental information matching the target video; and determining a processing result for a video processing task of the target video segment by providing the model input to the machine learning model and utilizing the machine learning model.
2. The method of claim 1, wherein determining a model input for a machine learning model based at least on the multi-modal information and reference information associated with the target video segment comprises: obtaining a user input indicating the video processing task; and determining the model input based on the user input, the multi-modal information, and the reference information.
3. The method of claim 1, wherein the machine learning model comprises at least an image encoder and an audio encoder, the machine learning model processing the model input via: encoding first information of an image type in the model input into a visual feature representation matching an input space of the machine learning model using the image encoder, the first information comprising at least video frames of the target video segment; encoding second information of an audio type in the model input into an audio feature representation matching the input space of the machine learning model using the audio encoder, the second information comprising at least audio corresponding to the target video segment; determining a text feature representation corresponding to third information of a text type in the model input, the third information comprising at least text presented in respective video frames and the reference information; and determining the processing result for the video processing task of the target video segment based on the visual feature representation, the audio feature representation, and the text feature representation.
4. The method of claim 1, wherein if the video processing task indicates tracking of at least one object in the target video segment, the processing result comprises at least coordinates of the at least one object in respective video frames of the target video segment; if the video processing task indicates obtaining a specific audio in the target video segment, the processing result indicates at least a start time and an end time of the specific audio in the target video segment; and / or if the video processing task indicates recognition of at least one object in the target video segment, the processing result comprises at least respective descriptions of the at least one object. 5. The method of claim 1, wherein the information associated with the at least one processed video segment of the plurality of video segments comprises a processing result of the at least one video segment by the machine learning model, and wherein the supplemental information comprises supplemental information determined from a knowledge base that matches the target video.
6. The method of claim 1, further comprising: storing the processing result for the target video segment for use in processing of a next video segment of the target video segment.
7. The method of claim 1, further comprising: determining a processing result for the video processing task of the target video based on the respective processing results of the plurality of video segments in the target video.
8. An apparatus for video processing, comprising: a segment obtaining module configured to obtain a target video segment of a target video to be processed, the target video comprising a plurality of video segments; an information determining module configured to determine multi-modal information corresponding to the target video segment, the multi-modal information comprising video frames of the target video segment and further comprising one of: audio corresponding to the target video segment, text in respective video frames of the target video segment; an input determining module configured to determine a model input for a machine learning model based at least on the multi-modal information and reference information associated with the target video segment, the reference information indicating information associated with at least one processed video segment of the plurality of video segments, and / or supplemental information matching the target video; and a result determining module configured to determine a processing result for a video processing task of the target video segment by the machine learning model by providing the model input to the machine learning model.
9. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method of any of claims 1-7.
10. A computer-readable storage medium having stored thereon a computer program, the computer program executable by a processor to implement the method of any of claims 1-7.
11. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any of claims 1-7.