Video processing method and apparatus, electronic device, and computer program product

By generating the temporal and spatial features of the video and utilizing the spatiotemporal query transformation module, the problem of alignment difficulties in multimodal models during video data processing is solved, thereby improving the performance and understanding capabilities of the video processing model.

WO2025218166A1PCT designated stage Publication Date: 2025-10-23DOUYIN VISION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133587
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2024-11-21
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Multimodal models struggle to effectively align the spatial and temporal features of video data, resulting in poor processing performance.

Method used

By generating the temporal and spatial features of the video and utilizing the spatiotemporal query transformation module, including time query and spatial expert network, text output corresponding to the video is generated.

Benefits of technology

It significantly enhances the visual-language alignment capability of the video processing model, improving the model's understanding of video content and processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133587_23102025_PF_FP_ABST
    Figure CN2024133587_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a video processing method and apparatus, an electronic device, and a computer program product. The method comprises: determining a temporal feature and a spatial feature of a video. Furthermore, the method comprises: on the basis of the temporal feature and the spatial feature, using a spatio-temporal query conversion module to generate a text output corresponding to the video, wherein the spatio-temporal query conversion module comprises a temporal query and a corresponding temporal expert network, and a spatial query and a corresponding spatial expert network.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, electronic device and computer program product for video processing

[0001] Cross Reference to Related Applications

[0002] This application claims priority to the Chinese patent application No. 202410479933.7, filed on April 19, 2024, and entitled “Method, apparatus, electronic device and computer program product for video processing”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present application relates to the technical field of computer, and more particularly to a method, apparatus, electronic device and computer program product for video processing. BACKGROUND

[0004] With the continuous development of artificial intelligence technology, generative language models have become increasingly important. These models can automatically generate high-quality, natural and fluent text, bringing new possibilities to various application fields. From natural language processing to creative content generation, generative language models play a key role in dialogue systems, text summarization, translation and creative copywriting.

[0005] With the growing demand for multi-modal intelligence, multi-modal models that integrate generative language models and image processing models have become increasingly important. Such models can handle multiple data types such as text and images simultaneously, thus more accurately understanding and expressing complex semantic information. In the fields of natural language processing, computer vision and intelligent dialogue, multi-modal models provide new possibilities for achieving more comprehensive and intelligent human-computer interaction. SUMMARY

[0006] Embodiments of the present disclosure provide a method, apparatus, electronic device, computer program product and medium for video processing.

[0007] According to a first aspect of the present disclosure, a method for video processing is provided. The method includes determining a temporal feature and a spatial feature of a video. In addition, the method further includes generating a text output corresponding to the video based on the temporal feature and the spatial feature using a spatio-temporal query conversion module, the spatio-temporal query conversion module including a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0008] According to a second aspect of the present disclosure, a device for video processing is provided. The device comprises a spatio-temporal feature determination module configured to determine a temporal feature and a spatial feature of a video. In addition, the device further comprises a text output generation module configured to generate a text output corresponding to the video based on the temporal feature and the spatial feature using a spatio-temporal query conversion module, the spatio-temporal query conversion module comprising a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0009] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises a processor and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.

[0010] In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and comprises computer executable instructions, which, when executed, cause a computer to perform the steps of the method of the first aspect of the present disclosure.

[0011] In a fifth aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium has stored thereon one or more computer instructions, wherein the one or more computer instructions, when executed by a processor, implement the method according to the first aspect.

[0012] The summary is presented to introduce some aspects of the concepts in a simplified form that are further described below in the detailed description. The summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0013] The above and other features, aspects and advantages of various embodiments of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals denote like elements, and wherein:

[0014] FIG. 1 shows a schematic diagram of an example environment in which devices and / or methods according to embodiments of the present disclosure can be implemented;

[0015] FIG. 2 shows a flowchart of a method for video processing according to embodiments of the present disclosure;

[0016] FIG. 3 shows a schematic diagram of a process of processing a video using a video processing model according to embodiments of the present disclosure;

[0017] FIG. 4A shows a schematic diagram of a structure of a spatio-temporal query conversion module according to embodiments of the present disclosure;

[0018] FIG. 4B shows a schematic diagram of a matrix of cross-attention masks, according to an embodiment of the present disclosure;

[0019] FIG. 5 shows a schematic diagram of a process of pre-training of a video processing model, according to an embodiment of the present disclosure;

[0020] FIG. 6 shows a schematic diagram of a process of interacting with a video processing system, according to an embodiment of the present disclosure;

[0021] FIG. 7 shows a block diagram of an apparatus for video processing, according to an embodiment of the present disclosure; and

[0022] FIG. 8 shows a block diagram of an electronic device, according to an embodiment of the present disclosure.

[0023] In all the drawings, the same or similar reference numerals indicate the same or similar elements. DETAILED DESCRIPTION

[0024] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.

[0026] In the description of the embodiments of the present disclosure, the term “comprising” and similar terms are to be understood as open-ended, i.e., “including but not limited to”. The term “based on” is to be understood as “based at least in part on”. The term “one embodiment” or “the embodiment” is to be understood as “at least one embodiment”. The terms “first”, “second”, etc. can refer to different or similar objects unless explicitly stated otherwise. Further explicit and implicit definitions can be included below.

[0027] As mentioned before, multi-modal models play an important role in many fields, and related multi-modal models can perform well when processing text and image data. However, when it comes to processing video data, multi-modal models perform poorly. After analysis, it is found that video data has both spatial features and temporal features, so there are often difficulties in semantic alignment of features, resulting in poor performance of multi-modal models for processing video.

[0028] To this end, embodiments of the present disclosure propose a scheme for video processing, which first generates temporal and spatial features of a video, and then generates a text output corresponding to the video using the temporal and spatial features of the video using a spatio-temporal query conversion module including a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0029] Embodiments of the present disclosure effectively generate spatio-temporal semantic alignment representations of input videos by using a spatio-temporal query conversion module, which utilizes a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network, significantly enhancing the visual language alignment capability of the video processing model, making it better understand the video content. With the improvement of the understanding capability of the model for the video input, the effect of the video processing model is also significantly improved. The spatio-temporal query conversion module not only improves the performance of the processing model, but also helps to lay the foundation for subsequent research and application of multi-modal video processing.

[0030] FIG. 1 shows a schematic diagram of an example environment in which devices and / or methods according to embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include a computing device 110, which can be a user terminal, a mobile device, a computer, etc., which can also be a computing system, a single server, a distributed server, or a cloud-based server. The computing device 110 can receive a video 120. Embodiments of the present disclosure can process multi-modal data, i.e., data of video modal and data of text modal. Multi-modal data refers to a collection containing multiple types or forms of data, which can be from different sensors, devices or sources, and usually includes at least two of multiple forms such as text, image, audio, video, etc.

[0031] The computing device 110 can include a video processing system 130 that can generate temporal features 132 and spatial features 134 of the video 120. Unlike static pictures, videos not only have spatial information of image frames, but also contain temporal information between image frames. Therefore, it can be desirable to capture not only spatial features 134 of static scenes of the video 120, but also temporal features 132 of changes over time. A spatio-temporal query conversion module 136 can semantically align the temporal features 132 and the spatial features of the video 120. The spatio-temporal query conversion module 136 can include temporal queries 138 and corresponding temporal expert networks 140, spatial queries 142 and corresponding spatial expert networks 144. The video processing system 130 can generate a text output 150 corresponding to the video 120. The video processing system 130 can generate the text output 150 corresponding to the video 120 based on the temporal features 132 and the spatial features 134 using the spatio-temporal query conversion module 136 that includes the temporal queries 138 and the corresponding temporal expert networks 140, and the spatial queries 142 and the corresponding spatial expert networks 144.

[0032] It should be understood that the architecture and functionality of the example environment 100 is described for illustrative purposes only and is not meant to limit the scope of the present disclosure. Embodiments of the present disclosure can be applied to other environments with different structures and / or functionalities.

[0033] Processes according to embodiments of the present disclosure will be described in detail below with reference to FIGS. 2-8. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of the present disclosure. It can be understood that the following described embodiments can also include additional actions not shown and / or can omit actions shown, and the scope of the present disclosure is not limited in this respect.

[0034] FIG. 2 illustrates a flowchart of a method 200 of video processing according to embodiments of the present disclosure. At block 202, temporal features and spatial features of a video can be determined. For example, with reference to FIG. 1, the video processing system 130 can determine the temporal features 132 and the spatial features 134 of the video 120.

[0035] At block 204, a text output corresponding to the video can be generated based on the temporal features and the spatial features using a spatio-temporal query conversion module, which includes a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network. For example, referring to FIG. 1, the video processing system 130 can generate a text output 150 corresponding to the video 120 based on the temporal features 132 and the spatial features 134 using the spatio-temporal query conversion module 136, which includes the temporal query 138 and the corresponding temporal expert network 140 and the spatial query 142 and the corresponding spatial expert network 144.

[0036] Thus, according to the method 200 of embodiments of the present disclosure, by utilizing the spatio-temporal query conversion module, the spatio-temporal semantic aligned representation of the input video can be effectively generated, significantly enhancing the visual language alignment capability of the video processing model, making it better understand the video content. With the improvement of the understanding capability of the model to the video input, the effect of the video processing model is also significantly improved. In addition, the spatio-temporal query conversion module not only improves the performance of the processing model, but also helps to lay the foundation for subsequent research and application of multi-modal video processing.

[0037] FIG. 3 shows a schematic diagram of a process 300 of processing a video using a video processing model according to embodiments of the present disclosure. As shown in FIG. 3, a video 302 can be input into a visual encoder. Embodiments of the present disclosure do not limit the format of the video 302, and can be applied to various video formats, including but not limited to MP4, AVI, MOV, etc., regardless of which device or software generates the video data, and a corresponding semantic aligned representation can be generated by the spatio-temporal query conversion module. The visual encoder 304 can receive the video 302 and generate corresponding visual features or visual representations. The visual encoder is a neural network model used for image processing and computer vision tasks, which can be used as a visual feature extractor, capable of converting input image data into high-dimensional feature vectors containing semantic and visual information of the image, which can be used for various tasks such as image classification, object detection, image semantic segmentation, etc.

[0038] The attention pooling module 306 can receive the visual features of the video 302 and generate the spatial features 308 and the temporal features 310. The attention pooling module 306 can be used to decouple the spatio-temporal features of the video. Explicit modeling of the spatio-temporal features of the video is crucial for the language model to effectively understand the video content, which allows the language model to capture rich semantic information, dynamic changes and contextual clues, thereby enhancing the language model's ability to understand the video content. The attention pooling module 306 consists of a cross-attention layer and a feedforward layer, which can obtain the spatial features 308 and the temporal features 310 of the video 302 through a learnable pooling process. For example, assuming that the input video is where T is the number of frames, H, W and C represent the height, width and number of channels of each image frame. When the video 302 passes through the visual encoder 304, an initial video embedding where N represents the number of image patches per frame, and D represents the feature dimension.

[0039] To extract spatio-temporal features from the video 302, the attention pooling module 306 can introduce two attention pooling queries to learn to extract corresponding features. Specifically, a temporal pooling query can be utilized to perform cross-attention operation on the video embedding to generate the temporal feature 310 This process can be represented by equation (1) and equation (2):

[0040] where CA(Q s , x, x) represents the cross-attention network, and FFN(·) represents the feed-forward neural network. Similarly, a spatial pooling query can be utilized to perform cross-attention operation on the transposed video embedding to generate the spatial feature 308, which can be represented by equation (3) and equation (4):

[0041] As shown in FIG. 3, the spatio-temporal query conversion module 312 can utilize the temporal query 314, the spatial query 316 and the fusion query 318 to process the spatial feature 308 and the temporal feature 310. For example, the temporal query 314, the spatial query 316 and the fusion query 318 can be regarded as feature extractors to extract corresponding features from the spatial feature 308 and the temporal feature 310 to achieve semantic alignment. In some embodiments, the feature dimension of the temporal query 314 and the spatial query 316 can be 32-dimension, and the dimension of the fusion query can be 1-dimension. In some embodiments, only the temporal query 314 and the spatial query 316 can be included, without the fusion query 318.

[0042] The spatio-temporal query conversion module 312 can generate video features 320 of the video 302, and align the dimension of the video features 320 with the dimension of the language model 324 through a multi-layer perception (MLP) layer 322. For example, the video feature dimension can be 1024, while the input dimension of the language model 324 is 4096, thus the MLP layer 322 can be utilized to convert the dimension of the video features 320 to the input dimension of the language model 324 for further processing of the video features 320 to generate the output 328. In addition, the video feature dimension and the input dimension of the language model can also be other numbers, and the embodiments of the present disclosure do not limit this. For example, the output 328 can be a summary or a description of the video 302. In some embodiments, the prompt content 326 can also be input to the language model 324 to instruct the language model 324 to generate the output indicated by the prompt content 326. For example, the prompt content 326 can be “what is the man doing in the video”, then the language model 324 can generate the corresponding output 328 according to its understanding of the video content. For example, the output 328 can be “a girl is feeding a boy a piece of pizza”, and the process can be regarded as a video question answering task. In addition, the video processing model of the present disclosure can also perform other tasks, including but not limited to multi-modal instruction following, video object localization, video description generation, and video summary generation, etc.

[0043] FIG. 4A shows a schematic diagram of a structure 400A of a spatio-temporal query conversion module according to an embodiment of the present disclosure. As shown in FIG. 4A, a video 402 can generate visual features through a visual encoder 404, and generate spatial features 408 and temporal features 410 through an attention pooling module 406. The spatio-temporal query conversion module 412 can connect the visual encoder and the language model, and bridge the gap between the visual representation and the language pattern. Specifically, the spatio-temporal query conversion module 412 can include three expert networks, a temporal expert network 414, a spatial expert network 416, and a fusion expert network 418. In the spatio-temporal query conversion module 412, each expert network is an independent sub-model, and each expert network is responsible for processing a specific subspace or subtask. In the training process, the spatio-temporal query conversion module 412 dynamically combines the outputs of the individual expert networks by selecting and weighting between different expert networks, to effectively process the temporal features, the spatial features, and the spatio-temporal fusion features in the video data.

[0044] The spatio-temporal query conversion module 412 can also include a learnable spatial query 420, a temporal query 422, and a fusion query 424. In some embodiments, the spatial query 420, the temporal query 422, and the fusion query 424 can extract a spatial representation, a temporal representation, and a spatio-temporal fusion representation of the video 402, respectively. In some embodiments, a spatial query 420 with a feature dimension of 32, a temporal query 422 with a feature dimension of 32, and a fusion query 424 with a feature dimension of 1 can be used. It should be appreciated that other dimensional query vectors can also be used, and embodiments of the present disclosure do not limit the dimension of the vectors. Further, in some embodiments, the spatio-temporal query conversion module 412 can not include the fusion query 424 and the corresponding fusion expert network 418.

[0045] The spatial query 420, the temporal query 422, and the fusion query 424 interact through a self-attention network 426 and cross-attention network 428 with the spatial features 408 and the temporal features 410 of the video 402. In addition, each query can also interact with the video description 430 through the same self-attention network 426. The video description 430 can be a description text corresponding to the video 402. For example, the video description 430 can be “a girl is feeding a boy a piece of pizza”. A feed-forward neural network 432 can receive the output of the self-attention network 426 to further encode the video description 430. To ensure that each query can perform cross-attention operation with the corresponding features, embodiments of the present disclosure design a matrix of cross-attention masks in the cross-attention network 428. The cross-attention masks will be described below in connection with FIG. 4B.

[0046] FIG. 4B illustrates a schematic diagram of a matrix 400B of cross-attention masks, according to an embodiment of the present disclosure. As shown in FIG. 4B, where a gray block indicates that the corresponding position of the matrix is “1”, and a white block indicates that the corresponding position of the matrix is “0”. The cross-attention masks can control the visibility of the spatial queries and the temporal queries on various spatio-temporal features, where the spatial queries and the temporal queries are limited to focus on their respective features, while the fusion queries can focus on all features. For example, the mask corresponding to the spatial query 450 and the spatial feature 452 is 1, indicating that the spatial query 450 is visible to the spatial feature 452. In addition, the mask corresponding to the spatial query 450 and the temporal feature 462 is 0, indicating that the spatial query 450 is not visible to the temporal feature 462. Likewise, the mask corresponding to the temporal query 460 and the temporal feature 462 is 1, but the mask corresponding to the temporal query 460 and the spatial feature 452 is 0. In addition, the masks corresponding to the fusion query 470 and the spatial feature 452 and the temporal feature 462 are both 1, because the fusion query 470 can focus on both the spatial feature 452 and the temporal feature 462 at the same time. The design of the cross-attention masks enables precise control of the allocation of attention when processing spatio-temporal features, thereby more effectively capturing important information in the video data.

[0047] Referring back to FIG. 4A, different expert networks are used to process different query vectors in parallel. In addition, the L VTM The loss function 434, L VTC The loss function 436, L VTG and the loss function 438 are used to train the spatio-temporal query conversion module 412. For example, the loss function 434 is used to represent a video text matching (VTM) loss function, the loss function 436 is used to represent a video text contrast (VTC) loss function, and the loss function 438 is used to represent a video localization text generation (VTG) loss function. VTM The loss function 434 is used to represent a video text matching (VTM) loss function, the loss function 436 is used to represent a video text contrast (VTC) loss function, and the loss function 438 is used to represent a video localization text generation (VTG) loss function. VTC The loss function 436 is used to represent a video text contrast (VTC) loss function, and the loss function 438 is used to represent a video localization text generation (VTG) loss function. VTG The loss function 438 is used to represent a video localization text generation (VTG) loss function.

[0048] FIG. 5 illustrates a schematic diagram of a process 500 of pre-training of a video processing model according to an embodiment of the present disclosure. As shown in FIG. 5, at block 502, the spatio-temporal query transformation module is trained independently, for a first stage of pre-training. In the first stage of pre-training, the spatio-temporal query transformation module is trained to extract spatio-temporal video embeddings that are most relevant to the video description. In connection with FIG. 4A, the spatio-temporal query transformation module 412 is trained to extract spatio-temporal video embeddings that are most relevant to the video description 430. In some embodiments, in the first stage of pre-training, the spatio-temporal query transformation module is trained by jointly optimizing the VTM loss function, the VTC loss function, and the VTG loss function. In computing the VTM loss function, the average of the spatial query and the temporal query can be computed separately, which can then be concatenated with the classification label and input into a binary classification task to predict whether the video and the video description match.

[0049] At block 504, the spatio-temporal query transformation module is connected with the language model, for a second stage of pre-training. For example, in connection with FIG. 3, the spatio-temporal query transformation module 312 can be connected with the language model 324 to train the spatio-temporal query transformation module 312. In some embodiments, the parameters of the language model 324 can remain unchanged, which can speed up the process of model training. In addition, to implement the connection of the spatio-temporal query transformation module with the language model, an MLP layer can be used to project the video features from the spatio-temporal query transformation module to the embedding space of the language model. For example, in connection with FIG. 3, the spatio-temporal query transformation module 312 and the language model 324 can be connected by using the MLP layer 332, which can project the video features 320 to the embedding space of the language model 324.

[0050] FIG. 6 illustrates a schematic diagram of a process 600 of interacting with a video processing system according to an embodiment of the present disclosure. As shown in FIG. 6, a user can input a video 602 to the video processing system, which can have a video processing model according to an embodiment of the present disclosure. For example, the video 602 can be a video related to a vehicle. In some embodiments, the video processing system can also display the image frames below the video 602. Then, the user can interact with the video processing system. For example, the user can send a dialogue 604 “What happened in this video?” and then the video processing system can generate a corresponding dialogue 606 “In this video, the rearview mirror of a car was damaged after the accident.” In addition, the video processing system can also support multi-round dialogues. For example, the user can input a dialogue 608 to further inquire about the content related to the video 602, and the video processing system can further generate a dialogue 610 to answer.

[0051] FIG. 7 shows a block diagram of an apparatus 700 for video processing according to an embodiment of the present disclosure. As shown in FIG. 7, the apparatus 700 includes a spatio-temporal feature determination module 702 configured to determine a temporal feature and a spatial feature of a video. In addition, the apparatus 700 further includes a text output generation module 704 configured to generate a text output corresponding to the video based on the temporal feature and the spatial feature using a spatio-temporal query conversion module, the spatio-temporal query conversion module including a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0052] FIG. 8 shows a block diagram of an electronic device 800 according to certain embodiments of the present disclosure. FIG. 8 shows a block diagram of an electronic device 800 that can be a device or apparatus described in embodiments of the present disclosure. As shown in FIG. 8, the device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage unit 808. Various programs and data required for operation of the device 800 can also be stored in the RAM 803. The CPU / GPU 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804. Although not shown in FIG. 8, the device 800 can further include a coprocessor.

[0053] A plurality of components in the device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc., an output unit 807, such as various types of displays, speakers, etc., a storage unit 808, such as a magnetic disk, a magneto-optical disk, etc., and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0054] The various methods or processes described above can be performed by the CPU / GPU 801. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU / GPU 801, one or more steps or actions of the methods or processes described above can be performed.

[0055] In some embodiments, the methods and processes described above can be tied to a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions thereon for performing various aspects of the present disclosure.

[0056] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0057] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0058] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0059] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include non- transitory computer readable storage media that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions for causing an apparatus to implement various aspects of the functions / acts specified in the flowchart and / or block diagram block or blocks can be utilized.

[0060] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0061] The computer program product of the present disclosure can have a computer readable medium storing instructions that, when executed by one or more processors, enable the performance of any of the methods disclosed herein. The computer readable medium can be a machine-readable storage medium, having stored thereon, a computer program, a piece of code, a instruction, or some combination thereof, to cause a machine to perform any of the features described herein. Examples of a machine-readable storage medium include, but are not limited to, any type of disk including floppy disks, optical disks, compact disk-read only memories (CD-ROMs), compact disk-read / write (CD-R / W) disks, magnetic tapes, floppy disks, optical disks, magneto-optical disks, read only memories (ROMs), random access memories (RAMs), erasable programmable read only memories (EPROMs), electrically erasable programmable read only memories (EEPROMs), flash memories, solid state drives (SSDs) or other types of storage media that can be used for storing data and / or computer program code that can direct a computer function. The computer program product or software program can cause a computer to perform any of the features disclosed herein, including any of the methods disclosed herein.

[0062] The above-described embodiments of the present disclosure have been described in connection with what is presently considered to be the most practical and preferred implementations. However, it should be understood that the present disclosure is in no way limited to the disclosed implementations. On the contrary, it is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the described embodiments, including other implementations that are now known or become known in the future, including, but not limited to, any improvements that are now known or become known. The choice of words in the description is intended to be construed most broadly and to include equivalents.

[0063] Some example implementations of the present disclosure are listed below.

[0064] Example 1. A method of video processing, comprising:

[0065] determining temporal features and spatial features of a video; and

[0066] generating, based on the temporal features and the spatial features, a text output corresponding to the video using a spatio-temporal query conversion module, the spatio-temporal query conversion module comprising a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0067] Example 2. The method of example 1, wherein the temporal features and the spatial features are generated by an attention pooling module, the method further comprising:

[0068] generating, based on the video, visual features of the video; and

[0069] generating, based on the visual features, the temporal features and the spatial features using a cross-attention network and a feed-forward neural network in the attention pooling module.

[0070] Example 3. The method of any one of examples 1-2, wherein the spatio-temporal query conversion module further comprises a fusion query and a corresponding fusion expert network.

[0071] Example 4. The method of any one of examples 1-3, further comprising:

[0072] obtaining prompt text; and

[0073] generating, using the spatio-temporal query conversion module, a text output indicated by the prompt text based on the temporal feature, the spatial feature, and the prompt text.

[0074] Example 5. The method of any one of examples 1-4, wherein generating the text output indicated by the prompt text comprises:

[0075] generating, using the spatio-temporal query conversion module, video features based on the temporal feature and the spatial feature; and

[0076] generating the text output indicated by the prompt text based on the video features and the prompt text.

[0077] Example 6. The method of any one of examples 1-5, wherein the spatio-temporal query conversion module further comprises a cross-attention mask, the cross-attention mask comprising a temporal mask, a spatial mask, and a fusion mask.

[0078] Example 7. The method of any one of examples 1-6, wherein the text output is generated by a video processing model, the video processing model comprising the spatio-temporal query conversion module, the method further comprising:

[0079] training the video processing model based on a training video, a video description text of the training video, and a training prompt text.

[0080] Example 8. The method of any one of examples 1-7, wherein training the video processing model comprises:

[0081] generating, by a visual encoder, training visual features based on the training video;

[0082] generating, by an attention pooling module, training temporal features and training spatial features based on the training visual features;

[0083] generating, by the spatio-temporal query conversion module, a conversion output based on the training temporal features, the training spatial features, and the video description text; and

[0084] training the attention pooling module and the spatio-temporal query conversion module based on the conversion output.

[0085] Example 9. The method of any one of examples 1-8, wherein training the attention pooling module and the spatio-temporal query conversion module comprises:

[0086] training the attention pooling module and the spatio-temporal query conversion module based on the video-text matching loss function, the video-text contrastive loss function, and the video-localization text generation loss function.

[0087] Example 10. The method of any one of examples 1-9, further comprising:

[0088] outputting, by a language model, a target output indicated by the training prompt text based on the conversion output and the training prompt text; and

[0089] training the attention pooling module and the spatio-temporal query conversion module based on the target output indicated by the training prompt text, parameters of the language model remaining unchanged in the training.

[0090] Example 11. An apparatus for video processing, comprising:

[0091] a spatio-temporal feature determination module configured to determine a temporal feature and a spatial feature of a video; and

[0092] a text output generation module configured to generate, based on the temporal feature and the spatial feature, a text output corresponding to the video using a spatio-temporal query conversion module, the spatio-temporal query conversion module comprising a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0093] Example 12. The apparatus of example 11, wherein the temporal feature and the spatial feature are generated by an attention pooling module, the apparatus further comprising:

[0094] a visual feature generation module configured to generate, based on the video, a visual feature of the video, and the spatio-temporal feature generation module configured to generate, based on the visual feature, the temporal feature and the spatial feature using a cross-attention network and a feed-forward neural network in the attention pooling module.

[0095] Example 13. The apparatus of any one of examples 11-12, wherein the spatio-temporal query conversion module further comprises a fusion query and a corresponding fusion expert network.

[0096] Example 14. The apparatus of any one of examples 11-13, the apparatus further comprising:

[0097] a prompt text acquisition module configured to acquire a prompt text; and

[0098] An instruction output generation module configured to generate, based on the temporal feature, the spatial feature, and the prompt text, a text output indicated by the prompt text using the spatio-temporal query conversion module.

[0099] Example 15. The apparatus of any one of examples 11-14, wherein the instruction output generation module comprises:

[0100] A video feature generation module configured to generate, based on the temporal feature and the spatial feature, a video feature using the spatio-temporal query conversion module; and

[0101] A second instruction output generation module configured to generate, based on the video feature and the prompt text, the text output indicated by the prompt text.

[0102] Example 16. The apparatus of any one of examples 11-15, wherein the spatio-temporal query conversion module further comprises a cross-attention mask, the cross-attention mask comprising a temporal mask, a spatial mask, and a fusion mask.

[0103] Example 17. The apparatus of any one of examples 11-16, wherein the text output is generated by a video processing model, the video processing model comprising the spatio-temporal query conversion module, the apparatus further comprising:

[0104] A video processing model training module configured to train the video processing model based on a training video, a video description text of the training video, and a training prompt text.

[0105] Example 18. The apparatus of any one of examples 11-17, wherein the video processing model training module comprises:

[0106] A training visual feature generation module configured to generate, based on the training video, a training visual feature by a visual encoder;

[0107] A training spatio-temporal feature generation module configured to generate, based on the training visual feature, a training temporal feature and a training spatial feature by an attention pooling module;

[0108] A conversion output generation module configured to generate, based on the training temporal feature, the training spatial feature, and the video description text, a conversion output by the spatio-temporal query conversion module; and

[0109] A processing model second training module configured to train the attention pooling module and the spatio-temporal query conversion module based on the conversion output.

[0110] Example 19. The apparatus of any one of examples 11-18, wherein the processing model second training module comprises:

[0111] The processing model third training module is configured to train the attention pooling module and the spatio-temporal query conversion module based on a video-text matching loss function, a video-text contrast loss function, and a video-localization text generation loss function.

[0112] Example 20. The apparatus of any of examples 11-19, the apparatus further comprising:

[0113] a target output shaping module configured to output, by a language model, a target output indicated by the training prompt text based on the conversion output and the training prompt text; and

[0114] a processing model fourth training module to train the attention pooling module and the spatio-temporal query conversion module based on the target output indicated by the training prompt text, parameters of the language model remaining unchanged in the training.

[0115] Example 21. An electronic device, comprising:

[0116] a processor; and

[0117] a memory coupled with the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform actions comprising:

[0118] determining temporal features and spatial features of a video; and

[0119] generating, using a spatio-temporal query conversion module, a text output corresponding to the video based on the temporal features and the spatial features, the spatio-temporal query conversion module comprising a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

[0120] Example 22. The electronic device of example 21, wherein the temporal features and the spatial features are generated by an attention pooling module, the method further comprising:

[0121] generating, based on the video, visual features of the video; and

[0122] generating, using a cross-attention network and a feed-forward neural network in the attention pooling module, the temporal features and the spatial features based on the visual features.

[0123] Example 23. The electronic device of any of examples 21-22, wherein the spatio-temporal query conversion module further comprises a fusion query and a corresponding fusion expert network.

[0124] Example 24. The electronic device of any of examples 21-23, the actions further comprising:

[0125] obtaining prompt text; and

[0126] generating, using the spatio-temporal query conversion module, the text output indicated by the prompt text based on the temporal feature, the spatial feature, and the prompt text.

[0127] Example 25. The electronic device of any of examples 21-24, wherein generating the text output indicated by the prompt text comprises:

[0128] generating, using the spatio-temporal query conversion module, video features based on the temporal feature and the spatial feature; and

[0129] generating the text output indicated by the prompt text based on the video features and the prompt text.

[0130] Example 26. The electronic device of any of examples 21-25, wherein the spatio-temporal query conversion module further comprises cross-attention masks, the cross-attention masks comprising temporal masks, spatial masks, and fusion masks.

[0131] Example 27. The electronic device of any of examples 21-26, wherein the text output is generated by a video processing model, the video processing model comprising the spatio-temporal query conversion module, the method further comprising:

[0132] training the video processing model based on a training video, a video description text for the training video, and a training prompt text.

[0133] Example 28. The electronic device of any of examples 21-27, wherein training the video processing model comprises:

[0134] generating, by a visual encoder, training visual features based on the training video;

[0135] generating, by an attention pooling module, training temporal features and training spatial features based on the training visual features;

[0136] generating, by the spatio-temporal query conversion module, a conversion output based on the training temporal features, the training spatial features, and the video description text; and

[0137] training the attention pooling module and the spatio-temporal query conversion module based on the conversion output.

[0138] Example 29. The electronic device of any of examples 21-28, wherein training the attention pooling module and the spatio-temporal query conversion module comprises:

[0139] The attention pooling module and the spatio-temporal query conversion module are trained based on a video text matching loss function, a video text contrast loss function, and a video localization text generation loss function.

[0140] Example 30. The electronic device of any of examples 21-29, the operations further comprising:

[0141] outputting, by a language model, a target output indicated by the training prompt text based on the conversion output and the training prompt text; and

[0142] training the attention pooling module and the spatio-temporal query conversion module based on the target output indicated by the training prompt text, parameters of the language model remaining unchanged in the training.

[0143] Example 31. A computer-readable storage medium having stored thereon one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method of any of examples 1-10.

[0144] Example 32. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any of examples 1-10.

[0145] Although the present disclosure has been described in some detail with specific reference to structure features and / or method logical steps, it is understood that the subject matter defined in the appended claims is not necessarily limited to the particular features or steps described above. The particular features and steps described above are merely examples of implementing the claims.

Claims

1. A method of video processing, comprising: determining temporal features and spatial features of a video; and generating, based on the temporal features and the spatial features, a text output corresponding to the video using a spatio-temporal query conversion module, the spatio-temporal query conversion module comprising temporal queries and corresponding temporal expert networks and spatial queries and corresponding spatial expert networks.

2. The method of claim 1, wherein the temporal features and the spatial features are generated by an attention pooling module, the method further comprising: generating, based on the video, visual features of the video; and generating, based on the visual features, the temporal features and the spatial features using a cross-attention network and a feed-forward neural network in the attention pooling module.

3. The method of claim 2, wherein the spatio-temporal query conversion module further comprises fusion queries and corresponding fusion expert networks.

4. The method of claim 1, further comprising: obtaining a prompt text; and generating, based on the temporal features, the spatial features, and the prompt text, a text output indicated by the prompt text using the spatio-temporal query conversion module.

5. The method of claim 4, wherein generating the text output indicated by the prompt text comprises: generating, based on the temporal features and the spatial features, video features using the spatio-temporal query conversion module; and generating, based on the video features and the prompt text, the text output indicated by the prompt text.

6. The method of claim 5, wherein the spatio-temporal query conversion module further comprises cross-attention masks, the cross-attention masks comprising temporal masks, spatial masks, and fusion masks.

7. The method of claim 1, wherein the text output is generated by a video processing model, the video processing model comprising the spatio-temporal query conversion module, the method further comprising: training the video processing model based on training videos, video description texts of the training videos, and training prompt texts.

8. The method of claim 7, wherein training the video processing model comprises: generating, based on the training videos, training visual features by a visual encoder; generating, based on the training visual features, training temporal features and training spatial features by an attention pooling module; generating, based on the training temporal features, the training spatial features, and the video description texts, conversion outputs by the spatio-temporal query conversion module; and training the attention pooling module and the spatio-temporal query conversion module based on the conversion outputs.

9. The method of claim 8, wherein training the attention pooling module and the spatio-temporal query conversion module comprises: training the attention pooling module and the spatio-temporal query conversion module based on a video text matching loss function, a video text contrastive loss function, and a video localization text generation loss function.

10. The method of claim 8, further comprising: outputting, based on the conversion outputs and the training prompt texts, target outputs indicated by the training prompt texts by a language model; and ​ ​ ​ ​ ​ Based on the target output indicated by the training prompt text, the attention pooling module and the spatio-temporal query conversion module are trained, and parameters of the language model remain unchanged in the training.

11. An apparatus of video processing, comprising: a spatio-temporal feature determination module configured to determine a temporal feature and a spatial feature of a video; and a text output generation module configured to generate a text output corresponding to the video based on the temporal feature and the spatial feature using a spatio-temporal query conversion module, the spatio-temporal query conversion module comprising a temporal query and a corresponding temporal expert network and a spatial query and a corresponding spatial expert network.

12. An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1-10.

13. A computer program product tangibly stored on a non-transitory computer readable medium and comprising computer executable instructions for performing the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Video multi-event clipping and text description method and device, equipment and medium

    CN111723238A

  • Video dialogue and model training method and device, equipment and storage medium

    CN117351387A

  • Video description generation method and system based on video space-time scene graph fusion reasoning

    CN117370604A

  • Video processing method and device, electronic equipment and computer program product

    CN118590707A

  • Concept-conditioned and pretrained language models based on time series to free-form text description generation

    US20240061998A1