Method, apparatus, electronic device, and computer program product for video processing
By compressing the visual markers of video image frames through gating processing, the problem of excessive computational burden in video processing is solved, and efficient fine-grained video processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-03-15
- Publication Date
- 2026-05-08
AI Technical Summary
Existing video processing technologies are computationally burdensome when processing at a fine-grained level, making it difficult to effectively extract useful information from image frames and resulting in low processing efficiency.
By using gating processing, visual markers for each image frame of the video are compressed, retaining highly relevant and useful visual markers, reducing the number of visual markers, lowering computational resource consumption, and achieving a balance between the number of visual markers and the processing effect.
It effectively reduces the consumption of computing resources while improving the effect of video processing, enabling the model to process a large number of video image frames and improving the efficiency of fine-grained video processing.
Smart Images

Figure CN118214918B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to methods, apparatus, electronic devices, and computer program products for video processing. Background Technology
[0002] With the rapid development of digitalization, video processing tasks have become increasingly crucial. Whether in the business world or in personal life, various video content processing is involved. Existing video processing tasks are divided into three categories based on processing granularity: the first is video-level tasks, which mainly focus on capturing global information within the video; the second is frame-level tasks, which mainly identify and analyze image frames in the video, emphasizing temporal awareness between image frames; and the third is object-level tasks, which require models to locate objects in each image frame and effectively distinguish and track objects over time.
[0003] With the continuous advancement of video processing technology, the importance of processing video content with finer granularity is becoming increasingly prominent. As technology develops and application scenarios diversify, more detailed video processing is becoming increasingly crucial and has become a significant trend in the video technology field, playing a vital role in promoting the development and innovation of video processing technology. Summary of the Invention
[0004] Embodiments of this disclosure provide a method, apparatus, electronic device, computer program product, and medium for video processing.
[0005] According to a first aspect of this disclosure, a method for video processing is provided. The method includes acquiring a video and a prompt text. Furthermore, the method includes generating an output indicating the prompt text based on the video and the prompt text, wherein multiple visual markers for each image frame of the video are compressed through gating processing.
[0006] According to a second aspect of this disclosure, a video processing apparatus is provided. The apparatus includes a text video acquisition module configured to acquire video and prompt text. Furthermore, the apparatus includes an indication output generation module configured to generate an output indicating the prompt text based on the video and the prompt text, wherein multiple visual markers of each image frame of the video are compressed through gating processing.
[0007] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.
[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer program product is tangibly stored on the non-transitory computer-readable medium and includes computer-executable instructions that, when executed, cause a computer to perform the steps of the method of the first aspect of this disclosure.
[0009] In a fifth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to the first aspect.
[0010] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0012] Figure 1 A schematic diagram of an example environment in which the apparatus and / or methods according to embodiments of the present disclosure may be implemented is shown;
[0013] Figure 2 A flowchart of a video processing method according to an embodiment of the present disclosure is shown;
[0014] Figure 3A A schematic diagram illustrating the process of performing a video processing task according to an embodiment of the present disclosure is shown;
[0015] Figure 3B A schematic diagram illustrating an implementation of a compression module according to an embodiment of the present disclosure is shown;
[0016] Figure 3C A schematic diagram showing prompt text and corresponding visual output according to embodiments of the present disclosure is provided.
[0017] Figure 4 A schematic diagram illustrating the process of constructing training data for an object-level video processing task according to an embodiment of the present disclosure is shown.
[0018] Figure 5 A schematic diagram illustrating the model training process according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A block diagram of a video processing apparatus according to some embodiments of the present disclosure is shown;
[0020] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0021] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation
[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the type, scope of use, and usage scenarios of the personal information (such as voice) involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.
[0025] As mentioned earlier, fine-grained video processing is becoming increasingly important in the field of video processing. However, the performance of related fine-grained video processing techniques is unsatisfactory. This is because fine-grained video processing often requires processing a large number of image frames, which creates a huge computational burden. Furthermore, video image frames contain a large amount of redundant information, making it difficult to extract the effective information from the image frames. To address this, embodiments of this disclosure provide a video processing scheme. The method first receives video and prompt text, and then generates an output indicating the prompt text based on the video and prompt text. The visual tokens of each image frame in the video are compressed through gating processing. Thus, the scheme according to embodiments of this disclosure can compress the visual tokens of image frames in the video through gating processing, retaining visual tokens with high relevance and providing effective information, thereby reducing the number of visual tokens, reducing computational resource consumption, and achieving a balance between the number of visual tokens and video processing performance. This allows the model to process a large number of video image frames while improving the video processing effect.
[0026] Figure 1 A schematic diagram of an example environment in which the apparatus and / or methods according to embodiments of the present disclosure may be implemented is shown. Figure 1 As shown, the example environment 100 may include a computing device 110, which may be a user terminal, mobile device, computer, etc., or it may be a computing system, a single server, a distributed server, or a cloud-based server. The computing device 110 can receive video 120 and prompt text 122. For example, video 120 may include multiple image frames, and each image frame may include multiple objects (e.g., animals, people, etc.). Prompt text 122 may be "Please find the puppy and provide its coordinates in each frame," instructing the user to find the object in the video (i.e., the puppy) and provide the object's coordinates in each image frame. The object's coordinates in each frame can be represented by text, such as "First frame: [23,45,46,72]," indicating the object's coordinates in the first image frame. Accordingly, the text "First frame: [23,45,46,72]" can also be converted into a visual form, such as displaying a bounding box corresponding to the coordinates on the first frame of video 120, which can visually display the object's position. It should be understood that the above is for illustrative purposes only, and the embodiments of this disclosure do not limit the content of the prompt text and video. Furthermore, embodiments of this disclosure can process multimodal data, namely video modal data and text modal data. Multimodal data refers to a collection of data containing multiple types or forms, which may come from different sensors, devices, or sources, and typically include at least two of the following forms: text, images, audio, video, etc.
[0027] The computing device 100 may include a video processing system 130, which can acquire an image frame set 132 from a video 120. The video may consist of a series of consecutive still image frames, which, when played continuously at a certain frame rate, create a continuous dynamic effect, thus presenting the video effect. For example, the image frame set 132 may include 16 or 32 image frames, or other numbers of image frames. In some embodiments, a predetermined number (e.g., 16) of image frame sets 132 can be obtained by uniformly sampling the video 120. The image frame set 132 can be converted into a visual tag set 134 by a visual encoder, where the visual tags in the visual tag set 134 can also be referred to as features of the image frames. For example, the image frame set 132 may include 16 image frames, each of which can be converted into multiple visual tags (e.g., 100 visual tags), each of which can be an embedding in vector form. That is, each image frame can be converted into multiple embeddings for processing by the video processing system 130. In some embodiments, image frames can be encoded to generate multiple visual tags using a pre-trained model of contrastive language-image learning.
[0028] The video processing system 130 can compress the visual tag set 134 using the gating module 136. For example, the visual tag set 134 may include visual tags for 16 image frames, with each image frame converted into 100 visual tags, resulting in a total of 16*100 visual tags. After compression by the gating module 136, a visual tag set 138 is generated. The visual tags for each image frame in the visual tag set 138 can be compressed to 50, resulting in a total of 16*50 visual tags. The ratio of the number of tags before and after compression is called the compression ratio. In some embodiments, the gating module 136 can use gating processing to determine the scores of multiple visual tags for each image frame and can determine the visual tag set 138 based on the scores.
[0029] Continue to refer to Figure 1 The video processing system 130 can generate preprocessed text 140 from the prompt text 122. In some embodiments, the video processing system 130 can perform word segmentation on the prompt text. For example, the prompt text is "Please find the puppy and provide its coordinates in each frame," which, after word segmentation, becomes "Please find the puppy and provide its coordinates in each frame." In some embodiments, the preprocessed text 140 can be converted into text tags 142 using a bag-of-words representation. For example, the preprocessed text after word segmentation is "Please find the puppy and provide its coordinates in each frame," and the index of each word in the bag-of-words representation can be looked up to generate text tags 142. It should be understood that the above processing is for illustrative purposes only, and this disclosure does not limit the way text tags are generated.
[0030] The video processing system 130 can combine the visual marker set 138 and the text marker 142 to generate output 150 via the processing module 144. In some embodiments, the prompt text 122 may be "Please track the girl in the red dress" and the video 120 may include the object indicated in the prompt text 122 (e.g., the girl in the red dress), and the output 150 may be "First frame: [23,45,46,72], Second frame: [23,45,46,72]...". That is, the video processing system 130 can track the coordinates of the object in each frame using the object described in the prompt text 122 without needing the object's initial coordinate information for localization. Furthermore, while the output 150 in the above example is described in text form, it should be understood that the object coordinates provided can be easily converted into video for visualization.
[0031] It should be understood that the architecture and functionality in example environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure. Embodiments of this disclosure can also be applied to other environments with different structures and / or functionalities.
[0032] The following will combine Figures 2 to 7 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0033] Figure 2 A flowchart of a video processing method 200 according to an embodiment of the present disclosure is shown. At block 202, video and prompt text can be acquired. For example, refer to... Figure 1 The video processing system 130 can acquire video 120 and prompt text 122.
[0034] At box 204, outputting a cue text instruction can be generated based on the video and the cue text. Multiple visual markers for each frame of the video are compressed through gating processing. For example, refer to... Figure 1 The video processing system 130 can generate an output 150 indicated by the prompt text 122 based on the video 120 and the prompt text 122. Multiple visual markers 134 of each image frame 132 of the video 120 are compressed through gating processing, which can be performed by the gating module 126.
[0035] Therefore, the method 200 according to the embodiments of this disclosure can compress the visual tags of image frames in a video through gating processing, retain visual tags with high relevance and providing effective information, thereby reducing the number of visual tags, reducing computational resource consumption, and achieving a balance between the number of visual tags and the video processing effect, so that the model can process a large number of video image frames while improving the video processing effect.
[0036] Figure 3A A schematic diagram of a process 300 for performing a video processing task according to an embodiment of the present disclosure is shown. As shown in FIG3, a video processing model 302 (e.g., as...) Figure 1The video processing model 130 shown can receive video 304 and extract multiple image frames from video 304, such as image frame 306, image frame 308, image frame 310, etc. Image frame 306 can be the first frame of video 304, image frame 308 can be the second frame of video 304, and image frame 306 can be the last frame of video 304. For ease of illustration, intermediate frames of the video are not shown; it should be understood that the video also includes several intermediate frames. As mentioned above, video processing tasks can be divided into video-level tasks, frame-level tasks, and object-level tasks. In some embodiments, for video-level tasks, video 304 can be uniformly sampled to obtain a predetermined number (e.g., 16) of image frames. In some embodiments, for object-level tasks, video 304 can be divided into multiple segments, each segment including a predetermined number (e.g., 8) of image frames, and the last frame of the previous segment and the first frame of the subsequent segment are the same overlapping frame. This allows the object position of the subsequent segment to be initialized by utilizing the object position of the last frame in the previous segment.
[0037] Continue to refer to Figure 3A Image frames can be input into the visual encoder 312. For example, the visual encoder 312 can process image frame 306 and generate a set of visual tags for image frame 304, which can be referred to as the visual features of image frame 306. Let v represent an image frame, where v represents video and i represents the i-th frame. Then, the visual features generated by the visual encoder 312 can be expressed by formula (1):
[0038]
[0039] in The visual features of the image frame are represented by the visual tag set, where N represents the number of visual tags in the visual tag set, and C represents the embedding dimension of each visual tag. The visual tag set can be compressed using the compression module 314, and the compression process can be represented by formula (2):
[0040]
[0041] Among them, T v Let be the compressed set of visual tags, α∈(0,1] be the compression ratio, and D be the embedding dimension of the compressed visual tags. See below for reference. Figure 3B This describes the implementation of the compression module 314.
[0042] Figure 3B A schematic diagram of an implementation 300B of a compression module according to an embodiment of the present disclosure is shown. (In conjunction with...) Figure 3A As shown, the visual marker set 342 can be Figure 3AThe output of the visual encoder 312. For example, the visual tag set 342 can be the output obtained by the visual encoder 312 processing image frame 306, that is, the image features of image frame 306. For illustrative purposes, Figure 3B The illustrated visual marker set 342 includes six visual markers. Embodiments of this disclosure do not limit the number of visual markers in the visual marker set; it may have more or fewer visual markers. For example... Figure 3B As shown, the visual tag set 342 can be input into the compression module 314, which may include a gating module 344 and an alignment module 346. The gating module 344 may include a multilayer perceptron (MLP) 348 and a SoftMax layer 350. The gating process performed by the gating module 344 can be expressed by formula (3):
[0043] G v =KeepTopK(Softmax(MLP(F v )), k, F v (3)
[0044] Where k = αN, representing the number of compressed visual tags G. v The compressed visual tags are represented. The gating process, denoted as KeepTopK, selects the k highest-scoring (e.g., 2) visual tags from the visual tag set 342 to compress the number of visual tags. For example, the gating process can generate a score for each visual tag in the visual tag set 342. In some embodiments, the k highest-scoring visual tags can be selected from the visual tag set based on their scores. The gating operation selection strategy can reduce the number of visual tags while retaining the most relevant and informative visual tags based on their scores. The alignment module 346 can adjust the embedding dimension of the visual tags to align with the processing dimension of subsequent modules. The alignment module 346 may include an MLP 348 and can be represented by Equation (4):
[0045] T v =MLP(G v (4)
[0046] Where T v For the visual tag set 350 generated by the alignment module 346, which includes two visual tags, the number of visual tags is compressed and the dimensions are aligned. It should be understood that in some implementations, if the embedding dimensions of the visual tags do not need to be adjusted, the compression module may only include the gating module and not the alignment module.
[0047] Return to reference Figure 3AThe compression module 314 can generate compressed visual markers 316, 318, and 320 for each image frame. For example, visual marker set 316 corresponds to image frame 306 and may include two visual markers, visual marker set 318 corresponds to image frame 308, and visual marker set 310 corresponds to image frame 320. The time stamps of the image frames can then be input into the timestamp module 322, which can add time-aware information to the visual markers of each image frame. For example, the text "Frame 1" can be added to visual marker set 316 to indicate that it is the visual marker for the first image frame. Similarly, the text "Frame 2" can be added to visual marker set 318, and the text "Frame N" can be added to visual marker set 320. As shown, the combination of the above time stamps and visual markers can be called video marker 330, which represents the video features of video 304. By adding time stamps, the video processing model 302 can distinguish consecutive frames of video 304.
[0048] Furthermore, the video processing model 302 can also receive prompt text 332 and preprocess the prompt text through the preprocessing module 334. For example, the preprocessing module 334 can perform word segmentation on the prompt text 332 and look up the index of each word using a bag-of-words model to generate text tags 336. Then, the text tags 336 and video tags 330 can be input into the processing module 338 to generate output 340. See below for reference. Figure 3C This section describes an example implementation of the prompt text and output.
[0049] Figure 3C A schematic diagram 300C illustrates a prompt text and corresponding visual output according to an embodiment of the present disclosure. In some embodiments, the prompt text may instruct a video processing model to perform a Referring Single Object Tracking (RSOT) task, i.e., to track a single object specified in the prompt text in each frame of the video. For example, the prompt text 360 is “Please find the hat on the puppy’s head and provide the detailed coordinates of the hat in each frame,” and the output (e.g., Figure 3A Output 340 in the output can be "First frame: [23,45,46,72], Second frame: [23,45,46,72]...". That is, the video processing model can track the position of an object in each frame using the object described in the prompt text 360, without needing the object's initial coordinate information for localization. Accordingly, a visual bounding box of the object in each frame can be generated based on the object's position coordinates provided in the output; box 362 shows the bounding box of the object in the first frame.
[0050] In some embodiments, the prompt text can instruct the video processing model to perform a Single Object Tracking (SOT) task. Compared to the RSOT task, which requires providing the object's position information (i.e., object coordinates) in the first frame, RSOT can directly indicate the object through the prompt text, without requiring the object's coordinate information. For example, prompt text 370 could read, "This video displays an object at coordinates [34,40,51,67] in the first frame. Please provide the detailed coordinates of this object in each frame." Box 372 shows the object's position in the first frame, and the corresponding output could be "First frame: [38,42,53,69], Second frame: [40,46,57,72]...". Accordingly, based on the object's position coordinates provided in the output for each frame, a visual bounding box of the object in each frame can be generated, as shown in box 374, which shows the object's bounding box in the second frame.
[0051] In some embodiments, the prompt text may instruct the video processing model to perform a Video Referring Expression Generation (Video-REG) task, which predicts a description of an object given its coordinates in any frame of a video. Since objects in the current frame may be affected by occlusion or motion blur, but can be identified in other frames, the model needs to consider the temporal context to generate an accurate description when performing the Video-REG task. For example, prompt text 380 could be "This video shows an object at coordinates [13,61,47,99] in the first frame. Please describe the object," box 382 shows the object's position in the first frame, and the corresponding output 384 could be "excavator," thus generating a description of the object at the coordinates indicated by prompt text 380.
[0052] Figure 4 A schematic diagram illustrating a process 400 for constructing training data for an object-level video processing task according to an embodiment of the present disclosure is shown. Figure 4 As shown, video 402 (shown as image frames) may include five image frames, and video description 404 can be parsed to generate text blocks 406. For example, video description 404 may be "a girl with a mobile phone and a computer," which can be parsed to generate text blocks "girl," "mobile phone," and "computer." In some embodiments, to reduce data noise, text blocks including virtual words (e.g., time, wind) may be removed. The predicted locations of the objects corresponding to the text blocks can then be generated in multiple specified image frames. For example, the first frame, middle frames, and last frame, along with text block 406, can be input into a pre-trained localization model 408 to generate video 410. Figure 4As shown, the objects corresponding to the text blocks in the first, middle, and last frames of video 410 all have bounding boxes with predicted locations. In some embodiments, data with confidence scores less than a confidence threshold (e.g., 0.6) can be filtered based on the confidence score of the predicted locations of the localization model 408.
[0053] Then, the predicted position of the object in the first frame of video 410 can be used as a template input into the pre-trained tracking model 470, and the remaining frames can also be input into the localization model, thereby generating the predicted position of the object in each frame of video 414. Since the predicted position of the object in each frame is generated, the trajectory of the object in the video is obtained, such as... Figure 4 As shown. In some embodiments, data with a trajectory confidence greater than a confidence threshold (e.g., 0.8) can be retained based on the prediction confidence of the tracking model 412. In some embodiments, a Kalman filter can be used to filter data with drift trajectories. In some embodiments, the intersection-over-union ratio (IoU) between the bounding boxes and tracking boxes in intermediate and last frames can be calculated to filter data with an IoU threshold.
[0054] Therefore, according to process 400 of an embodiment of the present disclosure, a large amount of labeled object-level training data can be generated to fully train the video processing model according to an embodiment of the present disclosure, enabling the model to perceive objects across multiple image frames and understand inter-frame relationships, thereby improving the effect of object-level video processing.
[0055] Figure 5 A schematic diagram of a model training process 500 according to an embodiment of the present disclosure is shown. Figure 5 As shown, at box 502, a training dataset can be prepared. In some embodiments, the training dataset can be divided into three categories: (1) image data; (2) video-level task data; and (3) object-level task data. As described below, in order to fully train the model, embodiments of this disclosure employ a progressive training strategy, and the entire training process includes a pre-training phase and a fine-tuning phase. At box 504, the compression module (e.g., compression module 314 shown in FIG. 3) can be pre-trained using image data. For example, when pre-training the compression module, the parameters of other modules of the video processing model (e.g., video encoder 312 and processing module 338 shown in FIG. 3) remain unchanged. By pre-training the compression module, the parameters of the compression module can be initialized, improving the stability of subsequent training. At box 506, the video processing model can be pre-trained using image data. Since image data generates fewer visual labels, pre-training using image data can speed up the pre-training process.
[0056] At box 508, high-quality data can be used to fine-tune the video processing model. For example, high-quality data can be used... Figure 4 The object-level training data generated by process 400 is used to fine-tune the video processing model. In some embodiments, the first fine-tuning phase (e.g., the first 20,000 steps) may employ a random sampling method, randomly selecting 2 to 8 frames from each video at random intervals ranging from 1 to 60. This random sampling process helps simulate different frame rates and motion speeds. In the second fine-tuning phase (e.g., the subsequent 20,000 steps), the number of frames per video may be increased (e.g., 32).
[0057] Therefore, according to process 500 of the embodiments of this disclosure, a progressive training strategy is adopted to perform multi-stage pre-training and fine-tuning of the video processing model, thereby improving the training effect of the model. Furthermore, since the gating process performed by the compression module can retain visual tags with high relevance and providing effective information, the number of visual tags is effectively reduced, thus reducing the training burden on the model and promoting sufficient training of the video processing model, thereby improving the training effect of the video processing model.
[0058] Figure 6 A block diagram of a video processing apparatus 600 according to some embodiments of the present disclosure is shown. Figure 6 As shown, the device 600 includes a text-video acquisition module 602, configured to acquire video and prompt text. Furthermore, the device 600 also includes an indication output generation module 604, configured to generate an output indicating the prompt text based on the video and the prompt text, wherein multiple visual markers of each image frame of the video are compressed through gating processing.
[0059] Figure 7 A block diagram of an electronic device 700 according to certain embodiments of the present disclosure is shown. Figure 7 A block diagram of an electronic device 700 according to certain embodiments of the present disclosure is shown. Device 700 may be the device or apparatus described in the embodiments of the present disclosure. Figure 7 As shown, device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. Various programs and data required for the operation of device 700 can also be stored in RAM 703. The CPU / GPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704. Although not shown in... Figure 7 As shown, device 700 may also include a coprocessor.
[0060] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0061] The various methods or processes described above can be executed by CPU / GPU 701. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU / GPU 701, one or more steps or actions in the methods or processes described above can be performed.
[0062] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.
[0063] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0064] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0065] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0066] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0067] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0069] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0070] The following are some example implementations of this disclosure.
[0071] Example 1. A video processing method, comprising:
[0072] Get the video and prompt text; and
[0073] Based on the video and the prompt text, an output indicating the prompt text is generated, wherein multiple visual markers of each image frame of the video are compressed through gating processing.
[0074] Example 2. According to the method of Example 1, wherein generating the output indicating the prompt text includes:
[0075] Based on the prompt text, generate multiple text tags; and
[0076] Based on each frame of the video, a first number of visual tags corresponding to the frame are generated;
[0077] The first number of visual tags is compressed through the gating process to generate a second number of visual tags, the second number being less than the first number; and
[0078] The output indicating the prompt text is generated based on the plurality of text tags and the second number of visual tags.
[0079] Example 3. The method according to any one of Examples 1-2 further includes:
[0080] Add the time stamp of the frame image to the second number of visual stamps.
[0081] Example 4. The method according to any one of Examples 1-3, wherein generating the second number of visual markers includes:
[0082] Through the gating process, a visual tag score is generated for the first number of visual tags; and
[0083] Based on the visual marker scores, a second number of visual markers are selected from the first number of visual markers.
[0084] Example 5. The method according to any one of Examples 1-4, wherein the prompt text indicates at least one of the following:
[0085] Generate the coordinates of the objects in the video; and
[0086] Generate a description of the object at a given coordinate in the video.
[0087] Example 6. The method according to any one of Examples 1-5, wherein the output is generated via an object processing model, and the method further includes:
[0088] Obtain the first training set, which includes image data;
[0089] Based on the first training set, the compression module in the video processing model used for the compression processing is pre-trained; and
[0090] The video processing model is pre-trained based on the first training set.
[0091] Example 7. The method according to any one of Examples 1-6 further includes:
[0092] Obtain a second training set that includes video-level data;
[0093] Generate a third training set that includes object-level data;
[0094] The model parameters of the video processing model are adjusted based on the second training set and the third training set.
[0095] Example 8. The method according to any one of Examples 1-7, wherein generating the third training set including the object-level data comprises:
[0096] Obtain the training video and the corresponding video description;
[0097] Based on the training video and the video description, generate predicted positions of objects in multiple specified image frames of the training video; and
[0098] Based on the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame of the training video is generated.
[0099] Example 9. The method according to any one of Examples 1-8, wherein generating the predicted position of the object in the plurality of specified image frames of the training video comprises:
[0100] Based on the video description, generate a text block corresponding to the object; and
[0101] Based on the text block and the training video, the predicted position of the object in the plurality of specified image frames of the training video is determined using a localization model.
[0102] Example 10. The method according to any one of Examples 1-9, wherein generating the predicted location of the object in each image frame of the training video comprises:
[0103] Based on the predicted position of the object in the first image frame of the training video, a location template is generated; and
[0104] Based on the location template and the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame is generated using a tracking model.
[0105] Example 11. A video processing apparatus, comprising:
[0106] The text-to-video acquisition module is configured to acquire video and prompt text; and
[0107] An instruction output generation module is configured to generate an output indicating the prompt text based on the video and the prompt text, wherein multiple visual markers of each image frame of the video are compressed through gating processing.
[0108] Example 12. The apparatus according to Example 11, wherein the indication output generation module comprises:
[0109] The text marker generation module is configured to generate multiple text markers based on the prompt text; and
[0110] The first visual marker generation module is configured to generate a first number of visual markers corresponding to each frame image of the video.
[0111] The second visual marker generation module is configured to compress the first number of visual markers through the gating process to generate a second number of visual markers, the second number being less than the first number; and
[0112] The second instruction output generation module is configured to generate the output of the prompt text instruction based on the plurality of text marks and the second number of visual marks.
[0113] Example 13. The apparatus according to any one of Examples 1-12 further includes:
[0114] A time stamp adding module is configured to add time stamps of the frame image to the second number of visual stamps.
[0115] Example 14. The apparatus according to any one of Examples 1-13, wherein the second visual marker generation module comprises:
[0116] A tag score generation module is configured to generate visual tag scores for the first number of visual tags through the gating process; and
[0117] The visual marker selection module is configured to select a second number of visual markers from a first number of visual markers based on the visual marker scores.
[0118] Example 15. The apparatus according to any one of Examples 1-14, wherein the prompt text indicates at least one of the following:
[0119] Generate the coordinates of objects in the video; and
[0120] Generate a description of the object at a given coordinate in the video.
[0121] Example 16. An apparatus according to any one of Examples 1-15, wherein the output is generated via an object processing model, and the apparatus further comprises:
[0122] The first training set acquisition module is configured to acquire a first training set including image data;
[0123] A compression module pre-training module is configured to pre-train a compression module for the compression processing in a video processing model based on the first training set; and
[0124] The video model pre-training module is configured to pre-train the video processing model based on the first training set.
[0125] Example 17. The apparatus according to any one of Examples 1-16 further includes:
[0126] The second training set acquisition module is configured to acquire a second training set that includes video-level data.
[0127] The third training set generation module is configured to generate a third training set that includes object-level data.
[0128] The video model adjustment module is configured to adjust the model parameters of the video processing model based on the second training set and the third training set.
[0129] Example 18. The apparatus according to any one of Examples 1-17, wherein the third training set generation module comprises:
[0130] The training data acquisition module is configured to acquire training videos and video descriptions corresponding to the training videos;
[0131] A first predicted location generation module is configured to generate predicted locations of objects in multiple specified image frames of the training video based on the training video and the video description; and
[0132] The second predicted position generation module is configured to generate the predicted position of the object in each image frame of the training video based on the predicted position of the object in the plurality of specified image frames.
[0133] Example 19. The apparatus according to any one of Examples 1-18, wherein the first predicted position generation module comprises:
[0134] The text block generation module is configured to generate text blocks corresponding to the object based on the video description; and
[0135] The third predicted location generation module is configured to determine the predicted location of the object in the plurality of specified image frames of the training video based on the text block and the training video using a localization model.
[0136] Example 20. The apparatus according to any one of Examples 1-19, wherein the second predicted position generation module comprises:
[0137] A location template generation module is configured to generate a location template based on the predicted location of the object in a first image frame of the training video; and
[0138] The fourth predicted position generation module is configured to generate the predicted position of the object in each image frame using a tracking model based on the position template and the predicted position of the object in the plurality of specified image frames.
[0139] Example 21. An electronic device comprising:
[0140] Processor; and
[0141] A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform actions, the actions including:
[0142] Get the video and prompt text; and
[0143] Based on the video and the prompt text, an output indicating the prompt text is generated, wherein multiple visual markers of each image frame of the video are compressed through gating processing.
[0144] Example 22. An electronic device according to any one of Example 21, wherein generating the output of the prompt text indication includes:
[0145] Based on the prompt text, generate multiple text tags; and
[0146] Based on each frame of the video, a first number of visual tags corresponding to the frame are generated;
[0147] The first number of visual tags is compressed through the gating process to generate a second number of visual tags, the second number being less than the first number; and
[0148] The output indicating the prompt text is generated based on the plurality of text tags and the second number of visual tags.
[0149] Example 23. The electronic device according to any one of Examples 21-22, further comprising:
[0150] Add the time stamp of the frame image to the second number of visual stamps.
[0151] Example 24. An electronic device according to any one of Examples 21-23, wherein generating the second number of visual tags includes:
[0152] Through the gating process, a visual tag score is generated for the first number of visual tags; and
[0153] Based on the visual marker scores, a second number of visual markers are selected from the first number of visual markers.
[0154] Example 25. An electronic device according to any one of Examples 21-24, wherein the prompt text indicates at least one of the following:
[0155] Generate the coordinates of the objects in the video; and
[0156] Generate a description of the object at a given coordinate in the video.
[0157] Example 26. An electronic device according to any one of Examples 21-25, wherein the output is generated via an object processing model, and the method further includes:
[0158] Obtain the first training set, which includes image data;
[0159] Based on the first training set, the compression module in the video processing model used for the compression processing is pre-trained; and
[0160] The video processing model is pre-trained based on the first training set.
[0161] Example 27. The electronic device according to any one of Examples 21-26, further comprising:
[0162] Obtain a second training set that includes video-level data;
[0163] Generate a third training set that includes object-level data;
[0164] The model parameters of the video processing model are adjusted based on the second training set and the third training set.
[0165] Example 28. An electronic device according to any one of Examples 21-27, wherein generating the third training set including the object-level data comprises:
[0166] Obtain the training video and the corresponding video description;
[0167] Based on the training video and the video description, generate predicted positions of objects in multiple specified image frames of the training video; and
[0168] Based on the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame of the training video is generated.
[0169] Example 29. An electronic device according to any one of Examples 21-28, wherein generating the predicted position of the object in the plurality of specified image frames of the training video includes:
[0170] Based on the video description, generate a text block corresponding to the object; and
[0171] Based on the text block and the training video, the predicted position of the object in the plurality of specified image frames of the training video is determined using a localization model.
[0172] Example 30. An electronic device according to any one of Examples 21-29, wherein generating the predicted location of the object in each image frame of the training video comprises:
[0173] Based on the predicted position of the object in the first image frame of the training video, a location template is generated; and
[0174] Based on the location template and the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame is generated using a tracking model.
[0175] Example 31. A computer-readable storage medium having stored thereon one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of Examples 1 to 10.
[0176] Example 32. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of Examples 1 to 10.
[0177] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video processing method, comprising: Get the video and prompt text; as well as Based on the video and the prompt text, the output indicated by the prompt text is generated; The prompt text specifies an object in each image frame of the video, and the output indicating the prompt text includes: By gating and compressing a first number of visual markers for each image frame of the video, a second number of visual markers is generated, the second number being less than the first number; as well as The multiple text markers of the prompt text are combined with the second number of visual markers to generate the output associated with the object indicated by the prompt text.
2. The method according to claim 1, further comprising: Add the time stamp of the image frame to the second number of visual stamps.
3. The method of claim 1, wherein generating the second number of visual markers comprises: The gating process generates a visual tag score for the first number of visual tags. as well as Based on the visual marker scores, a second number of visual markers are selected from the first number of visual markers.
4. The method of claim 1, wherein the prompt text indicates at least one of the following: Generate the coordinates of the objects in the video; and Generate a description of the object at a given coordinate in the video.
5. The method of claim 1, wherein the output is generated via an object processing model, and the method further comprises: Obtain the first training set, which includes image data; Based on the first training set, the compression module in the video processing model used for the compression processing is pre-trained; as well as The video processing model is pre-trained based on the first training set.
6. The method according to claim 5, further comprising: Obtain a second training set that includes video-level data; Generate a third training set that includes object-level data; The model parameters of the video processing model are adjusted based on the second training set and the third training set.
7. The method of claim 6, wherein generating the third training set including the object-level data comprises: Obtain the training video and the corresponding video description; Based on the training video and the video description, the predicted position of the object in multiple specified image frames of the training video is generated; as well as Based on the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame of the training video is generated.
8. The method of claim 7, wherein generating the predicted position of the object in the plurality of specified image frames of the training video comprises: Based on the video description, generate a text block corresponding to the object; as well as Based on the text block and the training video, the predicted position of the object in the plurality of specified image frames of the training video is determined using a localization model.
9. The method of claim 7, wherein generating the predicted location of the object in each image frame of the training video comprises: A location template is generated based on the predicted position of the object in the first image frame of the training video; as well as Based on the location template and the predicted position of the object in the plurality of specified image frames, the predicted position of the object in each image frame is generated using a tracking model.
10. A video processing apparatus, comprising: The text-to-video acquisition module is configured to acquire video and prompt text; as well as The output generation module is configured to generate the output of the prompt text indication based on the video and the prompt text; The prompt text specifies the object in each image frame of the video, and the indication output generation module is further configured to: By gating and compressing a first number of visual markers for each image frame of the video, a second number of visual markers is generated, the second number being less than the first number; as well as The multiple text markers of the prompt text are combined with the second number of visual markers to generate the output associated with the object indicated by the prompt text.
11. An electronic device, comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 9.
12. A computer program product tangibly stored on a non-transient computer-readable medium and comprising computer-executable instructions that are executed by a processor to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Cross-modal action positioning method and system based on interactive attention guidance and correction
CN115223086A