Video Understanding Method, System, Device and Medium Based on Instruction Condition Compression

By injecting instruction conditions at the local and global levels of the video language model and using attention mechanism to compress video tokens, the problem of difficulty in taking into account both compression rate and information loss in the prior art is solved, and efficient video comprehension capabilities are achieved.

CN119784861BActive Publication Date: 2025-05-27UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244373.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-27
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing video large language model is difficult to take into account high compression rate and low information loss during video token compression, resulting in a decline in video comprehension capabilities.

Method used

Using the compression method based on instruction conditions, instruction conditions are injected at both local and global levels, and conditional compression is performed using attention mechanisms to retain visual information related to instructions.

Benefits of technology

It realizes more efficient video token compression, while maintaining excellent video comprehension capabilities, and better completing video comprehension tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784861B_ABST
    Figure CN119784861B_ABST
Patent Text Reader

Abstract

The present invention discloses a video understanding method, system, device and medium based on instruction conditional compression, which are one-to-one corresponding schemes. In the scheme: starting from the perspective of conditional compression, the instruction content is introduced as a condition to perform targeted compression, that is, the instruction is injected at two levels of local and global mixing, and during the compression process, the visual information associated with the instruction is retained as much as possible, and irrelevant information loss is allowed to achieve conditional compression. During compression, the high compression rate and low information loss of visual features can be well taken into account, and the visual details required to complete the instruction task can be retained as much as possible, so as to better complete the video understanding task and achieve excellent video understanding capabilities while achieving more efficient compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video content understanding based on video large language models, and particularly relates to a video understanding method, system, device and medium based on instruction-conditioned compression. Background Art

[0002] A video large language model is an autoregressive text generation model with a large number of parameters, trained with a vast amount of data, and capable of understanding video content, supporting multi-modal data input including videos. Based on the understanding of video content, the video large language model can respond to the input text instructions, answer instruction questions, and has great application potential in fields such as intelligent human-computer interaction, augmented reality, autonomous driving, robotics, and intelligent healthcare, and is also an important step towards achieving general artificial intelligence.

[0003] A video large language model mainly consists of a visual encoder, a visual connector, and a large language model (LLM). Video data consists of multiple frames of images, and after visual encoding, a vast amount of visual tokens are generated, thus bringing a heavy computational burden to the training and inference of the large language model. Therefore, an efficient video token compression method is an urgently needed solution strategy for all current video large language models to balance the computational overhead. Compression often inevitably loses information details, resulting in a decline in the model's video understanding ability.

[0004] Existing compression methods include pooling, convolution, clustering in space and time, and using attention-based mechanisms such as Q-Former (a lightweight Transformer architecture designed for multi-modal models, and Transformer is a neural network based on the attention mechanism) or memory pool. However, their compression methods lack explicit guidance, so they are unconditional compressions and it is difficult to balance high compression rate and low information loss of video tokens.

[0005] In view of this, the present invention is specifically proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a video understanding method, system, device and medium based on instruction-conditioned compression. Starting from the perspective of conditional compression, instruction injection is performed at two hybrid levels, local and global, and conditional compression is achieved using the attention mechanism, realizing excellent video understanding ability while achieving more efficient compression.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] A video understanding method based on instruction-conditioned compression, implemented in combination with a video large language model, includes:

[0009] Visually encode the input video to obtain a sequence of visual features;

[0010] At the local level, divide the sequence of visual features into several groups, compress each group into a single visual feature, and calculate the local visual features through an attention mechanism in combination with the injected instruction conditions. Synthesize the local visual features corresponding to all groups as the locally compressed visual features; at the global level, perform positional embedding on the sequence of visual features, introduce learnable visual features and inject instruction conditions, and then calculate the globally compressed visual features through the attention mechanism; wherein, the instruction conditions are the text features of the text instruction;

[0011] Combine the input text instruction, as well as the locally compressed visual features and the globally compressed visual features, and output the answer result.

[0012] A video understanding system based on instruction-condition compression, in which a video large language model is configured to implement the foregoing method. The video large language model includes:

[0013] A visual encoder for visually encoding the input video to obtain a sequence of visual features;

[0014] A visual connector for, at the local level, dividing the sequence of visual features into several groups, compressing each group into a single visual feature, and calculating the local visual features through an attention mechanism in combination with the injected instruction conditions. Synthesize the local visual features corresponding to all groups as the locally compressed visual features; at the global level, perform positional embedding on the sequence of visual features, introduce learnable visual features and inject instruction conditions, and then calculate the globally compressed visual features through the attention mechanism; wherein, the instruction conditions are the text features of the text instruction;

[0015] A large language model for combining the input text instruction, as well as the locally compressed visual features and the globally compressed visual features, and outputting the answer result.

[0016] A processing device includes: one or more processors; a memory for storing one or more programs;

[0017] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0018] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0019] As can be seen from the technical solution provided by the present invention above, the instruction content is introduced as a condition in the compression stage for targeted compression, that is, the instructions are injected at two levels of local and global mixing. During the compression process, the visual information associated with the instructions is retained as much as possible, allowing irrelevant information to be lost, achieving conditional compression. When compressing, it can well balance the high compression ratio and low information loss of visual features, and can retain as many visual details required to complete the instruction task as possible, so as to better complete the video understanding task. Description of the Drawings

[0020] In order to more clearly illustrate the technical solution of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0021] Figure 1 Schematic diagram of a video understanding method based on instruction - conditional compression provided by an embodiment of the present invention;

[0022] Figure 2 Schematic diagram of the direct injection method of instruction conditions provided by an embodiment of the present invention;

[0023] Figure 3 Schematic diagram of the coarse - grained injection method of instruction conditions provided by an embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the fine - grained injection method of instruction conditions provided by an embodiment of the present invention;

[0025] Figure 5 Schematic diagram of attention in the instruction - conditional compression at the mixed level provided by an embodiment of the present invention;

[0026] Figure 6 Schematic diagram of a processing device provided by an embodiment of the present invention. Detailed Embodiments

[0027] The following combines the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0028] First, the following explanations are given for the terms that may be used in this article:

[0029] Descriptions using terms such as "comprising", "including", "containing", "having" or other similar semantics shall be construed as non-exclusive inclusion. For example, including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be construed as not only including the specifically listed certain technical feature element, but also including other technical feature elements well-known in the art that are not specifically listed.

[0030] The term "consisting of" means excluding any technical feature element that is not specifically listed. If this term is used in a claim, it will make the claim a closed type, so that it does not include technical feature elements other than the specifically listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements specifically listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.

[0031] Unless otherwise clearly specified or limited, terms such as "installed", "connected", "joined", "fixed", etc. shall be understood in a broad sense. For example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this text can be understood according to specific circumstances.

[0032] The following provides a detailed description of a video understanding method, system, device and medium based on instruction condition compression provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those of ordinary skill in the art. For those conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the instruments used in the embodiments of the present invention that are not marked with the manufacturer, they are all conventional products that can be obtained through commercial purchase.

[0033] Embodiment 1

[0034] The embodiment of the present invention provides a video understanding method based on instruction condition compression, which is implemented in combination with a video large language model, and mainly includes the following steps:

[0035] Step 1, perform visual encoding on the input video to obtain a visual feature sequence.

[0036] Step 2: At the local level, divide the visual feature sequence into several groups, compress each group into a single visual feature, and calculate the local visual features through an attention mechanism in combination with the injected instruction conditions (i.e., the text features of the text instruction). Synthesize the local visual features corresponding to all groups as the locally compressed visual features; at the global level, perform positional embedding on the visual feature sequence, introduce learnable visual features and inject instruction conditions, and then calculate the globally compressed visual features through the attention mechanism.

[0037] Step 3: Combine the input text instruction, the locally compressed visual features, and the globally compressed visual features to output the answer result. Specifically: map the locally compressed visual features and the globally compressed visual features to the embedding space of the large language model respectively, concatenate them, and then the large language model combines the concatenated embedding features with the text instruction to output the answer result.

[0038] As Figure 1 shown, the overall framework process of this method is presented; as Figure 1 shown at the bottom, the input information includes: vision and text instructions. The vision encoder is responsible for performing Step 1 above and outputting the visual feature sequence. The vision encoder is a well-pre-trained model. Step 2 is executed by the vision connector of the video large language model. The vision connector randomly initializes its parameters for training to achieve the compression of visual features and project the visual features into the embedding layer of the large language model. The present invention designs the vision connector part, injects instructions at two hybrid levels of local and global, and uses the attention mechanism (referred to as local attention and global attention respectively) to complete conditional compression. Specifically, at the local level, the visual feature sequence is divided into several groups, and according to the injected instruction conditions, each group of visual features is compressed into a single visual feature. Different from Q-Former, the grouped local attention can explicitly preserve the spatio-temporal structural visual features of the video while highlighting the relevant parts within each group. At the global level, the instruction conditions are injected into a small number of learnable visual features, and then attention processing is performed on the flattened visual features that are not grouped. Global compression pays more attention to searching for instruction-related information throughout the video, so it can be expressed with a small number of tokens as assistance; the instruction conditions input in this part are the text features extracted from the text instruction by the text encoder. Step 3 is implemented by the large language model, which is also a well-pre-trained model and is responsible for outputting the corresponding answer result, as Figure 1As shown in the right part, examples of input text instructions and answer results are provided. In this part, in addition to the above locally compressed visual features and globally compressed visual features, the large language model also needs to input text instructions. Specifically, in order for the large language model to be able to recognize the content of the text instructions, the text instructions need to be embedded and then the embedded text instructions are input into the large language model.

[0039] The above-mentioned scheme provided by the embodiment of the present invention can be loaded on a computer or a server and used for intelligent question and answering, and accurate answers are given according to the video and questions input by the user to help the user understand the video.

[0040] The above scheme provided by the embodiment of the present invention mainly has the following advantages: the existing methods usually adopt unconditional compression, that is, compression is not performed under the guidance of instruction information, which often makes it difficult to balance high compression rate and low information loss of visual features during compression. The visual details required to complete the instruction task may be ignored or suppressed, making it impossible for the large language model to complete the video understanding task of the given instruction. Starting from the perspective of conditional compression, the present invention performs instruction injection at the local and global mixed levels, and uses the attention mechanism to complete conditional compression, thereby achieving excellent video understanding capabilities while achieving more efficient compression.

[0041] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the detailed process of the method provided by the embodiment of the present invention is described in detail with specific embodiments below.

[0042] In many practical application scenarios, the visual area that the user is interested in completing the instruction may be ignored or suppressed by the unconditional compression strategy. Therefore, we explore introducing the instruction content that the user cares about as a condition in the compression stage to perform targeted compression, that is, injecting instructions at two levels of local and global mixing, retaining the visual information associated with the instruction as much as possible during the compression process, allowing irrelevant information to be lost, and realizing conditional compression. When processing complex visual information, a coarse-to-fine approach is usually adopted, first considering the relevant parts in the overall rough impression, and then looking for detailed information locally in the video. Inspired by this perception process, the present invention compresses at the local and global levels, retaining as much information related to the instruction as possible, while maintaining the spatiotemporal structure with a smaller number of features, thereby achieving a better balance between computational burden and video understanding.

[0043] 1. Local level instruction condition compression.

[0044] In order to achieve effective conditional compression while maintaining the spatiotemporal structure of visual features, the present invention performs instruction conditional compression at the local level. The visual feature sequence is denoted as ,in, is the symbol for the set of real numbers, T represents the length of the visual feature sequence, H and W represent the height and width of each visual feature (i.e., visual token), and D represents the dimension of the visual feature; at the same time, a downsampling rate parameter is given ; divide the visual feature sequence into groups, where the quantity , the quantity , the quantity , and the symbol represents rounding up.

[0045] Each group contains visual features, denoted as , where are three index symbols, , , . Average pool (downsample) each group into a single visual feature , and combine it with the injected instruction condition C (i.e., the text feature of the text instruction), and calculate the corresponding local visual feature through the attention mechanism , denoted as:

[0046] ;

[0047] where, Inj represents the injection process of the instruction condition, and Attn represents the attention mechanism; the inputs of Attn are the query, key, and value in sequence (i.e., the three parts in the square brackets). Therefore, the visual features in each group are compressed into only one visual feature, while highlighting the instruction-related parts conditionally. Concatenate the local visual features corresponding to all group visual features to obtain the locally compressed visual feature . Then, use an MLP (multi-layer perceptron) layer to map into the embedding space of the large language model.

[0048] 2. Instruction condition compression at the global level.

[0049] Although the spatio-temporal structure can be maintained at the local level, the attention is forced to focus on small sub-regions. However, each sub-region may contain information that is not equally effective for the problem, and some sub-regions may even be completely irrelevant. Therefore, the embodiments of the present invention also perform conditional compression at the global level to highlight the most relevant parts throughout the video. Specifically: Initialize a small set of learnable visual features , where represents the number of tokens, for example, 32; then inject the instruction condition C into it. Different from grouping the video frame features, in this part, for the visual feature sequence Apply three-dimensional position embeddings Pos, flatten them directly when compressing at the global level, and then calculate the globally compressed visual features through the attention mechanism, expressed as:

[0050] ;

[0051] Among them, represents the globally compressed visual features, and Flat represents the flattening operation.

[0052] Similar to the local level, an MLP layer is also used to map into the embedding space of the large language model. Finally, and the mapped results are concatenated together for further understanding by the large language model.

[0053] Considering that the training process of the learnable visual features can be implemented with reference to conventional techniques. For example, it is randomly initialized at the beginning of training and set to be gradient-updatable, so that the learnable visual features during the model training process will also be optimized to appropriate parameters under the supervision of the loss function, enabling it to have the potential ability to represent visual features. The model training process described here refers to the training process of the video large language model, which can also be implemented with reference to conventional techniques, so it will not be elaborated here.

[0054] 3. Instruction condition injection.

[0055] To achieve the guiding role of the instruction conditions, the present invention explores three different types of injection methods, which are respectively defined as direct injection, coarse-grained injection, and fine-grained injection.

[0056] (1) Direct injection.

[0057] As Figure 2 shown, the method of direct injection is used to inject the instruction condition C. The steps include: (1.1) Input the text instruction into the text encoder to obtain the pooled global text feature , and use it as the instruction condition C; (1.2) Process the global text through a linear layer to complete the injection of the instruction condition. In this way, the global text is converted into the injection output by the linear layer without any interaction with or L.

[0058] (2) Coarse-grained injection.

[0059] As Figure 3 shown, the method of coarse-grained injection is used to inject the instruction condition C. The steps include: (2.1) Input the text instruction into the text encoder to obtain the pooled global text feature , use it as the instruction condition C; (2.2) Calculate the product coefficient and the offset coefficient through linear layer regression; (2.3) At the local level, after performing layer normalization on each group of visual features, add the product coefficient and the offset coefficient to complete the injection of the instruction condition; at the global level, after performing layer normalization on the learnable visual features, add the product coefficient and the offset coefficient to complete the injection of the instruction condition.

[0060] Use A to represent uniformly or L, then the above process is expressed as:

[0061] ;

[0062] Among them, 、 The corresponding ones represent calculating the product coefficient and the offset coefficient from the global text , and LN is layer normalization processing.

[0063] (3) Fine-grained injection.

[0064] As Figure 4 shown, inject the instruction condition C in a fine-grained manner. The steps include: (4.1) Input the text instruction into the text encoder to obtain the last-layer hidden layer feature of the text encoder, which is called the fine-grained text feature , use it as the instruction condition C; (4.2) At the local level, after performing layer normalization on each group of visual features, perform an attention operation (such as cross-attention operation) with the fine-grained text to complete the injection of the instruction condition; at the global level, after performing layer normalization on the learnable visual features, perform a cross-attention operation with the fine-grained text to complete the injection of the instruction condition.

[0065] Use A to represent uniformly or L, then the above process is expressed as:

[0066] ;

[0067] Among them, Attn represents the attention mechanism.

[0068] Since direct injection directly converts the condition into a query of the attention mechanism at the local level and the global level, and the length of the query limits its application only to local-level compression because the local branch compresses each group into only one visual feature. On the contrary, coarse-grained and fine-grained injection are more flexible and can adapt to various situations because the condition is added to A.

[0069] The above Figures 2 to 4Among them, the rectangular box filled with a horizontal line on the right represents the corresponding instruction condition C, and the white rectangular box on the left represents the corresponding visual feature or L.

[0070] To verify the effectiveness of the present invention, verification and evaluation were carried out on 5 general evaluation benchmark datasets (Video MME, MV-Bench, Ego-Schema, Activity-Net, VCG-Bench). The training results of the 7B (7 billion) parameter Qwen2.5 model (an open-source large model launched by Alibaba Cloud) implemented by the present invention have all achieved advanced performance, as shown in Table 1.

[0071] Table 1: Video understanding performance on various benchmark datasets

[0072]

[0073] In Table 1, VideoMME w / o sub and VideoMME w / sub are two subsets of the benchmark dataset Video MME. The former subset consists of videos without subtitle information, and the latter subset consists of videos with time-synchronized subtitles.

[0074] In addition, the attention mechanism during the compression process was visualized, as Figure 5 shown, intuitively demonstrating that the video large language model noticed the regions related to the instructions during the compression process, thus better retaining the relevant information.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0076] Embodiment 2

[0077] The present invention also provides a video understanding system based on instruction condition compression. A video large language model is configured in this system, which is mainly used to implement the method provided in the foregoing embodiments. The video large language model includes:

[0078] A visual encoder for visually encoding the input video to obtain a visual feature sequence;

[0079] A visual connector is used to divide a sequence of visual features into several groups at the local level. Each group is compressed into a single visual feature, and local visual features are calculated through an attention mechanism in combination with injected instruction conditions. The local visual features corresponding to all groups are integrated as the locally compressed visual features. At the global level, positional embedding is performed on the sequence of visual features, and learnable visual features are introduced and instruction conditions are injected, and then the globally compressed visual features are calculated through the attention mechanism. Among them, the instruction conditions are the text features of text instructions.

[0080] A large language model is used to combine the input text instructions, as well as the locally compressed visual features and the globally compressed visual features, and output an answer result.

[0081] Considering that the technical details involved in the system have been introduced in detail in the previous embodiments, they will not be elaborated here.

[0082] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0083] Embodiment III

[0084] The present invention also provides a processing device, as Figure 6 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0085] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0086] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0087] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.

[0088] The output device can be a display terminal;

[0089] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0090] Embodiment 4

[0091] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.

[0092] In the embodiments of the present invention, the readable storage medium, as a computer-readable storage medium, may be disposed in the foregoing processing device. For example, it may be a memory in the processing device. In addition, the readable storage medium may also be various media capable of storing program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc.

[0093] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those skilled in the art.

Claims

1. A video understanding method based on instruction conditional compression, combined with a video large language model, characterized in that: include: Visually encode the input video to obtain a visual feature sequence; At the local level, the visual feature sequence is divided into several groups, each group is compressed into a visual feature, and the local visual features are calculated through the attention mechanism in combination with the injected instruction conditions, and the local visual features corresponding to all groups are integrated as the local compressed visual features, including: recording the visual feature sequence as ,in, is a real number set symbol, T represents the length of the visual feature sequence, H and W represent the height and width of each visual feature, and D represents the dimension of the visual feature; the visual feature sequence is divided into groups, each containing visual features, represented as , where the number ,quantity ,quantity , For a given downsampling rate parameter, the symbol Indicates rounding up; Average pooling as a visual feature , and combined with the injected instruction condition C, the corresponding local visual features are calculated through the attention mechanism , expressed as: ; Inj represents the injection process of instruction conditions, Attn represents the attention mechanism; the local visual features corresponding to all groups of visual features are spliced ​​to obtain the locally compressed visual features ; At the global level, the visual feature sequence is positionally embedded, and learnable visual features are introduced and injected into the instruction conditions, and then the globally compressed visual features are calculated through the attention mechanism; where the instruction conditions are the text features of the text instructions; The answer result is outputted by combining the input text instruction, the locally compressed visual features and the globally compressed visual features.

2. The video understanding method based on instruction conditional compression according to claim 1, characterized in that: At the global level, the visual feature sequence is positionally embedded, and learnable visual features are introduced and instruction conditions are injected. Then, the globally compressed visual features are calculated through the attention mechanism, which is expressed as: ; in, represents the visual feature sequence, Pos represents the position embedding, Flat represents the flattening operation, L represents the learnable visual feature, C represents the instruction condition, Inj represents the injection process of the instruction condition, Attn represents the attention mechanism, Represents the visual features after global compression.

3. A video understanding method based on instruction conditional compression according to claim 1 or 2, characterized in that: It also includes: injecting instruction condition C by direct injection, the steps include: Input the text instruction into the text encoder to obtain the pooled global text features , take it as instruction condition C; Through the linear layer, the global text Processing is performed to complete the injection of instruction conditions.

4. A video understanding method based on instruction conditional compression according to claim 1 or 2, characterized in that: It also includes: injecting instruction condition C in the following manner, the steps include: Input the text instruction into the text encoder to obtain the pooled global text features , take it as instruction condition C; The product coefficient and the offset coefficient are calculated by linear layer regression; At the local level, after the layer normalization processing is performed on the visual features of each group, the product coefficient and the offset coefficient are added to complete the injection of the instruction condition; at the global level, after the layer normalization processing is performed on the learnable visual features, the product coefficient and the offset coefficient are added to complete the injection of the instruction condition; The method of injecting instruction condition C using the above steps is called a coarse-grained method.

5. A video understanding method based on instruction conditional compression according to claim 1 or 2, characterized in that: It also includes: injecting instruction condition C in the following manner, the steps include: Input the text instruction into the text encoder to obtain the last hidden layer feature of the text encoder, which is called fine-grained text feature , take it as instruction condition C; At the local level, the visual features of each group are layer-normalized and compared with the fine-grained text Attention is calculated to complete the injection of instruction conditions; at the global level, the learnable visual features are layer-normalized and compared with the fine-grained text Perform cross-attention operations to complete the injection of instruction conditions; The method of injecting instruction condition C using the above steps is called a fine-grained method.

6. The video understanding method based on instruction conditional compression according to claim 1, characterized in that: The combined input text instruction, the locally compressed visual features and the globally compressed visual features, output answer results include: The locally compressed visual features and the globally compressed visual features are respectively mapped to the embedding space of the large language model and spliced; the large language model then combines the spliced ​​embedded features with the text instructions to output an answer result.

7. A video understanding system based on instruction conditional compression, characterized in that: The system is configured with a video language model for implementing the method described in any one of claims 1 to 6, and the video language model includes: A visual encoder is used to visually encode the input video to obtain a visual feature sequence; The visual connector is used to divide the visual feature sequence into several groups at the local level, each group is compressed into a visual feature, and the local visual features are calculated through the attention mechanism in combination with the injected instruction conditions, and the local visual features corresponding to all groups are integrated as the local compressed visual features; at the global level, the visual feature sequence is positionally embedded, and the learnable visual features are introduced and the instruction conditions are injected, and then the globally compressed visual features are calculated through the attention mechanism; wherein the instruction conditions are the text features of the text instructions; The large language model is used to combine the input text instructions, the locally compressed visual features and the globally compressed visual features to output an answer result.

8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.