Low-redundancy reasoning method for continuous video streams for Transformer
By combining the inference results of the previous frame and the current frame in the video stream, extracting the inference results of the region of interest and multiplexing the inference results of the area without moving objects, the problems of de-redundant inference accuracy and redundant calculation in the prior art are solved, and an efficient and low-redundant video stream inference method is realized.
Patent Information
- Application Number
- CN202410632601.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-05-21
AI Technical Summary
When implementing video data de-redundant inference within Transformer, the existing technology fails to effectively consider the computing characteristics of Transformer and its classic variants, resulting in a decrease in inference accuracy or difficulty in adapting, and ignoring semantic information leads to a large number of redundant calculations in some scenarios.
By combining the inference results of the previous frame in the video stream with the preliminary inference results of the current frame, the area containing the moving object is extracted, and the calculation characteristics of Transformer are used to multiplex the inference results of the area without the moving object, and only the area of interest is inferred.
It effectively reduces the calculation cost corresponding to the time and space redundant information in the video stream, maintains high accuracy, breaks through the bottlenecks of previous work, optimizes the video data processing process, and adapts to the deployment needs of different scenarios.
Smart Images

Figure CN118573914B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular to a low-redundancy reasoning method for continuous video streams oriented to Transformer. Background Art
[0002] There is a lot of temporal and spatial redundancy in videos. For example, there are usually only slight changes between consecutive frames, and there are also a lot of background areas that are irrelevant to the target object in a single frame. Traditional methods usually perform complete inference on each frame, which will cause a lot of redundant calculations.
[0003] Previous technical solutions for implementing redundant reasoning of video data within Transformer have two major flaws: 1) They do not take into account the computational characteristics of the special modules of Transformer and its classic variants, resulting in excessive decrease in reasoning accuracy or difficulty in adapting their methods to the classic variants of Transformer; 2) By ignoring semantic information, redundant calculations are performed at the purely arithmetic level to circumvent the computational characteristics of special modules, while ignoring the importance of semantic information for mining spatiotemporal redundant information in video data, which will result in a large amount of redundant calculations in some scenarios.
[0004] The prior art (Chinese patent with publication number CN109379625A) discloses "video processing method, device, electronic device and computer-readable medium", and provides a video stream processing method, which allows the processing of non-target areas to be performed on the CPU instead of the GPU, so that the non-target areas with relatively little change can be processed on the CPU while the target area is processed quickly, achieving a balance between processing efficiency and processing accuracy. Although this method can reduce redundant calculations on the GPU, it actually puts the redundant calculations on the CPU without essentially reducing the calculation of redundant data. In addition, because the computational characteristics of the algorithm used to process the target area are not taken into account, directly splitting the video frame for calculation may cause additional precision loss. Summary of the invention
[0005] To solve the above technical problems, the present invention provides a low-redundancy inference method for continuous video streams for Transformer, which combines the inference result of the previous frame in the video stream with the preliminary inference result of the current frame to extract the area containing the moving object, and reuses the inference results of the area without moving objects in combination with the computing characteristics of Transformer, which can effectively reduce the computational cost corresponding to the temporal and spatial redundant information in the video stream and maintain a high accuracy.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0007] A low-redundancy reasoning method for continuous video streams for Transformer, comprising: jointly determining the area containing a moving object in the current frame, i.e., the region of interest, through the reasoning result of the previous frame of the video stream and the preliminary reasoning result of the current frame; and in the subsequent reasoning process, only the region of interest of the current frame is input into the Transformer operator for reasoning to obtain the reasoning result of the region of interest of the current frame, and the reasoning result of the non-interest region obtained in the reasoning of the previous frame is combined with the reasoning result of the region of interest of the current frame to obtain the complete reasoning result of the current frame.
[0008] Furthermore, the output of the Transformer attention module can represent the importance of each region in the current frame. The part of the output of the attention module that is higher than the set threshold is taken as the region R that may contain the moving object obtained after preliminary reasoning in the current frame. s , that is, to obtain the preliminary inference result of the current frame; through target trajectory prediction, according to the inference result r of the previous frame prev Predict the area R where the object may move to in the current frame t , and R s ∪R t ∪r prev The region in the current frame containing the moving object is referred to as the region of interest.
[0009] Compared with the prior art, the beneficial technical effects of the present invention are:
[0010] (1) Breaking through the bottleneck of implementing de-redundant calculations within Transformer and its classic variants: The present invention combines the computational characteristics of the special modules of Transformer and its classic variants with the semantic information contained in the inference results of the previous frame in the video. It can perform de-redundant reasoning while reducing the loss of accuracy, reducing the amount of calculation, and breaking through the bottleneck of previous work.
[0011] (2) Optimizing the processing of video data: The present invention mines spatiotemporal redundant information and extracts regions of interest (ROIs) for reasoning, thereby significantly reducing redundant calculations in the reasoning process and improving the overall processing efficiency.
[0012] (3) High-precision video inference quality: The present invention ensures that only necessary video frame areas are processed through accurate ROI extraction and customized Transformer modules, and maintains high accuracy while reducing energy consumption and improving processing efficiency.
[0013] (4) Diversified deployment requirements: By inferring only the area of interest within the Transformer and selectively enabling customized implementations of internal special modules, the present invention can efficiently utilize computing resources in specific scenarios and achieve a balance between accuracy requirements and video memory overhead, thereby adapting to the deployment requirements of different scenarios. This allows the Transformer video inference system based on this framework to be deployed and run in a variety of environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic diagram of the reasoning framework used in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0016] The core concept of the low-redundancy inference method for continuous video streams for Transformer proposed in this invention is to extract the area containing moving objects by combining the inference results of the previous frame in the video stream with the preliminary inference results of the current frame, and reuse the inference results of the area without moving objects by combining the computing characteristics of Transformer. It can effectively reduce the computational cost corresponding to the temporal and spatial redundant information in the video stream and maintain a high accuracy.
[0017] The following is an introduction to the Transformer-oriented continuous video stream low-redundancy reasoning framework adopted by the present invention in several parts.
[0018] 1. Temporal and spatial redundant ROI extraction:
[0019] The output of the attention module inside the Transformer can represent the importance of each region in the current frame. The framework uses the part of the output of the shallow attention module that is higher than the set threshold as the region R that may contain moving objects obtained after preliminary reasoning in the current frame. s , that is, the preliminary inference result of the current frame is obtained. The framework uses the target trajectory prediction algorithm to calculate the inference result r of the previous frame. prev Predict the area R where the object may move to in the current frame t , and R s ∪R t ∪r prev The region of interest (ROI) is the area that should be inferred in the current frame and contains moving objects. In the subsequent inference process, only the input corresponding to the ROI is inferred, and the intermediate results that do not belong to the ROI in the original image obtained during the inference of the previous frame are reused. This method greatly reduces the GPU's repeated calculations on areas that do not belong to the ROI, effectively reducing the amount of calculations.
[0020] 2. Customized Transformer intermediate module:
[0021] Due to the computational characteristics of some modules in Transformer and its typical variants, only reasoning about the input corresponding to the ROI will cause additional loss in final accuracy. The framework designs customized implementations of these modules to reduce the accuracy loss caused by Transformer and its variants when only reasoning about the input corresponding to the ROI.
[0022] Specifically, special modules are mainly divided into three categories: the Attention module in the classic Transformer, the WindowAttention module in the Transformer variant, and the PatchMerging module. 1) The reason why Attention produces errors is that the number of summation items of matrix multiplication is reduced when only the input corresponding to the ROI is inferred. The framework caches the input of matrix multiplication when executing the Attention module, samples the matrix multiplication input corresponding to some non-interested regions from the cache during the next inference, and performs matrix multiplication after splicing with the input corresponding to the current region of interest, thereby increasing the number of summation items in the matrix multiplication and reducing errors. 2) WindowAttention is a variant that divides the input of Attention into multiple windows based on Attention, and then performs a smaller Attention inside each window to reduce video memory overhead and achieve acceleration. Due to the limitation of window division and matrix multiplication parallelism, when only the input corresponding to the ROI is inferred, the number of inputs belonging to the ROI contained in each window is mostly different, which makes it difficult to parallelize, and similar to the ordinary Attention in 1), there is a loss of precision in the window. Therefore, the framework divides the window into several categories according to the range of the number of inputs belonging to the ROI contained in the window, and performs sampling operations similar to the Attention calculation in 1) in each category of windows until the input size participating in the matrix multiplication calculation in all windows belonging to the category is the same. This can improve the parallelism of matrix multiplication while reducing the precision loss of WindowAttention. 3) PatchMerging divides the input into multiple square patches of the same size, splices all the inputs in the patch into a single input, and then performs matrix multiplication. This is a classic downsampling algorithm used in Transformer variants. If only part of the input in the patch belongs to the ROI, the length of the spliced data will be shortened when the data in the patch is spliced into a single data, which does not meet the requirement of the matrix multiplication sum dimension. Therefore, the framework caches the input of the PatchMerging module and adds the input corresponding to all non-regions of interest in the same patch to the splicing, so that the matrix multiplication can be performed correctly and zero loss of accuracy can be achieved.
[0023] 3. Adaptive multiplexing mechanism:
[0024] Although only reasoning about the area corresponding to the ROI can reduce the amount of calculation and memory usage, it also means a certain loss of accuracy. Although the accuracy loss can be reduced by modifying the intermediate module of Transformer, the need to cache additional data reduces the benefits of video memory. In order to adapt to the diverse computing resource constraints and accuracy requirements in various scenarios, the framework decides whether to enable customized modules and determines the sampling scheme in customized modules based on the accuracy loss and video memory benefits caused by enabling or not enabling customized modifications to the intermediate modules of Transformer and its variants, thereby achieving a trade-off between computing resources and accuracy requirements in the code execution phase.
[0025] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.
[0026] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A low-redundancy reasoning method for continuous video streams for Transformer, characterized in that: include: The area containing the moving object in the current frame, i.e., the region of interest, is determined by using the inference result of the previous frame of the video stream and the preliminary inference result of the current frame; In the subsequent reasoning process, only the region of interest of the current frame is input into the Transformer operator for reasoning to obtain the reasoning result of the region of interest of the current frame. The reasoning result of the non-region of interest obtained during the reasoning of the previous frame is combined with the reasoning result of the region of interest of the current frame to obtain the complete reasoning result of the current frame. The region containing the moving object in the current frame, i.e., the region of interest, is determined by combining the inference result of the previous frame of the video stream with the preliminary inference result of the current frame, specifically including: The output of the Transformer attention module can represent the importance of each region in the current frame. The part of the output of the attention module that is higher than the set threshold is taken as the region R that may contain the moving object obtained after preliminary reasoning in the current frame. s , that is, to obtain the preliminary inference result of the current frame; through target trajectory prediction, according to the inference result r of the previous frame prev Predict the area R where the object may move to in the current frame t , and R s ∪R t ∪r prev The region in the current frame containing the moving object is referred to as the region of interest.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and computer readable medium
CN109379625A
Video low-loss compression method for unmanned aerial vehicle crowd monitoring and monitoring system
CN116320427A
Video storage management method and system for target tracking query
CN116521934A