Monitoring video definition improving method and system

By using spatial attention mechanism and adaptive sub-window technology in the surveillance video clarity improvement method, the problem of limited clarity improvement caused by focusing on specific areas in the prior art is solved, and a more efficient improvement of clarity of surveillance video is achieved.

CN120050396AInactive Publication Date: 2025-05-27SICHUAN YUEDONG MIRACLE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510298329.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing methods of surveillance video clarity improvement in attention mechanism query, keys, and values ​​all come from the same feature map or the same area, resulting in insufficient attention to non-critical areas and limiting the clarity improvement effect.

Method used

The spatial attention mechanism is used to calculate the heat map of each channel of the low-resolution feature map, and adaptively divide it into multiple sub-windows. Queries and keys are obtained through the offset degree of the sub-windows to realize feature fusion between windows.

Benefits of technology

Through the feature fusion between windows, the effect of super-resolution image reconstruction is improved, the clarity of the surveillance video is improved, and the problem of limited clarity improvement caused by focusing on specific areas in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050396A_ABST
    Figure CN120050396A_ABST
Patent Text Reader

Abstract

The invention relates to a monitoring video definition improving method, which comprises the following steps: segmenting a low-resolution video into video frames, and carrying out feature extraction on the video frames to obtain a low-resolution feature map; calculating a heat map of each channel of the low-resolution feature map by adopting a space attention mechanism, adaptively dividing the low-resolution feature map into a plurality of sub-windows, and for each channel in the low-resolution feature map, determining the offset degree of the sub-windows in the channel based on the sub-windows and the heat maps of the channels, obtaining an area obtaining value of the sub-window on the channel, obtaining a query and a key by utilizing the offset degree of the sub-window on the channel, and calculating attention according to the query, the key and the value on the channel; and combining the attention results on all channels into a new feature map, and obtaining a super-resolution video frame corresponding to the video frame by using the new feature map. According to the method, query and keys are obtained through offset, so that fusion of features between windows is improved, and the monitoring video definition improvement effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and specifically to a method and system for improving the clarity of surveillance videos. Background Art

[0002] Video surveillance systems are a core component of modern urban security systems and play an important role in fields such as public safety, traffic management, and judicial evidence collection. High-definition videos can not only improve the extraction accuracy of key features such as face recognition and license plate identification, but also directly affect the reliability of intelligent analysis algorithms such as abnormal behavior detection and cross-camera tracking. However, due to hardware costs, storage pressure, and transmission bandwidth limitations, actually deployed surveillance devices often use low-resolution video streams, such as those below 720p, for continuous recording. Especially in the existing market where the proportion of old devices exceeds 40%, problems such as blurriness, excessive noise, and detail loss seriously restrict the effective utilization of video data. When it is necessary to recall on-site details or analyze small targets, the insufficient quality of the original video directly leads to the failure of relevant information extraction, which makes video enhancement technology a bottleneck problem for improving the effectiveness of security systems. Super-resolution reconstruction technology restores high-frequency details from low-resolution images through algorithmic means, providing a new way to solve the quality defects of surveillance videos. Traditional interpolation methods have the defect of edge blurriness due to ignoring the image degradation model. Deep learning-based methods, especially those applying the attention mechanism to the super-resolution field, have achieved an improvement in the quality of surveillance videos through end-to-end mapping. However, in existing methods, the query, key, and value of the attention mechanism all come from the same feature map or the same region, which easily focuses the attention on the most critical places and pays insufficient attention to other places with relatively lower importance, thus easily limiting the clarity improvement effect. Summary of the Invention

[0003] In view of the above problems, in the first aspect, the present invention provides a method for improving the clarity of surveillance videos, and the method includes the following steps: Segment a low-resolution video into video frames, and perform feature extraction on the video frames to obtain low-resolution feature maps; Use a spatial attention mechanism to calculate the heat map of each channel of the low-resolution feature map, adaptively divide the low-resolution feature map into multiple sub-windows, for each channel in the low-resolution feature map, determine the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel, obtain the value of the region of the sub-window on the channel, obtain the query and key using the offset degree of the sub-window on the channel, and calculate the attention based on the query, key, and value on the channel; Merge the attention results on all channels into a new feature map, and use the new feature map to obtain the super-resolution video frame corresponding to the video frame.

[0004] Preferably, the adaptive division of the low-resolution feature map into multiple sub-windows is specifically as follows: Regard the entire low-resolution feature map as an initial root node, and obtain the predefined minimum sub-window size; Starting from the root node, regard the root node as a sub-window, calculate the variance of the pixel values within the sub-window. If the variance exceeds the predefined division threshold, divide the node into four equal-sized sub-nodes; Recursively perform the same decomposition process on each sub-node until the size of the sub-node reaches the predefined minimum sub-window size or the variance of the sub-node is lower than the division threshold.

[0005] Preferably, the calculation of the heat map for each channel of the low-resolution feature map using the spatial attention mechanism is specifically as follows: For the low-resolution feature map, obtain the maximum pooling feature map by using channel maximum pooling, and activate the maximum pooling feature map through an activation function; Calculate the spatial attention of the channel feature map, and multiply the activated feature map and the spatial attention of the channel feature map element-wise to obtain the heat map of the channel.

[0006] Preferably, the determination of the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel is specifically as follows: For each sub-window, extract the heat value region corresponding to the sub-window position from the corresponding channel heat map, and calculate the average value of the heat value region as the average heat of the sub-window in the channel; Use a non-linear mapping function to calculate the offset degree corresponding to the average heat, or obtain the maximum average heat of the neighboring sub-windows of the sub-window, and determine the offset degree based on the maximum average heat and the average heat of the sub-window.

[0007] Preferably, the obtaining of the query and key using the offset degree of the sub-window in the channel is specifically as follows: Obtain the center of the sub-window, calculate the offset distance according to the offset degree, and move the offset distance in the direction of the center of the channel feature map or the center of the neighboring sub-window with the maximum average heat of the sub-window; Using the moved center as the center and the sub-window size as the size, obtain the query and key.

[0008] Preferably, the calculation of the offset distance according to the offset degree is specifically as follows: Calculate the product of half of the height of the sub-window and the offset degree to obtain the vertical movement distance, and calculate the product of half of the width of the sub-window and the offset degree to obtain the horizontal movement distance.

[0009] Second aspect, the present invention provides a system for improving the clarity of surveillance videos, and the system includes the following modules: A feature extraction module, configured to segment a low-resolution video into video frames, and perform feature extraction on the video frames to obtain low-resolution feature maps; An attention calculation module, configured to calculate the heat map of each channel of the low-resolution feature map by using a spatial attention mechanism, adaptively divide the low-resolution feature map into multiple sub-windows, for each channel in the low-resolution feature map, determine the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel, obtain the value of the area of the sub-window on the channel, use the offset degree of the sub-window on the channel to obtain queries and keys, and calculate the attention based on the queries, keys, and values on the channel; A super-resolution reconstruction module, configured to merge the attention results on all channels into a new feature map, and use the new feature map to obtain a super-resolution video frame corresponding to the video frame.

[0010] Preferably, the step of adaptively dividing the low-resolution feature map into multiple sub-windows is specifically: Regard the entire low-resolution feature map as an initial root node, and obtain a predefined minimum sub-window size; Starting from the root node, regard the root node as a sub-window, calculate the variance of the pixel values within the sub-window, if the variance exceeds the predefined division threshold, then divide the node into four equal-sized sub-nodes; Recursively perform the same decomposition process on each sub-node until the size of the sub-node reaches the predefined minimum sub-window size or the variance of the sub-node is lower than the division threshold.

[0011] Preferably, the step of calculating the heat map of each channel of the low-resolution feature map by using a spatial attention mechanism is specifically: For the low-resolution feature map, obtain the maximum pooling feature map by using channel maximum pooling, and activate the maximum pooling feature map through an activation function; Calculate the spatial attention of the channel feature map, and multiply the activated feature map and the spatial attention of the channel feature map element by element to obtain the heat map of the channel.

[0012] Preferably, the step of determining the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel is specifically: For each sub-window, extract the heat value area corresponding to the sub-window position from the corresponding channel heat map, and calculate the average value of the heat value area as the average heat of the sub-window on the channel; Calculate the offset degree corresponding to the average heat using a non - linear mapping function, or obtain the maximum average heat of the neighborhood sub - windows of the sub - window, and determine the offset degree according to the maximum average heat and the average heat of the sub - window.

[0013] Preferably, obtaining the query and key by using the offset degree of the sub - window on the channel is specifically as follows: Obtain the center of the sub - window, calculate the offset distance according to the offset degree, and move the offset distance in the direction of the center of the channel feature map or the center of the neighborhood sub - window with the maximum average heat of the sub - window; Use the moved center as the center and the size of the sub - window as the size to obtain the query and key.

[0014] Preferably, calculating the offset distance according to the offset degree is specifically as follows: Calculate the product of half of the height of the sub - window and the offset degree to obtain the vertical movement distance, and calculate the product of half of the width of the sub - window and the offset degree to obtain the horizontal movement distance.

[0015] In the existing video clarity improvement, the attention mechanism in the encoder and decoder is prone to focus on certain specific regions, which limits the model's ability to model long - distance dependencies, and since QKV comes from the same window, it is prone to block effects. The present invention first obtains the heat map of each channel based on spatial attention and the maximum - pooling feature map of the low - resolution feature map. Then, an adaptive sub - window method is adopted, and in combination with the offset method, the query (Q) and key (K) are obtained, realizing the fusion of features between windows, thereby improving the effect of super - resolution image reconstruction and enhancing the clarity of surveillance videos. Description of the Drawings

[0016] Figure 1 Is the flowchart of Embodiment 1; Figure 2 Is the schematic diagram of the sub - window; Figure 3 Is the schematic diagram of the sub - window offset; Figure 4 The clarity improvement framework adopting the encoder - decoder structure; Figure 5 Is the comparison diagram of clarity improvement effects; Figure 6 Is the structure diagram of Embodiment 2. Detailed Embodiments

[0017] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner for easy understanding.

[0018] It can be understood that the "embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It can be understood that in the various embodiments of the present application, the magnitudes of the serial numbers of the various processes do not mean the sequence of execution, and the execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0019] In the present invention, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of the present invention, as well as in each implementation manner / implementation method / realization method in each embodiment, if there is no special specification and logical conflict, the terms and / or descriptions between different embodiments, as well as between each implementation manner / implementation method / realization method in each embodiment, are consistent and can be referred to each other, and the technical features in different embodiments, as well as in each implementation manner / implementation method / realization method in each embodiment, can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their internal logical relationships. The implementation manners of the present application described below do not constitute a limitation to the protection scope of the present application.

[0020] Embodiment 1 provides a method for improving the clarity of a surveillance video, as Figure 1 shown, the method includes the following steps: S1, splitting the low-resolution video into video frames, and performing feature extraction on the video frames to obtain low-resolution feature maps; The surveillance video is composed of multiple video frames. When improving the clarity, each video frame needs to be reconstructed. The low-resolution video is divided into multiple video frames, and the splitting process separates the video stream in the video frame, for example, using OpenCV, etc. Feature extraction is performed on each video frame. Preferably, an encoder is used for feature extraction, and the encoder includes, but is not limited to, an encoder based on a convolutional neural network and an encoder based on Transformer. After feature extraction, the obtained features exist in the form of feature maps, and each channel of the feature map represents a specific feature. For example, one channel may represent the edge information of the image, and another channel may represent the texture information of the image.

[0021] S2. Calculate the heat map of each channel of the low-resolution feature map using a spatial attention mechanism, adaptively divide the low-resolution feature map into multiple sub-windows. For each channel in the low-resolution feature map, determine the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel, obtain the value of the region of the sub-window on the channel, use the offset degree of the sub-window on the channel to obtain the query and key, and calculate the attention according to the query, key, and value on the channel; Generate a single-channel heat map for each channel of the low-resolution feature map using a 1x1 convolution or a fully connected layer, and then apply Softmax or Sigmoid normalization to obtain the heat map of the channel. In one embodiment, the step of adaptively dividing the low-resolution feature map into multiple sub-windows is to divide the feature map into non-overlapping sub-windows of equal size. Specifically, the feature map is evenly divided into sub-windows of the same size and non-overlapping along the height and width of the channel feature map. For example, if the size of the channel feature map is 512×512 and the size of the sub-window is 64×64, then 64 sub-windows can be obtained. Figure 2 Fig. shows a partitioning method with 9 sub-windows.

[0022] Since using sub-windows of a fixed size cannot adapt to the detail differences in different regions of the image, in another embodiment, the step of adaptively dividing the low-resolution feature map into multiple sub-windows is specifically as follows: Regard the entire low-resolution feature map as an initial root node, and obtain a predefined minimum sub-window size; Starting from the root node, regard the root node as a sub-window, calculate the variance of the pixel values within the sub-window. If the variance exceeds the predefined partitioning threshold, divide the node into four equal-sized child nodes; Recursively perform the same decomposition process on each child node until the size of the child node reaches the predefined minimum sub-window size or the variance of the child node is lower than the partitioning threshold.

[0023] In this embodiment, in regions with rich details, smaller sub-windows are used, and in smooth regions, larger sub-windows are used to reduce redundant calculations. Specifically, the entire low-resolution feature map is used as a large initial block. Then, according to the variance, it is judged whether the feature map in this large block is complex. If it is complex, this large block is divided into four small blocks. Then, repeat this calculation and segmentation process for each small block until the feature variance in the small block is very small or reaches the minimum size. In this way, the complex regions in the feature map will be divided into smaller blocks, while the simple regions in the feature map will retain larger blocks.

[0024] Calculate the average heat value within the sub-window, generate an offset through a learnable linear layer, take the region of the sub-window on the channel feature map as value V, and use the offset to determine the sub-windows corresponding to query Q and key K. Take the sub-window regions corresponding to query Q and key K on the channel feature map as query Q and key K. Here, value V, query Q, and key K are all key-values and queries in the attention mechanism, that is, QKV in the attention mechanism. By directly obtaining the sub-window region as V, the original feature information within this region can be maximally retained, avoiding information loss caused by additional transformations or compressions; moreover, it can better capture the local details and textures of the image, enhance the feature expression ability, and thus improve the performance of tasks such as super-resolution reconstruction. In addition, obtaining Q and K by combining the sub-window with the offset can focus on the local regions of the image and capture the spatial relationships between different local regions according to the offset of the window.

[0025] After obtaining QKV, calculate the attention of each word window using the attention calculation method, and merge the attentions of all sub-windows into the attention result of the channel according to the position of the value sub-window on the channel feature map. Of course, the attention result of each channel can also be calculated using multi-head attention.

[0026] In one embodiment, calculating the heat map of each channel of the low-resolution feature map using the spatial attention mechanism is specifically as follows: For the low-resolution feature map, obtain the maximum pooling feature map by using channel maximum pooling, and activate the maximum pooling feature map through an activation function; Calculate the spatial attention of the channel feature map, and multiply the activated feature map and the spatial attention of the channel feature map element-wise to obtain the channel heat map.

[0027] For the low-resolution feature map, for example, the shape of the low-resolution feature map is C×H×W, where C is the number of channels, H is the height, and W is the width. Perform maximum pooling operation along the channel dimension C, that is, for each spatial position (h, w), select the maximum value at this position from all C channels to obtain a maximum pooling feature map with a shape of 1×H×W, which retains the strongest activation information spatially. Input the maximum pooling feature map into an activation function such as Sigmoid, ReLU, etc. For each channel's feature map, calculate using the spatial attention mechanism. The spatial attention includes calculating spatial attention using convolution and / or using the spatial attention in CBAM. For each channel, the result obtained by multiplying the activated feature map (1×H×W) and the spatial attention map (1×H×W) of this channel element-wise is the heat map of this channel.

[0028] In yet another embodiment, determining the offset degree of the sub-window in the channel based on the heat map of the sub-window and the channel is specifically as follows: For each sub-window, extract the heat value region corresponding to the position of the sub-window from the corresponding channel heat map, and calculate the average value of the heat value region as the average heat of the sub-window on the channel. Use a non-linear mapping function to calculate the degree of offset corresponding to the average heat, or obtain the maximum average heat of the neighboring sub-windows of the sub-window, and determine the degree of offset based on the maximum average heat and the average heat of the sub-window.

[0029] For each channel of the low-resolution feature map, divide it into multiple non-overlapping sub-windows of the same size. For each sub-window, extract the heat value region corresponding to the position of the sub-window from the corresponding channel heat map, and calculate the average value of the extracted heat value region. Use a predefined non-linear mapping function to map the average heat value of the sub-window to the degree of offset. Or, for each sub-window, define its neighborhood, such as the four adjacent sub-windows above, below, left, and right, calculate the average heat of all sub-windows in the neighborhood, and select the average heat of the sub-window with the maximum average heat in the neighborhood as the maximum average heat of the neighborhood of the sub-window. Determine the degree of offset based on the maximum average heat and the average heat of the sub-window. In one embodiment, calculate the difference between the maximum average heat of the neighborhood and the average heat of the sub-window, and determine the degree of offset based on the difference. For example, the larger the difference, the greater the degree of offset. Preferably, input the difference into the mapping function to obtain the degree of offset.

[0030] In one embodiment, obtaining the query and key by using the degree of offset of the sub-window on the channel is specifically as follows: Obtain the center of the sub-window, calculate the offset distance according to the degree of offset, and move the offset distance in the direction of the center of the channel feature map or the center of the neighboring sub-window with the maximum average heat of the sub-window according to the offset distance. Take the moved center as the center and the size of the sub-window as the size to obtain the query and key.

[0031] Determine the center coordinates of each sub-window. Specifically, calculate by adding half of its width and height to the upper left corner coordinates of the sub-window. For example, if the upper left corner coordinates of the sub-window are (x, y), the width is w, and the height is h, then the center coordinates are (x + w / 2, y + h / 2). According to the degree of offset of the sub-window on the channel calculated previously, convert it into an actual offset distance. In one embodiment, the degree of offset is a relative value, such as a ratio between 0 and 1. In another embodiment, the degree of offset is an absolute value, such as a pixel offset. If the degree of offset is a relative value, it needs to be multiplied by the width or height of the sub-window to obtain the actual offset distance. The offset distance includes offsets in both the horizontal and vertical directions.

[0032] Since the sub - windows do not overlap, after moving towards the center or the high - heat direction, the sub - windows will partially overlap, which increases the feature fusion between different sub - windows. In one embodiment, according to the calculated offset distance, the center coordinates of the sub - window are moved, either towards the center of the channel feature map or towards the center of the neighboring sub - window with the maximum average heat of the sub - window. The moved center coordinates are the new center coordinates. Taking the moved center coordinates as the center and the original size of the sub - window as the size, a new region is re - defined, such as Figure 3 shown, and the features within this region are extracted from the feature map as the query (Q) and key (K).

[0033] After obtaining the attention of each sub - window, if a pixel point in the low - resolution feature map is located in two or more sub - windows simultaneously, the average value of the attention of all sub - windows is used as the value of the attention calculated for this pixel point. For example, the point at position (15, 35) in the low - resolution feature map is located in sub - windows 3 and 6. The attention of this point will be calculated in both sub - window 3 and sub - window 6, and the average value of the attention values calculated for this point in window 3 and window 6 is used as the attention of this point. If a point in the low - resolution feature map is only located in one sub - window, the attention value of this point in the sub - window to which it belongs is directly used as the attention calculation result.

[0034] S3, merge the attention results on all channels into a new feature map, and use the new feature map to obtain the super - resolution video frame corresponding to the video frame.

[0035] After obtaining the attention output of each channel, these outputs are merged into a unified feature map for subsequent super - resolution reconstruction. In one embodiment, the merging operation is to concatenate the attention outputs of all channels in the channel dimension. Suppose there are C channels, and the attention output shape of each channel is H×W, then the shape of the merged feature map will be C×H×W. After obtaining the new feature map, up - sampling or a super - resolution reconstruction model such as SRCNN, ESPCN, etc. is used for super - resolution reconstruction. In one embodiment, a framework for improving the clarity of a video using an encoder - decoder structure is used, such as Figure 4 shown. After the video frame passes through the encoder, a low - resolution feature map is obtained, and then after processing the low - resolution feature map by the method provided by the present invention, it is input into the decoder, thereby completing the improvement of the video clarity. Figure 5 The comparison diagrams before and after the clarity improvement are shown.

[0036] Embodiment 2, the present invention provides a system for improving the clarity of a surveillance video, such as Figure 6 shown, the system includes the following modules: A feature extraction module, which is used to segment a low-resolution video into video frames, and perform feature extraction on the video frames to obtain low-resolution feature maps; An attention calculation module, which is used to calculate the heat map of each channel of the low-resolution feature map by adopting a spatial attention mechanism, adaptively divide the low-resolution feature map into multiple sub-windows, for each channel in the low-resolution feature map, determine the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel, obtain the value of the area of the sub-window on the channel, use the offset degree of the sub-window on the channel to obtain queries and keys, and calculate the attention according to the queries, keys and values on the channel; A super-resolution reconstruction module, which is used to merge the attention results on all channels into a new feature map, and use the new feature map to obtain the super-resolution video frame corresponding to the video frame.

[0037] Preferably, the calculating the heat map of each channel of the low-resolution feature map by adopting a spatial attention mechanism specifically includes: Adopting channel maximum pooling for the low-resolution feature map to obtain a maximum pooling feature map, and activating the maximum pooling feature map through an activation function; Calculating the spatial attention of the channel feature map, and multiplying the activated feature map and the spatial attention of the channel feature map element by element to obtain the heat map of the channel.

[0038] Preferably, the determining the offset degree of the sub-window in the channel based on the sub-window and the heat map of the channel specifically includes: For each sub-window, extract the heat value area corresponding to the sub-window position from the corresponding channel heat map, and calculate the average value of the heat value area as the average heat of the sub-window on the channel; Adopting a non-linear mapping function to calculate the offset degree corresponding to the average heat, or obtaining the maximum average heat of the neighboring sub-windows of the sub-window, and determining the offset degree according to the maximum average heat and the average heat of the sub-window.

[0039] Preferably, the obtaining the queries and keys by using the offset degree of the sub-window on the channel specifically includes: Obtain the center of the sub-window, calculate the offset distance according to the offset degree, and move the offset distance in the direction of the center of the channel feature map or the center of the neighboring sub-window with the maximum average heat of the sub-window; Taking the moved center as the center and the size of the sub-window as the size to obtain the queries and keys.

[0040] Preferably, the calculating the offset distance according to the offset degree specifically includes: Calculate the product of half of the height of the sub-window and the offset degree to obtain the vertical movement distance, and calculate the product of half of the width of the sub-window and the offset degree to obtain the horizontal movement distance.

[0041] The above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0042] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, a software unit executed by a processor, or a combination of the two. The software unit can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC.

[0043] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0044] Although the present application has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the specification and drawings are merely exemplary illustrations of the present application as defined by the appended claims and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for improving the clarity of surveillance video, characterized in that: The method comprises the following steps: Segment the low-resolution video into video frames, and extract features from the video frames to obtain low-resolution feature maps; A spatial attention mechanism is used to calculate a heat map of each channel of the low-resolution feature map, and the low-resolution feature map is adaptively divided into a plurality of sub-windows. For each channel in the low-resolution feature map, the offset degree of the sub-window in the channel is determined based on the heat map of the sub-window and the channel, and the area of ​​the sub-window on the channel is obtained to obtain a value. The offset degree of the sub-window on the channel is used to obtain a query and a key, and attention is calculated according to the query, key, and value on the channel. The attention results on all channels are combined into a new feature map, and the new feature map is used to obtain a super-resolution video frame corresponding to the video frame.

2. The method according to claim 1, characterized in that The adaptive division of the low-resolution feature map into multiple sub-windows is specifically as follows: Treat the entire low-resolution feature map as an initial root node and obtain the predefined minimum sub-window size; Starting from the root node, the root node is used as a sub-window to calculate the variance of the pixel values ​​in the sub-window. If the variance exceeds the predefined division threshold, the node is divided into four sub-nodes of equal size. The same decomposition process is recursively performed on each child node until the size of the child node reaches a predefined minimum sub-window size or the variance of the child node is below the split threshold.

3. The method according to claim 1, characterized in that The spatial attention mechanism is used to calculate the heat map of each channel of the low-resolution feature map, specifically: For the low-resolution feature map, the channel maximum pooling method is used to obtain the maximum pooling feature map, and the maximum pooling feature map is activated by the activation function; The spatial attention of the channel feature map is calculated, and the heat map of the channel is obtained by element-by-element multiplication of the activated feature map and the spatial attention of the channel feature map.

4. The method according to claim 1, characterized in that The heat map based on the sub-window and the channel determines the offset degree of the sub-window in the channel, specifically: For each sub-window, extract the heat value area corresponding to the sub-window position from the corresponding channel heat map, and calculate the average value of the heat value area as the average heat of the sub-window on the channel; A nonlinear mapping function is used to calculate the degree of deviation corresponding to the average heat, or the maximum average heat of the neighborhood subwindow of the subwindow is obtained, and the degree of deviation is determined according to the maximum average heat and the average heat of the subwindow.

5. The method according to claim 1, characterized in that The query and key are obtained by utilizing the offset degree of the sub-window on the channel, specifically: Obtain the center of the subwindow, calculate the offset distance according to the offset degree, and move the offset distance toward the center of the channel feature map or the center of the neighborhood subwindow of the maximum average heat of the subwindow; The query and key are obtained with the moved center as the center and the sub-window size as the size.

6. The method according to claim 1, characterized in that The offset distance is calculated according to the offset degree, specifically: The vertical moving distance is obtained by calculating the product of half of the height of the sub-window and the offset degree, and the horizontal moving distance is obtained by calculating the product of half of the width of the sub-window and the offset degree.

7. A surveillance video definition enhancement system, characterized in that: The system includes the following modules: A feature extraction module is used to segment the low-resolution video into video frames and extract features from the video frames to obtain a low-resolution feature map; An attention calculation module is used to calculate the heat map of each channel of the low-resolution feature map by using a spatial attention mechanism, adaptively divide the low-resolution feature map into multiple sub-windows, and for each channel in the low-resolution feature map, determine the offset degree of the sub-window in the channel based on the heat map of the sub-window and the channel, obtain the area of ​​the sub-window on the channel to obtain a value, use the offset degree of the sub-window on the channel to obtain a query and a key, and calculate attention according to the query, key and value on the channel; The super-resolution reconstruction module is used to merge the attention results on all channels into a new feature map, and use the new feature map to obtain a super-resolution video frame corresponding to the video frame.

8. The system according to claim 7, characterized in that The adaptive division of the low-resolution feature map into multiple sub-windows is specifically as follows: Treat the entire low-resolution feature map as an initial root node and obtain the predefined minimum sub-window size; Starting from the root node, the root node is used as a sub-window to calculate the variance of the pixel values ​​in the sub-window. If the variance exceeds the predefined division threshold, the node is divided into four sub-nodes of equal size. The same decomposition process is recursively performed on each child node until the size of the child node reaches a predefined minimum sub-window size or the variance of the child node is below the split threshold.

9. The system according to claim 7, characterized in that The spatial attention mechanism is used to calculate the heat map of each channel of the low-resolution feature map, specifically: For the low-resolution feature map, the channel maximum pooling method is used to obtain the maximum pooling feature map, and the maximum pooling feature map is activated by the activation function; The spatial attention of the channel feature map is calculated, and the heat map of the channel is obtained by element-by-element multiplication of the activated feature map and the spatial attention of the channel feature map.

10. The system according to claim 7, characterized in that The heat map based on the sub-window and the channel determines the offset degree of the sub-window in the channel, specifically: For each sub-window, extract the heat value area corresponding to the sub-window position from the corresponding channel heat map, and calculate the average value of the heat value area as the average heat of the sub-window on the channel; A nonlinear mapping function is used to calculate the degree of deviation corresponding to the average heat, or the maximum average heat of the neighborhood subwindow of the subwindow is obtained, and the degree of deviation is determined according to the maximum average heat and the average heat of the subwindow.