Video definition improving method and system

By using gradient intensity and gradient direction consistency to weight and quantize the feature vectors in video definition improvement, the problem of weaker detail improvement and lower performance in video definition improvement is solved, and higher quality high-resolution video reconstruction is achieved.

CN119946377AInactive Publication Date: 2025-05-06QINGDAO YAODONG TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510164613.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Variable autoencoder has problems with weaker details and lower performance in improving video clarity.

Method used

By analyzing the video frame image, the gradient intensity and gradient direction of each pixel point are obtained and normalized. Then, the frame image is encoded by an encoder, the effective receptive field of the feature vector on the feature map is obtained, the gradient intensity and gradient direction consistency corresponding to the feature vector are determined based on the effective receptive field, and the gradient direction consistency is weighted and quantized, and finally a high-resolution frame image is generated using the decoder.

Benefits of technology

By classifying and quantifying the image area through the consistency of gradient intensity and gradient direction, the model can better pay attention to the structural information and detailed information of the image, and reconstruct higher quality, clearer and sharper high-resolution frame images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946377A_ABST
    Figure CN119946377A_ABST
Patent Text Reader

Abstract

The invention relates to a video definition improving method and system, and specifically, the method comprises the steps: analyzing a video to obtain a frame image in the video, obtaining the gradient intensity and the gradient direction of each pixel point in the frame image, and carrying out the normalization of the gradient intensity and the gradient direction; the method comprises the following steps: encoding a frame image through an encoder to obtain a feature map, obtaining an effective receptive field of a feature vector on the feature map in the frame image, determining gradient intensity and gradient direction consistency corresponding to the feature vector according to the effective receptive field, weighting the feature vector by using the gradient intensity corresponding to the feature vector to obtain a weighted feature map, and obtaining a weighted image; quantizing the feature vector by using gradient intensity and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result; and decoding the quantized result of the feature map by using a decoder to obtain a high-resolution frame image, and synthesizing the high-resolution images of all the frame images in the video into the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for improving video clarity. Background Art

[0002] Limited by the resolution, storage and bandwidth costs of camera equipment, as well as the quality limitations of historical data, low-definition videos are still widely used in streaming media, security monitoring, medical imaging, remote sensing monitoring and other fields. Improving video clarity can not only improve the user's viewing experience, but also play a huge role in key applications. For example, in the field of security monitoring, high-definition images can help more accurate face recognition, license plate recognition and other tasks, and improve security prevention capabilities; in the field of medical imaging, high-definition videos can help doctors more accurately analyze dynamic physiological signals and improve the accuracy of disease diagnosis; in autonomous driving and intelligent transportation systems, clear video input can enhance target detection and environmental perception capabilities, and improve the decision-making ability of the system. In addition, the demand for high-definition video compression and streaming transmission has also promoted the development of video super-resolution technology, allowing users to watch higher-quality video content under limited bandwidth.

[0003] In recent years, deep learning-based methods have made breakthrough progress in video super-resolution (VSR) tasks. Compared with traditional interpolation methods and optimization-based super-resolution reconstruction methods, deep neural networks can better model the mapping relationship from low resolution to high resolution and improve the restoration quality. At present, the mainstream methods include methods based on convolutional neural networks, recurrent neural networks, variational autoencoders, generative adversarial networks, and Transformer structures. Among them, the method based on vector quantization variational autoencoder has received widespread attention in video reconstruction tasks due to its powerful data compression and discrete feature modeling capabilities. However, it is better at restoring low-frequency structural information, but its restoration ability in high-frequency details is weak; in addition, it relies on a discrete embedding space for feature encoding, and the size of the embedding space affects not only the performance of the model, but also the reconstruction quality. Summary of the invention

[0004] In order to solve the problem of weak detail improvement and low performance of the variational autoencoder in video definition improvement, a video definition improvement method is provided in a first aspect, the method comprising: Parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; The frame image is encoded by an encoder to obtain a feature map, and the effective receptive field of the feature vector on the feature map in the frame image is obtained. The gradient strength and gradient direction consistency corresponding to the feature vector are determined according to the effective receptive field. The feature vector is weighted by the gradient strength corresponding to the feature vector to obtain a weighted feature map. The feature vector is quantized by the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result. The decoder is used to decode the quantization result of the feature map to obtain a high-resolution frame image, and the high-resolution images of all frame images in the video are synthesized into a video.

[0005] Optionally, the step of determining the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field is specifically as follows: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

[0006] Optionally, the step of weighting the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

[0007] Optionally, the feature vector is quantized using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

[0008] Optionally, obtaining a target feature vector according to the most recent embedding vector and the feature vector is specifically: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

[0009] In a second aspect, a video definition enhancement system is provided, the system comprising: A feature extraction module is used to parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; The encoding and quantization module is used to encode the frame image through the encoder to obtain a feature map, obtain the effective receptive field of the feature vector on the feature map in the frame image, determine the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field, weight the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map, and quantize the feature vector using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a quantized result of the feature map; The decoding and generation module is used to use a decoder to decode the quantization result of the feature map to obtain a high-resolution frame image, and synthesize the high-resolution images of all frame images in the video into a video.

[0010] Optionally, the step of determining the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field is specifically as follows: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

[0011] Optionally, the step of weighting the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

[0012] Optionally, the feature vector is quantized using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

[0013] Optionally, obtaining a target feature vector according to the most recent embedding vector and the feature vector is specifically: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

[0014] The present invention uses gradient strength to enhance features, highlight image edges and texture details, and enhance the feature representations corresponding to these important areas in the feature map, so that the model pays more attention to the structural information and detail information of the image in subsequent processing. Image regions are classified by two dimensions, gradient strength and gradient direction consistency, and different quantization strategies are adaptively selected according to the characteristics of different regions. The decoder can make full use of these optimized feature representations to reconstruct higher quality, clearer and sharper high-resolution frame images. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of Embodiment 1; Figure 2 Visualize the frame image and gradient intensity; Figure 3 Visualize the frame image and gradient direction; Figure 4 This is a visualization of the gradient direction consistency with a window of 5×5; Figure 5 Improve contrast image for clarity; Figure 6 This is a structural diagram of the second embodiment. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0017] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed description of some known functions and known components.

[0018] Figure 1 A first embodiment of the present invention is shown. Figure 1 The video definition improvement method in includes the following steps: Step 1: parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; Video is composed of a series of continuous image frames at a certain frame rate. Video files usually use a specific encoding format. Video decoders are used to decode video files and extract the image frames that make up the video. Frame images are static images in the video. In image processing, gradient is the speed and direction of the change of image pixel values, reflecting the degree of grayscale or color change of the image. In the edge, texture and other areas of the image, the pixel value changes dramatically, and the gradient value, that is, the gradient intensity, is relatively large; while in the flat area of ​​the image, the pixel value changes slowly, and the gradient value is relatively small.

[0019] The gradient direction reflects the direction of the pixel value change. The gradient is calculated for each pixel position in the image to obtain a gradient intensity map and a gradient direction map of the same size as the original image. The gradient is preferably calculated using a Sobel operator, a Prewitt operator, or a Canny operator. Then, the gradient intensity and the gradient direction are normalized respectively. Preferably, the gradient intensity is normalized using a Sigmoid function, and the gradient direction map is normalized using a Z-score normalization or a linear normalization. In another embodiment, the gradient intensity is normalized to (0, 1) and the gradient direction is normalized to [0, 1].

[0020] Figure 2 The frame image and the gradient intensity visualization of the frame image are shown. Figure 2 The gradient intensity visualization in the figure is obtained by calculating the gradient of the pixel point and normalizing the gradient value to 0-255. Figure 3 The frame image and the gradient direction visualization of the frame image are shown. Figure 3 The gradient direction visualization diagram in is obtained by calculating the gradient direction of the pixel point and normalizing the gradient direction to 0-255. In order to observe the gradient strength and gradient direction, Figure 2 The gradient magnitude and gradient direction are normalized to 0-255.

[0021] Step 2: Encode the frame image through an encoder to obtain a feature map, obtain the effective receptive field of the feature vector on the feature map in the frame image, determine the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field, weight the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map, and quantize the feature vector using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result; The frame image is processed by the encoder to extract the feature map. Each feature vector corresponds to an effective receptive field in the frame image. The effective receptive field is the actual impact area of ​​the feature vector in the frame image. The effective receptive field reflects the core area of ​​feature extraction more accurately than the theoretical receptive field. In an alternative embodiment, the receptive field can also be used to obtain the area of ​​the feature vector in the frame image. The gradient strength and gradient direction consistency corresponding to the feature vector are calculated based on the pixel changes in the area. The gradient strength reflects the degree of edge change of the feature vector in the frame image, while the gradient direction consistency indicates the directional stability or degree of change of the area. The feature vector is weighted according to the gradient strength to make the features of the edge and details more prominent, while reducing its influence on the low gradient area, thereby obtaining a weighted feature map.

[0022] When quantizing the feature map, in addition to considering the value of the feature vector itself, the corresponding gradient strength and gradient direction consistency are also considered. The gradient strength affects the accuracy of quantization and ensures that the features of the edge area remain clear, while the gradient direction consistency determines the quantization strategy, so that different quantization methods are used for corner points, edges, and flat areas, and finally the optimized feature map quantization results are obtained, which improves the quality of the decoded frame image.

[0023] Step 3: Use a decoder to decode the feature map quantization result to obtain a high-resolution frame image, and synthesize the high-resolution images of all frame images in the video into a video.

[0024] The decoder receives the result of the quantized feature map and converts it back into a high-resolution frame image. All frame images are processed into high-resolution versions after decoding, and the details of each frame are enhanced. All decoded high-resolution frames are synthesized into a video in the original time sequence to restore the dynamic playback effect. The originally low-resolution or blurred video is optimized, so that the final output video has higher clarity, richer details and a more natural visual experience.

[0025] Each eigenvector on the feature map encodes certain features of the image, and the eigenvector is linked to the local content of the image through the effective receptive field. The effective receptive field is not uniformly sensitive, and its central area usually contains more important and representative information. Using Gaussian weights, the eigenvector can pay more attention to the central area of ​​the effective receptive field and reduce interference from the edge area, so as to more accurately extract and utilize key information within the effective receptive field. By calculating the gradient strength and gradient direction consistency of the pixel points in the effective receptive field, the edge features of the effective receptive field area are quantified as gradient strength and gradient direction consistency. In a specific embodiment, the gradient strength and gradient direction consistency corresponding to the eigenvector are determined according to the effective receptive field, specifically: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

[0026] Get the center point of the effective receptive field, and use this center point as a reference to calculate the distance from each pixel in the effective receptive field to the center point. Gaussian distribution and this distance are used to calculate the weights. Pixels near the center of the effective receptive field will be given higher weights, while the weights of pixels far from the center will be reduced, forming a weight distribution with high center and low edge. The gradient strength of the pixels in the effective receptive field is weighted by weight, and the weighted gradient strength values ​​are calculated, such as summing or averaging, to obtain the gradient strength of the feature vector. For gradient direction consistency, it is directly based on the gradient direction of all pixels in the effective receptive field. The smaller the gradient direction consistency value, the more chaotic the gradient direction distribution, the larger the gradient direction consistency value, and the more regular the gradient direction distribution. In a specific embodiment, the calculation of the gradient direction consistency value is to use the variance or standard deviation of the direction as the intermediate value, and then input the intermediate value into the negative exponential function to obtain, so that the better the gradient direction consistency, the higher the corresponding value. Figure 4 The figure shows a visualization of the gradient direction consistency obtained by calculating the standard deviation of the gradient direction within the window 5×5 and then performing negative exponential calculation. It can be seen that the gradient direction consistency of the sky part is higher, while the gradient direction consistency of the trees, rhinoceros and other parts is lower.

[0027] The gradient strength and gradient direction consistency of all eigenvectors are normalized separately to ensure that the gradient information between different eigenvectors can participate fairly in subsequent calculations and model learning.

[0028] The greater the gradient strength, the greater the probability that the feature vector includes an edge or corner point in the frame image. The importance of the feature vector is dynamically adjusted by weighting the feature vector in the feature map by the gradient strength, and the feature vector corresponding to the area with high edge strength will be enhanced. In one embodiment, the feature map obtained by weighting the feature vector by the gradient strength corresponding to the feature vector is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

[0029] In a more specific embodiment, the gradient strength is directly used as a scoring criterion. The higher the gradient strength value, the higher the score and the greater the weight; conversely, the lower the gradient strength value, the lower the score and the smaller the weight. For example, there is a feature vector A on the feature map, and its corresponding gradient strength is 0.8. 0.8 is directly used as the weight of feature vector A; another feature vector B has a corresponding gradient strength of 0.3, and 0.3 is used as the weight of feature vector B.

[0030] In another embodiment, instead of using the gradient strength directly as the weight, a function is used to calculate the weight. The function is parameter-learnable and can be automatically adjusted through learning during model training to find the best way to calculate the weight. For example, a simple linear function is used to calculate the weight: weight = a*gradient strength + b. During training, the model will automatically learn the best values ​​of a and b. For example, the model may learn a=2, b=0.5. For feature vector A with a gradient strength of 0.8, its weight will be calculated as 2.1; for feature vector B with a gradient strength of 0.3, its weight will be 1.1. It should be noted that parameter-learnable functions are not limited to simple linear relationships.

[0031] After obtaining the weight, apply the weight to the corresponding feature vector. Preferably, element-by-element multiplication is adopted, that is, the weight value is multiplied by each element of the feature vector. For example, if the feature vector A is [0.5, 0.2, 1.0, 0.1] and the weight is 0.8, the weighted feature vector A is [0.4, 0.16, 0.8, 0.08].

[0032] Feature map quantization is to convert continuous feature vectors into discrete, finite representations. In VQ-VAE, the quantization step decouples the encoder and decoder. The encoder encodes the image into a feature vector, and then converts it into discrete indexes through the quantization step. The decoder only needs to reconstruct based on these discrete indexes without directly processing the continuous feature vector output by the encoder. The embedding space is a codebook that is pre-learned or defined. The codebook contains a set of discrete embedding vectors. For each continuous feature vector output by the encoder, the quantization step searches in the codebook to find the embedding vector closest to the continuous feature vector and replaces the original continuous feature vector. However, if the embedding space is too large, the search space will be larger, the computational complexity of searching the nearest neighbor will increase significantly, and the feature distribution of the training data may be overfitted; if the embedding space is too small, due to the limited representation capability, the quantization process will cause serious information loss, and high-quality images cannot be reconstructed, which will eventually lead to poor video clarity improvement. The reconstructed image may be blurred, distorted, missing details, and other problems. Among them, the embedding space size is the number of embedding vectors contained in the embedding space. In one embodiment, the feature vector is quantized using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain the feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

[0033] For areas with high gradient intensity but low directional consistency, such as corner points and complex texture areas, in order to retain the information of these areas as much as possible, the embedding space with the most embedding vectors is used as the target embedding space of this feature vector to describe these complex features and reduce information loss.

[0034] For areas with high gradient intensity and high directional consistency, such as boundaries, these areas are the main structural edges of the image and are crucial for perceiving image content and clarity. Compared with corners and complex textures, the feature patterns of clear boundaries may be relatively simpler and more regular. Using an embedding space with a less rich vocabulary for quantization can ensure clear edge expression while also appropriately compressing information and avoiding excessive redundancy.

[0035] For areas with low gradient intensity, such as flat areas and background areas, which usually contain less information, the embedding space with the smallest vocabulary is used for quantization.

[0036] According to the above three conditional branches, a target embedding space is selected for each feature vector, and the embedding vector closest to the current feature vector is searched in the selected target embedding space. Preferably, the Euclidean distance is used to calculate the distance. The nearest neighbor embedding vector searched is used as the target feature vector, and the feature vector in the original feature map is replaced by the target feature vector, thereby completing the quantization of the current feature vector. For example, for the feature vector of the whisker area of ​​a cat, according to conditional branch 1, the embedding space with the richest vocabulary is selected as the target embedding space, and then in the target embedding space, an embedding vector that is most similar to the current feature vector is found, and this most similar embedding space is used to represent the original feature vector to complete the quantization.

[0037] Determining the target feature vector according to the nearest neighbor, that is, directly using the nearest embedding vector as the quantization result, may not be optimal in some cases. In one embodiment, obtaining the target feature vector according to the nearest embedding vector and the feature vector is specifically: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

[0038] The embedding vector that is most similar to the feature vector is searched in the target embedding space. The feature vector is the original, unquantized feature vector. Each element in the feature vector is rounded, and the distance from the feature vector after rounding to the feature vector before rounding is calculated to obtain the rounding distance, which measures the distortion caused by the use of rounding quantization. If the embedding distance is greater than the rounding distance, it means that the distortion of integer rounding quantization is smaller, and the feature vector after rounding is selected as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector. Furthermore, the quantization method is made more flexible, and the model can adaptively choose whether to use the embedded vector quantization of the embedding space or the integer rounding quantization according to the specific situation of each feature vector, so as to better adapt to different types of features. Figure 5 Improved comparison image for clarity.

[0039] Figure 6 The video definition enhancement system provided by the second embodiment is shown, comprising: A feature extraction module 11 is used to parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; The encoding and quantization module 12 is used to encode the frame image through an encoder to obtain a feature map, obtain the effective receptive field of the feature vector on the feature map in the frame image, determine the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field, weight the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map, and quantize the feature vector using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result; The decoding and generating module 13 is used to use a decoder to decode the quantization result of the feature map to obtain a high-resolution frame image, and synthesize the high-resolution images of all frame images in the video into a video.

[0040] Optionally, the step of determining the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field is specifically as follows: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

[0041] Optionally, the step of weighting the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

[0042] Optionally, the feature vector is quantized using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

[0043] Optionally, obtaining a target feature vector according to the most recent embedding vector and the feature vector is specifically: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

[0044] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.

[0045] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0046] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

Claims

1. A method for improving video clarity, characterized in that: The method comprises: Parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; The frame image is encoded by an encoder to obtain a feature map, and the effective receptive field of the feature vector on the feature map in the frame image is obtained. The gradient strength and gradient direction consistency corresponding to the feature vector are determined according to the effective receptive field. The feature vector is weighted by the gradient strength corresponding to the feature vector to obtain a weighted feature map. The feature vector is quantized by the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a feature map quantization result. The decoder is used to decode the quantization result of the feature map to obtain a high-resolution frame image, and the high-resolution images of all frame images in the video are synthesized into a video.

2. The method according to claim 1, characterized in that The step of determining the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field is specifically as follows: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

3. The method according to claim 1, characterized in that The weighted feature map obtained by weighting the feature vector using the gradient strength corresponding to the feature vector is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

4. The method according to claim 1, characterized in that The feature vector is quantized by using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain the feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

5. The method according to claim 4, characterized in that The obtaining of the target feature vector according to the most recent embedding vector and the feature vector is specifically as follows: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

6. A video definition enhancement system, characterized in that: The system comprises: A feature extraction module is used to parse the video to obtain a frame image in the video, obtain the gradient strength and gradient direction of each pixel in the frame image, and normalize the gradient strength and gradient direction respectively; The encoding and quantization module is used to encode the frame image through the encoder to obtain a feature map, obtain the effective receptive field of the feature vector on the feature map in the frame image, determine the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field, weight the feature vector using the gradient strength corresponding to the feature vector to obtain a weighted feature map, and quantize the feature vector using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain a quantized result of the feature map; The decoding and generation module is used to use a decoder to decode the quantization result of the feature map to obtain a high-resolution frame image, and synthesize the high-resolution images of all frame images in the video into a video.

7. The system according to claim 6, characterized in that The step of determining the gradient strength and gradient direction consistency corresponding to the feature vector according to the effective receptive field is specifically as follows: Obtaining the center of the effective receptive field, calculating the distance between the pixel point in the effective receptive field and the center, and using Gaussian distribution and the distance to calculate the weight of the pixel point in the effective receptive field; The weight is used to weight the gradient strength of the pixel points in the effective receptive field to obtain the gradient strength corresponding to the feature vector; The gradient direction consistency corresponding to the feature vector is obtained according to the gradient direction of the pixel point in the effective receptive field; The gradient strength and gradient direction consistency of all feature vectors are normalized separately.

8. The system according to claim 6, characterized in that The weighted feature map obtained by weighting the feature vector using the gradient strength corresponding to the feature vector is specifically: The gradient strength corresponding to the feature vector is used as the weight, or the gradient strength corresponding to the feature vector is input into a parameter learnable function to obtain the weight; Using the weights to weight the feature vectors to obtain weighted feature vectors; The feature vector in the feature map is replaced by the weighted feature vector to obtain the weighted feature map.

9. The system according to claim 6, characterized in that The feature vector is quantized by using the gradient strength and gradient direction consistency corresponding to the feature vector to obtain the feature map quantization result, specifically: If the gradient strength is greater than the threshold and the gradient direction consistency is less than the preset value, the embedding space with the most discrete embedding vectors is used as the target embedding space; If the gradient strength is greater than the threshold and the gradient direction consistency is not less than the preset value, the embedding space with the next most discrete embedding vector is used as the target embedding space; If the gradient strength is not greater than the threshold, the embedding space with the least discrete embedding vector is used as the target embedding space; Searching for an embedding vector closest to the feature vector from the target embedding space, obtaining a target feature vector according to the closest embedding vector and the feature vector, and replacing the feature vector in the feature map with the target feature vector; When all feature vectors in the feature map are replaced, the feature map quantization result is obtained.

10. The system according to claim 9, characterized in that The obtaining of the target feature vector according to the most recent embedding vector and the feature vector is specifically as follows: Calculating the distance between the nearest embedding vector and the feature vector to obtain an embedding distance; Rounding each element in the feature vector, and calculating the distance between the feature vector after rounding and the feature vector before rounding to obtain a rounding distance; If the embedding distance is greater than the rounding distance, the rounded feature vector is used as the target feature vector, otherwise the nearest embedding vector is used as the target feature vector.

Citation Information

Cited By

  • Mineral resource low-altitude monitoring method based on unmanned aerial vehicle vision

    CN122223600A

  • A mineral resource low-altitude monitoring method based on unmanned aerial vehicle vision

    CN122223600B