Video quality enhancement method and related equipment

By extracting and fusing the motion feature information between video frames, the video frames are quality enhanced, which solves the problem of video compression artifacts in low bit rate or low latency scenarios, and achieves efficient and real-time video quality improvement.

CN119967106APending Publication Date: 2025-05-09ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510121364.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In low bit rate or low latency scenarios, video compression often introduces serious compression artifacts, resulting in a significant reduction in user viewing experience. Traditional methods have poor flexibility and limited performance when improving video quality, making it difficult to adapt to low latency scenarios.

Method used

By acquiring multiple video frames, the position offset information of the moving object is extracted, the motion feature information is determined using the deformation operation of the lookup table, and fused it to obtain the fused feature information. Finally, the video frame is quality-enhanced based on this information.

Benefits of technology

This method can efficiently and quickly achieve video quality enhancement in low-latency scenarios, with high real-time performance, effectively solve the problem of video compression artifacts, and significantly improve video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967106A_ABST
    Figure CN119967106A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video quality enhancement method and related equipment, and relates to the field of image processing. For quality enhancement of video frames, firstly, feature modulation processing is performed on a first frame and a second frame so as to extract position offset information of a moving object on the two video frames; further, deformation operation of the lookup table is carried out according to the two pieces of position offset information, motion feature information corresponding to the two video frames is determined, the two pieces of motion feature information are fused to obtain fused feature information, and further, multiple pieces of fused feature information under different resolutions are obtained to carry out quality enhancement processing on the second frame, so that the quality of the second frame is improved. In other words, feature information of different receptive fields is fused when the video quality is enhanced until the quality of each frame in the video sequence is enhanced to obtain the whole target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a video quality enhancement method and related equipment. Background Art

[0002] With the rapid development of Internet technology, the amount of video data in the network has exploded. In order to cope with the transmission needs of massive video data in the network, video compression technology is widely used to reduce the consumption of bandwidth and storage space. For example, H.264 / AVC and H.265 / HEVC are the current mainstream video coding standards. However, in low bit rate or low latency scenarios, video compression often introduces severe compression artifacts (such as blur, ringing effect and block effect), which significantly reduces the user's viewing experience (Quality of Experience, QoE).

[0003] In response to various problems caused by video compression, although traditional methods have improved the quality of compressed videos to a certain extent, they are usually less flexible and have limited performance, making it difficult to adapt to compressed video quality enhancement (CVQE) tasks in low-latency scenarios. Summary of the invention

[0004] The embodiments of this specification provide a video quality enhancement method and related devices to solve the above problems. The technical solution is as follows:

[0005] In a first aspect, an embodiment of the present specification provides a method for enhancing video quality, the method comprising:

[0006] Acquire a plurality of video frames sorted based on a time sequence; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame;

[0007] Performing feature modulation processing on the first frame and the second frame to extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one;

[0008] Determine, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and fuse the first motion feature information and the second motion feature information to obtain fused feature information;

[0009] Acquire multiple video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the multiple video frame combinations respectively through the above steps; wherein the first frame and the second frame included in the video frame combination have the same resolution, and the resolutions of the multiple video frame combinations are different;

[0010] The second frame is subjected to quality enhancement processing according to the plurality of fused feature information corresponding to the moving object to obtain a target second frame.

[0011] In a second aspect, an embodiment of the present specification provides a video quality enhancement device, the device comprising:

[0012] A video frame acquisition module, used to acquire a plurality of video frames sorted in time order; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame;

[0013] a convolution processing module, configured to perform feature modulation processing on the first frame and the second frame, and extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one;

[0014] a lookup table processing module, configured to determine, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and to fuse the first motion feature information and the second motion feature information to obtain fused feature information;

[0015] A multi-scale fusion module, used to obtain a plurality of video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the plurality of video frame combinations respectively through the above steps; wherein the resolutions of the first frame and the second frame included in the video frame combination are the same, and the resolutions of the plurality of video frame combinations are different;

[0016] The feature fusion module is used to perform quality enhancement processing on the second frame according to the multiple fused feature information corresponding to the moving object to obtain a target second frame.

[0017] In a third aspect, an embodiment of the present specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0018] In a fourth aspect, an embodiment of the present specification provides a computer program product, wherein the computer program product stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0019] In a fifth aspect, an embodiment of the present specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0020] The beneficial effects brought by the technical solutions provided by some embodiments of this specification include at least:

[0021] For video frame quality enhancement, feature modulation processing is first performed on the first frame and the second frame to extract the position offset information of the moving object in the two video frames, so that the timing information between multiple video frames can be used during quality enhancement to fully capture the association between multiple video frames and improve the effect of video quality enhancement;

[0022] Furthermore, a deformation operation of the lookup table is performed according to the two position offset information to determine the motion feature information corresponding to the two video frames, and the two motion feature information are fused to obtain fused feature information. The low-complexity lookup table algorithm is used to replace the neural network algorithm, which greatly reduces the amount of calculation and the performance requirements for the computing device. In addition, the computing task can be completed efficiently and quickly in a low-latency scenario, and the real-time performance is high.

[0023] Furthermore, multiple fused feature information at different resolutions are obtained to perform quality enhancement processing on the second frame to obtain the target second frame. That is, feature information of different receptive fields is integrated into the video quality enhancement, which effectively solves the problem of insufficient performance when extending the single-frame quality enhancement method to the multi-frame quality enhancement processing scenario, and greatly improves the quality enhancement effect on the second frame and even the quality enhancement effect on the entire video. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 It is a schematic diagram of the architecture of a video quality enhancement method provided in an embodiment of this specification;

[0026] Figure 2 is a flowchart of a video quality enhancement method provided by an embodiment of this specification;

[0027] Figure 3 is a schematic diagram of a scene for obtaining a first frame and a second frame provided in an embodiment of this specification;

[0028] Figure 4is a flowchart of a video quality enhancement method provided by an embodiment of this specification;

[0029] Figure 5 is a schematic diagram of a process for obtaining first position offset information and second position offset information provided by an embodiment of this specification;

[0030] Figure 6 is a flowchart of a video quality enhancement method provided by an embodiment of this specification;

[0031] Figure 7 is a schematic diagram of a process for obtaining first motion feature information provided by an embodiment of this specification;

[0032] Figure 8 is a structural diagram of a video quality enhancement device provided in an embodiment of this specification;

[0033] Fig. 9 It is a structural schematic diagram of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0035] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise clearly specified and limited, "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood in specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are an "or" relationship.

[0036] The present specification is described in detail below with reference to specific embodiments.

[0037] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the features, information and data involved in this specification are all obtained with full authorization.

[0038] With the rapid development of Internet technology, the amount of video data in the network has exploded. In order to cope with the transmission needs of massive video data in the network, video compression technology is widely used to reduce the consumption of bandwidth and storage space. For example, H.264 / AVC and H.265 / HEVC are the current mainstream video coding standards. However, in low bit rate or low latency scenarios, video compression often introduces severe compression artifacts (such as blur, ringing effect and block effect), which significantly reduces the user's viewing experience (Quality of Experience, QoE).

[0039] In response to various problems caused by video compression, although traditional methods have improved the quality of compressed videos to a certain extent, they are usually less flexible and have limited performance, making it difficult to adapt to compressed video quality enhancement (CVQE) tasks in low-latency scenarios.

[0040] In recent years, the rapid development of deep learning has promoted the research of CVQE based on deep neural networks. Early methods mostly focused on the quality enhancement of single-frame images. In order to better utilize the temporal information of the video, subsequent researchers proposed a variety of compressed video quality enhancement methods based on multiple frames. However, these methods usually have high computational complexity and strong dependence on hardware devices, making it difficult to meet the real-time requirements of low-latency scenarios.

[0041] Lookup table (LUT)-based algorithms have gradually attracted attention due to their high efficiency and have shown great potential in the field of image / video super-resolution. However, most of the current LUT-based methods focus on single-frame image processing, while existing multi-frame methods only use simple timing information through LUT weighting. Therefore, there is still a technical gap in the research on low-latency CVQE tasks, and it is urgent to propose efficient and real-time solutions to meet the needs of actual engineering implementation.

[0042] Therefore, in view of the above problems, the embodiments of this specification propose a video quality enhancement method and related devices to solve them. Figure 1 As shown, Figure 1 It is a schematic diagram of the architecture of a video quality enhancement method provided in an embodiment of this specification. Figure 1 The system at least includes a server 101 for executing the video quality enhancement method, and also includes a plurality of electronic devices for uploading a plurality of video frames and an execution instruction for instructing the server to execute the video quality enhancement method. The plurality of electronic devices at least includes an electronic device 1021, an electronic device 1022, and an electronic device 1023. It can be understood that Figure 1 The number of servers and electronic devices shown in the figure is for illustration only, and the embodiments of this specification do not impose any limitation on this.

[0043] The above-mentioned server 101 includes but is not limited to physical or virtual processors, mobile stations (MS), mobile terminals (Mobile Terminal), mobile phones (Mobile Telephone), handsets and portable equipment (portable equipment), Bluetooth headsets, smart watches and other types. The server 101 can communicate with one or more core networks via a radio access network (Radio Access Network, RAN).

[0044] For example, the server 101 is a plurality of physical servers, and the plurality of physical servers are independent in hardware. Alternatively, the server 101 is a plurality of virtual servers, and the plurality of virtual servers are deployed in the same hardware resource pool, and the deployment methods of the virtual servers include but are not limited to: VMware, Virtual Box, and Virtual PC.

[0045] It is understandable that the server 10 also has other service capabilities and functions to complete the tasks in the following embodiments. For example, the server 1011 also provides portal services, resource management services, and CI / CD services.

[0046] In the embodiments of the present specification, electronic devices such as the electronic device 1021, the electronic device 1022, and the electronic device 1023 may also be equipped with a display device, and the display device may be any device capable of realizing a display function, for example, the display device may be a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. For example, a user may use the display device on the electronic device 1021 to send multiple video frames to the server 101, or view a target video with enhanced quality sent by the server 101.

[0047] Multiple electronic devices and servers can communicate through communication links established by communication protocols, such as the gRPC protocol. gRPC is a high-performance, general-purpose open source remote procedure call (RPC) framework, which is mainly designed for mobile application development and based on the HTTP / 2 protocol standard. It is developed based on the protocol buffer (PB) serialization protocol and supports many development languages. In addition, the communication link can also be a wireless communication link or a wired communication link, for example, a wired communication link includes an optical fiber, a twisted pair or a coaxial cable, and a wireless communication link includes a Bluetooth communication link, a wireless fidelity (WIreless-FIdelity, Wi-Fi) communication link or a microwave communication link.

[0048] In one embodiment, Figure 2 As shown, it is a flow chart of a video quality enhancement method provided by an embodiment of this specification, which can be implemented by a computer program and can be run on a video quality enhancement device based on a von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0049] Specifically, the video quality enhancement method includes:

[0050] S102: Acquire a plurality of video frames sorted in time sequence, wherein the plurality of video frames include a first frame and a second frame following the first frame.

[0051] The video can be obtained through a variety of methods such as local files, network streams, cameras, wireless devices, etc. For example, the server 101 receives video streams from other electronic devices, which can be received through network protocols such as RTSP, RTMP, HLS, etc. For another example, the server 101 is provided with an image acquisition device, and the video is acquired through the image acquisition device.

[0052] Video frames can be understood as dividing a video based on a preset frame rate to obtain multiple video frames corresponding to the video. Each video frame includes various patterns. For example, the preset frame rate is 10 frames per minute, and 20 video frames can be obtained based on a 2-minute video. The specific value of the frame rate can be a value set based on the video quality enhancement, and can also be set by relevant personnel as needed.

[0053] In another embodiment, multiple video frames may be directly acquired, and the multiple video frames may be merged to obtain a video. For example, the server 101 directly receives multiple video frames sent from the electronic device 1021 .

[0054] The multiple video frames are sorted in chronological order, specifically, from early to late according to the occurrence time, or can be understood as sorting from front to back based on the generation time corresponding to the multiple video frames. It can also be understood that the order of the multiple video frames is the order of the multiple video frames obtained by segmenting the video when the video is played in the forward order.

[0055] like Figure 3 As shown, Figure 3 201 is a schematic diagram of a scene for obtaining a first frame and a second frame provided by an embodiment of the present specification. The plurality of video frames are sorted in time order, and the generation time of the first frame 201 is before the generation time of the second frame 202. It can be understood that any two video frames can be extracted from the plurality of video frames as the first frame and the second frame, and two video frames that are adjacent in order can also be extracted as the first frame and the second frame (such as Figure 3 The latter extraction method can better capture the timing information between two video frames and improve the quality enhancement effect of the video.

[0056] S104, performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information of the moving object on the second frame; wherein the number of the moving object is at least one.

[0057] Feature modulation processing can be understood as a convolution processing, which is a common method used to extract image features in computer vision. Through feature modulation processing, features such as edges and textures on the first frame and the second frame can be extracted, especially the motion change features of at least one moving object from the first frame to the second frame. Usually, a convolution kernel (such as an edge detection kernel, a blur kernel, etc.) is used to perform feature modulation processing on the first frame and the second frame to obtain the first convolution information and the second convolution information, that is, the first position offset information and the second position offset information.

[0058] There is a moving object on both the first frame and the second frame, and the number of the moving object is at least one, and the processing steps for obtaining the first position offset information of each moving object on the first frame and the second position offset information on the second frame are the same.

[0059] For example, firstly, the first frame and the second frame are preprocessed, such as converting both the first frame and the second frame into grayscale images to reduce the complexity of the convolution processing; further, a convolution kernel is selected, which is a filter used to process images. Commonly used convolution kernels include edge detection convolution kernels (such as Sobel or Prewitt kernels), blur convolution kernels, sharpening convolution kernels, etc. The selected convolution kernel is applied to the first frame and the second frame and the convolution results are extracted, including convolution result conv1 obtained by convolution of the first frame and convolution result conv2 obtained by convolution of the second frame; the gradient amplitude of the first frame and the second frame is calculated by the two convolution results. If the gradient of some areas changes significantly between the two frames, those areas may contain moving objects; image processing methods (such as contour detection, area filling, etc.) can be further used to analyze the shape and position of the moving object, and obtain the first position offset information of the moving object in the first frame and the second position offset information in the second frame.

[0060] In one embodiment, the time domain data of the first frame and the second frame are aligned, and feature modulation processing is performed on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0061] The purpose of aligning the time domain data of the first frame and the second frame is to ensure that the pixels in the video frames corresponding to different times correspond to the same physical area (i.e. different moments in the same scene). If the first frame and the second frame are not aligned, the image content at the two moments may be offset due to the movement of the object, camera shake or other factors, resulting in inaccurate detection of the moving object, affecting subsequent analysis. In other words, aligning the time domain data can avoid false detection or missed detection of moving objects due to the disordered position of the video frame content, and extract the correct first position offset information and second position offset information of the moving object.

[0062] In this embodiment, the movement of the moving object can be detected more accurately by comparing the position offset information corresponding to the same area in the first frame and the second frame, aligning the time domain data of the first frame and the second frame. In another embodiment, the time domain data of the first frame and the second frame can also be aligned by an image alignment method of feature points or an image alignment method based on optical flow, which can be set by relevant personnel as needed.

[0063] like Figure 4 As shown, Figure 4 301 and the second frame 302 are used as input for encoding, and the encoded data corresponding to the first frame and the second frame are input into the motion feature modulation network for convolution processing to obtain the first position offset information of the moving object on the first frame and the second position offset information on the second frame. Figure 4 As shown, the moving object is a volleyball, and the pixel positions in the first frame Sample and the second frame Offsets are different.

[0064] S106: Determine first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information through a deformation operation of a lookup table, and fuse the first motion feature information and the second motion feature information to obtain fused feature information.

[0065] The table used in the Lookup Table (LUT) is a mapping table that returns corresponding predefined information based on certain input values ​​(such as convolution output, specific features, etc.). The LUT can be a simple array or matrix whose index is the feature value or activation value of the position offset information. The lookup table can be optimized through training to better predict and select appropriate motion feature information based on the features in the position offset information.

[0066] A first motion feature information corresponding to the first frame is determined using a lookup table (LUT) according to the first position offset information, and a second motion feature information corresponding to the second frame is determined using a lookup table according to the second position offset information. For example, the first position offset information is a convolution feature map, which includes a large number of activation values, and these activation values ​​are mapped to the index of the lookup table, so as to obtain the motion feature associated with each convolution value in the convolution feature map until the first motion feature information is obtained.

[0067] The first motion feature information represents the motion mode, motion direction or speed of the moving object in the first frame, and the second motion feature information represents the motion mode, motion direction or speed of the moving object in the second frame. After obtaining the first and second motion feature information, the above motion feature information is fused to form more comprehensive motion feature information. Fusion can be performed in many ways, and common methods include weighted fusion, concatenation, element-by-element summation, etc.

[0068] like Figure 4 As shown, further, the first position offset information and the second position offset information are matched with the first motion feature information corresponding to the first position offset information and the second motion feature information associated with the second position offset information by a lookup table in the temporal feature extraction module, and the two motion feature information are further fused to obtain fused feature information.

[0069] In this embodiment, the number of moving objects is taken as an example. In other embodiments, the number of moving objects can also be multiple, that is, the fused motion feature information corresponding to the multiple moving objects is obtained through the time domain feature extraction module TFEM, and each motion feature information represents the motion feature of the moving object based on the timing information of the first frame and the second frame.

[0070] S108. Acquire multiple video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the multiple video frame combinations through the above steps; wherein the resolutions of the first frame and the second frame included in the video frame combination are the same, and the resolutions of the multiple video frame combinations are different.

[0071] At the same time, the resolution of the first frame and the second frame are changed to obtain another set of video frame combinations including the first frame and the second frame. Figure 3 The resolution of the first frame and the second frame shown is 1080P, and the fused motion feature information corresponding to the moving object is obtained based on the first frame 201 and the second frame 202 with a resolution of 1080P. The resolution of the first frame 201 and the second frame 202 is changed to 720P, and the fused motion feature information corresponding to the moving object is obtained based on the above steps. The resolution of the first frame 201 and the second frame 202 is changed to 480P again, and the fused motion feature information corresponding to the moving object is obtained based on the above steps.

[0072] like Figure 4 As shown, the video frames corresponding to different resolutions are combined and fused feature information is obtained through the motion feature modulation network MFMN and the time domain feature extraction module TFEM.

[0073] S110, performing quality enhancement processing on the second frame according to multiple fusion feature information corresponding to the moving object to obtain a target second frame.

[0074] The multiple fusion feature information corresponding to the moving object is fused and the second frame is subjected to quality enhancement processing to obtain a target second frame with quality enhancement for the second frame. Figure 4 As shown, fusion feature information corresponding to multiple resolutions is obtained, and the multiple fusion feature information is fused through the Multi-Scale Fusion Module, that is, the quality of the second frame is enhanced by combining the feature information of different receptive fields to obtain the target second frame 303.

[0075] Multi-Scale Fusion Module (MSFM) is a structure in convolutional neural networks, which can effectively combine information of different receptive fields (resolutions), especially in processing videos, dynamic scenes, image segmentation and other tasks. Through this multi-scale fusion module, information from local details to global features can be taken into account at the same time, thereby improving the quality of the second frame. This is because in the video quality enhancement process, the details of each frame may be different at different resolutions. For example, high-resolution features can capture detailed information, while low-resolution features can capture global information. Through multi-scale fusion, these features of different receptive fields can be combined, thereby improving the model's understanding and processing capabilities of diverse information.

[0076] Understandably, Figure 4 This is an example in which the number of moving objects in the first frame and the second frame is one. In the case in which the number of moving objects is multiple, multiple fused motion feature information corresponding to the multiple moving objects are respectively obtained, and multiple fused motion feature information at different resolutions corresponding to each moving object is obtained, and the multiple fused motion feature information at different resolutions corresponding to each moving object is fused in sequence, until the multiple fused motion feature information corresponding to all moving objects is fused, and then the second frame is quality enhanced to obtain the target second frame.

[0077] For quality enhancement of video frames, feature modulation processing is first performed on the first frame and the second frame to extract the position offset information of the moving object on the two video frames, so that the timing information between multiple video frames can be used during quality enhancement to fully capture the relationship between multiple video frames and improve the effect of video quality enhancement; further, a deformation operation of the lookup table is performed according to the two position offset information to determine the motion feature information corresponding to the two video frames respectively, and the two motion feature information are fused to obtain fused feature information. By replacing the neural network algorithm with a low-complexity lookup table algorithm, the amount of calculation and the performance requirements for the computing device are greatly reduced, and the computing task can be completed efficiently and quickly in a low-latency scenario with high real-time performance; further, multiple fused feature information at different resolutions is obtained to perform quality enhancement processing on the second frame to obtain the target second frame, that is, feature information of different receptive fields is integrated during video quality enhancement, which effectively solves the problem of insufficient performance when the single-frame quality enhancement method is extended to a multi-frame quality enhancement processing scenario, and greatly improves the quality enhancement effect of the second frame and even the quality enhancement effect of the entire video.

[0078] based on Figure 2 - Figure 4 See also the embodiment shown in Figure 5 , Figure 5 is a schematic diagram of a process for obtaining first position offset information and second position offset information provided by an embodiment of this specification, step S104 includes:

[0079] S202, performing feature modulation processing on the first frame and the second frame according to the motion vector of the moving object to obtain a multi-scale convolution kernel offset.

[0080] Motion vectors (MVs) can reflect the displacement or direction of a moving object and can be used to guide the convolution process so that the multi-scale convolution kernel can adapt to the motion in multiple video frames. The purpose of feature modulation is to dynamically adjust the convolution process based on motion information to better capture the motion changes of the moving object between the first and second consecutive frames. The offset of the multi-scale convolution kernel is generated based on the horizontal and vertical components of the motion vector.

[0081] like Figure 4As shown, the encoded data corresponding to the first frame and the second frame are input into the motion feature modulation network MFMN. The motion feature modulation network MFMN processes the first frame and the second frame through motion modulation. Among them, the feature modulation processing uses a spatiotemporal convolutional network to extract the motion features of moving objects from multiple videos, and pays attention to the changes between frames. Motion modulation is a key step in convolution processing. At this stage, the motion feature modulation network uses a specific modulation mechanism based on the extracted motion features to enhance the perception of motion features and highlight the dynamic information in the video frame.

[0082] S204, performing feature modulation processing on the first frame and the second frame according to the multi-scale convolution kernel offset, and extracting first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0083] Specifically, the first frame and the second frame are sampled according to the multi-scale convolution kernel offset to achieve intra-frame feature refinement and inter-frame information alignment, and extract the first position offset information of the moving object in the first frame and the second position offset information in the second frame. The design of the multi-scale convolution kernel allows the motion feature modulation network MFMN to perform feature modulation processing on the first frame and the second frame at different resolutions. The convolution kernel at each scale can be resized or offset according to the motion vectors of the first frame and the second frame.

[0084] In this embodiment, feature modulation processing is performed on the first frame and the second frame through the motion vector of the moving object to obtain a multi-scale convolution kernel offset, and the first frame and the second frame are sampled through the multi-scale convolution kernel offset to obtain first position offset information of the moving object on the first frame and second position offset information on the second frame, so as to efficiently align the time domain data of the first frame and the second frame, and better focus can be placed on the motion features of the moving object obtained based on the timing information of continuous frames.

[0085] based on Figure 2 - Figure 4 The present invention also includes the following embodiments:

[0086] S104, before performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame, the method further includes:

[0087] Background separation processing is performed on the background area and the region of interest on the first frame to obtain a separated frame.

[0088] Background subtraction processing is used in tasks such as video analysis, object tracking and motion detection. In this embodiment, by performing background subtraction on the first frame, the background area and the region of interest (foreground object) including the moving object are extracted, and a separated frame after background subtraction processing is obtained. Figure 4 As shown, the video frames used as input include a first frame 301 , a separated frame on which the background separation process is performed on the first frame 301 , and a second frame 302 .

[0089] S104, performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame, specifically including:

[0090] The first frame and the second frame are subjected to feature modulation processing based on the separated frame, and first position offset information of the moving object on the first frame and second position offset information on the second frame are extracted.

[0091] Specifically, since the separated frame preliminarily distinguishes the background area and the region of interest including the moving object, the accuracy of capturing the moving object existing in both the first frame and the second frame can be improved. Figure 4 As shown, the first frame 301 and the second frame 302 and the separated frame corresponding to the first frame 301 are encoded and input into the motion feature modulation network MFMN to extract the first position offset information of the moving object in the first frame and the second position offset information in the second frame.

[0092] In this embodiment, a separated frame is obtained by performing background separation processing on the first frame, and first position offset information and second position offset information of the moving object are obtained through the separated frame, so that the moving object extraction is more accurate and the possibility of identifying the background area as a moving object for convolution processing is reduced.

[0093] based on Figure 2 The present invention also includes the following embodiments:

[0094] S110, after performing quality enhancement processing on the second frame according to the multiple fusion feature information corresponding to the moving object to obtain the target second frame, the method further includes:

[0095] Two adjacent video frames are sequentially obtained from multiple video frames as the first frame and the second frame, and a target second frame with quality enhancement corresponding to the second frame is obtained based on the above steps, until the target second frame corresponding to the reverse first frame is obtained; multiple target second frames are subjected to video fusion processing to obtain a target video.

[0096] For example, the plurality of video frames include video frame 01, video frame 02, video frame 03, and video frame 04, and the video frame 01, video frame 02, video frame 03, and video frame 04 are arranged from front to back in chronological order. In other words, video frame 01 is the first frame in the positive order, and video frame 04 is the first frame in the reverse order.

[0097] Extract video frame 01 and video frame 02 as the first frame and the second frame, and then Figure 2 As shown in S102-S110, the target second frame 02 for quality enhancement of the video frame 02 is obtained. The video frame 02 and the video frame 03 are extracted as the first frame and the second frame. Figure 2 As shown in S102-S110, the target second frame 03 for quality enhancement of the video frame 03 is obtained. The video frame 03 and the video frame 04 are extracted as the first frame and the second frame. Figure 2 As shown in S102 - S110 , a target second frame 04 with quality enhancement for the video frame 04 is obtained.

[0098] The above-mentioned video frame 01, the target second frame 02, the target second frame 03 and the target second frame 04 are subjected to video fusion processing to obtain a target video. Compared with the original video obtained by video frame 01, video frame 02, video frame 03 and video frame 04, the target video has enhanced quality and reduces problems such as blur, ringing and block effects existing in the original video.

[0099] In one embodiment, Figure 6 As shown, it is a flow chart of a video quality enhancement method provided by an embodiment of this specification, which can be implemented by a computer program and can be run on a video quality enhancement device based on a von Neumann system. The computer program can be integrated into an application or run as an independent tool application.

[0100] Specifically, the video quality enhancement method includes:

[0101] S302: Acquire multiple video frames sorted in time sequence.

[0102] Refer to the above S102, which will not be repeated here.

[0103] S304 , performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0104] See the above S104, which will not be described again here.

[0105] S306: Acquire a feature information table and a condition table; wherein the condition table includes a plurality of condition parameters used in the table lookup process.

[0106] The table used in the Lookup Table (LUT) is a feature information table that returns corresponding predefined information based on certain input values ​​(such as convolution output, specific features, etc.). The LUT can be a simple array or matrix whose index is the feature value or activation value of the position offset information, and stores the motion feature information that matches the position offset information.

[0107] In this embodiment, a conditional LUT is introduced to assist in finding motion feature information corresponding to position offset information in the feature information table. The conditional table contains multiple conditions used in the lookup table and multiple parameters involved in the conditions, which are effective when performing query, screening or calculation.

[0108] S308: Determine, according to the condition table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information in the feature information table by performing a deformation operation of a lookup table.

[0109] According to the conditions and parameters provided in the condition table, through the table lookup process, according to the target condition that the first position offset information meets in the condition table, at least one parameter corresponding to the target condition is obtained, and according to the target condition and the at least one parameter, the feature information table is matched to obtain the first motion feature information associated with the first position offset information. The step of obtaining the second motion feature information according to the second position offset information is as described above.

[0110] In one embodiment, during the table lookup process, the first position offset information and the second position offset information are matched in the feature information table respectively through a tetrahedron interpolation algorithm and a condition table to obtain first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information.

[0111] Tetrahedral interpolation is an algorithm for interpolation in three-dimensional space. It is often used in numerical simulation, computer graphics, geographic information system (GIS), physical simulation and other fields. It is especially effective when processing unstructured grids or three-dimensional data.

[0112] The 4D-LUT algorithm combines the speed of table lookup and the accuracy of interpolation, which can significantly improve computing efficiency, especially when processing video quality enhancement that requires fast calculations.

[0113] In one embodiment, during the table lookup process, the first position offset information is quantized using a tetrahedral interpolation algorithm to obtain first low-order data and first high-order data; a first interpolation parameter is obtained in the condition table based on the first low-order data, and an interpolation calculation is performed in the feature information table based on the first interpolation parameter being greater than the first high-order data to obtain first motion feature information associated with the first position offset information.

[0114] like Figure 7 As shown, Figure 7 This is a flow chart of obtaining the first motion feature information provided by an embodiment of this specification. Taking the first position offset information 401 as an example, the four-dimensional first position offset information 401 is quantized to a size of 17*17*17*17, that is, the four-dimensional data (usually floating point numbers) is quantized to integer values, and the calculation burden in the search process is reduced by reducing the precision. The value of each dimension is quantized into 17 discrete levels according to its range.

[0115] In the table lookup process, the conditional parameter Query including the interpolation weight and interpolation point is first obtained in the lookup table through the first low-order data LSB obtained by quantization processing. The interpolation weight and interpolation point are searched using the low four-bit (low-precision) data as an index. These interpolation points represent certain positions in the feature information table LUT, and the interpolation weight indicates how to weight these points. Further, the final table lookup result is obtained by weighting the interpolation point with the first high-order data MSB of the high four bits. The calculation formula is as follows: Figure 7 The expression of Output in , that is, the first motion feature information corresponding to the first position offset information 401 is obtained.

[0116] Furthermore, the second position offset information is quantized by a tetrahedron difference algorithm to obtain second low-order data and second high-order data; a second interpolation parameter is obtained in the condition table according to the second low-order data, and an interpolation calculation is performed on the second high-order data in the feature information table based on the second interpolation parameter to obtain second motion feature information associated with the second position offset information.

[0117] The second motion feature information is obtained according to the second position offset information by using a tetrahedron difference algorithm. Figure 7 As shown, the first motion feature information is obtained through the first position offset information 401, which will not be described in detail here.

[0118] In this embodiment, continuous input data is compressed into discrete indexes through quantization operations, which facilitates fast searching. The search of the lower four bits of data can be regarded as a fast positioning process for obtaining the position of the interpolation point and the preliminary interpolation weight, while the weighted adjustment of the upper four bits of data is a detailed adjustment of the interpolation point, ensuring a higher-precision interpolation result and improving the accuracy of the obtained motion feature information.

[0119] S310: Fuse the first motion feature information and the second motion feature information to obtain fused feature information.

[0120] The first motion feature information represents the motion mode, motion direction or speed of the moving object in the first frame, and the second motion feature information represents the motion mode, motion direction or speed of the moving object in the second frame. After obtaining the first and second motion feature information, the above motion feature information is fused to form more comprehensive motion feature information. Fusion can be performed in many ways, and common methods include weighted fusion, concatenation, element-by-element summation, etc.

[0121] S312. Acquire multiple video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving objects in the multiple video frame combinations through the above steps; wherein the resolutions of the first frame and the second frame included in the video frame combination are the same, and the resolutions of the multiple video frame combinations are different.

[0122] Refer to the above S108, which will not be repeated here.

[0123] S316: Perform quality enhancement processing on the second frame according to the multiple fusion feature information corresponding to the moving object to obtain the target second frame.

[0124] Please refer to the above S110 and I will not go into details here.

[0125] For quality enhancement of video frames, feature modulation processing is first performed on the first frame and the second frame to extract the position offset information of the moving object on the two video frames, so that the timing information between multiple video frames can be used during quality enhancement to fully capture the relationship between multiple video frames and improve the effect of video quality enhancement; further, a deformation operation of the lookup table is performed according to the two position offset information to determine the motion feature information corresponding to the two video frames respectively, and the two motion feature information are fused to obtain fused feature information. By replacing the neural network algorithm with a low-complexity lookup table algorithm, the amount of calculation and the performance requirements for the computing device are greatly reduced, and the computing task can be completed efficiently and quickly in a low-latency scenario with high real-time performance; further, multiple fused feature information at different resolutions is obtained to perform quality enhancement processing on the second frame to obtain the target second frame, that is, feature information of different receptive fields is integrated during video quality enhancement, which effectively solves the problem of insufficient performance when the single-frame quality enhancement method is extended to a multi-frame quality enhancement processing scenario, and greatly improves the quality enhancement effect of the second frame and even the quality enhancement effect of the entire video.

[0126] The following are device embodiments of this specification, which can be used to implement the method embodiments of this specification. For details not disclosed in the device embodiments of this specification, please refer to the method embodiments of this specification.

[0127] See also Figure 8 , which shows a schematic diagram of the structure of a video quality enhancement device provided by an exemplary embodiment of the present specification. The video quality enhancement device can be implemented as all or part of the device through software, hardware or a combination of both. The device includes a video frame acquisition module 501, a convolution processing module 502, a lookup table processing module 503, a multi-scale fusion module 504, and a feature fusion module 505.

[0128] The video frame acquisition module 501 is used to acquire a plurality of video frames sorted in time order; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame;

[0129] A convolution processing module 502, configured to perform feature modulation processing on the first frame and the second frame to extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one;

[0130] A lookup table processing module 503 is used to determine the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information through a deformation operation of a lookup table, and to fuse the first motion feature information and the second motion feature information to obtain fused feature information;

[0131] The multi-scale fusion module 504 is used to obtain a plurality of video frame combinations including a first frame and a second frame, and obtain the fusion feature information corresponding to the moving object in the plurality of video frame combinations respectively through the above steps; wherein the resolutions of the first frame and the second frame included in the video frame combination are the same, and the resolutions of the plurality of video frame combinations are different;

[0132] The feature fusion module 505 is used to perform quality enhancement processing on the second frame according to the multiple fused feature information corresponding to the moving object to obtain a target second frame.

[0133] In one embodiment, the video quality enhancement apparatus further comprises:

[0134] A table acquisition module, used to acquire a feature information table and a condition table; wherein the condition table includes a plurality of condition parameters used in the table lookup process;

[0135] The lookup table processing module 503 includes:

[0136] a table processing unit, configured to determine, according to the condition table, in the feature information table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information by a transformation operation of a lookup table;

[0137] The information fusion unit is used to fuse the first motion feature information and the second motion feature information to obtain fused feature information.

[0138] In one embodiment, the table processing unit includes:

[0139] A table lookup subunit is used to match the first position offset information and the second position offset information in the feature information table respectively through a tetrahedral interpolation algorithm and the condition table during the table lookup process to obtain first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information.

[0140] In one embodiment, the table lookup subunit is specifically configured to:

[0141] In the table lookup process, the first position offset information is quantized by a tetrahedral interpolation algorithm to obtain first low-order data and first high-order data;

[0142] Acquire a first interpolation parameter according to the first low-order data in the condition table, and perform interpolation calculation on the first high-order data in the feature information table according to the first interpolation parameter to obtain first motion feature information associated with the first position offset information;

[0143] quantizing the second position offset information by a tetrahedron difference algorithm to obtain second low-order data and second high-order data;

[0144] A second interpolation parameter is obtained in the condition table according to the second low-order data, and an interpolation calculation is performed on the second high-order data in the feature information table according to the second interpolation parameter to obtain second motion feature information associated with the second position offset information.

[0145] In one embodiment, the convolution processing module 502 includes:

[0146] An alignment processing unit is used to align the time domain data of the first frame and the second frame, and perform feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0147] In one embodiment, the convolution processing module 502 includes:

[0148] A feature modulation unit, configured to perform feature modulation processing on the first frame and the second frame by using a motion vector of a moving object to obtain a multi-scale convolution kernel offset;

[0149] A convolution processing unit is used to perform feature modulation processing on the first frame and the second frame according to the multi-scale convolution kernel offset, and extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0150] In one embodiment, a video quality enhancement device comprises:

[0151] A separation unit, configured to perform background separation processing on the background area and the region of interest on the first frame to obtain a separated frame;

[0152] The convolution processing module 502 includes:

[0153] The separation convolution unit is used to perform feature modulation processing on the first frame and the second frame based on the separation frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0154] In one embodiment, a video quality enhancement device comprises:

[0155] a frame enhancement unit, configured to sequentially obtain two adjacent video frames from a plurality of video frames as the first frame and the second frame, and obtain a target second frame with quality enhancement corresponding to the second frame based on the above steps, until the target second frame corresponding to the reversed first frame is obtained;

[0156] The video enhancement unit is used to perform video fusion processing on the multiple target second frames to obtain a target video.

[0157] For quality enhancement of video frames, feature modulation processing is first performed on the first frame and the second frame to extract the position offset information of the moving object on the two video frames, so that the timing information between multiple video frames can be used during quality enhancement to fully capture the relationship between multiple video frames and improve the effect of video quality enhancement; further, a deformation operation of the lookup table is performed according to the two position offset information to determine the motion feature information corresponding to the two video frames respectively, and the two motion feature information are fused to obtain fused feature information. By replacing the neural network algorithm with a low-complexity lookup table algorithm, the amount of calculation and the performance requirements for the computing device are greatly reduced, and the computing task can be completed efficiently and quickly in a low-latency scenario with high real-time performance; further, multiple fused feature information at different resolutions is obtained to perform quality enhancement processing on the second frame to obtain the target second frame, that is, feature information of different receptive fields is integrated during video quality enhancement, which effectively solves the problem of insufficient performance when the single-frame quality enhancement method is extended to a multi-frame quality enhancement processing scenario, and greatly improves the quality enhancement effect of the second frame and even the quality enhancement effect of the entire video.

[0158] It should be noted that the video quality enhancement device provided in the above embodiment only uses the division of the above functional modules as an example when executing the video quality enhancement method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the video quality enhancement device provided in the above embodiment and the video quality enhancement method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0159] The serial numbers of the embodiments of this specification are for description only and do not represent the advantages or disadvantages of the embodiments.

[0160] The present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figure 1 - Figure 7 The video quality enhancement method of the embodiment shown in the figure can be specifically implemented by referring to Figure 1 - Figure 7 The specific description of the illustrated embodiment will not be repeated here.

[0161] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figure 1 - Figure 7 The video quality enhancement method of the embodiment shown in the figure can be specifically implemented by referring to Figure 1 - Figure 7The specific description of the illustrated embodiment will not be repeated here.

[0162] See also Fig. 9 , is a schematic diagram of the structure of an electronic device provided in the embodiment of this specification. Fig. 9 As shown, the electronic device 600 may include: at least one processor 601 , at least one network interface 604 , a user interface 603 , a memory 605 , and at least one communication bus 602 .

[0163] The communication bus 602 is used to realize the connection and communication between these components.

[0164] The user interface 603 may include a display screen (Display) and a camera (Camera), and the optional user interface 603 may also include a standard wired interface and a wireless interface.

[0165] The network interface 604 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0166] Among them, the processor 601 may include one or more processing cores. The processor 601 uses various interfaces and lines to connect various parts within the entire server 600, and executes various functions and processes data of the server 600 by running or executing instructions, programs, code sets or instruction sets stored in the memory 605, and calling data stored in the memory 605. Optionally, the processor 601 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 601 can integrate one or more combinations of a processor (Central Processing Unit, CPU), an image processor (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 601, and it can be implemented separately through a chip.

[0167] Among them, the memory 605 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 605 includes a non-transitory computer-readable storage medium. The memory 605 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 605 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments, etc. The memory 605 may optionally be at least one storage device located away from the aforementioned processor 601. As Fig. 9 As shown, the memory 605 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a video quality enhancement application.

[0168] exist Fig. 9 In the electronic device 600 shown, the user interface 603 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 601 can be used to call the video quality enhancement application stored in the memory 605 and specifically perform the following operations:

[0169] Acquire a plurality of video frames sorted based on a time sequence; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame;

[0170] Performing feature modulation processing on the first frame and the second frame to extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one;

[0171] Determine, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and fuse the first motion feature information and the second motion feature information to obtain fused feature information;

[0172] Acquire multiple video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the multiple video frame combinations respectively through the above steps; wherein the first frame and the second frame included in the video frame combination have the same resolution, and the resolutions of the multiple video frame combinations are different;

[0173] The second frame is subjected to quality enhancement processing according to the plurality of fused feature information corresponding to the moving object to obtain a target second frame.

[0174] In one embodiment, before the processor 601 performs the deformation operation of the lookup table to determine the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information, and fuses the first motion feature information and the second motion feature information to obtain the fused feature information, it further performs:

[0175] Acquire a feature information table and a condition table; wherein the condition table includes a plurality of condition parameters used in the table lookup process;

[0176] The processor 601 performs the deformation operation through the lookup table to determine the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information, and fuses the first motion feature information and the second motion feature information to obtain fused feature information, specifically performing:

[0177] According to the condition table, determining the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information by a transformation operation of a lookup table in the feature information table;

[0178] The first motion feature information and the second motion feature information are fused to obtain fused feature information.

[0179] In one embodiment, the processor 601 performs the step of determining, according to the condition table, the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information by a transformation operation of a lookup table in the feature information table, specifically performing:

[0180] In the table lookup process, the first position offset information and the second position offset information are matched in the feature information table respectively through the tetrahedron interpolation algorithm and the condition table to obtain the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information.

[0181] In one embodiment, the processor 601 performs the table lookup process, matches the first position offset information and the second position offset information in the feature information table respectively through the tetrahedron interpolation algorithm and the condition table, and obtains the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information, specifically performing:

[0182] In the table lookup process, the first position offset information is quantized by a tetrahedral interpolation algorithm to obtain first low-order data and first high-order data;

[0183] Acquire a first interpolation parameter according to the first low-order data in the condition table, and perform interpolation calculation on the first high-order data in the feature information table according to the first interpolation parameter to obtain first motion feature information associated with the first position offset information;

[0184] quantizing the second position offset information by a tetrahedron difference algorithm to obtain second low-order data and second high-order data;

[0185] A second interpolation parameter is obtained in the condition table according to the second low-order data, and an interpolation calculation is performed on the second high-order data in the feature information table according to the second interpolation parameter to obtain second motion feature information associated with the second position offset information.

[0186] In one embodiment, the processor 601 performs the feature modulation processing on the first frame and the second frame to extract the first position offset information of the moving object on the first frame and the second position offset information on the second frame, specifically performing:

[0187] The time domain data of the first frame and the second frame are aligned, and feature modulation processing is performed on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

[0188] In one embodiment, the processor 601 performs the feature modulation processing on the first frame and the second frame to extract the first position offset information of the moving object on the first frame and the second position offset information on the second frame, specifically performing:

[0189] Performing feature modulation processing on the first frame and the second frame by using a motion vector of the moving object to obtain a multi-scale convolution kernel offset;

[0190] The first frame and the second frame are subjected to feature modulation processing according to the multi-scale convolution kernel offset, and first position offset information of the moving object on the first frame and second position offset information on the second frame are extracted.

[0191] In one embodiment, before the processor 601 performs the feature modulation processing on the first frame and the second frame to extract the first position offset information of the moving object on the first frame and the second position offset information on the second frame, it further performs:

[0192] Performing background separation processing on the background area and the region of interest on the first frame to obtain a separated frame;

[0193] The processor 601 performs the feature modulation processing on the first frame and the second frame to extract the first position offset information of the moving object on the first frame and the second position offset information on the second frame, specifically performing:

[0194] The first frame and the second frame are subjected to feature modulation processing based on the separated frame, and first position offset information of the moving object on the first frame and second position offset information on the second frame are extracted.

[0195] In one embodiment, the processor 601 performs the quality enhancement processing on the second frame according to the plurality of fusion feature information corresponding to the moving object to obtain the target second frame, and specifically performs:

[0196] Sequentially acquiring two adjacent video frames from a plurality of video frames as the first frame and the second frame, and obtaining a quality-enhanced target second frame corresponding to the second frame based on the above steps, until obtaining a target second frame corresponding to the reversed first frame;

[0197] The plurality of target second frames are subjected to video fusion processing to obtain a target video.

[0198] For quality enhancement of video frames, feature modulation processing is first performed on the first frame and the second frame to extract the position offset information of the moving object on the two video frames, so that the timing information between multiple video frames can be used during quality enhancement to fully capture the relationship between multiple video frames and improve the effect of video quality enhancement; further, a deformation operation of the lookup table is performed according to the two position offset information to determine the motion feature information corresponding to the two video frames respectively, and the two motion feature information are fused to obtain fused feature information. By replacing the neural network algorithm with a low-complexity lookup table algorithm, the amount of calculation and the performance requirements for the computing device are greatly reduced, and the computing task can be completed efficiently and quickly in a low-latency scenario with high real-time performance; further, multiple fused feature information at different resolutions is obtained to perform quality enhancement processing on the second frame to obtain the target second frame, that is, feature information of different receptive fields is integrated during video quality enhancement, which effectively solves the problem of insufficient performance when the single-frame quality enhancement method is extended to a multi-frame quality enhancement processing scenario, and greatly improves the quality enhancement effect of the second frame and even the quality enhancement effect of the entire video.

[0199] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only storage memory, or a random access memory, etc.

[0200] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0201] The above disclosure is only the preferred embodiment of this specification, which certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for enhancing video quality, the method comprising: Obtain multiple video frames sorted based on time order; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame; Performing feature modulation processing on the first frame and the second frame to extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one; Determine, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and fuse the first motion feature information and the second motion feature information to obtain fused feature information; Acquire multiple video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the multiple video frame combinations respectively through the above steps; wherein the first frame and the second frame included in the video frame combination have the same resolution, and the resolutions of the multiple video frame combinations are different; The second frame is subjected to quality enhancement processing according to the plurality of fused feature information corresponding to the moving object to obtain a target second frame.

2. The video quality enhancement method according to claim 1, before determining the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information through a deformation operation of a lookup table, and fusing the first motion feature information and the second motion feature information to obtain fused feature information, further comprising: Acquire a feature information table and a condition table; wherein the condition table includes a plurality of condition parameters used in the table lookup process; The determining, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and fusing the first motion feature information and the second motion feature information to obtain fused feature information, includes: According to the condition table, determining the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information by a transformation operation of a lookup table in the feature information table; The first motion feature information and the second motion feature information are fused to obtain fused feature information.

3. The video quality enhancement method according to claim 2, wherein determining, according to the condition table, the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information in the feature information table by a deformation operation of a lookup table comprises: In the table lookup process, the first position offset information and the second position offset information are matched in the feature information table respectively through the tetrahedron interpolation algorithm and the condition table to obtain the first motion feature information associated with the first position offset information and the second motion feature information associated with the second position offset information.

4. The video quality enhancement method according to claim 3, wherein in the table lookup process, the first position offset information and the second position offset information are matched in the feature information table respectively by using a tetrahedron interpolation algorithm and the condition table to obtain first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, including: In the table lookup process, the first position offset information is quantized by a tetrahedral interpolation algorithm to obtain first low-order data and first high-order data; Acquire a first interpolation parameter according to the first low-order data in the condition table, and perform interpolation calculation on the first high-order data in the feature information table according to the first interpolation parameter to obtain first motion feature information associated with the first position offset information; quantizing the second position offset information by a tetrahedron difference algorithm to obtain second low-order data and second high-order data; A second interpolation parameter is obtained in the condition table according to the second low-order data, and an interpolation calculation is performed on the second high-order data in the feature information table according to the second interpolation parameter to obtain second motion feature information associated with the second position offset information.

5. The video quality enhancement method according to claim 1, wherein the step of performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame comprises: The time domain data of the first frame and the second frame are aligned, and feature modulation processing is performed on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame.

6. The video quality enhancement method according to claim 1, wherein the step of performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame comprises: Performing feature modulation processing on the first frame and the second frame by using a motion vector of the moving object to obtain a multi-scale convolution kernel offset; The first frame and the second frame are subjected to feature modulation processing according to the multi-scale convolution kernel offset, and first position offset information of the moving object on the first frame and second position offset information on the second frame are extracted.

7. The video quality enhancement method according to claim 1, before performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame, further comprising: Performing background separation processing on the background area and the region of interest on the first frame to obtain a separated frame; The performing feature modulation processing on the first frame and the second frame to extract first position offset information of the moving object on the first frame and second position offset information on the second frame includes: The first frame and the second frame are subjected to feature modulation processing based on the separated frame, and first position offset information of the moving object on the first frame and second position offset information on the second frame are extracted.

8. The video quality enhancement method according to claim 1, wherein after performing quality enhancement processing on the second frame according to the plurality of fusion feature information corresponding to the moving object to obtain the target second frame, the method further comprises: Sequentially acquiring two adjacent video frames from a plurality of video frames as the first frame and the second frame, and obtaining a quality-enhanced target second frame corresponding to the second frame based on the above steps, until obtaining a target second frame corresponding to the reversed first frame; The plurality of target second frames are subjected to video fusion processing to obtain a target video.

9. A video quality enhancement device, the device comprising: A video frame acquisition module, used to acquire multiple video frames sorted based on time sequence; wherein the plurality of video frames include a first frame and a second frame arranged after the first frame; a convolution processing module, configured to perform feature modulation processing on the first frame and the second frame, and extract first position offset information of a moving object on the first frame and second position offset information on the second frame; wherein the number of the moving object is at least one; a lookup table processing module, configured to determine, by a deformation operation of a lookup table, first motion feature information associated with the first position offset information and second motion feature information associated with the second position offset information, and to fuse the first motion feature information and the second motion feature information to obtain fused feature information; A multi-scale fusion module, used to obtain a plurality of video frame combinations including a first frame and a second frame, and obtain fusion feature information corresponding to the moving object in the plurality of video frame combinations respectively through the above steps; wherein the resolutions of the first frame and the second frame included in the video frame combination are the same, and the resolutions of the plurality of video frame combinations are different; The feature fusion module is used to perform quality enhancement processing on the second frame according to the multiple fused feature information corresponding to the moving object to obtain a target second frame.

10. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 8.

11. A computer program product, wherein the computer program product stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 8.

12. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video quality enhancement method and device, equipment and storage medium

    CN116309173A

  • Video enhancement method and device, equipment, storage medium and program product

    CN117061683A

  • Video enhancement method and device, electronic equipment, storage medium and program product

    CN118014862A

  • Video quality enhancement method and device, equipment and storage medium

    CN118485610A