Video filtering method and device based on MCTF, computing equipment and readable storage medium
Through the MCTF-based video filtering method, the reference frame set of target frames is determined and the filtering function is optimized, which solves the problem of poor existing video filtering effect, and achieves better video frame filtering and encoding effects, reducing data volume and transmission cost.
Patent Information
- Application Number
- CN202510567175.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-24
AI Technical Summary
The existing video filtering methods still need to be improved in terms of noise elimination and image quality, especially in video frames, the filtering effect is not ideal.
Using a method based on motion compensation time domain filtering (MCTF), the initial filtering function is adjusted to improve the filtering effect by determining the reference frame set of target frames, and inter-frame filtering is performed using motion estimation and compensation techniques to optimize the global correlation between the filtering results and the reference frame.
The filtering effect of video frames is improved, making the filtered target frame more uniform and smooth, improving the overall filtering quality of the video, and helping to better video encoding effect, reducing the amount of encoded data, and reducing the code rate required for transmission.
Smart Images

Figure CN120201204A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of video processing, and in particular, to a video filtering method based on MCTF, and also relate to a video filtering device based on MCTF, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of digital video technology, video processing has become increasingly important in the fields of multimedia communication and streaming media services.
[0003] Video filtering is one of the core links of video processing, which is used to eliminate noise, enhance image quality, suppress compression artifacts, etc., so as to improve the visual quality of the video and the accuracy of subsequent analysis. In the usual video filtering process, based on the pixel information within the video frame, the noise in the video frame is removed to achieve video filtering. However, the filtering effect of this method still needs to be improved. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a video filtering method based on MCTF, which can improve the video filtering effect. One or more embodiments of this specification also relate to a video filtering device based on MCTF, a computing device, a computer-readable storage medium, and a computer program product.
[0005] According to the first aspect of the embodiments of this specification, a video filtering method based on MCTF is provided, including:
[0006] Determining a reference frame set for a target frame in a target video; wherein, the target frame is any video frame in the target video, and the reference frame set includes the reference frames that directly and indirectly reference the target frame in the case of encoding of the target video;
[0007] Performing motion compensated temporal filtering (MCTF) processing on the target frame based on an initial filtering function, and adjusting the initial filtering function to obtain a target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set;
[0008] Performing MCTF processing on the target frame based on the target filtering function to obtain a target filtering result of the target frame.
[0009] According to the second aspect of the embodiments of this specification, a video filtering device based on MCTF is provided, including:
[0010] A reference frame determination module for determining a set of reference frames for a target frame in a target video; wherein, the target frame is any video frame in the target video, and the set of reference frames includes reference frames that directly and indirectly reference the target frame in the case of encoding the target video.
[0011] A filtering optimization module for performing motion compensated temporal filtering (MCTF) processing on the target frame based on an initial filtering function, and adjusting the initial filtering function to obtain a target filtering function based on the global correlation between the obtained filtering result and each reference frame in the set of reference frames.
[0012] A filtering module for performing MCTF processing on the target frame based on the target filtering function to obtain a target filtering result of the target frame.
[0013] According to a third aspect of the embodiments of the present specification, a computing device is provided, including: a memory and a processor;
[0014] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps in the above method are implemented.
[0015] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps in the above method are implemented.
[0016] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed in a processor, the steps in the above method are implemented.
[0017] In an embodiment of the present specification, the manner of the reference relationship of each frame in the encoding stage is used in the filtering process based on MCTF. For each video frame to be filtered, the reference frames that reference the video frame during encoding are determined, and the initial filtering function is adjusted based on the global correlation between the filtering result obtained by performing MCTF processing on the video frame based on the initial filtering function and each reference frame. Furthermore, MCTF processing is performed on the video frame based on the obtained target filtering function to obtain a better target filtering result. In this way, a filtering result with a higher degree of correlation with the reference frames can be obtained, and the filtering result can contain more information of other frames, which can make the filtered target frame more uniform and smooth, ensure a better filtering effect of the target frame, and correspondingly improve the overall filtering effect of the video. Description of the Drawings
[0018] Figure 1 It is a schematic structural diagram of an information interaction system provided by an embodiment of the present specification;
[0019] Figure 2 It is a flowchart of a video filtering method based on MCTF provided by an embodiment of this specification;
[0020] Figure 3 It is a schematic diagram of a GOP provided by an embodiment of this specification;
[0021] Figure 4 It is a schematic diagram of a pyramid search structure provided by an embodiment of this specification;
[0022] Figure 5 It is a flowchart of a video encoding method based on MCTF provided by an embodiment of this specification;
[0023] Figure 6 It is a timing diagram of an information interaction method provided by an embodiment of this specification;
[0024] Figure 7 It is a schematic diagram of the structure of a video filtering device based on MCTF provided by an embodiment of this specification;
[0025] Figure 8 It is a schematic diagram of the structure of a video encoding device based on MCTF provided by an embodiment of this specification;
[0026] Figure 9 It is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0027] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0028] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "based on a determination".
[0030] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0031] First, the meanings of some terms mentioned in this specification will be explained. As used in this specification, the term "in response to" means the state in which a corresponding event occurs or a condition is satisfied. The execution timing of the subsequent actions executed in response to this event or condition and the time when this event occurs or this condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent actions can be immediately executed when the event occurs or the condition is established; while in other cases, the subsequent actions can be executed after a period of time after the event occurs or the condition is established.
[0032] The term "published content" as used in this specification refers to content information published on a network platform, which may include information such as text, images, audio, and video, and can be displayed in the form of notes, articles, video files, or live broadcasts. The specific content information included in the published content applicable to different scenarios can be adjusted accordingly, and the content information in the published content can be determined by the publisher himself.
[0033] The term "trigger operation" as used in this specification refers to an operation performed on the information (such as a control) provided by a computer device. Through this operation, a corresponding instruction can be sent to the computer device to trigger the computer device to execute the next task. The next tasks triggered by performing different trigger operations on different information can all be preset in the program. This trigger operation can be manually executed by the user. For example, operations such as clicking, double-clicking, long-pressing, and swiping performed by the user on the content displayed on the screen of the computer device can all belong to the trigger operation. In some cases, this trigger operation can also be executed by the computer device based on a set program.
[0034] The "control" refers to any element with which a user can directly interact in a computer user interface. These elements allow users to input data, select options, or trigger certain actions. Controls can be graphical (such as buttons, text boxes, checkboxes, etc.) or in more abstract forms (such as voice commands). A button is typically used to perform an operation or open a new interface. A text box allows users to input or edit text information. A checkbox is used to represent a boolean value (yes / no), and users can select or deselect it. A dropdown list provides a list of options for users to choose from, usually only showing one selected value. A list box displays multiple options, and users can select one or more items. A label displays static text, usually used to explain the function of other controls and can also be used as a button. A scrollbar enables users to navigate through a large dataset, such as browsing a long document or list. An image displays a static picture and can sometimes also be used as a button. A combobox combines the functions of a text box and a dropdown list, allowing users to both input text and select from the list. These controls usually have standard styles provided by the operating system or development toolkits, and developers can customize their appearance and behavior according to needs. Different operating systems and programming languages (such as Java, C#, Python, etc.) have their own sets of controls and provide corresponding APIs to create and manage these controls.
[0035] With the development of information technology, various media contents (such as videos, texts, audios, or images) are gradually increasing, and the requirements for the presentation and transmission of media contents are getting higher and higher. For example, for videos, it is required to have good presentation effects, fast transmission rates, and ensure high video quality. To meet these requirements, video processing is usually needed, and the effects of this video processing are crucial for the presentation and transmission of videos.
[0036] Video filtering is a very important part of the video processing process. It mainly extracts useful information or removes unwanted components (such as noise) from the original video information to improve video quality, remove noise, enhance specific features, or achieve certain visual effects. For the video to be filtered, it can be filtered in an intra-frame filtering manner or an inter-frame filtering manner. Intra-frame filtering means filtering each frame in the video separately, and inter-frame filtering means considering the relationships between multiple frames (such as temporal smoothing or motion compensation) to filter each frame.
[0037] The information in a video frame can include high-frequency information and low-frequency information. Usually, people pay more attention to the low-frequency information. In one case, a low-pass filter can be used for video filtering to allow the low-frequency information in the video frame to pass through and suppress the high-frequency information. This filtering method can be used to smooth the image and remove noise. In another case, a high-pass filter can also be used for video filtering to allow the high-frequency information in the video frame to pass through and suppress the low-frequency information. This filtering method can be used for edge detection and image sharpening. In some cases, a band-pass filter can be used for video filtering to allow only the information within a certain frequency range to pass through. This filtering method can be used for extracting information of a specific frequency. In some cases, a band-stop filter can be used for video filtering to block the information within a certain frequency range. This filtering method can be used for removing interference of a specific frequency.
[0038] After filtering the video, it can be directly stored and played to adjust the video display effect. Video filtering can also be a part of other video processing processes. For example, video filtering can be a pre-step in the video encoding process. In many cases, videos need to be transmitted between different devices, such as in communication platforms, content consumption platforms, or data storage platforms, where video transmission between different devices is required. To ensure the video transmission efficiency and quality, the video needs to be encoded first to reduce the data volume while trying to ensure the video quality. In some cases, to reduce the occupation of storage space, video encoding is also required before storing the video. During the video encoding process, video filtering needs to be performed first to remove the redundant information in each video frame, and then the actual encoding process is executed, representing each video frame with the residual information between video frames. This video filtering also has a great impact on the overall video encoding effect. If the video filtering effect is good, it is easier to achieve a better video encoding effect, and then better compression of the video can be achieved, reducing the bit rate and bandwidth required for video transmission.
[0039] The embodiments of this specification provide a video filtering method based on Motion Compensation Temporal Filtering (MCTF), which can achieve a good video filtering effect. When this video filtering method is applied to the video encoding scenario, it can bring a good video encoding effect. The specific filtering process based on MCTF will be introduced in detail later and will not be elaborated here for the time being. This specification also relates to a video filtering device based on MCTF, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0040] In one way, the MCTF-based video filtering method provided by the embodiments of this specification can be applied to an information interaction system based on a content application platform. Figure 1 It is a schematic structural diagram of an information interaction system provided by an embodiment of this specification. As Figure 1 shown, the information interaction system may include a server device 101 and multiple terminal devices 102. The server device 101 and the terminal devices 102 can communicate bidirectionally through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks.
[0041] The server device 101 may correspond to a content application platform, and the content application platform corresponds to a target application program. The server device 101 may be the background server of the target application program, and is used to provide support for the operation of the target application program. The content application platform may be an application platform integrating various contents. The content application platform provides users with rich and diverse multimedia contents, including but not limited to graphic information, live broadcast, on-demand video, audio, graphic information, social interaction, etc. These multimedia contents may be published by users and may also be referred to as published contents. The content application platform may be a note publishing and browsing platform. Users can publish notes on this platform. That is, the published contents in the content application platform may be notes, and the notes may include one or more of text information, image information, and video information. Users can view the notes published by other users and can also perform information interaction with other users based on the notes. For example, different users can perform information interaction by liking, commenting, or sharing the notes. The notes viewed by users may be recommended by the content application platform or may be searched by users.
[0042] The terminal device 102 may be a device that simultaneously includes information receiving and information sending functions, that is, a device having hardware for receiving and sending in a two-way communication link to perform two-way communication. The target application program may be installed in the terminal device 102, and the communication connection with the server device 101 is realized by running the target application program, and then the functions of the content application platform are used. Information interaction or communication can be performed between different terminal devices 102 through the server device 101. The multiple terminal devices 102 in the information interaction system respectively correspond to multiple users. For example, each terminal device 102 corresponds to a user, and there may also be a situation where multiple terminal devices 102 correspond to one user. The users in the embodiments of this specification are represented by user accounts, and one user account represents one user. The terminal device 102 may log in to the content application platform based on the user account, and correspondingly, it is considered that the terminal device 102 corresponds to the user represented by the user account. If the same user account is logged in on multiple terminal devices 102, it is considered that the multiple terminal devices 102 correspond to the same user.
[0043] For example, in a scenario where a user uses the terminal device 102 to publish a video in a content application platform, video filtering and encoding can be performed. The user specifies the video to be published, and uses the terminal device 102 to filter and encode the video (such as executing the video filtering method and video encoding method provided in the embodiments of this specification), and uploads the encoded video to the server device 101 to improve the efficiency of video upload. After receiving the video, the server device 101 distributes the video to other terminal devices 102, and the terminal device 102 that receives the video can decode and play the video. Optionally, the server device 101 can transcode the received encoded video into another format or multiple formats suitable for distribution and playback, and then distribute the transcoded video to other terminal devices 102. In some ways, the terminal device 102 can also upload the video to the server device 101, and the server device 101 performs video filtering and encoding.
[0044] In some embodiments, the terminal device 102 can also only perform filtering processing on the video, and play or locally store the filtered video. In this way, the terminal device 102 can execute the video filtering method provided in the embodiments of this specification. Optionally, when only performing video filtering processing, the terminal device 102 may not be connected to the server device 101.
[0045] The server device 101 is a server responsible for processing tasks such as receiving, storing, transcoding, and distributing live streams. It can be an independent server or a server network or server cluster composed of servers, including but not limited to computers, network hosts, single network servers, multiple network server sets, or cloud servers composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing.
[0046] The terminal device 102 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (such as tablet computers, personal digital assistants, laptop computers, object content notebooks, netbooks, etc.), mobile phones (such as smart phones), wearable computing devices (such as smart watches, smart glasses, etc.) or other types of mobile devices, smart home devices or in-vehicle devices, etc., or stationary computing devices such as desktop computers or personal computers.
[0047] Figure 2 is a flowchart of a video filtering method based on MCTF provided in an embodiment of this specification. This video filtering method can be applied to Figure 1 the information interaction system shown, such as specifically applicable to the terminal device 102 in this information interaction system. As Figure 2As shown, the video filtering method includes the following steps 202 to 206. Hereinafter, the device that executes the video filtering method will be described as a video filtering device.
[0048] Step 202: Determine a reference frame set for a target frame in the target video; wherein, the target frame is any video frame in the target video, and the reference frame set includes reference frames that directly and indirectly reference the target frame in the case of encoding of the target video.
[0049] In the embodiments of this specification, the target video is any video that needs to be filtered, and for any video that needs to be filtered, the relevant introduction for the target video can be referred to. The user can specify the target video that needs to be filtered. Exemplarily, the target video can be a video taken by the user using a terminal device, or can also be a video stored in the terminal device, or a video downloaded from the Internet. Here, the format and acquisition method of the target video are not limited.
[0050] The video filtering device can perform filtering processing on each video frame in the target video. By performing filtering processing on each video frame in the target video, the filtering of the target video is achieved. The target frame is any video frame in the target video, and the filtering of each video frame in the target video can refer to the introduction for the target frame.
[0051] For the target frame, the video filtering device can determine a reference frame set composed of reference frames that satisfy the reference logic with the target frame in the case of encoding, and then perform filtering on the target frame based on this reference frame set. In the embodiments of this specification, the reference frame that satisfies the reference logic with the target frame refers to the video frame in the target video that directly or indirectly references the target frame in the case of encoding; correspondingly, the reference frame set can include each reference frame that directly and indirectly references the target frame in the target video. In one way, the video filtering device can use all the video frames in the target video that reference the target frame in the case of encoding as reference frames to form the reference frame set.
[0052] When the video is encoded, the corresponding encoding standard will stipulate the reference relationship between each video frame. For example, the encoding standard can be H.266 / VVC (Versatile Video Coding), which is a video compression standard used to provide more efficient video compression technology for different types of network environments and various resolutions, especially improving the video compression efficiency for high-resolution and high-quality video content. The encoding standard can also be other types of standards, such as H.265 / VVC, R266 and other encoding standards.
[0053] When encoding, for the current frame to be encoded, the correlation between video frames will be used to predict part or all of the content of the current frame, so that for the current frame, only the difference data between it and the predicted value needs to be stored, without storing the complete image data, so as to achieve the encoding of the current frame. According to the correlation between video frames, such as based on the scene changes in the video, each video frame in the video will be calibrated as a certain encoding prediction type. The encoding prediction type of each video frame can indicate other video frames that the video frame (such as the target frame) needs to refer to when encoding. The reference frame of the target frame during encoding is a video frame that is more relevant to the target frame. The encoding prediction type can indicate the direction of the video frame that the target frame needs to refer to relative to the target frame in the video, such as calling this direction the encoding prediction direction.
[0054] Coding prediction types can include key frames (I frames, Intra Frame), forward prediction frames (P frames, PredictiveFrame) and bidirectional prediction frames (B frames, Bidirectional Frame). Among them, I frames independently store complete image data, do not rely on other frames for encoding, and do not need to rely on other frames when decoding I frames. Coding prediction for I frames is called intra-frame prediction. P frames can refer to their forward I frames or P frames when encoding. P frames can store the parts that have changed compared to the forward frames, thereby significantly reducing the amount of data. B frames refer to multiple forward and backward frames (such as I frames and P frames) when encoding. B frames record the differences between the previous and next frames, and their data volume is usually the smallest. B frames also support time domain level (that is, the coding time level introduced later) prediction. The forward direction here refers to the earlier direction in the time dimension, and the backward direction refers to the later direction in the time dimension.
[0055] According to the encoding prediction type of each video frame in the target video, the video filtering device can determine the video frame in the target video that needs to directly refer to the target frame when encoding (such as called the first reference frame). The video filtering device can also determine the second reference frame that needs to be encoded with reference to the first reference frame, and the second reference frame indirectly refers to the target frame, and the second reference frame is also determined as the reference frame of the target frame. The video filtering device can also determine other video frames that are encoded with reference to the second reference frame for the second reference frame, and these video frames also indirectly refer to the target frame, and then determine these video frames as the reference frames of the target frame. By analogy, the video frames that directly and indirectly refer to the target frame in the target video can be found to form a reference frame set. In some embodiments, the video filtering device can also set the reference hierarchy of the selected reference frames, such as selecting only the first reference frame and the second reference frame to form a reference frame set.
[0056] In one embodiment, determining a reference frame set for a target frame in a target video in step 202 includes: determining a reference direction of the target frame in the target video based on the encoding prediction type of each video frame in the target video; and selecting a reference frame in the reference direction in the target video based on a limited number of reference frames to obtain a reference frame set.
[0057] The video filtering device can also obtain a limited number of reference frames to select reference frames that do not exceed the limited number to form a reference frame set. According to the aforementioned method of determining the first reference frame and the second reference frame, the video filtering device can determine, for each video frame (such as a target frame), the direction in which its corresponding reference frame is relative to the video frame based on the encoding prediction type of each video frame. In the embodiment of this specification, this direction is referred to as the referenced direction. If the video frames encoded with reference to the target frame are all located behind the target frame, the referenced direction of the target frame is backward. For example, if the video frames encoded with reference to the target frame exist both in front of and behind the target frame, the referenced direction of the target frame includes forward and backward.
[0058] For example, the target frame is a B frame, which is referenced by other video frames in the front and rear directions (that is, the referenced directions include forward and backward), and the limit number of reference frames is set to 8, then 4 video frames adjacent to the target frame can be selected as reference frames in the forward and backward directions. The limit number can also be 6, 10, 12, 14 or other arbitrary values, and the number of video frames selected in the forward and backward directions can also be different. If the target frame is referenced by other video frames in only one direction, the limit number of reference frames corresponding to the target frame can be different from the limit number for bidirectional reference. In the embodiments of the present specification, a limit number of reference frames is set for the target frame, so that calculations for too many reference frames can be avoided, reducing calculation time and resource consumption.
[0059] Optionally, the video filtering device can select reference frames according to the similarity between each video frame and the target frame. Generally, video frames that are closer to the target frame are more likely to be similar to the target frame, and reference frames can be selected based on the distance to the target frame, and reference frames close to the target frame can be selected preferentially. The distance between video frames can be determined based on the number of frames between video frames. For example, if the distance between adjacent video frames is 1, if there is one video frame between two video frames, the distance between the two video frames can be 2, and so on.
[0060] During video encoding, the video is divided into multiple groups of pictures (GOPs). Each GOP includes a series of continuous frames (i.e., images). These frames are organized together in a certain order and type. The type refers to the aforementioned coding prediction type. The correlation between the frames in a GOP is high. For example, in a video captured by a panning lens, the background may remain almost unchanged, while the position of the moving object changes slightly. The similarity between frames is called inter-frame redundancy. Video encoding can be performed by eliminating this redundancy, significantly reducing the amount of data that needs to be stored or transmitted. The video filtering device can determine the frame that each frame needs to refer to during encoding based on the GOP structure in the video, and then encode each video frame.
[0061] During video encoding, the video to be encoded can be repeatedly encoded according to its GOP structure. The GOP structure can utilize inter-frame redundancy (P frames and B frames) to significantly reduce the amount of video data, thereby improving compression efficiency. In addition, I frames, as key frames, allow random access to videos at any I frame, which is very important for applications such as streaming media and video on demand. In addition, since I frames are independent, if some frames are lost or damaged, when decoding the encoded video, the decoder can restart decoding from the next I frame to reduce error propagation.
[0062] In the coding standard, images in a GOP are divided into multiple coding time levels based on the reference relationship between video frames. Images in different coding time levels have different coding times and corresponding decoding times. Figure 3 is a schematic diagram of a GOP provided in an embodiment of this specification, and Figure 3 Take the example that the GOP includes 16 video frames. Figure 3 The frames are displayed from left to right in sequence. Each frame is represented by a serial number above it, which indicates the order in which the image is displayed. Figure 3 The shown GOP includes frames 0 to 15. Figure 3 The first video frame (frame 0) and the last video frame (frame 16) in the video are both I frames, and the other video frames are B frames. Figure 3 The last video frame (frame 16) in belongs to the next GOP. Figure 3 In the text, the coding time layer is represented by layer. Below, the coding time layer is referred to as the time layer.
[0063] Among them, layer0 includes frame 0, layer1 includes frame 8, layer2 includes frame 4 and frame 12, layer3 includes frame 2, frame 6, frame 10 and frame 14, and layer4 includes frame 1, frame 3, frame 5, frame 7, frame 9, frame 11, frame 13 and frame 15. In the encoding process, the frames of the latter temporal level in a GOP are encoded based on their coding prediction type with reference to the frames of the previous temporal level. For example, frame 8 is encoded with reference to frame 0, frame 4 and frame 12 are encoded with reference to frame 8, frame 2 and frame 6 are encoded with reference to frame 4, frame 10 and frame 14 are encoded with reference to frame 12, and so on. In the GOP structure, the lower the temporal level of the video frame, the more likely it is to be directly or indirectly referenced by more video frames. The video frames of the lower temporal level are more important, and the encoding and decoding time is also earlier.
[0064] For example, for frame 0, frame 8 directly references frame 0, and other frames directly or indirectly reference frame 8, which also indirectly reference frame 0. Therefore, frames 1 to 15 in the GOP are all video frames encoded with reference to frame 0, and can all be reference frames of frame 0. Similarly, for frame 8, it is directly or indirectly referenced by frames 1 to 7 and frames 9 to 15, and these frames can be reference frames of frame 8. The same applies to other frames.
[0065] Corresponding to the introduction of the first reference frame and the second reference frame, for frame 0, frame 8 directly refers to frame 0, and frame 8 belongs to the first reference frame of frame 0. Frame 8 is directly referenced by frames 1 to 7 and frames 9 to 15, and frames 1 to 7 and frames 9 to 15 can belong to the first reference frame of frame 8 and the second reference frame of frame 0. The same applies to other frames.
[0066] A GOP can be a relatively independent unit, with an I frame as the starting frame. During encoding, only video frames within the same GOP refer to each other. Modern coding standards have introduced a more flexible reference mechanism that allows video frames to refer to each other across GOPs. For example, in an open GOP structure, the first frame of a GOP (usually a P frame) can refer to a frame in the previous GOP. This design can improve compression efficiency, but it will increase decoding complexity and the difficulty of random access. In some implementations, a B frame can refer to a frame in the previous GOP to further optimize compression performance. These reference relationships can also be obtained based on the GOP structure of the video.
[0067] In the embodiment of the present specification, the video filtering device can determine the reference relationship between each video frame in the case of encoding based on the GOP structure of the video, and then determine the reference frame set of each video frame based on the reference relationship to achieve filtering of each video frame. Accordingly, in step 202, the reference frame set is determined for the target frame in the target video, including: for the target frame in the target video, based on the GOP structure of the group of pictures, determining the target GOP where the video frame encoded with reference to the target frame when the target video is encoded, and the encoding time level where the target frame is located; based on the video frame in the target GOP whose encoding time level is higher than the encoding time level where the target frame is located, the reference frame set is determined.
[0068] Based on the GOP structure in the target video, the video filtering device can determine the GOP where each video frame in the target video is located, the time level where each video frame is located, and the encoding prediction type of each video frame. The video filtering device can also determine the reference relationship between GOPs based on the GOP structure, such as whether cross-GOP reference is allowed. Based on the reference relationship between GOPs, the GOP where the video frame required to be referenced by each video frame is located can be determined. Accordingly, for any video frame in the target video (such as the target frame), the GOP where the video frame encoded with reference to the target frame is located is determined. In the embodiments of this specification, this GOP is referred to as the target GOP.
[0069] Since video frames at a higher temporal level will refer to video frames at a lower temporal level, video frames that directly or indirectly refer to the target frame can be determined in the target GOP based on the temporal level of each video frame. Video frames at a temporal level higher than the temporal level of the target frame directly or indirectly refer to the target frame, and the video filtering device can determine video frames at a temporal level higher than the temporal level of the target frame in the target GOP as reference frames of the target frame, thereby obtaining a reference frame set.
[0070] For example, please continue to refer to Figure 3 , each video frame in the GOP only references each other in the GOP. For frame 0, it belongs to the lowest temporal level layer0, and the temporal levels of other video frames are higher than this temporal level. Therefore, other video frames in the GOP (i.e., frames 1 to 15) can be determined as reference frames of frame 0 to obtain a corresponding reference frame set. For frame 8, it belongs to temporal level layer1, and the temporal levels of frames 1 to 7 and frames 9 to 15 are higher than this temporal level. Frames 1 to 7 and frames 9 to 15 can be determined as reference frames of frame 8, that is, the 7 video frames before and after frame 8 are used as reference frames to obtain a corresponding reference frame set.
[0071] In the embodiment of this specification, for the frame at the highest time level in the GOP, such as Figure 3For the video frame in layer 4, there is no frame in the GOP that references this frame for encoding. Therefore, for this video frame, the reference frame set can be left undetermined, and filtering can be directly performed based on its own information. Correspondingly, in the embodiments of this specification, the reference frame sets of video frames at other temporal levels (i.e., layer 0 to layer 3) outside the highest temporal level can be determined to adaptively improve the filtering performance and effect of these video frames based on these reference frames. The video filtering device can determine the number of frames at each temporal level, and accordingly, can obtain the number of reference frames included in the reference frame sets of the video frames at each temporal level.
[0072] The above division of temporal levels is only an example. In other ways, I-frames may not be included in the temporal levels, and the video frames that directly reference I-frames (such as Figure 3 frame No. 8 in
[0073] In some embodiments, the video filtering device can also screen among video frames at temporal levels higher than the temporal level of the target frame, such as screening based on a set limit number of reference frames, or screening based on the similarity or distance to the target frame. By way of example, for Figure 3 frame No. 8 in
[0074] In the embodiments of this specification, for the target frame to be filtered, the reference frame set that directly and indirectly references the target frame is determined. In this way, more reference frames can be determined, and the image data in these reference frames is similar to the target frame. Filtering the target frame based on these reference frames can make the filtered target frame more uniform and smooth, thereby achieving a better filtering effect. In the embodiments of this specification, based on the encoding temporal levels of each video frame indicated by the GOP structure, the frames at higher levels in the GOP where the video frames that reference the target frame for encoding are located are used as the reference frames of the target frame. In this way, the video frames that directly and indirectly reference the target frame for encoding can be determined relatively simply and conveniently, so as to improve the filtering efficiency.
[0075] Step 204: Perform motion compensation temporal filtering (MCTF) processing on the target frame based on the initial filtering function, and adjust the initial filtering function to obtain the target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set.
[0076] In the embodiments of this specification, the filtering process for videos can adopt a time-domain filtering method. The main purpose of time-domain filtering is to remove video noise. The noise in a video may come from the photon noise of the shooting device or the noise introduced during the transmission process, etc. Such noise not only affects the subjective viewing experience but also, due to the lack of correlation of the noise, it is difficult to effectively encode it subsequently, which is not conducive to video compression. The core idea of time-domain filtering is frame averaging, that is, performing weighted averaging on the current frame and the temporally adjacent frames to smooth the noise. However, directly using the method of "averaging corresponding positions" will cause a "trailing" phenomenon in the moving parts of the video.
[0077] In the embodiments of this specification, the video filtering device can use motion estimation (ME) and motion compensation (MC) and adopt the MCTF technology to perform filtering processing to enhance the compression performance of the video sequence and improve the quality. The main idea of MCTF is to transform the original video sequence into the time domain, generate a residual signal (i.e., the difference between adjacent frames) by predicting and compensating for the motion between adjacent frames, so as to achieve more efficient compression. During the motion estimation and motion compensation process, the video frames and the corresponding reference frames will be divided into blocks. For example, the video frames will be divided into multiple blocks according to a set block size (such as 16 pixels × 16 pixels). For each block (such as the target block), the best matching block can be found among the blocks in its corresponding reference frame, and the best matching block and the target block are weighted averaged to achieve better filtering of the target block, thus avoiding the introduction of the "trailing" phenomenon of motion and achieving a noise reduction effect in the spatial domain, thereby improving the subsequent coding efficiency.
[0078] For the target block to be filtered in the target frame in MCTF, motion estimation is first performed using each reference frame to determine the motion vector (MV, Motion Vector) between adjacent frames. The vector value of the motion vector usually includes information in the horizontal and vertical directions, indicating the displacement of the intra-frame block on the time axis. Among them, the pixel blocks between the target frame and the reference frame can be compared, and the reference block that best matches the target block is searched in the reference frame, so as to estimate the motion trajectory of the target block between adjacent frames and record it as the motion vector of MCTF. Then, according to the estimated motion vector, the pixel values of the target frame are adjusted, and the motion vector is used to locate and compensate the pixel positions within the frame, so that the homogeneous blocks (i.e., the target block and the reference block) on the target frame and the reference frame are aligned. After that, actual temporal filtering can be performed. The weights of the homogeneous blocks on different frames are calculated through a series of similarity metrics, and the motion-compensated frames are filtered using bidirectional or unidirectional filtering to remove short-term fluctuations and noise, thereby improving the visual quality of the video. Through the selection of the reference frames for the video frames to be filtered in the embodiments of this specification, MCTF can be performed based on these reference frames, and the performance and effect of MCTF can be adaptively maximized.
[0079] Exemplarily, when searching for the reference block during the motion estimation process in MCTF, block matching algorithms and search patterns such as Full Search, Three-Step Search, or Pyramid Search can be used to determine the reference block in the reference frame. To balance the search accuracy and computational complexity, a pyramid structure constructed with different downsampling ratios is generally used for the search, that is, the search is extended from a high downsampling ratio to high precision.
[0080] Figure 4 It is a schematic diagram of a pyramid search structure provided by an embodiment of this specification. Please refer to Figure 4 , first, the current frame to be filtered (such as the target frame) and its respective reference frames need to be downsampled twice to form a pyramid structure. This pyramid structure consists of the original resolution version of the video frame and two downsampled versions. Taking Figure 4 the pyramid structure of the target frame shown as an example, L0 is the target frame with the original resolution, L1 is obtained by performing 1 / 2 downsampling in both the horizontal and vertical directions, and L2 is obtained by performing another 1 / 2 downsampling on the basis of L1, that is, the 1 / 4 downsampled version of the original image.
[0081] The hierarchical motion estimation for the target frame starts from the L2 layer with the smallest resolution, and performs motion estimation with integer-pixel accuracy layer by layer. The processing unit for each layer is the same. For example, the processing unit for each layer is a 16×16 block, and the search start point for each layer depends on the optimal motion vector obtained from the search in the previous layer. Exemplarily, if the optimal motion vector obtained from the search in the previous layer corresponds to the third block in the first row, then during the search in the next layer, the search is performed in the area corresponding to this block and a certain area around it. Search layer by layer in this way until the most similar pixel or block is found in the reference frame with the original resolution.
[0082] After obtaining the set of reference frames for the target frame, the video filtering device can perform filtering processing on the video frame based on this set of reference frames. The video filtering device can perform MCTF processing on the target frame based on a filtering function to achieve filtering of the target frame. The filtering function indicates the MCTF processing procedure to be performed. The filtering function can include various filtering parameters, and different filtering parameters can be used at different stages of the MCTF processing procedure. Exemplarily, the filtering function includes motion detection parameters, filtering limit parameters, motion estimation and compensation parameters, and other parameters. The motion detection parameters can be used to detect the motion area, control the performance consumption brought by motion detection, and control the sensitivity of motion detection, etc. The filtering limit parameters can be used to control the ability of spatial noise reduction, limit the filtering amplitude, judge the weights of the spatial and temporal filtering results, limit the amplitudes of spatial and temporal filtering, modify the hybrid weights of the spatial and temporal domains, etc. The motion estimation and compensation parameters can limit and control the search order (such as from the L2 layer to the L1 layer and then to the L0 layer), control the unit of sub-pixel search, control the coefficients used during motion compensation, etc. The other parameters can be used to adjust the weight of temporal filtering and can also implement other auxiliary functions. These parameters work together to achieve noise reduction and quality improvement of the video frame through steps such as motion estimation and compensation, bilateral filtering, etc. The specific parameter settings can be adjusted according to the actual application scenario and requirements.
[0083] In the embodiments of this specification, the video filtering device can perform optimization analysis on the filtering result of the target frame to adjust the filtering result to obtain a better target filtering result. Exemplarily, it can be adjusted by adjusting the filtering influencing factors such as the method adopted in the filtering process, the parameters used, or the data considered, etc., and then perform filtering processing again based on the adjusted filtering influencing factors to obtain the adjusted filtering result. In an alternative manner, the video filtering device can also directly adjust the filtering result without adjusting the filtering process.
[0084] In the embodiments of this specification, an initial filtering function may be preset, and the parameter values of the filtering parameters in the initial filtering function are initial values. The video filtering device first performs MCTF processing on the target frame based on the initial filtering function to obtain a corresponding filtering result. Then, it determines whether the filtering result meets the requirements. If it does not meet the requirements, the initial filtering function is adjusted to obtain a better target filtering function, ensuring that a better filtering result can be obtained based on the target filtering function. In this way, the filtering result of the target frame can be optimized. In this specification, adjusting the initial filtering function may be to adjust the parameter values of the filtering parameters it contains, and the type of the filtering parameters included in the initial filtering function may not be adjusted.
[0085] For the target frame, the video filtering device may adjust the current filtering function (such as the initial filtering function) based on the global correlation between its filtering result (such as obtained based on the initial filtering function) and each reference frame in the reference frame set. In one implementation, the video filtering device determines the global correlation degree between the filtering result of the target frame and each reference frame in the reference frame set. The above global correlation can be represented by this global correlation degree. The higher the global correlation degree, the better the corresponding filtering result, and the more likely the filtering result meets the requirements. The video filtering device adjusts the initial filtering function with the goal of improving the global correlation degree.
[0086] The video filtering device may compare and analyze the filtering result of the target frame with each reference frame to determine the overall correlation degree between the filtering result of the target frame and each reference frame, and obtain the global correlation degree. In an optional implementation, the target frame and the reference frame set are integrated into a new set (such as called the target set), and the video filtering device determines the overall correlation degree between the filtering result of the target frame and this set, and uses this correlation degree as the global correlation degree. In the embodiments of this specification, the mutual information amount may be used as the above global correlation degree.
[0087] Correspondingly, before adjusting the initial filtering function to obtain the target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set in step 204 above, the video filtering method provided in the embodiments of this specification further includes: integrating the target frame and the reference frame set into a target set; for the target frame, determining the mutual information amount between the obtained filtering result and the target set; where the mutual information amount is used to characterize the global correlation between the filtering result and each reference frame in the reference frame set. In the embodiments of this specification, for any type of video frame, the target set can be determined to determine the global correlation degree based on this target set. In an optional manner, since only itself is referred to when encoding an I frame, the above method of determining the target set can be only applied to the case where the target frame is an I frame; when the target frame is not an I frame, the target set may not be determined, and the global correlation degree between the target frame and the reference frame set is directly determined.
[0088] Suppose the i-th video frame in the target video is denoted as X i , where the i-th video frame can be any video frame in the target video and can represent the target frame. The filtering function for filtering the i-th video frame in the target video is P i , and the filtered frame (i.e., the filtering result) is C i = P i (X i ).
[0089] In one implementation, the mutual information is used to represent the global correlation between the filtering result of the target frame and each reference frame in the reference frame set. The mutual information represents the amount of information shared by two variables. Taking the video filtering device that integrates the reference frame set and the target frame X i into the target set and then determining the global correlation based on the filtering result C i of the target frame and the target set as an example, this global correlation can be the mutual information This mutual information can measure the degree of independence between the filtering result C i and the target set from each other. The larger it is, the more information of each video frame in the target set is contained in C i . During the process of the video filtering device optimizing and analyzing the filtering result of the target frame, this should be maximized as much as possible so that C i retains the part most relevant to the target frame in
[0090] to improve the correlation. The mutual information can measure the statistical dependence between two random variables. The filtering result of the target frame can also be regarded as an image (such as called the target image). Determining the mutual information between the filtering result of the target frame and the target set is also to determine the mutual information between the target image and the target set. This mutual information can be jointly determined based on the entropy of the target image and the joint entropy of the target image and the target set.
[0091] Variables X and Y can be defined for the target image and the target set respectively. Variable X represents a certain feature of the target image, and variable Y represents the statistical feature of the target set. For the target set, each image in it can be analyzed and processed to determine the information that can represent the overall feature of the target set. This information can be similar to the feature information of an image. Based on this information, the variable Y corresponding to the target set can be defined. The mutual information I(X;Y) = H(X) + H(Y) - H(X,Y). Among them, H(X) represents the entropy of variable X, corresponding to the entropy of the target image; H(Y) represents the entropy of variable Y, corresponding to the entropy of the target set; H(X,Y) represents the joint entropy of variables X and Y, corresponding to the joint entropy of the target image and the target set. Exemplarily, the mutual information between the target image and the target set can be determined by successively performing five steps: defining variables, data discretization, estimating probability distributions, calculating entropy and joint entropy, and calculating mutual information. The mutual information between two images can also be calculated in a similar way, and only variable Y needs to be directly defined based on the features of the image.
[0092] In the process of defining variables, features (such as pixel values, gradients, etc.) are extracted from the target image and regarded as variable X. There are multiple ways to determine the random variable Y for the target set. In one way, the features of the average image of the target set are used as variable Y. The average image of the target set refers to the arithmetic average of the values of all images in the target set at each pixel position, and finally a single image representing the overall feature of the set is generated. In another way, the principal components of the target set are used as variable Y. The principal components refer to a set of orthogonal basis vectors extracted from the target set by the principal component analysis (PCA, Principal Component Analysis) method. These basis vectors are arranged in descending order of variance and can efficiently represent the main variation patterns of the images in the image set. In yet another way, if the images in the target set are ordered, the target set can be regarded as a multi-dimensional variable Y = {Y1, Y2,..., Yn}; correspondingly, the mutual information between the target image and the target set can be obtained by taking the average after calculating the mutual information I(X;Yi) between the target image and each image in the target set.
[0093] In data discretization, continuous pixel values are discretized (such as binning), and the value ranges of variable X and variable Y are divided into multiple intervals (such as histogram binning). During the process of estimating the probability distribution, the marginal distribution information P(X) of variable X and the marginal distribution information P(Y) of variable Y are determined through histogram or kernel density estimation. Based on the joint histogram of the target image and the target set, its joint distribution information P(X, Y) is determined. The joint histogram is used to count the joint occurrence frequency of pixel pairs in two images. The joint histogram is a two-dimensional table, where the horizontal axis and the vertical axis respectively represent the pixel values of the two images (or the intervals after binning), and each cell in the table records the number of times a pair of pixel values at the same position in the two images appear simultaneously. If the two images have been binned into K bins, the joint histogram is a K×K matrix. P(X) and P(Y) are obtained by summing the rows or columns of the joint histogram, and P(X, Y) = joint histogram / total number of pixels.
[0094] After that, based on the obtained marginal distribution information, using the entropy calculation formula, the entropy H(X) of the target image, the entropy H(Y) of the target set, and the joint entropy H(X, Y) of the target image and the target set can be determined. Entropy is a measure of the amount of information, which represents the uncertainty of information or the degree of surprise of information. If an image has a high entropy, it means that the pixel distribution of the image is relatively random and the amount of information is large; on the contrary, if the entropy is low, it means that the pixel distribution of the image is relatively regular and the amount of information is small. During the process of calculating the entropy for an image, it can be first converted into a grayscale image, and the frequency of each pixel value appearing in the image is counted (that is, the histogram of the image is determined). The probability of each pixel value is calculated according to the occurrence frequency of each pixel value, and then the entropy of the image is calculated using the entropy formula based on the probabilities of each pixel value. Among them, P(x i ) represents the probability of the i-th pixel value in the target image, and P(y j ) represents the probability of the j-th pixel value in the average image of the target set.
[0095] For example, for an 8-bit grayscale image, the pixel value range is 0 to 255, with a total of 256 possible values. Count the number of times each pixel value appears in the image. Suppose there is a 5×5 grayscale image. The pixel values in the first row are 0, 10, 10, 20, 20 in sequence. The pixel values in the second row are 10, 20, 30, 30, 40 in sequence. The pixel values in the third row are 20, 30, 40, 50, 50 in sequence. The pixel values in the fourth row are 30, 40, 50, 255, 255 in sequence. The pixel values in the fifth row are 40, 50, 255, 255, 255 in sequence. It can be counted that the pixel value "0" appears 1 time, "10" appears 3 times, "20" appears 3 times, and "255" appears 4 times. Here, all are not listed one by one. After that, the histogram is normalized to probability, and the occurrence probability of each pixel value is calculated, usually the frequency of the pixel value appearance divided by the total number of pixels. The total number of pixels in the above 5×5 image is 25. The corresponding probabilities of pixel values are P(0) = 1 / 25 = 0.04, P(10) = 3 / 25 = 0.12, P(255) = 4 / 25 = 0.16, and so on. The sum of the probabilities of all pixel values is 1. If some pixel values do not appear, then P(x) = 0, and they need to be ignored in subsequent calculations. Then, according to the entropy calculation formula, calculate P(x)log(P(x)) for each non-zero P(x), sum them up and take the negative to obtain the entropy of the image.
[0096] For the joint entropy, construct a joint histogram to count the frequency of each pixel pair appearance in two images, and calculate the probability that any pixel pair appears jointly in the two images. For example, assume that two images A and B have the same size (such as both M×N pixels). The probability can be determined based on the following steps 1 to 3. Step 1: Initialize a two-dimensional array (joint histogram). Assume the pixel value range is 0 to 255 (8-bit grayscale image), then the joint histogram is a 256×256 two-dimensional array with all initial values being 0. Step 2: Traverse each pixel of the two images. For the pixel pair at the corresponding positions in image A and image B (for example, A(x, y) = 50, B(x, y) = 100), find the corresponding cell in the joint histogram (abscissa = 50, ordinate = 100) and increment the value of that cell by 1. Repeat this step until all pixels are traversed. Then each cell of the joint histogram records the number of times a certain pixel pair appears. For example, the value of the histogram (50, 100) is 10, indicating that the pixel pair (A = 50, B = 100) appears 10 times. Step 3: Calculate the probability of each pixel pair. For each cell (i, j) in the joint histogram, its corresponding probability P(i, j) = histogram(i, j) / total number of pixels. P(i, j) represents the probability that the pixel pair (A = i, B = j) appears jointly. In this way, a joint probability distribution is obtained, representing the probabilities of all possible pixel pairs appearing in the two images. After that, use the definition formula of joint entropy Calculate and sum for each pixel pair to obtain the joint entropy of the two images.
[0097] After determining the entropies and joint entropy of the target image and the image set respectively, the mutual information I(X; Y) = H(X) + H(Y) - H(X, Y) can be calculated. In some ways, the mutual information can also be directly determined based on the marginal distribution information and joint distribution information of the target image and the image set respectively, such as the mutual information
[0098] After obtaining the mutual information based on the foregoing method, the global correlation degree can be characterized by this mutual information, and then the initial filtering function can be adjusted based on this global correlation degree so that the global correlation degree corresponding to the adjusted filtering function is higher. For example, an adjustment method that can obtain better filtering results can be set according to experience, such as specifying the adjustment order of each filtering parameter and the adjustment step size of the parameter value of each filtering parameter, and the initial filtering function can be adjusted accordingly based on this adjustment method. Optionally, the initial filtering function can also be randomly adjusted by the video filtering device. In the embodiments of the present application, the filtering function obtained after adjusting the initial filtering parameters is referred to as the target filtering function.
[0099] Step 206: Perform MCTF processing on the target frame based on the target filtering function to obtain the target filtering result of the target frame.
[0100] After adjusting the filtering function, the MCTF processing can be performed on the target frame again based on the obtained target filtering function. For the relevant introduction of this MCTF processing, please refer to the relevant introduction in the foregoing step 204 and will not be elaborated here. After re-performing the MCTF processing on the target frame based on the adjusted target filtering function, it can continue to be determined whether the obtained filtering result meets the requirements. If it does not meet the requirements, the filtering parameters are adjusted continuously until the obtained filtering result meets the requirements, and then the filtering result at this time is determined as the final target filtering result. That is, after step 204, the process of determining the global correlation degree corresponding to the current filtering function, adjusting the filtering parameters based on this global correlation degree, and performing MCTF processing again based on the adjusted filtering function can be executed multiple times.
[0101] The above content only takes the adjustment of the filtering function based on the global relevance to optimize the filtering result as an example. In the embodiments of this specification, an optimization index and the optimization goal to be achieved can be set. The optimization index is positively correlated with the above-mentioned global relevance. For the filtering result obtained by performing MCTF processing on the target frame based on any filtering function, the video filtering device calculates this global relevance, and then determines the optimization index based on this global relevance to determine whether the obtained optimization index meets the requirements, that is, to determine whether the optimization goal is achieved. When the optimization index does not meet the requirements, the filtering function is adjusted, and the target frame is re-performed MCTF processing based on the adjusted filtering function. For the filtering result obtained after each MCTF processing, the global relevance can be re-determined, and then the optimization index can be recalculated until the optimization index corresponding to the filtering result meets the requirements. The filtering function when the optimization index meets the requirements is used as the final target filtering function, and the filtering result obtained by performing MCTF processing on the target frame based on this target filtering function is used as the final target filtering result.
[0102] The requirements (i.e., the optimization goal) that the optimization index needs to meet can be preset. For example, a target value can be set, and taking the optimization index reaching the target value as the optimization goal. The target value can be a preset fixed value, or the target value can also not be a fixed value and can be adaptively changed during the optimization process. Exemplarily, the optimization goal can be that the optimization index reaches the maximum value that can be achieved. In this way, the target value can first be a smaller value. If the value of the optimization index obtained after each optimization is greater than the target value, the target value is updated to the value of the optimization index, and then the optimization continues until a larger optimization index cannot be obtained anymore. If Z represents the optimization index, the function for determining the target filtering result of the target frame can be expressed as That is, to determine the filtering function P that can maximize the optimization index Z i , and use the filtering result obtained based on this filtering function as the target filtering result.
[0103] In one implementation manner of the optimization index, the optimization index Optionally, a certain weight can be set for this global relevance, such as ω i , so that the optimization index includes This weight can be a preset empirical value or can be obtained through a certain way of debugging.
[0104] In the embodiments of this specification, the video filtering device optimizes the filtering result of the target frame based on the filtering result of the target frame and the global correlation degrees of the reference frames in the reference frame set, which can ensure that the obtained target filtering result has a relatively high global correlation degree with each reference frame. In this way, the target filtering result for the target frame can contain more information in each reference frame, that is, it contains more information that can be mutually referred to in adjacent frames. Filtering the target frame with more reference frames in this way can make the filtered target frame more uniform and smooth, improve the filtering effect on the target frame, and correspondingly improve the overall filtering effect on the target video.
[0105] In addition, after filtering in this way, it is easier to perform video encoding based on the obtained filtering result. When encoding these filtered reference frames, these reference frames can refer to the target frame, and then the same reconstruction quality effect can be restored with fewer bits, filtering out more redundant results for the entire video and improving the overall encoding effect on the target video.
[0106] When encoding a video, the bitrate (rate) usually needs to be considered. The bitrate represents the transmission rate or storage requirement of the encoded data, usually in bits per second (bps). The bitrate is related to the efficiency of the compression algorithm and the content complexity of the video (such as the degree of motion and texture details). Filtering optimization based on the global correlation degree facilitates better compression of the video during encoding, can optimize the bitrate loss, and correspondingly, only a smaller residual bitrate is required for encoding, reducing the bandwidth overhead required for video transmission.
[0107] When filtering or encoding a video, the distortion situation of the video can also be considered. The distortion situation represents the difference between the reconstructed signal and the original signal, which can be quantified by the mean squared error (MSE) or other similarity metrics. The distortion situation is related to the information loss during the pre-quantification process and the accuracy of the prediction model (such as the accuracy of motion estimation). In the embodiments of this specification, the video filtering device can also consider this distortion situation to determine the optimization metric to impose distortion constraints on the filtering result and avoid excessive distortion of the target frame due to filtering.
[0108] In one implementation, in step 204, based on the global correlation between the obtained filtering result and the reference frame set, adjusting the initial filtering function to obtain the target filtering function includes: calculating the global correlation degrees between the obtained filtering result and each reference frame in the reference frame set, and the distortion information of the filtering result relative to the target frame; aiming to improve the global correlation degree and / or reduce the distortion information, adjusting the initial filtering function to obtain the target filtering function. For the determination of the global correlation degree and the method of adjusting the filtering function based on the global correlation degree, please refer to the foregoing relevant introduction and will not be elaborated here. Since distortion affects the filtering effect, the filtering optimization should minimize the distortion information to reduce the influence of distortion. Accordingly, the initial filtering function can be adjusted based on this distortion information so that the distortion information of the filtering result obtained based on the adjusted filtering function is relatively low.
[0109] After the video filtering device performs filtering processing on the target frame to obtain the filtering result, in addition to determining the global correlation degrees between the filtering result and each reference frame in the reference frame set, it can also determine the distortion situation of the filtering result relative to the original information of the target frame, and this distortion situation is represented by the distortion information. For this global correlation degree, please refer to the foregoing relevant introduction and will not be elaborated here. For this distortion information, the video filtering device can calculate it in various ways and can select the calculation method according to actual requirements.
[0110] In an alternative way, the mean squared error (MSE) calculation method can be used to calculate the distortion information, and this method can measure the average value of the squares of the pixel value differences between the original image and the filtered image. As where M×N represents the image size, I o and I f represent the original image and the filtered image respectively. The smaller the value of MSE, the smaller the distortion. Due to the square amplification, this kind of distortion information is more sensitive to large errors.
[0111] In another alternative way, the peak signal-to-noise ratio (PSNR) calculation method can be used for calculation. This method is a quality evaluation index calculated based on MSE and can objectively measure the image quality. where MAX I represents the maximum pixel value (such as 255 for an 8-bit image). If PSNR > 30dB in this kind of distortion information, it is generally considered that the image quality is good.
[0112] In yet another alternative approach, the Structural Similarity Index (SSIM) calculation method can be adopted. This method takes into account the brightness, contrast, and structural information of the image, and is an image quality assessment method that is closer to human visual perception. where μ represents the local mean, σ represents the standard deviation, and σ xy represents the covariance; C1 and C2 represent stable constants. The value range of this type of distortion information is [-1, 1], where 1 indicates that the filtered result is exactly the same as the original image, and this representation of the distortion information is more in line with human visual perception.
[0113] For example, use E[d(X i , C i )] to represent the distortion information of the filtered result of the target frame relative to the target frame. In the embodiments of this specification, the optimization metric Z can be combined with and E[d(X i , C i )] to be jointly determined. Distortion will have a negative impact on the filtering effect. Therefore, when determining the optimization metric based on the distortion information, the optimization metric should be negatively correlated with this distortion information. For example, when filtering a video, it should be ensured that the original information of the video is not changed as much as possible, and the distortion should be minimized. In the embodiments of this specification, the video filtering device optimizes the filtering result by combining the global correlation and the distortion information, and uses the optimization metric to constrain the distortion, so as to obtain a target filtering result with a relatively high global correlation and a relatively low distortion information as much as possible, which can ensure that the obtained target filtering result has a small distortion and a good filtering effect.
[0114] In some embodiments, the global correlation and the distortion information can be weighted to obtain the optimization metric. In the embodiments of this specification, corresponding weights can be set for both the global correlation and the distortion information, and then the optimization metric can be determined in combination with this weight. For the weight of the global correlation, please refer to the foregoing introduction and will not be elaborated here. For the distortion information corresponding to the target frame, if the weight set for it is α i , so that the optimization metric includes α i ·E[d(X i , C i )]. This weight can be a preset empirical value or can be obtained by a certain method of debugging. In one implementation of the optimization metric, the optimization metric Z =
[0115]
[0116] The weights of the global relevance and the distortion information can also be adjusted according to requirements. If more consideration needs to be given to the relevance between the filtering result and the reference frame during the filtering process, the weight of the global relevance can be increased, and correspondingly, the allowable degree of distortion will be smaller. If more consideration needs to be given to the impact of the distortion information during the filtering process, the weight of the distortion information can be increased, and correspondingly, the allowable degree of distortion will also be smaller. If it is determined that the impact of the distortion information is small, the weight of the distortion information can be reduced to allow a certain degree of distortion in video filtering. In the embodiments of this specification, corresponding weights are set for the global relevance and the distortion information, and correspondingly, the impact of the reference frame and the filtering distortion on the filtering result can be considered according to requirements, which can improve the flexibility of video filtering and the video filtering effect.
[0117] During video filtering and encoding, it is usually necessary to balance the bitrate and the distortion. A lower bitrate will result in higher distortion and degrade the video quality. A higher bitrate can reduce the distortion but will increase the transmission bandwidth and storage requirements. In the embodiments of this specification, by optimizing the metrics, efforts are made to find a better balance between the bitrate and the distortion, maximize the video quality at a given bitrate, or minimize the bitrate under a given quality requirement. In the embodiments of this specification, for different video frames, the bitrate and distortion performance corresponding to the video frame can be dynamically optimized based on its actual reference frame set and filtering situation.
[0118] In the embodiments of this specification, the bitrate required for the filtered video can also be optimized based on the simplicity of the video frame itself. The simplicity of the video frame can be reflected by the complexity of the video frame and characterized by the noise situation in the video frame. Exemplarily, the simplicity can be represented by the entropy value, which is used to measure the complexity and information content of the image. The smaller the entropy value, the smaller the noise in the video frame, the smaller the complexity of the video frame, and the easier it is to be encoded and compressed. The video filtering device can use the entropy value of the video frame as a constraint condition to optimize the filtering result of the video frame, so that the obtained target filtering result can minimize the entropy value of the video frame as much as possible. Regarding the entropy value, reference can be made to the relevant introduction of entropy in the content introduced for the mutual information mentioned above.
[0119] In one embodiment, in step 204, based on the global correlation between the obtained filtering result and the reference frame set, adjusting the initial filtering function to obtain the target filtering function includes: calculating the global correlation degree between the obtained filtering result and each reference frame in the reference frame set, and the entropy value of the filtering result; aiming to improve the global correlation degree and / or reduce the entropy value, adjusting the initial filtering function to obtain the target filtering function. Regarding the determination of the global correlation degree and the method of adjusting the filtering function based on the global correlation degree, please refer to the foregoing related introduction, which will not be elaborated here. Since the entropy value characterizes the complexity and information amount of the image and can affect the filtering effect to a certain extent, the filtering optimization should reduce the entropy value to a certain extent to improve the overall filtering effect. Accordingly, the initial filtering function can be adjusted based on this entropy value so that the entropy value of the filtering result obtained based on the adjusted filtering function is relatively low.
[0120] After the video filtering device performs filtering processing on the target frame and obtains the filtering result, in addition to determining the global correlation degree between the filtering result and each reference frame in the reference frame set, the entropy value of the filtering result can also be determined. Regarding this global correlation degree, please refer to the foregoing related introduction, which will not be elaborated here. For this entropy value, in one way, the video filtering device can count the frequency of each pixel value appearing in the image to estimate its probability distribution, and obtain the entropy value based on this probability distribution.
[0121] For example, use H(C i ) to represent the entropy value of the filtering result of the target frame. The optimization index Z can be jointly determined in combination with and H(C i ). If the filtering result is difficult to compress, it will have a negative impact on the filtering effect. Therefore, when determining the optimization index based on the entropy value, the optimization index should be negatively correlated with this entropy value. As For the filtering of the video, it should be ensured that the video information is relatively concise, the compression complexity of the filtering result is controllable, the noise in the filtering result is small, and the entropy value is as small as possible. In the embodiments of this specification, the video filtering device optimizes the filtering result by combining the global correlation degree and the entropy value, and uses the optimization index to constrain the compression complexity, so as to obtain a target filtering result with a relatively high global correlation degree and a relatively low entropy value as much as possible, which can ensure a better filtering effect.
[0122] In some embodiments, the global correlation degree and the entropy value can be weighted to obtain the optimization index. In the embodiments of this specification, corresponding weights can be set for both the global correlation degree and the entropy value, and then the optimization index is determined in combination with this weight. Regarding the weight of the global correlation degree, please refer to the foregoing introduction, which will not be elaborated here. For the entropy value corresponding to the target frame, if the set weight is λ i , the optimization index includes λ i ·H(C i)。The weight can be a preset empirical value or can be obtained through a certain way of debugging. In one implementation of the optimization metric, the optimization metric
[0123] The weights of the global relevance and the entropy value can also be adjusted according to requirements. If more consideration needs to be given to the relevance between the filtering result and the reference frame during the filtering process, the weight of the global relevance can be increased, and correspondingly, the allowable complexity of the video frame will be smaller. If more consideration needs to be given to the impact brought by the complexity of the video frame during the filtering process, the weight of the entropy value can be increased, and correspondingly, the allowable complexity will also be smaller. In the embodiments of this specification, corresponding weights are set for the global relevance and the entropy value, and correspondingly, the impact of the reference frame and the filtering distortion on the filtering result can be considered according to requirements, which can improve the flexibility of video filtering and the video filtering effect.
[0124] In the embodiments of this specification, the optimization metric Z can be jointly determined by combining the above global relevance, distortion information, and entropy value. Corresponding weights can also be set for each type of information for weighted processing to obtain the optimization metric. Exemplarily, Correspondingly, the function for optimizing the filtering result of the target frame can be expressed as where, is used to optimize the bitrate required after video encoding, and α i ·E[d(X i , C i )] is used to constrain the distortion situation of the filtering result. In the embodiments of this specification, the video filtering device optimizes the filtering result of the target frame by combining the global relevance, distortion information, and entropy value to obtain a target filtering result with relatively high global relevance, relatively low distortion information, and relatively low entropy value as much as possible, which can ensure that the obtained target filtering result contains more reference information of adjacent frames, has less distortion, and lower complexity, and can achieve a better filtering effect.
[0125] The filtering function is a relatively complex algorithm, and its function form can be set. In the embodiments of this specification, based on the optimization metric corresponding to any video frame in the target video, the parameter values of the filtering function are adjusted to optimize the filtering function and ensure a better filtering result. In the embodiments of this specification, the same filtering method can be adopted for each video frame in the same video. For example, the filtering function can be adjusted and optimized for any video frame in the target video to determine the corresponding target filtering function, and the MCTF processing is performed on each video frame of the target video based on this target filtering function to obtain the corresponding target filtering result. This can avoid optimizing the filtering function for each video frame, reduce the computational resource consumption brought by the optimization, and balance performance and computational complexity. Optionally, the video filtering device can also select a more representative video frame (such as an I-frame) in the target video and determine the target filtering function based on this video frame. In some embodiments, if the computational performance of the video filtering device is strong enough, the filtering function can be adjusted and optimized for each video frame in the target video, and then the MCTF processing is performed on different video frames using filtering functions with different parameter values.
[0126] The implementation of the video filtering method provided in the embodiments of this specification will be introduced below in combination with a specific example. Continuing with the GOP shown in the Figure 3 illustration as an example. Assume that the target currently needs to filter the 8th frame in this GOP. Then, for this frame, the time layer it belongs to can be determined as layer1, and all frames in this GOP with a higher time layer are determined as the reference frames of the 8th frame (that is, frames 1 to 7 and frames 9 to 15), obtaining a reference frame set. After that, the 8th frame is filtered according to the set filtering method to obtain an initial filtering result. Calculate the global correlation between this filtering result and the reference frame set (such as represented by I1), the distortion information of this filtering result relative to the 8th frame (such as represented by E1), and the entropy value of this filtering result (such as represented by H1). Determine the optimization metric corresponding to this filtering result based on these three parameters obtained. For example, the current optimization metric is Z1 = I1 - H1 - E1. Compare this optimization metric with the set target value. If the value of this optimization metric is less than the target value, adjust the parameter values used in the filtering process. Based on the adjusted parameter values, filter the 8th frame again to obtain a new filtering result. Recalculate the global correlation between this new filtering result and the reference frame set (such as represented by I2), the distortion information of this new filtering result relative to the 8th frame (such as represented by E2), and the entropy value of this new filtering result (such as represented by H2), and then calculate the current optimization metric as Z2 = I2 - H2 - E2. Compare this optimization metric Z2 with the target value again. Iteratively optimize in this way multiple times until the determined optimization metric is greater than the target value, and use the filtering result here as the final filtering result for the 8th frame.
[0127] In the embodiments of this specification, for each frame to be filtered (such as the target frame) during the filtering process, the reference frame of the target frame can be selected according to the GOP structure during encoding, and each video frame that references the target frame is used as a reference frame. Then, based on the optimization metric determined by the global correlation between the filtering result and the reference frame, the filtering result is optimized to obtain a better target filtering result. In this way, a filtering result with a higher correlation with the reference frame can be obtained. This filtering result can contain more information of other frames, making the filtered target frame more uniform and smooth, ensuring a better filtering effect for the target frame, and correspondingly improving the overall filtering effect of the video.
[0128] In the embodiments of this specification, after filtering the target video, video encoding can continue based on the obtained filtering result. In some other embodiments, the filtering result of the target video can also be directly stored, or the filtered target video can be played.
[0129] Figure 5 is a flowchart of a video encoding method based on MCTF provided by an embodiment of this specification. This video encoding method can be applied to Figure 1 the information interaction system shown in the figure. For example, it can be specifically applied to the terminal device 102 in this information interaction system. As Figure 5 shown, this video encoding method specifically includes the following steps 502 and 504. Hereinafter, the device that executes this video encoding method is taken as an example of a video encoding device for explanation.
[0130] Step 502: Obtain the target filtering result of each video frame in the target video, where the target filtering result of each video frame is obtained based on the above-mentioned video filtering method.
[0131] In the embodiments of this specification, the video encoding device here and the aforementioned video filtering device can be the same device. For example, they are both Figure 1 the terminal device or the server device in the information interaction shown in the figure. After filtering the target video, this device directly performs encoding based on the obtained filtering result. The filtering result of the target video includes the target filtering result of each video frame in the target video.
[0132] In some embodiments, the video encoding device and the aforementioned video filtering device may also be different devices. After the video filtering device filters the target video, it sends the filtering result to the video encoding device for encoding. For example, the video filtering device may first store the obtained filtering result and send the filtering result to the corresponding video encoding device when video encoding is required. Exemplarily, the video encoding device may send a request for obtaining the filtering result of the target video to the video filtering device so that the video filtering device sends the filtering result of the target video based on this request. In some alternative ways, the video filtering device may be Figure 1 the terminal device in the information interaction shown in, and the video encoding device may be the server device in the information interaction.
[0133] Step 504: Encode the target video based on the target filtering results of the video frames in the target video.
[0134] The video encoding device may, according to actual requirements, encode the target video by using a suitable encoding standard or method based on the target filtering results of the video frames in the target video. This encoding is also to encode the filtering result of the target video.
[0135] Exemplarily, the video encoding device may perform inter prediction (Inter Prediction) or intra prediction (Intra Prediction) on the video frames based on the encoding prediction types of the video frames in the target video. Exemplarily, intra prediction is performed on I frames, and inter prediction is performed on P frames and B frames. For inter prediction, temporal redundancy information may be utilized to generate the predicted pixel values of the current frame by performing motion estimation and compensation on the previously encoded frames. For intra prediction, spatial redundancy information is utilized to obtain the predicted pixel values by predicting the information of adjacent blocks within the current frame.
[0136] After prediction, the video encoding device may calculate the difference between the original pixel values and the predicted pixel values. This difference is called the residual or prediction error. Since the residual data has lower energy and higher correlation, the reference data is generally easier to compress than the original data. Then, the residual data may be transformed (such as discrete cosine transform DCT or discrete sine transform DST), converted from the spatial domain to the frequency domain, and then the transformed data is quantized to reduce the amount of data. The quantized data may be compressed using entropy coding to further reduce redundancy and achieve efficient data compression to generate the final bitstream.
[0137] In the embodiments of this specification, through the foregoing filtering method, the filtered target frame can be made more uniform and smooth, ensuring a better filtering effect for the target frame, and correspondingly achieving a better filtering effect for the video. After obtaining the video filtering result by using this filtering method, video encoding is further performed. In this way, the amount of data to be processed for encoding is less, the correlation between data is higher, and the compression complexity of the data is lower. The same reconstruction quality effect can be restored with fewer bits, and more redundant results can be filtered out for the entire video. Therefore, more efficient and fast video encoding can be achieved, ensuring that the amount of video data obtained after encoding is less, the required bit rate is lower, and a better video encoding effect is obtained.
[0138] The video encoded by using the video encoding method provided in the embodiments of this specification can be transmitted between multiple devices. Exemplarily, it can be applied to Figure 1 the information interaction system shown, and the encoded video can be transmitted in this information interaction system. When transmitting this encoded video, the required bandwidth overhead is also lower, which can improve the user consumption experience and reduce the operating cost of the information interaction system.
[0139] Next, in combination with the transmission process of this video, an information interaction method is used to introduce the operation process of the video in an exemplary application scenario. Figure 6 FIG. is a timing diagram of an information interaction method provided by an embodiment of this specification, in which the operation processes of each device in the information interaction system for the video are introduced. The server device in the information interaction system can correspond to a content application platform for content-consuming users, and multiple terminal devices in the information interaction system can include terminal devices corresponding to content-consuming users. Figure 6 Taking two terminal devices (the first terminal device and the second terminal device respectively) in the information interaction system as an example for introduction, and taking the first terminal device as the device for uploading the video to the content application platform, with the first terminal device performing video filtering and encoding, and the second terminal device as the device for obtaining the video from the content application platform as an example. The first terminal device and the second terminal device can be different devices or the same device. For other terminal devices, the introduction for the first terminal device and the second terminal device can be referred to. In the embodiments of this specification, the operations for the video can all be triggered based on the content application platform. As Figure 6 shown, this information interaction method can include the following steps 602 to step 614. Figure 6 The content of each step in can be mutually referred to the similar content introduced previously.
[0140] Step 602, the first terminal device obtains the target video to be published.
[0141] In the embodiments of this specification, step 602 can be referred to in conjunction with the relevant introduction to the target video in the foregoing step 202.
[0142] The first terminal device can obtain the target video to be filtered when accessing the content application platform. For example, if an application corresponding to the content application platform is installed on the first terminal device, the first terminal device can, when running this application, obtain the target video based on the relevant operations performed in this application. Alternatively, the first terminal device can access the web page corresponding to the content application platform and obtain the target video based on the relevant operations performed in this web page.
[0143] Exemplarily, when the first terminal device accesses the content application platform, the page displayed on the screen may include a video publishing control. The user can perform a triggering operation on this video publishing control, such as clicking on this video publishing control, to trigger the first terminal device to obtain the target video to be published. For example, the first terminal device can display the stored videos for the user to select the target video from these videos; the first terminal device can also display a shooting control for the user to trigger this shooting control to capture the target video in real time. Before the target video is actually published, filtering and encoding need to be performed, and correspondingly, the video filtering method and video encoding method provided in the embodiments of this specification can be executed.
[0144] Step 604: The first terminal device determines a reference frame set for the target frame in the target video; wherein, the target frame is any video frame in the target video, and the reference frame set includes the reference frames that directly and indirectly reference the target frame in the case of encoding the target video.
[0145] In the embodiments of this specification, step 604 can refer to the relevant introduction in the foregoing step 202, and details are not described here.
[0146] Step 606: The first terminal device performs motion compensated temporal filtering (MCTF) processing on the target frame based on the initial filtering function, and adjusts the initial filtering function to obtain the target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set; based on the target filtering function, MCTF processing is performed on the target frame to obtain the target filtering result of the target frame.
[0147] In the embodiments of this specification, step 606 can refer to the relevant introduction in the foregoing steps 204 and 206, and details are not described here.
[0148] Step 608: The first terminal device encodes the target video based on the target filtering results of each video frame in the target video.
[0149] In the embodiments of this specification, step 608 can refer to the relevant introduction in the foregoing step 504, and details are not described here.
[0150] Step 610: The first terminal device sends the encoded target video to the server device.
[0151] During the process of the first terminal device accessing the content application platform, it can establish a communication connection with the server device. Then, it can send information to the server device and receive information sent by the server device based on this communication connection. The first terminal device sends the encoded target video to the server device. Correspondingly, the server device receives the target video from the first terminal device, thus realizing the upload of the target video to the content application platform. After the successful sending, the first terminal device can receive a prompt message indicating successful publication.
[0152] Step 612: The server device sends the encoded target video to the second terminal device.
[0153] The server device can send the distributable content (such as including the target video) to the corresponding terminal device based on the set content distribution policy. For example, it sends the encoded target video to the second terminal device. Correspondingly, the second terminal device receives the target video.
[0154] Step 614: The second terminal device decodes and plays the received target video.
[0155] When the second terminal device accesses the content application platform, it can send a content acquisition request to the server device. The server device can send the corresponding content (such as the target video) to the second terminal device based on this request. The second terminal device can first display a viewing entry for the received content on the page corresponding to the content application platform. When detecting a trigger operation for this viewing entry, it specifically displays the content. For example, when the second terminal device detects a click operation by the user on the downloaded target video, it decodes and plays the target video.
[0156] Exemplarily, the target video received by the second terminal device is encoded data, which may include information such as a prediction mode, motion vectors, quantization coefficients, etc. The second terminal device can reconstruct a prediction signal based on this encoded data, restore the residual data through inverse quantization and inverse transformation, add the prediction signal to the restored residual data to obtain the final reconstructed frame of the target video, thus realizing the decoding and restoration of the target video, and then it can play the restored video.
[0157] In the embodiments of this specification, the method of the reference relationship of each frame in the encoding stage is used in the filtering process based on MCTF. For each video frame to be filtered, determine the reference frames that refer to this video frame during encoding, and based on the global correlation between the filtering result obtained by performing MCTF processing on this video frame using the initial filtering function and each reference frame in the reference frame set, adjust the initial filtering function, and then perform MCTF processing on the video frame based on the obtained target filtering function to obtain a better target filtering result. In this way, the first terminal device can obtain a filtering result with a higher degree of correlation with the reference frames, and this filtering result can contain more information of other frames, so the filtered target frame can be made more uniform and smooth, ensuring a better filtering effect of the target frame, and correspondingly improving the video filtering effect. After obtaining the video filtering result using this filtering method, the first terminal device can perform video encoding. In this way, the amount of data to be processed for encoding is less, and the correlation between the data is higher, and the compression complexity of the data is lower. The same reconstruction quality effect can be restored using fewer bits, and more redundant results can be filtered out for the entire video. Therefore, more efficient and fast video encoding can be achieved, ensuring that the amount of video data obtained after encoding is less, the required bit rate is lower, and a better video encoding effect is obtained.
[0158] After video encoding, the first terminal device can transmit the encoded video. Since the amount of video data obtained by encoding is small, the bandwidth overhead required to transmit the video is also low, and the bandwidth overhead required for the server device to receive this video and transmit this video to the second terminal device is also low, and the bandwidth overhead required for the second terminal device to receive this video is also low. In this way, the user consumption experience can be improved, and the operating cost of the information interaction system can be reduced.
[0159] Corresponding to the above method-based MCTF video filtering embodiments, this specification also provides video filtering device embodiments. Figure 7 It is a schematic structural diagram of a video filtering device based on MCTF provided by an embodiment of this specification. As Figure 7 shown, the video encoding device includes:
[0160] A reference frame determination module 702, configured to determine a reference frame set for a target frame in a target video; wherein, the target frame is any video frame in the target video, and the reference frame set includes the reference frames that directly and indirectly refer to the target frame in the case of encoding the target video;
[0161] A filtering optimization module 704, configured to perform motion compensation temporal filtering MCTF processing on the target frame based on an initial filtering function, and adjust the initial filtering function to obtain a target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set;
[0162] The filtering module 706 is configured to perform MCTF processing on a target frame based on a target filtering function to obtain a target filtering result of the target frame.
[0163] Optionally, the reference frame determination module 702 is configured to: for a target frame in a target video, based on the GOP structure of the group of pictures, determine the target GOP where the video frame encoded with reference to the target frame is located and the coding time level where the target frame is located when the target video is encoded; and determine a reference frame set based on the video frames in the target GOP whose coding time levels are higher than the coding time level where the target frame is located.
[0164] Optionally, the reference frame determination module 702 is configured to: determine the reference direction of the target frame in the target video based on the coding prediction type of each video frame in the target video; and select reference frames based on the limited number of reference frames in the reference direction in the target video to obtain a reference frame set.
[0165] Optionally, the filtering optimization module 704 is configured to: calculate the global correlation between the obtained filtering result and each reference frame in the reference frame set, and the distortion information of the filtering result relative to the target frame; and adjust the initial filtering function to obtain a target filtering function with the goal of improving the global correlation and / or reducing the distortion information.
[0166] Optionally, the filtering optimization module 704 is configured to: calculate the global correlation between the obtained filtering result and each reference frame in the reference frame set, and the entropy value of the filtering result; and adjust the initial filtering function to obtain a target filtering function with the goal of improving the global correlation and / or reducing the entropy value.
[0167] Optionally, the video filtering device based on MCTF provided in the embodiments of this specification further includes:
[0168] An integration module, configured to integrate the target frame and the reference frame set into a target set before adjusting the initial filtering function to obtain a target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set;
[0169] A determination module, configured to determine the mutual information between the obtained filtering result and the target set for the target frame; where the mutual information is used to characterize the global correlation between the filtering result and each reference frame in the reference frame set.
[0170] In summary, in the embodiments of this specification, the method of the reference relationship of each frame in the encoding stage is used in the filtering process based on MCTF. For each video frame to be filtered, determine the reference frames that refer to this video frame during encoding, and adjust the initial filtering function based on the global correlation between the filtering result obtained by performing MCTF processing on this video frame using the initial filtering function and each reference frame. Then, perform MCTF processing on the video frame based on the obtained target filtering function to obtain a better target filtering result. In this way, a filtering result with a higher degree of correlation with the reference frames can be obtained. This filtering result can contain more information of other frames, can make the target frame after filtering more uniform and smooth, ensure that the filtering effect of the target frame is better, and correspondingly can improve the overall filtering effect of the video.
[0171] Corresponding to the above embodiments of the video encoding method based on MCTF, this specification also provides embodiments of a video encoding apparatus based on MCTF. Figure 8 It is a schematic structural diagram of a video encoding apparatus provided in an embodiment of this specification. As Figure 8 shown, the video encoding apparatus includes:
[0172] An acquisition module 802, configured to acquire the target filtering results of each video frame in the target video, where the target filtering result of each video frame is obtained based on the above video filtering method;
[0173] An encoding module 804, configured to encode the target video based on the target filtering results of each video frame in the target video.
[0174] The above is a schematic solution of the video encoding apparatus in this embodiment. It should be noted that the technical solution of this video encoding apparatus and the technical solution of the corresponding video encoding method above belong to the same concept. For the details not described in the technical solution of the video encoding apparatus, reference can be made to the description of the technical solution of the video encoding method above.
[0175] Figure 9 It is a structural block diagram of a computing device provided in an embodiment of this specification. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data. Among them, the processor 920 is configured to execute a computer program / instructions, and when the computer program / instructions are executed by the processor, the steps in the above video filtering method or the steps in the video encoding method are implemented.
[0176] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0177] In one embodiment of the present specification, the above components of the computing device 900, as well as Figure 9 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed. The computing device 900 may be a mobile or stationary server or server cluster, or some stationary or mobile computing devices with large data storage capabilities. For example, mobile computers or mobile computing devices may include tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc., and stationary computing devices may include desktop computers or personal computers (PCs).
[0178] For the embodiment of the computing device, since it is basically similar to the above embodiments of the video filtering method and video coding method, the description is relatively simple. For related parts, refer to the partial description of the embodiments of the video filtering method and video coding method.
[0179] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the above video filtering method or the steps of the video encoding method. The computer programs / instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0180] An embodiment of this specification also provides a computer program product including computer programs / instructions, which, when executed in a processor, implement the steps in the above video filtering method or the steps in the video encoding method.
[0181] For the embodiments of the computer-readable storage medium and the computer program product, since they are basically similar to the embodiment of the video encoding method, the description is relatively simple. For the relevant parts, please refer to the partial description of the embodiments of the video filtering method and the video encoding method.
[0182] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0183] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0184] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0185] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and variations can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can understand and utilize the present specification well. The present specification is only limited by the claims and their full scope and equivalents.
Claims
1. A video filtering method based on MCTF, characterized in that: include: Determine a reference frame set for a target frame in a target video; wherein the target frame is any video frame in the target video, and the reference frame set includes reference frames that the target video directly and indirectly references to the target frame when encoding; Performing motion compensated time domain filtering MCTF processing on the target frame based on the initial filtering function, and adjusting the initial filtering function to obtain a target filtering function based on the global correlation between the obtained filtering result and each reference frame in the reference frame set; Based on the target filtering function, MCTF processing is performed on the target frame to obtain a target filtering result of the target frame.
2. The method according to claim 1, characterized in that The step of determining a reference frame set for a target frame in a target video includes: For a target frame in a target video, based on a GOP structure, determining, when the target video is encoded, a target GOP where a video frame encoded with reference to the target frame is located, and a coding time level where the target frame is located; A reference frame set is determined based on video frames in the target GOP whose coding temporal level is higher than the coding temporal level where the target frame is located.
3. The method according to claim 1, characterized in that The step of determining a reference frame set for a target frame in a target video includes: Determining a reference direction of a target frame in the target video based on a coding prediction type of each video frame in the target video; In the referenced direction in the target video, reference frames are selected based on a limited number of reference frames to obtain a reference frame set.
4. The method according to claim 1, characterized in that The adjusting the initial filter function to obtain a target filter function based on the obtained filter result and the global correlation of the reference frame set includes: The calculated global correlation between the filtering result and each reference frame in the reference frame set, and the distortion information of the filtering result relative to the target frame; With the goal of improving the global correlation and / or reducing the distortion information, the initial filter function is adjusted to obtain a target filter function.
5. The method according to any one of claims 1 to 4, characterized in that: The adjusting the initial filter function to obtain a target filter function based on the obtained filter result and the global correlation of the reference frame set includes: The calculated global correlation between the filtering result and each reference frame in the reference frame set, and the entropy value of the filtering result; With the goal of improving the global correlation and / or reducing the entropy value, the initial filter function is adjusted to obtain a target filter function.
6. The method according to any one of claims 1 to 4, characterized in that: Before adjusting the initial filter function to obtain the target filter function based on the global correlation between the obtained filter result and each reference frame in the reference frame set, the method further includes: integrating the target frame and the reference frame set into a target set; For the target frame, the mutual information between the obtained filtering result and the target set is determined; wherein the mutual information is used to characterize the global correlation between the filtering result and each reference frame in the reference frame set.
7. A video filtering device based on MCTF, characterized in that: include A reference frame determination module, configured to determine a reference frame set for a target frame in a target video; wherein the target frame is any video frame in the target video, and the reference frame set includes reference frames that the target video directly and indirectly references to the target frame when encoding; A filtering optimization module, configured to perform motion compensated time domain filtering MCTF processing on the target frame based on an initial filtering function, and adjust the initial filtering function to obtain a target filtering function based on a global correlation between the obtained filtering result and each reference frame in the reference frame set; The filtering module is used to perform MCTF processing on the target frame based on the target filtering function to obtain a target filtering result of the target frame.
8. A computing device, characterized in that include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps in any one of the methods described in claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: A computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps in any one of the methods described in claims 1 to 6 are implemented.
10. A computer program product, characterized in that The method comprises a computer program / instruction, and when the computer program / instruction is executed in a processor, the steps in any one of the methods of claims 1 to 6 are implemented.