Video encoding method and apparatus, device, medium and program product
Patent Information
- Application Number
- PCT/CN2026/083882
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-17
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026083882_24092026_PF_FP_ABST
Abstract
Description
Video encoding methods, apparatus, equipment, media and software products
[0001] This application claims priority to Chinese Patent Application No. 2025103363819, filed on March 19, 2025, entitled "Video Coding Method, Apparatus, Device, Media and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of encoders, and more particularly to a video encoding method, apparatus, device, medium, and program product. Background Technology
[0003] An encoder is a program or device used to compress video, reducing its size and bandwidth. A group of pictures (GOP) is a collection of multiple video frames. An encoder can divide video frames into at least one group of pictures and encode the video frames in each group sequentially.
[0004] In related technologies, the encoder determines the optimal image group structure by minimizing the overall cost of video frames in the pre-analysis queue. Based on this optimal image group structure, the video frames are divided into image groups, and this optimal image group structure needs to be continuously updated.
[0005] However, the related techniques require the encoder to traverse every possible image group structure, resulting in high algorithm complexity and poor encoder performance. Summary of the Invention
[0006] This application provides a video encoding method, apparatus, device, medium, and program product. The technical solution is as follows:
[0007] On the one hand, a video encoding method is provided, the method comprising:
[0008] Obtain a video frame sequence with a preset number of frames;
[0009] Based on the video features of the video frame sequence, the last frame is determined from the video frame sequence;
[0010] The first frame to the last frame in the video frame sequence are determined as an image group;
[0011] Each video frame in the image group is encoded to obtain an encoded bitstream.
[0012] On the other hand, a video encoding apparatus is provided, the apparatus comprising:
[0013] The acquisition module is used to acquire a video frame sequence with a preset number of frames from the video to be encoded;
[0014] The processing module is used to determine the last frame from the video frame sequence based on the video features of the video frame sequence;
[0015] The determining module is used to determine the first frame to the last frame in the video frame sequence as an image group;
[0016] The encoding module is used to encode each video frame in the image group to obtain an encoded bitstream.
[0017] On the other hand, an encoder is provided, the encoder comprising: an encoding unit storing a computer program, the computer program being loaded and executed by the encoding unit to implement the video encoding method as described above.
[0018] On the other hand, a computer device is provided, the computer device comprising: a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the video encoding method as described above.
[0019] On the other hand, a computer-readable storage medium is provided that stores a computer program, which is loaded and executed by a processor to implement the video encoding method described above.
[0020] On the other hand, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium, wherein a processor retrieves the computer instructions from the computer-readable storage medium, causing the processor to load and execute them to implement the video encoding method as described above. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a structural block diagram of a computer system provided in an exemplary embodiment of this application;
[0023] Figure 2 is a schematic diagram of a video encoding method provided in an exemplary embodiment of this application;
[0024] Figure 3 is a flowchart of a video encoding method provided in an exemplary embodiment of this application;
[0025] Figure 4 is a flowchart of a video encoding method provided in an exemplary embodiment of this application;
[0026] Figure 5 is a flowchart of a video encoding method provided in an exemplary embodiment of this application;
[0027] Figure 6 is a schematic diagram of a video encoding method provided in an exemplary embodiment of this application;
[0028] Figure 7 is a block diagram of a video encoding apparatus provided in an exemplary embodiment of this application;
[0029] Figure 8 is a structural block diagram of an encoder provided in an exemplary embodiment of this application;
[0030] Figure 9 is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0033] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0035] It should be noted that this application may display a prompt interface, pop-up window, or output voice prompts before and during the collection of user-related data (e.g., videos, videos to be encoded, video frames). These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their relevant data is being collected. This ensures that the application only begins the steps related to collecting user-related data after receiving confirmation from the user regarding the prompt interface or pop-up window; otherwise, if no confirmation is received from the user, the steps to collect user-related data end, i.e., no user-related data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of relevant user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0036] First, let me briefly introduce the terms used in the embodiments of this application:
[0037] An encoder is a program or device used to compress video, reducing its size and bandwidth. An encoder divides video frames into at least one group of images, and then encodes each frame in each group sequentially, resulting in a encoded bitstream.
[0038] A Group of Pictures (GOP) is a collection of at least one video frame; it is a sequence of video frames. The size of a GOP is the number of frames in the group. For example, a GOP with a size of 16 means that the group contains 16 frames.
[0039] In this embodiment, the image group may include: a current image group and an image group to be encoded. The current image group is an image group that includes a video frame sequence with a preset number of frames. For example, if the preset number of frames is 32, then the current image group is the first 32 consecutive frames in the video to be encoded. The image group to be encoded is the final image group that needs to be encoded, determined by further dividing the video frame sequence. For example, if the number of frames in the image group to be encoded is 16, then the image group to be encoded is the first 16 consecutive frames in the video frame sequence.
[0040] The first candidate frame, also known as the maximum end (max_index) frame, has a frame number (or simply number) equal to the upper limit of the number of frames in the image group to be encoded. This can also be understood as the number equal to the maximum number of frames supported by the image group. Specifically, the number of the first candidate frame in the video frame sequence indicates the upper limit of the number of frames in the image group to be encoded. For example, if the video frame sequence has 32 frames and the first candidate frame is the 28th frame, then the upper limit of the number of frames supported by the image group to be encoded is 28, meaning the maximum size of the image group to be encoded is 28. In this embodiment, the first candidate frame is one of the following video frames in the video frame sequence: a video frame that meets the resolution requirement, or the end frame of a scene before a scene change.
[0041] Video frames that meet the resolution criteria are also understood as high-quality frames rich in detail and texture information. Whether a frame meets the resolution criteria is determined based on its intra-frame mode complexity. A high intra-frame mode complexity indicates that the video frame contains more detail and texture information, making it a high-quality frame. When high-quality frames are encoded with high quality, they can provide a better reference for other blurry frames, improving the overall encoding quality of the video being encoded. For example,
[0042] When a video frame sequence involves two or more scenes, the last frame of the scene before the scene switch is the last frame of the first scene. Specifically, it is the last frame of the video frame corresponding to scene 1 before switching from scene 1 to scene 2. For example, if the video frame sequence includes 32 frames, with frames 1 to 16 involving scene 1 and frames 17 to 32 involving scene 2, then the last frame of the scene before the scene switch is frame 16.
[0043] In this invention, scene switching can be detected using any method known in the art. For example, it can be determined by calculating the difference in brightness histograms between adjacent frames: if the difference exceeds a preset threshold, it is determined to be a scene switching. Alternatively, it can be detected by utilizing abrupt changes in inter-frame prediction complexity: if the prediction complexity (SATD) with reference to adjacent frames is significantly higher than the normal range, it indicates that a scene switching has occurred. Furthermore, it can be determined by statistically analyzing the proportion of intra-frame mode blocks: if the proportion of pre-analysis blocks using intra-frame mode in a certain frame increases sharply, it also indicates that a scene switching has occurred.
[0044] Last frame: Used to indicate the optimized frame number of the group of images to be encoded, which is less than or equal to the maximum number of frames supported by the group. Specifically, the sequence number of the last frame in the video frame sequence is used to indicate the optimized frame number of the group of images to be encoded. The last frame is determined from the first candidate frame and at least one preset frame.
[0045] At least one preset frame is a preset frame in the video frame sequence that satisfies the image group structure conditions and is located before the first candidate frame. The image group structure can be characterized by the size (number of frames) of the image group to be encoded, and the image group structure conditions refer to the conditions that the size of the image group to be encoded must satisfy. In this embodiment, the size of the image group to be encoded is set to a first size, a second size, or a third size. Then, at least one preset frame is a preset frame in the video frame sequence that is located before the first candidate frame and makes the size of the image group to be encoded the first size, the second size, or the third size. The sequence number of each preset frame corresponds to a predefined image group size candidate value. Each preset frame is located before the first candidate frame in the video frame sequence, and its sequence number is equal to a predefined image group size candidate value. Different preset frames are set with different image group size candidate values. In one example, the predefined image group size candidate value is, for example, 16, 8, or 4. At least one preset frame includes at least one of the following: a first preset frame, a second preset frame, and a third preset frame. In one example, the first preset frame is frame 16, the second preset frame is frame 8, and the third preset frame is frame 4. In other words, the preset sequence numbers of the first, second, and third preset frames are 16, 8, and 4, respectively.
[0046] For example, if the video frame sequence has 32 frames, and the first candidate frame is the 28th frame in the video frame sequence, then the maximum number of frames supported by the image group is 28. For example, if the last frame can be the first candidate frame, i.e., the 28th frame, then the image group to be encoded consists of frames 1 to 28 of the video frame sequence, and the size of the image group is 28. As another example, if the last frame can be the first preset frame, i.e., the 16th frame, then the image group to be encoded consists of frames 1 to 16 of the video frame sequence, and the size of the image group is 16. As yet another example, if the last frame can be the second preset frame, i.e., the 8th frame, then the image group to be encoded consists of frames 1 to 8 of the video frame sequence, and the size of the image group is 8. As yet another example, if the last frame can be the third preset frame, i.e., the 4th frame, then the image group to be encoded consists of frames 1 to 4 of the video frame sequence, and the size of the image group is 4.
[0047] Prediction complexity: In the field of video coding, the prediction complexity of a current video frame measures the ease with which the pixel value at each position of that current video frame is predicted. In some embodiments, prediction complexity reflects the spatial and temporal variation characteristics of the current video frame and its correlation with other video frames (e.g., adjacent video frames). For the pixel value at each position of the current video frame, prediction can be performed using at least one of intra-frame mode and inter-frame mode. Accordingly, the prediction complexity in this embodiment can include the prediction complexity corresponding to the intra-frame mode (i.e., intra-frame mode complexity) and the prediction complexity corresponding to the inter-frame mode (i.e., inter-frame mode complexity).
[0048] Intra-frame mode: Also known as intra-prediction mode or intra-coding mode, it uses the pixel values in the current video frame and / or the relationships between pre-analysis blocks for encoding. Specifically, intra-frame mode predicts the pixel value at each position in each pre-analysis block using only information from the current video frame itself, without relying on other video frames. The prediction complexity of intra-frame mode is called intra-frame mode complexity. It should also be noted that intra-frame mode complexity refers to the prediction complexity when using intra-frame mode to predict the pixel value at each position in each pre-analysis block; that is, each pre-analysis block of the video frame uses intra-frame mode.
[0049] Inter-frame mode, also known as inter-frame prediction mode or inter-frame coding mode, utilizes the temporal correlation of video frames for encoding. Specifically, inter-frame mode uses a video frame as a reference frame and, through a motion search algorithm, determines the best-matching reference block in the reference frame for each pre-analyzed block in the current video frame. It then predicts the pixel value at each position of each pre-analyzed block and calculates the motion vector (MV) for each pre-analyzed block, which indicates the positional changes of the reference block in the reference frame. In some examples, the motion vector is also called the displacement vector. The prediction complexity of inter-frame mode is called inter-frame mode complexity.
[0050] A pre-analysis block is a pre-divided block of pixels within a video frame; that is, a video frame includes at least one pre-analysis block. Each pre-analysis block can use either intra-frame or inter-frame mode. When using inter-frame mode, each pre-analysis block has a corresponding motion vector.
[0051] A preanalysis block is an N×N block. N can take at least one of the following values: 16, 8, or 4. The value of N is related to at least one of the following factors: prediction accuracy, computational load, signal-to-noise ratio (SNR), video coding standard, and computer equipment performance. To achieve higher prediction accuracy, smaller preanalysis blocks can capture finer motions. However, smaller preanalysis blocks may increase computational load; therefore, a balance needs to be struck between prediction accuracy and computational load to determine the preanalysis block size. When the SNR is higher than a threshold, smaller preanalysis blocks can be used to capture details in the video frame; when the SNR is lower than a threshold, larger preanalysis blocks can be used to smooth noise. Video coding standards can specify the size of the preanalysis block and its partitioning method. Higher computer equipment performance allows for the processing of smaller preanalysis blocks. Therefore, the size of the preanalysis block can be set according to actual technical needs.
[0052] In this embodiment, we take an 8×8 block as an example for each pre-analysis block. The number of 8×8 blocks in a video frame is calculated as follows: scale the video frame to its original preset size, for example, 1 / 4 (scale the width and height to 1 / 2 of their original size), and divide the scaled video frame by 64 to obtain the number of 8×8 blocks. Scaling can be achieved through downsampling. 64 is the pixel size of an 8×8 block. For example, if the video frame size is 1920×1080, divide by 4, then divide by 64 to obtain the number of 8×8 blocks.
[0053] How to calculate the intra-frame complexity of frame A:
[0054] Divide the A-frame into at least one pre-analysis block.
[0055] Traverse at least one pre-analysis block. For the current pre-analysis block in at least one pre-analysis block, refer to the 35 intra prediction modes in the High Efficiency Video Coding (HEVC) coding standard to predict the pixel value at each position of the current pre-analysis block and obtain the predicted pixel value.
[0056] Subtract the predicted pixel value from the original pixel value at each position of the current pre-analysis block to obtain the residual value at each position. Perform a Hadamard Transform (HT) on the residual values and sum them (satd) to obtain the intra-mode complexity of the current pre-analysis block. Here, performing the Hadamard Transform and summing (stad) involves summing the absolute values of the residual values after the Hadamard Transform. Repeat this process until the intra-mode complexity of each pre-analysis block of at least one pre-analysis block is obtained. Accumulate the intra-mode complexity of each pre-analysis block of at least one pre-analysis block to obtain the intra-mode complexity of frame A using intra mode.
[0057] In some embodiments, dividing an A-frame into at least one pre-analysis block can be based on the size of the pre-analysis block, dividing the A-frame into at least one pre-analysis block of the same size in a left-to-right, top-to-bottom order. Optionally, dividing an A-frame into at least one pre-analysis block can be done by first scaling the A-frame to a preset ratio, for example, 1 / 4 (scaling the width and height to 1 / 2 of their original values), and then dividing the scaled A-frame into at least one pre-analysis block. In other words, the A-frame can first be downsampled to a preset ratio, for example, 1 / 4 (scaling the width and height to 1 / 2 of their original values), and then the scaled A-frame can be divided into at least one pre-analysis block.
[0058] Frame A is any frame in a video frame sequence. The algorithm in this embodiment is applicable to calculating the intra-frame mode complexity of any frame in a video frame sequence, such as the intra-frame mode complexity of the first candidate frame, the intra-frame mode complexity of the first preset frame, the intra-frame mode complexity of the second preset frame, and the intra-frame mode complexity of each video frame in the video frame sequence, etc.
[0059] How to calculate the prediction complexity of frame A compared to frame B:
[0060] Divide the A-frame into at least one pre-analysis block.
[0061] Using B-frame as reference frame, traverse at least one pre-analysis block. For the current pre-analysis block in at least one pre-analysis block, determine the corresponding reference block in B-frame, determine the inter-frame mode complexity of the current pre-analysis block, and record the motion vector (mv) of the current pre-analysis block.
[0062] If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, then the current pre-analysis block is determined to use inter-frame mode, and its prediction complexity is equal to the inter-frame mode complexity. Otherwise, the current pre-analysis block is determined to use intra-frame mode, and its prediction complexity is equal to the intra-frame mode complexity. This process is repeated until the prediction complexity of each pre-analysis block of at least one pre-analysis block is determined. In this embodiment, the prediction complexity of each pre-analysis block can be understood as the smaller of the intra-frame mode complexity and the inter-frame mode complexity. The prediction complexity of each pre-analysis block of at least one pre-analysis block is accumulated to obtain the prediction complexity of frame A compared to frame B.
[0063] The calculation method for determining the inter-frame mode complexity of the current pre-analysis block is as follows: Based on the reference block corresponding to the current pre-analysis block in frame A, the pixel value at each position of the current pre-analysis block is predicted using inter-frame mode. The predicted pixel value at each position is then subtracted from the original pixel value of the reference block to obtain the residual value at each position. A Hadamard transform is performed on the residual values at each position, and the summation (stad) yields the inter-frame mode complexity of the current pre-analysis block. The Hadamard transform and summation (stad) involves summing the absolute values of the residuals after the Hadamard transform. In some possible implementations, the motion vector (mv) of the current pre-analysis block can be multiplied by a preset coefficient and added to the inter-frame mode complexity to obtain the final inter-frame mode complexity.
[0064] In some embodiments, dividing the A-frame into at least one pre-analysis block can be based on the size of the pre-analysis block, dividing the A-frame into at least one pre-analysis block of the same size in a left-to-right, top-to-bottom order. Optionally, dividing the A-frame into at least one pre-analysis block can be done by first scaling the A-frame to its original preset size, for example, 1 / 4 (scaling the width and height to 1 / 2 of their original sizes), scaling the B-frame to its original preset size as well, for example, 1 / 4 (scaling the width and height to 1 / 2 of their original sizes), and then dividing the scaled A-frame into at least one pre-analysis block.
[0065] Frame A and Frame B are any two frames in a video frame sequence. It should also be noted that the algorithm in this embodiment is applicable to calculating the prediction complexity of any A frame compared to any B frame in a video frame sequence. For example, the following refers to: the prediction complexity of the first candidate frame compared to the starting frame, the prediction complexity of the first candidate frame compared to the previous frame, the prediction complexity of the first preset frame compared to the starting frame, and the prediction complexity of the second preset frame compared to the starting frame.
[0066] How to calculate the prediction complexity of frame A compared to frames B and C:
[0067] Divide the A-frame into at least one pre-analysis block.
[0068] Using only B-frames as reference frames, the at least one pre-analysis block is traversed. For the current pre-analysis block in the at least one pre-analysis block, the corresponding reference block in the B-frame is determined, the inter-frame mode complexity of the current pre-analysis block is determined, and the motion vector (mv) of the current pre-analysis block is recorded. If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, it is determined that the current pre-analysis block adopts inter-frame mode and the prediction complexity of the current pre-analysis block is inter-frame mode complexity; otherwise, it is determined that the current pre-analysis block adopts intra-frame mode and the prediction complexity of the current pre-analysis block is intra-frame mode complexity. This process is repeated until the first prediction complexity of each pre-analysis block of the at least one pre-analysis block is determined.
[0069] Using only the C-frame as a reference frame, traverse at least one pre-analysis block. For the current pre-analysis block in at least one pre-analysis block, determine the corresponding reference block in the C-frame, determine the inter-frame mode complexity of the current pre-analysis block, and record the motion vector (mv) of the current pre-analysis block. If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, determine that the current pre-analysis block adopts inter-frame mode and the prediction complexity of the current pre-analysis block is inter-frame mode complexity; otherwise, determine that the current pre-analysis block adopts intra-frame mode and the prediction complexity of the current pre-analysis block is intra-frame mode complexity. Repeat this process until the second prediction complexity of each pre-analysis block of at least one pre-analysis block is determined.
[0070] Simultaneously, using B-frames and C-frames as reference frames, at least one pre-analysis block is traversed. For the current pre-analysis block in the at least one pre-analysis block, the corresponding reference block in the B-frame is determined, and the same steps as in the aforementioned embodiment using B-frames as reference frames are performed to obtain a first prediction complexity for each pre-analysis block of the at least one pre-analysis block. Then, the corresponding reference block in the C-frame is determined, and the same steps as in the aforementioned embodiment using C-frames as reference frames are performed to obtain a second prediction complexity for each pre-analysis block of the at least one pre-analysis block. The first and second prediction complexities of each pre-analysis block of the at least one pre-analysis block are weighted and summed to obtain a third prediction complexity for each pre-analysis block of the at least one pre-analysis block. The weights of the first and second prediction complexities can be pre-configured, for example, 0.5, 0.5, or 0.4, 0.6, respectively.
[0071] In this embodiment, the reference block corresponding to the current pre-analysis block in frame B is the same as or different from the reference block when frame B is the only reference frame. The reference block corresponding to the current pre-analysis block in frame C is the same as or different from the reference block when frame C is the only reference frame. When both frame B and frame C are used as reference frames, each pre-analysis block has two reference blocks.
[0072] The minimum prediction complexity among the first, second, and third prediction complexities of each pre-analysis block is determined as the final prediction complexity for each pre-analysis block. Simultaneously, the reference block (which can be one or two) corresponding to the minimum complexity is determined as the final reference block, and the motion vector corresponding to the minimum complexity is determined as the final motion vector. The final prediction complexities of each pre-analysis block are summed to obtain the prediction complexity of frame A compared to frames B and C.
[0073] In some embodiments, dividing the A-frame into at least one pre-analysis block can be based on the size of the pre-analysis block, dividing the A-frame into at least one pre-analysis block of the same size in a left-to-right, top-to-bottom order. Optionally, dividing the A-frame into at least one pre-analysis block can be done by first scaling the A-frame to its original preset size, for example, 1 / 4 (scaling the width and height to 1 / 2 of their original sizes), scaling the B-frame and C-frame to their original preset sizes, 1 / 4 (scaling the width and height to 1 / 2 of their original sizes), and then dividing the scaled A-frame into at least one pre-analysis block.
[0074] Frame A, frame B, and frame C are any three frames in a video frame sequence. It should also be noted that the algorithm in this embodiment is applicable to calculating the prediction complexity of any frame A compared to any frame B and any frame C in a video frame sequence, for example, the prediction complexity of the first preset frame compared to the starting frame and the first candidate frame, as discussed below.
[0075] Figure 1 is a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system 100 can be implemented as a system architecture for a video encoding method. The computer system 100 includes: a terminal 120, an encoder 130, and a server 140.
[0076] Terminal 120 can be at least one of the following: mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, PC (Personal Computer), unmanned reservation terminal, smart home appliance, smart voice interaction device, or unmanned vending terminal. A client application for the target application can be installed and run on terminal 120. This target application can be a video encoding / decoding application or an application that provides video encoding / decoding functions; this application embodiment does not limit the specific form of the target application. Furthermore, this application embodiment does not limit the form of the target application, including but not limited to Apps (Applications), mini-programs, etc., installed on terminal 120, and can also be in web page form.
[0077] The encoder 130 can be an electronic device dedicated to video encoding / decoding, and can be a standalone device or built into a professional video processing device. The video processing device includes, but is not limited to, at least one of the following: video editing workstation, high-definition camera / webcam, professional server, broadcast television transmission platform, and video transcoding platform. The encoder 130 includes an encoding unit for implementing video encoding. The encoding unit can be implemented as at least one of the following: chip, processor, hardware circuit, and logic circuit; this application embodiment does not limit this.
[0078] Server 140 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Server 140 can be the backend server for the target application of terminal 120 or the backend server for encoder 130, used to provide backend services. Optionally, server 140 can also be implemented as a node in a blockchain system.
[0079] Terminal 120, encoder 130 and server 140 communicate with each other via wired or wireless network.
[0080] The video encoding method provided in this application can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Taking the implementation environment shown in Figure 1 as an example, the video encoding method can be executed by a terminal 120, such as a client of a target application installed and running on the terminal 120. Alternatively, the video encoding method can be executed by an encoder 130, a server 140, or through interactive cooperation between the terminal 120 and the server 140, or through interactive cooperation between the terminal 120 and the encoder 130, or through interactive cooperation between the encoder 130 and the server 140, or through interactive cooperation between all three: the terminal 120, the encoder 130, and the server 140. This application does not limit the scope of the implementation.
[0081] Those skilled in the art will understand that the number of terminals 120, encoders 130, and servers 140 can be more or less. For example, there may be only one terminal 120, encoder 130, and server 140, or dozens, hundreds, or more terminals 120, encoders 130, and servers 140. This application does not limit the number or type of terminals 120, encoders 130, and servers 140.
[0082] In related technologies, video coding and image grouping methods include the following:
[0083] Method 1: Divide the input video into at least one image group with a fixed number of frames, and encode each video frame in each image group sequentially. The fixed number of frames is a pre-configured upper limit (gop_size). This method cannot adaptively divide the image groups based on the video content.
[0084] Method 2: Use the Viterbi algorithm to find the optimal picture group structure by minimizing the total cost of video frames in the pre-analysis queue. The optimal picture group structure is defined as the one that minimizes the cumulative cost of all video frames in the pre-analysis queue. Minimizing the cumulative cost means that temporal correlation is fully utilized, which minimizes the residual value in the formal encoding process, thereby reducing the bit consumption of encoding.
[0085] Assuming there are a total of n frames in the pre-analysis queue, the Viterbi algorithm is used to calculate the cumulative minimum cost, which involves splitting the overall minimum cost into three parts: 1. The minimum cost of the first i frames; 2. The cost of inserting B frames from frame i to frame (n-1); 3. The cost of using P frames in frame n. The optimal image group structure is found through dynamic programming. The image group structure is represented by a list of frame types, with each image group consisting of the next B frame from the current P frame to the next P frame. The steps are briefly described below:
[0086] Suppose we need to determine the optimal structure for the first i frames, where 2 <= i <= n.
[0087] 1. The cost of initializing the optimal frame type list is 1LL<<62.
[0088] 2. Traverse j within the range of 2 to i-1 to find the list of frame types with the lowest cost.
[0089] 2.1) Construct a list of candidate frame types.
[0090] 2.1.1) Copy the optimal frame type of the first j frames to the candidate frame type list;
[0091] 2.1.2) Set the frame type from j+1 to i-1 to B-frame;
[0092] 2.1.3) Set the frame type of the i-th frame to P frame.
[0093] 2.2) The cost of calculating the list of candidate frame types.
[0094] 2.2.1) Calculate the cost of the next P-frame with the current P-frame as a reference, according to the order of the pre-analysis queue from front to back, and add the cost to the candidate frame type list;
[0095] 2.2.2) Calculate the cost of B-frames between two P-frames and add it to the cost of the candidate frame type list:
[0096] 2.2.2.1) If there is only one B-frame, calculate the cost of that B-frame with reference to the current P-frame and the next P-frame;
[0097] 2.2.2.2) If there are multiple B-frames, then:
[0098] a) calculating the cost when the intermediate frame between two P-frames uses the current P-frame and the next P-frame as references. The index of the intermediate frame in the pre-analysis queue is middle=cur_p+(next_p-cur_p) / 2; wherein, cur_p is the index of the current P-frame, and next_p is the index of the next P-frame;
[0099] b) for a B-frame between the current P-frame and the intermediate frame, calculating the cost when the B-frame uses the current P-frame and the intermediate frame as references;
[0100] c) for a B-frame between the next P-frame and the intermediate frame, calculating the cost when the B-frame uses the next P-frame and the intermediate frame as references.
[0101] 2.3) if the cost of the current candidate frame type list is less than the optimal cost, updating the optimal cost to the current cost, and recording the current candidate frame type list.
[0102] Mode 3: dividing group of pictures according to the similarity of video frames. The steps are briefly described as follows:
[0103] 1. Traversing from the 0-th frame to the (n-2)-th frame in the pre-analysis queue, and setting a frame type for each frame.
[0104] Assume the currently analyzed frame is the i-th frame:
[0105] 1.1) calculating the cost of the (i+2)-th frame relative to the i-th frame, denoted as cost2p1; if the proportion of pre-analysis blocks in intra mode adopted by the (i+2)-th frame compared with the i-th frame exceeds a threshold, setting the (i+1)-th frame and the i-th frame as P-frames, ending the current loop, and then returning to step 1 to continue judging frame types starting from the (i+2)-th frame; otherwise, proceeding to 1.2);
[0106] 1.2) calculating the cost when the (i+1)-th frame uses the i-th frame and the (i+2)-th frame as references, denoted as cost1b1. calculating the cost of the (i+1)-th frame compared with the i-th frame, denoted as cost1p0; calculating the cost of the (i+2)-th frame compared with the (i+1)-th frame, denoted as cost2p0. If cost1p0+cost2p0<cost1b1+cost2p1, setting the (i+1)-th frame as a P-frame, ending the current loop, and then returning to step 1 to continue judging frame types starting from the (i+1)-th frame; otherwise, proceeding to 1.3);
[0107] 1.3) setting the (i+1)-th frame as a B-frame, traversing from the (i+2)-th frame to the end of the pre-analysis queue, and judging the frame type of each frame. Denoting the current frame as j;
[0108] 1.3.1) Calculate the cost of frame j+1 compared to frame i; if the cost exceeds the threshold, or the proportion of intra-pre-analysis blocks used exceeds the threshold, then set frame j as P-frame and end the current loop. Then return to step 1 and continue to determine the frame type starting from frame j.
[0109] 1.3.2) Otherwise, set the j-th frame as frame B and repeat 1.3).
[0110] For methods two and three, it is necessary to traverse every possible image group structure, which results in high algorithm complexity and limited overall performance. Table 1 shows the results on the MSU2022 test set, with the maximum image group size set to 16, adaptive image grouping disabled with x265, using the algorithm of method three as test 1 and the algorithm of method two as test 2.
[0111] The first row of Table 1 shows the results of Test 1, and the second row shows the results of Test 2. A negative speed value in Table 1 indicates a speed reduction. For the video quality assessment metrics in Table 1—Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Visual Multimethod Assessment Fusion (VMAF)—negative values indicate gain, while positive values indicate loss. Table 1 shows that the algorithm of Method 3 experiences an overall speed reduction of 4.95%, and a significant loss in all video quality assessment metrics; the algorithm of Method 2 experiences an overall speed reduction of 6.84%, with some loss in video quality assessment metrics, but the overall gain is not significant.
[0112] Table 1
[0113] This application provides a video encoding method that adaptively divides a group of images to be encoded based on video features, and then encodes each video frame in that group. Figure 2 is a schematic diagram of a video encoding method provided in an exemplary embodiment of this application. The method is executed by a computer device, which includes the terminal 120 and / or encoder 130 and / or server 140 of Figure 1. In this embodiment, taking the execution by encoder 130 as an example, the steps are briefly described as follows:
[0114] Obtain a video frame sequence with a preset number of frames from the video to be encoded 10 to obtain a video frame sequence 20. The preset number of frames is a pre-configured number of frames in the video frame sequence 20. For example, if the preset number of frames is 32, then the video frame sequence is the first 32 frames of the video to be encoded 10. Based on the video features of the video frame sequence 20, determine the last frame from the video frame sequence 20. For example, the last frame is the 4th frame in the video frame sequence 20. Determine the image group 30 to be encoded from the start frame to the end frame in the video frame sequence 20. Encode each video frame in the image group 30 to be encoded to obtain an encoded bitstream 40. And output the encoded bitstream 40.
[0115] It should also be noted that the steps in this embodiment can be performed iteratively. For example, each video frame in the image group 30 to be encoded is derived from the video frame sequence 20, resulting in a derived video frame sequence. After encoding each video frame in the image group 30 to obtain the encoded bitstream 40, video frames in the video 10 to be encoded are acquired to supplement the derived video frame sequence into a new video frame sequence 20 with a preset number of frames. The step of determining the last frame from the video frame sequence 20 based on the video features of the video frame sequence 20 is then re-executed, thereby continuing iterative encoding until the entire video 10 to be encoded has been encoded. Detailed examples are provided below.
[0116] Figure 3 is a flowchart of a video encoding method provided in an exemplary embodiment of this application. The method is executed by a computer device, which includes the terminal 120 and / or encoder 130 and / or server 140 of Figure 1. Taking execution by encoder 130 as an example, the method includes at least some of the steps 220, 240, 260, and 280:
[0117] Step 220: Obtain a video frame sequence with a preset number of frames.
[0118] In some embodiments, this application can obtain a video frame sequence with a preset number of frames from the video to be encoded to obtain the current image group. In other words, the video frame sequence can be referred to as the current image group.
[0119] Here, the video is the video to be encoded, that is, the video that needs to be encoded. The video to be encoded includes at least one video frame, which together form a video frame sequence. That is, at least one video frame in the video to be encoded is arranged in order, for example, in chronological order.
[0120] A video frame sequence is a group of images consisting of a preset number of video frames. These preset frame sequences are continuous and arranged chronologically within the video to be encoded. The preset frame number is a pre-configured number of frames in the video frame sequence, and its maximum value is the upper limit of the pre-configured frame number, which can be set according to actual technical needs. For example, the maximum preset frame number can be set to 32.
[0121] It should also be noted that when determining the video frame sequence, if the number of remaining uncoded frames in the video to be encoded is greater than or equal to this maximum value, then the preset number of frames corresponding to the video frame sequence is this maximum value. At the end of the video to be encoded, if the number of remaining uncoded frames in the video to be encoded is less than this maximum value, then the preset number of frames corresponding to the video frame sequence is the number of remaining uncoded frames.
[0122] The video frame sequence can be determined directly based on the video to be encoded. For example, the video to be encoded is acquired, and its frames are read frame by frame to obtain a video frame sequence with a preset number of frames. Alternatively, the video to be encoded can be pre-divided into at least one video frame, at least one video frame is acquired, and a video frame sequence with a preset number of frames is extracted from that at least one video frame to obtain the video frame sequence.
[0123] The video frame sequence can also be obtained based on a pre-analysis queue. The pre-analysis queue is a fixed-size queue used to store the video frames to be analyzed frame by frame. The size of the pre-analysis queue can be set according to actual technical needs. For example, the size of the pre-analysis queue can be set to at least one of the following: 48, 64, or 80, that is, the pre-analysis queue is used to store 48 video frames, 64 video frames, or 80 video frames. In some embodiments, the size of the pre-analysis queue is consistent with the maximum number of frames pre-configured by the encoder.
[0124] For example, the process involves reading the first number of video frames from the video to be encoded; storing this first number of video frames in a pre-analysis queue; and determining the preset number of video frames in the pre-analysis queue as the final video frame sequence. The maximum value of the first number of frames is the pre-configured maximum number of frames to read. For instance, if the size of the pre-analysis queue is set to 48 and the preset number of frames is set to 32, then the first 48 video frames from the video to be encoded are read and stored in the pre-analysis queue, and the first 32 video frames in the pre-analysis queue are determined as the final video frame sequence.
[0125] Step 240: Based on the video features of the video frame sequence, determine the last frame from the video frame sequence.
[0126] In some embodiments, the last frame is determined from the current image group based on video features of the current image group.
[0127] Video features refer to the characteristics associated with each video frame in a video frame sequence.
[0128] In some embodiments, video features represent quantized features associated with video coding prediction. Video features are used to measure the coding complexity and inter-frame correlation of video frames. These features include, but are not limited to, intra-frame pattern complexity of the video frame, prediction complexity with reference to other frames, and the noise level of the frame. Video features are used for adaptively grouping images.
[0129] In some embodiments, video features can be sequence features of a video frame sequence formed by multiple consecutive or non-consecutive video frames, or frame features of a single video frame in a video frame sequence. Sequence features can be features related to sharpness or features related to scene. In some embodiments, for sequence features in the video features of a video frame sequence, sequence features are used to characterize at least one of the following: motion information, scene change information, and temporal correlation information. Motion information includes at least one of the following: the trajectory, speed, and direction of objects in the video frame sequence. Scene change information is used to characterize scene transitions. Temporal correlation information is used to characterize the temporal order and dependencies between video frame sequences, which is beneficial for capturing the development and logic of plot content. For frame features in the video features of a video frame sequence, frame features are used to characterize at least one of the following: texture information, edge information, and brightness and color change information. Regions with rich texture information have more complex pixel value changes, and texture information can characterize the complexity of video frames. Edge regions are areas in video frames where the degree of pixel value change is greater than a threshold; edge information is used to indicate the structure of objects in video frames. Brightness and color change information is used to indicate the visual dynamics and visual richness of video frames. In some embodiments, the sequence features include at least one of the following: the maximum intra-frame mode complexity of all video frames in the video frame sequence, the sequence number of the video frame with the maximum value, the proportion of large motion vector blocks in the video frame with the maximum value when referenced to the starting frame, and the sequence number of the frame before the scene transition. The proportion of large motion vector blocks in the video frame with the maximum value when referenced to the starting frame refers to the proportion of pre-analyzed blocks whose motion vectors exceed a threshold when performing inter-frame prediction with reference to the starting frame in the frame with the highest complexity. This proportion serves to quantify the intensity of motion between the frame with the highest complexity and the starting frame.
[0130] In one possible implementation, a video frame sequence is input into a feature extraction model to obtain video features of the video frame sequence. These video features include sequence features of the video frame sequence and frame features of each individual video frame. The feature extraction model is a pre-trained neural network model, which may include at least one of the following: Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Long Short-Term Memory (LSTM) networks. CNNs are used to extract information such as the shape, texture, and color of objects in the video frames, as well as information such as the object's category and scene type. Recurrent Neural Networks and LSTM networks are used to capture the temporal order and dependencies between video frames and to extract motion-related information of objects.
[0131] The last frame is a video frame determined from the video frame sequence that can serve as the last frame of a group of images to be encoded. A group of images to be encoded is the final group of images to be encoded, determined by further dividing the video frame sequence. The method for determining the last frame will be described in detail in the following embodiments.
[0132] In some embodiments, the size of the group of images to be encoded is less than or equal to the size of the video frame sequence; that is, the number of frames in the group of images to be encoded is less than or equal to the number of frames in the video frame sequence. For example, the video frame sequence is the first 32 consecutive frames in the video to be encoded, and the group of images to be encoded is the first 16 consecutive frames in the video frame sequence.
[0133] Step 260: Determine the first frame to the last frame in the video frame sequence as an image group.
[0134] A picture group is a final group of images to be encoded, determined by further dividing a video frame sequence. The start frame is the first video frame in the video frame sequence, and can also be understood as the first video frame of the picture group to be encoded. After determining the end frame from the video frame sequence, the subsequence from the start frame to the end frame in the video frame sequence is determined as a picture group.
[0135] Step 280: Encode each video frame in the image group to obtain the encoded bitstream.
[0136] For example, each video frame in the image group to be encoded is encoded sequentially to obtain the encoded bitstream. The encoding method can be at least one of the following: entropy coding and lossy coding. Among them, entropy coding is a lossless compression method. Lossy coding can be implemented through intra-frame and inter-frame compression, motion compensation, etc. The encoding in this embodiment also conforms to the encoding and decoding standards, such as the HEVC standard, specifically the H.26x series and MPEG series standards.
[0137] In summary, the video encoding method provided in this application involves obtaining a video frame sequence with a preset number of frames from the video to be encoded, where the preset number of frames is a pre-configured number of frames in the video frame sequence; determining the last frame from the video frame sequence based on its video features; defining the image group to be encoded from the start frame to the last frame in the video frame sequence; and encoding each video frame in the image group to be encoded to obtain the encoded bitstream. Accordingly, a novel adaptive method for dividing the image group to be encoded is provided. After determining the video frame sequence from the video to be encoded, the last frame can be determined from the video frame sequence based on its video features, and then the image group to be encoded from the start frame to the last frame can be defined to achieve the encoding of each video frame in the image group to be encoded. On the one hand, compared to related technologies, it is not necessary to traverse every possible image group structure; the image group to be encoded can be determined by determining the last frame, greatly reducing the algorithm complexity. On the other hand, the last frame is determined based on the video features of the video frame sequence, which conforms to the encoding characteristics of the video to be encoded, improving the accuracy of the last frame, thereby improving compression efficiency and encoding performance.
[0138] It should also be noted that the following embodiments may use terms such as first, second, third, etc. to describe various information, but these terms should not be limited to these terms and do not indicate the order of these information or the sequence of calculation steps. They are only used to distinguish information of the same type from each other.
[0139] In some embodiments, at least one preset frame is a preset frame in the video frame sequence that satisfies the picture group structure conditions and is located before the first candidate frame. The picture group structure can be characterized by the size (number of frames) of the picture group to be encoded, and the picture group structure conditions refer to the conditions that the size of the picture group to be encoded must satisfy. In this embodiment, the size of the picture group to be encoded is set to a first size, a second size, or a third size, then at least one preset frame is a preset frame in the video frame sequence that is located before the first candidate frame and makes the size of the picture group to be encoded the first size, the second size, or the third size. In one example, the size of the picture group to be encoded is set to 16, 8, or 4.
[0140] At least one preset frame includes at least one of the following: a first preset frame, a second preset frame, and a third preset frame. In one example, the first preset frame is frame 16, the second preset frame is frame 8, and the third preset frame is frame 4. When frame 16 is the last frame, the size of the image group to be encoded is 16; when frame 8 is the last frame, the size of the image group to be encoded is 8; and when frame 4 is the last frame, the size of the image group to be encoded is 4.
[0141] To reduce algorithm complexity, the process of adaptively dividing the images to be encoded into groups is split into two parts.
[0142] Part 1: The size of the image group to be encoded is divided into a numerical value (max_index). This value can be a non-power of two (except for 32, 16, 8, 4, 2) or a power of two. This value is used to indicate the upper limit of the number of frames supported by the image group to be encoded.
[0143] Part Two: When the value (max_index) obtained in Part One is greater than or equal to 16, determine whether the size of the image group to be encoded needs to be further divided into 16, 8 or 4. The value of 16, 8 or 4 here is used to indicate the optimal number of frames for the image group to be encoded.
[0144] The first part of the process is also referred to as: determining whether it is necessary to divide the images to be encoded into groups according to sharpness / scene. The second part of the process is also referred to as: determining whether it is necessary to divide the size of the image groups to be encoded into 16, 8, or 4. It should also be noted that, according to actual needs and / or experimental verification, 16, 8, or 4 in this embodiment can be set to other values, which are not limited here.
[0145] In some other possible implementations, when the value (max_index) obtained in the first part is greater than or equal to 8 and less than 16, execution can directly begin by determining whether the size of the image group to be encoded needs to be further divided into 8 or 4. Alternatively, when the value (max_index) obtained in the first part is greater than or equal to 4 and less than 8, execution can directly begin by determining whether the size of the image group to be encoded needs to be further divided into 4. That is, based on the sequence number of at least one preset frame, the closest preset frame that is less than the value (max_index) is determined, and execution begins by determining whether the size of the image group to be encoded needs to be divided into the sequence number corresponding to the preset frame, thereby determining the image group to be encoded. Alternatively, in some other possible implementations, when the value (max_index) obtained in the first part is less than 16, subsequent judgments can be skipped, and the first candidate frame can be directly determined as the last frame, thereby determining the image group to be encoded.
[0146] For example, step 240 specifically includes steps 300 and 400:
[0147] Step 300: Based on the sequence features of the video frame sequence, determine the first candidate frame from the video frame sequence;
[0148] Step 400: Based on the frame features of the first candidate frame and at least one preset frame, determine the last frame from the first candidate frame and at least one preset frame; wherein, the sequence number of the first candidate frame in the video frame sequence is used to indicate the upper limit of the number of frames of the image group to be encoded, the sequence number of the last frame in the video frame sequence is used to indicate the optimized number of frames of the image group to be encoded, the optimized number of frames is less than or equal to the upper limit of the number of frames, and at least one preset frame is a preset frame in the video frame sequence that satisfies the image group structure condition and is located before the first candidate frame.
[0149] Video features include sequence features and frame features. Sequence features refer to the features of at least one consecutive or non-consecutive video frame in a video frame sequence; this embodiment uses the features of consecutive video frames as an example. Frame features refer to the features of a single video frame in a video frame sequence. As an example, sequence features can be features related to sharpness or features related to scene. For example, sharpness can be related to at least one of resolution, bitrate, and number of pixels; scene can be related to at least one of objects, scenery, people, lighting, color, composition, and shooting method.
[0150] The last frame refers to the last frame of the image group. The first candidate frame is the frame corresponding to the maximum number of frames in the image group to be encoded. For example, if the maximum number of frames in the image group to be encoded is 16, then the first candidate frame is the 16th frame. The last frame is the last frame corresponding to the optimized number of frames in the image group to be encoded. For example, after optimization, if the optimized number of frames in the image group to be encoded is 8, then the last frame is the 8th frame. The sequence number of the first candidate frame in the video frame sequence indicates the maximum number of frames in the image group to be encoded. The sequence number of the last frame in the video frame sequence indicates the optimized number of frames in the image group to be encoded, which is less than or equal to the maximum number of frames.
[0151] At least one preset frame is a preset frame in the video frame sequence that satisfies the image group structure conditions and is located before the first candidate frame. The image group structure can be characterized by the size (number of frames) of the image group to be encoded, and the image group structure conditions refer to the conditions that the size of the image group to be encoded must satisfy. In this embodiment, the size of the image group to be encoded is set to a first size, a second size, or a third size. Then, at least one preset frame is a preset frame in the video frame sequence that is located before the first candidate frame and makes the size of the image group to be encoded the first size, the second size, or the third size. At least one preset frame includes at least one of the following: a first preset frame, a second preset frame, and a third preset frame. In one example, the size of the image group to be encoded is set to 16, 8, or 4, the first preset frame is the 16th frame, the second preset frame is the 8th frame, and the third preset frame is the 4th frame.
[0152] For example, based on the sequence characteristics of the video frame sequence, the maximum end frame (max_index) is determined from the video frame sequence. Based on the frame characteristics of the first candidate frame and at least one preset frame, the end frame is determined from the first candidate frame and at least one preset frame. Here, the index of the first candidate frame in the video frame sequence indicates the upper limit of the number of frames in the image group to be encoded, and the index of the end frame in the video frame sequence indicates the optimized number of frames in the image group to be encoded, which is less than or equal to the upper limit. That is, the size of the image group to be encoded is less than or equal to the size of the video frame sequence.
[0153] This embodiment describes a method for adaptively dividing the image group to be encoded. By determining the first candidate frame, the upper limit of the number of frames supported by the image group to be encoded can be determined. By determining the last frame, the upper limit of the number of frames of the image group to be encoded is further optimized, thereby improving the accuracy of the image group to be encoded, which is beneficial to the subsequent encoding of the image group to be encoded.
[0154] Specifically, step 400 is implemented as steps 420 and 440:
[0155] Step 420: Statistical analysis is performed on the frame features of the first candidate frame and the frame features of at least one preset frame to obtain statistical features;
[0156] Step 440: Based on the relationship between statistical features and threshold, determine the last frame from the first candidate frame and at least one preset frame; wherein the threshold is related to the statistical features and at least one preset frame.
[0157] This application's embodiments decompose the adaptive GOP partitioning problem into two levels: "determining the maximum possible range" and "optimizing the final range," by introducing a "first candidate frame" as an intermediate decision variable and "at least one preset frame." This hierarchical decision structure reduces the overall algorithm's complexity (simplifying from a global search to a two-step selection process) while retaining the ability to make fine adjustments based on video content. Sequence features are used to capture macroscopic changes (such as sharpness peaks and scene boundaries), while frame features are used for microscopic analysis (such as motion intensity and prediction efficiency). The combination of both makes GOP partitioning more closely reflect the actual characteristics of the video content.
[0158] Statistical features are obtained by performing operations on the frame features of the first candidate frame and at least one preset frame. In some embodiments, statistical features include features related to prediction complexity, features related to motion vectors, and noise features. For example, features related to prediction complexity may be prediction complexity or the ratio between different prediction complexities; noise features may be noise level; and features related to motion vectors may be motion vectors or the proportion of motion vectors. The noise level of an image is typically measured by estimating the intensity of noise in the image (e.g., the standard deviation). The proportion of motion vectors refers to the proportion of the number of pre-analyzed blocks whose motion vectors exceed a vector threshold to the total number of pre-analyzed blocks in that frame. This proportion is used to quantify the intensity of motion between video frames—the higher the proportion, the more intense the motion and the weaker the inter-frame correlation. The ratio between prediction complexities refers to the ratio of the prediction complexity (SATD value) produced by two different prediction methods (e.g., inter-frame prediction versus intra-frame prediction, bidirectional prediction versus unidirectional prediction). This ratio is used to measure the relative efficiency of different prediction methods—the smaller the ratio, the more effective the prediction method being compared.
[0159] The threshold is pre-configured and is related to the statistical feature and at least one preset frame. Specifically, the type, value, and range of the threshold are related to the type of the statistical feature and to each of the at least one preset frame. For example, if the statistical feature is a noise feature, then the threshold is a pre-configured noise threshold. If the statistical feature is a feature related to a first preset frame, then the value and range of the threshold are related to the first preset frame.
[0160] In this embodiment, by statistically analyzing the frame features of the first candidate frame and at least one preset frame, the last frame can be determined from the first candidate frame and at least one preset frame based on the relationship between these statistical features and the threshold. This determines the final size of the image group to be encoded, simplifies the algorithm steps, reduces the algorithm complexity, and improves the accuracy of the image group to be encoded, which is beneficial for subsequent encoding of the image group to be encoded.
[0161] Determine whether the size of the image group to be encoded needs to be divided into 16.
[0162] In some embodiments, at least one preset frame includes a first preset frame, and the statistical features include prediction complexity-related features of the first candidate frame and the first preset frame, respectively, and noise features of the first candidate frame; wherein, the prediction complexity of the first candidate frame is used to measure the difficulty of predicting the pixel value at each position of the first candidate frame. The prediction complexity of the first preset frame is used to measure the difficulty of predicting the pixel value at each position of the first preset frame. Step 420 is implemented as steps 4211, 4212, 4213, 4214, 4215, and 4216:
[0163] Step 4211: Determine the prediction complexity of the first candidate frame compared to the starting frame; determine the prediction complexity of the first candidate frame compared to the previous frame; determine the prediction complexity of the first preset frame compared to the starting frame; determine the prediction complexity of the first preset frame compared to both the starting frame and the first candidate frame.
[0164] Step 4212: Calculate the first proportion of the specified pre-analysis block in the first candidate frame;
[0165] Step 4213: Calculate the first ratio of the prediction complexity of the first candidate frame compared to the starting frame to the first ratio of the intra-frame mode complexity of the first candidate frame.
[0166] Step 4214: Calculate the second ratio of the prediction complexity of the first candidate frame compared to the previous frame to the intra-frame mode complexity of the first candidate frame.
[0167] Step 4215: Calculate the third ratio of the prediction complexity of the first preset frame compared to the starting frame and the first candidate frame, and the prediction complexity of the starting frame compared to the first preset frame.
[0168] Step 4216: Calculate the noise level of the first candidate frame; wherein, the intra-frame mode complexity is the prediction complexity when predicting the pixel value at each position of each pre-analysis block using the intra-frame mode; the specified pre-analysis block is a pre-analysis block using the inter-frame mode and whose motion vector is greater than the vector threshold, the vector threshold is pre-configured, and the pre-analysis block is a pre-divided pixel block.
[0169] The start frame is the first video frame in the video frame sequence, also understood as the first video frame of the group of images to be encoded. The first candidate frame indicates the upper limit of the number of frames supported by the group of images to be encoded from the video frame sequence, also understood as the last frame when the group of images to be encoded is at its maximum size. The preceding frame is a video frame that precedes and is adjacent to the first candidate frame. For example, the first candidate frame is the max_indexx frame, and the preceding frame is the max_index-1 frame. Or, the first candidate frame is the nth frame, and the preceding frame is the (n-1)th frame.
[0170] The first preset frame is a preset frame in the video frame sequence that precedes the first candidate frame. The designated pre-analysis block is a pre-analysis block using inter-frame mode and with motion vectors greater than a vector threshold, which is pre-configured. The pre-analysis block is a pre-divided pixel block; in this embodiment, the pre-analysis block is set to 8×8 blocks. In this embodiment, since the prediction complexity of the first candidate frame compared to the starting frame is determined, the designated pre-analysis block is the designated pre-analysis block in the first candidate frame when the starting frame is used as the reference frame.
[0171] The first ratio (large_mv_frac) is the proportion of the specified pre-analysis blocks in the first candidate frame when the starting frame is used as the reference frame. The first cost ratio is the ratio of the prediction complexity of the first candidate frame compared to the starting frame (when the starting frame is used as the reference frame), and the ratio of the intra-frame mode complexity of the first candidate frame. It can also be understood as the ratio of the prediction complexity of the first candidate frame compared to the starting frame to the intra-frame mode complexity of the first candidate frame. The second ratio (inter_with_prev) is the ratio of the prediction complexity of the first candidate frame compared to the previous frame (when the previous frame is used as the reference frame), and the ratio of the intra-frame mode complexity of the first candidate frame. It can also be understood as the ratio of the prediction complexity of the first candidate frame compared to the previous frame to the intra-frame mode complexity of the first candidate frame. The third ratio (poc16_inter) is the ratio of the prediction complexity of the first preset frame compared to the starting frame and the first candidate frame (when both the starting frame and the first candidate frame are used as reference frames), to the prediction complexity of the first preset frame compared to the starting frame (when the starting frame is used as reference frame). It can also be understood as the ratio of the prediction complexity of the first preset frame compared to the starting frame and the first candidate frame, to the prediction complexity of the first preset frame compared to the starting frame. The noise level (noise_level) is obtained based on gradient calculation.
[0172] In some embodiments, the threshold includes a noise threshold, a first threshold, a second threshold, a third threshold, and a fourth threshold; step 440 is implemented as step 441:
[0173] Step 441: If the noise level of the first candidate frame is less than the noise threshold, and the first ratio is greater than the first threshold, and the first ratio is greater than the second threshold, and the second ratio is less than the third threshold, and the third ratio is greater than the fourth threshold, then the first preset frame is determined as the last frame; otherwise, the first candidate frame is determined as the last frame; wherein, the noise threshold, the first threshold, the second threshold, the third threshold, and the fourth threshold are pre-configured and used to indicate the characteristics of the first candidate frame and the first preset frame.
[0174] The noise level (noise_level), the first ratio (large_mv_frac), the first cost ratio (cost_ratio), the second cost ratio (inter_with_prev), and the third cost ratio (poc16_inter) each have corresponding pre-configured thresholds. The value of these thresholds is used to indicate the characteristics of the first candidate frame and the first preset frame; that is, the value of the threshold can be understood as being related to the first candidate frame and the first preset frame. In some embodiments, the value of the threshold is specifically related to the statistical characteristics of the first candidate frame and the first preset frame; that is, the value of the threshold is used to indicate the statistical characteristics of the first candidate frame and the first preset frame. For example, the first threshold, second threshold, third threshold, and fourth threshold are 0.3, 0.8, 0.6, and 0.7, respectively.
[0175] For example, if the noise level of the first candidate frame is less than the noise threshold, and the first ratio (large_mv_frac) is greater than the first threshold, and the first cost ratio (cost_ratio) is greater than the second threshold, and the second ratio (inter_with_prev) is less than the third threshold, and the third ratio (poc16_inter) is greater than the fourth threshold, then it means that the features of the first candidate frame and the first preset frame meet the pre-configured conditions, and the size of the image group to be encoded can be divided into 16, and the first preset frame can be determined as the last frame, that is, the 16th frame is determined as the last frame, and the size of the image group to be encoded is 16; otherwise, the first candidate frame is determined as the last frame, that is, the maximum last frame (max_index) is determined as the last frame, that is, the size of the image group to be encoded is the maximum supported size.
[0176] This embodiment provides a method for determining whether the size of the image group to be encoded needs to be divided into 16. The first preset frame is predetermined as the 16th frame, which simplifies the algorithm steps and reduces the algorithm complexity.
[0177] To determine the prediction complexity of the first candidate frame compared to the starting frame, the calculation method for the prediction complexity of frame A compared to frame B can be referenced, where frame A is the first candidate frame and frame B is the starting frame. Step 4211 is implemented as steps 42111, 42112, 42113, 42114, and 42115:
[0178] Step 42111: Divide the first candidate frame into at least one pre-analysis block;
[0179] Step 42112: Using the starting frame as the reference frame, traverse at least one pre-analysis block. For the current pre-analysis block in at least one pre-analysis block, determine the corresponding reference block in the starting frame.
[0180] Step 42113: Based on the reference block corresponding to the current pre-analysis block in the starting frame, determine the inter-frame mode complexity of the current pre-analysis block, and determine the motion vector of the current pre-analysis block;
[0181] Step 42114: If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, determine that the current pre-analysis block adopts inter-frame mode and the prediction complexity of the current pre-analysis block is the inter-frame mode complexity; otherwise, determine that the current pre-analysis block adopts intra-frame mode and the prediction complexity of the current pre-analysis block is the intra-frame mode complexity, and repeat the process until the prediction complexity of each pre-analysis block of at least one pre-analysis block is determined.
[0182] Step 42115: Accumulate the prediction complexity of each pre-analysis block of at least one pre-analysis block to obtain the prediction complexity of the first candidate frame compared to the starting frame.
[0183] For example, the first candidate frame is divided into at least one pre-analysis block. In one example, the pre-analysis block is an 8×8 block. Using the start frame as a reference frame, at least one pre-analysis block is traversed. For the current pre-analysis block in at least one pre-analysis block, the corresponding reference block in the start frame is determined. Based on the reference block corresponding to the current pre-analysis block in the start frame, the inter-frame mode complexity of the current pre-analysis block is determined, and the motion vector of the current pre-analysis block is determined.
[0184] If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, then the current pre-analysis block is determined to use inter-frame mode, and its prediction complexity is equal to the inter-frame mode complexity. Otherwise, the current pre-analysis block is determined to use intra-frame mode, and its prediction complexity is equal to the intra-frame mode complexity. This process is repeated until the prediction complexity of each pre-analysis block of at least one pre-analysis block is determined. The prediction complexity of each pre-analysis block of at least one pre-analysis block is accumulated to obtain the prediction complexity of the first candidate frame compared to the starting frame.
[0185] In some embodiments, dividing the first candidate frame into at least one pre-analysis block can be based on the size of the pre-analysis block, dividing the first candidate frame into at least one pre-analysis block of the same size in a left-to-right, top-to-bottom order. Optionally, dividing the first candidate frame into at least one pre-analysis block can be done by first scaling the first candidate frame to 1 / 4 of its original size (scaling the width and height to 1 / 2 of their original sizes), scaling the starting frame to 1 / 4 of its original size (scaling the width and height to 1 / 2 of their original sizes), and then dividing the scaled first candidate frame into at least one pre-analysis block.
[0186] This embodiment provides a method for calculating the prediction complexity of the first candidate frame compared to the starting frame. It should also be noted that the first candidate frame and the starting frame in this embodiment can be replaced with any two other video frames, thereby obtaining the prediction complexity of one video frame compared to another video frame, which improves data processing efficiency.
[0187] Regarding the calculation method of the inter-frame mode complexity of the current pre-analysis block, step 42113 is implemented as steps 42117, 42118, and 42119:
[0188] Step 42117: Based on the reference block corresponding to the current pre-analysis block in the starting frame, use the inter-frame mode to predict the pixel value at each position of the current pre-analysis block to obtain the predicted pixel value;
[0189] Step 42118: Subtract the original pixel value of the current analysis block from the predicted pixel value of each position in the current pre-analysis block to obtain the residual value of each position;
[0190] Step 42119: Perform Hadamard transform on the residual values at each position and sum them to obtain the inter-frame mode complexity of the current pre-analysis block.
[0191] For example, based on the reference block corresponding to the current pre-analysis block in the starting frame, an inter-frame mode is used to predict the pixel value at each position of the current pre-analysis block, resulting in a predicted pixel value. The predicted pixel value at each position of the current pre-analysis block is then subtracted from the original pixel value of the reference block corresponding to the current analysis block in the starting frame, yielding a residual value at each position. Finally, a Hadamard transform is performed on the residual values at each position, and the summation (stad) is performed to obtain the inter-frame mode complexity of the current pre-analysis block.
[0192] This embodiment provides a method for calculating the inter-frame mode complexity of the current pre-analysis block, which can improve the accuracy of data calculation and improve data processing efficiency.
[0193] • Determine whether the size of the image group to be encoded needs to be divided into 8 parts.
[0194] In some embodiments, at least one preset frame includes a first preset frame, and the statistical features include features of the first preset frame related to prediction complexity; step 420 is implemented as steps 4221, 4222, 4223, 4224, and 4225:
[0195] Step 4221: Determine the prediction complexity of the first preset frame compared to the starting frame;
[0196] Step 4222: Calculate the second proportion of the specified pre-analysis block in the first preset frame;
[0197] Step 4223: Calculate the third proportion of pre-analysis blocks using intra-frame mode in the first preset frame;
[0198] Step 4224: Calculate the fourth ratio of the prediction complexity of the first preset frame compared to the starting frame and the intra-frame mode complexity compared to the first preset frame.
[0199] Step 4225: Calculate the fifth ratio of the intra-frame mode complexity of the starting frame to the intra-frame mode complexity of the first preset frame.
[0200] The start frame is the first video frame in the video frame sequence, also understood as the first video frame of the group of images to be encoded. The first preset frame is the preset frame in the video frame sequence that precedes the first candidate frame.
[0201] The designated pre-analysis block is a pre-analysis block that uses inter-frame mode and whose motion vector is greater than a vector threshold, which is pre-configured. In this embodiment, since the prediction complexity of the first preset frame compared to the starting frame is determined, the designated pre-analysis block is the designated pre-analysis block in the first preset frame when the starting frame is used as the reference frame.
[0202] The second ratio (large_mv_frac) is the proportion of specified pre-analysis blocks in the first preset frame when the starting frame is used as the reference frame. The third ratio (intra_frac) is the proportion of pre-analysis blocks using intra-frame mode in the first preset frame when the starting frame is used as the reference frame. The fourth ratio (cost_ratio16) is the ratio of the prediction complexity of the first preset frame compared to the starting frame (when the starting frame is used as the reference frame), and the intra-frame mode complexity of the first preset frame. This can also be understood as the ratio of the prediction complexity of the first preset frame compared to the starting frame, to the intra-frame mode complexity of the first preset frame. The fifth ratio (similarity) is the ratio of the total intra-frame mode complexity of the starting frame compared to the total intra-frame mode complexity of the first preset frame. This can also be understood as the ratio of the total intra-frame mode complexity of the starting frame, and the total intra-frame mode complexity of the first preset frame. The closer the fifth ratio (similarity) is to 1, the more similar the starting frame and the first preset frame are.
[0203] In some embodiments, at least one preset frame further includes a second preset frame, which is a preset frame in the video frame sequence that precedes the first preset frame; the threshold includes a fifth threshold, a sixth threshold, a first interval range, and a second interval range; step 440 is implemented as step 442:
[0204] Step 442: If the second ratio is less than the fifth threshold, the third ratio is less than the sixth threshold, the fourth ratio is within the first interval range, and the fifth ratio is within the second interval range, then the second preset frame is determined as the last frame; otherwise, the first preset frame remains the last frame. The fifth threshold, the sixth threshold, the first interval range, and the second interval range are pre-configured and used to indicate the characteristics of the first preset frame. For example, the fifth threshold and the sixth threshold are 0.2 and 0.1, respectively. The first interval range is [0.5, 0.9], and the second interval range is [0.8, 1.2].
[0205] The second ratio (large_mv_frac), the third ratio (intra_frac), the fourth ratio (cost_ratio16), and the fifth ratio (similarity) each have corresponding pre-configured thresholds and / or ranges. The value of the threshold and the size of the range are used to indicate the characteristics of the first preset frame; that is, the value of the threshold and the size of the range are related to the first preset frame. In some embodiments, the value of the threshold and the size of the range are specifically related to the statistical characteristics of the first preset frame; that is, the value of the threshold is used to indicate the statistical characteristics of the first preset frame.
[0206] For example, if the second ratio (large_mv_frac) is less than the fifth threshold, the third ratio (intra_frac) is less than the sixth threshold, the fourth ratio (cost_ratio16) is within the first interval range, and the fifth ratio (similarity) is within the second interval range, then it means that the features of the first preset frame meet the pre-configured conditions, and the size of the image group to be encoded can be divided into 8, and the second preset frame can be determined as the last frame, that is, the 8th frame is determined as the last frame, and the size of the image group to be encoded is 8; otherwise, the first preset frame is kept as the last frame, that is, the size of the image group to be encoded is kept as 16.
[0207] This embodiment provides a method for determining whether the size of the image group to be encoded needs to be divided into 8. The second preset frame is predetermined as the 8th frame, which simplifies the algorithm steps and reduces the algorithm complexity.
[0208] Determine whether the size of the image group to be encoded needs to be divided into 4.
[0209] In some embodiments, at least one preset frame includes a second preset frame, and the statistical features include features of the second preset frame related to prediction complexity; step 420 is implemented as steps 4231, 4232, and 4233:
[0210] Step 4231: Determine the prediction complexity of the second preset frame compared to the starting frame;
[0211] Step 4232: Calculate the fourth proportion of the specified pre-analysis block in the second preset frame;
[0212] Step 4233: Calculate the sixth ratio of the prediction complexity of the second preset frame compared to the starting frame and the intra-frame mode complexity compared to the second preset frame.
[0213] The start frame is the first video frame in the video frame sequence, also understood as the first video frame of the group of images to be encoded. The second preset frame is a preset frame in the video frame sequence that precedes the first preset frame.
[0214] The designated pre-analysis block is a pre-analysis block that uses inter-frame mode and whose motion vector is greater than a vector threshold, which is pre-configured. In this embodiment, since the prediction complexity of the second preset frame compared to the starting frame is determined, the designated pre-analysis block is the designated pre-analysis block in the second preset frame when the starting frame is used as the reference frame.
[0215] The fourth ratio (large_mv_frac) is the proportion of the specified pre-analysis blocks in the second preset frame when the starting frame is used as the reference frame. The sixth ratio (cost_ratio) is the ratio of the prediction complexity of the second preset frame compared to the starting frame (when the starting frame is used as the reference frame), and the intra-frame mode complexity of the second preset frame. It can also be understood as the ratio of the prediction complexity of the second preset frame compared to the starting frame, and the intra-frame mode complexity of the second preset frame.
[0216] In some embodiments, at least one preset frame further includes a third preset frame, which is a preset frame in the current image frame that precedes the second preset frame; the threshold includes a seventh threshold and a third interval range; step 440 is implemented as step 443:
[0217] Step 443: If the fourth ratio is greater than the seventh threshold and the sixth ratio is within the third interval range, the third preset frame is determined as the last frame; otherwise, the second preset frame remains the last frame. The seventh threshold and the third interval range are pre-configured and used to indicate the characteristics of the second preset frame. For example, the seventh threshold is 0.4. The third interval range is, for example, [0.6, 0.95].
[0218] The fourth ratio (large_mv_frac) and the sixth ratio (cost_ratio) each have corresponding pre-configured thresholds and / or ranges. The value of the threshold and the size of the range are used to indicate the characteristics of the second preset frame; that is, the value of the threshold and the size of the range can also be understood as being related to the second preset frame. In some embodiments, the value of the threshold and the size of the range are specifically related to the statistical characteristics of the second preset frame; that is, the value of the threshold is used to indicate the statistical characteristics of the second preset frame.
[0219] For example, if the fourth ratio (large_mv_frac) is greater than the seventh threshold and the sixth ratio (cost_ratio) is within the third interval range, it means that the features of the second preset frame meet the pre-configured conditions, and the size of the image group to be encoded can be divided into 4, and the third preset frame can be determined as the last frame, that is, the fourth frame is determined as the last frame, and the size of the image group to be encoded is 4; otherwise, the second preset frame is kept as the last frame, that is, the size of the image group to be encoded is kept as 8.
[0220] This embodiment provides a method for determining whether the size of the image group to be encoded needs to be divided into 4. The third preset frame is predetermined as the fourth frame, which simplifies the algorithm steps and reduces the algorithm complexity.
[0221] Determine whether it is necessary to divide the images to be encoded into groups based on sharpness / scene.
[0222] In some embodiments, step 300 is implemented as steps 310, 320, 330, and 340:
[0223] Step 310: Determine the intra-frame mode complexity of each video frame in the video frame sequence;
[0224] Step 320: Determine the maximum complexity in the intra-frame mode complexity of each video frame, and determine the video frame with the maximum complexity as the maximum complexity frame of the video frame sequence.
[0225] Step 330: Determine the proportion of designated pre-analysis blocks in the maximum complexity frame, with the starting frame as the reference frame. Using the starting frame as the reference frame, perform inter-frame prediction on each pre-analysis block in the maximum complexity frame to determine its motion vector, and determine whether the pre-analysis block meets the criteria for a designated pre-analysis block—that is, using inter-frame prediction mode and having a motion vector greater than a vector threshold. Then, the proportion of the designated pre-analysis block in all predetermined analysis blocks of the maximum complexity frame can be determined.
[0226] Step 340: If the maximum complexity is less than the eighth threshold, the sequence number of the frame with the maximum complexity in the video frame sequence is greater than the ninth threshold, and the proportion of the specified pre-analysis block is greater than the tenth threshold, then the frame with the maximum complexity is determined as the first candidate frame. Here, the eighth threshold is calculated based on the image resolution, and the ninth and tenth thresholds are pre-configured. For example, the ninth and tenth thresholds are 4 and 0.25, respectively.
[0227] For a video frame sequence, determine the intra-mode complexity of each video frame. From the intra-mode complexity of each video frame, determine the maximum complexity (max_intra), and designate the video frame with the maximum complexity as the maximum complexity frame of the video frame sequence. This maximum complexity frame has a corresponding index (max_index) in the video frame sequence. Using the starting frame as a reference frame, determine the proportion (large_mv_ratio) of a specified pre-analysis block in the maximum complexity frame. The specified pre-analysis block is a pre-analysis block using inter-mode and whose motion vector is greater than a vector threshold.
[0228] If the maximum complexity (max_intra) is less than the eighth threshold, and the index of the maximum complexity frame in the video frame sequence (max_index) is greater than the ninth threshold, and the proportion of the specified pre-analysis block (large_mv_ratio) is greater than the tenth threshold, the maximum complexity frame is determined as the first candidate frame. That is, the upper limit of the number of frames supported by the image group to be encoded can be represented by this first candidate frame.
[0229] The eighth threshold is calculated based on the image resolution, while the ninth and tenth thresholds are pre-configured. In some embodiments, the formula for calculating the eighth threshold is: Eighth threshold = Preset coefficient × Video frame width × Video frame height. Since the video frames in the video frame sequence are obtained from the same video to be encoded, the eighth threshold is the same for each video frame. In one example, the preset coefficient is 0.65. For example, the eighth threshold is equal to 0.65 × width × height (e.g., for 1080p video, it is approximately 0.65 × 1920 × 1080 ≈ 1,347,840).
[0230] In some embodiments, when two or more scenes are involved in the video frame sequence, scene switching is also taken into account, and the video frames before the scene switching are divided into the same video frame sequence and the same group of images to be encoded as much as possible. Step 300 is also implemented as steps 350 and 360:
[0231] Step 350: Determine the last frame of the first scene in the video frame sequence;
[0232] Step 360: The frame that appears first in the order of the last frame of the scene and the frame with the highest complexity is determined as the first candidate frame.
[0233] The first scene is the scene that appears first in the video frame sequence, or it can be understood as the scene before the scene switch. The last frame of the scene is the last frame of the first scene before the scene switch. For example, if the video frame sequence includes 32 frames, frames 1 to 16 involve scene 1, and frames 17 to 32 involve scene 2, then the last frame of the first scene is frame 16. For example, based on at least one of the following: inter-frame prediction complexity change, brightness histogram difference, or intra-frame mode block ratio, it is detected whether there is a scene switch in the current image group; if there is a scene switch, then the frame before the scene switch is determined to be the last frame of the scene; the frame that appears first in the sequence between the last frame of the scene and the frame with the highest complexity is determined to be the highest last frame.
[0234] A scene can be associated with at least one of the following: objects, scenery, people, lighting, color, composition, and shooting method. In some embodiments, different scenes in a video frame sequence are pre-labeled, or determined by transition video frames in the individual video frames of the video frame sequence, or obtained by performing scene recognition on each video frame in the video frame sequence using a scene recognition model. For example, the scene recognition model is a pre-trained neural network model used to recognize the scene in each video frame of the video frame sequence. The scene recognition model includes at least one of the following: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and time series models.
[0235] For example, the last frame of the first scene in the video frame sequence is determined; the frame that appears first in the sequence between the last frame and the frame with the highest complexity is determined as the first candidate frame. For example, if the video frame sequence includes 32 frames, frames 1 to 16 relate to scene 1, frames 17 to 32 relate to scene 2, the frame with the highest complexity is frame 18, and the last frame of scene 1 is frame 16, then frame 16 is determined as the first candidate frame.
[0236] This embodiment provides a method for determining whether to divide the images to be encoded into groups based on sharpness / scene. Specifically, by using the frame with the highest complexity as the first candidate frame, it ensures the identification of high-quality frames containing more detail and texture information. When high-quality frames are encoded with high quality, they provide a better reference for other blurry frames, improving the overall encoding quality of the video, especially suitable for videos with significant motion and blurriness. By using the frame at the end of the scene as the first candidate frame, it ensures that video frames before the scene change are grouped into the same video frame sequence and the same image group to be encoded, improving encoding efficiency.
[0237] To calculate the intra-mode complexity of each video frame in the video frame sequence, refer to the method for calculating the intra-mode complexity of frame A, where frame A is each video frame in the video frame sequence. For the current video frame, step 310 is implemented as steps 311, 312, 313, 314, and 315:
[0238] Step 311: Divide the current video frame into at least one pre-analysis block;
[0239] Step 312: Traverse at least one pre-analysis block. For the current pre-analysis block in at least one pre-analysis block, use intra-frame mode to predict the pixel value at each position of the current pre-analysis block to obtain the predicted pixel value.
[0240] Step 313: Subtract the predicted pixel value from the original pixel value at each position of the current pre-analysis block to obtain the residual value at each position;
[0241] Step 314: Perform Hadamard transform on the residual value at each position and sum them to obtain the intra-mode complexity of the current pre-analysis block. Repeat this process until the intra-mode complexity of each pre-analysis block of at least one pre-analysis block is obtained.
[0242] Step 315: Accumulate the intra-frame mode complexity of each pre-analysis block of at least one pre-analysis block to obtain the intra-frame mode complexity of the current video frame.
[0243] For example, the current video frame is divided into at least one pre-analysis block; the at least one pre-analysis block is traversed, and for the current pre-analysis block in the at least one pre-analysis block, the intra mode is used to predict the pixel value at each position of the current pre-analysis block to obtain the predicted pixel value.
[0244] Subtract the predicted pixel value from the original pixel value at each position of the current pre-analysis block to obtain the residual value at each position. Perform a Hadamard transform on the residual values at each position and sum them (stad) to obtain the intra-frame mode complexity of the current pre-analysis block. Repeat this process until the intra-frame mode complexity of each pre-analysis block of at least one pre-analysis block is obtained. Accumulate the intra-frame mode complexity of each pre-analysis block of at least one pre-analysis block to obtain the intra-frame mode complexity of the current video frame.
[0245] This embodiment provides a method for calculating the intra-frame complexity of the current video frame, which can improve the accuracy of data calculation and increase data processing efficiency.
[0246] In some embodiments, step 330 is implemented as steps 331 and 332:
[0247] Step 331: Determine the specified pre-analysis block with the starting frame as the reference frame for the frame with the maximum complexity.
[0248] Step 332: Determine the proportion of the specified pre-analysis blocks in the maximum complexity frame based on the number of specified pre-analysis blocks and the number of pre-analysis blocks in the maximum complexity frame.
[0249] For example, using the starting frame as a reference frame, the specified pre-analysis of the maximum complexity (max_intra) frame compared to the starting frame is determined. Based on the number of specified pre-analysis blocks in the maximum complexity (max_intra) frame and the total number of pre-analysis blocks, the ratio (large_mv_ratio) of the specified pre-analysis blocks in the maximum complexity (max_intra) frame is determined by dividing the number of specified pre-analysis blocks by the total number of pre-analysis blocks. Here, the number of pre-analysis blocks refers to the total number of pre-analysis blocks in the maximum complexity frame.
[0250] Specifically, step 331 includes steps 3311, 3312, 3313, and 3314:
[0251] Step 3311: Using the starting frame as the reference frame, determine the reference block in the starting frame for each pre-analysis block in the maximum complexity frame.
[0252] Step 3312: Based on the reference block corresponding to each pre-analysis block in the starting frame, determine the inter-frame mode complexity of each pre-analysis block in the maximum complexity frame, and determine the motion vector of each pre-analysis block in the maximum complexity frame.
[0253] Step 3313: If the inter-frame mode complexity of the current pre-analysis block of the maximum complexity frame is less than the intra-frame mode complexity of the current pre-analysis block, determine that the current pre-analysis block adopts inter-frame mode; otherwise, determine that the current pre-analysis block adopts intra-frame mode.
[0254] Step 3314: The pre-analysis block in the frame with the highest complexity that adopts inter-frame mode and whose motion vector is greater than the vector threshold is identified as the specified pre-analysis block.
[0255] For example, taking the starting frame as the reference frame, the reference block corresponding to each pre-analysis block in the maximum complexity (max_intra) frame is determined in the starting frame; based on the reference block corresponding to each pre-analysis block in the maximum complexity frame, the inter-frame (inter) mode complexity of each pre-analysis block in the maximum complexity frame is determined, and the motion vector (mv) of each pre-analysis block in the maximum complexity frame is determined.
[0256] If the inter-frame complexity of the current pre-analysis block in the frame with the highest complexity is less than the intra-frame complexity, the current pre-analysis block is determined to use inter-frame mode; otherwise, the current pre-analysis block is determined to use intra-frame mode. Pre-analysis blocks in the frame with the highest complexity that use inter-frame mode and whose motion vectors are greater than a vector threshold are identified as designated pre-analysis blocks.
[0257] This embodiment provides a method for calculating the specified pre-analysis block of the maximum complexity frame compared to the starting frame. It should also be noted that the maximum complexity frame and the starting frame in this embodiment can be replaced with any two other video frames, thereby obtaining the specified pre-analysis block of one video frame compared to another video frame, which improves data processing efficiency.
[0258] In some embodiments, the method further includes step 500:
[0259] Step 500: If the starting frame meets the segmentation conditions, perform the step of determining the last frame from the first candidate frame and at least one preset frame based on the frame features of the first candidate frame and at least one preset frame; otherwise, determine the first candidate frame as the last frame.
[0260] The starting frame is the first video frame in the video frame sequence, also understood as the first video frame of the group of images to be encoded. The partitioning condition is the condition that the video frame sequence needs to be further divided into 16, 8, or 4 frames. When the starting frame meets the partitioning condition, the step of determining the ending frame from the first candidate frame and at least one preset frame based on frame features is performed; otherwise, the first candidate frame is directly determined as the ending frame, thereby determining the group of images to be encoded.
[0261] In some embodiments, the method for determining whether the starting frame meets the segmentation conditions is further implemented as steps 510 and 520:
[0262] Step 510: Determine the intra-frame mode complexity of the starting frame;
[0263] Step 520: If the complexity of the entire frame of the starting frame is less than or equal to the complexity threshold, the starting frame satisfies the partitioning condition; wherein, the complexity threshold is related to the number of pre-analysis blocks in the starting frame.
[0264] For example, the intra-frame mode complexity of the starting frame is determined. If the intra-frame mode complexity of the starting frame is less than or equal to a complexity threshold, the starting frame meets the partitioning condition, and the steps to partition the video frame sequence into 16, 8, or 4 frames need to be performed. Otherwise, if the intra-frame mode complexity of the starting frame is greater than the complexity threshold, no further judgment is performed, and the first candidate frame is directly determined as the last frame, thus determining the group of images to be encoded. The complexity threshold is related to the number of pre-analyzed blocks in the starting frame. In one example, the complexity threshold is expressed as F(N). F is a pre-configured function, and N is the number of pre-analyzed blocks in the starting frame. For example, F(N) = N × 500.
[0265] In this embodiment, if the starting frame meets the partitioning conditions, the subsequent judgments are performed; otherwise, the subsequent judgments are not performed, and the first candidate frame is directly determined as the last frame, which can improve data processing efficiency.
[0266] In some embodiments, the video frame sequence is obtained based on a pre-analysis queue. The pre-analysis queue is a fixed-size queue used to store the video frames to be analyzed frame by frame. Specifically, step 220 is implemented as steps 221, 222, and 223:
[0267] Step 221: Read the first frame of the video to be encoded;
[0268] Step 222: Save the first number of video frames in the pre-analysis queue;
[0269] Step 223: Determine the video frames of the preset number of frames in the pre-analysis queue as a video frame sequence; wherein, the maximum value of the first frame number is the pre-configured maximum number of read frames.
[0270] For example, video frames of the first frame number in the video to be encoded are read sequentially according to time. The maximum value of the first frame number is a pre-configured maximum number of frames to read. The video frames of the first frame number are stored in a pre-analysis queue, and the video frames of the preset frame number in the pre-analysis queue are determined as a video frame sequence. Optionally, the preset frame number is less than or equal to the first frame number.
[0271] In some embodiments, the size of the video frame sequence is represented as: gop_size = min(cfg_gop_size, num_frames). Here, cfg_gop_size is the pre-configured maximum image group size, and num_frames is the number of the first frame in the pre-analysis queue. The maximum number of the first frame is pre-configured, representing the maximum number of frames that can be read into the pre-analysis queue at one time. In one example, at the end of the video to be encoded, the number of the first frame in the pre-analysis queue is the number of remaining unencoded frames.
[0272] This embodiment introduces a pre-analysis queue, which can determine the video frame sequence and improve the accuracy of the video frame sequence.
[0273] In some embodiments, the video frame sequence is obtained based on a pre-analysis queue. After determining the group of images to be encoded from the video frame sequence, each video frame of the group of images to be encoded is dequeued from the pre-analysis queue and encoded. This embodiment supports iterative execution of encoding. For example, step 292 is included after step 260, and step 294 is included after step 280:
[0274] Step 292: Remove each video frame from the pre-analysis queue in the image group to be encoded to obtain the pre-analysis queue after removal. The pre-analysis queue after removal contains the video frames of the second frame number.
[0275] Step 294: After encoding each video frame in the image group to be encoded and obtaining the encoded bitstream, the pre-analysis queue is replenished with the first number of video frames, and the step of determining the video frames of the preset number of frames in the pre-analysis queue as a video frame sequence is repeated.
[0276] After determining the image group to be encoded from the video frame sequence, each video frame in the image group to be encoded is dequeued from the pre-analysis queue, resulting in a dequeued pre-analysis queue. This dequeued pre-analysis queue contains a second number of video frames. This second number of video frames represents the remaining video frames after dequeuing the video frames from the image group to be encoded from the pre-analysis queue.
[0277] After encoding each video frame in the image to be encoded to obtain the encoded bitstream, the pre-analysis queue is replenished with the first number of video frames, forming a new pre-analysis queue. The newly added video frames are placed after the second number of video frames in chronological order. After forming the new pre-analysis queue, the step of determining the preset number of video frames in the pre-analysis queue as a video frame sequence is repeated, thus continuing to iterate and determine new video frame sequences and new groups of images to be encoded, and performing encoding, until the entire video to be encoded has been encoded.
[0278] This embodiment provides an iterative execution method for the video encoding method. By iteratively executing the method, it is possible to ensure that the entire video to be encoded has been encoded, thereby improving the accuracy of encoding the video.
[0279] The following is a general description of the video encoding method provided in this application, executed by an encoder, using flowcharts and diagrams as examples. In this embodiment, the first preset frame in the aforementioned embodiments is the 16th frame, the second preset frame is the 8th frame, and the third preset frame is the 4th frame.
[0280] Application scenarios
[0281] The video encoding method provided in this application can be integrated into various types of encoders for adaptively dividing image groups. It is suitable for scenarios insensitive to latency and buffering, such as offline compression. In other embodiments, it is also suitable for videos with significant motion and blurriness to be encoded, as well as videos with scene transitions. Specifically, based on video characteristics, the size of the image group to be encoded is divided into a numerical value. This value supports a non-power-two value (values other than 32, 16, 8, 4, and 2), and also supports a power-two value. Furthermore, the size of the image group to be encoded is adaptively divided into 16, 8, or 4 to improve the encoder's compression efficiency and encoding performance.
[0282] • Technical Implementation
[0283] Referring to Figure 4, the overall steps of the video encoding method performed by the encoder are as follows:
[0284] 1. Input the video to be encoded into the encoder;
[0285] 2. Read the video frame sequence and store it in the pre-analysis queue;
[0286] 3. Calculate video features to determine the size of the image group to be encoded;
[0287] 4. Remove each video frame from the pre-analysis queue and encode it;
[0288] 5. Output encoded bitstream;
[0289] 6. Determine the pre-analysis queue after the exit, and repeat steps 2-4.
[0290] For example, if a pre-analysis queue stores 48 video frames, and the size of the video frame sequence is 32, then the first 32 video frames in the pre-analysis queue are determined as the video frame sequence. Based on video features, the size of the image to be encoded is determined to be 16. Therefore, the image group to be encoded consists of the first 16 video frames in the video frame sequence. These first 16 video frames are then extracted from the pre-analysis queue, encoded, and the resulting encoded bitstream is obtained. In the next iteration, video frames are read, and the pre-analysis queue is replenished to 48 video frames, forming a new pre-analysis queue. The steps related to determining the size of the image group to be encoded based on video features are then executed again.
[0291] Referring to Figure 5, the steps for adaptively dividing the image into groups are as follows:
[0292] 1. Calculate the maximum number of frames supported by the group of images to be encoded;
[0293] And; 2. Determine whether it is necessary to divide the images to be encoded into groups according to sharpness / scene;
[0294] 3. Determine whether the size of the image group to be encoded needs to be divided into 16;
[0295] If yes, divide the size of the image group to be encoded into 16 and execute 4; otherwise, execute 6 directly.
[0296] 4. Determine whether the size of the image group to be encoded needs to be divided into 8 parts;
[0297] If yes, divide the size of the image group to be encoded into 8 and execute 5; otherwise, execute 6 directly.
[0298] 5. Determine whether the size of the image group to be encoded needs to be divided into 4;
[0299] If yes, the size of the image group to be encoded is divided into 4 and 6 is executed; otherwise, 6 is executed directly.
[0300] 6. Determine the group of images to be encoded.
[0301] The steps for 1. calculating the maximum number of frames supported by the image group to be encoded; and 2. determining whether the image group to be encoded needs to be divided according to sharpness / scene are as follows:
[0302] 1. The size of the video frame sequence is represented as: gop_size = min(cfg_gop_size, num_frames). Where cfg_gop_size is the pre-configured maximum image group size, and num_frames is the number of the first frame in the pre-analysis queue. The maximum value of the first frame number is pre-configured, representing the maximum number of frames read into the pre-analysis queue at once. For example, the first frame number of the pre-analysis queue can be set to at least one of the following: 48, 64, or 80. The size of the video frame sequence is the preset number of frames, for example, the preset number of frames could be 32. At the end of the video, the first frame number in the pre-analysis queue is the number of remaining unencoded frames.
[0303] 2. For video frames in the pre-analysis queue ranging from the starting frame to `gop_size`, i.e., for a video frame sequence, determine the intra-frame mode complexity of each video frame in the sequence. Calculate the maximum intra-frame mode complexity (`max_intra`) for each video frame. Determine the index (`max_index`) of the frame with the maximum complexity in the pre-analysis queue.
[0304] 3. Determine the ratio (large_mv_ratio) of the specified pre-analysis block in the max_index frame (maximum complexity frame) compared to the starting frame. The specified pre-analysis block is a pre-analysis block that uses inter-frame mode and whose motion vector (mv) exceeds a pre-configured vector threshold.
[0305] 4. If the maximum complexity (max_intra) is less than the threshold calculated based on the image resolution, and the index (max_index) is greater than the pre-configured threshold, and the proportion of the specified pre-analysis block (large_mv_ratio) is greater than the pre-configured threshold, the upper limit of the number of frames supported by the image group to be encoded is determined as the index (max_index) of the maximum complexity frame. That is, the maximum complexity frame can be used as the first candidate frame of the image group to be encoded, and the maximum size supported by the image group to be encoded is the index (max_index).
[0306] Here, the index (max_index) can be a non-power-of-two value or a power-of-two value. The threshold calculated based on image resolution is: preset coefficient × width of video frame × height of video frame. In one example, the preset coefficient is 0.65.
[0307] In some embodiments, when the video frame sequence involves two or more scenes, the last frame of the first scene in the video frame sequence is determined; the frame that appears first in the sequence between the last frame of the scene and the frame with the highest complexity is determined as the first candidate frame. That is, the maximum size supported by the group of images to be encoded is determined as the sequence number corresponding to the last frame of the scene.
[0308] For step 3, determining whether the size of the image group to be encoded needs to be divided into 16, the implementation steps are as follows:
[0309] If the maximum supported size of the image group to be encoded is greater than or equal to 16, the optimal number of frames for the image group to be encoded is determined, thereby determining the optimal size of the image group to be encoded.
[0310] 1. If the intra-complexity of the starting frame in the video frame sequence is greater than the complexity threshold, then no further judgment is performed, and the image group to be encoded at this time is directly removed from the pre-analysis queue and encoding is performed. If the intra-complexity of the starting frame in the video frame sequence is less than or equal to the complexity threshold, then further judgment is performed to determine the optimized number of frames for the image group to be encoded.
[0311] The complexity threshold is related to the number of pre-analyzed blocks in the starting frame. In one example, the complexity threshold is denoted as F(N), where F is a pre-configured function and N is the number of pre-analyzed blocks in the starting frame.
[0312] 2. Based on the maximum size supported by the image group to be encoded, determine the first candidate frame of the image group to be encoded. Calculate the prediction complexity of the first candidate frame compared to the starting frame (frame 1); calculate the prediction complexity of frame 16 compared to the starting frame; calculate the prediction complexity of frame 16 compared to both the starting frame and the first candidate frame.
[0313] 3. Analyze the features of the first candidate frame related to prediction complexity:
[0314] The first proportion (large_mv_frac) of the specified pre-analysis block in the first candidate frame is calculated, where the specified pre-analysis block refers to the pre-analysis block that uses inter-frame mode and whose motion vector (mv) is greater than the vector threshold;
[0315] The first ratio (cost_ratio) is calculated between the prediction complexity of the first candidate frame when the starting frame is used as the reference frame and the intra-frame mode complexity when it uses the full intra-frame mode.
[0316] The prediction complexity of the first candidate frame when the previous frame is used as the reference frame is calculated as the second ratio (inter_with_prev) to the intra-frame mode complexity when it adopts the full intra mode; where the first candidate frame is the nth frame, and the previous frame is the (n-1)th frame.
[0317] The third ratio (poc16_inter) is calculated between the prediction complexity of frame 16 with the start frame and the maximum end frame as reference frames and the prediction complexity of frame 16 with only the start frame as reference frame.
[0318] 4. Statistically analyze the noise characteristics of the first candidate frame: noise level.
[0319] 5. If the noise level of the first candidate frame is less than the pre-configured threshold, the second ratio (inter_with_prev) is less than the pre-configured threshold, the first ratio (large_mv_frac) is greater than the pre-configured threshold, the first ratio (cost_ratio) is greater than the pre-configured threshold, and the third ratio (poc16_inter) is greater than the pre-configured threshold, then the size of the image group to be encoded is set to 16, that is, the 16th frame is determined as the last frame of the image group to be encoded.
[0320] For step 4, determining whether the size of the image group to be encoded needs to be divided into 8 parts, the implementation steps are as follows:
[0321] 1. Calculate the prediction complexity of the 16th frame in a video frame sequence compared to the starting frame (the 1st frame).
[0322] 2. Analyze the features related to prediction complexity in frame 16:
[0323] The second proportion (large_mv_frac) of a specified pre-analysis block in frame 16 is calculated; where the specified pre-analysis block refers to a pre-analysis block that uses inter-frame mode and whose motion vector (mv) is greater than the vector threshold.
[0324] The third proportion (intra_frac) of the pre-analysis block using intra mode in frame 16 is calculated.
[0325] The fourth ratio (cost_ratio16) is calculated between the prediction complexity of frame 16 with the starting frame as the reference frame and the complexity of frame 16 in intra mode.
[0326] The fifth similarity is calculated by comparing the intra-frame complexity of the starting frame with the intra-frame complexity of the 16th frame.
[0327] 3. If the third ratio (intra_frac) is less than the pre-configured threshold, the second ratio (large_mv_frac) is less than the pre-configured threshold, the fourth ratio (cost_ratio16) is within the pre-configured range, and the fifth ratio (similarity) is within the pre-configured range, then the size of the image group to be encoded is divided into 8, that is, the 8th frame is determined as the last frame of the image group to be encoded; otherwise, the size of the image group to be encoded remains unchanged at 16.
[0328] For step 5, determining whether the size of the image group to be encoded needs to be divided into 4 parts; the implementation steps are as follows:
[0329] 1. Calculate the prediction complexity of the 8th frame in a video frame sequence compared to the starting frame (the 1st frame).
[0330] 2. Analyze the features related to prediction complexity in frame 8:
[0331] The fourth proportion (large_mv_frac) of the specified pre-analysis block in frame 8 is calculated; where the specified pre-analysis block refers to the pre-analysis block that uses inter-frame mode and whose motion vector (mv) is greater than the vector threshold;
[0332] The sixth ratio (cost_ratio) is calculated between the prediction complexity of frame 8 when the starting frame is used as the reference frame and the intra-frame mode complexity when using the full intra-frame mode.
[0333] 3. If the fourth ratio (large_mv_frac) is greater than the pre-configured threshold and the sixth ratio (cost_ratio) is within the pre-configured range, the size of the image group to be encoded is divided into 4, that is, the 4th frame is determined as the last frame of the image group to be encoded; otherwise, the size of the image group to be encoded remains unchanged at 8.
[0334] The steps for calculating the intra-frame mode complexity of a video frame in the aforementioned embodiments are as follows:
[0335] The current video frame is scaled to its original preset size, for example, 1 / 4 (width and height are scaled to 1 / 2 of their original size). The scaled current video frame is divided into at least one pre-analysis block, each pre-analysis block being an 8×8 block. Iterating through at least one pre-analysis block, for the current pre-analysis block, the pixel value at each position in the current pre-analysis block is predicted using the 35 intra-prediction modes in the HEVC coding standard. The predicted pixel value is then subtracted from the original pixel value at each position in the current pre-analysis block to obtain the residual value at each position. A Hadamard transform is performed on the residual values, and the summation (satd) is calculated to obtain the intra-prediction mode complexity of the current pre-analysis block. This process is repeated until the intra-prediction mode complexity of each pre-analysis block in at least one pre-analysis block is obtained. The intra-prediction mode complexity of each pre-analysis block in at least one pre-analysis block is accumulated to obtain the intra-prediction mode complexity of the current video frame using full intra-prediction mode.
[0336] The steps for calculating the prediction complexity of frame A compared to frame B in the aforementioned embodiments are as follows:
[0337] Scale both A-frame and B-frame to their original preset size, for example, 1 / 4 (width and height scaled to 1 / 2 of their original size). Divide the scaled A-frame into at least one pre-analysis block, each block being 8×8. Iterate through at least one pre-analysis block. For the current pre-analysis block within the at least one pre-analysis block, determine the best-matching reference block in the scaled B-frame through motion search, and record the motion vector (mv) and inter-frame mode complexity at this point. The inter-frame mode complexity is calculated as follows: subtract the original pixel value of the reference block from the original pixel value of each position in the current pre-analysis block in the scaled A-frame to obtain the residual value at each position. Perform a Hadamard transform on the residual values at each position and sum them (satd) to obtain the inter-frame mode complexity of the current pre-analysis block.
[0338] If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity, then the current pre-analysis block is considered to be using inter-frame mode; otherwise, the current pre-analysis block is considered to be using intra-frame mode. The prediction complexity of each pre-analysis block is summed up to obtain the prediction complexity of frame A compared to frame B.
[0339] For example, taking a video frame sequence size of 32 as an example, refer to Figure 6(1) to determine the video frame sequence. Assuming that the upper limit of the number of frames supported by the image group to be encoded is 28, if the intra-complexity of the starting frame is greater than the complexity threshold, no further judgment is made, and the size of the image group to be encoded is 28. If the intra-complexity of the starting frame is less than or equal to the complexity threshold, it is further determined whether the size of the image group to be encoded needs to be divided into 16. If it is necessary to divide the size of the image group to be encoded into 16, refer to Figure 6(2); otherwise, it remains unchanged at 28. It is further determined whether it is necessary to divide the size of the image group to be encoded into 8. If it is necessary to divide the size of the image group to be encoded into 8, refer to Figure 6(3); otherwise, it remains unchanged at 16. It is further determined whether it is necessary to divide the size of the image group to be encoded into 4. If it is necessary to divide the size of the image group to be encoded into 4, refer to Figure 6(4); otherwise, it remains unchanged at 8. By encoding each video frame of the image group to be encoded, the encoded bitstream can be obtained. Beneficial effects:
[0340] Dividing the images to be encoded into groups based on sharpness allows for the selection of sharper, more textured video frames from videos with significant motion and blurriness. Referring to Table 2, the method described in this embodiment achieves a performance gain of over 20% on the MSU2022 test set. Further dividing the image groups into sizes of 16, 8, or 4, as shown in Table 3, yields an average performance gain of 0.5% on the MSU2022 test set, with a speeddown of 2.11%.
[0341] Table 2
[0342] Table 3
[0343] Figure 7 is a block diagram of a video encoding apparatus provided in an exemplary embodiment of this application. The video encoding apparatus 800 includes at least some of the following modules: an acquisition module 820, a processing module 840, a determination module 860, and an encoding module 880.
[0344] The acquisition module 820 is used to acquire a preset number of video frames from the video.
[0345] Processing module 840 is used to determine the last frame from the video frame sequence based on the video features of the video frame sequence;
[0346] The determining module 860 is used to determine the first frame to the last frame in the video frame sequence as an image group;
[0347] The encoding module 880 is used to encode each video frame in the image group to obtain an encoded bitstream.
[0348] In some embodiments, the processing module 840 is configured to:
[0349] Based on the sequence features of the video frame sequence, a first candidate frame is determined from the video frame sequence;
[0350] Based on the frame features of the first candidate frame and at least one preset frame, the last frame is determined from the first candidate frame and the at least one preset frame;
[0351] Wherein, the sequence number of the first candidate frame in the video frame sequence is used to indicate the upper limit of the number of frames of the image group to be encoded, the sequence number of the last frame in the video frame sequence is used to indicate the optimized number of frames of the image group to be encoded, the optimized number of frames is less than or equal to the upper limit of the number of frames, and the at least one preset frame is a preset frame in the video frame sequence that satisfies the image group structure condition and is located before the first candidate frame.
[0352] In some embodiments, the processing module 840 is configured to:
[0353] Statistical features are analyzed in the frame features of the first candidate frame and the at least one preset frame;
[0354] Based on the relationship between the statistical features and the threshold, the last frame is determined from the first candidate frame and the at least one preset frame;
[0355] The threshold is related to the statistical features and the at least one preset frame.
[0356] In some embodiments, the at least one preset frame includes a first preset frame, which is a preset frame in the video frame sequence preceding the first candidate frame; the statistical features include prediction complexity-related features of the first candidate frame and the first preset frame, respectively, and noise features of the first candidate frame; the processing module 840 is configured to:
[0357] Determine the prediction complexity of the first candidate frame compared to the starting frame; determine the prediction complexity of the first candidate frame compared to the previous frame; determine the prediction complexity of the first preset frame compared to the starting frame; determine the prediction complexity of the first preset frame compared to both the starting frame and the first candidate frame.
[0358] Calculate the first proportion of the specified pre-analysis block in the first candidate frame;
[0359] The prediction complexity of the first candidate frame compared to the reference frame (starting frame) is calculated as a first ratio to the intra-frame mode complexity of the first candidate frame.
[0360] The prediction complexity of the first candidate frame compared to the previous frame is calculated as a second ratio to the intra-frame mode complexity of the first candidate frame.
[0361] The prediction complexity of the first preset frame compared to the starting frame and the first candidate frame is calculated, and a third ratio is calculated between the prediction complexity of the first preset frame and the starting frame.
[0362] Analyze the noise level of the first candidate frame;
[0363] The intra-frame mode complexity is the prediction complexity when predicting the pixel value at each position of each pre-analysis block using intra-frame mode; the specified pre-analysis block is a pre-analysis block using inter-frame mode and whose motion vector is greater than a vector threshold, wherein the vector threshold is pre-configured and the pre-analysis block is a pre-divided pixel block.
[0364] In some embodiments, the threshold includes a noise threshold, a first threshold, a second threshold, a third threshold, and a fourth threshold; the processing module 840 is configured to:
[0365] If the noise level of the first candidate frame is less than the noise threshold, and the first proportion is greater than the first threshold, and the first ratio is greater than the second threshold, and the second ratio is less than the third threshold, and the third ratio is greater than the fourth threshold, then the first preset frame is determined as the last frame; otherwise, the first candidate frame is determined as the last frame.
[0366] The noise threshold, the first threshold, the second threshold, the third threshold, and the fourth threshold are all pre-configured and used to indicate the characteristics of the first candidate frame and the first preset frame.
[0367] In some embodiments, the at least one preset frame includes a first preset frame, and the statistical features include features of the first preset frame related to prediction complexity; the processing module 840 is configured to:
[0368] Determine the prediction complexity of the first preset frame compared to the starting frame;
[0369] Calculate the second proportion of the specified pre-analysis block in the first preset frame;
[0370] The third proportion of pre-analysis blocks using intra-frame mode in the first preset frame is calculated.
[0371] The fourth ratio of the prediction complexity of the first preset frame with the starting frame as the reference frame to the intra-frame mode complexity of the first preset frame is calculated.
[0372] The fifth ratio is calculated between the intra-frame mode complexity of the starting frame and the intra-frame mode complexity of the first preset frame.
[0373] In some embodiments, the at least one preset frame further includes a second preset frame, the second preset frame being a preset frame in the video frame sequence preceding the first preset frame; the threshold includes a fifth threshold, a sixth threshold, a first interval range, and a second interval range; the processing module 840 is configured to:
[0374] If the second ratio is less than the fifth threshold, the third ratio is less than the sixth threshold, the fourth ratio is within the first interval range, and the fifth ratio is within the second interval range, then the second preset frame is determined as the last frame; otherwise, the first preset frame remains unchanged as the last frame.
[0375] The fifth threshold, the sixth threshold, the first interval range, and the second interval range are pre-configured and used to indicate the characteristics of the first preset frame.
[0376] In some embodiments, the at least one preset frame includes a second preset frame, and the statistical features include features of the second preset frame related to prediction complexity; the processing module 840 is configured to:
[0377] Determine the prediction complexity of the second preset frame compared to the starting frame;
[0378] Determine the fourth proportion of the specified pre-analysis block in the second preset frame;
[0379] Determine the sixth ratio of the prediction complexity of the second preset frame when the starting frame is used as the reference frame to the intra-frame mode complexity of the second preset frame.
[0380] In some embodiments, the at least one preset frame further includes a third preset frame, the third preset frame being a preset frame in the current image frame preceding the second preset frame; the threshold includes a seventh threshold and a third interval range; the processing module 840 is configured to:
[0381] If the fourth ratio is greater than the seventh threshold and the sixth ratio is within the third interval range, the third preset frame is determined as the last frame; otherwise, the second preset frame remains unchanged as the last frame.
[0382] The seventh threshold and the third interval range are pre-configured and used to indicate the characteristics of the second preset frame.
[0383] In some embodiments, the processing module 840 is configured to:
[0384] The first candidate frame is divided into at least one pre-analysis block;
[0385] Using the starting frame as a reference frame, the at least one pre-analysis block is traversed, and for the current pre-analysis block in the at least one pre-analysis block, the corresponding reference block in the starting frame is determined;
[0386] Based on the reference block corresponding to the current pre-analysis block in the starting frame, determine the inter-frame mode complexity of the current pre-analysis block, and determine the motion vector of the current pre-analysis block;
[0387] If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, it is determined that the current pre-analysis block adopts inter-frame mode and the prediction complexity of the current pre-analysis block is inter-frame mode complexity; otherwise, it is determined that the current pre-analysis block adopts intra-frame mode and the prediction complexity of the current pre-analysis block is intra-frame mode complexity, and the process is repeated until the prediction complexity of each pre-analysis block of the at least one pre-analysis block is determined.
[0388] The prediction complexity of the first candidate frame is obtained by summing the prediction complexity of each pre-analysis block of the at least one pre-analysis block, compared to the prediction complexity of the starting frame.
[0389] In some embodiments, the processing module 840 is configured to:
[0390] Based on the reference block corresponding to the current pre-analysis block in the starting frame, the pixel value at each position of the current pre-analysis block is predicted using an inter-frame mode to obtain the predicted pixel value.
[0391] The predicted pixel value at each position of the current pre-analysis block is subtracted from the original pixel value at the corresponding position of the current analysis block to obtain the residual value at each position;
[0392] Perform a Hadamard transform on the residual value at each position and sum them to obtain the inter-frame mode complexity of the current pre-analysis block.
[0393] In some embodiments, the processing module 840 is configured to:
[0394] Determine the intra-frame mode complexity of each video frame in the video frame sequence;
[0395] Determine the maximum complexity among the intra-frame mode complexities of each video frame, and determine the video frame with the maximum complexity as the maximum complexity frame of the video frame sequence.
[0396] Determine the ratio of the maximum complexity frame to a specified pre-analysis block of the starting frame;
[0397] If the maximum complexity is less than the eighth threshold, the sequence number of the maximum complexity frame in the video frame sequence is greater than the ninth threshold, and the proportion of the specified pre-analysis block is greater than the tenth threshold, the maximum complexity frame is determined as the first candidate frame.
[0398] The eighth threshold is calculated based on the image resolution, while the ninth and tenth thresholds are pre-configured.
[0399] In some embodiments, the processing module 840 is further configured to:
[0400] Determine the last frame of the first scene in the video frame sequence;
[0401] The first candidate frame is determined from the last frame of the scene and the frame with the highest complexity.
[0402] In some embodiments, the processing module 840 is configured to:
[0403] The current video frame is divided into at least one pre-analysis block;
[0404] Traverse the at least one pre-analysis block, and for the current pre-analysis block in the at least one pre-analysis block, use intra-frame mode to predict the pixel value at each position of the current pre-analysis block to obtain the predicted pixel value;
[0405] Subtract the predicted pixel value from the original pixel value at each position of the current pre-analysis block to obtain the residual value at each position;
[0406] Perform a Hadamard transform on the residual value at each position and sum them to obtain the intra-mode complexity of the current pre-analysis block. Repeat this process until the intra-mode complexity of each pre-analysis block of the at least one pre-analysis block is obtained.
[0407] The intra-frame mode complexity of the current video frame is obtained by summing the intra-frame mode complexity of each pre-analysis block of the at least one pre-analysis block.
[0408] In some embodiments, the processing module 840 is configured to:
[0409] Determine the specified pre-analysis block of the maximum complexity frame compared to the starting frame;
[0410] The proportion of the specified pre-analysis blocks in the maximum complexity frame is determined based on the number of specified pre-analysis blocks and the number of pre-analysis blocks in the maximum complexity frame.
[0411] In some embodiments, the processing module 840 is configured to:
[0412] Using the starting frame as a reference frame, determine the reference block in the starting frame for each pre-analysis block in the maximum complexity frame;
[0413] Based on the reference block corresponding to each pre-analysis block in the maximum complexity frame, the inter-frame mode complexity of each pre-analysis block in the maximum complexity frame is determined, and the motion vector of each pre-analysis block in the maximum complexity frame is determined.
[0414] If the inter-frame mode complexity of the current pre-analysis block of the frame with the highest complexity is less than the intra-frame mode complexity of the current pre-analysis block, it is determined that the current pre-analysis block adopts inter-frame mode; otherwise, it is determined that the current pre-analysis block adopts intra-frame mode.
[0415] The pre-analysis block in the frame with the highest complexity that adopts the inter-frame mode and whose motion vector is greater than the vector threshold is determined as the designated pre-analysis block.
[0416] In some embodiments, the processing module 840 is further configured to:
[0417] If the starting frame meets the segmentation conditions, the step of determining the ending frame from the first candidate frame and the at least one preset frame based on the frame features of the first candidate frame and at least one preset frame is executed; otherwise, the first candidate frame is determined as the ending frame.
[0418] In some embodiments, the processing module 840 is configured to:
[0419] Determine the intra-frame mode complexity of the starting frame;
[0420] If the total complexity of the starting frame is less than or equal to the complexity threshold, the starting frame satisfies the partitioning condition.
[0421] The complexity threshold is related to the number of pre-analysis blocks in the starting frame.
[0422] In some embodiments, the acquisition module 820 is used for:
[0423] Read the video frame sequence of the first frame number in the video to be encoded;
[0424] The video frame sequence of the first frame number is stored in the pre-analysis queue;
[0425] The video frame sequence with the preset number of frames in the pre-analysis queue is determined as the video frame sequence;
[0426] The maximum value of the first frame number is the pre-configured maximum number of read frames.
[0427] In some embodiments, the acquisition module 820 is further configured to:
[0428] Each video frame in the image group to be encoded is removed from the pre-analysis queue to obtain the removed pre-analysis queue, which stores the video frame sequence of the second frame number.
[0429] After encoding each video frame in the image group to be encoded to obtain the encoded bitstream, the pre-analysis queue after being pushed out is supplemented with a video frame sequence of the first number of frames, and the step of determining the video frames of the preset number of frames in the pre-analysis queue as the video frame sequence is repeated.
[0430] It should be noted that the specific limitations of the one or more video encoding devices 800 provided above can be found in the limitations of the video encoding method above, and will not be repeated here. Each module of the above device can be implemented entirely or partially by software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0431] This application also provides an encoder, which includes an encoding unit storing a computer program. The computer program is loaded and executed by the encoding unit to implement the video encoding methods provided in the above-described method embodiments.
[0432] For example, FIG8 is a structural block diagram of an encoder provided in an exemplary embodiment of this application. The encoder 900 includes: an encoding unit 901, a storage unit 902, an input interface 903, and an output interface 904.
[0433] Input interface 903 is the input section of encoder 900, used to receive raw video to be encoded. The source of the video to be encoded includes at least one of the following: various sensors, audio and video devices, and other data sources.
[0434] The encoding unit 901 is the core component of the encoder 900, used to encode the video to be encoded. During the encoding process, the encoding unit 901 converts the video to be encoded into a specific format according to a preset encoding standard, such as the HEVC standard, specifically the H.26x series or MPEG series standards, to obtain the encoded bitstream. After encoding, the video to be encoded can not only compress the data size, but also improve its anti-interference ability to a certain extent, and improve the stability and reliability during transmission or storage.
[0435] In some embodiments, the encoding unit 901 may be implemented in at least one hardware form selected from chips, hardware circuits, logic circuits, processors, digital signal processing (DSP), field-programmable gate arrays (FPGA), and programmable logic arrays (PLA).
[0436] Storage unit 902 is used to temporarily store the encoded bitstream, which facilitates the rapid reading of the encoded bitstream for subsequent transmission or storage. The capacity and read / write speed of storage unit 902 can be set according to actual technical needs.
[0437] In some embodiments, storage unit 902 may include one or more computer-readable storage media, which may be non-transitory. Storage unit 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices.
[0438] Output interface 904 is the output section of encoder 900, used to output the encoded bitstream to external devices or networks. Output interface 904 needs to be compatible with various standards and protocols to facilitate the output of the encoded bitstream to external devices or networks.
[0439] In some embodiments, the encoding unit 901, storage unit 902, input interface 903, and output interface 904 can be connected via a bus or signal lines. Various external devices can be connected to the input interface 903 and output interface 904 via a bus, signal lines, or circuit board. The input interface 903 and output interface 904 can be used to connect at least one input / output (I / O) related external device to the encoding unit 901 and storage unit 902. In some examples, the encoding unit 901, storage unit 902, input interface 903, and output interface 904 are integrated on the same chip or circuit board; in other examples, any one or two of the encoding unit 901, storage unit 902, input interface 903, and output interface 904 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.
[0440] Those skilled in the art will understand that the structure shown in FIG8 does not constitute a limitation on the encoder and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0441] This application also provides a computer device, which includes a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the video encoding method provided in the above-described method embodiments.
[0442] For example, Figure 9 is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Optionally, the computer device is a server 1000.
[0443] Typically, server 1000 includes a processor 1001 and memory 1002.
[0444] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0445] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one instruction, which is executed by the processor 1001 to implement the video encoding method provided in the above-described method embodiments.
[0446] In some embodiments, the server 1000 may optionally include an input interface 1003 and an output interface 1004. The processor 1001, memory 1002, and input interfaces 1003 and 1004 can be connected via a bus or signal lines. Various external devices can be connected to the input interfaces 1003 and 1004 via a bus, signal lines, or a circuit board. The input interfaces 1003 and 1004 can be used to connect at least one input / output (I / O) related external device to the processor 1001 and memory 1002. In some embodiments, the processor 1001, memory 1002, and input interfaces 1003 and 1004 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, memory 1002, and input interfaces 1003 and 1004 can be implemented on separate chips or circuit boards, and this application does not limit this.
[0447] Those skilled in the art will understand that the structure shown in Figure 9 does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.
[0448] In an exemplary embodiment, this application also provides a chip, which includes programmable logic circuits and / or computer instructions, and is used to implement the video encoding methods provided in the above-described method embodiments when the chip is running on a computer device.
[0449] This application also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the video encoding methods provided in the above-described method embodiments.
[0450] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the processor of the computer device to load and execute the video encoding method provided in the above-described method embodiments.
[0451] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0452] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0453] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0454] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A video encoding method, executed in an electronic device, the method comprising: Obtain a video frame sequence with a preset number of frames; Based on the video features of the video frame sequence, the last frame is determined from the video frame sequence; The first frame to the last frame in the video frame sequence are determined as an image group; Each video frame in the image group is encoded to obtain an encoded bitstream.
2. The method according to claim 1, characterized in that, The step of determining the last frame from the video frame sequence based on video features of the video frame sequence includes: Based on the sequence characteristics of the video frame sequence, a first candidate frame is determined from the video frame sequence; wherein the sequence number of the first candidate frame in the video frame sequence is equal to the upper limit of the number of frames in the image group; Based on the frame features of the first candidate frame and the frame features of at least one preset frame, the last frame is determined from the first candidate frame and the at least one preset frame; wherein each preset frame is located before the first candidate frame in the video frame sequence, and its sequence number is equal to a predefined candidate value for image group size.
3. The method according to claim 2, characterized in that, The step of determining the last frame from the first candidate frame and the at least one preset frame based on the frame features of the first candidate frame and the frame features of at least one preset frame includes: The frame features of the first candidate frame and the frame features of the at least one preset frame are statistically analyzed to obtain statistical features; Based on the relationship between the statistical features and the threshold, the last frame is determined from the first candidate frame and the at least one preset frame; The threshold is related to the statistical features and the at least one preset frame.
4. The method according to claim 3, characterized in that, The at least one preset frame includes a first preset frame; the statistical features include features related to prediction complexity of the first candidate frame and the first preset frame, respectively, and noise features of the first candidate frame; The statistical features in the frame features of the first candidate frame and the at least one preset frame include: Determine the prediction complexity of the first candidate frame compared to the starting frame; Determine the prediction complexity of the first candidate frame compared to the previous frame; Determine the prediction complexity of the first preset frame compared to the starting frame; Determine the prediction complexity of the first preset frame compared to the starting frame and the first candidate frame; Calculate the first proportion of the specified pre-analysis block in the first candidate frame; Determine the first ratio of the prediction complexity of the first candidate frame compared to the starting frame to the first ratio of the intra-frame prediction complexity of the first candidate frame; The second ratio of the prediction complexity of the first candidate frame compared to the previous frame to the intra-frame mode complexity of the first candidate frame is calculated. The prediction complexity of the first preset frame compared to the starting frame and the first candidate frame is calculated as a third ratio to the prediction complexity of the starting frame compared to the first preset frame. Analyze the noise level of the first candidate frame; The specified pre-analysis block is a pre-analysis block that uses an inter-frame mode and whose motion vector is greater than a vector threshold. The vector threshold is pre-configured, and the pre-analysis block is a pre-divided pixel block.
5. The method according to claim 4, characterized in that, The thresholds include a noise threshold, a first threshold, a second threshold, a third threshold, and a fourth threshold; The step of determining the last frame from the first candidate frame and the at least one preset frame based on the relationship between the statistical features and the threshold includes: If the noise level of the first candidate frame is less than the noise threshold, and the first proportion is greater than the first threshold, and the first ratio is greater than the second threshold, and the second ratio is less than the third threshold, and the third ratio is greater than the fourth threshold, then the first preset frame is determined as the last frame; otherwise, the first candidate frame is determined as the last frame. The noise threshold, the first threshold, the second threshold, the third threshold, and the fourth threshold are all pre-configured.
6. The method according to any one of claims 3 to 5, characterized in that, The at least one preset frame includes a first preset frame, and the statistical features include features of the first preset frame related to prediction complexity. The statistical features in the frame features of the first candidate frame and the at least one preset frame further include: Determine the prediction complexity of the first preset frame compared to the starting frame; Calculate the second proportion of the specified pre-analysis block in the first preset frame; The third proportion of pre-analysis blocks using intra-frame mode in the first preset frame is calculated. The fourth ratio of the prediction complexity of the first preset frame compared to the starting frame to the intra-frame mode complexity of the first preset frame is calculated. The fifth ratio is calculated between the intra-frame mode complexity of the starting frame and the intra-frame mode complexity of the first preset frame.
7. The method according to claim 6, characterized in that, The at least one preset frame further includes a second preset frame, which is a preset frame in the video frame sequence that precedes the first preset frame; the threshold includes a fifth threshold, a sixth threshold, a first interval range, and a second interval range. The step of determining the last frame from the first candidate frame and the at least one preset frame based on the relationship between the statistical features and the threshold includes: If the second ratio is less than the fifth threshold, the third ratio is less than the sixth threshold, the fourth ratio is within the first interval range, and the fifth ratio is within the second interval range, then the second preset frame is determined as the last frame; otherwise, the first preset frame remains unchanged as the last frame. The fifth threshold, the sixth threshold, the first interval range, and the second interval range are pre-configured.
8. The method according to any one of claims 3 to 7, characterized in that, The at least one preset frame includes a second preset frame, and the statistical features include features of the second preset frame related to prediction complexity; The statistical features in the frame features of the first candidate frame and the at least one preset frame further include: Determine the prediction complexity of the second preset frame compared to the starting frame; The fourth proportion of the specified pre-analysis block in the second preset frame is statistically analyzed; The sixth ratio of the prediction complexity of the second preset frame to that of the starting frame and the intra-frame mode complexity of the second preset frame is calculated.
9. The method according to claim 8, characterized in that, The at least one preset frame further includes a third preset frame, which is a preset frame in the current image frame that precedes the second preset frame; the threshold includes a seventh threshold and a third interval range; The step of determining the last frame from the first candidate frame and the at least one preset frame based on the relationship between the statistical features and the threshold includes: If the fourth ratio is greater than the seventh threshold and the sixth ratio is within the third interval range, the third preset frame is determined as the last frame; otherwise, the second preset frame remains unchanged as the last frame. The seventh threshold and the third interval range are pre-configured.
10. The method according to claim 4, characterized in that, Determining the prediction complexity of the first candidate frame compared to the starting frame includes: The first candidate frame is divided into at least one pre-analysis block; Using the starting frame as a reference frame, the at least one pre-analysis block is traversed, and for the current pre-analysis block in the at least one pre-analysis block, the corresponding reference block in the starting frame is determined; Based on the reference block corresponding to the current pre-analysis block in the starting frame, determine the inter-frame mode complexity of the current pre-analysis block, and determine the motion vector of the current pre-analysis block; If the inter-frame mode complexity of the current pre-analysis block is less than the intra-frame mode complexity of the current pre-analysis block, it is determined that the current pre-analysis block adopts inter-frame mode and the prediction complexity of the current pre-analysis block is inter-frame mode complexity; otherwise, it is determined that the current pre-analysis block adopts intra-frame mode and the prediction complexity of the current pre-analysis block is intra-frame mode complexity, and the process is repeated until the prediction complexity of each pre-analysis block of the at least one pre-analysis block is determined. The prediction complexity of the first candidate frame is obtained by summing the prediction complexity of each pre-analysis block of the at least one pre-analysis block, compared to the prediction complexity of the starting frame.
11. The method according to claim 10, characterized in that, The step of determining the inter-frame mode complexity of the current pre-analysis block based on the reference block corresponding to the current pre-analysis block in the starting frame includes: Based on the reference block corresponding to the current pre-analysis block in the starting frame, the pixel value at each position of the current pre-analysis block is predicted using an inter-frame mode to obtain the predicted pixel value. The predicted pixel value at each position of the current pre-analysis block is subtracted from the original pixel value at the corresponding position of the current analysis block to obtain the residual value at each position; Perform a Hadamard transform on the residual value at each position and sum them to obtain the inter-frame mode complexity of the current pre-analysis block.
12. The method according to any one of claims 2 to 11, characterized in that, The step of determining the first candidate frame from the video frame sequence based on the sequence features of the video frame sequence includes: Determine the intra-frame mode complexity of each video frame in the video frame sequence; Determine the maximum complexity among the intra-frame mode complexities of the video frames in the video frame sequence, and determine the video frame with the maximum complexity as the maximum complexity frame of the video frame sequence. Determine the proportion of the specified pre-analysis block with the starting frame as the reference frame for the maximum complexity frame; If the maximum complexity is less than the eighth threshold, the sequence number of the maximum complexity frame in the video frame sequence is greater than the ninth threshold, and the proportion of the specified pre-analysis block is greater than the tenth threshold, the maximum complexity frame is determined as the first candidate frame. The eighth threshold is calculated based on the image resolution, while the ninth and tenth thresholds are pre-configured.
13. The method according to claim 12, characterized in that, The method further includes: Determine the last frame of the first scene in the video frame sequence; The first candidate frame is determined from the last frame of the scene and the frame with the highest complexity.
14. The method according to claim 12, characterized in that, Determining the intra-frame mode complexity of each video frame in the video frame sequence includes: Divide the current video frame into at least one pre-analysis block; Traverse the at least one pre-analysis block, and for the current pre-analysis block in the at least one pre-analysis block, use intra-frame mode to predict the pixel value at each position of the current pre-analysis block to obtain the predicted pixel value; Subtract the predicted pixel value from the original pixel value at each position of the current pre-analysis block to obtain the residual value at each position; Perform a Hadamard transform on the residual value at each position and sum them to obtain the intra-mode complexity of the current pre-analysis block. Repeat this process until the intra-mode complexity of each pre-analysis block of the at least one pre-analysis block is obtained. The intra-frame mode complexity of the current video frame is obtained by summing the intra-frame mode complexity of each pre-analysis block of the at least one pre-analysis block.
15. The method according to claim 12, characterized in that, The determination of the proportion of the specified pre-analysis block with the starting frame as the reference frame for the maximum complexity frame includes: Determine the specified pre-analysis block with the starting frame as the reference frame for the frame with the maximum complexity frame; The proportion of the specified pre-analysis blocks in the maximum complexity frame is determined based on the number of specified pre-analysis blocks and the number of pre-analysis blocks in the maximum complexity frame.
16. A video encoding apparatus, characterized in that, The device includes: The acquisition module is used to acquire a video frame sequence with a preset number of frames from the video to be encoded; The processing module is used to determine the last frame from the video frame sequence based on the video features of the video frame sequence; The determining module is used to determine the first frame to the last frame in the video frame sequence as an image group; The encoding module is used to encode each video frame in the image group to obtain an encoded bitstream.
17. An encoder, characterized in that, The encoder includes an encoding unit that stores a computer program, which is loaded and executed by the encoding unit to implement the video encoding method as described in any one of claims 1 to 15.
18. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the video encoding method as described in any one of claims 1 to 15.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is loaded and executed by a processor to implement the video encoding method as described in any one of claims 1 to 15.
20. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, from which a processor retrieves the computer instructions, causing the processor to load and execute them to implement the video encoding method as described in any one of claims 1 to 15.