Video encoding method and device, electronic equipment and storage medium
By dynamically adjusting the frame structure and picture group division according to the video content, the problem of reduced encoding quality caused by the fixed GOP encoding scheme in the existing technology is solved, and a more efficient video encoding effect is achieved.
Patent Information
- Application Number
- CN202310101276.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-28
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-01-28
AI Technical Summary
In existing video coding technologies, the fixed-length and fixed-frame-structure GOP coding scheme cannot adapt to the actual changes in video content, resulting in a decrease in coding quality.
Based on the actual content of the target video, determine the minimum cost frame structure of the video frame sequence, and divide the video frame sequence into coded picture groups based on the frame structure. Ensure that the number of video frames in each picture group does not exceed the number threshold, and encode the reference relationship between video frames indicated by the target frame structure.
It improves video encoding quality by dynamically adjusting frame structure and picture group division, thereby enhancing the similarity matching between video frames and optimizing encoding efficiency.
Smart Images

Figure CN116095326B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a video encoding method and device, electronic equipment and storage medium. BACKGROUND
[0002] A video is composed of continuous video frames. Since there is similarity between the continuous video frames, in order to facilitate storage and transmission of the video, the video needs to be encoded to reduce the memory or bandwidth occupied by the video. The scheme of encoding the video is generally as follows: the continuous video frames in the video are usually divided into multiple GOPs (Group Of Pictures), and the GOPs are encoded based on the reference relationship between the video frames indicated by the frame structure of the GOP.
[0003] In the related art, the GOP usually adopts a fixed length and a fixed frame structure to determine the reference relationship between the video frames. For example, the HEVC (High Efficiency Video Coding) standard generally encodes the GOP with a fixed length of 16 frames and a fixed frame structure of a two-way symmetrical frame structure.
[0004] However, the content of the video is not presented in the above fixed length, for example, there are 12 continuous frames in a picture group that are similar, but the 13th-16th frames are not similar to the previous 12 frames, and if the above scheme is used for encoding, the quality of video encoding will be reduced. SUMMARY
[0005] The present disclosure provides a video encoding method, device, electronic equipment and storage medium, which can encode a target video according to the actual content of the target video. The technical scheme of the present disclosure is as follows:
[0006] According to an aspect of an embodiment of the present disclosure, a video encoding method is provided, comprising:
[0007] For any video frame sequence in the multiple video frame sequences of the target video, a target frame structure with minimum cost is determined from multiple frame structures of the video frame sequence, the video frame sequence includes no more than a target number of video frames, and the cost is used to represent the similarity between the video frames;
[0008] An encoding picture group of the video frame sequence is determined from multiple picture groups determined based on the target frame structure, the number of video frames in the encoding picture group is not greater than a number threshold, and the cost of the encoding picture group is less than the cost of other picture groups in the multiple picture groups;
[0009] The target video is encoded based on the encoding picture group of each video frame sequence in the multiple video frame sequences of the target video.
[0010] In some embodiments, the target frame structure is used to indicate positions of bidirectional prediction frames and forward prediction frames in the video frame sequence; and the coded group of pictures of the video frame sequence is determined from the multiple groups of pictures determined based on the target frame structure, including:
[0011] Based on the positions of bidirectional prediction frames and forward prediction frames in the video frame sequence indicated by the target frame structure, multiple groups of pictures are determined in the video frame sequence, the first video frame and the last video frame in the group of pictures are forward prediction frames, the video frames in the group of pictures except the first video frame and the last video frame are bidirectional prediction frames, the first video frame in the group of pictures and the last video frame in the previous group of pictures are the same forward prediction frame; and the number of video frames in the group of pictures is not greater than the number threshold.
[0012] The first group of pictures and the second group of pictures in the multiple groups of pictures are spliced to obtain a first target group of pictures, the first video frame and the last video frame in the first target group of pictures are forward prediction frames, and the video frames in the first target group of pictures except the first video frame and the last video frame are bidirectional prediction frames.
[0013] In a case where the number of video frames in the first target group of pictures is greater than the number threshold, and the cost of the first group of pictures is less than the cost of the first target group of pictures, the first group of pictures is taken as the coded group of pictures of the video frame sequence.
[0014] In some embodiments, the splicing of the first group of pictures and the second group of pictures in the multiple groups of pictures to obtain the first target group of pictures includes:
[0015] The same prediction frames in the first group of pictures and the second group of pictures are replaced by a target bidirectional prediction frame;
[0016] The first group of pictures after replacement and the second group of pictures after replacement are merged to obtain the first target group of pictures, and the first target group of pictures includes one target bidirectional prediction frame.
[0017] In some embodiments, the method further includes:
[0018] In a case where the number of video frames in the first target group of pictures is not greater than the number threshold, the first n groups of pictures are spliced to obtain a second target group of pictures.
[0019] In a case where a number of video frames in the second target group of pictures is greater than the number threshold, a third target group of pictures is determined as a coded group of pictures of the sequence of video frames, the third target group of pictures being spliced from a first n-1 groups of pictures, n being a positive integer greater than 2.
[0020] In some embodiments, the case where the number of video frames in the second target group of pictures is greater than the number threshold includes:
[0021] In the case where the number of video frames in the second target group of pictures is greater than the number threshold, a cost of a fourth target group of pictures is obtained, the fourth target group of pictures being spliced from a first n-2 groups of pictures;
[0022] In the case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures, the third target group of pictures is determined as the coded group of pictures of the sequence of video frames.
[0023] In some embodiments, the method further includes:
[0024] In the case where the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures, the fourth target group of pictures is determined as the coded group of pictures of the sequence of video frames.
[0025] In some embodiments, the determining the target frame structure with the minimum cost from the multiple frame structures of the sequence of video frames includes:
[0026] The sequence of video frames is determined from the target video;
[0027] The sequence of video frames is divided into multiple groups of video frames, different groups of video frames including different numbers of continuous video frames, a first video frame of the groups of video frames being a first video frame of the sequence of video frames;
[0028] For any group of video frames, a frame structure with a minimum cost from multiple frame structures of the group of video frames is determined as a target cost frame structure of the group of video frames;
[0029] Based on the number threshold and the target cost frame structure of each group of video frames in the multiple groups of video frames, multiple frame structures of the sequence of video frames are determined, a number of the multiple frame structures of the sequence of video frames not being greater than the number threshold;
[0030] A frame structure with a minimum cost from the multiple frame structures of the sequence of video frames is determined as a target frame structure of the sequence of video frames.
[0031] In some embodiments, the determining, for any video frame group, the target cost frame structure of the video frame group from the plurality of frame structures of the video frame group, comprises:
[0032] In a case where the number of video frames in the video frame group is greater than the number threshold, from the plurality of first video frame groups, determining a number threshold of second video frame groups, the number of video frames in the first video frame groups is less than the number of video frames in the video frame group, and the number of video frames in the second video frame groups is greater than the number of video frames in the remaining first video frame groups;
[0033] Based on the target cost frame structures of the number threshold of second video frame groups, determining the number threshold of cost frame structures of the video frame group;
[0034] From the number threshold of cost frame structures, determining the target cost frame structure of the video frame group.
[0035] In some embodiments, the method further comprises:
[0036] In a case where the number of video frames in the video frame group is not greater than the number threshold, based on the target cost frame structures of the plurality of first video frame groups, determining a plurality of cost frame structures of the video frame group;
[0037] From the plurality of cost frame structures, determining the target cost frame structure of the video frame group.
[0038] In some embodiments, the video frame sequence comprises m consecutive video frames, m is a positive integer greater than 1;
[0039] The determining, based on the number threshold and the target cost frame structure of each video frame group in the plurality of video frame groups, the plurality of frame structures of the video frame sequence, comprises:
[0040] In a case where m is greater than the number threshold, based on the number threshold and the target cost frame structure of each video frame group in the plurality of video frame groups, determining a number threshold of frame structures of the video frame sequence;
[0041] In a case where m is not greater than the number threshold, based on the number threshold and the target cost frame structure of each video frame group in the plurality of video frame groups, determining m-1 frame structures of the video frame sequence.
[0042] In some embodiments, the method further comprises:
[0043] For any frame structure, the residual value between any video frame in the frame structure and the video frame closest to the reference distance of the video frame is taken as the residual value of the video frame;
[0044] Summing residual values of the plurality of video frames in the frame structure as a cost of the frame structure.
[0045] According to another aspect of the embodiments of the present disclosure, a video encoding apparatus is provided, comprising:
[0046] A first determining unit configured to, for any video frame sequence in a plurality of video frame sequences of a target video, determine a target frame structure with a minimum cost from a plurality of frame structures of the video frame sequence, the video frame sequence comprising no more than a target number of video frames, the cost being used to represent similarity between video frames;
[0047] A second determining unit configured to determine a coded group of pictures of the video frame sequence from a plurality of groups of pictures determined based on the target frame structure, a number of video frames in the coded group of pictures being no more than a number threshold, and a cost of the coded group of pictures being less than costs of other groups of pictures in the plurality of groups of pictures;
[0048] An encoding unit configured to encode the target video based on the coded group of pictures of each video frame sequence in the plurality of video frame sequences of the target video.
[0049] In some embodiments, the target frame structure is used to indicate positions of bi-predictive frames and forward-predictive frames in the video frame sequence; and the second determining unit comprises:
[0050] A first determining sub-unit configured to determine a plurality of groups of pictures in the video frame sequence based on the positions of bi-predictive frames and forward-predictive frames in the video frame sequence indicated by the target frame structure, a first video frame and a last video frame in the group of pictures being forward-predictive frames, and video frames other than the first video frame and the last video frame in the group of pictures being bi-predictive frames, and the first video frame in the group of pictures and a last video frame in a previous group of pictures being the same forward-predictive frame; and a number of video frames in the group of pictures being no more than the number threshold;
[0051] A splicing sub-unit configured to splice a first group of pictures and a second group of pictures in the plurality of groups of pictures to obtain a first target group of pictures, a first video frame and a last video frame in the first target group of pictures being forward-predictive frames, and video frames other than the first video frame and the last video frame in the first target group of pictures being bi-predictive frames;
[0052] A second determining sub-unit configured to, in a case where a number of video frames in the first target group of pictures is greater than the number threshold, and a cost of the first group of pictures is less than a cost of the first target group of pictures, take the first group of pictures as the coded group of pictures of the video frame sequence.
[0053] In some embodiments, the splicer unit is configured to replace the same predicted frame in the first group of pictures and the second group of pictures with a target bi-predicted frame; and merge the replaced first group of pictures and the replaced second group of pictures to obtain a first target group of pictures, the first target group of pictures comprising the target bi-predicted frame.
[0054] In some embodiments, the splicer unit is configured to splice the first n groups of pictures to obtain a second target group of pictures, in a case that a number of video frames in the first target group of pictures is not greater than the number threshold; and the splicer unit is configured to take a third target group of pictures as the coding group of pictures of the video frame sequence, in a case that a number of video frames in the second target group of pictures is greater than the number threshold, the third target group of pictures being spliced from the first n-1 groups of pictures, n being a positive integer greater than 2.
[0055] In some embodiments, the second determination unit is configured to obtain a cost of a fourth target group of pictures spliced from the first n-2 groups of pictures, in a case that the number of video frames in the second target group of pictures is greater than the number threshold; and take the third target group of pictures as the coding group of pictures of the video frame sequence, in a case that the cost of the third target group of pictures is less than the cost of the fourth target group of pictures.
[0056] In some embodiments, the second determination unit is configured to take the fourth target group of pictures as the coding group of pictures of the video frame sequence, in a case that the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures.
[0057] In some embodiments, the first determination unit comprises:
[0058] A third determination unit is configured to determine the video frame sequence from the target video.
[0059] A division unit is configured to divide the video frame sequence into a plurality of groups of video frames, different groups of video frames comprising different numbers of continuous video frames, a first video frame of the groups of video frames being a first video frame of the video frame sequence.
[0060] A fourth determination unit is configured to take a frame structure with a minimum cost in a plurality of frame structures of any group of video frames as a target cost frame structure of the group of video frames.
[0061] The fifth determining sub-unit is configured to determine a plurality of frame structures of the video frame sequence based on the quantity threshold and the target cost frame structure of each of the plurality of video frame groups, wherein the number of the plurality of frame structures of the video frame sequence is not greater than the quantity threshold.
[0062] The sixth determining sub-unit is configured to determine, as the target frame structure of the video frame sequence, a frame structure with the minimum cost among the plurality of frame structures of the video frame sequence.
[0063] In some embodiments, the fourth determining sub-unit is configured to, in a case where the number of video frames in the video frame group is greater than the quantity threshold, determine, from a plurality of first video frame groups, a quantity threshold number of second video frame groups, wherein the number of video frames in the first video frame groups is less than the number of video frames in the video frame group, and the number of video frames in the second video frame groups is greater than the number of video frames in the remaining first video frame groups; determine a quantity threshold number of cost frame structures of the video frame group based on the target cost frame structure of the quantity threshold number of second video frame groups; and determine, from the quantity threshold number of cost frame structures, the target cost frame structure of the video frame group.
[0064] In some embodiments, the fourth determining sub-unit is configured to, in a case where the number of video frames in the video frame group is not greater than the quantity threshold, determine a plurality of cost frame structures of the video frame group based on the target cost frame structure of the plurality of first video frame groups; and determine, from the plurality of cost frame structures, the target cost frame structure of the video frame group.
[0065] In some embodiments, the video frame sequence includes m consecutive video frames, m being a positive integer greater than 1; the fifth determining sub-unit is configured to, in a case where m is greater than the quantity threshold, determine a quantity threshold number of frame structures of the video frame sequence based on the quantity threshold and the target cost frame structure of each of the plurality of video frame groups; and in a case where m is not greater than the quantity threshold, determine m-1 frame structures of the video frame sequence based on the quantity threshold and the target cost frame structure of each of the plurality of video frame groups.
[0066] In some embodiments, the apparatus further includes:
[0067] The third determining unit is configured to, for any frame structure, determine, as the residual value of any video frame in the frame structure, a residual value between the video frame and a video frame closest to the video frame in reference distance.
[0068] The fourth determining unit is configured to determine, as the cost of the frame structure, a sum of residual values of a plurality of video frames in the frame structure.
[0069] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, which includes:
[0070] one or more processors;
[0071] a memory for storing the processor-executable program code;
[0072] wherein the processor is configured to execute the program code to implement the above-mentioned video encoding method.
[0073] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which, when the program code in the computer-readable storage medium is executed by a processor of an electronic device, enables the electronic device to perform the above-mentioned video encoding method.
[0074] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, which includes computer programs / instructions that, when executed by a processor, implement the above-mentioned video encoding method.
[0075] The embodiments of the present disclosure provide a video encoding method, which can determine a target frame structure with minimum cost from a plurality of frame structures of a sequence of video frames of a target video according to actual content of the target video. The target frame structure can indicate a reference relationship between video frames with high similarity in the target video. Based on the target frame structure, a coding picture group with a length not exceeding a quantity threshold is determined from the sequence of video frames, and the target video is encoded by using the coding picture groups of a plurality of sequences of video frames of the target video, which can encode the target video according to actual content of the target video, and improve the quality of encoding of the target video.
[0076] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not intended to limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0077] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, and do not limit the present disclosure.
[0078] Figure 1 is a schematic diagram of an implementation environment according to an exemplary embodiment.
[0079] Figure 2 is a flowchart of a video encoding method according to an exemplary embodiment.
[0080] Figure 3 is a flowchart of another video encoding method according to an exemplary embodiment.
[0081] Figure 4FIG. 1 is a diagram of a picture group of a two-part symmetrical frame structure according to an example embodiment.
[0082] Figure 5 FIG. 2 is a diagram of a picture group of an asymmetrical frame structure according to an example embodiment.
[0083] Figure 6 FIG. 3 is a block diagram of an apparatus for video coding according to an example embodiment.
[0084] Figure 7 FIG. 4 is another block diagram of an apparatus for video coding according to an example embodiment.
[0085] Figure 8 FIG. 5 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION
[0086] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings.
[0087] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0088] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the target video involved in the present application is obtained under sufficient authorization.
[0089] The electronic device can be provided as a terminal or a server, when the electronic device is provided as a terminal, the operations performed by the method of video coding can be implemented by the terminal; when provided as a server, the operations performed by the method of video coding can be implemented by the server; the method can also be implemented by the server and the terminal interacting; the method can also be implemented by the terminal sending a video coding request to the server, and the server coding the target video.
[0090] Figure 1 is a schematic diagram of an implementation environment according to an example embodiment. Referring to Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.
[0091] The terminal 101 can be at least one of a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop computer. The terminal 101 can have an application installed and running thereon, and a user can log in to the application through the terminal 101 to obtain services provided by the application. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.
[0092] The terminal 101 can be referred to as one of a plurality of terminals, and the terminal 101 is used as an example in the embodiment. It can be understood by those skilled in the art that the number of terminals can be more or less. For example, the terminals can be several, or the terminals can be tens or hundreds, or more, and the number of terminals and the type of equipment are not limited in the embodiment.
[0093] The server 102 is a stand-alone physical server, and can also be a server cluster or a distributed system composed of a plurality of physical servers, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and basic cloud computing services such as big data and artificial intelligence platforms. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or the server 102 and the terminal 101 cooperatively compute in a distributed computing architecture. The server 102 can be connected to the terminal 101 and other terminals through a wireless network or a wired network, and the number of servers can be more or less, which is not limited in the embodiment. Of course, the server 102 can also include other functional servers to provide more comprehensive and diversified services.
[0094] Figure 2 is a flowchart of a video encoding method according to an example embodiment, as shown in Figure 2As shown, the method is performed by an electronic device, and includes the following steps:
[0095] In step S201, for any video frame sequence in the plurality of video frame sequences of the target video, the electronic device determines a target frame structure with a minimum cost from a plurality of frame structures of the video frame sequence, the video frame sequence includes no more than a target number of video frames, and the cost is used to represent the similarity between the video frames.
[0096] In the embodiments of the present disclosure, the target video is a video to be encoded. The target video includes a plurality of continuous video frames. The electronic device determines no more than a target number of continuous video frames from the target video, and takes the determined video frames as a video frame sequence. The target number can be set according to the encoding buffer of the electronic device. The larger the encoding buffer, the more the target number; the smaller the encoding buffer, the less the target number.
[0097] For example, the target video includes 500 continuous video frames, and the target number is 90. The electronic device determines a video frame sequence including no more than 90 video frames from the target video.
[0098] The electronic device can determine a plurality of frame structures and the cost of each frame structure based on the video frame sequence. The frame structure can indicate the reference relationship between the video frames in the video frame sequence, and the electronic device can perform video encoding according to the reference relationship between the video frames. The cost of the frame structure can indicate the similarity between the video frames in the frame structure. The smaller the cost, the higher the similarity between the video frames, and the more similar the video content presented by the plurality of video frames in the frame structure. The electronic device takes the frame structure with the minimum cost in the plurality of frame structures as the target frame structure. By constructing a video frame sequence with no more than a target number of video frames, the electronic device can determine the target frame structure of the video frame sequence, and encode the video frame sequence based on the target frame structure.
[0099] In step S202, the electronic device determines an encoding picture group of the video frame sequence from a plurality of picture groups determined based on the target frame structure, the number of video frames in the encoding picture group is not greater than a number threshold, and the cost of the encoding picture group is less than the cost of other picture groups in the plurality of picture groups.
[0100] In the embodiment of the present disclosure, the electronic device divides the video frames in the video frame sequence into a plurality of picture groups according to the order of the video frames based on the reference relationship between the video frames indicated by the target frame structure, can determine the plurality of picture groups, and ensure that the video frames in each picture group are continuous and the number of video frames in each picture group is not greater than the number threshold. The electronic device splices the first picture group with the plurality of subsequent picture groups in turn to obtain a coding picture group with a cost smaller than the cost of other picture groups in the plurality of picture groups. When the number of video frames in the spliced picture group is greater than the number threshold, the electronic device stops splicing the picture groups. The video frames in the coding picture group are a group of video frames with the highest similarity in the target frame structure.
[0101] In step S203, the electronic device encodes the target video based on the coding picture group of each video frame sequence in the plurality of video frame sequences of the target video.
[0102] In the embodiment of the present disclosure, after the electronic device determines the coding picture group of the above-mentioned video frame sequence, the electronic device encodes the coding picture group based on the reference relationship between the video frames in the coding picture group. The electronic device starts from the next video frame of the coding picture group, determines no more than the target number of video frames from the target video again, and takes the determined video frames as a new video frame sequence. The electronic device determines the coding picture group of the new video frame sequence by performing steps S201 and S202. In this way, the electronic device determines the coding picture group of each video frame sequence in the plurality of video frame sequences of the target video, and encodes the coding picture groups of each video frame sequence in sequence.
[0103] The embodiment of the present disclosure provides a video encoding method, which can determine a target frame structure with the minimum cost from a plurality of frame structures of a video frame sequence of a target video according to the actual content of the target video. The target frame structure can indicate the reference relationship between video frames with high similarity in the target video. Based on the target frame structure, a coding picture group with a length not greater than a number threshold is determined from the video frame sequence, and the target video is encoded through the coding picture groups of the plurality of video frame sequences of the target video, which can encode the target video according to the actual content of the target video and improve the quality of the encoding of the target video.
[0104] In some embodiments, the target frame structure is used to indicate positions of bidirectional prediction frames and forward prediction frames in the video frame sequence; a coded group of pictures of the video frame sequence is determined from a plurality of groups of pictures determined based on the target frame structure, including determining a plurality of groups of pictures in the video frame sequence based on positions of bidirectional prediction frames and forward prediction frames in the video frame sequence indicated by the target frame structure, a first video frame and a last video frame in each group of pictures being forward prediction frames, and video frames in the group of pictures other than the first video frame and the last video frame being bidirectional prediction frames, the first video frame in each group of pictures and a last video frame in a previous group of pictures being the same forward prediction frame; a number of video frames in each group of pictures is not greater than a number threshold; a first group of pictures and a second group of pictures in the plurality of groups of pictures are spliced to obtain a first target group of pictures, the first video frame and the last video frame in the first target group of pictures being forward prediction frames, and video frames in the first target group of pictures other than the first video frame and the last video frame being bidirectional prediction frames; in a case where the number of video frames in the first target group of pictures is greater than the number threshold, and a cost of the first group of pictures is less than a cost of the first target group of pictures, the first group of pictures is taken as the coded group of pictures of the video frame sequence.
[0105] In the embodiments of the present disclosure, by the positions of the bidirectional prediction frames and the forward prediction frames, the video frame sequence is divided in the forward prediction frames to obtain a plurality of groups of pictures, and the groups of pictures are spliced to obtain a coded group of pictures, which can divide video frames with high similarity in the video frame sequence into one coded group of pictures, and encode the group of pictures composed of the video frames with high similarity, to encode the target video based on specific content of the target video.
[0106] In some embodiments, the first group of pictures and the second group of pictures in the plurality of groups of pictures are spliced to obtain the first target group of pictures, including replacing the same forward prediction frames in the first group of pictures and the second group of pictures with target bidirectional prediction frames; the first group of pictures after replacement and the second group of pictures after replacement are combined to obtain the first target group of pictures, the first target group of pictures including one target bidirectional prediction frame.
[0107] In the embodiments of the present disclosure, by replacing the same forward prediction frames in the first group of pictures and the second group of pictures with target bidirectional prediction frames, and combining the first group of pictures after replacement and the second group of pictures after replacement, the splicing of the group of pictures can be realized, and the first target group of pictures including a number of video frames not greater than a number threshold can be obtained.
[0108] In some embodiments, the method further includes, in a case where the number of video frames in the first target group of pictures is not greater than the number threshold, splicing the first n groups of pictures to obtain a second target group of pictures; and in a case where the number of video frames in the second target group of pictures is greater than the number threshold, taking a third target group of pictures as the coded group of pictures of the sequence of video frames, the third target group of pictures being obtained by splicing the first n-1 groups of pictures, n being a positive integer greater than 2.
[0109] In the embodiments of the present disclosure, by splicing the first n groups of pictures, the third target group of pictures including video frames with a number not exceeding the number threshold can be obtained.
[0110] In some embodiments, in a case where the number of video frames in the second target group of pictures is greater than the number threshold, taking the third target group of pictures as the coded group of pictures of the sequence of video frames includes: in a case where the number of video frames in the second target group of pictures is greater than the number threshold, obtaining a cost of a fourth target group of pictures, the fourth target group of pictures being obtained by splicing the first n-2 groups of pictures; and in a case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures, taking the third target group of pictures as the coded group of pictures of the sequence of video frames.
[0111] In the embodiments of the present disclosure, by obtaining the cost of the fourth target group of pictures, in a case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures, the third target group of pictures with a smaller cost is taken as the coded group of pictures of the sequence of video frames, which can improve the similarity between video frames in the coded group of pictures.
[0112] In some embodiments, the method further includes, in a case where the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures, taking the fourth target group of pictures as the coded group of pictures of the sequence of video frames.
[0113] In the embodiments of the present disclosure, the fourth target group of pictures with a smaller cost can be taken as the coded group of pictures of the sequence of video frames, which improves the similarity between video frames in the coded group of pictures.
[0114] In some embodiments, determining the target frame structure with the minimum cost from the plurality of frame structures of the video frame sequence of the target video comprises: determining the video frame sequence from the target video; dividing the video frame sequence into a plurality of video frame groups, different video frame groups comprising different numbers of continuous video frames, a first video frame of the video frame group being a first video frame of the video frame sequence; for any video frame group, determining a target cost frame structure of the video frame group from the plurality of frame structures of the video frame group, the target cost frame structure of the video frame group being a frame structure with the minimum cost from the plurality of frame structures of the video frame group; determining the plurality of frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each video frame group in the plurality of video frame groups, the number of the plurality of frame structures of the video frame sequence being not greater than the number threshold; and determining the target frame structure of the video frame sequence from the plurality of frame structures of the video frame sequence, the target frame structure of the video frame sequence being a frame structure with the minimum cost from the plurality of frame structures of the video frame sequence.
[0115] In the embodiments of the present disclosure, the plurality of frame structures of the video frame sequence are determined based on the target cost frame structure of each video frame group in the plurality of video frame groups, and the target frame structure of the video frame sequence is determined from the plurality of frame structures of the video frame sequence, which can determine the reference relationship between the video frames with high similarity in the video frame sequence.
[0116] In some embodiments, determining the target cost frame structure of the video frame group from the plurality of frame structures of the video frame group comprises: in a case where the number of video frames in the video frame group is greater than the number threshold, determining a number threshold of second video frame groups from a plurality of first video frame groups, the number of video frames in the first video frame group being less than the number of video frames in the video frame group, the number of video frames in the second video frame group being greater than the number of video frames in the remaining first video frame groups; determining the number threshold of cost frame structures of the video frame group based on the target cost frame structure of the number threshold of second video frame groups; and determining the target cost frame structure of the video frame group from the number threshold of cost frame structures.
[0117] In the embodiments of the present disclosure, the plurality of cost frame structures of the video frame group are determined based on the target cost frame structure of the number threshold of second video frame groups, and the target cost frame structure of the video frame group is determined from the number threshold of cost frame structures, which can accurately determine the cost frame structure with the minimum cost for each video frame group when the plurality of video frame groups comprise video frames with a number exceeding the number threshold.
[0118] In some embodiments, the method further comprises, in a case where the number of video frames in the video frame group is not greater than the number threshold, determining the plurality of cost frame structures of the video frame group based on the target cost frame structure of the plurality of first video frame groups; and determining the target cost frame structure of the video frame group from the plurality of cost frame structures.
[0119] In the embodiments of the present disclosure, the target cost frame structure of each of the plurality of video frame groups is determined based on the target cost frame structure of the plurality of first video frame groups, and for determining the plurality of video frame groups each of which includes a number of video frames not exceeding the number threshold, the cost frame structure with the minimum cost of each video frame group can be determined more accurately.
[0120] In some embodiments, the video frame sequence includes m consecutive video frames, m being a positive integer greater than 1; the plurality of frame structures are determined based on the number threshold and the target cost frame structure of each of the plurality of video frame groups, including: in a case where m is greater than the number threshold, determining m-1 frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each of the plurality of video frame groups; in a case where m is not greater than the number threshold, determining m-1 frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each of the plurality of video frame groups.
[0121] In the embodiments of the present disclosure, the frame structures with different numbers can be determined based on the number of video frames included in the video frame sequence.
[0122] In some embodiments, the method further includes, for any frame structure, taking a residual value between any video frame in the frame structure and a video frame closest to the video frame in a reference distance as a residual value of the video frame; and taking a sum of residual values of a plurality of video frames in the frame structure as a cost of the frame structure.
[0123] In the embodiments of the present disclosure, by taking the sum of residual values of the plurality of video frames as the cost of the frame structure, the cost can reflect the similarity between the video frames in the frame structure, so that the frame structure of the video frame sequence can be determined more accurately.
[0124] The above Figure 2 The basic flow of the present disclosure is shown, and the video coding scheme provided by the present disclosure is further described below, Figure 3 is a flowchart of another method of video coding according to an exemplary embodiment, which is performed by an electronic device, see Figure 3 The method includes:
[0125] In step S301, the electronic device determines a video frame sequence from a target video, the video frame sequence including no more than a target number of video frames.
[0126] In the embodiments of the present disclosure, the target video is a video to be encoded. The target video includes a plurality of continuous video frames. In a case where the number of video frames included in the target video is greater than a target number, the electronic device determines a first sequence of video frames as a first group of coded pictures, starting from a first video frame in the target video. In a case where the number of video frames included in the target video is not greater than the target number, the electronic device determines all the video frames in the target video as the first sequence of video frames. The target number can be set according to the encoding buffer of the electronic device. The larger the encoding buffer, the greater the target number. The smaller the encoding buffer, the smaller the target number.
[0127] It should be noted that after the electronic device determines the encoding pictures of the first sequence of video frames, the electronic device starts to determine a second sequence of video frames. In a case where the number of video frames in the target video other than the first group of coded pictures is greater than the target number, the electronic device determines a second sequence of video frames as the target number of continuous video frames, starting from the next video frame of the first group of coded pictures. In a case where the number of video frames in the target video other than the first group of coded pictures is not greater than the target number, the electronic device determines the video frames in the target video other than the first group of coded pictures as the second sequence of video frames. This is repeated until the corresponding group of coded pictures is determined for all the sequences of video frames in the target video.
[0128] In step S302, the electronic device divides the sequence of video frames into a plurality of groups of video frames, different groups of video frames including different numbers of continuous video frames, and the first video frame of a group of video frames being the first video frame of the sequence of video frames.
[0129] In the embodiments of the present disclosure, the electronic device divides the sequence of video frames into a plurality of groups of video frames according to the number of video frames included. The electronic device takes the first video frame in the sequence of video frames as the first group of video frames, takes the first two video frames in the sequence of video frames as the second group of video frames, and so on, and the last group of video frames includes all the video frames in the sequence of video frames.
[0130] In step S303, for any group of video frames, the electronic device takes the frame structure with the minimum cost in the plurality of frame structures of the group of video frames as the target cost frame structure of the group of video frames, the cost being used to represent the similarity between video frames.
[0131] In the embodiments of the present disclosure, the electronic device determines, based on the sequence of video frames, a plurality of frame structures of each video frame group and a cost of each frame structure. The frame structure of the video frame group can indicate a reference relationship between the video frames in the video frame group, and the costs of different frame structures are different. The cost of the frame structure of the video frame group can indicate the similarity between the video frames in the video frame group. The electronic device takes the frame structure with the minimum cost in the plurality of frame structures of the video frame group as the target cost frame structure of the video frame group. The ways of determining the target cost frame structure of different video frame groups are not completely the same, and the electronic device can determine the target cost frame structures of a plurality of video frame groups in the following two ways.
[0132] In the first way, for any video frame group, if the number of video frames in the video frame group is greater than a number threshold, the electronic device determines, from a plurality of first video frame groups, a number threshold of second video frame groups. The number of video frames in the first video frame group is less than the number of video frames in the video frame group. The number of video frames in the second video frame group is greater than the number of video frames in the remaining first video frame group. The electronic device determines a number threshold of cost frame structures of the video frame group based on the target cost frame structures of the number threshold of second video frame groups. The electronic device determines the cost of each cost frame structure and takes the cost frame structure with the minimum cost in the number threshold of cost frame structures as the target cost frame structure of the video frame group. By determining the plurality of cost frame structures of the video frame group through the target cost frame structures of the number threshold of second video frame groups, the cost frame structure with the minimum cost of each video frame group can be accurately determined for a plurality of video frame groups including video frames whose number exceeds the number threshold.
[0133] In the second way, if the number of video frames in the video frame group is not greater than the number threshold, the electronic device determines a plurality of cost frame structures of the video frame group based on the target cost frame structures of a plurality of first video frame groups. The electronic device takes the cost frame structure with the minimum cost from the plurality of cost frame structures as the target cost frame structure. By determining the plurality of cost frame structures of the video frame group through the target cost frame structures of the plurality of first video frame groups, the cost frame structure with the minimum cost of each video frame group can be accurately determined for a plurality of video frame groups including video frames whose number does not exceed the number threshold.
[0134] To make the process of determining the target cost frame structure of the plurality of video frame groups by the electronic device clearer, the target cost frame structure can be represented by bestpath. bestpath[i] can represent the target cost frame structure of the ith video frame group. In the case that i is greater than the number threshold, the electronic device can determine Nbframes bestpath[i] based on bestpath[i-Nbframes] to bestpath[i-1], and take the bestpath[i] with the minimum cost as the final bestpath[i], where Nbframes is the number threshold. In the case that i is not greater than the number threshold, the electronic device can determine i-1 bestpath[i] based on bestpath[i-1] to bestpath[1], and take the bestpath[i] with the minimum cost as the final bestpath[i].
[0135] For example, the video frame sequence includes 90 video frames, and the number threshold is 32. The electronic device divides the video frame sequence into 90 video frame groups. i is a parameter traversed in the process of determining the target cost frame structure of the plurality of video frame groups by the electronic device, and the length of the video frame sequence is 90 frames, so i is at most 90.
[0136] When i=1, the electronic device determines the target cost frame structure bestpath[1] of the first video frame group as I, that is, determines the video frame in the first video frame group as an I frame.
[0137] When i=2, the electronic device determines 1 bestpath[2] based on bestpath[1], and bestpath[2] is bestpath[1]+P, that is, (I, P), that is, determines the second video frame in the second video frame group as a P frame.
[0138] When i=3, the electronic device determines 2 bestpath[3] based on bestpath[1] and bestpath[2]. One bestpath[3] is bestpath[1]+BP, that is, (I, B, P), that is, determines the second video frame in the third video frame group as a B frame. The other bestpath[3] is bestpath[2]+P, that is, (I, P, P). The electronic device determines the bestpath[3] with the minimum cost from the 2 bestpath[3] as the final bestpath[3].
[0139] When i = 4, the electronic device determines 3 bestpath[4] based on bestpath[1], bestpath[2] and bestpath[3]. One of the bestpath[4] is bestpath[1]+BBP, i.e., (I, B, B, P). Another of the bestpath[4] is bestpath[2]+BP, i.e., (I, P, B, P). Another of the bestpath[4] is bestpath[3]+P. The electronic device determines the bestpath[4] with the minimum cost from the 3 bestpath[4] as the final bestpath[4].
[0140] When i = 32, the electronic device determines 31 bestpath
[32] based on bestpath[1] to bestpath
[31] . One of the bestpath
[32] is bestpath[1]+BB…BP, the number of B is 30. One of the bestpath
[32] is bestpath[2]+BB…BP, the number of B is 29. One of the bestpath
[32] is bestpath
[31] +P. The electronic device determines the bestpath
[32] with the minimum cost from the 31 bestpath
[32] as the final bestpath
[32] .
[0141] When i = 50, the electronic device determines 32 bestpath
[50] based on bestpath
[18] to bestpath
[49] . One of the bestpath
[50] is bestpath
[18] +BB…BP, the number of B is 31, and the number of B is the largest. One of the bestpath
[50] is bestpath
[19] +BB…BP, the number of B is 30. One of the bestpath
[50] is bestpath
[49] +P. The electronic device determines the bestpath
[50] with the minimum cost from the 32 bestpath
[50] as the final bestpath
[50] .
[0142] In step S304, the electronic device determines a plurality of frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each of the plurality of video frame groups, and the number of the plurality of frame structures of the video frame sequence is not greater than the number threshold.
[0143] In the embodiments of the present disclosure, since the last video frame group of the video frame sequence includes all the video frames in the video frame sequence, the electronic device can determine a plurality of frame structures of the last video frame group based on the number threshold and the target cost frame structure of the plurality of video frame groups before the last video frame group.
[0144] In some embodiments, the electronic device can determine the number of frame structures according to the number of video frames included in the sequence of video frames. The sequence of video frames includes m consecutive video frames, m being a positive integer greater than 1; in the case that m is greater than the number threshold, the electronic device determines m-1 frame structures of the sequence of video frames based on the number threshold and the target cost frame structure of each of the plurality of video frame groups. In the case that m is not greater than the number threshold, the electronic device determines 32 frame structures of the sequence of video frames based on the number threshold and the target cost frame structure of each of the plurality of video frame groups. In this way, different numbers of frame structures can be determined based on the number of video frames included in the sequence of video frames.
[0145] For example, in the case that the number threshold is 32 and m is 90, the first video frame group to the 89th video frame group are all first video frame groups of the sequence of video frames, and the electronic device determines the 58th video frame group to the 89th video frame group as 32 second video frame groups from the first video frame group to the 89th video frame group. The electronic device determines 32 frame structures of the sequence of video frames based on the target cost frame structure of the 32 second video frame groups. In the case that the number threshold is 32 and m is 3, the first two video frame groups are both first video frame groups of the sequence of video frames, and the electronic device determines 2 frame structures of the sequence of video frames based on the target cost frame structure of the first two video frame groups.
[0146] In step S305, the electronic device determines the frame structure with the minimum cost among the plurality of frame structures of the sequence of video frames as the target frame structure of the sequence of video frames.
[0147] In the embodiments of the present disclosure, the electronic device determines the cost of each frame structure and determines the frame structure with the minimum cost as the target frame structure of the sequence of video frames.
[0148] In some embodiments, the electronic device can determine the cost of a frame structure based on the residual values between video frames. For any frame structure, the electronic device determines the residual value between any video frame in the frame structure and the video frame closest to the reference distance of the video frame as the residual value of the video frame. The electronic device determines the sum of the residual values of the plurality of video frames in the frame structure as the cost of the frame structure. The higher the similarity between video frames, the smaller the residual value between video frames, and the smaller the cost of the frame structure. The lower the similarity between video frames, the greater the residual value between video frames, and the greater the cost of the frame structure. By determining the sum of the residual values of the plurality of video frames as the cost of the frame structure, the cost can reflect the degree of similarity between the video frames in the frame structure, thereby more accurately determining the frame structure of the sequence of video frames.
[0149] In some embodiments, the electronic device can determine the cost by calculating a SATD (Sum of Absolute Transformed Difference) loss. In the process of calculating the SATD loss, D(p0, b, p1) is used to represent the inter-frame SATD loss of the video frame with the sequence number b referring to the frames p0 and p1, and D(p0, p1, p1) is used to represent the inter-frame SATD loss of the frame p1 referring to the frame p0. The sum of the SATD loss between each video frame in the frame structure and the video frame closest to the reference of the video frame is the cost of the frame structure.
[0150] In some embodiments, the electronic device can also determine the cost by calculating an SSE (the Sum of Squares due to Error), an MSE (Mean Squared Error), an SAD (Sum of Absolute Error), or the like. The method of calculating the cost is not limited in the embodiments of the present disclosure.
[0151] In step S306, the electronic device determines a plurality of groups of pictures in the sequence of video frames based on the positions of the bidirectional prediction frames and the forward prediction frames in the sequence of video frames indicated by the target frame structure, the first video frame and the last video frame in the group of pictures are forward prediction frames, the video frames in the group of pictures other than the first video frame and the last video frame are bidirectional prediction frames, the first video frame in the group of pictures and the last video frame in the previous group of pictures are the same forward prediction frame, and the number of video frames in the group of pictures is not greater than a threshold.
[0152] In the embodiments of the present disclosure, the target frame structure can indicate the reference relationship between the video frames in the sequence of video frames. The reference relationship of the video frames can be determined by the type of the video frames. The type of the video frames includes B frames and P frames. The B frame records the difference with the previous video frame and the difference with the next video frame, and accordingly, the electronic device encodes the difference between the video frame and the previous and next video frames when encoding the B frame. The P frame records the difference with the previous video frame, and accordingly, the electronic device encodes the difference between the video frame and the previous video frame when encoding the P frame.
[0153] The target frame structure indicates the positions of bidirectional and forward-predicted frames in a video frame sequence. Based on the positions of these frames, the electronic device segments the video frame sequence within the forward-predicted frames, resulting in multiple Groups of Pictures (GOPs). Since the target frame structure is determined based on a number threshold and the target cost frame structure of multiple video frame groups, the interval between two forward-predicted frames in the target frame sequence will not exceed the number threshold of video frames. Therefore, the number of video frames in a GOP obtained from the forward-predicted frame segmentation is no greater than the number threshold.
[0154] For example, for a video frame sequence with a target frame structure of (P0, B, B, B, P1, B, B, B, P2, B, B, B, B, B, P3), the electronic device segments the video frame sequence according to the position of the P frames, obtaining three GOPs: (P0, B, B, B, P1), (P1, B, B, B, P2), and (P2, B, B, B, B, B, P3). Here, 1, 2, and 3 represent the sequence numbers of the P frames.
[0155] It should be noted that the multiple GOPs identified above are bipartite symmetric frame structures. Figure 4 This is a schematic diagram of a frame group with a bipartite symmetrical frame structure. (Example) Figure 4 As shown, this GOP is 16 frames long, with video frame 8 being the lowest-level B-frame, determined by the bipartite symmetric frame structure. Video frames 1-7 and 9-15 can also be divided according to the bipartite symmetric frame structure.
[0156] In step S307, the electronic device splices the first and second video groups from multiple video groups to obtain a first target video group. The first and last video frames in the first target video group are forward prediction frames, and the video frames in the first target video group other than the first and last video frames are bidirectional prediction frames.
[0157] In this embodiment of the disclosure, after the electronic device determines multiple screen groups, it splices the video frames in the first screen group and the second screen group to obtain the first target screen group.
[0158] In some embodiments, the electronic device can combine the first group of pictures and the second group of pictures. The electronic device replaces the same forward predicted frames in the first group of pictures and the second group of pictures with the target bi-predicted frames; and combines the replaced first group of pictures and the replaced second group of pictures to obtain a first target group of pictures, the first target group of pictures including one target bi-predicted frame. By replacing the same forward predicted frames in the first group of pictures and the second group of pictures with the target bi-predicted frames and combining the replaced first group of pictures and the second group of pictures, the first target group of pictures including video frames with a number not exceeding the number threshold and the video frames in the first target group of pictures have a high similarity can be obtained.
[0159] For example, the video frames in the first group of pictures GOP1 are (P0, B, B, …, Pn), and the video frames in the second group of pictures GOP2 are (Pn, B, B, …, Pm). The electronic device splices the GOP1 and the GOP2, replaces the same forward predicted frame Pn with the target bi-predicted frame Bn, and combines the replaced GOP1 and the GOP2 to obtain the video frames in the first target group of pictures as (P0, B, B, …, Bn, B, B, …, Pm). Wherein, n and m represent the serial numbers of the video frames.
[0160] The electronic device performs the following step S308 or steps S309-S310 based on the number of the video frames in the first target group of pictures and the number threshold.
[0161] In step S308, in a case where the number of the video frames in the first target group of pictures is greater than the number threshold and the cost of the first group of pictures is less than the cost of the first target group of pictures, the electronic device takes the first group of pictures as the coding group of pictures of the video frame sequence, the number of the video frames in the coding group of pictures is not greater than the number threshold, and the cost of the coding group of pictures is less than the cost of other groups of pictures in the plurality of groups of pictures.
[0162] In the embodiments of the present disclosure, in a case where the number of the video frames in the first target group of pictures is greater than the number threshold and the cost of the first group of pictures is less than the cost of the first target group of pictures, the electronic device does not splice the first group of pictures and the second group of pictures, but directly takes the first group of pictures as the coding group of pictures of the video frame sequence. Wherein, the coding group of pictures is a group of pictures to be pushed out for encoding.
[0163] In step S309, in a case where the number of the video frames in the first target group of pictures is not greater than the number threshold, the electronic device splices the first n groups of pictures to obtain a second target group of pictures, and n is a positive integer greater than 2.
[0164] In the embodiments of the present disclosure, when the number of video frames in the first target picture group is not greater than the number threshold, the electronic device continues to splice the first n picture groups in the front-back order of the picture groups with the first target picture group to obtain a second target picture group.
[0165] For example, the video frames in the first picture group GOP1 are (P0, B, B, …, Pn), the video frames in the second picture group GOP2 are (Pn, B, B, …, Pm), the video frames in the third picture group GOP3 are (Pm, B, B, …, Pk), and the video frames in the fourth picture group GOP4 are (Pk, B, B, …, Pj). Wherein, n, m, k and j represent the sequence numbers of the video frames.
[0166] When n = 3, the electronic device continues to splice the third picture group with the first target picture group, replaces the same forward prediction frame Pm with a target bidirectional prediction frame Bm, and merges the replaced first target picture group and GOP3 to obtain a second target picture group, and the video frames in the second target picture group are (P0, B, B, …, Bn, B, B, …, Bm, B, …, B, Pk).
[0167] When n = 4, the electronic device continues to splice the third and fourth picture groups in the front-back order of the picture groups with the first target picture group, replaces the same forward prediction frames Pm and Pk with target bidirectional prediction frames Bm and Bk, and merges the replaced GOP3, GOP4 and the first target picture group to obtain a second target picture group, and the video frames in the second target picture group are (P0, B, B, …, Bn, B, B, …, Bm, B, …, B, Bk, B, …, Pj).
[0168] In step S310, when the number of video frames in the second target picture group is greater than the number threshold, the electronic device takes a third target picture group as the encoding picture group of the video frame sequence, and the third target picture group is spliced from the first n-1 picture groups.
[0169] In the embodiments of the present disclosure, the electronic device splices the first n picture groups in the front-back order of the picture groups with the first target picture group until the number of video frames in the obtained second target picture group is greater than the number threshold, and the electronic device takes a third target picture group spliced from the first n-1 picture groups as the encoding picture group of the video frame sequence.
[0170] It should be noted that the encoding picture group spliced from the multiple picture groups is a non-symmetrical frame structure. Figure 5 FIG. 1 is a schematic diagram of a picture group of a non-symmetrical frame structure. Figure 5As shown, the GOP length is 16 frames, in which the No. 12 video frame is the replaced target bidirectional prediction frame. The No. 1-No. 11 video frames and the No. 13-No. 15 video frames can also be divided according to the asymmetric frame structure.
[0171] In some embodiments, the electronic device is capable of determining the coded group of pictures according to the cost of the group of pictures. In a case where the number of video frames in the second target group of pictures is greater than the number threshold, the electronic device obtains the cost of a fourth target group of pictures, which is spliced from the first n-2 groups of pictures. In a case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures, the electronic device takes the third target group of pictures as the coded group of pictures of the sequence of video frames. In a case where the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures, the electronic device takes the fourth target group of pictures as the coded group of pictures of the sequence of video frames. By judging the cost of the spliced group of pictures and the cost of the group of pictures before splicing, and taking the group of pictures with smaller cost as the coded group of pictures, the similarity between the video frames in the determined coded group of pictures can be improved, and the quality of video coding can also be improved.
[0172] For example, in a case where the number of video frames in the second target group of pictures determined by the electronic device is greater than the number threshold when n = 4, the electronic device determines the cost of a fourth target group of pictures spliced from the first n-2 groups of pictures (i.e., the first 2 groups of pictures) and the cost of a third target group of pictures spliced from the first n-1 groups of pictures (i.e., the first 3 groups of pictures), and determines the coded group of pictures based on the above-mentioned costs. In a case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures, the electronic device takes the third target group of pictures as the coded group of pictures of the sequence of video frames. In a case where the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures, the electronic device takes the fourth target group of pictures as the coded group of pictures of the sequence of video frames.
[0173] In step S311, the electronic device encodes the target video based on the coded group of pictures of each of the plurality of sequences of video frames of the target video.
[0174] In the embodiments of the present disclosure, after determining the coded picture group of the first video frame sequence, the electronic device encodes the coded picture group of the first video frame sequence according to the reference relationship between the video frames. The electronic device determines the second video frame sequence. In the case that the number of video frames in the target video except the coded picture group of the first video frame sequence is greater than the target number, the electronic device determines the target number of continuous video frames from the next frame of the coded picture group as the second video frame sequence. In the case that the number of video frames in the target video except the coded picture group of the first video frame sequence is not greater than the target number, the electronic device takes the video frames in the target video except the coded picture group of the first video frame sequence as the second video frame sequence. The electronic device determines the coded picture group of the second video frame sequence and encodes the coded picture group according to the reference relationship between the video frames. In this way, until the corresponding coded picture group of all the video frame sequences in the target video is determined, the electronic device encodes the corresponding coded picture group of the video frame sequence in sequence based on the order of the video frame sequence.
[0175] The embodiments of the present disclosure provide a video encoding method, which can determine a target frame structure with minimum cost from multiple frame structures of a video frame sequence of a target video according to the actual content of the target video. The target frame structure can indicate the reference relationship between video frames with high similarity in the target video. Based on the target frame structure, a coded picture group with a length not exceeding a number threshold is determined in the video frame sequence, and the target video is encoded by using the coded picture groups of multiple video frame sequences of the target video, so that the target video can be encoded according to the actual content of the target video, and the quality of the encoding of the target video is improved.
[0176] All the optional technical solutions described above can be combined to form optional embodiments of the present disclosure, which will not be described one by one here.
[0177] Figure 6 is a block diagram of a video encoding device according to an exemplary embodiment. As shown in Figure 6 the device includes a first determining unit 601, a second determining unit 602, and an encoding unit 603.
[0178] The first determining unit 601 is configured to determine, for any video frame sequence of multiple video frame sequences of a target video, a target frame structure with minimum cost from multiple frame structures of the video frame sequence, the video frame sequence including no more than a target number of video frames, and the cost being used to represent the similarity between the video frames.
[0179] The second determining unit 602 is configured to determine an encoded picture group of a video frame sequence from multiple picture groups determined based on the target frame structure, wherein the number of video frames in the encoded picture group is not greater than a number threshold, and the cost of the encoded picture group is less than the cost of other picture groups in the multiple picture groups.
[0180] The encoding unit 603 is configured to encode the target video based on the encoded frame group of each video frame sequence in a plurality of video frame sequences of the target video.
[0181] In some embodiments, the target frame structure is used to indicate the positions of bidirectional prediction frames and forward prediction frames in a video frame sequence; Figure 7 This is a block diagram of another video encoding apparatus provided in an embodiment of this disclosure. See also... Figure 7 As shown, the second determining unit 602 includes:
[0182] The first determining subunit 6021 is configured to determine multiple picture groups in the video frame sequence based on the positions of bidirectional prediction frames and forward prediction frames indicated by the target frame structure. The first and last video frames in the picture group are forward prediction frames, and the video frames in the picture group other than the first and last video frames are bidirectional prediction frames. The first video frame in the picture group is the same forward prediction frame as the last video frame in the previous picture group. The number of video frames in the picture group is not greater than a number threshold.
[0183] The splicing subunit 6022 is configured to splice the first and second video groups of multiple video groups to obtain a first target video group. The first and last video frames in the first target video group are forward prediction frames, and the video frames in the first target video group other than the first and last video frames are bidirectional prediction frames.
[0184] The second determining subunit 6023 is configured to, when the number of video frames in the first target frame group is greater than a number threshold and the cost of the first frame group is less than the cost of the first target frame group, use the first frame group as the encoded frame group of the video frame sequence.
[0185] In some embodiments, the splicing subunit 6022 is configured to replace the same predicted frames in the first frame group and the second frame group with target bidirectional predicted frames; and to merge the replaced first frame group and the replaced second frame group to obtain a first target frame group, wherein the first target frame group includes a target bidirectional predicted frame.
[0186] In some embodiments, the splicing subunit 6022 is configured to splice the first n groups of pictures to obtain a second target group of pictures in a case where a number of video frames in the first target group of pictures is not greater than the number threshold; and the second determining subunit 6023 is configured to take a third target group of pictures as the coding group of pictures of the sequence of video frames in a case where the number of video frames in the second target group of pictures is greater than the number threshold, the third target group of pictures being spliced from the first n-1 groups of pictures, and n being a positive integer greater than 2.
[0187] In some embodiments, the second determining subunit 6023 is configured to obtain a cost of a fourth target group of pictures spliced from the first n-2 groups of pictures in a case where the number of video frames in the second target group of pictures is greater than the number threshold; and take the third target group of pictures as the coding group of pictures of the sequence of video frames in a case where the cost of the third target group of pictures is less than the cost of the fourth target group of pictures.
[0188] In some embodiments, the second determining subunit 6023 is configured to take the fourth target group of pictures as the coding group of pictures of the sequence of video frames in a case where the cost of the third target group of pictures is not less than the cost of the fourth target group of pictures.
[0189] In some embodiments, referring to FIG. 6, Figure 7 The first determining unit 601 includes:
[0190] The third determining subunit 6011 is configured to determine the sequence of video frames from the target video;
[0191] The dividing subunit 6012 is configured to divide the sequence of video frames into a plurality of groups of video frames, different groups of video frames including different numbers of continuous video frames, and a first video frame of a group of video frames being a first video frame of the sequence of video frames;
[0192] The fourth determining subunit 6013 is configured to, for any group of video frames, take a frame structure with a minimum cost in a plurality of frame structures of the group of video frames as a target cost frame structure of the group of video frames;
[0193] The fifth determining subunit 6014 is configured to determine a plurality of frame structures of the sequence of video frames based on the number threshold and the target cost frame structure of each group of video frames in the plurality of groups of video frames, a number of the plurality of frame structures of the sequence of video frames being not greater than the number threshold;
[0194] The sixth determining subunit 6015 is configured to take a frame structure with a minimum cost in the plurality of frame structures of the sequence of video frames as a target frame structure of the sequence of video frames.
[0195] In some embodiments, the fourth determining subunit 6013 is configured to, in a case where the number of video frames in the video frame group is greater than the number threshold, determine, from the plurality of first video frame groups, a number threshold of second video frame groups, the number of video frames in the first video frame group being less than the number of video frames in the video frame group, the number of video frames in the second video frame group being greater than the number of video frames in the remaining first video frame group; determine a number threshold of cost frame structures of the video frame group based on the target cost frame structure of the number threshold of second video frame groups; and determine the target cost frame structure of the video frame group from the number threshold of cost frame structures.
[0196] In some embodiments, the fourth determining subunit 6013 is configured to, in a case where the number of video frames in the video frame group is not greater than the number threshold, determine a plurality of cost frame structures of the video frame group based on the target cost frame structure of the plurality of first video frame groups; and the fourth determining subunit 6013 is configured to determine the target cost frame structure of the video frame group from the plurality of cost frame structures.
[0197] In some embodiments, the video frame sequence includes m consecutive video frames, m being a positive integer greater than 1; the fifth determining subunit 6014 is configured to, in a case where m is greater than the number threshold, determine a number threshold of frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each of the plurality of video frame groups; and in a case where m is not greater than the number threshold, determine m-1 frame structures of the video frame sequence based on the number threshold and the target cost frame structure of each of the plurality of video frame groups.
[0198] In some embodiments, referring to FIG. 6, Figure 7 As shown in FIG. 6, the apparatus further includes:
[0199] The third determining unit 604 is configured to, for any frame structure, take a residual value between any video frame in the frame structure and a video frame closest to the video frame in reference distance as a residual value of the video frame.
[0200] The fourth determining unit 605 is configured to take a sum of the residual values of the plurality of video frames in the frame structure as a cost of the frame structure.
[0201] Embodiments of the present disclosure provide a video encoding apparatus, which can determine a target frame structure with minimum cost from a plurality of frame structures of a video frame sequence of a target video according to actual content of the target video. The target frame structure can indicate a reference relationship between video frames with high similarity in the target video. Based on the target frame structure, a number of coded picture groups with a length not exceeding a number threshold in the video frame sequence are determined, and the target video is encoded through the coded picture groups of a plurality of video frame sequences of the target video, which can encode the target video according to actual content of the target video, thereby improving the quality of encoding of the target video.
[0202] It should be noted that the video encoding apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional units when encoding a target video. In actual applications, the above functions can be completed by different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the above described functions. In addition, the video encoding apparatus and the video encoding method provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be described here.
[0203] As to the video encoding apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.
[0204] Figure 8 is a block diagram of an electronic device according to an example embodiment. Generally, the electronic device 800 includes a processor 801 and a memory 802.
[0205] The processor 801 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 801 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 801 can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content needed to be displayed on the display screen. In some embodiments, the processor 801 can also include an AI (Artificial Intelligence) processor for processing machine learning related computing operations.
[0206] The memory 802 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 802 can also include high-speed random access memory and can include non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. In some embodiments, the non-transitory computer-readable storage medium of the memory 802 is used for storing the at least one program code for being executed by the processor 801 to implement the video coding method provided by the method embodiments of the present disclosure.
[0207] In some embodiments, the electronic device 800 can further optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 803 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.
[0208] The peripheral device interface 803 can be used to connect at least one peripheral device related to input / output (I / O) to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on a separate chip or circuit board, and the present embodiment is not limited in this regard.
[0209] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 804 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the radio frequency circuit 804 can also include NFC (Near Field Communication) related circuitry, and the present disclosure is not limited in this regard.
[0210] The display screen 805 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 is further configured to capture touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. In this case, the display screen 805 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 805 can be one, configured on the front panel of the electronic device 800; in other embodiments, the display screen 805 can be at least two, respectively configured on different surfaces of the electronic device 800 or in a folding design; in yet other embodiments, the display screen 805 can be a flexible display screen, configured on a curved surface or a folding surface of the electronic device 800. Even, the display screen 805 can also be configured in an irregular shape other than a rectangle, i.e., a notched screen. The display screen 805 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.
[0211] The camera assembly 806 is configured to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is configured on the front panel of the electronic device, and the rear camera is configured on the back of the electronic device. In some embodiments, the rear camera is at least two, respectively configured as any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 806 can further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0212] The audio circuit 807 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 801 for processing, or input to the radio frequency circuit 804 to realize voice communication. The microphone can be multiple for the purpose of stereo sound collection or noise reduction, and arranged at different parts of the electronic device 800 respectively. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert an electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a conventional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert an electrical signal into a sound wave audible to humans, but also convert an electrical signal into an inaudible sound wave to humans for ranging purposes, etc. In some embodiments, the audio circuit 807 can also include a headphone jack.
[0213] The power supply 808 is configured to supply power to each component of the electronic device 800. The power supply 808 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0214] Those skilled in the art can understand that the structure shown in the above description is not a limitation on the electronic device 800, and the electronic device 800 can include more or fewer components than those shown in the figure, or combine certain components, or use different component arrangements. Figure 8 Those skilled in the art can understand that the structure shown in the above description is not a limitation on the electronic device 800, and the electronic device 800 can include more or fewer components than those shown in the figure, or combine certain components, or use different component arrangements.
[0215] In an exemplary embodiment, a computer readable storage medium including instructions, such as the memory 802 including instructions, is also provided, which can be executed by the processor 801 of the terminal 800 to complete the above-mentioned video encoding method. Alternatively, the computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0216] A computer program product includes a computer program that is executed by a processor to implement the above-mentioned video encoding method.
[0217] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The disclosure is intended to cover any variations, uses, or adaptations of the disclosure following the general principles thereof and including such departures from the present disclosure that come within known use or custom in the art to which the disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the following claims.
[0218] It should be understood that the present disclosure is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method of video coding, the method comprising: include: For any video frame sequence in a plurality of video frame sequences of a target video, the video frame sequence is divided into a plurality of video frame groups. The video frame sequence includes no more than the target number of video frames. Different video frame groups include different numbers of consecutive video frames. The first video frame of each video frame group is the first video frame of the video frame sequence. For any group of video frames, the frame structure with the lowest cost among the multiple frame structures of the video frame group is taken as the target cost frame structure of the video frame group, where the cost is used to represent the similarity between video frames. Based on the quantity threshold and the target cost frame structure of each video frame group in the plurality of video frame groups, a plurality of frame structures of the video frame sequence are determined, wherein the number of the plurality of frame structures of the video frame sequence is not greater than the quantity threshold. The frame structure with the lowest cost among multiple frame structures in the video frame sequence is taken as the target frame structure of the video frame sequence. From multiple frame groups determined based on the target frame structure, an encoded frame group of the video frame sequence is determined, wherein the number of video frames in the encoded frame group is not greater than a number threshold, and the cost of the encoded frame group is less than the cost of other frame groups in the multiple frame groups. The target video is encoded based on the encoded frame group of each video frame sequence in multiple video frame sequences of the target video.
2. The video coding method of claim 1, wherein, The target frame structure is used to indicate the positions of bidirectional prediction frames and forward prediction frames in the video frame sequence; Determining the encoded frame group of the video frame sequence from multiple frame groups determined based on the target frame structure includes: Based on the positions of bidirectional prediction frames and forward prediction frames in the video frame sequence indicated by the target frame structure, multiple frame groups are determined in the video frame sequence. The first and last video frames in each frame group are forward prediction frames, and all video frames in each frame group except the first and last video frames are bidirectional prediction frames. The first video frame in each frame group is the same forward prediction frame as the last video frame in the previous frame group. The number of video frames in each frame group is not greater than the number threshold. The first and second video groups in the plurality of video groups are stitched together to obtain a first target video group. The first and last video frames in the first target video group are forward prediction frames, and the video frames in the first target video group other than the first and last video frames are bidirectional prediction frames. If the number of video frames in the first target frame group is greater than the number threshold, and the cost of the first frame group is less than the cost of the first target frame group, then the first frame group is used as the encoded frame group of the video frame sequence.
3. The video coding method of claim 2, wherein, The step of stitching together the first and second frame groups from the plurality of frame groups to obtain the first target frame group includes: Replace the same forward prediction frames in the first and second frame groups with target bidirectional prediction frames; The first target picture group is obtained by merging the first replaced picture group and the second replaced picture group, and the first target picture group includes one target bidirectional prediction frame.
4. The video coding method of claim 2, wherein, The method further comprises: In a case where the number of video frames in the first target picture group is not greater than the number threshold, splicing the first n picture groups to obtain a second target picture group; In a case where the number of video frames in the second target picture group is greater than the number threshold, taking a third target picture group as the coding picture group of the video frame sequence, the third target picture group being spliced from the first n-1 picture groups, n being a positive integer greater than 2.
5. The video coding method of claim 4, wherein, The case where the number of video frames in the second target picture group is greater than the number threshold comprises: In a case where the number of video frames in the second target picture group is greater than the number threshold, obtaining a cost of a fourth target picture group, the fourth target picture group being spliced from the first n-2 picture groups; In a case where the cost of the third target picture group is less than the cost of the fourth target picture group, taking the third target picture group as the coding picture group of the video frame sequence.
6. The video coding method of claim 5, wherein, The method further comprises: In a case where the cost of the third target picture group is not less than the cost of the fourth target picture group, taking the fourth target picture group as the coding picture group of the video frame sequence.
7. The video coding method of claim 1, wherein, The case where, for any video frame group, the frame structure with the minimum cost in the multiple frame structures of the video frame group is taken as the target cost frame structure of the video frame group comprises: In a case where the number of video frames in the video frame group is greater than the number threshold, determining, from the multiple first video frame groups, a number threshold of second video frame groups, the number of video frames in the first video frame groups being less than the number of video frames in the video frame group, the number of video frames in the second video frame groups being greater than the number of video frames in the remaining first video frame groups; Determining, based on the target cost frame structures of the number threshold of second video frame groups, a number threshold of cost frame structures of the video frame group; Determining, from the number threshold of cost frame structures, the target cost frame structure of the video frame group.
8. The video coding method of claim 7, wherein, The method further comprises: In a case where the number of video frames in the video frame group is not greater than the number threshold, determining, based on the target cost frame structures of the multiple first video frame groups, multiple cost frame structures of the video frame group; Determining, from the multiple cost frame structures, the target cost frame structure of the video frame group.
9. The video coding method of claim 1, wherein, The video frame sequence comprises m continuous video frames, m being a positive integer greater than 1; The case where, based on the number threshold and the target cost frame structure of each video frame group in the multiple video frame groups, a plurality of frame structures of the video frame sequence are determined comprises: In a case where m is greater than the number threshold, determining, based on the number threshold and the target cost frame structure of each video frame group in the multiple video frame groups, a number threshold of frame structures of the video frame sequence; In a case that m is not greater than the number threshold, m-1 frame structures of the video frame sequence are determined based on the number threshold and the target cost frame structure of each of the plurality of video frame groups.
10. The video coding method of claim 1, wherein, The method further comprises: For any frame structure, a residual value between any video frame in the frame structure and a video frame closest to the reference distance of the video frame is taken as a residual value of the video frame; A sum of residual values of a plurality of video frames in the frame structure is taken as a cost of the frame structure.
11. A video encoding apparatus, comprising: The apparatus comprises: A first determining unit configured to, for any video frame sequence in a plurality of video frame sequences of a target video, divide the video frame sequence into a plurality of video frame groups, the video frame sequence comprising no more than a target number of video frames, different video frame groups comprising different numbers of consecutive video frames, a first video frame of each of the video frame groups being a first video frame of the video frame sequence; The first determining unit is further configured to, for any video frame group, take a frame structure with a minimum cost in a plurality of frame structures of the video frame group as a target cost frame structure of the video frame group, the cost being used to represent a similarity between video frames; The first determining unit is further configured to determine a plurality of frame structures of the video frame sequence based on a number threshold and the target cost frame structure of each of the plurality of video frame groups, a number of the plurality of frame structures of the video frame sequence being not greater than the number threshold; The first determining unit is further configured to take a frame structure with a minimum cost in the plurality of frame structures of the video frame sequence as a target frame structure of the video frame sequence; A second determining unit configured to determine, from a plurality of picture groups determined based on the target frame structure, a coded picture group of the video frame sequence, a number of video frames in the coded picture group being not greater than a number threshold, a cost of the coded picture group being less than a cost of other picture groups in the plurality of picture groups; An encoding unit configured to encode the target video based on the coded picture group of each of the plurality of video frame sequences of the target video.
12. An electronic device, comprising: The electronic device comprises: one or more processors; a memory for storing program codes executable by the processors; wherein the processors are configured to execute the program codes to implement the video encoding method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processors of the electronic device, the electronic device is enabled to perform the video encoding method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video information processing method and device, multimedia information processing method and device and electronic equipment
CN112788341A
Video coding method and device, electronic equipment and storage medium
CN115514960A