Video generation method and device, electronic equipment, medium and product
By acquiring target image frames during video generation and setting the skip encoding type, the problems of high computational resource consumption and large file size of static video images are solved, achieving more efficient encoding and lower bandwidth usage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies consume a lot of computational resources and produce large video files when encoding static video images, resulting in high bandwidth usage.
By responding to the video generation mode switching operation during the video generation process, the target image frame is acquired and used as an intra-frame coded frame. The minimum coding unit of the inter-frame predictive coding frame is set to the skip coding type, which simplifies the coding process.
It improves encoding efficiency, reduces video file size, and lowers bandwidth usage.
Smart Images

Figure CN121750858A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more particularly to a video generation method, apparatus, electronic device, medium, and product. Background Technology
[0002] In real-time audio and video, live video streaming, and other application systems, there may be situations where the input video footage contains static, fixed content. These static video frames are also directly encoded using a video encoder designed for dynamic video content.
[0003] However, encoding the original static image content also requires encoding operations such as prediction, quantization, and entropy encoding of the various input images before and after, which consumes a lot of computing resources and cannot fully compress the static image. Summary of the Invention
[0004] This disclosure provides a video generation method, apparatus, electronic device, medium, and product that can simplify the video generation process by encoding image frames based on the characteristics of the target video content, thereby simplifying the encoding process, improving encoding efficiency, and obtaining video encoding results with smaller data volume, thus reducing bandwidth usage.
[0005] In a first aspect, embodiments of this disclosure provide a video generation method, the method comprising:
[0006] During the transmission of the video stream generated by the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired.
[0007] Based on the video file structure corresponding to the video stream and the target image frames, a video file that meets the video file structure is generated to obtain the target video.
[0008] In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
[0009] Secondly, embodiments of this disclosure also provide a video generation apparatus, the apparatus comprising:
[0010] The video generation trigger module is used to acquire the target image frame associated with the video generation mode switching operation in response to the video generation mode switching operation during the transmission of the video stream generated based on the video encoder.
[0011] The video encoding generation module is used to generate a video file that meets the video file structure according to the video file structure corresponding to the video stream and the target image frame, so as to obtain the target video;
[0012] In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any embodiment of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any of the embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described in any of the embodiments of the present invention.
[0019] In this embodiment, during the transmission of a video stream generated by a video encoder, in response to a video generation mode switching operation, a target image frame associated with the video generation mode switching operation is acquired. Based on the video file structure corresponding to the video stream and the target image frame, a video file satisfying the video file structure is generated, resulting in a target video. In the target video, the target image frame is used as an intra-coded frame, and the encoding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to a skipped encoding type. This embodiment solves the problems of high computational resource consumption and large file size of static video generated by conventional encoders. It achieves simplified encoding by encoding the image frames of the target video based on the characteristics of the target video content, simplifying the encoding process, improving encoding efficiency and compression ratio, and simultaneously obtaining a smaller video encoding result, reducing bandwidth usage. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 This is a schematic flowchart of a video generation method provided in an embodiment of this disclosure;
[0022] Figure 2 This is a schematic flowchart of a video generation method provided in an embodiment of this disclosure;
[0023] Figure 3 This is a schematic flowchart of a video generation method provided in an embodiment of this disclosure;
[0024] Figure 4 This is a schematic diagram of a video generation method under the Advanced Video Coding Standard (H.264) provided in an embodiment of this disclosure;
[0025] Figure 5 This is a schematic diagram of a video generation method under the efficient video coding standard (H.265) provided in an embodiment of this disclosure;
[0026] Figure 6 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this disclosure;
[0027] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0029] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0030] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0032] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0038] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of the present disclosure. This embodiment is applicable to scenarios where static video content appears in real-time video stream encoding. The method can be executed by a video generation device, which can be implemented in software and / or hardware. Optionally, the video generation device can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0039] like Figure 1 As shown, the video generation methods include:
[0040] S110. During the transmission of the video stream generated based on the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired.
[0041] The transmission of a video stream generated by a video encoder can occur in scenarios such as real-time audio-visual interaction and live video streaming. It involves the video stream processing server sending the video stream data generated by the video encoder at the video stream source to the corresponding content delivery network or video stream receiver. The video stream can be the result of video encoding the captured audio and images at the video stream source using a video encoder. The video encoding process requires encoding intra-frame coded frames and inter-frame predictive coded frames. The encoding of inter-frame predictive coded frames requires image processing processes such as quantization decomposition. The encoded video stream can contain either dynamic or static images.
[0042] However, for videos with static images, using the same encoding process as for dynamic images is relatively cumbersome, resulting in larger video file sizes for static images and excessive bandwidth consumption during transmission.
[0043] In this embodiment, video generation methods can be switched to improve the efficiency of static video generation and reduce the file size of static video content. The process of triggering the video generation method switch can be either passive or active. Passive triggering occurs when the video stream processing server cannot continuously acquire a transmittable (forwarded) real-time video stream due to network congestion or insufficient bandwidth. Active triggering occurs when, during real-time audio-visual interaction or live streaming, the user leaves the screen. In conventional video stream processing, the encoder at the video stream source encodes the captured audio and image data to obtain a static video stream with unchanged visuals, which is then sent as a real-time video stream to the video stream processing server for forwarding. In this case, the user can actively trigger the video generation method switch using a preset control, eliminating the need for the encoder at the video stream source to encode and generate a static video stream, and preventing the uploading of the real-time video stream. The video generation device at the video stream processing server quickly generates a static video stream and forwards it, thereby reducing bandwidth consumption at both the video stream source and server ends.
[0044] In one optional implementation, the same video generation device as the video stream processing server can be configured at the video stream source. During the encoding and generation of the video stream, the user can trigger a video generation mode switch. A target image frame is obtained through conventional encoding by the encoder, and then the video generation device encodes the target video based on the target image frame. The generated target video is then sent to the video stream processing server, reducing the file size of the uploaded video stream and also reducing bandwidth consumption between the video stream source and the video stream processing server to some extent. The target image frame can be an intra-coded frame in the video stream that has already been encoded by the encoder.
[0045] When the video generation mode is switched, either actively or passively, the video generation device on the video stream processing server will initiate a video encoding process based on the target image frame. The target image frame is the image frame associated with the video generation mode switching operation. Typically, the user can pre-set a video frame interpolation map as the target image frame. When the target video needs to be generated, i.e., when the video generation mode switch is triggered, the video frame interpolation map can be sent synchronously to the video generation device so that the content of the final encoded target video is based on this interpolation map.
[0046] S120. Based on the video file structure corresponding to the video stream and the target image frame, generate a video file that meets the video file structure to obtain the target video.
[0047] In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
[0048] Due to differences in video encoding types (encoding standards) used in video streams generated by video encoders, the corresponding video file structures differ. This means the encoding parameter information within the video files varies. In the target video, the target image frame is used as the intra-coded frame; that is, the video encoding information corresponding to the target image frame is used as the encoding information for the intra-coded frames in the target video. The target image frame, also known as an Intra-coded Picture (I-frame) or keyframe, is used as the intra-coded frame in the target video's group of pictures. It contains complete image information and can be decoded independently. It can serve as a reference for inter-frame coding prediction frames within the target video's group of pictures (GOP).
[0049] The video coding information corresponding to the target image frame can also be the video coding information of the inter-frame predictive coding frame, such as the number of slices in the image frame and the distribution of macroblocks. Then, by setting the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-frame coded frame to the skip coding type, the video file of the target video can be edited to obtain the target video. The inter-frame predictive coding frame can be a forward predictive coded picture (P-frame) or a bidirectionally-predicted picture (B-frame).
[0050] Image groups can be formed from intra-frame coded frames and inter-frame predictive coded frames, thus obtaining the target video.
[0051] In this embodiment, during the transmission of the video stream generated by the video encoder, in response to a video generation mode switching operation, a target image frame associated with the video generation mode switching operation is acquired. Based on the video file structure corresponding to the video stream and the target image frame, a video file satisfying the video file structure is generated, resulting in the target video. In the target video, the target image frame is used as the intra-coded frame, and the encoding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to a skipped encoding type. The technical solution of this embodiment solves the problems of high computational resource consumption and large file size of static video generated by conventional encoders. It achieves the simplification of the video generation process by encoding the image frames of the target video based on the characteristics of the target video content, simplifying the encoding process, improving encoding efficiency and compression ratio, and simultaneously obtaining a smaller video encoding result, reducing bandwidth usage.
[0052] Figure 2 This is a flowchart illustrating a video generation method provided in an embodiment of the present disclosure. Based on the above embodiment, it further explains the process of determining the target image frame and encoding the inter-frame predictive coding frame during the implementation of the video generation method. This method can be executed by a video generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.
[0053] like Figure 2 As shown, the video generation methods include:
[0054] S210. During the transmission of the video stream generated based on the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired.
[0055] In one possible scenario, such as when the target image frame associated with the switching operation is not available, the target image frame can be determined from the keyframes in the video stream.
[0056] For example, the target video stream data corresponding to the operation time of the video generation mode switching operation can be obtained; the intra-coded frame in any image group in the target video stream data can be used as the target image frame.
[0057] Alternatively, a preset scene image that matches an intra-coded frame in the image group of the target video stream data can be determined. The preset scene image is then encoded using the video encoding type corresponding to the target video stream data to obtain the scene image encoding result, and the scene image encoding result is determined as the target image frame.
[0058] For example, if the target video stream data contains images of an outdoor tourism scene, a scene image with a corresponding theme can be selected, such as an image including mountains and water or an image including the sky. Then, the preset scene image is encoded using the video encoding type corresponding to the target video stream data to obtain the scene image encoding result, and the scene image encoding result is determined as the target image frame.
[0059] S220. Based on the coding parameter information of the target image frame, determine the video coding syntax elements and corresponding syntax element content of the intra-frame coded frames and inter-frame predictive coded frames in the video file structure.
[0060] The encoding parameter information of the target image frame can be obtained by parsing the target image frame. The parsed encoding parameter information includes at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), and Video Parameter Set (VPS).
[0061] SPS contains global information about the video sequence, which is essential for decoding the entire video sequence. This mainly includes image format information, encoding parameter information, information related to the reference image, level and grade correlation parameters, temporal classification information, visualization availability information, and other information.
[0062] PPS contains parameters associated with a single image / frame that guide the decoder on how to process each frame. These include availability flags for the encoding tool, quantization process-related syntax elements, tile-related syntax elements, and other information that can be shared when encoding an image.
[0063] VPS is a concept in the H.265 / HEVC (High Efficiency Video Coding) standard, primarily used for transmitting video grading information and multi-view video extensions. It contains the overall structural information of the entire video sequence, such as syntax elements shared by multiple sublayers and operation points, and key information about operation points required by the session. Each piece of video coding information corresponds to a video coding syntax element. If the target image frame is taken as an intra-coded frame in the target video, then the video coding syntax elements and their corresponding content can be identified for the intra-coded frame.
[0064] Based on the encoding parameter information of the target image frame, determine the video coding syntax elements and syntax element content of the inter-frame predictive coding frames in the video file structure. This can be done by determining the video coding syntax elements of the inter-frame predictive coding frames in the video file structure based on the encoding type in the encoding parameter information; and by determining the syntax element content of the video coding syntax elements corresponding to the number and distribution of the partitioning units of the inter-frame predictive coding frames based on the image segmentation parameters in the encoding parameters.
[0065] For example, the number and distribution of the partitioning units of the inter-frame predictive coding frame can be determined based on the coding parameter information, and the partitioning units can be divided to obtain the corresponding minimum coding unit.
[0066] During the generation of the target video, the inter-frame predictive coding frames remain consistent with the intra-frame coding frames. The minimum decoding unit may differ depending on the video coding type.
[0067] By analyzing the encoding parameter information of the target image frame, the number of slices and the number and distribution of macroblocks or coding units (CUs) contained in each slice are obtained. The same number of slices and macroblocks or CUs need to be generated when encoding P-frames and / or B-frames.
[0068] For example, in H.264, a video frame can be divided into one or more slices. Each slice contains at least one macroblock and can contain data for the entire frame. The macroblock is the most basic processing unit in H.264 video coding. Each macroblock typically contains 16x16 luma pixels and the corresponding chroma pixels. Macroblocks can be further subdivided into smaller subblocks during encoding to accommodate different image content and coding requirements.
[0069] In H.265, the concept of macroblocks is replaced by coding tree units (CTUs). CTUs can have larger sizes (e.g., 64x64) and can be recursively divided into smaller coding units (CUs) by setting `split_cu_flags`. CUs are the basic prediction units in H.265 and can be adaptively segmented based on the complexity and detail of the image.
[0070] S230. Set the syntax element content of the video coding syntax element corresponding to the coding type of the smallest coding unit of the inter-frame predictive coding frame to the skip coding type.
[0071] Each minimum coding unit is encoded according to the syntax elements of the video coding type corresponding to the video stream, and the coding type of the minimum coding unit is set to the skip coding type to obtain the encoding result of the inter-frame predictive coding frame, so as to meet the requirements of the corresponding video file structure.
[0072] Specifically, after determining the smallest coding units that need to be encoded in each inter-frame predictive coding frame, encoding can be performed according to the syntax elements of the corresponding video coding type, writing the information corresponding to each syntax element. The information corresponding to each syntax element can be the information of the syntax element corresponding to the smallest coding unit in the target image frame. The difference lies in that, during the encoding process, the coding type of the smallest coding unit of the inter-frame predictive coding frame is set to a skip coding type, resulting in the encoded inter-frame predictive coding frame.
[0073] Skip-type macroblocks can improve coding efficiency and reduce coding complexity. Their main characteristic is that they do not perform actual coding; instead, they reconstruct image content through inter-frame prediction. Similarly, a CU can specify its own prediction type. If it is a SKIP-type CU, it can directly reuse a coding unit of the reference frame to obtain the current frame's image.
[0074] S240. Obtain at least one image group based on the coding results of intra-frame coded frames and inter-frame predictive coded frames according to the coding parameter information, and obtain the target video.
[0075] Based on the preset image group length in the coding parameter information, intra-coded frames and a corresponding number of inter-predictive coded frames are combined into an image group. The inter-predictive coded frames are forward predictive coded frames, meaning only P-frames are encoded. The number of image groups is determined based on the preset video frame rate in the coding parameter information to generate a target video containing at least one image group. When the number of groups is greater than one, the image group composed of intra-coded frames and a corresponding number of inter-predictive coded frames can be directly copied to obtain multiple image groups.
[0076] The technical solution of this disclosure, in response to a video generation mode switching operation during the transmission of a video stream generated by a video encoder, acquires a target image frame associated with the video generation mode switching operation; determines the video coding syntax elements and corresponding syntax element content of intra-coded frames and inter-predictive coded frames in the video file structure based on the coding parameter information of the target image frame; sets the syntax element content of the video coding syntax element corresponding to the coding type of the smallest coding unit of the inter-predictive coded frame to a skip coding type; and obtains at least one image group composed of the encoding results of intra-coded frames and inter-predictive coded frames based on the coding parameter information, thereby obtaining the target video. The technical solution of this disclosure solves the problems of high computational resource consumption and large video file size resulting from encoding static video based on conventional encoders. It also solves the problem of determining the target image frame when the associated target image frame is not acquired, realizing encoding of image frames of the target static video based on the characteristics of static content video, simplifying the encoding process, improving encoding efficiency and compression ratio, and obtaining video encoding results with smaller data volume, thus reducing bandwidth usage.
[0077] Figure 3 This is a flowchart illustrating a video generation method provided in an embodiment of this disclosure. Based on the above embodiment, the entropy reduction encoding process is further explained during the static video generation process, which can further improve the static video encoding efficiency. This method can be executed by a video generation device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server.
[0078] like Figure 3 As shown, the video generation methods include:
[0079] S310. During the transmission of the video stream generated based on the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired.
[0080] S320. Determine whether the target image frame has undergone entropy coding during the encoding process.
[0081] Entropy coding is a lossless data compression technique in video coding, designed to reduce data redundancy and improve coding efficiency. Its basic idea is to utilize the probability distribution of source symbols, allowing common symbols to be represented with shorter codewords and uncommon symbols with longer codewords, thereby reducing the overall coding length. Whether entropy coding is used in video coding is determined by parameters in PPS (Programmable Scripting System).
[0082] Whether the image has undergone entropy coding can be determined by the value of the parameters in the PPS corresponding to the target image frame.
[0083] S330. If the target image frame has undergone entropy encoding during the encoding process, decode the target image frame and re-encode it to obtain the target image frame after de-entropy encoding.
[0084] The target image frame is used as the I-frame of the GOP in the target still video. If entropy coding is enabled for the I-frame, the subsequent P-frames or B-frames generated by encoding also need to be consistent with it and use entropy coding as well. Since the encoding process of entropy coding is relatively complex, this disclosure chooses not to use entropy coding in the process of encoding still video in order to further reduce the complexity of still video encoding.
[0085] Therefore, if it is known that the target image frame has undergone entropy coding during the encoding process, the target image frame can be decoded and re-encoded to obtain a de-entropy coded target image frame. That is, entropy coding is not used during the re-coding of the target video frame.
[0086] S340. Parse the encoding parameter information of the target image frame, and generate a video file that meets the video file structure according to the video file structure corresponding to the video stream and the encoding parameter information of the target image frame, thereby obtaining the target video.
[0087] In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
[0088] The video file structure corresponding to the video stream can be any video encoding type used in video coding. Different video encoding types may have the same or different syntax elements in their video file structures. In this step, instead of directly using the encoder of the corresponding video encoding type to encode the static video, the inter-frame predictive coded frames in the image group of the target video are encoded based on the syntax elements in the encoding parameter information of the corresponding video encoding type and the target image frames. These inter-frame predictive coded frames can be forward predictive coded frames (P-frames) and / or bidirectionally-predicted frames (B-frames).
[0089] Specifically, when the syntax elements of the inter-frame predictive coding frame are set with reference to the syntax elements of the target image frame, the coding type of the smallest coding unit of the inter-frame predictive coding frame is set to the skip coding type to obtain the coding result of the inter-frame predictive coding frame. That is, instead of actually coding each coding unit of the inter-frame predictive coding frame, the image content is reconstructed through inter-frame prediction, thereby improving coding efficiency and reducing coding complexity.
[0090] Based on the preset video frame rate and preset image group length, an image group consisting of the encoding results of intra-frame coded frames and inter-frame predictive coded frames is obtained to generate the target static video.
[0091] The preset group size length (GOP) represents the number of image frames in a GOP. Knowing the preset GOP length, we can determine how many inter-frame predictive coding (IPC) frames need to be encoded in addition to the intra-coded frames corresponding to the target image frame. The encoding results of the corresponding number of IPC frames obtained in the previous step can be used to form a GOP of the preset GOP length. During continuous static video encoding, the GOP, based on the encoding results of the intra-coded frames and IPC frames, can be copied according to the preset video frame rate to generate the target video. The encoded target video presents the static image corresponding to the target image frame, significantly reducing the complexity of target video encoding.
[0092] S350, in response to the video generation mode restoration operation, ends the process of generating the target video.
[0093] Corresponding to the switching of video generation methods, the restoration of video generation methods can be triggered actively by the user. For example, during a live broadcast, when the host returns to the live broadcast screen, the process of generating the target video based on the target image frame will end, and the process of generating the target video based on the video encoder will be restored.
[0094] The video generation mode restoration operation can also be triggered when the network bandwidth is restored to normal so that dynamic video stream data can be transmitted normally. In this case, the video generation device can automatically trigger the video generation mode restoration operation and end the process of video encoding based on the target image frame.
[0095] The technical solution of this disclosure, in the transmission of a video stream generated by a video encoder, in response to a video generation mode switching operation, acquires a target image frame associated with the video generation mode switching operation; determines whether the target image frame has undergone entropy encoding processing during the encoding process; if the target image frame has undergone entropy encoding processing during the encoding process, decodes and re-encodes the target image frame to obtain an entropy-de-encoded target image frame; parses the encoding parameter information of the target image frame, and generates a video file that meets the video file structure according to the video file structure corresponding to the video stream and the encoding parameter information of the target image frame, thus obtaining the target video; in response to a video generation mode restoration operation, ends the process of generating the target video. The technical solution of this disclosure solves the problems of high computational resource consumption and large file size of static video generated by conventional encoders, and also solves the problem of how to stop the generation of static video. It realizes the encoding of image frames of the target static video based on the characteristics of static content video, and the entropy-de-encoded encoding process, further simplifying the encoding process, improving encoding efficiency and compression ratio, and obtaining video encoding results with smaller data volume, thus reducing bandwidth usage.
[0096] Figure 4 and Figure 5 The application process of video generation methods is demonstrated in real-time video stream encoding under different video encoding types (video encoding standards). Among them, Figure 4 This is a flowchart illustrating the video generation method under the High-Level Video Coding Standard (H.264). Figure 5 A flowchart illustrating the video generation method under the High Efficiency Video Coding Standard (H.265).
[0097] Specifically, Figure 4 The process shown generates an H.264 static video file, including the following video encoding process:
[0098] After the video generation mode switch is triggered, the target image frame (I-frame) and the corresponding SPS / PPS are acquired.
[0099] 1. Analyze SPS / PPS to determine whether entropy encoding is enabled in the video.
[0100] 2. Parse the I-frame to obtain the number of slices and the number of macroblocks contained in each slice. The same number of slices and macroblocks will be needed when generating P-frames later.
[0101] 3. Calculate the total number of GOPs based on the expected video duration, frame rate, and GOP length.
[0102] 4. For each GOP, copy the input I-frame and use it as the starting I-frame of the GOP.
[0103] 5. Generate the remaining P frames based on the GOP length.
[0104] 6. Based on the slice information obtained in step 2, generate all slices within a P-frame.
[0105] 7. Generate the slice header, with the structure set according to the H.264 syntax elements, and write it to the slice type as a P-frame.
[0106] 8. Based on the number of macroblocks in each slice, write the mb_skip_run syntax element, indicating that all macroblocks in the slice are of type SKIP.
[0107] 9. Write the remaining syntax elements within the slice, such as rbsp_trailing_bits at the end.
[0108] Figure 5 The process shown generates an H.265 still video file; the main flow is similar to... Figure 4 The steps are similar, except that the special structure of H.265 needs to be considered when generating P-frames. Specifically, the video encoding process includes the following:
[0109] After the video generation mode switch is triggered, the target image frame (I-frame) and the corresponding SPS / PPS / VPS are acquired.
[0110] 1. Determine whether tile division is enabled based on the tiles_enabled_flag field in PPS.
[0111] If tile division is enabled, CTU needs to be generated by dividing each tile according to the number of rows and columns in PPS, the width and height of each tile, and the overall screen resolution.
[0112] 2. Based on the current CTU size, log2_min_cb_size_y, and log2_diff_max_min_luma_coding_block_size in SPS, determine whether the CTU can be further divided into smaller quadtrees.
[0113] 3. If further division is possible, recursively traverse the four regions within the CTU (top left, top right, bottom left, and bottom right) to determine whether further division is possible.
[0114] 4. Once the minimum segment size is reached, begin writing the contents of the CU. The specific structure is detailed in the syntax elements of H.265. The main step is to set CU_SKIP_FLAG to 1.
[0115] 5. Write other syntax elements into CU.
[0116] Figure 4 and Figure 5 The implementation process uses methods other than encoders, directly referencing the file structure in the video encoding type to generate the encoded video file. It fully utilizes the characteristics of static content in the encoding standard to specifically generate the target video file.
[0117] Taking the generation of a 1000-frame video (25 frames per GOP) as an example, the encoder generates a video of 32,452,840 bytes in size, taking 12,387 ms. In contrast, the video generated using the scheme of this embodiment is 26,202,800 bytes in size (file size -19.25%), taking 131 ms (time reduction -98.9%, including the time to generate the first I-frame).
[0118] Figure 6 This disclosure provides a video generation apparatus suitable for scenarios where static video content appears in real-time video stream encoding. The video generation apparatus can be implemented in software and / or hardware and can be configured on an electronic device, such as a mobile terminal, a PC, or a server.
[0119] like Figure 6 As shown, the video generation device includes a video generation trigger module 410 and a video encoding generation module 420.
[0120] The video generation trigger module 410 is used to acquire the target image frame associated with the video generation mode switching operation in response to the video generation mode switching operation during the transmission of the video stream generated based on the video encoder; the video encoding generation module 420 is used to generate a video file that meets the video file structure according to the video file structure corresponding to the video stream and the target image frame, thereby obtaining the target video; wherein, in the target video, the target image frame is used as the intra-coded frame, and the encoding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to the skip coding type.
[0121] The technical solution of this disclosure, in response to a video generation mode switching operation during the transmission of a video stream generated by a video encoder, acquires a target image frame associated with the video generation mode switching operation; and generates a video file that satisfies the video file structure according to the video file structure corresponding to the video stream and the target image frame, thus obtaining a target video. In the target video, the target image frame is used as an intra-coded frame, and the encoding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to a skipped encoding type. This technical solution solves the problems of high computational resource consumption and large file size of static video generated by conventional encoders. It achieves the simplification of the video generation process by encoding the image frames of the target video based on the characteristics of the target video content, simplifying the encoding process, improving encoding efficiency and compression ratio, and simultaneously obtaining a smaller video encoding result, thus reducing bandwidth usage.
[0122] In one alternative implementation, the video encoding generation module 420 is specifically used for:
[0123] Based on the coding parameter information of the target image frame, determine the video coding syntax elements and corresponding syntax element content of the intra-frame coded frames and inter-frame predictive coded frames in the video file structure;
[0124] Set the syntax element content of the video coding syntax element corresponding to the coding type of the smallest coding unit of the inter-frame predictive coding frame to the skip coding type.
[0125] Based on the coding parameter information, at least one image group is obtained, consisting of the coding results of intra-frame coded frames and inter-frame predictive coded frames, to obtain the target video.
[0126] In one alternative implementation, the video encoding generation module 420 is specifically used for:
[0127] Based on the preset image group length in the coding parameter information, the intra-frame coded frames and the corresponding number of inter-frame predictive coded frames are combined into an image group; among them, the inter-frame predictive coded frames are forward predictive coded frames;
[0128] The number of image groups is determined based on the preset video frame rate in the encoding parameter information, so as to generate a target video containing at least one image group.
[0129] In one alternative implementation, the video encoding generation module 420 is specifically used for:
[0130] Based on the encoding type in the encoding parameter information, determine the video coding syntax elements of the inter-frame predictive coding frames in the video file structure;
[0131] Based on the image segmentation parameters in the coding parameters, determine the number and distribution of the partitioning units of the inter-frame predictive coding frames, and the corresponding syntax element content of the video coding syntax elements.
[0132] In an alternative implementation, the video generation trigger module 410 can also be used for:
[0133] If the target image frame associated with the switching operation is not obtained, the target image frame is determined from the keyframes in the video stream.
[0134] In an optional implementation, the video generation trigger module 410 may further be used to:
[0135] Obtain the target video stream data corresponding to the operation time of the video generation method switching operation;
[0136] The intra-coded frame in any image group in the target video stream data is used as the target image frame; or, a preset scene image that matches the intra-coded frame in the image group in the target video stream data is determined, and the preset scene image is encoded using the video encoding type corresponding to the target video stream data to obtain the scene image encoding result, and the scene image encoding result is determined as the target image frame.
[0137] In an alternative implementation, the video encoding generation module 420 can also be used for:
[0138] After acquiring the target image frame, determine whether the target image frame has undergone entropy coding processing during the encoding process;
[0139] If the target image frame has undergone entropy encoding during the encoding process, the target image frame is decoded and re-encoded to obtain the de-entropy encoded target image frame.
[0140] In one optional implementation, the video generation pause module of the video generation device is used for:
[0141] In response to the video generation method restoration operation, the process of generating the target video ends.
[0142] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0143] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0144] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0145] like Figure 7 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 505.
[0146] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0147] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0148] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0149] The electronic device provided in this embodiment and the video generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0150] This disclosure also provides a computer storage medium storing a computer program that, when executed by a processor, implements the video generation method provided in the above embodiments.
[0151] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0152] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0153] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0154] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0155] During the transmission of the video stream generated by the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired.
[0156] Based on the video file structure corresponding to the video stream and the target image frames, a video file that meets the video file structure is generated to obtain the target video.
[0157] In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
[0158] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0160] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0161] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0163] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method provided in any embodiment of this disclosure.
[0164] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0165] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0166] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0167] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video generation method, characterized in that, include: During the transmission of the video stream generated by the video encoder, in response to the video generation mode switching operation, the target image frame associated with the video generation mode switching operation is acquired. Based on the video file structure corresponding to the video stream and the target image frame, a video file that satisfies the video file structure is generated to obtain the target video; In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
2. The method according to claim 1, characterized in that, Based on the video file structure corresponding to the video stream and the target image frame, a video file satisfying the video file structure is generated to obtain the target video, including: Based on the encoding parameter information of the target image frame, determine the video coding syntax elements and corresponding syntax element content of the intra-frame coded frames and inter-frame predictive coded frames in the video file structure; Set the syntax element content of the video coding syntax element corresponding to the coding type of the smallest coding unit of the inter-frame predictive coding frame to the skip coding type. Based on the coding parameter information, at least one image group is obtained, which consists of the coding results of the intra-coded frames and the inter-predictive coded frames, to obtain the target video.
3. The method according to claim 2, characterized in that, The step of obtaining at least one image group based on the coding results of the intra-frame coded frames and the inter-frame predictive coded frames according to the coding parameter information, to obtain the target video, includes: Based on the preset image group length in the encoding parameter information, the intra-frame coded frames and the corresponding number of inter-frame predictive coded frames are combined into an image group; wherein, the inter-frame predictive coded frames are forward predictive coded frames; The number of image groups is determined based on the preset video frame rate in the encoding parameter information, so as to generate a target video containing at least one of the image groups.
4. The method according to claim 2, characterized in that, Based on the encoding parameter information of the target image frame, the video coding syntax elements and syntax element content of the inter-frame predictive coding frames in the video file structure are determined, including: Based on the encoding type in the encoding parameter information, determine the video coding syntax elements of the inter-frame predictive coding frames in the video file structure; Based on the image segmentation parameters in the encoding parameters, the syntax element content of the video coding syntax elements corresponding to the number and distribution of the partitioning units of the inter-frame predictive coding frames is determined.
5. The method according to claim 1, characterized in that, The method further includes: If the target image frame associated with the switching operation is not obtained, the target image frame is determined from the keyframes in the video stream.
6. The method according to claim 5, characterized in that, Determining the target image frame from the keyframes in the video stream includes: Obtain the target video stream data corresponding to the operation time of the video generation method switching operation; The intra-coded frame in any image group of the target video stream data is taken as the target image frame; or, a preset scene image that matches the intra-coded frame in the image group of the target video stream data is determined, the preset scene image is encoded using the video encoding type corresponding to the target video stream data to obtain a scene image encoding result, and the scene image encoding result is determined as the target image frame.
7. The method according to any one of claims 1-6, characterized in that, After acquiring the target image frame, the process also includes: Determine whether the target image frame has undergone entropy coding processing during the encoding process; If the target image frame has undergone entropy encoding during the encoding process, the target image frame is decoded and re-encoded to obtain a de-entropy encoded target image frame.
8. The method according to claim 1, characterized in that, Also includes: In response to the video generation method restoration operation, the process of generating the target video ends.
9. A video generation apparatus, characterized in that, include: The video generation triggering module is used to acquire the target image frame associated with the video generation mode switching operation in response to the video generation mode switching operation during the transmission of the video stream generated based on the video encoder. The video encoding generation module is used to generate a video file that satisfies the video file structure according to the video file structure corresponding to the video stream and the target image frame, thereby obtaining the target video; In the target video, the target image frame is used as the intra-coded frame, and the coding type of the smallest coding unit of the inter-frame predictive coding frame corresponding to the intra-coded frame is set to skip coding type.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the video generation method as described in any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-8.