An AI-based video encoding method, device, equipment and storage medium

CN116582686BActive Publication Date: 2026-10-09BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310396495.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2026-10-09
Estimated Expiration
2043-04-12

AI Technical Summary

Technical Problem

[0005]本申请实施例提供了一种基于AI的视频编码方法、装置、设备和存储介质,解决了利用现有AI视频编码技术对视频进行编码时,存在编码模型适应性较差、模型训练难度较大以及传输码率较高的问题,通过对待编码视频的源参考帧和驱动帧采用关键点检测网络输出关键点信息,并根据源参考帧的关键点信息以及预设压缩规则,确定各驱动帧的关键点信息压缩结果,进而基于所述源参考帧、源参考帧的关键点信息以及各驱动帧的关键点信息压缩结果生成码流数据进行传输,可以实现超低码率视频编码的效果,通过对视频数据的深度压缩,并且结合AI模型来实现数据的编码过程,可以获得更好的视频压缩效果,同时,本方案也提高了AI视频编码模型的适用范围、降低了模型开发训练的难度与成本

Benefits of technology

[0021] Fifthly, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the AI-based video coding method described in embodiments of this application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116582686B_ABST
    Figure CN116582686B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses an AI-based video coding method, device, equipment and storage medium, which comprises the following steps: acquiring a video to be coded; adopting a key point detection network to output key point information of a source reference frame and a driving frame of the video to be coded; determining key point information compression results of each driving frame according to the key point information of the source reference frame and a preset compression rule; and generating code stream data based on the source reference frame, the key point information of the source reference frame and the key point information compression results of each driving frame to realize transmission. The present scheme realizes the effect of ultra-low code rate video coding, realizes the coding process of data by deep compression of video data and combination of an AI model, and can obtain better video compression effect. Meanwhile, the present scheme also improves the application range of the AI video coding model and reduces the difficulty and cost of model development and training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to an AI-based video encoding method, apparatus, device, and storage medium. Background Technology

[0002] With the continuous development of the internet industry, the application scenarios for video are becoming increasingly widespread, such as smart security, autonomous driving, smart cities, and the industrial internet. More and more fields are beginning to use video as an auxiliary means of processing work, while more and more people are gradually incorporating video sharing and viewing into their entertainment. Because video data is inherently very large, it needs to be encoded to ensure successful transmission and presentation on device terminals. Therefore, research on video encoding technology, especially AI-based video encoding technology, has become extremely important.

[0003] Existing AI video coding methods mainly include Deep Video Coding (DVC) and machine vision coding. DVC is an end-to-end video coding model where the entire video compression framework is implemented by a neural network and can be trained uniformly. By replacing individual modules in the traditional coding framework with deep neural networks, it achieves both video encoding and image frame reconstruction. Machine vision coding is a video coding technology aimed at intelligent applications. It combines video encoding and decoding with machine vision analysis, utilizing an end-to-end network system to perform video encoding and decoding tasks, enabling machines to complete visual tasks.

[0004] Because the Deep Video Coding (DVC) framework is implemented using neural networks and the deep learning methods employed are primarily optimized offline, it suffers from poor model adaptability, complex model deployment and implementation, and high transmission bitrate. Furthermore, machine vision coding models, in addition to video encoding and decoding, are more importantly used to perform machine vision tasks using the decoded video, resulting in highly specific models and significant training difficulties. Moreover, existing AI video coding technologies typically require large floating-point data representations, which do not offer a significant advantage over high-efficiency P-frame compression in video coding. Therefore, using existing AI video coding technologies to encode video presents challenges such as poor model adaptability, high training difficulty, and high transmission bitrate. Summary of the Invention

[0005] This application provides an AI-based video encoding method, apparatus, device, and storage medium, which solves the problems of poor adaptability of encoding models, high training difficulty, and high transmission bitrate when encoding videos using existing AI video encoding technologies. By using a keypoint detection network to output keypoint information for the source reference frame and driving frame of the video to be encoded, and determining the keypoint information compression result of each driving frame based on the keypoint information of the source reference frame and preset compression rules, a bitstream data is generated and transmitted based on the source reference frame, the keypoint information of the source reference frame, and the keypoint information compression result of each driving frame. This achieves the effect of ultra-low bitrate video encoding. By performing deep compression of video data and combining it with an AI model to implement the data encoding process, better video compression effect can be obtained. At the same time, this solution also improves the applicability of AI video encoding models and reduces the difficulty and cost of model development and training.

[0006] In a first aspect, embodiments of this application provide an AI-based video encoding method, the method comprising:

[0007] Obtain the video to be encoded;

[0008] The source reference frame and driving frame of the video to be encoded are used to output key point information using a key point detection network;

[0009] Based on the key point information of the source reference frame and the preset compression rules, determine the compression result of the key point information of each driving frame;

[0010] Based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, a bitstream data is generated and transmitted.

[0011] Secondly, embodiments of this application also provide an AI-based video encoding device, comprising:

[0012] The video acquisition module is used to acquire the video to be encoded.

[0013] The key point information output module is used to output key point information from the source reference frame and the driving frame of the video to be encoded using a key point detection network.

[0014] The key point information compression module is used to determine the key point information compression result of each driving frame based on the key point information of the source reference frame and the preset compression rules.

[0015] The bitstream data transmission module is used to generate bitstream data for transmission based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame.

[0016] Thirdly, embodiments of this application also provide an AI-based video encoding device, which includes:

[0017] One or more processors;

[0018] Storage device for storing one or more programs.

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the AI-based video encoding method described in the embodiments of this application.

[0020] Fourthly, embodiments of this application also provide a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to execute the AI-based video coding method described in embodiments of this application.

[0021] Fifthly, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the AI-based video coding method described in embodiments of this application.

[0022] In this embodiment, a video to be encoded is obtained; a key point detection network is used to output key point information for the source reference frame and driving frame of the video to be encoded; the key point information compression result of each driving frame is determined according to the key point information of the source reference frame and a preset compression rule; and bitstream data is generated and transmitted based on the source reference frame, the key point information of the source reference frame and the key point information compression result of each driving frame. The aforementioned AI-based video coding method addresses the problems of poor model adaptability, high training difficulty, and high transmission bitrate when using existing AI video coding technologies. By employing a keypoint detection network to output keypoint information from the source reference frame and driving frame of the video to be encoded, and determining the keypoint compression result of each driving frame based on the keypoint information of the source reference frame and preset compression rules, a bitstream data is generated for transmission based on the source reference frame, its keypoint information, and the compression results of each driving frame. This achieves ultra-low bitrate video coding. Through deep compression of video data and the integration of an AI model into the data encoding process, better video compression results can be obtained. Furthermore, this solution expands the applicability of the AI ​​video coding model and reduces the difficulty and cost of model development and training. Attached Figure Description

[0023] Figure 1 A flowchart illustrating an AI-based video encoding method provided in this application embodiment;

[0024] Figure 2A schematic diagram of an AI-based video coding system framework provided in this application embodiment;

[0025] Figure 3 A flowchart illustrating an AI-based video encoding method provided in this application embodiment;

[0026] Figure 4 A flowchart illustrating an AI-based video encoding method provided in this application embodiment;

[0027] Figure 5 A flowchart illustrating key point coordinate compression provided in an embodiment of this application;

[0028] Figure 6 A flowchart of a reconstructed image evaluation method provided in an embodiment of this application;

[0029] Figure 7 A structural block diagram of an AI-based video encoding device provided in this application embodiment;

[0030] Figure 8 This is a schematic diagram of the structure of an AI-based video encoding device provided in an embodiment of this application. Detailed Implementation

[0031] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the scope of the embodiments. Furthermore, it should be noted that, for ease of description, only the parts relevant to the embodiments of this application are shown in the accompanying drawings, not the entire structure.

[0032] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0033] For some face video generation and compression schemes, based on the First Order Motion Model (FOM) framework, the first frame (reference frame) is first encoded and transmitted using a traditional encoder. Then, key points of subsequent frames are extracted at the encoding end and transmitted to the decoding end. The decoding end relies on the relative displacement of key points between the current frame and the reference frame to distort the face in the reference frame and fill the background to generate the current frame.

[0034] However, this method is only suitable for simple head-motion videos with minimal movement. Once the video experiences large head shifts, facial obstruction, or the addition of background objects, the quality of the generated current frame will significantly decrease. Issues such as facial distortion, shakiness, and missing background objects severely impact the viewing experience. In such cases, frequent updates of the reference frame are required to resolve the problem, which undoubtedly increases the bitrate burden.

[0035] To explore the compression limits and application scenarios of generative compression schemes, this invention proposes an ultra-low bitrate generative encoding method. This method can compress the original keypoint data and reuse the keypoint data, further reducing the keypoint information that needs to be transmitted at the encoding end. It achieves a bitrate reduction of at least 1 / 10 compared to HEVC, while ensuring no significant decrease in subjective quality.

[0036] Existing FOM-based generative compression schemes are still in the exploratory stage, with significant room for improvement in bitrate compression and driving strategies. Firstly, key information typically requires 240 bytes of floating-point data for representation, which is not advantageous compared to P-frame compression in HEVC. Therefore, given the stringent requirements of transmission bitrate and subjective quality in real-world applications, further research is needed to explore the compression limits achievable by generative compression schemes and their effective application scenarios.

[0037] This invention addresses the problems existing in the FOM generation and compression scheme, and, combined with the needs of mobile live streaming services, constructs an end-to-end generation and encoding system framework. Through pre-encoding processing and the FOM generation and compression link, it provides new ideas for frame loss regeneration and generating frames to replace traditional encoded frames for the transcoding engine side and the decoding side of the audience side, respectively.

[0038] Figure 1 The flowchart provided in this application illustrates an AI-based video encoding method, applicable to video image transmission scenarios, particularly those involving video transmission through video encoding compression. This method can be executed by servers and smart terminals possessing encoding and computing capabilities. Specifically, the method includes the following steps:

[0039] S101, Obtain the video to be encoded;

[0040] In one embodiment, the video to be encoded can be video data to be transmitted, which may consist of multiple identical or different frames. By encoding the video, video compression can be achieved, enabling the video to be successfully transmitted to the user terminal device.

[0041] In this solution, the video to be encoded can be acquired using a video recording device, such as a camera. After acquiring the video, data acquisition and image preprocessing can be performed on the images in the video, such as classification, judgment, and enhancement, which can improve the robustness of the video encoding model training and the efficiency of video information acquisition.

[0042] S102, the key point detection network is used to output key point information for the source reference frame and driving frame of the video to be encoded;

[0043] The source reference frame can be the first frame of the video to be encoded. The driving frame can be any frame in the video following the source reference frame. The image content and image data of each frame in the driving frame can be the same or different, and the image content and image data of the driving frame can be the same or different from those of the source reference frame.

[0044] In one embodiment, keypoints can be points that represent the image features of each frame in the video to be encoded, including points representing the overall shape or content distribution of the source reference frame and the driving frame, as well as points that change in the next frame or multiple driving frame images. For example, in a teaching video mainly featuring a teacher, the teacher's head and facial features can be used as keypoints in the video to be encoded. Keypoint information can be information used for video encoding of the video to be encoded, including the content information, location information, and data information of the keypoints, such as pixel data, position coordinate data, and data information within a certain range around the coordinates of the keypoints.

[0045] The keypoint detection network can be a network used to acquire the keypoint information, including a keypoint detector such as a Kp-detector. By applying a sparse motion field network within the keypoint detection network to the source reference frame and the driving frame in the video to be encoded, the keypoint information of the source reference frame and the driving frame can be acquired.

[0046] S103, Based on the key point information of the source reference frame and the preset compression rules, determine the compression result of the key point information of each driving frame;

[0047] The preset compression rules can be rules for compressing each driving frame in the video to be encoded to reduce the size of the video according to the video compression requirements, including intra-frame compression rules and inter-frame compression rules. The key point information compression result of each driving frame can be the result of compressing the key point data of each driving frame into encoded data, including the intra-frame compression result of different key point information data in the same driving frame and the inter-frame compression result of the same key point data in different driving frames.

[0048] In one embodiment, an end-to-end generative coding system model based on FOM (First Order Motion Model) can be used. A keypoint data compression method is employed, performing intra-frame compression of keypoint information for each driving frame according to preset compression rules, and inter-frame compression of keypoint information for each driving frame according to the relationship between the source reference frame and the keypoint information of each driving frame, as well as the preset compression rules. The keypoint data compression may include pixel data compression, position coordinate data compression, and data type conversion, etc., without further limitation here. By performing intra-frame and inter-frame compression on each driving frame, the compression result of the keypoint information for each driving frame can be determined.

[0049] This solution further compresses key information using preset compression rules, achieving a significant improvement in compression efficiency compared to the original compressed key information. This results in deep data compression and ultra-low bitrate encoding.

[0050] S104, based on the source reference frame, the key point information of the source reference frame, and the compression result of the key point information of each driving frame, a bitstream data is generated and transmitted.

[0051] The bitstream data can be the encoded and compressed data of the source reference frame and each driving frame in the video to be encoded. The video to be encoded can be transmitted to the user terminal device in the form of bitstream data. The bitstream data is generated based on the image encoding result of the source reference frame, the key point information encoding result of the source reference frame, and the key point information compression result of each driving frame.

[0052] In one embodiment, taking a video primarily featuring a face as an example, when the user terminal device receives the bitstream, for videos with no facial obstruction, minimal head movement, and simple backgrounds, the content, position, and data information of the key points in the source reference frame and each driving frame change relatively little. That is, the key point information of the source reference frame and each driving frame has a high correlation. Therefore, the driving frames can be directly generated using the unidirectional reference of the source reference frame to reconstruct the video to be encoded. For example, based on the key point information of the source reference frame and preset compression rules, dense motion field calculations are performed to generate driving frame images that are adjacent to or not adjacent to the source reference frame.

[0053] In another embodiment, when the user terminal device receives the bitstream, for videos with occasional facial occlusion, moderate motion, and simple backgrounds, the content, position, and data information of key points in the source reference frame and each driving frame vary significantly. That is, the correlation between the key point information of the source reference frame and each driving frame is low. After video compression, it is impossible to accurately reconstruct the driving frame image based on the unidirectional reference of the source reference frame. Therefore, a backward reference can be added to the FOM (Form of Image) to adopt a bidirectional reference driving strategy to generate the final image of each driving frame. For example, the key point information of the current driving frame d and the driving frames at positions n frames before and after the current driving frame d are obtained through a sparse motion field, and the driving frames at positions n frames before and after the current driving frame are used as the forward reference frame d. -n and backward reference frame d +n , forward reference frame d -n and backward reference frame d +n After being subjected to a dense motion field, reference images with their respective reference directions are generated. Finally, the reference images of the forward and backward reference frames are fused to obtain a bidirectional fused reference image M, using the formula:

[0054] output = d -n ×M+d +n ×(1-M);

[0055] This guides the final generation of the current driving frame image. Compared to using only unidirectional reference, bidirectional reference significantly improves subjective quality and reduces facial jitter. Better subjective quality is achieved by adjusting the update frequency of the reference source, and a preprocessing module is added to detect scene changes. Scene change detection methods include a scene detection network and calculation of mutual information between adjacent frames. By detecting scene changes, it is determined whether the estimated update probability of the current frame is greater than the average predicted update probability of previously encoded frames, and whether the mutual information between the current frame and adjacent frames is greater than the average mutual information calculated from previously encoded frames. This indicates a more accurate source reference update location, ultimately achieving a certain degree of improvement in subjective quality and improving poor generation results. For videos without facial occlusion, with minimal head movement, and with simple backgrounds, the same bidirectional reference driving strategy can be used according to the user's requirements for video restoration quality; no further limitations are imposed here.

[0056] In one embodiment, optionally, after transmitting the bitstream data generated based on the source reference frame, the key point information of the source reference frame, and the key point information of each driving frame, the method further includes:

[0057] The transcoding engine is used to identify whether there are terminals that cannot recognize the bitstream data.

[0058] If it exists, the bitstream data is converted based on the target terminal parameters, and the converted bitstream data is transmitted to the target terminal;

[0059] If it does not exist, the bitstream data will be directly transmitted to the corresponding terminal.

[0060] The transcoding engine can receive the bitstream data generated by the above scheme and perform transcoding processing as needed, taking into account server requirements such as resolution. The transcoding engine can have multiple output ports for forwarding the bitstream data to different user terminal devices. The transcoding engine can identify whether there are user terminal devices whose data transmission conditions are inconsistent with the bitstream data, i.e., user terminal devices that cannot identify or transmit the bitstream data.

[0061] In one embodiment, the target terminal may be a user terminal device that cannot transmit the bitstream. The target terminal parameters may be the bitstream transmission parameters of the target terminal. If the transcoding engine identifies the existence of the target terminal, it decodes the bitstream data and re-encodes it according to the target terminal parameters, then forwards the converted bitstream data to the corresponding target terminal. Since the transcoding engine can connect to multiple different user terminal devices simultaneously, the encoding method and encoding standard used by the transcoding engine for re-encoding can be the same or different for different target terminals. If the transcoding engine does not identify the existence of the target terminal, the bitstream data is directly transmitted to the corresponding user terminal device for video regeneration.

[0062] In another embodiment, the transcoding engine and the user terminal device can decode the source reference frame bitstream to regenerate the source reference frame image, and regenerate each driving frame image according to the driving frame order based on the regenerated source reference frame and bitstream data for video restoration. Simultaneously, the transcoding engine and the user terminal device can determine whether frame loss exists based on the sequence number of each driving frame in the bitstream; if so, they can regenerate the lost frames by encoding adjacent driving frames.

[0063] This solution uses a transcoding engine to identify whether there are terminals that cannot recognize the bitstream data, and converts the bitstream data based on the target terminal parameters. Then, the converted bitstream data is transmitted to the target terminal. This can avoid the problem of not being able to receive bitstream data due to different configurations of user terminal devices, and improve the accuracy and efficiency of user terminal devices in receiving and restoring transmitted video.

[0064] The technical solution provided in this application embodiment obtains a video to be encoded; uses a keypoint detection network to output keypoint information for the source reference frame and driving frame of the video to be encoded; determines the keypoint information compression result of each driving frame based on the keypoint information of the source reference frame and a preset compression rule; and generates bitstream data for transmission based on the source reference frame, the keypoint information of the source reference frame, and the keypoint information compression result of each driving frame. The aforementioned AI-based video coding method addresses the problems of poor model adaptability, high training difficulty, and high transmission bitrate when using existing AI video coding technologies. By employing a keypoint detection network to output keypoint information from the source reference frame and driving frame of the video to be encoded, and determining the keypoint compression result of each driving frame based on the keypoint information of the source reference frame and preset compression rules, a bitstream data is generated for transmission based on the source reference frame, its keypoint information, and the compression results of each driving frame. This achieves ultra-low bitrate video coding. Through deep compression of video data and the integration of an AI model into the data encoding process, better video compression results can be obtained. Furthermore, this solution expands the applicability of the AI ​​video coding model and reduces the difficulty and cost of model development and training.

[0065] Figure 2 A schematic diagram of an AI-based video encoding system framework is provided for an embodiment of this application, as shown below. Figure 2 As shown, it specifically includes: acquisition / preprocessing module, AI Enc module, transcoding engine module and AI Dec module.

[0066] The acquisition / preprocessing module is used to perform preprocessing operations such as classification, enhancement, and highlighting on the video to be encoded, and then transmits the processed video to be encoded to the AI ​​Enc module.

[0067] The AI ​​Enc module is used to encode and compress the video to be encoded, including one traditional encoding link and two AI encoding links. The traditional encoding link connects the Codec Engine module to the mixing module. The Codec Engine module encodes the source reference frame image in the video to be encoded, enabling the source reference frame image to enter the mixing module and be transmitted to the user terminal. The mixing module mixes the image encoding results of the source reference frame with the image encoding results of each driving frame to generate a bitstream for transmission of the encoded video.

[0068] One part of the AI ​​encoding chain connects to the Codec Engine module and is used to encode and compress the source reference frame image. This includes a DPB (Decoded Picture Buffer) module, a keypoint detection Net, and a generative information compression module. The DPB module decodes and buffers the source reference frame image encoded by the Codec Engine module, and then transmits the decoded image to the keypoint detection Net for keypoint information detection. Alternatively, the source reference frame image can be directly input to the keypoint detection Net. The generative information compression module encodes and compresses the keypoint information detected by the keypoint detection Net and transmits the compressed keypoint information to the mixing module in the traditional encoding chain.

[0069] Another AI encoding link includes a keypoint detection network (Keypoint Detection Network) and an AI Dec (AI Decomposition Network). The AI ​​Dec includes a sparse motion field network (SRF) and a dense motion field network (DNF). The Keypoint Detection Network performs keypoint information detection on each driving frame image and transmits the detected keypoint information to the generation information compression module. The AI ​​Dec regenerates each frame image based on the detected keypoint information and transmits the generated frame image to a discrimination network (Discrimination Network). Together with the Discrimination Network, they form a generative adversarial network (GAN) to evaluate the reliability and accuracy of the generated frame. Specifically, the Sparse Motion Field Network in the AI ​​Dec is used to obtain specific keypoint data from the keypoint information of each frame image, while the DNF is used to regenerate each frame image based on this data.

[0070] The transcoding engine module is used for transcoding the bitstream data and regenerating dropped frame data, including Codec Dec and AI Dec. Codec Dec is used to generate a source reference frame image based on the bitstream data and the source reference frame image encoding data. AI Dec is used to regenerate dropped frames based on the bitstream data and the generated adjacent reference frames or driving frame images.

[0071] The AI ​​Dec module receives bitstream data directly transmitted by the AI ​​Enc module or bitstream data forwarded by the transcoding engine module, including Codec Dec and AI Dec. Codec Dec generates a source reference frame image based on the bitstream data and source reference frame image encoding data. AI Dec performs frame regeneration based on the bitstream data and already generated adjacent reference frames or driving frame images.

[0072] Figure 3 A flowchart of an AI-based video encoding method provided in this application embodiment is shown below. Figure 3 As shown, the specific steps include the following:

[0073] S201, Obtain the video to be encoded;

[0074] S202, the key point detection network is used to output the key point coordinates and the corresponding Jacobian matrix for the source reference frame and driving frame of the video to be encoded.

[0075] The keypoint coordinates can be the coordinates of each keypoint in the source reference frame and each driving frame in the same coordinate system, representing the position of each keypoint. The changes in the image of each frame can be determined by the changes in the keypoint coordinates.

[0076] The Jacobian matrix can be the Jacobian matrix corresponding to the keypoint coordinates. It is a matrix formed by arranging the first-order partial derivatives of the keypoint coordinates in a certain way, and its determinant is called the Jacobian determinant. It embodies the optimal linear approximation of a differentiable equation with a given point. The keypoint coordinates are obtained and output using a sparse motion field in a keypoint detection network on the source reference frame and driving frame of the video to be encoded, and then the Jacobian matrix corresponding to the keypoint coordinates is calculated.

[0077] S203, read the source Jacobian matrix corresponding to the coordinates of each source key point of the source reference frame, and read the driving Jacobian matrix corresponding to the coordinates of each driving key point of each driving frame.

[0078] The source Jacobian matrix can be the Jacobian matrix corresponding to the coordinates of each source keypoint in the source reference frame, and the driving Jacobian matrix can be the Jacobian matrix corresponding to the coordinates of each driving keypoint in each driving frame. The coordinates of each driving keypoint can be the same or different, therefore the driving Jacobian matrices can also be the same or different accordingly. The source Jacobian matrix corresponding to the coordinates of each source keypoint in the source reference frame and the driving Jacobian matrix corresponding to the coordinates of each driving keypoint in each driving frame are read by the keypoint detection network outputting the keypoint coordinates and their corresponding Jacobian matrices.

[0079] S204, Based on the source Jacobian matrix and the first preset compression rule, the driving Jacobian matrix of each driving frame is compressed to obtain the driving Jacobian matrix compression result.

[0080] The first preset compression rule can be a rule for compressing the driving Jacobian matrix of each driving frame, it can be a residual calculation of the driving Jacobian matrices of adjacent driving frames, or it can be a residual calculation of each driving Jacobian matrix with the source Jacobian matrix, etc., without further limitations here. Based on the source Jacobian matrix and the first preset compression rule, the driving Jacobian matrix of each driving frame can be compressed to obtain the driving Jacobian matrix compression result.

[0081] In one embodiment, optionally, before compressing the driving Jacobian matrix of each driving frame according to the source Jacobian matrix and the first preset compression rule to obtain the driving Jacobian matrix compression result, the method further includes:

[0082] The source Jacobian matrix and the driving Jacobian matrix are converted from Float32 data to Float16 data.

[0083] In one embodiment, to maximize the compression of the driving Jacobian matrix, before compressing the driving Jacobian matrix of each driving frame according to the source Jacobian matrix and a first preset compression rule to obtain the compressed driving Jacobian matrix, the precision of the source Jacobian matrix and the driving Jacobian matrix can be reduced first. Specifically, the precision of the matrix can be reduced by converting the format of the source Jacobian matrix and the driving Jacobian matrix from Float32 data to Float16 data. In this case, the amount of data required to calculate each driving frame is 100 bytes, and the compression is approximately 1 / 5.09 of the source video.

[0084] In one embodiment, by performing format conversion on the source Jacobian matrix and the driving Jacobian matrix, the precision of the Jacobian matrix can be reduced, thereby achieving the purpose of preliminary video compression.

[0085] S205, based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, a bitstream data is generated and transmitted.

[0086] The technical solution provided in this application embodiment uses a keypoint detection network to output keypoint coordinates and corresponding Jacobian matrices for the source reference frame and driving frame of the video to be encoded. It reads the source Jacobian matrix corresponding to the source keypoint coordinates of the source reference frame and the driving Jacobian matrix corresponding to the driving keypoint coordinates of each driving frame. Then, according to the source Jacobian matrix and the first preset compression rule, it compresses the driving Jacobian matrix of each driving frame to obtain the driving Jacobian matrix compression result. This can improve the simplicity and efficiency of Jacobian matrix compression of driving frames and reduce the difficulty and cost of developing and training video coding models.

[0087] Figure 4 A flowchart of an AI-based video encoding method provided in this application embodiment is shown below. Figure 4 As shown, the specific steps include the following:

[0088] S301, Obtain the video to be encoded;

[0089] S302, the key point detection network is used to output the key point coordinates and the corresponding source Jacobian matrix for the source reference frame and driving frame of the video to be encoded.

[0090] S303, read the source Jacobian matrix corresponding to the coordinates of each source key point of the source reference frame, and read the driving Jacobian matrix corresponding to the coordinates of each driving key point of each driving frame.

[0091] S304. Based on the frame sequence of the source reference frame and each driving frame, the driving Jacobian matrix of each driving frame is subtracted from the Jacobian matrix of the source reference frame frame by frame according to the frame sequence to obtain the Jacobian residual matrix of each driving frame, which is used as the driving Jacobian matrix compression result.

[0092] In one embodiment, the frame sequence can be the arrangement sequence number of the source reference frame and each driving frame, determined according to the playback order of the video image frames. The frame sequence can be stored in each frame encoding in the form of numbers, symbols, and strings. Since the overall fluctuation range of the Jacobian matrix corresponding to the same key point position in the source reference frame and each driving frame in the video does not exceed [-1, 1], in order to further compress the Jacobian matrix of each driving frame, the driving Jacobian matrix of each driving frame can be subtracted from the Jacobian matrix of the source reference frame frame by frame according to the frame sequence of the source reference frame and each driving frame to obtain the Jacobian residual matrix of each driving frame. The Jacobian residual matrix is ​​then used as the compression result of the driving Jacobian matrix. The Float16 type Jacobian matrix can be compressed into a Uint8 type Jacobian residual matrix. At this time, the size of each driving frame is 60 bytes, and the compression is about 1 / 8.28 of the source video.

[0093] S305, based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, a bitstream data is generated and transmitted.

[0094] The technical solution provided in this application embodiment, by subtracting the Jacobian matrix of each driving frame from the Jacobian matrix of the source reference frame frame by frame according to the frame sequence of the source reference frame, obtains the Jacobian residual matrix of each driving frame, which is used as the compression result of the driving Jacobian matrix. This can achieve the effect of further compressing the Jacobian matrix of each driving frame, thereby improving the degree of video compression and the effect of video encoding.

[0095] Figure 5 A flowchart of key point coordinate compression provided in an embodiment of this application is shown below. Figure 5 As shown, the specific steps include the following:

[0096] S401, Read the coordinates of each source key point in the source reference frame, and read the coordinates of each drive key point in each drive frame.

[0097] In one embodiment, the source keypoint coordinates can be the coordinates of each keypoint in the source reference frame; the driving keypoint coordinates can be the coordinates of each keypoint in each driving frame. The coordinates of each keypoint can be obtained through the sparse motion field in the keypoint detection network.

[0098] S402, using the second preset compression rule, the source key point coordinates and the driving key point coordinates are format-converted from Float32 type data to Uint8 type data, to obtain the compression result of the source key point coordinates and the compression result of each driving key point coordinate.

[0099] In one embodiment, the second preset compression rule can be a rule for compressing the source keypoint coordinates and the driving keypoint coordinates, such as encoding redundancy of keypoint coordinates within or between frames, without further limitation. Since keypoint coordinates are normalized to [-1, 1], they can be mapped from [-1, 1] to [0, 255] to achieve the purpose of compressing keypoint coordinates. Specifically, the second preset compression rule can be used to convert the format of the source keypoint coordinates and the driving keypoint coordinates, so that the data type of the keypoint coordinates is converted from Float32 to Uint8, thereby obtaining the compression result of the source keypoint coordinates and the compression result of the driving keypoint coordinates. Furthermore, keypoint data can be compressed to the maximum extent by reusing keypoint data. Specifically, keypoints can be sampled at a fixed interval of 4 frames. In this case, the compressed video is about 1 / 27.67 of the original video, achieving a greater degree of compression and improving the video encoding effect.

[0100] The technical solution provided in this application, by reading the coordinates of each source keypoint in the source reference frame and the coordinates of each driving keypoint in each driving frame, and using a second preset compression rule to convert the source keypoint coordinates and the driving keypoint coordinates from Float32 data to Uint8 data, obtains the compression results of the source keypoint coordinates and the compression results of each driving keypoint coordinate. This can increase the degree of video compression, reduce the compression bitrate, improve the video encoding effect, and at the same time reduce the difficulty of developing and training the video encoding model.

[0101] Figure 6 A flowchart of a reconstructed image evaluation method provided in this application embodiment is shown below. Figure 6 As shown, the specific steps include the following:

[0102] S501, the source reference frame and each of the driving frames are input into the generator to generate key point information of the source reference frame and key point information of the driving frame through the sparse motion field network of the generator, and intermediate products are output through the dense motion field network of the generator, and the reconstructed image of the current frame is generated through the image generation unit.

[0103] In one embodiment, the generator can be a model that maps a noisy signal (typically a random number) to a sample similar to real data. Specifically, it can be a model that outputs an image, such as a fully connected neural network or a deconvolutional network. The generator is used to generate a reconstructed image of the current frame, including a sparse motion field network and a dense motion field network. The sparse motion field network operates on the source reference frame image and each driving frame image to obtain keypoint information for each frame; the dense motion field network operates on the keypoint information received by the generator and the source reference frame to obtain an intermediate product. The intermediate product can be the change in keypoint information between the source reference frame and the corresponding keypoint information of the driving frame, or other data, used to guide the generator in generating the reconstructed image of the current frame. The image generation unit can be a unit used to generate the reconstructed image of the current frame, having functions such as data receiving, data processing, and image processing. The reconstructed image of the current frame can be an image generated by the image generation unit based on the source reference frame image or the driving frame image generated in the previous frame and the keypoint information of the current frame. The source reference frame and each of the driving frames are input into the generator to generate key point information of the source reference frame and key point information of the driving frames through the sparse motion field network of the generator, and intermediate products are output through the dense motion field network of the generator, and the reconstructed image of the current frame is generated through the image generation unit.

[0104] In one embodiment, optionally, the intermediate product includes: a motion optical flow field and an occlusion map;

[0105] Accordingly, the reconstructed image of the current frame is generated by the image generation unit, including:

[0106] The motion optical flow field, occlusion map, and source image are input into the image generation unit to output the reconstructed image of the current frame.

[0107] In one embodiment, the motion optical flow field can be the motion of an object between consecutive frames, caused by the relative motion between the object and the camera. The motion optical flow field includes keypoint information of the observed object in the video and the motion trend of these keypoints, used to calculate pixel motion information of each keypoint between adjacent frames based on the temporal changes of pixels in the source reference frame and each driving frame image sequence, as well as the correlation between adjacent frames. The occlusion map is used to indicate to the image generation unit which parts can be obtained from pixel displacements of the source image and which parts need to be obtained through context padding when generating the reconstructed image of the current frame. The motion optical flow field, occlusion map, and source reference frame image are input to the image generation unit of the generator to output the reconstructed image of the current frame.

[0108] In one embodiment, by inputting the motion optical flow field, occlusion map, and source image to the image generation unit to output the reconstructed image of the current frame, the efficiency and accuracy of regenerating the driving frame image can be improved.

[0109] S502, the reconstructed image is input to the discriminator to obtain the evaluation result of the reconstructed image; wherein, the generator and the discriminator constitute an adversarial neural network.

[0110] In one embodiment, the discriminator can be used to evaluate the generation quality of the reconstructed image of the current frame generated by the generator, and can be determined by the similarity between the reconstructed image of the current frame and the actual current frame image. The input of the discriminator is the reconstructed image of the current frame, and the output is the true / false label of the image, i.e., the similarity. Therefore, by inputting the reconstructed image into the discriminator, the discriminator's evaluation result of the reconstructed image can be obtained. An adversarial neural network is a generative model consisting of a generator and a discriminator. They compete and challenge each other, generating high-quality data through this adversarial process. During the training of the adversarial neural network, through continuous training and updating of the discriminator and the generator, the discriminator network gradually becomes more accurate, and the data ultimately generated by the generator becomes closer to real data.

[0111] The technical solution provided in this application generates key point information of the source reference frame and the key point information of the driving frame through the sparse motion field network of the generator, outputs intermediate products through the dense motion field network of the generator, and generates a reconstructed image of the current frame through the image generation unit. The reconstructed image is input to the discriminator to obtain the evaluation result of the reconstructed image, which can improve the accuracy of the current frame image regeneration and thus improve the reliability of the coded video restoration.

[0112] Figure 7This is a structural block diagram of an AI-based video encoding device provided in an embodiment of this application. The device is used to execute the AI-based video encoding method provided in the above embodiments, and has corresponding functional modules and beneficial effects for executing the method. Figure 7 As shown, the device specifically includes: a video acquisition module 601, a key point information output module 602, a key point information compression module 603, and a bitstream data transmission module 604, wherein...

[0113] Video acquisition module 601 is used to acquire the video to be encoded;

[0114] The key point information output module 602 is used to output key point information from the source reference frame and the driving frame of the video to be encoded using a key point detection network.

[0115] The key point information compression module 603 is used to determine the key point information compression result of each driving frame based on the key point information of the source reference frame and the preset compression rules.

[0116] The bitstream data transmission module 604 is used to generate bitstream data for transmission based on the source reference frame, the key point information of the source reference frame, and the compression result of the key point information of each driving frame.

[0117] In one possible embodiment, the key point information includes: key point coordinates and the corresponding source Jacobian matrix;

[0118] Accordingly, the key point information compression module 603 includes:

[0119] The Jacobian matrix reading unit is used to read the source Jacobian matrix corresponding to the coordinates of each source key point in the source reference frame, and to read the driving Jacobian matrix corresponding to the coordinates of each driving key point in each driving frame.

[0120] The Jacobian matrix compression unit is used to compress the driving Jacobian matrix of each driving frame according to the source Jacobian matrix and the first preset compression rule to obtain the driving Jacobian matrix compression result.

[0121] In one possible embodiment, the Jacobian matrix compression unit is specifically used for:

[0122] Based on the frame sequence of the source reference frame and each driving frame, the driving Jacobian matrix of each driving frame is subtracted from the Jacobian matrix of the source reference frame frame by frame according to the frame sequence to obtain the Jacobian residual matrix of each driving frame, which is used as the driving Jacobian matrix compression result.

[0123] In one possible embodiment, the Jacobian matrix compression unit is further used for:

[0124] The source Jacobian matrix and the driving Jacobian matrix are converted from Float32 data to Float16 data.

[0125] In one possible embodiment, the key point information compression module 603 further includes:

[0126] The key point coordinate reading unit is used to read the coordinates of each source key point in the source reference frame, and to read the coordinates of each driving key point in each driving frame.

[0127] The key point coordinate compression unit is used to perform format conversion on the source key point coordinates and the driving key point coordinates using a second preset compression rule, converting the data from Float32 type to Uint8 type, to obtain the compression result of the source key point coordinates and the compression result of each driving key point coordinate.

[0128] In one possible embodiment, the video acquisition module 601 includes:

[0129] The reconstructed image generation unit is used to input the source reference frame and each of the driving frames into the generator, so as to generate key point information of the source reference frame and key point information of the driving frames through the sparse motion field network of the generator, output intermediate products through the dense motion field network of the generator, and generate the reconstructed image of the current frame through the image generation unit.

[0130] A reconstructed image evaluation unit is used to input the reconstructed image into a discriminator to obtain an evaluation result of the reconstructed image; wherein the generator and the discriminator constitute an adversarial neural network.

[0131] In one possible embodiment, the intermediate products include: a motional optical flow field and an occlusion map;

[0132] Accordingly, the reconstructed image generation unit is specifically used for:

[0133] The motion optical flow field, occlusion map, and source image are input into the image generation unit to output the reconstructed image of the current frame.

[0134] In one possible embodiment, the code stream data transmission module 604 is further configured to:

[0135] The transcoding engine is used to identify whether there are terminals that cannot recognize the bitstream data.

[0136] If it exists, the bitstream data is converted based on the target terminal parameters, and the converted bitstream data is transmitted to the target terminal;

[0137] If it does not exist, the bitstream data will be directly transmitted to the corresponding terminal.

[0138] The technical solution provided in this application includes a video acquisition module for acquiring a video to be encoded; a key point information output module for outputting key point information from the source reference frame and driving frame of the video to be encoded using a key point detection network; a key point information compression module for determining the key point information compression result of each driving frame based on the key point information of the source reference frame and a preset compression rule; and a bitstream data transmission module for generating bitstream data for transmission based on the source reference frame, the key point information of the source reference frame, and the key point information compression result of each driving frame. The aforementioned AI-based video encoding device solves the problems of poor model adaptability, high training difficulty, and high transmission bitrate when using existing AI video encoding technologies to encode videos. By using a keypoint detection network to output keypoint information for the source reference frame and driving frame of the video to be encoded, and determining the compression result of keypoint information for each driving frame based on the keypoint information of the source reference frame and preset compression rules, a bitstream data is generated and transmitted based on the source reference frame, the keypoint information of the source reference frame, and the compression result of keypoint information for each driving frame. This achieves the effect of ultra-low bitrate video encoding. By performing deep compression of video data and combining it with an AI model to implement the data encoding process, better video compression effect can be obtained. At the same time, this solution also improves the applicability of the AI ​​video encoding model and reduces the difficulty and cost of model development and training.

[0139] Figure 8 A schematic diagram of the structure of an AI-based video encoding device provided in this application embodiment is shown below. Figure 8 As shown, the device includes a processor 701, a memory 702, an input device 703, and an output device 704; the number of processors 701 in the device can be one or more. Figure 7 Taking a processor 701 as an example; the processor 701, memory 702, input device 703, and output device 704 in the device can be connected via a bus or other means. Figure 7 Taking a bus connection as an example, the memory 702, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the AI-based video encoding method in this embodiment. The processor 701 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 702, thereby implementing the aforementioned AI-based video encoding method. The input device 703 can be used to receive input digital or character information and generate key signal inputs related to user settings and function control of the device. The output device 704 may include a display screen or other display device.

[0140] This application embodiment also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform an AI-based video coding method described in the above embodiments, wherein the method includes:

[0141] Obtain the video to be encoded;

[0142] The source reference frame and driving frame of the video to be encoded are used to output key point information using a key point detection network;

[0143] Based on the key point information of the source reference frame and the preset compression rules, determine the compression result of the key point information of each driving frame;

[0144] Based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, a bitstream data is generated and transmitted.

[0145] It is worth noting that in the above embodiments of the AI-based video encoding device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0146] In some possible implementations, various aspects of the methods provided in this application can also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps of the methods according to the various exemplary embodiments of this application described above. For example, the computer device can execute the AI-based video encoding method described in the embodiments of this application. The program product can be implemented using any combination of one or more readable media.

Claims

1. An AI-based video coding method, characterized in that, include: Obtain the video to be encoded; The source reference frame and driving frame of the video to be encoded are used to output key point information by a key point detection network. The key point information includes: key point coordinates and the corresponding Jacobian matrix. Based on the key point information of the source reference frame and the preset compression rules, the compression result of the key point information of each driving frame is determined, including: reading the source Jacobian matrix corresponding to the coordinates of each source key point of the source reference frame, and reading the driving Jacobian matrix corresponding to the coordinates of each driving key point of each driving frame. The residual is calculated based on the source Jacobian matrix and the driving Jacobian matrix to compress the driving Jacobian matrix of each driving frame, and the calculated residual matrix is ​​used as the compression result of the driving Jacobian matrix. Based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, a bitstream data is generated and transmitted.

2. The AI-based video coding method according to claim 1, characterized in that, Residual calculations are performed based on the source Jacobian matrix and the driving Jacobian matrix to compress the driving Jacobian matrix of each driving frame. The calculated residual matrix is ​​used as the compression result of the driving Jacobian matrix, including: Based on the frame sequence of the source reference frame and each driving frame, the driving Jacobian matrix of each driving frame is subtracted from the source Jacobian matrix frame by frame according to the frame sequence to obtain the Jacobian residual matrix of each driving frame, which is used as the driving Jacobian matrix compression result.

3. The AI-based video coding method according to claim 1, characterized in that, Before performing residual calculation based on the source Jacobian matrix and the driving Jacobian matrix, the method further includes: The source Jacobian matrix and the driving Jacobian matrix are converted from Float32 data to Float16 data.

4. The AI-based video coding method according to claim 1, characterized in that, Based on the key point information of the source reference frame and the preset compression rules, the compression results of the key point information of each driving frame are determined, including: Read the coordinates of each source key point in the source reference frame, and read the coordinates of each drive key point in each drive frame; Using a second preset compression rule, the source keypoint coordinates and the driving keypoint coordinates are format-converted from Float32 type data to Uint8 type data, resulting in the compression results of the source keypoint coordinates and the compression results of each driving keypoint coordinate.

5. The AI-based video coding method according to claim 1, characterized in that, After acquiring the video to be encoded, the method further includes: The source reference frame and each of the driving frames are input into the generator to generate key point information of the source reference frame and key point information of the driving frame through the sparse motion field network of the generator, and intermediate products are output through the dense motion field network of the generator, and the reconstructed image of the current frame is generated through the image generation unit. The reconstructed image is input into a discriminator to obtain an evaluation result of the reconstructed image; wherein the generator and the discriminator constitute an adversarial neural network.

6. The AI-based video coding method according to claim 5, characterized in that, The intermediate products include: a moving optical flow field and an occlusion map; Accordingly, the reconstructed image of the current frame is generated by the image generation unit, including: The motion optical flow field, occlusion map, and source image are input into the image generation unit to output the reconstructed image of the current frame.

7. The AI-based video coding method according to claim 1, characterized in that, After transmitting the bitstream data generated based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame, the method further includes: The transcoding engine is used to identify whether there are terminals that cannot recognize the bitstream data. If it exists, the bitstream data is converted based on the target terminal parameters, and the converted bitstream data is transmitted to the target terminal; If it does not exist, the bitstream data will be directly transmitted to the corresponding terminal.

8. An AI-based video encoding device, characterized in that, include: The video acquisition module is used to acquire the video to be encoded. The key point information output module is used to output key point information from the source reference frame and driving frame of the video to be encoded using a key point detection network. The key point information includes: key point coordinates and the corresponding Jacobian matrix. The key point information compression module is used to determine the key point information compression result of each driving frame based on the key point information of the source reference frame and the preset compression rules. Specifically, the key point information compression module is used to: read the source Jacobian matrix corresponding to the coordinates of each source key point of the source reference frame, and read the driving Jacobian matrix corresponding to the coordinates of each driving key point of each driving frame; perform residual calculation based on the source Jacobian matrix and the driving Jacobian matrix to compress the driving Jacobian matrix of each driving frame; and use the calculated residual matrix as the driving Jacobian matrix compression result. The bitstream data transmission module is used to generate bitstream data for transmission based on the source reference frame, the key point information of the source reference frame, and the compression results of the key point information of each driving frame.

9. An AI-based video encoding device, the video encoding device comprising: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the AI-based video coding method according to any one of claims 1-7.

10. A storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the AI-based video coding method according to any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the AI-based video coding method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video encoding and decoding method and device, equipment and storage medium

    CN114257818A

  • Image super-resolution reconstruction method

    CN115797183A