Video decoding method, video encoding method, device, computer program and equipment

The video encoding method optimizes encoding by determining target sampling parameters based on application and content characteristics, addressing inefficiencies in conventional methods to improve encoding efficiency and quality.

JP7820027B2Active Publication Date: 2026-02-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024563130
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-07-07
Publication Date
2026-02-25
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

Conventional video encoding methods result in large amounts of redundant information due to inefficient encoding, leading to low video data encoding efficiency, especially under limited bandwidth conditions.

Method used

A video encoding method that determines target sampling parameters based on media application scenarios and video content characteristics, performing sampling and encoding processes to optimize video data for efficient encoding and decoding, reducing data volume while maintaining quality.

Benefits of technology

The method enhances encoding efficiency by adaptively sampling video data, reducing data volume without compromising the quality of the restored video, suitable for various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007820027000005
    Figure 0007820027000005
  • Figure 0007820027000006
    Figure 0007820027000006
  • Figure 0007820027000007
    Figure 0007820027000007
Patent Text Reader

Abstract

A video decoding method executed by a computer device, comprising the steps of: acquiring a media application scenario and video content characteristics of original video data to be encoded (S101); determining target sampling parameters for performing a sampling process on the original video data based on the media application scenario and the video content characteristics (S102); performing a sampling process on the original video data based on the target sampling parameters to obtain sampled video data (S103); and encoding the sampled video data to obtain video encoding data corresponding to the original video data (S104).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to a Chinese patent application filed with the China Patent Office on October 11, 2022, bearing application number 2022112635838 and entitled "Video decoding method, video encoding method, device, storage medium and apparatus," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the technical field of data processing, and in particular to a video decoding method, a video encoding method, an apparatus, a storage medium, an appliance and a computer program product. [Background technology]

[0003] With the development of digital media technology and computer technology, video has been applied to various fields, such as mobile communication, network awareness, network television, etc., bringing great convenience to people's entertainment and life. Under the condition of limited bandwidth, if a video frame is encoded by a conventional encoder, there may be a large amount of redundant information in the video code stream, resulting in low video data encoding efficiency. Summary of the Invention [Problem to be solved by the invention]

[0004] According to the embodiments provided in the present application, a video decoding method, a video encoding method, an apparatus, a storage medium, a device, and a computer program product are provided. [Means for solving the problem]

[0005] One aspect of an embodiment of the present application provides a video encoding method executed by a computer device, comprising: obtaining a media application scenario and video content characteristics of original video data to be encoded; determining target sampling parameters for performing sampling processing on the original video data based on a media application scenario and video content characteristics; performing a sampling process on the original video data based on the target sampling parameters to obtain sampled video data; encoding the sampled video data to obtain video encoded data corresponding to the original video data.

[0006] One aspect of an embodiment of the present application provides a video decoding method executed by a computer device, comprising: obtaining video encoding data to be decoded and target sampling parameters corresponding to the video encoding data, the video encoding data being obtained by encoding sampled video data, the sampled video data being obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters, and the target sampling parameters being determined based on a media application scenario and video content features of the original video data; decoding the video encoding data to obtain sampled video data; performing a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data.

[0007] One aspect of an embodiment of the present application provides a video decoding device, a first acquisition module for acquiring video encoding data to be decoded and target sampling parameters corresponding to the video encoding data, the video encoding data being obtained by encoding sampled video data, and the sampled video data being obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters, the target sampling parameters being determined based on a media application scenario and video content features of the original video data; a decoding module for decoding the video encoding data to obtain sampled video data; and a sampling restoration module for performing a sampling restoration process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data.

[0008] One aspect of an embodiment of the present application provides a video encoding device, a second acquiring module for acquiring media application scenarios and video content features of original video data to be encoded; a second determining module for determining target sampling parameters for performing sampling processing on the original video data based on a media application scenario and video content characteristics; a sampling processing module for performing sampling processing on the original video data based on the target sampling parameters to obtain sampled video data; an encoding module for encoding the sampled video data to obtain video encoding data corresponding to the original video data.

[0009] One aspect of an embodiment of the present application provides a computer device, comprising a processor and a memory; The processor and the memory are coupled together, and the memory stores computer-readable instructions that, when executed by the processor, cause the computing device to perform the methods provided by the embodiments of the present application.

[0010] One aspect of the embodiments of the present application provides a computer-readable storage medium having computer-readable instructions stored therein, the computer-readable instructions being read and executed by a processor to cause a computing device including the processor to perform a method provided by the embodiments of the present application.

[0011] One aspect of the embodiments of the present application provides a computer program product, the computer program product including computer-readable instructions, the computer-readable instructions being stored in a computer-readable storage medium, a processor of a computing device reading the computer-readable instructions from the computer-readable storage medium and executing the computer-readable instructions, thereby causing the computing device to perform the method provided by the embodiments of the present application.

[0012] The details of one or more embodiments of the present application are set forth in the following drawings and description, from which other features, objects, and advantages of the present application will become apparent. [Brief explanation of the drawings]

[0013] In order to more clearly explain the technical solutions of the embodiments of the present invention or the prior art, the following briefly introduces the drawings necessary for the description of the embodiments or the prior art. The drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings on the premise that they do not perform any work worthy of inventive step.

[0014] [Figure 1] 1 is a schematic diagram of a video data processing process provided by an embodiment of the present application; [Figure 2] 1 is a schematic diagram of an encoding unit provided by an embodiment of the present application; [Figure 3]1 is a flow diagram of a video encoding method provided by an embodiment of the present application; [Figure 4] 1 is a schematic diagram of a time sampling configuration provided by an embodiment of the present application; [Figure 5] 1 is a schematic diagram of a spatial sampling configuration provided by an embodiment of the present application. [Figure 6] 1 is a schematic diagram of a video decoding method provided by an embodiment of the present application; [Figure 7] 1 is a structural schematic diagram of a video decoding device provided by an embodiment of the present application; [Figure 8] 1 is a structural schematic diagram of a video encoding device provided by an embodiment of the present application; [Figure 9] 1 is a structural schematic diagram of a computer device provided by an embodiment of the present application; [Figure 10] 1 is a structural schematic diagram of a computer device provided by an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0015] The following clearly and completely describes the technical solutions of the embodiments of the present invention, combined with the drawings of the embodiments of the present invention, and the described embodiments are not all embodiments but only some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without any inventive effort are within the scope of protection of the present invention.

[0016] This application relates to the cloud technology field. This application relates to cloud computing in the cloud technology field. Cloud computing is a computing mode in which computing tasks are distributed across a resource pool consisting of a large number of computers, allowing various application systems to obtain computing power, storage space, and information services according to their needs. The network that provides these resources is called a "cloud." For users, resources in the "cloud" are infinitely scalable and available at any time, and can be used and expanded according to needs. This application relates to the encoding and decoding of video data using cloud computing.

[0017] The embodiments of this application relate to video data processing technology, and the complete video data processing process specifically includes video acquisition, video encoding, video file packaging, video transmission, video file unpackaging, video decoding, and finally video display. Note that video acquisition converts analog video into digital video and stores it according to a digital video file format, that is, video acquisition converts video signals into binary digital information. The binary information converted from the video signal is a binary data stream, which is also called the code stream or bitstream of the video signal. Video encoding is the use of compression technology to convert a file in an original video format into a file in another video format. The generation of video media content referred to in the embodiments of the present application includes real scenarios collected and generated by cameras, and screen content scenarios generated by computers. From the perspective of video signal acquisition form, video signals can be divided into two forms: camera-captured and computer-generated. Due to different statistical characteristics, the corresponding compression encoding forms may also be different. For the current mainstream video encoding technology, the international video encoding standard HEVC (High Efficiency Video Coding, international video encoding standard HEVC / H.265), VVC (versatile video coding, international video encoding standard VVC / H.266), and video encoding standard AVS (Audio Video Coding Standard, video encoding standard AVS) or AVS3 (the third generation video encoding standard developed by the AVS working group) are taken as examples. Using a mixed encoding framework, the following series of operations and processes are performed on the input original video signal, as shown in Figure 1. Figure 1 is a schematic diagram of the video data processing process provided by the embodiments of the present application. Specifically, refer to Figure 1:

[0018] (1) Block partition structure: An input image (e.g., a video frame in video data) is divided into several non-overlapping processing units according to a size, and each processing unit performs a similar compression operation. These processing units are called CTUs (coding tree units) or LCUs (largest coding units). Generally, coding tree units are divided downward from the largest coding unit. Continuing downward from the CTU, more detailed divisions are obtained, resulting in one or more basic coding units, called CUs (coding units). Each CU is the most basic element in an encoding process. The following describes various coding forms available for each CU, as shown in FIG. 2, which is a schematic diagram of a coding unit provided by an embodiment of the present application. For the relationship between LCUs (or CTUs) and CUs, please refer to FIG. 2.

[0019] (2) Predictive Coding: Including forms such as intraframe prediction and interframe prediction, the original video signal is predicted by a selected reconstructed video signal to obtain a residual video signal. The encoding side needs to select the most appropriate one from multiple possible predictive coding modes for the current CU and notify the decoding side. a. Intra (picture) Prediction: The predicted signal comes from a coded and reconstructed region within the same picture. b. Inter (picture) Prediction: The signal to be predicted comes from another picture (called a reference picture) that is different from the current picture being coded.

[0020] (3) Transform & Quantization: The residual video signal undergoes transformation operations such as DFT (Discrete Fourier Transform) and DCT (Discrete Cosine Transform, a subset of DFT) to transform the signal into a transform domain, which is called a transform coefficient. The signal in the transform domain is further subjected to an irreversible quantization operation to discard certain information, making the quantized signal suitable for compressed representation.

[0021] In some video coding standards, one or more transform forms can be selected, and therefore the encoding side must also select one of the transforms for the currently coded CU and notify the decoding side. Generally, the degree of quantization precision is determined by a quantization parameter (QP), and a larger QP value generally results in larger distortion and a lower bit rate because coefficients in a larger value range are quantized and output identically. Conversely, a smaller QP value generally results in smaller distortion and a higher bit rate because coefficients in a smaller value range are quantized and output identically.

[0022] (4) Entropy or statistical coding: The quantized transform domain signal undergoes statistical compression coding based on the frequency of occurrence of each value, ultimately outputting a binary (0 or 1) compressed code stream. In addition, other coding information, such as selection mode and motion vectors, also needs to be entropy coded to reduce the bit rate.

[0023] Statistical coding is a form of lossless coding that can effectively reduce the bit rate for representing the same signal. Common forms of statistical coding include Variable Length Coding (VLC) and Content Adaptive Binary Arithmetic Coding (CABAC).

[0024] (5) Loop Filtering: A reconstructed decoded image is obtained by performing inverse quantization, inverse transform, and predictive compensation (the inverse operations of (2) to (4) above) on the coded image. Because the reconstructed image is affected by quantization compared to the original image, some information differs from the original image, resulting in distortion. Filtering operations, such as deblocking, SAO (Sample Adaptive Offset), or ALF (Adaptive Loop Filter), can be performed on the reconstructed image to effectively reduce the degree of quantization-induced distortion. These filtered reconstructed images are used as references for subsequent coded images to predict subsequent signals, so the above filtering operations are also called loop filtering or in-coding loop filtering operations.

[0025] Figure 1 shows the basic flow of the video encoder, and the kth CU (S k [×, y]) is taken as an example, where k is a positive integer greater than or equal to 1 and less than or equal to the number of CUs in the current input image, and S k [×, y] denotes the pixel point in the kth CU whose coordinate is [×, y], where × denotes the horizontal coordinate of the pixel point, y denotes the middle coordinate of the pixel point, and S k A prediction signal is generated by performing a suitable process such as motion compensation or intra-frame prediction on [x, y].

number

number

number

[0026] After the video data is encoded, the encoded data stream must be packaged and transmitted to the user. A video file package refers to storing the encoded and compressed video and audio in a single file in a specific format according to a package format (or container, or file container). Common package formats include AVI (Audio Video Interleaved) and ISOBMFF (ISO Based Media File Format, an ISO (International Standard Organization) based media file format). ISOBMFF is a media file packaging standard, and the most typical ISOBMFF file is the MP4 (Moving Picture Experts Group 4) file. The packaged file is transmitted via video to a decoding device, which then performs the reverse operations of unpackaging and decoding, allowing the final video content to be displayed on the decoding device.

[0027] Here, the file unpackaging process by the decoding device is the opposite of the file packaging process described above. The decoding device unpacks the packaged file according to the file format requirements of the packaging process to obtain a video code stream. The decoding process by the decoding device is also the opposite of the encoding process. The decoding device decodes the video code stream to restore video data. As can be seen from the encoding process described above, on the decoding side, for each CU, the decoder first obtains the compressed code stream, then performs entropy decoding to obtain various mode information and quantized transform coefficients. Each coefficient is then inversely quantized and inversely transformed to obtain a residual video signal. Furthermore, a prediction signal corresponding to the CU is obtained based on the known coding mode information, and the two are added to obtain a reconstructed signal. Finally, loop filtering is performed on the reconstructed value of the decoded image to generate a final output signal.

[0028] As shown in Figure 3, Figure 3 is a flow diagram of a video encoding method provided in an embodiment of the present application, where the method is performed by a computer device, and the computer device may be an encoding device. As shown in Figure 3, the method specifically includes, but is not limited to, the following steps:

[0029] S101: Obtain a media application scenario and video content characteristics of original video data to be encoded.

[0030] Specifically, after obtaining original video data to be encoded, the encoding device obtains media application scenarios of the original video data, including user viewing scenarios, machine recognition scenarios, etc. The user viewing scenario is a scenario in which a target user views video data, and the machine recognition scenario is a scenario in which a machine interprets video data to complete related tasks (e.g., detection, recognition, etc.). The video perception characteristics of the target object for video data in different media application scenarios are different. For example, the video perception characteristics of the target user for video data in a user viewing scenario are different from the video perception characteristics of the target machine for video data in a machine recognition scenario. Therefore, the requirements for video data quality and resolution in a user viewing scenario are different from the requirements for video data quality and resolution in a machine recognition scenario. Thus, different encoding methods are used for different media application scenarios to meet the needs of the corresponding scenarios. The encoding device also obtains video content characteristics of the original video data, including the video content change rate in the original video data, the amount of video content information, the video resolution of video frames in the original video data, and the number of video frames played per unit time in the original video data.

[0031] S102: Determine target sampling parameters for performing sampling processing on the original video data based on a media application scenario and video content characteristics.

[0032] Specifically, the media application scenario reflects the required video data quality requirements of the target object (e.g., content change rate requirements and resolution requirements, etc.), and the video content characteristics of the original video data reflect the video content change rate and video content information amount of the original video data. The encoding device determines target sampling parameters for sampling the original video data based on the media application scenario and the video content characteristics. The target sampling parameters include a target sampling type and a target sampling rate for the target sampling type. Specifically, the target sampling type includes a temporal sampling type and a spatial sampling type. The temporal sampling type performs video frame sampling on the video data, and the spatial sampling type performs sampling according to the video resolution of the video data. The target sampling rate in the target sampling mode includes the target sampling rate in the time sampling mode and the target sampling rate in the spatial sampling mode. For example, the target sampling rate in the time sampling mode may be frame extraction with a factor of 2 (i.e., sampling one frame at an interval of one frame), frame extraction with a factor of 3 (i.e., sampling one frame at an interval of two frames), etc. The target sampling rate in the spatial sampling mode may be any value greater than 0, such as 0.5 factor (i.e., reducing the resolution by 0.5 times), 0.75 factor (i.e., reducing the resolution by 0.75 times), 2 factor (i.e., increasing the resolution by 2 times), etc.

[0033] Preferably, a specific form in which an encoding device determines target sampling parameters for performing sampling processing on original video data includes the steps of: determining a target sampling form for performing sampling processing on the original video data based on video content features; determining video sensing features for video data of a target object in a media application scenario, where the target object is an object for performing sensing processing on the original video data; determining a target sampling rate in the target sampling form based on the video sensing features and video content features; and determining the target sampling rate and target sampling form as target sampling parameters for performing sampling processing on the original video data.

[0034] Specifically, the encoding device determines a target sampling format for performing sampling processing on the original video data based on video content characteristics. In this way, the target sampling format of the original video data can be self-adaptively determined to improve the accuracy of sampling the original video data. The encoding device determines video sensing characteristics for video data of a target object in a media application scenario of the original video data, where the target object is an object for performing sensing processing on the original video data, and the video sensing characteristics reflect information such as the target object's quality requirements and resolution requirements for the video data. Furthermore, the encoding device determines a target sampling rate for the target sampling format based on the video sensing characteristics and video content characteristics, and determines the target sampling rate and target sampling format as target sampling parameters for performing sampling processing on the original video data. In this way, the target sampling format and the target sampling rate for the target sampling format can be self-adaptively determined based on the media application scenario and video content characteristics to improve the sampling accuracy of the original video data. This ensures that when a decoding device restores the original video data based on the video encoding data, it does not affect application (e.g., user viewing or machine recognition), and can reduce the data amount of video encoding data obtained by encoding the original video data. In other words, after sampling the original video data using the target sampling parameters, the decoding device can restore the viewing quality of the original video data based on the video encoding data while reducing the data amount of the video encoding data.

[0035] Here, the target sampling type includes one of a temporal sampling type, a spatial sampling type, a temporal sampling type, and a spatial sampling type, where the temporal sampling type performs frame extraction sampling on the original video data, and the spatial sampling type performs video resolution sampling on the original video data. The target sampling rate in the temporal sampling type is the ratio of the number of extracted video frames to the number of original video frames when frame extraction sampling is performed on the original video data, and the target sampling rate in the spatial sampling type is the ratio of the video resolution resulting from sampling to the original video resolution when video resolution sampling is performed on the original video data.

[0036] Preferably, a specific manner in which the encoding device determines the target sampling form includes a step of determining a duplication rate of video content in the original video data based on a video content change rate included in the video content characteristics, and a step of determining a target sampling form for performing sampling processing on the original video data based on the duplication rate of video content in the original video data.

[0037] Specifically, the video content characteristics include the video content change rate of the original video data (i.e., the change rate of the screen content in the video), which may be the moving speed of a moving object in the video content or the change rate of pixels in the video content. The encoding device determines the video content duplication rate of the original video data based on the video content change rate of the original video data, which is the duplication rate between two adjacent video frames in an arbitrary playback order in the original video data. Furthermore, the encoding device determines a target sampling form for performing sampling on the original video data based on the video content duplication rate of the original video data. For example, if the video content duplication rate of the original video data is too low, it is not suitable to perform frame extraction sampling on the original video data, and performing frame extraction sampling on the original video data will affect the display effect of the original video data restored by the decoding device based on the video encoding data (e.g., problems such as discontinuity of the video content or large jumps in the video content may occur). In this way, by determining the target sampling pattern based on the duplication rate of the video content in the original video data, the accuracy of the target sampling pattern is improved, and the display effect of the original video data restored by the decoding device based on the video encoding data is not affected, while the data volume of the video encoding data obtained by encoding the original video data is reduced.

[0038] Preferably, a specific manner in which the encoding device determines the target sampling format based on the overlap rate includes: if the overlap rate of the video content in the original video data is greater than a first overlap rate threshold, determining the temporal sampling format and the spatial sampling format as the target sampling format for performing sampling on the original video data; if the overlap rate of the video content in the original video data is equal to or less than the first overlap rate threshold and greater than a second overlap rate threshold, determining the temporal sampling format as the target sampling format for performing sampling on the original video data, where the second overlap rate threshold is smaller than the first overlap rate threshold; and if the overlap rate of the video content in the original video data is equal to or less than the second overlap rate threshold, determining the spatial sampling format as the target sampling format for performing sampling on the original video data. The first overlap rate threshold and the second overlap rate threshold may be set based on the detection needs of the target object or on specific circumstances, and the embodiments of the present application are not limited to the first overlap rate threshold and the second overlap rate threshold.

[0039] Specifically, when the encoding device determines that the duplication rate of the video content in the original video data is greater than a first duplication rate threshold, performing temporal sampling and spatial sampling on the original video data will not affect the display effect of the sampled video data, and therefore determines the temporal sampling type and the spatial sampling type as target sampling types for performing sampling processing on the original video data. This significantly reduces the data amount of the video encoded data obtained by encoding the original video data. Specifically, when the encoding device determines that the duplication rate of the video content in the original video data is equal to or less than the first duplication rate threshold and greater than a second duplication rate threshold, it determines the temporal sampling type as the target sampling type for performing sampling processing on the original video data, and the second duplication rate threshold is smaller than the first duplication rate threshold. This reduces the data amount of the video encoded data of the original video data when sampling the original video data using only the temporal sampling type, and avoids a large loss of information in the sampled video data after performing time sampling on the original video data using the temporal sampling type and the spatial sampling type, which would affect the display effect of the original video data restored based on the sampled video data.

[0040] Of course, even if the encoding device determines that the overlap rate of the video content in the original video data is equal to or less than the first overlap rate threshold and greater than the second overlap rate threshold, the encoding device still determines the spatial sampling format as the target sampling format for performing sampling on the original video data. In other words, if the overlap rate of the video content in the original video data is equal to or less than the first overlap rate threshold and greater than the second overlap rate threshold, the encoding device may determine either the temporal sampling format or the spatial sampling format as the target sampling format for performing sampling on the original video data. Specifically, if the encoding device determines that the overlap rate of the video content in the original video data is equal to or less than the second overlap rate threshold, sampling the original video data using the temporal sampling format will result in a large amount of information loss, so the encoding device determines the spatial sampling format as the target sampling format for performing sampling on the original video data. In this way, an appropriate target sampling format is determined based on the overlap rate of the video content in the original video data to improve sampling accuracy. Of course, the manner in which the encoding device determines the target sampling form based on the overlap rate can be applied to video frames in the original video data, for example, the encoding device determines the target sampling form for performing sampling processing on the current video frame based on the overlap rate between the current video frame and a reference video frame (a video frame that is the frame before the current video frame in the playback order, or a video frame that is the frame after the current video frame in the playback order) in the original video data.

[0041] Preferably, a specific manner in which the encoding device determines the target sampling form further includes a step of determining the complexity of the video content in the original video data based on the amount of video content information contained in the video content characteristics, and a step of determining a target sampling form for performing sampling processing on the original video data based on the complexity of the video content in the original video data.

[0042] Preferably, the video content characteristics of the original video data include a video content information amount, which reflects the complexity of the content of the original video data. That is, the higher the video content information amount of the original video data, the more complex the video content, and the lower the video content information amount of the original video data, the simpler the video content. For example, the more scenarios and scenes related to the video content, the more information the video content contains. For example, a video scene of a playground contains multiple human scenarios, so the video content information amount is high. A video scene of a monotonous scene, such as an ocean or lake, contains monotonous elements, so the video content information amount is high. For example, for video A and video B both containing text, the smaller the font size and the more text there is in video A compared to video B, the higher the video content information amount of video A. The encoding device determines the complexity of the video content of the original video data based on the video content information amount included in the video content characteristics of the original video data. Furthermore, the encoding device determines a target sampling form for performing sampling on the original video data based on the complexity of the video content of the original video data. For example, if the video content of the original video data is highly complex, using the spatial sampling format may cause the original video data to lose key information, which may disrupt the video content of the sampled video data due to sampling, making it unsuitable to use the spatial sampling format to sample the original video data. Thus, by determining a target sampling format for performing sampling processing on the original video data based on the complexity of the video content of the original video data, an appropriate target sampling format is determined to improve sampling accuracy.Of course, the manner in which the encoding device determines the target sampling form based on complexity can be applied to video frames in the original video data, for example, determining the target sampling form for performing sampling processing on a current video frame based on the complexity of the current video frame in the original video data.

[0043] Preferably, the encoding device's determining the target sampling format based on complexity includes: determining the temporal sampling format and the spatial sampling format as the target sampling format for performing sampling on the original video data if the complexity of the video content of the original video data is lower than a first complexity threshold; determining the spatial sampling format as the target sampling format for performing sampling on the original video data if the complexity of the video content of the original video data is equal to or greater than the first complexity threshold and lower than a second complexity threshold, where the second complexity threshold is higher than the first complexity threshold; and determining the temporal sampling format as the target sampling format for performing sampling on the original video data if the complexity of the video content of the original video data is higher than the second complexity threshold. The first and second complexity thresholds may be set based on the detection needs of the target object or specific circumstances, and the embodiments of the present application are not limited to the first and second complexity thresholds.

[0044] Specifically, if the encoding device determines that the complexity of the video content of the original video data is less than a first complexity threshold, the encoding device determines that the video content of the original video data is simple, and can then sample the original video data using a temporal sampling format and a spatial sampling format, without affecting the display effect of the original video data restored by the decoding device based on the sampled video data, and significantly reducing the data volume of the resulting video encoded data. If the complexity of the video content of the original video data is equal to or greater than the first complexity threshold and less than a second complexity threshold, the encoding device determines the spatial sampling format as the target sampling format for performing sampling processing on the original video data, and the second complexity threshold is greater than the first complexity threshold. In this way, by sampling the original video data using only the spatial sampling format, the data volume of the video encoded data corresponding to the original video data is reduced, and it is possible to avoid a large amount of information loss of the sampled video data after time sampling the original video data using the temporal sampling format and the spatial sampling format, which would affect the display effect of the original video data restored based on the sampled video data.

[0045] Of course, if the complexity of the video content in the original video data is equal to or greater than the first complexity threshold but less than the second complexity threshold, the temporal sampling format is determined as the target sampling format for performing sampling on the original video data. In other words, if the complexity of the video content in the original video data is equal to or greater than the first complexity threshold but less than the second complexity threshold, one of the temporal sampling format and the spatial sampling format is determined as the target sampling format for performing sampling on the original video data. Furthermore, if the complexity of the video content in the original video data is greater than the second complexity threshold, it means that the video content of the original video data is highly complex. If the spatial sampling format is used, the original video data may lose key information, which may cause confusion in the video content of the sampled video data due to sampling. Therefore, it is not suitable to use the spatial sampling format to sample the original video data, and the temporal sampling format is determined as the target sampling format for performing sampling on the original video data.

[0046] Preferably, the encoding device determines the target sampling format in the following manner: the encoding device determines the duplication rate of the video content in the original video data based on the video content change rate included in the video content characteristics, and determines the complexity of the video content in the original video data based on the video content information amount included in the video content characteristics. Furthermore, the encoding device determines the target sampling format for sampling the original video data based on the duplication rate of the video content in the original video data and the complexity of the video content in the original video data. The encoding device determines whether to sample the original video data using a temporal sampling format based on the duplication rate of the video content in the original video data, and determines whether to sample the original video data using a spatial sampling format based on the complexity of the video content in the original video data. Specifically, if the duplication rate of the video content in the original video data is greater than a third duplication rate threshold, the temporal sampling format is selected as the target sampling format for sampling the original video data; and if the duplication rate of the video content in the original video data is equal to or less than the third duplication rate threshold, the temporal sampling format is prohibited from being selected as the target sampling format for sampling the original video data. If the complexity of the video content in the original video data is less than a third complexity threshold, the spatial sampling form is set as the target sampling form for performing sampling processing on the original video data, and if the overlap rate of the video content in the original video data is equal to or greater than the third complexity threshold, the temporal sampling form is prohibited from being set as the target sampling form for performing sampling processing on the original video data.The third overlap rate threshold may be set based on the detection needs of the target object or based on specific circumstances, and the embodiments of the present application are not limited to the third overlap rate threshold; the third complexity threshold may be set based on the detection needs of the target object or based on specific circumstances, and the embodiments of the present application are not limited to the third complexity threshold.

[0047] Preferably, a specific form in which the encoding device determines a target sampling rate in the target sampling format includes, if the target sampling format is a time sampling format, determining a limiting number of video frames corresponding to video data sensed by the target object within a unit time based on the video sensing characteristics, and determining a target sampling rate in the time sampling format based on a ratio between the limiting number of video frames and the number of playback video frames, where the number of playback video frames is the number of video frames to be played within a unit time in the original video data indicated by the video content characteristics.

[0048] Specifically, when the encoding device determines to sample the original video data using the time sampling format, it determines a limit number of video frames of video data perceived by the target object within a unit time based on the video sensing characteristics. The unit time may be one second, one minute, etc. Here, the number of video frames that a user or machine can sense within a unit time is finite. For example, the user's eye can sense a video frame rate of 55 frames / second. The human eye cannot distinguish between a video with a frame rate exceeding 55 frames / second and a video with a frame rate of 55 frames / second. Only when the frame rate is too low can the human eye perceive the problem of a video with a frame rate that is too low, that the video screen appears unsmooth. Furthermore, since different objects have different video sensing characteristics, the limit number of video frames corresponding to different objects differs. The limit number of video frames of video data perceived by the target object within a unit time is equal to or less than the video frame rate that the target object can sense, and this limit number of video frames is the lowest frame number that meets the sensing needs of the target object. Specifically, the encoding device determines a limit number of video frames of video data sensed by the target object within a unit time based on the video sensing characteristics and the video content characteristics of the original video data, samples the original video data based on the limit number of video frames to obtain sampled video data, and then restores original video data based on the sampled video data, the video quality and resolution of which meet the sensing needs of the target object.

[0049] Furthermore, the number of video frames to be played is the number of video frames to be played within a unit time in the original video data indicated by the video content features, and the encoding device determines the target sampling rate in the time sampling mode based on the ratio between the limited number of video frames and the number of video frames to be played within a unit time in the original video data indicated by the video content features. Specifically, since the sampling rate of video frame sampling must be a positive integer, the encoding device obtains the ratio between the limited number of video frames and the number of video frames to be played within a unit time in the original video data indicated by the video content features, and if the ratio is a positive integer, determines the ratio as the target sampling rate in the time sampling mode. If the ratio is not a positive integer, the encoding device rounds the ratio to obtain a rounded ratio, and determines the rounded ratio as the target sampling rate in the time sampling mode. In this way, since the limit number of video frames that different target objects can detect is different, the target sampling rate in the time sampling form is self-adaptively determined based on the limit number of video frames corresponding to the target object, so that when the original video data is restored based on the sampled video data obtained by sampling, the quality and resolution of the restored original video data can meet the detection needs of the target object.

[0050] Preferably, a specific form in which the encoding device determines a target sampling rate in the target sampling format based on video sensing characteristics and video content characteristics includes, if the target sampling format is a spatial sampling format, determining a limiting video resolution associated with the target object based on the video sensing characteristics, and determining a ratio between the limiting video resolution and the video frame resolution as the target sampling rate in the spatial sampling format, wherein the video frame resolution is the video resolution of the video frame in the original video data indicated by the video content characteristics.

[0051] Specifically, if the target sampling format is a spatial sampling format, the encoding device determines a limiting video resolution associated with the target object based on the video sensing characteristics. Different target objects have different limiting video resolutions, which may be the lowest resolution that meets the sensing needs of the target object. For example, the video resolution for a user (i.e., the human eye) to view video data differs from the video resolution for a machine to process a recognition task. When a user views video data, a rich display effect is required, so the required video resolution is high. When a machine processes a recognition task, only information related to the object to be recognized needs to be recognized, so the required video resolution is low. Furthermore, the video frame resolution is the video resolution of a video frame in the original video data indicated by the video content characteristics. The encoding device determines the ratio between the limiting video resolution and the video resolution of the video frame in the original video data indicated by the video content characteristics as the target sampling rate in the spatial sampling format. In this way, since the required limiting video resolution of different target objects is different, the target sampling rate in the spatial sampling form is self-adaptively determined based on the limiting video resolution of the target object, so that when the original video data is restored based on the sampled video data by sampling, the quality and resolution of the restored original video data can meet the sensing needs of the target object.

[0052] Preferably, a specific form in which the encoding device determines a target sampling rate in a target sampling format based on video sensing characteristics and video content characteristics includes, if the target sampling format is a temporal sampling format and a spatial sampling format, determining a limiting video frame number corresponding to video data sensed by a target object within a unit time based on the video sensing characteristics and determining a limiting video resolution associated with the target object; determining a target sampling rate in the temporal sampling format based on a ratio between the limiting video frame number and the number of playback video frames, where the number of playback video frames is the number of video frames played within a unit time in the original video data indicated by the video content characteristics; and determining a target sampling rate in the spatial sampling format based on the ratio between the limiting video resolution and the video frame resolution, where the video frame resolution is the video resolution of the video frames in the original video data indicated by the video content characteristics.

[0053] Specifically, if the target sampling format is a temporal sampling format or a spatial sampling format, the encoding device determines a limiting number of video frames of video data sensed by the target object per unit time based on the video sensing characteristics. Because different objects have different video sensing characteristics, the limiting number of video frames corresponding to different objects will be different. Furthermore, the encoding device determines a target sampling rate for the temporal sampling format based on the ratio between the limiting number of video frames and the number of video frames to be played per unit time in the original video data indicated by the video content characteristics. Specifically, since the sampling rate for video frame sampling must be a positive integer, the encoding device obtains the ratio between the limiting number of video frames and the number of video frames to be played per unit time in the original video data indicated by the video content characteristics. If the ratio is a positive integer, the encoding device rounds the ratio to obtain a rounded ratio, and determines the rounded ratio as the target sampling rate for the temporal sampling format. If the ratio is not a positive integer, the encoding device rounds the ratio to obtain a rounded ratio, and determines the rounded ratio as the target sampling rate for the temporal sampling format. In this way, since the limit number of video frames that different target objects can detect is different, the target sampling rate in the time sampling form is self-adaptively determined based on the limit number of video frames corresponding to the target object, so that when the original video data is restored based on the sampled video data obtained by sampling, the quality and resolution of the restored original video data can meet the detection needs of the target object.

[0054] Furthermore, the encoding device determines a limiting video resolution associated with the target object based on the video sensing characteristics. Different target objects have different limiting video resolutions, and the limiting video resolution is the lowest resolution that meets the sensing needs of the target object. The encoding device determines a target sampling rate in spatial sampling format as a ratio between the limiting video resolution and the video resolution of the video frame in the original video data indicated by the video content characteristics. In this way, since different target objects require different limiting video resolutions, the target sampling rate in spatial sampling format is self-adaptively determined based on the limiting video resolution of the target object, so that when the original video data is restored based on the sampled video data, the quality and resolution of the restored original video data can meet the sensing needs of the target object.

[0055] S103: Perform sampling processing on the original video data based on the target sampling parameters to obtain sampled video data.

[0056] Specifically, the encoding device performs sampling on the original video data according to the target sampling parameters to obtain sampled video data. The target sampling parameters include a target sampling format and a target sampling rate for the target sampling format, and the encoding device performs sampling on the original video data according to the target sampling format and the target sampling rate for the target sampling format to obtain sampled video data. In this way, the original video data is sampled to obtain the sampled video data, and the sampled video data is then encoded to obtain video encoded data corresponding to the original video data, thereby reducing the amount of video encoded data, improving the transmission efficiency of the video encoded data, and reducing the storage space for the video encoded data.

[0057] Preferably, a specific form in which the encoding device performs sampling processing on the original video data based on the target sampling parameters to obtain sampled video data includes, if the target sampling form is a time sampling form, a step of obtaining the playback number of a video frame in the original video data and the total number of video frames included in the original video data; a step of determining the number of video frames to be extracted from the original video data as a first video frame number based on the target sampling rate and the total video frame number in the time sampling form; and a step of extracting the first video frame number of video frames from the original video data as sampled video data according to the playback number of the video frame in the original video data.

[0058] Specifically, if the target sampling format is the time sampling format, the encoding device obtains the playback numbers of video frames in the original video data and the total number of video frames included in the original video data. Based on the target sampling rate and the total number of video frames in the time sampling format, the encoding device determines the number of video frames to be extracted from the original video data as a first number of video frames. Specifically, the encoding device obtains the ratio between the total number of video frames and the target sampling rate in the time sampling format (i.e., total number of video frames / target sampling rate in the time sampling format) as the first number of video frames. For example, if the total number of video frames included in the original video data is 100 frames and the target sampling rate in the time sampling format is 2 times, the first number of video frames is 100 / 2=50. Furthermore, the encoding device extracts the first number of video frames from the original video data as sampled video data according to the playback numbers of the video frames in the original video data.

[0059] Specifically, the encoding device extracts video frames from the original video data at intervals according to the playback numbers of the video frames in the original video data and based on the target sampling rate in the time sampling format, and the extracted video frames are sampled video data. As shown in Figure 4, Figure 4 is a schematic diagram of the time sampling format provided by an embodiment of the present application. Referring to Figure 4, the total number of video frames included in the original video data is 10 frames, the target sampling rate in the time sampling format is 2 times, and the original video data includes video frame 0, video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 6, video frame 7, video frame 8, video frame 9, etc. The encoding device extracts one video frame from the original video data at intervals of one video frame, that is, extracts video frame 0, video frame 2, video frame 4, video frame 6, video frame 8, etc. as sampled video data. In other words, based on a 2x magnification in the time sampling format, sampling is performed on video frame 0, video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 6, video frame 7, video frame 8, video frame 9, etc. contained in the original video data to obtain sampled video data, i.e., video frame 0, video frame 2, video frame 4, video frame 6, video frame 8, etc.

[0060] Preferably, after the encoding device samples the original video data using a time sampling format and a target sampling rate in the time sampling format to obtain sampled video data, in order to ensure that the decoding device can restore the total number of video frames of the original video data, the encoding device transmits the total number of video frames and the target sampling rate in the time sampling format to the decoding device, and the decoding device restores the number of frames of the original video data by performing a sampling recovery process on the sampled video data corresponding to the video encoding data based on the total number of video frames and the target sampling rate in the time sampling format.

[0061] Preferably, when the encoding device samples the original video data using the temporal sampling format, to ensure that the decoding device can restore the total number of video frames of the original video data, the encoding device transmits the number of tail-discarded frames whose tails are discarded after sampling and the target sampling rate in the temporal sampling format to the decoding device, and the decoding device restores the number of frames of the original video data by performing a sampling recovery process on the sampled video data corresponding to the video encoding data based on the number of tail-discarded frames and the target sampling rate in the temporal sampling format. Note that the number of tail-discarded frames may be the number of video frames whose tails are discarded after the original video data is temporally sampled.

[0062] Specifically, when the encoding device transmits the number of tail-discarded frames whose tails are discarded after sampling and the target sampling rate for the temporal sampling type to the decoding device, it generates TemporalScaleFlag (temporal sampling label), TemporalRatio (temporal sampling type target sampling rate label), and DroppedFrameNumber (tail-discarded frame number label). TemporalScaleFlag may be set to 0 or 1. A value of 1 indicates that the encoding device samples the original video data using the temporal sampling type, while a value of 0 indicates that the encoding device does not sample the original video data using the temporal sampling type. When TemporalScaleFlag is set to 1, TemporalRatio is set to the target sampling rate for the temporal sampling type, and the value of TemporalRatio may be 2, 3, 4, etc. When TemporalScaleFlag is set to 1, DroppedFrameNumber is set to the number of video frames whose tails are discarded after temporal sampling of the original video data. For example, after sampling video frame 0, video frame 1, video frame 2, video frame 3, video frame 4, video frame 5, video frame 6, video frame 7, video frame 8, and video frame 9 included in the original video data using the time sampling form and a target sampling rate of 2 for the time sampling form, the number of video frames whose tails are discarded is 1 (i.e., the tail video frame 9 is discarded). By performing sampling processing on the original video data according to the first video frame number determined based on the total number of video frames and the target sampling rate based on the playback numbers of the video frames in the original video data, it is possible to avoid omissions and deviations and ensure the accuracy of video frame sampling.

[0063] Preferably, the encoding device performs sampling on the original video data according to the target sampling parameters. The specific form of obtaining the sampled video data is as follows: if the target sampling form is a spatial sampling form, the video frame M in the original video data is sampled; i obtaining an original video resolution of i, where i is a positive integer less than or equal to M, and M is the number of video frames in the original video data; i Based on the original video resolution of i to perform resolution conversion on a video frame M i and performing resolution sampling on all video frames in the original video data, and then determining the original video data with the converted resolution as the sampled video data.

[0064] Specifically, if the target sampling format is the spatial sampling format, the encoding device i , and the original video resolution reflects the number of pixels in the original video data. Specifically, i The higher the original video resolution, the faster the video frame M i There are many pixels in the video frame M i is clear, and the video frame M i The lower the original video resolution, the more i contains fewer pixels, and the video frame M iFor example, the pixel points included in a video frame with a video resolution of 1920*1080 are larger than the pixel points included in a video frame with a video resolution of 720*480, but the amount of data obtained by encoding the video frame with a video resolution of 1920*1080 is larger than the amount of data obtained by encoding the video frame with a video resolution of 720*480. Note that M is the number of video frames in the original video data and is a positive integer. For example, the value of M may be 1, 2, 3, etc., and i is a positive integer equal to or less than M.

[0065] Furthermore, the encoding device may select a target sampling rate and a video frame rate in spatial sampling format. i Based on the original video resolution of i to perform resolution conversion on a video frame M i After performing resolution sampling on all video frames in the original video data, the resolution-converted original video data is determined as sampled video data. In this way, sampling is performed on the video resolution of the original video data using the spatial sampling form and the target sampling rate in the spatial sampling form, ensuring that the detection needs of the target object are met, while reducing the amount of video encoding data corresponding to the original video data, and further improving the transmission efficiency of the video encoding data, so that the decoding device can quickly obtain and decode the video encoding data, thereby improving the decoding efficiency.

[0066] Specifically, the encoding device uses one of the spatial sampling methods, such as nearest neighbor interpolation, resampling filtering, bilinear interpolation, and sampling model prediction (e.g., video or image super-resolution neural network), to generate a video frame M having the original video resolution. i to perform resolution conversion on a video frame M iA video frame M with the original video resolution is obtained. i includes Q original pixel points and pixel values ​​corresponding to the Q original pixel points, where Q is a positive integer, and the encoding device uses nearest neighbor interpolation to generate a video frame M i to perform resolution conversion on a video frame M i The specific manner of obtaining the initial video resolution includes the following: the product of the target sampling rate in the spatial sampling form and the original video resolution is the initial video resolution, which includes P sampling pixel points, where P is a positive integer. Furthermore, the encoding device obtains the sampling pixel points P from the Q original pixel points. j The reference pixel point corresponding to the sampling pixel point P is determined. j belongs to P sampling pixel points, and j is a positive integer less than or equal to P. The pixel value of the reference pixel point is j and pixel values ​​corresponding to the P sampling pixel points are acquired, a video frame M having an initial video resolution is generated based on the P sampling pixel points and the pixel values ​​corresponding to the P sampling pixel points. i Generate.

[0067] Preferably, the encoding device generates a video frame M having a target video resolution. i The specific form of obtaining the target sampling rate in the spatial sampling form and the video frame M i and setting the initial video resolution as a product of the original video resolution of the video frame M i to perform resolution conversion on the video frame M i and obtaining a video frame M having an initial video resolution. i does not satisfy the encoding condition, the video frame M i Pixel filling is performed on the filled video frame M iThe video resolution of the filled video frame M is determined as the target video resolution. i , a video frame M having the target video resolution i and determining a video frame M having an initial video resolution. i satisfies the encoding condition, the initial video resolution is determined as the target video resolution, and the video frame M having the initial video resolution is generated. i , a video frame M having the target video resolution i and determining:

[0068] Specifically, the encoding device determines the target sampling rate in the spatial sampling format and the video frame M i The initial video resolution is the product of the original video resolution of the video frame M i to perform resolution conversion on the video frame M i The encoding device obtains the video frame M iThe original video resolution of video frame 50a is denoted as width*height (i.e., width*height), and the target sampling rate in the spatial sampling format is denoted as q, so that the initial video resolution is width*q*height*q. As shown in FIG. 5, FIG. 5 is a schematic diagram of the spatial sampling format provided by an embodiment of the present application. Referring to FIG. 5, if the target sampling rate in the spatial sampling format is denoted as 1 / 2 and the original resolution of video frame 50a is width*height, the encoding device obtains the initial video resolution of width / 2*height / 2 by multiplying the 1 / 2 sampling rate in the spatial sampling format by the original resolution of video frame 50a. Furthermore, the encoding device performs resolution conversion on video frame 50a having the original video resolution of width*height using any one of the spatial sampling methods including nearest neighbor interpolation, resampling filtering, bilinear interpolation, and sampling model prediction, to obtain video frame 50b having the initial video resolution of width / 2*height / 2. Since the encoder in the encoding device can only encode video frames having a fixed resolution format, the encoding device encodes a video frame M having an initial video resolution. i It is detected whether the video frame M of the initial video resolution satisfies the encoding condition. The encoding condition is that the resolution is a multiple of 8, that is, the width and height of the video resolution are a multiple of 8. For example, i is 720*480, where 720 is a multiple of 8 and 480 is also a multiple of 8, so the encoding device can generate a video frame M i satisfies the encoding condition, and a video frame M i is 727*483, 727 is not a multiple of 8, and 483 is not a multiple of 8, so the encoding device must generate a video frame M i does not satisfy the encoding condition.

[0069] Furthermore, the encoding device generates a video frame M having the initial video resolution. idoes not satisfy the encoding condition, the video frame M i Pixel filling is performed on the initial video resolution so that the video resolution satisfies the encoding conditions, and the filled video frame M i The encoding device obtains the filled video frame M i The video resolution of the filled video frame M is determined as the target video resolution. i , a video frame M having the target video resolution i Specifically, the encoding device determines the video frame M having the initial video resolution. i When pixel filling is performed on a video frame M at the initial video resolution, the video encoding resolution that satisfies the encoding conditions and has the smallest difference from the initial video resolution is obtained as the target video resolution. i is 717*479, it can be determined that the difference between the video resolution of 720*480 and the initial video resolution of 714*475 is the smallest, and therefore the video frame M i Pixel filling is performed on the video frame M with a resolution of 720*480. i Get.

[0070] Specifically, the encoding device uses the target pixel values ​​to encode video frame M i The target pixel value may be any pixel value, such as 0 or 255, and the embodiments of the present application do not limit the target pixel value. For example, the encoding device may fill three pixel points with a pixel value of 0 in the horizontal direction (i.e., width direction) of the initial video resolution, and fill one pixel point with a pixel value of 0 in the vertical direction (i.e., height direction) of the initial video resolution, to generate a video frame M with a video resolution of 720*480 after filling. i A video frame M with the initial video resolution is obtained. i satisfies the encoding condition, the initial video resolution is determined as the target video resolution, and the video frame M having the initial video resolution is generated. i, a video frame M having the target video resolution i is decided.

[0071] Preferably, the encoding device samples the original video data using a spatial sampling scheme, and then generates a video frame M having a target video resolution. i is obtained by filling pixels in a video frame M with the target video resolution i The pixel filling position information of the video frame M having the target video resolution is obtained, and the target sampling rate and pixel filling position information in the spatial sampling format are sent to the decoding device. The decoding device performs a sampling recovery process on the sampled video data obtained by decoding the video encoding data according to the target sampling rate and pixel filling position information in the spatial sampling format to restore the original video data. i is obtained without pixel filling, the target sampling rate in the spatial sampling form is sent to the decoding device, and the decoding device performs sampling recovery processing on the sampled video data corresponding to the video encoding data according to the target sampling rate in the spatial sampling form to restore the original video data.

[0072] Specifically, the encoding device can generate SpatialScaleFlag (temporal and spatial sampling label), SpatialScaleRatio (target sampling rate label for spatial sampling), PaddingFlag (pixel filling label), PaddingX (horizontal filling pixel value, i.e., width), and PaddingY (vertical filling pixel value, i.e., height). SpatialScaleFlag can be set to 0 or 1. Setting SpatialScaleFlag to 1 indicates that the encoding device samples the original video data using the spatial sampling format, while setting SpatialScaleFlag to 0 indicates that the encoding device does not sample the original video data using the spatial sampling format. Setting SpatialScaleFlag to 1 sets SpatialScaleRatio to the target sampling rate for the spatial sampling format. For example, SpatialScaleRatio can be set to any value greater than 0, such as 0.5, 0.75, or 2. When the value of SpatialScaleFlag is set to 1, the value of PaddingFlag is set to 0 or 1. When the value of PaddingFlag is set to 1, it indicates that a pixel filling operation is performed during the spatial sampling process of the original video data (i.e., the video frame M having the target video resolution). i is obtained by filling pixels), and if the value of PaddingFlag is 0, it indicates that there is no pixel filling operation in the spatial sampling process of the original video data (i.e., the video frame M with the target video resolution i was obtained without padding pixels). Additionally, if the value of PaddingFlag is 1, the encoding device may set PaddingX (horizontal fill pixel value) and PaddingY (vertical fill pixel value).

[0073] In this embodiment, an initial video resolution is determined, and then a target video resolution is determined by filling in the initial video resolution for video frames that do not satisfy the encoding conditions. By directly determining the target video resolution for video frames that have the initial video resolution and satisfy the encoding conditions, the amount of video encoding data corresponding to the original video data is reduced, and the transmission efficiency of the video encoding data is improved, and the decoding device can quickly obtain and decode the video encoding data, thereby improving the decoding efficiency.

[0074] Preferably, the specific form of the encoding device performing sampling processing on the original video data includes, if the target sampling form is a time sampling form and a space sampling form, determining the number of video frames to be extracted from the original video data as a second number of video frames based on the target sampling rate of the time sampling form and the total number of video frames in the original video data; extracting the second number of video frames from the original video data as initial sampled video data according to the playback numbers of the video frames in the original video data; and extracting the N video frames in the initial sampled video data. j obtaining an original video resolution of j, where j is a positive integer less than or equal to N, and N is the number of video frames in the initial sampling video data; j Based on the original video resolution of j , and perform resolution conversion on the video frame N jand performing resolution sampling on all video frames in the initial sampling video data, and then determining the initial sampling video data after resolution sampling as the sampled video data. For specific details, please refer to the above-mentioned content when the target sampling format is the temporal sampling format and the target sampling format is the spatial sampling format, and no further details will be provided in the embodiments of the present application. By performing sampling processing on the original video data using a target sampling format that combines the temporal sampling format and the spatial sampling format, the sampling accuracy of the original video data can be improved, ensuring video viewing quality and effectively reducing redundancy in the video encoding data.

[0075] If the target sampling format is a temporal sampling format and a spatial sampling format, the encoding device transmits the total number of video frames in the original video data and the target sampling rate in the temporal sampling format to the decoding device. j is obtained by filling pixels, the encoding device generates a video frame N having a target sampling rate and a target video resolution in spatial sampling form. j and transmitting pixel filling position information in the video frame N having the target video resolution to the decoding device. j is obtained by filling pixels, the encoding device sends the target sampling rate in spatial sampling form to the decoding device, and the decoding device restores the video frame rate and video resolution of the original video data.

[0076] S104: Encode the sampled video data to obtain video encoding data corresponding to the original video data.

[0077] Specifically, the encoding device predicts sampled video data using a prediction method such as intraframe prediction or interframe prediction to obtain a residual video signal of the sampled video data. The encoding device then transforms the residual video signal of the sampled video data to obtain a transform domain signal of the sampled video data, and quantizes the transform domain signal of the sampled video data to obtain a quantized transform domain signal. The encoding device then performs entropy coding on the quantized transform domain signal to output binary (0 or 1) encoded video data. Of course, the encoding device may also perform entropy coding on parameters such as the target sampling parameter, the total number of video frames of the original video data, and pixel filling position information to reduce the bit rate. Thus, in the embodiment of the present application, encoding the sampled video data reduces the bit rate of the encoded video data, thereby improving the transmission efficiency of the encoded video data and reducing the storage space of the encoded video data.

[0078] Preferably, the encoding device determines key video region information (e.g., a region of interest (ROI)) for the original video data and sends the key video region information and the video encoding data to the decoding device, where the key video region information instructs the decoding device to perform image enhancement processing on the key video region in the original video data. In this way, the decoding device performs image enhancement processing on the key video region in the original video data, thereby enriching the display effect of the original video data restored by the decoding device and further improving the accuracy of the restoration of the original video data. Furthermore, when an object recognition task is subsequently performed on the restored original video data, the recognition accuracy is improved.

[0079] Preferably, the encoding device determines key video region information for the original video data in a specific manner, including: inputting the original video data into a target detection model, converting an embedding vector for the original video data using an embedding layer in the target detection model to obtain a media embedding vector for the original video data; extracting an object from the media embedding vector using an object extraction layer in the target detection model to obtain a video object in the original video data; determining a region in the original video data to which the video object belongs as a key video region; and generating key video region information for describing the location of the key video region in the original video data. Of course, the encoding device may extract key video regions in the original video data using other key video region determination methods (e.g., target detection, target recognition, etc.), and the embodiments of the present application are not limited thereto. Specifically, after the encoding device extracts key video regions in the original video data, it generates an ROINumber (region number label) and an ROIInformation (region feature information label, e.g., region coordinate information). ROINumber indicates the number of the key video region of the current original video data or the current video frame, and if ROINumber is greater than 0, ROIInformation whose number is ROINumber is transmitted, and the ROIInformation indicates information about the key video region (i.e., ROI region of interest), such as coordinates, the type of video object contained, etc. As shown in Table 1, the encoding device performs temporal sampling and spatial sampling on the original video data and then transmits the parameters in Table 1, so that the decoding device recovers the original video data based on the parameters in Table 1. [Table 1]

[0080] [Table 1]

[0081] In an embodiment of the present application, target sampling parameters for the original video data are self-adaptively determined based on the media application scenario and video content of the original video data, and sampling processing is performed on the original video data based on the target sampling parameters to obtain sampled video data, thereby improving the sampling accuracy of the original video data, ensuring video viewing quality, and effectively reducing the redundancy of the video encoding data.Furthermore, the sampled video data is encoded to obtain video encoding data, and the video encoding data is sent to a decoding device, thereby reducing the data volume of the video encoding data and further improving the transmission efficiency of the video encoding data, allowing the decoding device to quickly obtain the video encoding data, thereby improving the encoding efficiency of the original video data.

[0082] As shown in Figure 6, Figure 6 is a video decoding method provided by the embodiments of the present application. Below, in combination with Figure 6, the video decoding method provided by the embodiments of the present application will be described in detail, and the method is performed by a computer device, and the computer device may be a decoding device. As shown in Figure 6, the method specifically includes, but is not limited to, the following steps:

[0083] S201: Obtain video encoding data to be decoded and target sampling parameters corresponding to the video encoding data.

[0084] Specifically, the decoding device obtains video encoding data to be decoded and target sampling parameters corresponding to the video encoding data, the video encoding data is obtained by encoding sampled video data, and the sampled video data is obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters, and the target sampling parameters are determined based on the media application scenario and video content features of the original video data.

[0085] Preferably, the target sampling parameters are transmitted from the encoding device, and include a target sampling type and a target sampling rate for the target sampling type, where the target sampling type is determined based on video content characteristics of the original video data. The target sampling rate for the target sampling type is determined based on video sensing characteristics and video content characteristics, where the video sensing characteristics are sensing characteristics of a target object in a media application scenario for video data, and the target object is an object for performing sensing processing on the original video data. The target sampling type includes a temporal sampling type, a target sampling rate for the temporal sampling type, a spatial sampling type, and a target sampling rate for the spatial sampling type. For determining the target sampling parameters, see step S102 in FIG. 3 above. The video encoding data to be decoded is obtained by encoding based on the self-adaptively determined target sampling type and target sampling rate for the target sampling type, where the target sampling type and target sampling rate for the target sampling type are determined based on a media application scenario and video content characteristics, thereby reducing the data amount of the video encoding data obtained by encoding the original video data and further reducing the data amount of the video decoding process, thereby improving the decoding efficiency of the video data.

[0086] S202: The video encoding data to be decoded is decoded to obtain sampled video data.

[0087] Specifically, the decoding process by the decoding device is opposite to the encoding process by the encoding device. The decoding device receives video coded data to be decoded transmitted from the encoding device, performs entropy decoding on the video coded data to obtain parameters and quantized transform coefficients, performs inverse quantization and inverse transform on the quantized transform coefficients to obtain a residual video signal, obtains a corresponding video prediction signal based on known coding mode information transmitted from the encoding device, and adds the video prediction signal and the residual video signal to obtain a reconstructed video signal. Finally, loop filtering is performed on the reconstructed video signal to generate sampled video data. Note that the sampled video data may be sampled video data obtained by sampling original video data in the encoding device.

[0088] S203: Perform a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data.

[0089] Specifically, the decoding device performs a sampling restoration process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video-encoded data. In this way, the decoding device quickly obtains the low-bit-rate video-encoded data, and performs decoding and sampling restoration on the low-bit-rate video-encoded data to obtain original video data corresponding to the video-encoded data, thereby improving the decoding efficiency of the video-encoded data and reducing the storage space of the video-encoded data.

[0090] Preferably, a specific form of the decoding device performing sampling recovery processing on the sampled video data includes the steps of: if the target sampling format is a time sampling format, determining a third number of video frames between the first decoded video frame and the second decoded video frame based on a target sampling rate of the time sampling format, where the first decoded video frame and the second decoded video frame are video frames in the sampled video data that have adjacent playback numbers, and the third number of video frames is the number of recovery video frames to be recovered between the first decoded video frame and the second decoded video frame; inserting the third number of recovery video frames between the first decoded video frame and the second decoded video frame; and after inserting the recovery video frames between any two adjacent decoded video frames in the sampled video data, determining original video data corresponding to the video encoding data based on the restored sampled video data.

[0091] Specifically, when the parameters transmitted from the encoding device and received by the decoding device indicate that the original video data should be sampled using a temporal sampling format, the decoding device determines that the target sampling format is the temporal sampling format, obtains a target sampling rate for the temporal sampling format, and obtains a third number of restored video frames to be restored between the first decoded video frame and the second decoded video frame based on the target sampling rate for the temporal sampling format. Specifically, if the target sampling rate for the temporal sampling format is n times (i.e., collecting one video frame at an interval of one video frame), the third number of restored video frames to be restored between the first decoded video frame and the second decoded video frame is n-1. For example, if the target sampling rate for the temporal sampling format is 2 times (i.e., collecting one video frame at an interval of one video frame), the third number of restored video frames to be restored between the first decoded video frame and the second decoded video frame is 1. If the target sampling rate in the time sampling format is 3 times (i.e., one video frame is collected every two video frames), the third number of restored video frames to be restored between the first decoded video frame and the second decoded video frame is 2. Furthermore, the encoding device inserts the third number of restored video frames between the first decoded video frame and the second decoded video frame. After inserting the restored video frames between any two adjacent decoded video frames in the sampled video data, the original video data corresponding to the video encoding data is determined based on the restored sampled video data.

[0092] Preferably, the restored video frame of the third video frame number is generated based on the first decoded video frame, or the restored video frame of the third video frame number is generated based on the second decoded video frame, or the restored video frame of the third video frame number is generated based on the first decoded video frame and the second decoded video frame.

[0093] Specifically, the restored video frame of the third video frame number may be the same as the first decoded video frame, i.e., the first decoded video frame of the third video frame number is inserted between the first decoded video frame and the second decoded video frame. Of course, the restored video frame of the third video frame number may be the same as the second decoded video frame, i.e., the second decoded video frame of the third video frame number is inserted between the first decoded video frame and the second decoded video frame. In other words, a duplicated video frame of the third video frame number (the duplicated video frame may be the first decoded video frame or the second decoded video frame) is inserted between the first decoded video frame and the second decoded video frame. Of course, the restored video frame of the third video frame number is obtained by performing network prediction based on the first decoded video frame and the second decoded video frame. For example, the decoding device predicts the object driving information of the repaired video frames between the first decoded video frame and the second decoded video frame based on the object driving information in the first decoded video frame and the object motion information in the second decoded video frame, and inserts some repaired video frames between video frames whose playback numbers are adjacent to each other, thereby realizing accurate and fast decoding of the video encoding data.

[0094] Preferably, a specific form of the decoding device determining original video data corresponding to the video encoding data based on the restored sampled video data includes the steps of: inserting restored video frames between any two adjacent decoded video frames in the sampled video data, and then obtaining a fourth video frame number of the video frames included in the restored sampled video data and a total video frame number of the video frames included in the original video data; if the fourth video frame number and the total video frame number are the same, determining the restored sampled video data as original video data corresponding to the video encoding data; if the fourth video frame number and the total video frame number are different, obtaining a difference between the fourth video frame number and the total video frame number as a fifth video frame number, and inserting restored video frames of the fifth video frame number after the restored sampled video data to obtain original video data corresponding to the video encoding data.

[0095] Specifically, if the parameters received by the decoding device from the encoding device include the total number of video frames included in the original video data, the decoding device inserts restored video frames between any two adjacent decoded video frames in the sampled video data, and then determines the number of video frames included in the restored sampled video data as a fourth video frame number. Furthermore, the decoding device detects whether the fourth video frame number of the video frames included in the restored sampled video data is the same as the total video frame number of the video frames included in the initial video data. If the fourth video frame number and the total video frame number are different, the difference between the fourth video frame number and the total video frame number is determined as a fifth video frame number, and the restored video frames of the fifth video frame number are inserted after the restored sampled video data to obtain original video data corresponding to the video encoding data. This allows the number of video frames of the original video data to be accurately restored. The restored video frame of the fifth video frame number is determined based on the video frame that is last in the playback order in the restored sampled video data. For example, the restored video frame of the fifth video frame number may be the video frame that is last in the playback order in the restored sampled video data, or may be obtained by performing network prediction based on the video frame that is last in the playback order in the restored sampled video data. If the fourth video frame number and the total video frame number are the same, the restored sampled video data may be determined as original video data corresponding to the video encoding data.

[0096] Specifically, when the parameters acquired by the decoding device and transmitted from the encoding device include the number of trailing discarded video frames of the video frames included in the original video data, the decoding device directly inserts the restored video frames of the number of trailing discarded video frames after the restored sampled video data to obtain the original video data corresponding to the video encoded data. In this way, the number of video frames of the original video data can be accurately restored according to the total number of video frames included in the original video data, thereby ensuring the quality of the video data decoding. Of course, the restored video frames of the number of trailing discarded video frames can be determined based on the video frame that is last to be played in the restored sampled video data.

[0097] Preferably, a specific form of the decoding device performing sampling restoration processing on the sampled video data includes: if the target sampling format is a spatial sampling format, obtaining a current target video resolution of a third decoded video frame in the sampled video data; performing a resolution restoration processing on the third decoded video frame having the target video resolution according to the target sampling rate and the target video resolution in the spatial sampling format to obtain a third decoded video frame having an original video resolution; and after restoring all the decoded video frames in the sampled video data, determining the restored sampled video data as original video data corresponding to the sampled video data.

[0098] Specifically, if the target sampling format is spatial sampling format, taking the third video frame in the sampled video data as an example, the third decoded video frame belongs to any video frame in the sampled video data, and the encoding device obtains a current target video resolution of the third decoded video frame in the sampled video data and a target sampling rate in the spatial sampling format. The decoding device then obtains an initial video resolution based on the ratio between the current target video resolution of the third decoded video frame and the target sampling rate in the spatial sampling format (i.e., target video resolution / target sampling rate in the spatial sampling format). The encoding device then performs a restoration process on the third decoded video frame having the target video resolution using a spatial sampling restoration method to obtain a third decoded video frame having the initial video resolution, and then generates a third decoded video frame having the original video resolution based on the third decoded video frame having the initial video resolution. The spatial sampling restoration method may be any one of nearest neighbor interpolation, resampling filtering, bilinear interpolation, and sampling model prediction, and may be the same as or different from the spatial sampling method in the encoding device. After restoring all the decoded video frames in the sampled video data, the restored sampled video data is determined as original video data corresponding to the video encoding data, thereby performing resolution restoration processing at the target video resolution on the video encoding data obtained by encoding in a spatial sampling format, thereby accurately restoring the original video resolution of the original video data and improving decoding accuracy.

[0099] Preferably, a specific manner in which the decoding device obtains the third decoded video frame having the original video resolution includes: when receiving pixel filling position information for the third decoded video frame, performing pixel cropping on the third decoded video frame having the target video resolution according to the pixel filling position information to obtain the third decoded video frame having the initial video resolution, determining a ratio between the initial video resolution and a target sampling rate in spatial sampling form as the original video resolution to be restored for the third decoded video frame, and performing a resolution restoration process on the third decoded video frame having the initial video resolution to obtain the third decoded video frame having the original video resolution; and when not receiving pixel filling position information for the third decoded video frame, determining a ratio between the target sampling rate in spatial sampling form and the target video resolution as the original video resolution to be restored for the third decoded video frame, and performing a resolution restoration process on the third decoded video frame having the target video resolution to obtain the third decoded video frame having the original video resolution.

[0100] For example, if the target sampling rate in the spatial sampling format is 0.5, the target video resolution of the third decoded video frame is 960*544, and the received pixel filling position information corresponding to the third decoded video frame indicates that four pixels are filled vertically, the decoding device performs pixel cropping of four pixels on the vertical pixel values ​​of the target video resolution of the third decoded video frame of 960*544 to obtain a third decoded video frame with an initial video resolution of 960*540. Furthermore, the decoding device obtains the ratio between the initial video resolution of 960*540 and 0.5 (i.e., the target sampling rate in the spatial sampling format) (i.e., (960*540) / 0.5) to obtain the original video resolution to be restored for the third decoded video frame, i.e., 1920*1080. Furthermore, the decoding device performs a resolution restoration process on the third decoded video frame having the target video resolution based on any one of nearest neighbor interpolation, resampling filtering, bilinear interpolation, and sampling model prediction to obtain a third decoded video frame having the original video resolution. In this way, the resolution is restored based on the pixel filling position information or by directly determining the original video resolution, thereby accurately restoring the original video resolution of the original video data and improving the decoding accuracy.

[0101] Preferably, a specific form of the decoding device performing sampling restoration processing on the sampled video data includes the steps of: if the target sampling format is a temporal sampling format and a spatial sampling format, obtaining a target video resolution of a third decoded video frame in the sampled video data; and performing a resolution restoration processing on the target video resolution of the third decoded video frame according to the target sampling rate and the target video resolution in the spatial sampling format to obtain a third decoded video frame having an original video resolution; and after restoring all the decoded video frames in the sampled video data, determining the restored sampled video data as initial video data corresponding to the video encoding data. and determining a sixth number of video frames of restoration video frames to be restored between the first initial video frame and the second initial video frame based on the target sampling rate in the time sampling form, where the first initial video frame and the second initial video frame belong to video frames in the initial video data that have adjacent playback numbers; and inserting the sixth number of restoration video frames between the first initial video frame and the second initial video frame, and inserting the restoration video frames between any two adjacent initial video frames in the initial video data, and then determining original video data corresponding to the video encoding data based on the restored initial video data.

[0102] Specifically, when the decoding device receives parameters such as TemporalScaleFlag (temporal sampling label), TemporalRatio (target sampling rate label in temporal sampling form), DroppedFrameNumber (label for the number of discarded frames at the end), SpatialScaleFlag (temporal-spatial sampling label), SpatialScaleRatio (target sampling rate label in spatial sampling form), PaddingFlag (pixel filling label), PaddingX (horizontal filling pixel value, i.e., width direction), and PaddingY (vertical filling pixel value, i.e., height direction) transmitted from the encoding device, the decoding device determines whether spatial sampling has been performed on the decoded sampled video data based on the SpatialScaleFlag in the parameters, and if the value of SpatialScaleFlag is 1, continuously obtains SpatialScaleRatio and PaddingFlag from the parameters. If the value of PaddingFlag is 1, continuously obtains PaddingX and PaddingY to trim horizontal and vertical filling pixels in the sampled video data. An original video resolution is calculated based on SpatialScaleRatio and the target video resolution of the cropped sampled video data, and a restoration process is performed on the cropped sampled video data using a spatial sampling method (e.g., nearest neighbor interpolation, bilinear interpolation, video or image super-resolution neural network, etc.) to obtain initial video data corresponding to the sampled video data.

[0103] Furthermore, the decoding device determines whether the sampled video data is temporally sampled according to the TemporalScaleFlag in the parameter, and continuously obtains TemporalRatio and DroppedFrameNumber from the parameter if the value of TemporalScaleFlag is 1. The decoding device determines a sixth video frame number of the restored video frames between the first and second initial video frames in the initial video data according to the TemporalRatio, where the first and second initial video frames belong to video frames in the initial video data that have adjacent playback numbers, and determines the video frames to be restored between the first and second initial video frames in the form of, for example, duplicated frames, video frame insertion network, etc. (See the above description of determining the video frames to be restored between the first and second decoded video frames). A restored video frame of the sixth video frame number is inserted between the first initial video frame and the second initial video frame, and a restored video frame is inserted between any two adjacent initial video frames in the initial video data. Then, the missing video frames are supplemented in the restored initial video data according to the DroppedFrameNumber to obtain the corresponding original video data after trimming.

[0104] Preferably, the decoding device receives key video region information related to the video encoding data transmitted from the encoding device, and determines key video regions in the original video data based on the key video region information. Then, the decoding device performs image enhancement on the key video regions in the original video data to obtain image-enhanced original video data. Specifically, after obtaining the ROI Number (region number label) and ROI Information (region feature information label, e.g., region coordinate information), the decoding device performs image enhancement on all key video regions in the original video data based on the ROI Number (region number label) and ROI Information (region feature information label, e.g., region coordinate information). In this way, the display effect of the restored original video data and the accuracy of the restored original video data are improved. Furthermore, when an object recognition task is subsequently performed on the restored original video data, the recognition accuracy is improved.

[0105] Preferably, the specific form in which the decoding device performs image enhancement on the key video region in the original video data to obtain the image-enhanced original video data includes the steps of inputting the key video region in the original video data into an image enhancement model, generating image enhancement coefficients for the key video region in the original video data using an enhancement coefficient generation layer in the image enhancement model, and performing image enhancement on the key video region in the original video data using the image enhancement layer in the image enhancement model based on the image enhancement coefficients to obtain the image-enhanced original video data. In this way, by performing image enhancement processing on the key video region in the original video data using the image enhancement model, the efficiency and accuracy of image enhancement can be improved. Of course, the decoding device may use other image enhancement methods to perform image enhancement processing on the key video region in the original video data, and the embodiments of the present application are not limited thereto.

[0106] In an embodiment of the present application, video encoding data to be decoded is decoded to obtain sampled video data, and target sampling parameters are determined based on the media application scenario and video content characteristics of the original video data corresponding to the video encoding data. Thus, the video encoding data is obtained by performing sampling encoding on the original video data based on the target sampling parameters, and the target sampling parameters are determined based on the media application scenario and video content characteristics of the original video data corresponding to the video encoding data. Since the video encoding data is obtained by encoding sampled video data, i.e., this video encoding data is obtained by encoding a portion of the video content in the original video data, only the encoded data of a portion of the video content needs to be decoded during the decoding process, thereby improving the decoding efficiency of the video data. Furthermore, by performing sampling recovery processing on the sampled video data based on the target sampling parameters, the original video data can be recovered to a certain extent, thereby improving the quality of the video data.

[0107]

[0023] Referring to Figure 7, Figure 7 is a structural schematic diagram of a video decoding device provided by an embodiment of the present application. The video decoding device may be computer-readable instructions (including program code) executed by a computer device, for example, the video decoding device is application software, and the video decoding device can perform corresponding steps in the video decoding method provided by an embodiment of the present application. As shown in Figure 7, the video decoding device includes a first acquisition module 11, a decoding module 12, a sampling recovery module 13, a receiving module 14, a first determination module 15 and an image enhancement module 16.

[0108] The first acquisition module 11 acquires video encoding data to be decoded and target sampling parameters corresponding to the video encoding data, the video encoding data is obtained by encoding sampled video data, and the sampled video data is obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters, the target sampling parameters being determined based on the media application scenario and video content features of the original video data, the decoding module 12 decodes the video encoding data to obtain the sampled video data, and the sampling recovery module 13 performs a sampling recovery process on the sampled video data based on the target sampling parameters to obtain the original video data corresponding to the video encoding data.

[0109] The target sampling parameters are transmitted from the encoding device, and the target sampling parameters include a target sampling type and a target sampling rate in the target sampling type, the target sampling type is determined based on the video content characteristics of the original video data, and the target sampling rate in the target sampling type is determined based on the video sensing characteristics and the video content characteristics, the video sensing characteristics are the sensing characteristics of the target object in the media application scenario for the video data, and the target object is an object that performs sensing processing on the original video data.

[0110] If the target sampling format is a time sampling format, the sampling restoration module 13 includes: a first determining unit 1301 for determining a third number of video frames between a first decoded video frame and a second decoded video frame based on a target sampling rate of the time sampling format, where the first decoded video frame and the second decoded video frame are video frames in the sampled video data that have adjacent playback numbers, and the third number of video frames is the number of restored video frames to be restored between the first decoded video frame and the second decoded video frame; an inserting unit 1302 for inserting the third number of restored video frames between the first decoded video frame and the second decoded video frame; and a second determining unit 1303 for inserting the restored video frames between any two adjacent decoded video frames in the sampled video data and then determining original video data corresponding to the video encoding data based on the restored sampled video data.

[0111] The recovered video frame of the third video frame number is generated based on the first decoded video frame, or the recovered video frame of the third video frame number is generated based on the second decoded video frame, or the recovered video frame of the third video frame number is generated based on the first decoded video frame and the second decoded video frame.

[0112] The second determining unit 1303 specifically inserts a restored video frame between any two adjacent decoded video frames in the sampled video data, and then obtains a fourth video frame number of the video frames included in the restored sampled video data and a total video frame number of the video frames included in the original video data. If the fourth video frame number and the total video frame number are the same, the restored sampled video data is determined to be the original video data corresponding to the video encoding data. If the fourth video frame number and the total video frame number are different, the difference between the fourth video frame number and the total video frame number is obtained as a fifth video frame number. The restored video frame of the fifth video frame number is inserted after the restored sampled video data to obtain the original video data corresponding to the video encoding data.

[0113] The sampling restoration module 13 further includes: a first obtaining unit 1304 for obtaining a current target video resolution of a third decoded video frame in the sampled video data if the target sampling format is a spatial sampling format; a resolution restoration unit 1305 for performing a resolution restoration process on the third decoded video frame having the target video resolution based on the target sampling rate and the target video resolution in the spatial sampling format to obtain a third decoded video frame having an original video resolution; and a third determining unit 1306 for determining the restored sampled video data as original video data corresponding to the video encoding data after restoring all the decoded video frames in the sampled video data.

[0114] The resolution restoration unit 1305 specifically performs at least one of the following steps: when receiving the pixel filling position information for the third decoded video frame, performing pixel cropping on the third decoded video frame having the target video resolution according to the pixel filling position information to obtain the third decoded video frame having an initial video resolution, determining a ratio between the initial video resolution and the target sampling rate in spatial sampling form as the original video resolution to be restored for the third decoded video frame, and performing a resolution restoration process on the third decoded video frame having the initial video resolution to obtain the third decoded video frame having the original video resolution; when not receiving the pixel filling position information for the third decoded video frame, determining a ratio between the target sampling rate in spatial sampling form and the target video resolution as the original video resolution to be restored for the third decoded video frame, and performing a resolution restoration process on the third decoded video frame having the target video resolution to obtain the third decoded video frame having the original video resolution.

[0115] The sampling restoration module 13 includes a fourth determining unit 1307 for obtaining a target video resolution of a third decoded video frame in the sampled video data if the target sampling format is a temporal sampling format and a spatial sampling format, and performing a resolution restoration process on the target video resolution of the third decoded video frame according to the target sampling rate and the target video resolution in the spatial sampling format to obtain a third decoded video frame with the original video resolution; a fifth determining unit 1308 for determining the restored sampled video data as initial video data corresponding to the video encoding data after restoring all the decoded video frames in the sampled video data; and a fifth determining unit 1309 for determining the target video resolution of the third decoded video frame in the temporal sampling format according to the target sampling rate and the target video resolution in the spatial sampling format to obtain a third decoded video frame with the original video resolution. The video encoding data further includes a sixth determining unit 1309 for determining a sixth number of video frames of restored video frames to be restored between the first initial video frame and the second initial video frame based on a sampling rate, where the first initial video frame and the second initial video frame belong to video frames in the initial video data that have adjacent playback numbers; and a seventh determining unit 1310 for inserting the sixth number of restored video frames between the first initial video frame and the second initial video frame, and inserting the restored video frames between any two adjacent initial video frames in the initial video data, and then determining original video data corresponding to the video encoding data based on the restored initial video data.

[0116] The video decoding device further includes a receiving module 14 that receives key video region information related to the video encoding data transmitted from the encoding device, a first determination module 15 that determines a key video region in the original video data based on the key video region information, and an image enhancement module 16 that performs image enhancement on the key video region in the original video data to obtain original video data after image enhancement.

[0117] The image enhancement module 16 includes a first generation unit 1601 that inputs a key video region in the original video data into an image enhancement model and generates an image enhancement coefficient for the key video region in the original video data using an enhancement coefficient generation layer in the image enhancement model, and an image enhancement unit 1602 that performs image enhancement on the key video region in the original video data based on the image enhancement coefficient using the image enhancement layer in the image enhancement model, and obtains original video data after image enhancement.

[0118] According to one embodiment of the present application, each module in the video decoding device of Figure 7 may be individually or entirely configured to be combined into one or several units, or some of the units may be further decomposed into multiple functionally smaller sub-units, which may perform the same operations without affecting the technical effect of the embodiment of the present application. The above modules are divided based on logical functions, and in actual applications, the functions of one module may be realized by multiple units, or the functions of multiple modules may be realized by one unit. In other embodiments of the present application, the video decoding device may include other units, and in actual applications, these functions may be realized by other units cooperating with each other, or by multiple units cooperating with each other.

[0119] The decoding device decodes the video encoding data to obtain the sampled video data, which is obtained by encoding the sampled video data, i.e., the video encoding data is obtained by encoding a portion of the video content in the original video data, so that only a portion of the encoded video content needs to be decoded during the decoding process, thereby improving the decoding efficiency of the video data.Furthermore, by performing a sampling recovery process on the sampled video data based on the target sampling parameters, the original video data can be recovered to a certain extent, thereby improving the quality of the video data.

[0120]

[0023] Referring to Figure 8, Figure 8 is a structural schematic diagram of a video encoding device provided by an embodiment of the present application. The video encoding device may be computer-readable instructions (including program code) executed by a computer device, for example, the video encoding device is application software, and the video encoding device can perform corresponding steps in the video encoding method provided by an embodiment of the present application. As shown in Figure 8, the video encoding device includes a second acquisition module 21, a second determination module 22, a sampling processing module 23, an encoding module 24, a first sending module 25, a second sending module 26, a third sending module 27, a fourth determination module 28 and a fourth sending module 29.

[0121] The second acquisition module 21 acquires the media application scenario and video content features of the original video data to be encoded.

[0122] The second determination module 22 determines target sampling parameters for performing sampling processing on the original video data based on the media application scenario and the video content characteristics.

[0123] The sampling processing module 23 performs sampling processing on the original video data based on the target sampling parameters to obtain sampled video data.

[0124] Encoding module 24 encodes the sampled video data to obtain video encoding data that corresponds to the original video data.

[0125] The second determination module 22 includes: an eighth determination unit 2201 for determining a target sampling form for performing sampling processing on original video data based on video content features; a ninth determination unit 2202 for determining video sensing features for video data of a target object in a media application scenario, where the target object is an object for performing sensing processing on the original video data; a tenth determination unit 2203 for determining a target sampling rate in the target sampling form based on the video sensing features and video content features; and an eleventh determination unit 2204 for determining the target sampling rate and the target sampling form as target sampling parameters for performing sampling processing on the original video data.

[0126] The eighth determination unit 2201 specifically determines the duplication rate of the video content in the original video data based on the video content change rate included in the video content characteristics, and determines a target sampling form for performing sampling processing on the original video data based on the duplication rate of the video content in the original video data.

[0127] The step of determining a target sampling form for performing sampling processing on the original video data based on the overlap rate of video content in the original video data includes at least one of the following steps: if the overlap rate of video content in the original video data is greater than a first overlap rate threshold, determining the temporal sampling form and the spatial sampling form as the target sampling form for performing sampling processing on the original video data; if the overlap rate of video content in the original video data is equal to or less than the first overlap rate threshold and greater than a second overlap rate threshold, determining the temporal sampling form as the target sampling form for performing sampling processing on the original video data, where the second overlap rate threshold is smaller than the first overlap rate threshold; and if the overlap rate of video content in the original video data is equal to or less than the second overlap rate threshold, determining the spatial sampling form as the target sampling form for performing sampling processing on the original video data.

[0128] The eighth determination unit 2201 specifically determines the complexity of the video content in the original video data based on the amount of video content information contained in the video content features, and determines a target sampling form for performing sampling processing on the original video data based on the complexity of the video content in the original video data.

[0129] The step of determining a target sampling form for performing sampling processing on the original video data based on the complexity of the video content in the original video data includes at least one of the steps of: determining a temporal sampling form and a spatial sampling form as target sampling forms for performing sampling processing on the original video data if the complexity of the video content of the original video data is less than a first complexity threshold; determining a spatial sampling form as the target sampling form for performing sampling processing on the original video data if the complexity of the video content in the original video data is equal to or greater than the first complexity threshold and less than a second complexity threshold, where the second complexity threshold is greater than the first complexity threshold; and determining a temporal sampling form as the target sampling form for performing sampling processing on the original video data if the complexity of the video content in the original video data is greater than the second complexity threshold.

[0130] The tenth determination unit 2203 specifically determines, if the target sampling form is a time sampling form, a limiting video frame number corresponding to the video data sensed by the target object within a unit time based on the video sensing characteristics, and determines a target sampling rate in the time sampling form based on the ratio between the limiting video frame number and the playback video frame number, where the playback video frame number is the number of video frames to be played within a unit time in the original video data indicated by the video content characteristics.

[0131] The tenth determination unit 2203 specifically determines a limiting video resolution associated with the target object based on the video sensing features if the target sampling format is a spatial sampling format, and determines a ratio between the limiting video resolution and the video frame resolution as a target sampling rate in the spatial sampling format, where the video frame resolution is the video resolution of the video frame in the original video data indicated by the video content features.

[0132] The tenth determination unit 2203 specifically determines, if the target sampling form is a time sampling form and a spatial sampling form, a limiting video frame number corresponding to the video data sensed by the target object within a unit time based on the video sensing features, determines a limiting video resolution associated with the target object, and determines a target sampling rate in the time sampling form based on the ratio between the limiting video frame number and the number of playback video frames, where the number of playback video frames is the number of video frames played within a unit time in the original video data indicated by the video content features, and determines a target sampling rate in the spatial sampling form based on the ratio between the limiting video resolution and the video frame resolution, where the video frame resolution is the video resolution of the video frames in the original video data indicated by the video content features.

[0133] The sampling processing module 23 includes: a second acquisition unit 2301 that acquires, if the target sampling format is a time sampling format, the playback number of a video frame in the original video data and the total number of video frames included in the original video data; a twelfth determination unit 2302 that determines, based on the target sampling rate and the total number of video frames in the time sampling format, the number of video frames to be extracted from the original video data as a first number of video frames; and a first extraction unit 2303 that extracts the first number of video frames from the original video data as sampling video data according to the playback number of the video frame in the original video data.

[0134] If the target sampling format is a spatial sampling format, the sampling processing module 23 performs the sampling of the video frame M in the original video data. i a third acquisition unit 2304 for acquiring an original video resolution of i, where i is a positive integer less than or equal to M, and M is the number of video frames in the original video data; and a third acquisition unit 2304 for acquiring a target sampling rate in a spatial sampling format and a video frame M i Based on the original video resolution of i to perform resolution conversion on a video frame M i and a thirteenth determining unit 2306, which performs resolution sampling on all video frames in the original video data and then determines the resolution-converted original video data as sampling video data.

[0135] The first resolution conversion unit 2305 specifically converts the target sampling rate in the spatial sampling format into the video frame M i The initial video resolution is the product of the original video resolution of the video frame M ito perform resolution conversion on the video frame M i and obtain a video frame M with the initial video resolution i does not satisfy the encoding condition, the video frame M i Pixel filling is performed on the filled video frame M i The video resolution of the filled video frame M is determined as the target video resolution. i , a video frame M having the target video resolution i and a video frame M having an initial video resolution i satisfies the encoding condition, the initial video resolution is determined as the target video resolution, and the video frame M having the initial video resolution is generated. i , a video frame M having the target video resolution i is decided.

[0136] The sampling processing module 23 includes a fourteenth determining unit 2307 for determining, if the target sampling format is a time sampling format and a spatial sampling format, a second video frame number as a number of video frames to be extracted from the original video data according to a target sampling rate in the time sampling format and a total number of video frames in the original video data; a second extracting unit 2308 for extracting, according to the playback numbers of the video frames in the original video data, the second video frame number as initial sampling video data; and a video frame N in the initial sampling video data. j a fourth acquisition unit 2309 for acquiring an original video resolution of j, where j is a positive integer less than or equal to N, and N is the number of video frames in the initial sampling video data; and a fourth acquisition unit 2309 for acquiring a target sampling rate in spatial sampling form and a video frame N j Based on the original video resolution of j, and perform resolution conversion on the video frame N j and a fifteenth determination unit 2311, which performs resolution sampling on all video frames in the initial sampling video data and then determines the initial sampling video data after resolution sampling as the sampling video data.

[0137] The video encoding device further includes a first transmitting module 25 for transmitting the total number of video frames and the target sampling rate in the time sampling form to a decoding device, and the decoding device performs sampling recovery processing on the sampled video data corresponding to the video encoding data based on the total number of video frames and the target sampling rate in the time sampling form.

[0138] The video encoder generates a video frame M having a target video resolution. i is obtained by filling pixels in a video frame M with the target video resolution i a second transmitting module 26 for obtaining pixel filling position information of the pixel filled in the target sampling rate in the spatial sampling format and transmitting the pixel filling position information to a decoding device, and the decoding device performs sampling restoration processing on sampled video data corresponding to the video encoding data according to the target sampling rate in the spatial sampling format and the pixel filling position information; i If the target sampling rate in the spatial sampling form is obtained without filling pixels, the third sending module 27 sends the target sampling rate in the spatial sampling form to the decoding device, and the decoding device performs sampling recovery processing on the sampled video data corresponding to the video encoding data based on the target sampling rate in the spatial sampling form.

[0139] The video encoding device further includes a fourth determination module 28 for determining key video region information for the original video data, and a fourth transmission module 29 for transmitting the key video region information and the video encoding data to a decoding device, where the key video region information instructs the decoding device to perform image enhancement processing on the key video region in the original video data.

[0140] The fourth determination module 28 includes a vector conversion unit 2801 that inputs original video data into a target detection model and performs embedding vector conversion on the original video data using an embedding layer in the target detection model to obtain media embedding vectors of the original video data; an object extraction unit 2802 that extracts objects from the media embedding vectors using an object extraction layer in the target detection model to obtain video objects in the original video data; and a second generation unit 2803 that determines the area to which the video object belongs in the original video data as a key video area and generates key video area information that describes the position of the key video area in the original video data.

[0141] According to one embodiment of the present application, each module in the video encoding device of Figure 8 may be individually or entirely configured to be combined into one or several units, or some of the units may be further decomposed into multiple functionally smaller sub-units, which can perform the same operations without affecting the technical effect of the embodiment of the present application. The above modules are divided based on logical functions, and in actual applications, the functions of one module may be realized by multiple units, or the functions of multiple modules may be realized by one unit. In other embodiments of the present application, the video encoding device may include other units, and in actual applications, these functions may be realized by the cooperation of other units or multiple units.

[0142] The encoding device self-adaptively determines target sampling parameters for the original video data based on the media application scenario and video content of the original video data, and performs sampling on the original video data based on the target sampling parameters to obtain sampled video data, thereby improving the sampling accuracy of the original video data, ensuring video viewing quality, and effectively reducing redundancy of the video-encoded data. Furthermore, the sampled video data is encoded to obtain video-encoded data, which is then transmitted to the decoding device, thereby reducing the amount of video-encoded data and improving transmission efficiency of the video-encoded data. The decoding device can quickly obtain the video-encoded data, thereby improving the encoding efficiency of the original video data.

[0143] Referring to FIG. 9, FIG. 9 is a structural schematic diagram of a computer device provided by an embodiment of the present application. As shown in FIG. 9, the computer device 1000 includes a processor 1001, a network interface 1004, and a memory 1005. The computer device 1000 may further include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 realizes communication between these components. The user interface 1003 may include a display and a keyboard. Preferably, the user interface 1003 may further include a standard wired interface and a wireless interface. Preferably, the network interface 1004 may include a standard wired interface and a wireless interface (e.g., a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one magnetic disk memory. Preferably, the memory 1005 may further include at least one storage device remote from the processor 1001. As shown in FIG. 9, the memory 1005 as a computer-readable storage medium includes an operating system, a network communication module, a user interface module, and a device control application program.

[0144] In the computer device 1000 of FIG. 9, the network interface 1004 provides a network communication function, the user interface 1003 mainly provides an input interface for the user, and the processor 1001 invokes a device control application program stored in the memory 1005 to The method includes the steps of: obtaining video encoding data to be decoded and target sampling parameters corresponding to the video encoding data, wherein the video encoding data is obtained by encoding sampled video data, and the sampled video data is obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters, and the target sampling parameters are determined based on a media application scenario and video content features of the original video data; decoding the video encoding data to obtain the sampled video data; and performing a sampling recovery process on the sampled video data based on the target sampling parameters to obtain the original video data corresponding to the video encoding data.

[0145] Here, the computer device 1000 described in the embodiment of the present application may implement the video decoding method described in the embodiment corresponding to Fig. 6 above, or the video decoding device described in the embodiment corresponding to Fig. 7 above, which will not be described here in detail, and the beneficial effects of the same method will not be described in detail.

[0146] Referring to FIG. 10, FIG. 10 is a structural schematic diagram of a computer device provided by an embodiment of the present application. As shown in FIG. 10, the computer device 2000 includes a processor 2001, a network interface 2004, and a memory 2005. The computer device 2000 also includes a user interface 2003 and at least one communication bus 2002. The communication bus 2002 realizes communication between these components. The user interface 2003 may include a display and a keyboard. Preferably, the user interface 2003 may further include a standard wired interface and a wireless interface. Preferably, the network interface 2004 includes a standard wired interface and a wireless interface (e.g., a Wi-Fi interface). The memory 2005 may be a high-speed RAM memory or a non-volatile memory, such as at least one magnetic disk memory. Preferably, the memory 2005 may be at least one storage device remote from the processor 2001. As shown in FIG. 10, the memory 2005 as a computer-readable storage medium includes an operating system, a network communication module, a user interface module, and a device control application program.

[0147] In the computer device 2000 of FIG. 10, the network interface 2004 provides a network communication function, the user interface 2003 mainly provides an input interface for the user, and the processor 2001 invokes a device control application program stored in the memory 2005, The method includes the steps of acquiring a media application scenario and video content characteristics of original video data to be encoded, determining target sampling parameters for performing sampling processing on the original video data based on the media application scenario and the video content characteristics, performing sampling processing on the original video data based on the target sampling parameters to obtain sampled video data, and encoding the sampled video data to obtain video encoding data corresponding to the original video data.

[0148] Here, the computer device 2000 described in the embodiment of the present application may implement the video encoding method described in the embodiment corresponding to Figure 3 above, or the video encoding device described in the embodiment corresponding to Figure 8 above, and no further details will be given here. In addition, no further details will be given about the beneficial effects of the same method.

[0149] In addition, the embodiments of the present application further provide a computer-readable storage medium, which stores computer-readable instructions to be executed by the video decoding device mentioned above, and the computer-readable instructions include program instructions. When the processor executes the program instructions, the video decoding method in the embodiment corresponding to Figure 6 above or the video encoding method in the embodiment corresponding to Figure 3 can be performed, so no further details will be provided here.

[0150] Furthermore, the beneficial effects of the same method will not be described in detail. For technical details not disclosed in the embodiments of the computer-readable storage medium of the present application, please refer to the description of the method embodiments of the present application. For example, the program instructions may be executed by one computer device, or by multiple computer devices located in one location, or by multiple computer devices distributed across multiple locations and connected via a communication network, where the multiple computer devices distributed across multiple locations and connected via a communication network constitute a blockchain system.

[0151]

[0023] In addition, an embodiment of the present application further provides a computer program product, the computer program product including computer-readable instructions stored in a computer-readable storage medium, a processor of a computer device reading and executing the computer-readable instructions from the computer-readable storage medium, and causing the computer device to perform the video decoding method in the embodiment corresponding to Figure 6 or the video encoding method in the embodiment corresponding to Figure 3.

[0152] For the sake of simplicity, the above method embodiments are each expressed as a series of combinations of operations. However, those skilled in the art will appreciate that the present application is not limited to the order of operations described, because some steps may be performed in other orders or simultaneously according to the present application. Furthermore, those skilled in the art will appreciate that the embodiments described in the specification are preferred embodiments, and that such operations and modules are not necessarily required for the present application.

[0153] The steps in the method of the present application may be rearranged, merged, or deleted based on actual needs.

[0154] The modules in the embodiment device of the present application may be merged, divided and deleted based on actual needs.

[0155] As can be understood by those skilled in the art, all or part of the steps in the above-described method embodiments can be realized by instructing relevant hardware with computer-readable instructions, and the program is stored in a computer-readable storage medium, which may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0156] The above disclosure is merely a preferred embodiment of the present invention, and does not limit the scope of the claims of the present invention. Therefore, equivalent modifications according to the claims of the present invention still belong to the scope of the present invention. [Explanation of symbols]

[0157] 11 First Acquisition Module 12 Decryption Module 13 Sampling Recovery Module 14 Receiver Module 15 First Decision Module 16 Image Enhancement Module 21 Second Acquisition Module 22 Second Decision Module 23 Sampling Processing Module 24 Encoding Module 25 First Transmitting Module 26 Second Transmitting Module 27 Third Transmitting Module 28 Fourth Decision Module 29 Fourth Transmitting Module 1000 Computer Equipment 1001 processor 1002 communication bus 1003 User Interface 1004 Network Interface 1005 memory 1301 First Decision Unit 1302 Insertion Unit 1303 Second Decision Unit 1304 First Acquisition Unit 1305 Resolution Recovery Unit 1306 Third Decision Unit 1307 Fourth Decision Unit 1308 5th Decision Unit 1309 6th Decision Unit 1310 Seventh Decision Unit 1601 First Generating Unit 1602 Image Enhancement Unit 2000 Computer Equipment 2001 processor 2002 communication bus 2003 User Interface 2004 Network Interface 2005 Memory 2201 8th Decision Unit 2202 9th Decision Unit 2203 10th Decision Unit 2204 11th Decision Unit 2301 Second Acquisition Unit 2302 12th Decision Unit 2303 First Extraction Unit 2304 Third Acquisition Unit 2305 First Resolution Conversion Unit 2306 13th Decision Unit 2307 14th Decision Unit 2308 Second Extraction Unit 2309 4th Acquisition Unit 2310 Second Resolution Conversion Unit 2311 15th Decision Unit 2801 Vector Transformation Unit 2802 Object Extraction Unit 2803 Second Generating Unit

Claims

1. 1. A video encoding method implemented by a computing device, comprising: obtaining a media application scenario and video content characteristics of original video data to be encoded; determining target sampling parameters for performing sampling processing on the original video data based on the media application scenario and the video content characteristics, determining a video content overlap rate in the original video data based on a video content change rate included in the video content characteristics; determining a target sampling form for performing sampling processing on the original video data based on a video content overlap rate in the original video data; determining video sensing characteristics of a target object in the media application scenario for video data, the target object being an object that performs sensing processing on the original video data; determining a target sampling rate for the target sampling configuration based on the video sensing characteristics and the video content characteristics; determining the target sampling rate and the target sampling form as target sampling parameters for performing sampling processing on the original video data; and performing a sampling process on the original video data based on the target sampling parameters to obtain sampled video data; encoding the sampled video data to obtain video encoding data corresponding to the original video data; A method comprising:

2. determining a target sampling form for performing sampling processing on the original video data based on a video content overlap rate in the original video data, If the overlap rate of the video content in the original video data is greater than a first overlap rate threshold, determining a time sampling form and a spatial sampling form as target sampling forms for performing sampling processing on the original video data; a step of determining the time sampling form as a target sampling form for performing sampling processing on the original video data if the overlap rate of video content in the original video data is equal to or less than the first overlap rate threshold and greater than a second overlap rate threshold, wherein the second overlap rate threshold is smaller than the first overlap rate threshold; If the overlap rate of the video content in the original video data is equal to or less than the second overlap rate threshold, determining the spatial sampling form as a target sampling form for performing sampling processing on the original video data; The method of claim 1 , comprising at least one of:

3. determining a target sampling form for performing sampling processing on the original video data based on the video content characteristics, determining a complexity of video content in the original video data based on an amount of video content information included in the video content characteristics; determining a target sampling form for performing sampling processing on the original video data based on the complexity of video content in the original video data; The method of claim 1 , comprising:

4. determining a target sampling form for performing sampling processing on the original video data based on the complexity of video content in the original video data, If the complexity of the video content of the original video data is less than a first complexity threshold, determining a time sampling form and a spatial sampling form as target sampling forms for performing sampling processing on the original video data; determining the spatial sampling form as a target sampling form for performing sampling processing on the original video data if the complexity of the video content in the original video data is equal to or greater than the first complexity threshold and less than a second complexity threshold, wherein the second complexity threshold is greater than the first complexity threshold; If the complexity of the video content in the original video data is greater than the second complexity threshold, determining the time sampling form as a target sampling form for performing sampling processing on the original video data; The method of claim 3 , comprising at least one of:

5. determining a target sampling rate for the target sampling configuration based on the video sensing characteristics and the video content characteristics, If the target sampling format is a time sampling format, determining a limit number of video frames corresponding to video data sensed by the target object within a unit time based on the video sensing characteristics; determining a target sampling rate for the time sampling configuration based on a ratio between the limited number of video frames and the number of playback video frames; Including, the number of video frames to be reproduced indicates the number of video frames to be reproduced within a unit time in the original video data indicated by the video content characteristics; The method of claim 1.

6. determining a target sampling rate for the target sampling configuration based on the video sensing characteristics and the video content characteristics, if the target sampling format is a spatial sampling format, determining a limiting video resolution associated with the target object based on the video sensing characteristics; determining a ratio between the limited video resolution and a video frame resolution to a target sampling rate for the spatial sampling configuration; Including, the video frame resolution indicates a video resolution of a video frame in the original video data indicated by the video content feature; The method of claim 1.

7. determining a target sampling rate for the target sampling configuration based on the video sensing characteristics and the video content characteristics, If the target sampling format is a time sampling format and a spatial sampling format, determining a limit video frame number corresponding to video data sensed by the target object within a unit time based on the video sensing characteristics, and determining a limit video resolution associated with the target object; determining a target sampling rate for the time sampling format based on a ratio between the limited number of video frames and a number of playback video frames, wherein the number of playback video frames indicates a number of video frames to be played within a unit time in the original video data indicated by the video content characteristics; determining a ratio between the limited video resolution and a video frame resolution to a target sampling rate for the spatial sampling configuration; Including, the video frame resolution indicates a video resolution of a video frame in the original video data indicated by the video content feature; The method of claim 1.

8. The step of obtaining sampled video data by performing a sampling process on the original video data based on the target sampling parameters includes: If the target sampling format is a time sampling format, obtaining a playback number of a video frame in the original video data and a total number of video frames included in the original video data; determining a first number of video frames to be extracted from the original video data based on a target sampling rate in the time sampling format and the total number of video frames; extracting the first number of video frames from the original video data according to playback numbers of the video frames in the original video data as the sampling video data; The method of claim 1 , comprising:

9. The step of obtaining sampled video data by performing a sampling process on the original video data based on the target sampling parameters includes: If the target sampling format is a spatial sampling format, then the video frame M in the original video data i obtaining an original video resolution of i, where i is a positive integer less than or equal to M, and M is the number of video frames in the original video data; The target sampling rate for the spatial sampling form and the video frame M i Based on the original video resolution of i performing resolution conversion on M to obtain a video frame M having a target video resolution; performing resolution sampling on all video frames in the original video data, and then determining the resolution-converted original video data as the sampled video data; The method of claim 1 , comprising:

10. The target sampling rate for the spatial sampling form and the video frame M i Based on the original video resolution of i performing resolution conversion on the video frame Mi to obtain a video frame Mi having a target video resolution, The target sampling rate in the spatial sampling format and the video frame M i and setting the initial video resolution as the product of the original video resolution of the A video frame M having the original video resolution i to perform resolution conversion on the video frame M having the initial video resolution. i obtaining a Video frame M having the initial video resolution i does not satisfy the encoding condition, the video frame M i Pixel filling is performed on the filled video frame M i The video resolution of the filled video frame M is determined as the target video resolution. i , a video frame M having a target video resolution. i and determining Video frame M having the initial video resolution i satisfies the encoding condition, the initial video resolution is determined to be the target video resolution, and a video frame M having the initial video resolution is generated. i a video frame M having the target video resolution i and determining 10. The method of claim 9, comprising:

11. The step of obtaining sampled video data by performing a sampling process on the original video data based on the target sampling parameters includes: If the target sampling format is a time sampling format and a spatial sampling format, determining a second number of video frames to be extracted from the original video data based on a target sampling rate of the time sampling format and a total number of video frames in the original video data; extracting the second number of video frames from the original video data according to playback numbers of the video frames in the original video data as initial sampling video data; Video frame N in the initial sampling video data j obtaining an original video resolution of j, where j is a positive integer less than or equal to N, and N is the number of video frames in the initial sampled video data; The target sampling rate for the spatial sampling format and the video frame N j Based on the original video resolution of j performing resolution conversion on N to obtain a video frame Nj having a target video resolution; performing resolution sampling on all video frames in the initial sampling video data, and then determining the initial sampling video data after resolution sampling as the sampling video data; The method of claim 1 , comprising:

12. 1. A video decoding method implemented by a computer device, comprising: obtaining video encoding data to be decoded and a target sampling parameter corresponding to the video encoding data, the video encoding data being obtained by encoding sampled video data, the sampled video data being obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameter, and the target sampling parameter being determined based on a media application scenario and video content features of the original video data; decoding the video encoding data to obtain the sampled video data; performing a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data; the target sampling parameters are transmitted from an encoding device, and the target sampling parameters include a target sampling format and a target sampling rate in the target sampling format; the target sampling configuration is determined based on the video content characteristics; a target sampling rate in the target sampling manner is determined based on video sensing characteristics and the video content characteristics, the video sensing characteristics being sensing characteristics of a target object in the media application scenario with respect to video data, and the target object being an object that performs sensing processing on the original video data; if the target sampling format is a time sampling format, determining a third number of video frames between a first decoded video frame and a second decoded video frame based on a target sampling rate of the time sampling format, wherein the first decoded video frame and the second decoded video frame belong to video frames in the sampled video data that have adjacent playback numbers, and the third number of video frames is the number of recovery video frames to be recovered between the first decoded video frame and the second decoded video frame; inserting the third number of recovery video frames between the first decoded video frame and the second decoded video frame; inserting a restored video frame between any two adjacent decoded video frames in the sampled video data, and then determining original video data corresponding to the video encoding data based on the restored sampled video data; Including steps and A method comprising:

13. the step of inserting any restored video frame between any two adjacent decoded video frames in the sampled video data, and then determining original video data corresponding to the video encoding data based on the restored sampled video data, comprises: After inserting a restored video frame between any two adjacent decoded video frames in the sampled video data, obtaining a fourth video frame number of the video frames included in the restored sampled video data and a total video frame number of the video frames included in the original video data; if the fourth number of video frames and the total number of video frames are equal, determining the recovered sampled video data as original video data corresponding to the video encoding data; if the fourth number of video frames is different from the total number of video frames, obtaining a difference between the fourth number of video frames and the total number of video frames as a fifth number of video frames, and inserting restored video frames of the fifth number of video frames after the restored sampled video data to obtain original video data corresponding to the video encoding data; 13. The method of claim 12, comprising:

14. performing a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data, if the target sampling format is a spatial sampling format, obtaining a current target video resolution of a third decoded video frame in the sampled video data; performing a resolution restoration process on a third decoded video frame having the target video resolution according to a target sampling rate in the spatial sampling format and the target video resolution to obtain a third decoded video frame having an original video resolution; after recovering all the decoded video frames in the sampled video data, determining the recovered sampled video data as original video data corresponding to the video encoding data; 13. The method of claim 12, comprising:

15. performing a resolution restoration process on the target video resolution of the third decoded video frame according to a target sampling rate in the spatial sampling form and the target video resolution to obtain the third decoded video frame having an original video resolution; when pixel filling position information for the third decoded video frame is received, performing pixel cropping on the third decoded video frame having the target video resolution according to the pixel filling position information to obtain a third decoded video frame having an initial video resolution; determining a ratio between the initial video resolution and a target sampling rate in the spatial sampling format as an original video resolution to be restored for the third decoded video frame; and performing a resolution restoration process on the third decoded video frame having the initial video resolution to obtain a third decoded video frame having the original video resolution; if pixel filling position information for the third decoded video frame has not been received, determining a ratio between a target sampling rate in the spatial sampling format and the target video resolution as an original video resolution to be restored for the third decoded video frame, and performing a resolution restoration process on the third decoded video frame having the target video resolution to obtain the third decoded video frame having the original video resolution; The method of claim 14 , comprising at least one of:

16. performing a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data, if the target sampling format is a temporal sampling format and a spatial sampling format, obtaining a target video resolution of a third decoded video frame in the sampled video data, and performing a resolution restoration process on the target video resolution of the third decoded video frame according to the target sampling rate in the spatial sampling format and the target video resolution to obtain a third decoded video frame having an original video resolution; after recovering all decoded video frames in the sampled video data, determining the recovered sampled video data as initial video data corresponding to the video encoding data; determining a sixth video frame number of recovery video frames between a first initial video frame and a second initial video frame according to a target sampling rate in the time sampling format, the sixth video frame number being a recovery target video frame, the first initial video frame and the second initial video frame belonging to video frames in the initial video data having adjacent playback numbers; inserting the sixth number of restored video frames between the first initial video frame and the second initial video frame, and inserting the restored video frames between any two adjacent initial video frames in the initial video data, and then determining original video data corresponding to the video encoding data based on the restored initial video data; 13. The method of claim 12, comprising:

17. 1. A video decoding device, comprising: a first acquisition module for acquiring video encoding data to be decoded and a target sampling parameter corresponding to the video encoding data, the first acquisition module comprising: the video encoding data is obtained by encoding sampled video data; the sampled video data is obtained by performing a sampling process on original video data corresponding to the video encoding data based on the target sampling parameters; The target sampling parameter is determined based on a media application scenario and video content characteristics of the original video data; the target sampling parameters are transmitted from an encoding device, and the target sampling parameters include a target sampling format and a target sampling rate in the target sampling format; the target sampling configuration is determined based on the video content characteristics; a target sampling rate in the target sampling manner is determined based on video sensing characteristics and the video content characteristics, the video sensing characteristics being sensing characteristics of a target object in the media application scenario with respect to video data, and the target object being an object that performs sensing processing on the original video data; a first determining unit for determining, if the target sampling format is a time sampling format, a third number of video frames between a first decoded video frame and a second decoded video frame according to a target sampling rate of the time sampling format, wherein the first decoded video frame and the second decoded video frame belong to video frames in the sampled video data that have adjacent playback numbers, and the third number of video frames is the number of recovery video frames to be recovered between the first decoded video frame and the second decoded video frame; an insertion unit for inserting the third number of recovery video frames between the first decoded video frame and the second decoded video frame; a second determining unit for determining original video data corresponding to the video encoding data based on the restored sampled video data after inserting a restored video frame between any two adjacent decoded video frames in the sampled video data; a first acquisition module comprising: a decoding module for decoding the video encoding data to obtain the sampled video data; a sampling recovery module that performs a sampling recovery process on the sampled video data based on the target sampling parameters to obtain original video data corresponding to the video encoding data; a video decoding device comprising:

18. 1. A video encoding device, comprising: a second obtaining module for obtaining media application scenarios and video content features of original video data to be encoded; a second determination module for determining target sampling parameters for performing sampling processing on the original video data based on the media application scenario and the video content characteristics, an eighth determining unit for determining a target sampling form for performing sampling processing on the original video data based on the video content characteristics, determining a video content duplication rate in the original video data based on a video content change rate included in the video content characteristics; determining a target sampling form for performing sampling processing on the original video data based on the overlap rate of video content in the original video data; an eighth decision unit; and a ninth determining unit for determining a video sensing feature of a target object in the media application scenario with respect to video data, the target object being an object that performs sensing processing on the original video data; a tenth determining unit for determining a target sampling rate for the target sampling configuration based on the video sensing characteristics and the video content characteristics; an eleventh determining unit for determining the target sampling rate and the target sampling form as target sampling parameters for performing sampling processing on the original video data; a second determination module comprising: a sampling processing module that performs sampling processing on the original video data based on the target sampling parameters to obtain sampled video data; an encoding module for encoding the sampled video data to obtain video encoding data corresponding to the original video data; 1. A video encoding device comprising:

19. A computing device comprising a processor and a memory, A computer device, the processor and a memory coupled thereto, the memory storing computer readable instructions, the processor invoking the computer readable instructions to cause the computer device to perform the method of any one of claims 1 to 16.

20. A computer program comprising computer readable instructions which, when executed by a processor, perform the method of any one of claims 1 to 16.

Citation Information

Patent Citations

  • Video encoding and decoding method and device, computer device, and storage medium

    US20200374514A1

  • Method and apparatus for video encoding and decoding

    US20200382792A1

  • Digital foveation for machine vision

    US20210089803A1