Adaptive coding and decoding optimization method and device, equipment, medium and program product

Through the adaptive codec optimization method, the codec parameters are dynamically adjusted using network status information and device capability matrix, which solves the network adaptability and delay problems of traditional fixed-code rate solutions when bandwidth fluctuations are fluctuating, and realizes efficient and low-latency audio and video streaming.

CN120378415APending Publication Date: 2025-07-25启朔(深圳)科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510505181.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The traditional fixed-code rate scheme has a significant decrease in peak signal-to-noise ratio when bandwidth fluctuates, poor network adaptability, high end-to-end delay at high loads, and the fixed-resolution grading strategy leads to an increase in the lag rate in network switching scenarios.

Method used

By obtaining network status information, a device capability matrix is built, the optimal codec algorithm and parameters are determined based on the preset user experience quality and the maximum available code rate, and the network state is obtained using dual channels of active detection and passive analysis, and the time series model and reinforcement learning model are combined to optimize codec parameters.

Benefits of technology

Effectively respond to bandwidth fluctuations, maintain high efficiency and high quality of audio and video stream processing, significantly reduce latency, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378415A_ABST
    Figure CN120378415A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive coding and decoding optimization method and device, equipment, a medium and a program product, and relates to the technical field of cloud computing and multimedia transmission crossing. The method comprises the following steps: acquiring network state information, and performing bandwidth prediction based on the network state information to obtain a bandwidth prediction result; acquiring hardware capability information of the equipment through the hardware abstraction layer interface, and constructing an equipment capability matrix based on the hardware capability information; calculating the maximum available code rate based on the bandwidth prediction result and the equipment capability matrix; determining an optimal coding and decoding algorithm and optimal coding and decoding parameters based on preset user experience quality and the maximum available code rate; and optimizing the target audio and video based on the optimal coding and decoding algorithm and the optimal coding and decoding parameters to obtain an optimized audio and video stream. According to the embodiment of the invention, the bandwidth fluctuation condition can be effectively handled, the high efficiency and high quality of audio and video stream processing are kept, the delay of audio and video stream transmission is remarkably reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to the cross - field of cloud computing and multimedia transmission, and particularly to an adaptive codec optimization method, apparatus, device, medium, and program product. Background Art

[0002] In traditional fixed - bitrate schemes, the Peak Signal - to - Noise Ratio (PSNR) drops significantly when the bandwidth fluctuates. For example, in scenarios such as 3G / 4G network handovers, the frame rate fluctuation range of the fixed - bitrate scheme reaches ±30%, and the stuttering rate rises to 15%, indicating poor network adaptability. Moreover, the fixed - resolution grading strategy of the fixed bitrate also has a relatively high end - to - end delay under high load. Summary of the Invention

[0003] Embodiments of the present disclosure propose an adaptive codec optimization method, apparatus, device, medium, and program product.

[0004] In a first aspect, embodiments of the present disclosure propose an adaptive codec optimization method, including: obtaining network status information, predicting the bandwidth based on the network status information to obtain a bandwidth prediction result; obtaining the hardware capability information of the device through a hardware abstraction layer interface, and constructing a device capability matrix based on the hardware capability information; calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix; determining an optimal codec algorithm and optimal codec parameters based on a preset quality of user experience and the maximum available bitrate; and optimizing the target audio - video based on the optimal codec algorithm and optimal codec parameters to obtain an optimized audio - video stream.

[0005] Further, the obtaining of the network status information includes:

[0006] Obtaining the network status information through a dual - channel of active probing and passive analysis.

[0007] Further, the network status information includes bandwidth, round - trip time, and jitter rate. The obtaining of the network status information through a dual - channel of active probing and passive analysis includes:

[0008] Measuring the round - trip time based on an active probing message, and calculating the jitter rate based on the round - trip time;

[0009] Parsing the TCP ACK packet to calculate the bandwidth.

[0010] Further, the network status information includes bandwidth, round - trip time, and jitter rate. The predicting of the bandwidth based on the network status information to obtain a bandwidth prediction result includes:

[0011] Use a time series model to perform bandwidth prediction based on the bandwidth, round-trip time, and jitter rate to obtain the bandwidth prediction result.

[0012] Further, the device capability matrix includes: GPU decoding throughput, NPU computing power, and memory bandwidth.

[0013] Calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix includes:

[0014] Determine the target frame rate and the maximum GPU decoding capability based on the GPU decoding throughput, NPU computing power, and memory bandwidth.

[0015] Calculate the maximum available bitrate based on the bandwidth prediction result, the target frame rate, and the maximum GPU decoding capability.

[0016] Further, it also includes:

[0017] In response to the maximum available bitrate not being within the preset threshold range, adjust the device parameters in the device capability matrix, and return to execute the steps of obtaining network status information, performing bandwidth prediction based on the network status information to obtain the bandwidth prediction result, to the step of calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix, until the maximum available bitrate is within the preset threshold range.

[0018] Further, determining the optimal codec parameters based on the preset quality of user experience and the maximum available bitrate includes:

[0019] Construct a state space based on image quality metrics, latency, and user ratings.

[0020] Construct a reward function based on the image quality metrics and the preset quality of user experience.

[0021] Adjust the codec parameters according to the state space through a reinforcement learning model to optimize the reward function.

[0022] Determine the codec parameters corresponding to the maximized reward function as the optimal codec parameters.

[0023] Second aspect, embodiments of the present disclosure propose an adaptive codec optimization device, including: a network probe module configured to obtain network status information, perform bandwidth prediction based on the network status information, and obtain a bandwidth prediction result; a device fingerprint module configured to obtain hardware capability information through a hardware abstraction layer interface and construct a device capability matrix based on the hardware capability information; a bitrate control engine configured to calculate a maximum available bitrate based on the bandwidth prediction result and the device capability matrix; an algorithm decision module configured to determine an optimal codec algorithm and optimal codec parameters based on a preset quality of user experience and the maximum available bitrate; and an audio and video output module configured to optimize target audio and video based on the optimal codec algorithm and the optimal codec parameters to obtain an optimized audio and video stream.

[0024] Third aspect, embodiments of the present invention provide a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding implementation manner thereof.

[0025] Fourth aspect, embodiments of the present invention provide a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding implementation manner thereof.

[0026] Fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, and the computer program can implement the adaptive codec optimization method described in any implementation manner in the first aspect when executed by a processor.

[0027] The adaptive codec optimization method, device, device, medium, and program product provided by the embodiments of the present disclosure can effectively cope with the situation of bandwidth fluctuations, maintain high efficiency and high quality in the processing of audio and video streams, significantly reduce the delay of audio and video stream transmission, and improve the user experience.

[0028] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present disclosure will become more apparent:

[0030] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0031] Figure 2Flowchart of an adaptive encoding and decoding optimization method provided by an embodiment of the present disclosure;

[0032] Figure 3 Flowchart of another adaptive encoding and decoding optimization method provided by an embodiment of the present disclosure;

[0033] Figure 4 Schematic flowchart of another adaptive encoding and decoding optimization method provided by an embodiment of the present disclosure;

[0034] Figure 5 Block diagram of the structure of an adaptive encoding and decoding optimization device provided by an embodiment of the present disclosure;

[0035] Figure 6 Schematic diagram of the structure of an electronic device suitable for executing the adaptive encoding and decoding optimization method provided by an embodiment of the present disclosure. Detailed implementation manners

[0036] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0037] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0038] Figure 1 Exemplary system architecture 100 showing embodiments to which the adaptive encoding and decoding optimization method, device, electronic device, and computer-readable storage medium of the present disclosure can be applied.

[0039] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0040] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for implementing information communication between the two can be installed on terminal devices 101, 102, 103 and server 105, such as instant messaging applications, etc.

[0041] Terminal devices 101, 102, 103 and server 105 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.; when terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, and no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can also be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, and no specific limitation is made here.

[0042] Server 105 can provide various services through various built-in applications. It should be noted that the data or information, etc. required to provide relevant services can be obtained from terminal devices 101, 102, 103 via network 104, and can also be pre-stored locally on server 105 in various ways. Therefore, when server 105 detects that these data have been stored locally, it can choose to directly obtain these data from the local. In this case, exemplary system architecture 100 may not include terminal devices 101, 102, 103 and network 104.

[0043] Since performing encoding and decoding processing may require a large amount of computing resources and strong computing capabilities, the adaptive encoding and decoding optimization method provided in the subsequent embodiments of the present disclosure is generally executed by the server 105 with strong computing capabilities and a large amount of computing resources. Correspondingly, the adaptive encoding and decoding optimization device is generally also set in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have computing capabilities and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the above operations that were originally performed by the server 105 through relevant applications installed thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing capabilities at the same time, when the relevant application determines that the terminal device where it is located has strong computing capabilities and a large amount of remaining computing resources, the terminal device can be allowed to execute the above operations, thereby appropriately reducing the computing pressure on the server 105. Correspondingly, the adaptive encoding and decoding optimization device can also be set in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may not include the server 105 and the network 104 either.

[0044] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0045] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 2 , Figure 2 Please refer to

[0046] Step S201: Obtain network status information, and perform bandwidth prediction based on the network status information to obtain a bandwidth prediction result.

[0047] In the embodiments of the present application, the execution entity of the adaptive encoding and decoding optimization method (such as Figure 1 the server 105 shown) obtains the network status information of the current user side (which can be, for example, the terminal device used by the user or a virtual platform such as a cloud phone), and performs bandwidth prediction based on the network status information to obtain a bandwidth prediction result. Among them, the network status information can, for example, include bandwidth, latency, jitter, packet loss rate, signal strength, etc. The execution entity can perform bandwidth prediction based on the network status information in a historical period to obtain the bandwidth prediction result.

[0048] Step S202: Obtain the hardware capability information of the target device through the hardware abstraction layer interface, and construct a device capability matrix based on the hardware capability information.

[0049] In the embodiments of the present application, the above-mentioned execution entity obtains the hardware capability information of the target device through the hardware abstraction layer interface. The hardware abstraction layer interface refers to the key intermediate layer connecting the software system and the hardware device. Through this interface, the execution entity can obtain the hardware capability information of the target device (which can be, for example, the terminal device used by the user). The hardware capability information refers to the detailed data describing the performance, functions, and limitations of the hardware components of the target device. On this basis, the execution entity can construct a device capability matrix based on the hardware capability information to structure and visualize the hardware capability information for quickly comparing device performance, supporting resource scheduling decisions, or optimizing application compatibility. For example, the device capability matrix can be constructed for device types (such as servers, mobile phones, IoT devices, edge computing nodes, etc.), hardware categories (such as CPUs, GPUs, storage, networks, sensors, etc.), and capability indicators (such as performance indicators: computing power (TOPS), bandwidth (Gbps), latency (ms), function support: whether key technologies (such as ray tracing, 5G) are supported, capacity limitations: memory size, storage space, maximum number of connections), etc.

[0050] Step S203: Calculate the maximum available bitrate based on the bandwidth prediction result and the device capability matrix.

[0051] In the embodiments of the present application, the above-mentioned execution entity calculates the maximum available bitrate according to the bandwidth prediction result and the device capability matrix. The maximum available bitrate refers to the maximum bitrate that can be used for data transmission or storage under specific conditions such as a communication system, network environment, or storage medium. This parameter is used to characterize the data transmission ability in the network environment.

[0052] Step S204: Determine the optimal codec algorithm and optimal codec parameters based on the preset quality of experience of the user and the maximum available bitrate.

[0053] In the embodiments of the present application, the quality of experience (QoE) of the user is a comprehensive indicator for measuring the subjective satisfaction of the user with a service or product. In the embodiments of the present application, the candidate codec algorithms can be initially screened by the maximum available bitrate, and the candidate codec algorithms that do not meet the preset requirements for the maximum available bitrate are screened out. Then, based on the screened candidate codec algorithms, the QoE scores of the candidate codec algorithms are calculated in combination with the QoE standard, and the optimal codec algorithm and optimal codec parameters are determined based on the QoE scores.

[0054] Step S205: Optimize the target audio and video based on the optimal codec algorithm and optimal codec parameters to obtain an optimized audio and video stream.

[0055] In the embodiments of the present application, the original audio-visual data is compressed and encoded using an encoding and decoding algorithm. On the premise of ensuring the integrity of the audio-visual content, the encoding details are adjusted according to the optimal parameters, such as reducing the video resolution to the level suitable for the device and network, setting an appropriate frame rate to balance the smoothness and data volume, and optimizing the audio sampling rate. After encoding, an optimized audio-visual stream is generated. While satisfying the maximum available bitrate limit, this audio-visual stream maximizes the quality of user experience and can provide users with a clear, smooth, and high-quality audio-visual playback experience under the current network environment and target device conditions.

[0056] Through the above process, the adaptive encoding and decoding optimization method provided by the embodiments of the present disclosure can effectively cope with the situation of bandwidth fluctuations, maintain high efficiency and quality in the processing of audio-visual streams, significantly reduce the latency of audio-visual stream transmission, and improve the user experience.

[0057] In some optional implementation manners of this embodiment, in the above step 201, the process of obtaining the network status information mainly obtains the network status information through a dual-channel of active detection (ICMP) and passive analysis. Among them, the active detection ICMP (Internet Control Message Protocol) is a technical means used for fault diagnosis, host reachability detection, and network information acquisition in a network. TCP ACK passive analysis is a network analysis technology that obtains information about network connections and application behaviors by listening to and analyzing TCP ACK (acknowledgment) packets in the network.

[0058] Exemplarily, the network status information may include: bandwidth, round-trip time, and jitter rate. Further, the process of obtaining the network status information through the dual-channel of active detection and passive analysis mainly includes: measuring the round-trip time based on the active detection packet, and calculating the jitter rate based on the round-trip time; parsing the TCP ACK packet to calculate the bandwidth.

[0059] Obtaining the network status information through the dual-channel method of the above process can comprehensively understand the network status from different perspectives, complement each other's deficiencies, and provide more complete network status information.

[0060] In some alternative embodiments of this embodiment, in the above step 201, the process of performing bandwidth prediction based on network status information to obtain a bandwidth prediction result mainly includes: using a time series model to perform bandwidth prediction based on bandwidth, round-trip time, and jitter rate to obtain a bandwidth prediction result. Specifically, based on historical bandwidth data and real-time network status (round-trip time RTT, jitter rate), a time series model or machine learning algorithm is used to predict the available bandwidth within a preset period. Exemplarily, historical bandwidth data refers to bandwidth values recorded at fixed intervals (such as every second). RTT and jitter rate are real-time sampled values obtained at a fixed frequency (such as once every 100 ms). In the embodiments of the present application, the time series model can use an Autoregressive Integrated Moving Average Model (ARIMA model), which is a time series prediction method. The ARIMA model refers to a model established by regressing the dependent variable only on its lag values and the present and lag values of the random error term during the process of transforming a non-stationary time series into a stationary time series. Through this time series model, bandwidth prediction can be performed to obtain a bandwidth prediction result.

[0061] In some alternative embodiments of this embodiment, the device capability matrix mainly includes: GPU decoding throughput, NPU computing power (the computing power of the NPU (Neural Network Processor) is an important indicator to measure its ability to process neural network calculations, usually expressed in Operations Per Second (OPS) or TeraOperations Per Second (TOPS)), and memory bandwidth. Correspondingly, in the above step 203, the process of calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix mainly includes:

[0062] Step 1: Determine the target frame rate and the maximum decoding capability of the GPU based on the GPU decoding throughput, NPU computing power, and memory bandwidth.

[0063] First, relevant parameters of the device will be collected, that is, the key data of GPU decoding throughput, NPU computing power, and memory bandwidth will be extracted from the device capability matrix. These parameters respectively represent the ability of the GPU to decode data per unit time, the computing ability of the NPU, and the data transmission ability of the memory.

[0064] Next, a comprehensive analysis of these parameters will be performed. It will consider the relationship between GPU decoding throughput and the image data processing speed, the support of NPU computing power for complex algorithm processing, and whether the memory bandwidth can provide the required data for the GPU and NPU in a timely manner.

[0065] By establishing an appropriate mathematical model or adopting a preset algorithm rule, the execution entity will, based on the mutual constraints and collaborative relationships among these parameters, determine a target frame rate that can not only fully exert the performance of the device but also ensure smooth playback of audio and video. Meanwhile, in combination with the hardware characteristics of the GPU itself and the current operating environment, these parameters are used to calculate the maximum decoding capacity that the GPU can achieve under the current conditions, that is, the maximum amount of decoded data that the GPU can process per unit time.

[0066] Step 2: Calculate the maximum available bitrate based on the bandwidth prediction result, the target frame rate, and the maximum decoding capacity of the GPU.

[0067] Obtain the previously obtained bandwidth prediction result, which reflects the available bandwidth that the network can provide in the next period of time. Then, the execution entity will conduct a comprehensive consideration in combination with the target frame rate and the maximum decoding capacity of the GPU. The target frame rate determines the number of frames that need to be processed per second, while the maximum decoding capacity of the GPU limits the upper limit of data processing by the GPU per unit time. The execution entity will consider that the network bandwidth needs to meet the audio and video data transmission requirements at the target frame rate and, at the same time, not exceed the maximum decoding capacity of the GPU to avoid problems such as stuttering caused by decoding bottlenecks.

[0068] Through a specific calculation formula, factors such as the bandwidth prediction result, the target frame rate, and the maximum decoding capacity of the GPU are incorporated, weighing various limitations and requirements, and calculating the maximum available bitrate that can be used under the current network conditions and device performance. This maximum available bitrate can ensure that during the transmission and processing of audio and video data, there will be no problems due to the bitrate being too high exceeding the network bandwidth and GPU processing capabilities, nor will the quality of audio and video be affected due to the bitrate being too low.

[0069] For example: After determining the target frame rate and the maximum decoding capacity of the GPU, the execution entity can calculate the maximum available bitrate based on these two parameters. Exemplarily, the maximum available bitrate can be calculated by the following formula (1):

[0070] Q max =min(BW est ×0.8 / FPS target ,GPUd ec_max ), (1)

[0071] where Q max represents the maximum available bitrate, BW est is the bandwidth prediction result, FPS target is the target frame rate, GPU dec_max is the maximum decoding capacity of the GPU, and 0.8 is the safety factor.

[0072] In the embodiment of the present application, the Maximum Available Bitrate (MAB) refers to the highest video bitrate that can be stably transmitted without causing stuttering or packet loss under the current network conditions, hardware performance, and system load. Its range (fluctuation interval) is jointly determined by factors such as dynamically changing network bandwidth, decoding ability, and computing power resources. That is to say, for different network environments and application scenarios, the requirements for the reasonable range of the maximum available bitrate are also different. In some alternative embodiments of this embodiment, the value range of the maximum available bitrate can be further screened. Specifically, the adaptive codec optimization method may further include:

[0073] Judge whether the maximum available bitrate is within the preset threshold range. If the maximum available bitrate is not within the preset threshold range, it indicates that one or several parameters of the current device may not meet the requirements of the current application scenario, and the device parameters in the device capability matrix of the device need to be adjusted accordingly, and return to execute steps 201 to 203 to judge whether the adjusted maximum available bitrate meets the requirements of the preset threshold range; if the maximum bitrate is within the threshold range, it indicates that it can be applied to the current application scenario, and then steps 204 to 205 can be continued to optimize the target audio and video.

[0074] In another embodiment of the present application, the capability matrix information of each terminal device is collected, and hardware performance data such as the CPU, GPU computing power, memory bandwidth, and storage capacity of the device, as well as function information such as the device's support for different codec formats, are obtained through the hardware abstraction layer interface. At the same time, the network status of each device is monitored in real time, including bandwidth, latency, packet loss rate, etc.

[0075] Then, the server evaluates the audio and video playback requirements and bearing capacity of each device based on these device capability matrix and network status data. For devices with high configuration and good network conditions, a higher bitrate audio and video stream can be allocated to ensure a high-definition playback experience; for devices with low configuration or poor network, a lower bitrate but smooth-playing video stream is allocated. At the same time, some codec tasks are sunk to the edge computing nodes of the cell. The server analyzes the complexity of the audio and video content and the processing capabilities of each device, and assigns some codec tasks that can be processed in parallel or have low real-time requirements to the edge server.

[0076] After receiving the task, the edge server performs preliminary decoding and optimization on the audio and video, such as performing resolution adjustment, frame rate adaptation, etc., and then transmits the processed audio and video stream to the corresponding user device. This not only reduces the processing pressure on the cloud server, but also reduces the distance and latency of data transmission, thus significantly improving the efficiency of multi-device collaborative playback of audio and video in the whole home.

[0077] Please refer to Figure 3 , Figure 3 , which is a flowchart of an adaptive codec optimization method provided by an embodiment of the present disclosure. That is, a specific implementation manner is provided for the process of determining the optimal codec parameters in step 204 of the process 200 shown in Figure 2 . Other steps in the process 200 are not adjusted, and a new complete embodiment is obtained by replacing step 204 with the specific implementation manner provided by this embodiment. The process 300 includes the following steps:

[0078] Step S301: Construct a state space based on the picture quality index, latency, and user rating.

[0079] In the embodiment of the present application, based on the deep reinforcement learning (DRL) framework, a reinforcement learning model (such as the DQN (Deep Q-Network) model) is adopted, and through continuous interaction with the environment (audio-visual transmission system), the optimal bitrate control and algorithm selection strategy are learned. Specifically, the picture quality index may refer to the SSIM (Structural Similarity Index), which is an index used to measure the visual quality similarity between two images (or video frames). The user rating is the rating given by the user in the preset quality of experience. Therefore, based on the picture quality index, latency, and user rating, the state space of the reinforcement learning model can be constructed. The main goal of the state space is to abstract the dynamically changing system environment into a multi-dimensional state vector that can be understood by reinforcement learning, providing a decision-making basis for the model. Specifically in implementation, each dimension index can be scaled to a unified range (for example, [0, 1]), so as to construct the state vector in the state space based on a unified dimension.

[0080] Step S302: Construct a reward function based on the picture quality index and the preset quality of experience.

[0081] In the embodiment of the present application, the reward function of the reinforcement learning model is mainly constructed based on the picture quality index and the preset quality of experience. Exemplarily, the reward function can be expressed as:

[0082] Reward = 0.7×QoE + 0.3×SSIM,

[0083] where QoE is the quality of experience, and SSIM is the picture quality index. The core elements of QoE include the above-mentioned latency picture quality index, latency, and user rating.

[0084] Step S303: Adjust the codec parameters according to the state space through the reinforcement learning model to optimize the reward function.

[0085] In the embodiments of the present application, by interacting with the environment to learn the optimal mapping strategy from state to action, the encoding and decoding parameters are dynamically adjusted. The reinforcement learning model outputs the corresponding action probability based on the state vector in the state space. Exemplarily, the encoding and decoding parameters include: bitrate change, Group of Pictures (GOP) length, and quantization parameter (QP). The GOP length refers to the number of frames within a GOP, which usually includes one key frame (I-frame) and several predicted frames (P-frames and / or B-frames). GOP is a concept in video coding used to describe the structure and length of a group of frames. The GOP length has a significant impact on video compression efficiency, quality, and editability. The quantization parameter is a parameter used to control the image compression quality during video coding.

[0086] Based on the action probability of the reinforcement learning model, the reward function will feedback a reward value to the reinforcement learning model. The reinforcement learning model learns whether it has received a high score reward or a low score penalty based on this reward value, and adjusts the change trend of the model parameters and output results based on this reward value to make the final reward function reach the optimal.

[0087] Step S304: Determine the encoding and decoding parameters corresponding to the maximized reward function as the optimal encoding and decoding parameters.

[0088] The encoding and decoding parameters corresponding to the maximized reward function determined through the above process are the optimal encoding and decoding parameters.

[0089] In some alternative embodiments of this embodiment, the execution subject of the above method for realizing adaptive encoding and decoding optimization can determine the optimal encoding and decoding algorithm to be used through an algorithm decision tree, and optimize the encoding and decoding parameters of the optimal encoding and decoding algorithm through the above process. The algorithm decision tree is used to represent the mapping relationship between conditional branches and encoding and decoding algorithms. The process implemented by the algorithm decision tree is as Figure 4 shown.

[0090] Step 401: Start.

[0091] This step is the starting point of the decision-making process, indicating the start of the algorithm selection process.

[0092] Step 402: Determine whether the device supports H.265.

[0093] This step aims to check whether the device supports the H.265 encoding and decoding algorithm. Specifically, it can be determined whether the device supports H.265 according to the device capability matrix. If so, execute Step 403; otherwise, execute Step 404.

[0094] Step 403: Select H.265. Then execute Step 409.

[0095] This step aims to determine that the device supports H.265. Then select the H.265 codec algorithm and use H.265 as the current codec algorithm.

[0096] Step 404: Determine whether the device supports SVC.

[0097] This step aims to determine that the device does not support H.265. Then check whether it supports SVC. Determine whether the device supports SVC according to the device capability matrix. If so, execute step 405; otherwise, execute step 406.

[0098] Step 405: Select SVC. Then execute step 409.

[0099] This step aims to determine that the device supports SVC. Then select the SVC codec algorithm and use SVC as the current codec algorithm.

[0100] Step 406: Determine whether the device supports AV1.

[0101] This step aims to determine that the device does not support H.265 and SVC. Then check whether it supports AV1. Determine whether the device supports AV1 according to the device capability matrix. If so, execute step 407; otherwise, execute step 408.

[0102] Step 407: Select AV1.

[0103] This step aims to determine that the device supports AV1. Then select the AV1 codec algorithm and use AV1 as the current codec algorithm. Then execute step 409.

[0104] Step 408: Select H.264 Baseline Profile.

[0105] This step aims to determine that the device does not support H.265, SVC, and AV1. Then select H.264 Baseline Profile and use H.264 Baseline Profile as the current codec algorithm.

[0106] Step 409: Check the network status.

[0107] This step aims to adjust the codec parameters according to the current network status (such as bandwidth, latency, etc.) to ensure that the codec algorithm can operate efficiently under the current network conditions. If the network status is good, execute step 410; otherwise, execute step 411.

[0108] Step 410: Output the optimized audio and video stream.

[0109] This step aims to determine that the network status is good. Then output the optimized audio and video stream and transfer the optimized audio and video stream to the user device.

[0110] Step 411: Adjust the bit rate and re - select the algorithm.

[0111] This step aims to determine the network state fluctuation, then adjust the bit rate and re - enter the algorithm selection process, dynamically adjusting the encoding and decoding parameters to adapt to network changes.

[0112] Step 412: End.

[0113] This step aims to represent the end point of the decision - making process, indicating the end of the algorithm selection process.

[0114] To deepen the understanding and clarify the advantages and effects of the adaptive encoding and decoding optimization method provided in this embodiment, the present disclosure also provides an illustration in combination with specific comparison examples.

[0115] First, it can clearly achieve delay optimization.

[0116] In the bandwidth fluctuation scenario, the end - to - end delay is reduced from 142 ms in the traditional scheme to 68 ms (a decrease of 52%). The calculation formula is:

[0117] ΔT = T static - T dynamic = 142 ms - 68 ms = 74 ms.

[0118] Among them, T static represents the end - to - end delay under the traditional fixed strategy. In the embodiment of the present application, the delay of the traditional scheme is 142 milliseconds (ms).

[0119] T dynamic represents the end - to - end delay after adopting the adaptive encoding and decoding optimization method provided in this embodiment. In the embodiment of the present application, the optimized delay is 68 milliseconds (ms).

[0120] ΔT represents the amount of delay reduction, that is, the difference between the optimized delay and the traditional delay.

[0121] Percentage of decrease: The percentage of delay reduction can be calculated by the following formula:

[0122]

[0123] In the embodiment of the present application, the percentage of decrease is:

[0124]

[0125] Secondly, it can clearly achieve image quality improvement.

[0126] The SSIM is increased from 0.85 to 0.95. The calculation formula is:

[0127] ΔS = S new - Sold = 0.95 - 0.85 = 0.10

[0128] Among them, S old represents the Structural Similarity Index (SSIM) under the traditional fixed strategy. In the embodiments of the present application, the SSIM value of the traditional scheme is 0.85.

[0129] S new represents the Structural Similarity Index (SSIM) after adopting the adaptive codec optimization method provided in this embodiment. In the embodiments of the present application, the optimized SSIM value is 0.95.

[0130] ΔS represents the amount of image quality improvement, that is, the difference between the optimized SSIM value and the traditional SSIM value.

[0131] Improvement percentage: The percentage of image quality improvement can be calculated by the following formula:

[0132]

[0133] In this case, the improvement percentage is:

[0134]

[0135] Actual improvement effect description

[0136] Through theoretical analysis and simulation verification, the adaptive codec optimization method provided in this embodiment can achieve the following technical effects:

[0137] 1. Latency optimization:

[0138] In the bandwidth fluctuation scenario, the end-to-end latency is reduced from 142 ms of the traditional scheme to 68 ms, and the latency reduction amount is:

[0139] ΔT = T static - T dynamic = 142 ms - 68 ms = 74 ms.

[0140] The percentage of latency reduction is:

[0141]

[0142] It can be seen that in the embodiments of the present application, by dynamically adjusting the bit rate and optimizing the codec algorithm, the end-to-end latency is significantly reduced, improving the user experience.

[0143] 2. Image quality improvement:

[0144] The image quality index (SSIM) is improved from 0.85 to 0.95, and the amount of image quality improvement is:

[0145] ΔS = S new-S old = 0.95 - 0.85 = 0.10

[0146] The percentage of image quality improvement is:

[0147]

[0148] It can be seen that in the embodiments of this application, by optimizing the encoding and decoding algorithms and dynamically adjusting the bit rate, the image quality is significantly improved, especially the visual effect in key areas (such as ROI).

[0149] 3. Device compatibility: Support smooth adaptation from low-end devices (Snapdragon 4 series, 2GB RAM) to high-end devices (M1 chip, 8GB RAM).

[0150] Specific application examples:

[0151] Application Example 1: Optimization in weak network environment

[0152] · Scenario: The user switches from WiFi (20Mbps) to 4G (2Mbps).

[0153] · Steps:

[0154] 1. The network probe detects a sudden drop in bandwidth, and ARIMA predicts BW_est = 2.5Mbps.

[0155] 2. The device capability matrix shows that the maximum decoding capability of the GPU is 1080p@30fps (GPU_dec_max = 4Mbps).

[0156] 3. Calculate Q_max = min(2.5×0.8 / 30, 4) = 0.067Mbps, and dynamically adjust the bit rate to 670Kbps.

[0157] 4. The decision tree selects H.265 + ROI encoding, and the bit rate in the central area is increased to 804Kbps.

[0158] Application Example 2: Multi-device collaboration

[0159] · Scenario: Cross-screen collaboration between a mobile phone (H.265 hardware decoding) and a tablet (AV1 hardware decoding).

[0160] · Steps:

[0161] 1. The device fingerprint module identifies that the main screen supports H.265 and the secondary screen supports AV1.

[0162] 2. The cloud generates dual bitstreams (H.265@3Mbps + AV1@3.6Mbps).

[0163] 3. The edge node dynamically distributes the corresponding versions, and the end-to-end delay < 80ms.

[0164] Experimental data

[0165] 1. Performance comparison (with the WebRTC solution):

[0166] Average latency: 68 ms (this embodiment) vs. 142 ms (WebRTC), with a 52% improvement.

[0167] SSIM: 0.95 (this embodiment) vs. 0.85 (WebRTC), with a 12% improvement.

[0168] Frame rate of low-end devices: 24 fps (this embodiment) vs. 15 fps (WebRTC), with a 60% improvement.

[0169] Further reference Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an adaptive codec optimization device. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0170] As Figure 5 shown, the adaptive codec optimization device 500 in this embodiment may include: a network probe module 501, a device fingerprint module 502, a bitrate control engine 503, an algorithm decision module 504, and an audio-video output module 505. Among them, the network probe module 501 is configured to obtain network status information, perform bandwidth prediction based on the network status information, and obtain a bandwidth prediction result; the device fingerprint module 502 is configured to obtain hardware capability information of a target device through a hardware abstraction layer interface, and construct a device capability matrix based on the hardware capability information; the bitrate control engine 503 is configured to calculate the maximum available bitrate based on the bandwidth prediction result and the device capability matrix; the algorithm decision module 504 is configured to determine an optimal codec algorithm and optimal codec parameters based on a preset user experience quality and the maximum available bitrate; the audio-video output module 505 is configured to optimize the target audio-video based on the optimal codec algorithm and optimal codec parameters to obtain an optimized audio-video stream.

[0171] In the embodiment of the present application, in the adaptive codec optimization device 500, the specific processing of the network probe module 501, the device fingerprint module 502, the bitrate control engine 503, the algorithm decision module 504, and the audio-video output module 505 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201-205 in the corresponding embodiment, which will not be elaborated here.

[0172] This embodiment exists as a device embodiment corresponding to the above method embodiment. The adaptive codec optimization device provided in this embodiment can effectively handle the situation of bandwidth fluctuation through the above process, maintain high efficiency and high quality in the processing of audio and video streams, significantly reduce the latency of audio and video stream transmission, and improve the user experience.

[0173] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the adaptive codec optimization method described in any of the above embodiments.

[0174] According to an embodiment of the present disclosure, the present disclosure also provides a readable storage medium, which stores computer instructions for enabling a computer to implement the adaptive codec optimization method described in any of the above embodiments when executed.

[0175] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which can implement the adaptive codec optimization method described in any of the above embodiments when executed by a processor.

[0176] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0177] Please refer to Figure 6 , Figure 6 is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 6As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system).

[0178] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0179] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0180] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 20 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0181] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above-mentioned types of memories.

[0182] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0183] Embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading via a network from an original storage in a remote storage medium or a non-transitory machine-readable storage medium and will be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0184] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An adaptive codec optimization method, comprising: Obtaining network status information, predicting the bandwidth based on the network status information to obtain a bandwidth prediction result; Obtaining the hardware capability information of the target device through the hardware abstraction layer interface, and constructing a device capability matrix based on the hardware capability information; Calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix; Determining an optimal codec algorithm and optimal codec parameters based on a preset quality of experience for users and the maximum available bitrate; Optimizing the target audio and video based on the optimal codec algorithm and optimal codec parameters to obtain an optimized audio and video stream.

2. The adaptive encoding and decoding optimization method according to claim 1, wherein The obtaining of the network status information includes: Obtaining the network status information through a dual-channel of active detection and passive analysis.

3. The adaptive encoding and decoding optimization method according to claim 2, wherein, The network status information includes: bandwidth, round-trip time, and jitter rate. The obtaining of the network status information through a dual-channel of active detection and passive analysis includes: Measuring the round-trip time based on an active detection message, and calculating the jitter rate based on the round-trip time; Parsing the TCP ACK packet to calculate the bandwidth.

4. The adaptive encoding and decoding optimization method according to claim 1, wherein The network status information includes: bandwidth, round-trip time, and jitter rate. The predicting of the bandwidth based on the network status information to obtain a bandwidth prediction result includes: Using a time series model to perform bandwidth prediction based on the bandwidth, round-trip time, and jitter rate to obtain the bandwidth prediction result.

5. The adaptive encoding and decoding optimization method according to claim 1, wherein The device capability matrix includes: GPU decoding throughput, NPU computing power, and memory bandwidth. The calculating of the maximum available bitrate based on the bandwidth prediction result and the device capability matrix includes: Determining a target frame rate and the maximum GPU decoding capability based on the GPU decoding throughput, NPU computing power, and memory bandwidth; Calculating the maximum available bitrate based on the bandwidth prediction result, the target frame rate, and the maximum GPU decoding capability.

6. The adaptive encoding and decoding optimization method according to claim 1, characterized in that It further includes: In response to the maximum available bitrate not being within a preset threshold range, adjusting the device parameters in the device capability matrix, and returning to execute the steps of obtaining the network status information, predicting the bandwidth based on the network status information to obtain a bandwidth prediction result, to the step of calculating the maximum available bitrate based on the bandwidth prediction result and the device capability matrix until the maximum available bitrate is within the preset threshold range.

7. The adaptive codec optimization method according to any one of claims 1-6, characterized in that Determining the optimal codec parameters based on a preset quality of experience for users and the maximum available bitrate includes: Constructing a state space based on image quality metrics, latency, and user ratings; Constructing a reward function based on the image quality metrics and a preset quality of experience for users; Adjusting the codec parameters according to the state space through a reinforcement learning model to optimize the reward function; Determining the codec parameters corresponding to the maximized reward function as the optimal codec parameters.

8. An adaptive encoding and decoding optimization device, characterized in that It includes: A network probe module configured to obtain network status information, predict the bandwidth based on the network status information to obtain a bandwidth prediction result; A device fingerprint module configured to obtain the hardware capability information of the target device through the hardware abstraction layer interface, and construct a device capability matrix based on the hardware capability information; A bitrate control engine, configured to calculate a maximum available bitrate based on the bandwidth prediction result and the device capability matrix; An algorithm decision module, configured to determine an optimal codec algorithm and optimal codec parameters based on a preset quality of user experience and the maximum available bitrate; An audio and video output module, configured to optimize a target audio and video based on the optimal codec algorithm and optimal codec parameters to obtain an optimized audio and video stream.

9. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the adaptive codec optimization method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the adaptive codec optimization method according to any one of claims 1-7.