Image encoding method, apparatus and device for video call, storage medium, and product

By dynamically adjusting the bitrate during video calls, increasing the minimum and maximum bitrates of preset levels, and determining the encoding resolution, the problem of decreased call quality caused by maintaining a high image encoding resolution in scenarios with average or fluctuating network conditions is solved, resulting in smoother video playback and a better user experience.

WO2026108772A1PCT designated stage Publication Date: 2026-05-28BIGO TECH PTE LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIGO TECH PTE LTD
Filing Date
2025-11-17
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

In scenarios with poor network conditions or significant network fluctuations, the image encoding resolution remains at a high level, making it difficult for the video playback end to trigger video super-resolution processing, resulting in a decrease in call quality.

Method used

By determining the detection bandwidth, video playback frame rate, and video packet loss rate for real-time video calls with the second terminal, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table. Based on the detection bandwidth and the target code table, the encoding resolution is determined, and video encoding is performed to trigger the downslicing of the image encoding resolution, thereby reducing encoding and decoding time and improving the smoothness of video playback.

Benefits of technology

In scenarios with average network conditions or significant network fluctuations, image encoding is more likely to trigger resolution downscaling and less likely to trigger resolution upscaling, thus improving call quality and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025135396_28052026_PF_FP_ABST
    Figure CN2025135396_28052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image encoding method, apparatus and device for a video call, a storage medium, and a product. The technical solution provided in the embodiments of the present application comprises: determining an estimated bandwidth for performing a real-time video call with a second terminal, a video playback frame rate of the second terminal, and a video packet loss rate of a first terminal; when the video playback frame rate and the video packet loss rate meet a bitrate table adjustment condition, increasing the minimum bitrate and the maximum bitrate of a preset level in a preset bitrate table, to obtain a target bitrate table; on the basis of the estimated bandwidth and the target bitrate table, determining an encoding resolution; and on the basis of the encoding resolution, performing video encoding on images to be encoded in the video call to obtain encoded video information. In scenarios with average network conditions or significant network fluctuations, increasing the minimum and maximum bitrates of a preset level in the preset bitrate table makes it easier for video encoding to trigger downscaling of the image encoding resolution, and also makes it easier for the second terminal to trigger video super‑resolution processing, thereby improving call quality and enhancing user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Image encoding methods, devices, equipment, storage media, and products for video calls

[0001] This application claims priority to Chinese Patent Application No. 202411684112.3, filed on November 22, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing technology, and in particular to image encoding methods, apparatus, devices, storage media, and products for video calls. Background Technology

[0003] Real-Time Communication (RTC) is a technology that enables real-time audio and video communication over a network. With the rise and development of the Internet era, RTC has addressed people's urgent need for real-time communication.

[0004] A key characteristic of real-time communication scenarios is low latency. Users have higher demands for the real-time performance of audio and video and lower tolerance for latency and stuttering. Therefore, in addition to providing a clear and aesthetically pleasing video playback experience, real-time communication also needs to ensure smooth call quality. The quality of the encoded video during a user's call is closely related to network conditions. When the user's network conditions are good, the encoding resolution is high, resulting in better smoothness. When the user's network conditions are poor, super-resolution algorithms can be used to provide better image quality support.

[0005] However, in scenarios where network conditions are generally poor or fluctuate significantly, even if a user's network is not bad enough to trigger video super-resolution processing for video encoding, and the image encoding resolution is maintained at a relatively high level, making it difficult for the video playback end to trigger video super-resolution processing, their calls may already be affected by the network, resulting in stuttering and a decline in call quality. Summary of the Invention

[0006] This application provides an image encoding method, apparatus, device, storage medium, and product for video calls to solve the technical problem in related technologies where, in scenarios with average network conditions or large network fluctuations, the image encoding resolution is maintained at a high level, making it difficult for the video playback end to trigger video super-resolution processing, thus leading to a decrease in call quality. This allows for easier triggering of image encoding resolution downscaling and video super-resolution processing in scenarios with average network conditions or large network fluctuations, thereby improving call quality and enhancing user experience.

[0007] In a first aspect, embodiments of this application provide an image encoding method for video calls applied to a first terminal, comprising:

[0008] Determine the detection bandwidth for real-time video calls with the second terminal;

[0009] Determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal;

[0010] When the video playback frame rate and the video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table.

[0011] The encoding resolution is determined based on the detection bandwidth and the target code table, and the image to be encoded in the video call is encoded based on the encoding resolution to obtain encoded video information.

[0012] In a second aspect, embodiments of this application provide an image encoding device for video calls applied to a first terminal, including a bandwidth detection module, an information acquisition module, a code table adjustment module, and a video encoding module, wherein:

[0013] The bandwidth detection module is configured to determine the detection bandwidth for real-time video calls with the second terminal;

[0014] The information acquisition module is configured to determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal.

[0015] The code table adjustment module is configured to increase the minimum and maximum bit rates of a preset level in the preset code table to obtain a target code table when the video playback frame rate and the video packet loss rate meet the code table adjustment conditions.

[0016] The video encoding module is configured to determine the encoding resolution based on the detection bandwidth and the target code table, and to perform video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information.

[0017] In a third aspect, embodiments of this application provide an image encoding device for video calls, including: a memory and one or more processors;

[0018] The memory is used to store one or more programs;

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the image encoding method for video calls as described in the first aspect.

[0020] In a fourth aspect, embodiments of this application provide a non-volatile storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform the image encoding method for video calls as described in the first aspect.

[0021] In a fifth aspect, embodiments of this application provide a computer program product comprising a computer program stored in a computer-readable storage medium, wherein at least one processor of a device reads from and executes the computer program from the computer-readable storage medium, causing the device to perform the image encoding method for video calls as described in the first aspect.

[0022] This application embodiment determines the detection bandwidth for real-time video calls with a second terminal, the video playback frame rate of the second terminal, and the video packet loss rate of the first terminal. When the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of preset levels in the preset code table are increased to obtain the target code table. The encoding resolution is determined based on the detection bandwidth and the target code table, and the image to be encoded in the video call is encoded based on the encoding resolution to obtain encoded video information. In scenarios with general network conditions or large network fluctuations, increasing the minimum and maximum bit rates of preset levels in the preset code table makes it easier to trigger the down-cutting of image encoding resolution and more difficult to trigger the up-cutting of image encoding resolution. This makes it easier to trigger video super-resolution processing on the second terminal, improving call quality and enhancing user experience. Attached Figure Description

[0023] Figure 1 is a flowchart of an image encoding method for video calls provided in an embodiment of this application;

[0024] Figure 2 is a flowchart of another image encoding method for video calls provided in an embodiment of this application;

[0025] Figure 3 is a schematic diagram of the structure of an image encoding device for video calls provided in an embodiment of this application;

[0026] Figure 4 is a schematic diagram of the structure of an image encoding device for video calls provided in an embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but additional steps not included in the drawings may also be present. The above processes can correspond to methods, functions, procedures, subroutines, subroutines, etc.

[0028] The image encoding method for video calls provided in this application can be applied to image encoding during real-time video calls. It aims to increase the minimum and maximum bitrates of preset levels in the preset code table under scenarios with average network conditions or large network fluctuations, making it easier to trigger the down-slicing of image encoding resolution and more difficult to trigger the up-slicing of image encoding resolution. This makes it easier to trigger video super-resolution processing on the second terminal, thereby improving call quality and enhancing user experience.

[0029] In existing real-time video calls, bandwidth probing algorithms are typically used to determine the available bandwidth (i.e., bitrate) for image encoding during the video call. Then, a corresponding resolution (width and height) is selected for image encoding based on a code table. When the probed bandwidth falls within the minimum and maximum bitrate range of a certain resolution within the code table, it indicates that the current resolution can be selected for video encoding. Due to the limited computing power of mobile devices commonly used for real-time video calls, super-resolution algorithms based on neural networks need to limit the image resolution of the input model for use on mobile devices. On the other hand, some downsampling operators offer a combined advantage in image quality and computing power when scaled to lower resolutions. When the user's network conditions are good, the image encoding resolution is high, resulting in smooth video playback. When the user's network conditions are poor, super-resolution and downsampling algorithms can be used to provide image quality support, effectively alleviating the poor user experience. However, due to the limited computing power of mobile devices commonly used in real-time call scenarios, many practical image pre- and post-processing operators cannot process images of all resolutions equally. Furthermore, a fixed code table lacks flexibility and cannot be adjusted in a timely manner to alleviate stuttering in scenarios with average or fluctuating network conditions. In scenarios where network conditions are average or fluctuate significantly, even if a user's network is not bad enough to trigger super-resolution encoding at a low resolution, their calls may still be affected by network issues and experience stuttering. The image encoding resolution remains at a high level, making it difficult for the video playback end to trigger video super-resolution processing, thus leading to a decrease in call quality.

[0030] Based on this, an image encoding method for video calls is provided according to an embodiment of this application. By adjusting the code table for pre- and post-processing of the video call at both ends, stuttering can be accurately detected, and the code table can be dynamically adjusted in real time to make the encoding resolution tend to be lower. The encoding is performed at a resolution that can trigger pre- and post-processing operators, so as to solve the technical problem that existing image encoding schemes maintain the image encoding resolution at a high level in scenarios with average network conditions or large network fluctuations, making it difficult for the video playback end to trigger video super-resolution processing, resulting in a decrease in call quality.

[0031] Figure 1 shows a flowchart of an image encoding method for video calls provided in an embodiment of this application. This image encoding method for video calls can be applied to a first terminal. The image encoding method for video calls provided in this embodiment of the application can be executed by an image encoding device for video calls. This image encoding device for video calls can be implemented in hardware and / or software and integrated into the image encoding equipment for video calls.

[0032] The following description uses a video call image encoding device to perform an image encoding method for a video call as an example. Referring to Figure 1, the image encoding method for this video call includes:

[0033] S110: Determine the detection bandwidth for real-time video calls with the second terminal.

[0034] This solution provides two terminal devices for video calls: a first terminal and a second terminal. The first terminal is used to capture (e.g., images captured via a camera or desktop image information), encode the images to generate encoded video information, and send the encoded video information to the second terminal (e.g., by forwarding the encoded video information to the second terminal via a preset server). The second terminal device is used to receive, decode, and play the encoded video information. One terminal device can simultaneously function as both the first and second terminal; for example, while capturing, encoding, and sending encoded video information, one terminal device can simultaneously receive and decode encoded video information sent by other terminal devices.

[0035] For example, during a real-time video call between the first terminal and the second terminal, the bandwidth of the real-time video call between the first terminal and the second terminal can be detected in real time based on a preset bandwidth estimation algorithm (such as the BWE algorithm, Bandwidth Estimation), and the detected bandwidth can be determined based on the detection results.

[0036] S120: Determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal.

[0037] For example, the system can acquire the video playback frame rate of the video played on the second terminal in real time (the video playback frame rate can be collected and provided by the second terminal), and also calculate the video packet loss rate of the data (encoded video information) sent from the first terminal to the second terminal in real time. For instance, when the second terminal decodes and plays the received encoded video information, it can collect the frame rate information of the player playing the encoded video information in real time, use the current frame rate information as the video playback frame rate, and upload it to a preset server (e.g., a media server). The first terminal can then obtain the corresponding video playback frame rate of the second terminal from the preset server.

[0038] S130: If the video playback frame rate and video packet loss rate meet the code table adjustment conditions, increase the minimum and maximum bit rates of the preset levels in the preset code table to obtain the target code table.

[0039] For example, before encoding the image to be encoded in a video call, it is determined whether the video playback frame rate and video packet loss rate meet the code table adjustment conditions. The code table records multiple preset bitrates, each configured with a corresponding resolution, minimum bitrate, and maximum bitrate. The minimum and maximum bitrates within a bitrate form a bitrate range. When encoding an image, the bitrate range corresponding to the latest determined probe bandwidth is used, and encoding is performed based on a resolution within the same bitrate range.

[0040] In one embodiment, when the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of preset positions in the preset code table (i.e., the original code table) can be increased to obtain the target code table. This makes it easier to switch down to a preset position and more difficult to switch up to a preset position. The preset position can be one or more positions in the preset code table. At this time, the minimum and maximum bit rates of each preset position in the target code table are improved compared to the minimum and maximum bit rates of the same position in the preset code table. Encoding with the target code table, compared to encoding with the original preset code table, makes the overall image encoding resolution more difficult to switch up to a preset position and easier to switch down to a preset position. This effectively reduces the resolution of image encoding based on the target code table, makes it easier to trigger pre- and post-processing algorithms with resolution limitations, reduces encoding and decoding time, improves video playback smoothness, and can be combined with pre- and post-processing operators with resolution limitations to process image frames, effectively maintaining or even improving the image quality of real-time video calls.

[0041] In one embodiment, if the video playback frame rate and video packet loss rate do not meet the code table adjustment conditions, the current preset code table is maintained, and the video call image to be encoded is encoded according to the current preset code table to obtain encoded video information.

[0042] In one possible embodiment, if the newly determined video playback frame rate and video packet loss rate do not meet the code table adjustment conditions, and / or after a gear switch based on the target code table, image encoding can be restored to be based on a preset code table.

[0043] S140: Determine the encoding resolution based on the detection bandwidth and target code table, and perform video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information.

[0044] For example, after increasing the minimum and maximum bit rates of the preset levels in the preset code table to obtain the target code table, the corresponding bit rate range in the target code table is determined according to the probe bandwidth determined above (i.e., the probe bandwidth is between the corresponding minimum and maximum bit rates), and the resolution corresponding to the determined bit rate range is determined as the encoding resolution.

[0045] It needs to be explained that by increasing the minimum and maximum bitrates of the preset code table, increasing the minimum bitrate makes it easier to trigger a down-switching during the downward migration of encoding resolution, while increasing the maximum bitrate makes it more difficult to trigger an up-switching during the upward migration of encoding resolution. In scenarios with average network conditions or significant network fluctuations, image encoding based on the target code table is more likely to switch to a lower resolution. When encoding the image to be encoded based on a lower resolution, the first terminal has more computing power remaining, allowing for the use of downsampling algorithms and / or image preprocessing methods that require higher computing power but produce better encoding results, achieving higher encoding quality at low encoding resolution. By increasing the minimum and maximum bitrates of the preset code table to tend towards a smaller encoding resolution, the computing power consumption of the first terminal's encoding and decoding ends is reduced, increasing the encoding frame rate and playback frame rate, thereby improving call smoothness. It also allows switching from a critical resolution that cannot trigger pre- and post-processing operators to a resolution that can trigger them, improving the image quality of video calls.

[0046] After determining the encoding resolution, the image to be encoded in the video call can be encoded to obtain encoded video information based on the encoding resolution. After obtaining the encoded video information, it can be sent to a second terminal (for example, forming a video stream based on continuously generated encoded video information and sending it to the second terminal). After receiving the encoded video information, the second terminal can decode and play it. Because the minimum and maximum bitrates of the preset levels in the target code table have been increased, in scenarios with average network conditions or large network fluctuations, image encoding will more easily switch to a lower resolution level. When the second terminal decodes the encoded video information obtained from encoding based on the lower resolution level, it is more likely to trigger video super-resolution processing on the encoded video information, improving the smoothness of the video call without sacrificing image quality or even slightly improving the image quality, thus enhancing the user's video call experience.

[0047] As described above, by determining the detection bandwidth for real-time video calls with the second terminal, the video playback frame rate of the second terminal, and the video packet loss rate of the first terminal, when the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table. The encoding resolution is determined based on the detection bandwidth and the target code table, and the image to be encoded in the video call is encoded based on the encoding resolution to obtain encoded video information. In scenarios with average network conditions or large network fluctuations, the minimum and maximum bit rates of the preset levels in the preset code table are increased, making it easier to trigger the down-cutting of the image encoding resolution and more difficult to trigger the up-cutting of the image encoding resolution. This makes it easier to trigger video super-resolution processing on the second terminal, improving call quality and enhancing user experience.

[0048] Based on the above embodiments, Figure 2 shows a flowchart of another image encoding method for video calls provided by an embodiment of this application. This image encoding method for video calls is a specific embodiment of the above-described image encoding method for video calls. Referring to Figure 2, the image encoding method for video calls includes:

[0049] S210: Determine the detection bandwidth for real-time video calls with the second terminal.

[0050] S220: Determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal.

[0051] S230: When the video playback frame rate and video packet loss rate meet the code table adjustment conditions, increase the minimum bit rate of the preset level in the preset code table according to the preset minimum bit rate adjustment parameter, and increase the maximum bit rate of the preset level in the preset code table according to the preset maximum bit rate adjustment parameter to obtain the target code table.

[0052] In one possible embodiment, whether the video playback frame rate and video packet loss rate provided by this solution meet the code table adjustment conditions can be determined based on the comparison results of the video playback frame rate and video packet loss rate with the corresponding preset thresholds.

[0053] For example, the video playback frame rate is compared with a preset frame rate threshold, and the video packet loss rate is also compared with a preset packet loss threshold. Based on the comparison results, it is determined whether the video playback frame rate and video packet loss rate meet the code table adjustment conditions. For instance, if the video playback frame rate is less than the preset frame rate threshold and the video packet loss rate is greater than the preset packet loss threshold, it can be considered that the video playback frame rate and video packet loss rate meet the code table adjustment conditions; alternatively, if the video playback frame rate is less than the preset frame rate threshold or the video packet loss rate is greater than the preset packet loss threshold, it can be considered that the video playback frame rate and video packet loss rate meet the code table adjustment conditions. This solution accurately determines whether the code table adjustment conditions are met by comparing the video playback frame rate with the preset frame rate threshold and the video packet loss rate with the preset packet loss threshold, thus accurately determining the timing for code table adjustment, ensuring call quality, and improving user experience.

[0054] In one embodiment, after determining that the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table can be increased based on preset adjustment parameters to obtain the target code table. The adjustment parameters can be a proportional coefficient greater than 1 or a preset bit rate increase.

[0055] For example, the target code table can be obtained by increasing the minimum bitrate of a preset level in the preset code table according to a preset minimum bitrate adjustment parameter, and increasing the maximum bitrate of a preset level in the preset code table according to a preset maximum bitrate adjustment parameter. The preset minimum bitrate adjustment parameter and the preset maximum bitrate adjustment parameter can be proportional coefficients greater than 1, or preset bitrate increase amounts; the preset minimum bitrate adjustment parameter and the preset maximum bitrate adjustment parameter can be the same or different; the preset minimum bitrate adjustment parameter and the preset maximum bitrate adjustment parameter corresponding to different levels can be the same or different.

[0056] In one embodiment, when the preset minimum bitrate adjustment parameter and the preset maximum bitrate adjustment parameter are proportional coefficients greater than 1, the minimum bitrate and maximum bitrate of the preset level in the preset code table can be multiplied by the corresponding preset minimum bitrate adjustment parameter and preset maximum bitrate adjustment parameter respectively to obtain the minimum bitrate and maximum bitrate after the preset level is increased.

[0057] In one embodiment, when the preset minimum bitrate adjustment parameter and the preset maximum bitrate adjustment parameter are the bitrate increase range, the minimum bitrate and the maximum bitrate of the preset level in the preset code table can be respectively added to the corresponding preset minimum bitrate adjustment parameter and preset maximum bitrate adjustment parameter to obtain the minimum bitrate and the maximum bitrate after the preset level is increased.

[0058] This solution increases the minimum bitrate of the preset code table based on the preset minimum bitrate adjustment parameter, and increases the maximum bitrate of the preset code table based on the preset maximum bitrate adjustment parameter. This allows for quick and accurate adjustment of the preset code table to obtain the target code table, ensuring call quality and improving user experience.

[0059] In one embodiment, when adjusting the preset code table using the applied bitrate adjustment parameters, the level corresponding to the critical resolution of the pre- and post-processing operators can be used as the preset level. This makes it easier to cut down between critical resolutions and more difficult to cut up. By adjusting the code table of the pre- and post-processing operations in conjunction with both ends, the stuttering of video calls can be accurately detected. Furthermore, the preset code table can be dynamically adjusted in real time, causing the encoding resolution to tend to cut down. Encoding with a resolution that can trigger the pre- and post-processing operators can not only alleviate the stuttering of video calls to a certain extent, but also maintain or even improve the image quality by triggering the pre- and post-processing operators more flexibly.

[0060] The following table provides one example of a preset code table:

[0061] For example, the resolution corresponding to gear 1 is R1, and its minimum and maximum bitrates are BR1min and BR1max respectively. Among them, the resolutions corresponding to gears 1-5 increase sequentially. Assume that the preset gears are R2 and R3, and the preset minimum bitrate adjustment parameter Fmin and the preset maximum bitrate adjustment parameter Fmax are proportionality coefficients greater than 1. Then, the target code table obtained by increasing the minimum and maximum bitrates of the preset gears in the preset code table based on the adjustment parameter is shown in the following table:

[0062] Among them, the minimum and maximum bitrates of the preset gear R2 are BR2min*Fmin and BR2max*Fmax respectively, and the minimum and maximum bitrates of the preset gear R3 are BR3min*Fmin and BR3max*Fmax respectively. The minimum and maximum bitrates of other gears remain unchanged.

[0063] It can be seen that when the minimum bitrate adjustment coefficient Fmin ∈ (1, +∞), the minimum bitrate of the preset gear becomes larger. Taking gear 2 as an example, before, the detection bandwidth BR < BR2min was required to downshift to gear 1 and encode using a smaller resolution R1. However, after the minimum bitrate of the code table becomes larger, it is possible to downshift to gear 1 when and only when the detection bandwidth BR < BR2min*Fmin. And the encoding bitrates within the range of detection bandwidth BR ∈ (BR2min, BR2min*Fmin), which should not have been downshifted under the original logic, will start to downshift and then encode using resolution R1. When the maximum bitrate adjustment coefficient Fmax ∈ (1, +∞), the maximum bitrate of the preset gear becomes larger. Taking gear 2 as an example, before, the detection bandwidth BR > BR2max was required to upshift to gear 3 and encode using a larger resolution R3. However, after the maximum bitrate of the code table becomes larger, it is possible to upshift to gear 3 when and only when the detection bandwidth BR > BR2max*Fmax. And the encoding bitrates within the range of BR ∈ (BR2max, BR2max*Fmax), which should have been upshifted under the original logic, will continue to encode using resolution R2. It can be seen that in this solution, by increasing the minimum bitrate of the preset code table, it is easier to trigger downshifting during the downward migration of the encoding resolution, and by increasing the maximum bitrate of the preset code table, it is more difficult to trigger upshifting during the upward migration of the encoding resolution. On the one hand, this can reduce the computing power consumption at the sending end for encoding and at the decoding end for decoding, improve the encoding frame rate and the playback frame rate, thereby enhancing the call fluency. On the other hand, it can switch the critical resolution that cannot trigger the pre- and post-processing operators to a resolution that can trigger the pre- and post-processing operators, improve the image quality of the video call, and improve the fluency of the video call without sacrificing image quality or even slightly improving the image quality, thus enhancing the user experience.

[0064] S240: Determine the encoding resolution based on the detection bandwidth and target code table, and perform video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information.

[0065] In one possible embodiment, the image encoding method for video calls provided by this solution performs video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information. This can be:

[0066] S2401: When the encoding resolution meets the processing computing power requirements, the first downsampling processing algorithm and the image preprocessing algorithm are used to perform first downsampling processing and image preprocessing on the image to be encoded in the video call, and the image to be encoded after first downsampling processing and image preprocessing is then used for video encoding to obtain encoded video information.

[0067] S2402: When the encoding resolution does not meet the processing computing power requirements, the image to be encoded in the video call is subjected to second downsampling processing based on the second downsampling processing algorithm, and the image to be encoded after the second downsampling processing is then video encoded to obtain encoded video information.

[0068] For example, after determining the encoding resolution in the target code table based on the detection bandwidth, it can be determined whether the encoding resolution meets the preset processing power conditions. Meeting the preset processing power conditions means that the first terminal has surplus computing power to perform more data processing when encoding the image based on that encoding resolution. Optionally, the processing power conditions can be a fixed configuration or dynamically configured based on different terminal models or processor models.

[0069] In one embodiment, whether the encoding resolution meets the processing power requirements can be determined based on a comparison between the encoding resolution and a first preset resolution. For example, if the encoding resolution is less than or equal to the preset resolution, it can be considered that the encoding resolution meets the processing power requirements. For instance, the first preset resolution includes a first preset width and a first preset height. The encoding resolution's width and the first preset width, as well as its height and the first preset height, are compared respectively. If the encoding resolution's width is less than or equal to the first preset width, and / or its height is less than or equal to the first preset height, the encoding resolution is considered to meet the processing power requirements.

[0070] For example, assuming the width and height of the encoded resolution are Ws and Hs respectively, and the first preset width and height of the first preset resolution are w1 and h1 respectively, when Ws ≤ w1 and Hs ≤ h1, the current encoded resolution is considered sufficiently small. The first terminal can use a downsampling processing algorithm with higher computational requirements but better performance, as well as an image preprocessing algorithm, for image processing. In this case, the encoded resolution meets the processing computational requirements. If Ws is greater than w1 or Hs > h1, the current encoded resolution is considered too large, and the first terminal's load cannot support high-resolution image processing. A faster but less effective downsampling processing algorithm can be used, and image preprocessing is abandoned. In this case, the encoded resolution does not meet the processing computational requirements.

[0071] This solution accurately determines whether the encoding resolution meets the processing power requirements by comparing the encoding resolution width with the first preset width, and the encoding resolution height with the first preset height. It then adopts a more suitable image processing method to effectively balance image processing efficiency and image processing quality, thereby improving the user's video call experience.

[0072] In one embodiment, when the encoding resolution meets the processing computing power requirements, the image to be encoded in the video call can be processed by a first downsampling processing algorithm (which may be a downsampling processing algorithm with high computing power requirements but good effect, such as bicubic interpolation, lanczos interpolation, etc.) and an image preprocessing algorithm (such as noise reduction, contrast enhancement, color transformation, etc.). The image to be encoded after the first downsampling processing and image preprocessing is then video encoded to obtain encoded video information.

[0073] In one embodiment, when the encoding resolution does not meet the processing power requirements, the image to be encoded in the video call can be downsampled based on a second downsampling processing algorithm (which may be a downsampling processing algorithm with low computing power requirements but high processing speed, such as the bilinear downsampling algorithm). The image to be encoded after the second downsampling processing is then video encoded to obtain encoded video information. This eliminates the need for image preprocessing of the image to be encoded, thus ensuring image encoding efficiency.

[0074] This solution performs a first downsampling and image preprocessing on the video call image to be encoded when the encoding resolution meets the processing power requirements. Then, it performs video encoding on the image to be encoded after the first downsampling and image preprocessing to obtain encoded video information. When the encoding resolution does not meet the processing power requirements, it performs a second downsampling on the video call image to be encoded. Then, it performs video encoding on the image to be encoded after the second downsampling to obtain encoded video information. This effectively balances image processing efficiency and image processing quality, improving the user's video call experience.

[0075] In one possible embodiment, the video call image encoding method provided by this solution, after encoding the image to be encoded in the video call based on the encoding resolution to obtain encoded video information, can also send the encoded video information to a second terminal, so that the second terminal can decode the encoded video information to obtain a decoded video frame, and if the decoded video frame meets the preset post-processing conditions, perform image post-processing on the decoded video frame based on the preset image post-processing algorithm, and play the image post-processed decoded video frame.

[0076] For example, after encoding the image to be encoded in a video call to obtain encoded video information, the encoded video information can be sent to a second terminal in the real-time video call. Upon receiving the encoded video information, the second terminal decodes it to obtain a decoded video frame and determines whether the decoded video frame meets preset post-processing conditions. Meeting the preset post-processing conditions means that the second terminal has excess computing power to perform further data post-processing on the decoded video frame. Optionally, the preset post-processing conditions can be a fixed configuration or dynamically configured based on different terminal models or processor models.

[0077] In one embodiment, the resolution of the decoded video frame can be determined based on a comparison between the decoded video frame's resolution and a second preset resolution to determine whether the resolution meets preset post-processing conditions. For example, if the decoded video frame's resolution is less than or equal to the second preset resolution, it can be considered that the decoded video frame meets the preset post-processing conditions. For instance, the second preset resolution includes a second preset width and a second preset height. The decoded video frame's resolution width and second preset width, as well as its resolution height and second preset height, are compared. If the decoded video frame's resolution width is less than or equal to the second preset width, and / or its resolution height is less than or equal to the second preset height, the decoded video frame is considered to meet the preset post-processing conditions.

[0078] For example, assuming the resolution width and resolution height of the decoded video frame are Wr and Hr respectively, and the second preset width and second preset height of the second preset resolution are w2 and h2 respectively, when Wr ≤ w2 and Hr ≤ h2, the resolution of the current decoded video frame is considered sufficiently small, and the second terminal can use a post-processing algorithm with higher computational requirements but better results for image post-processing. In this case, the decoded video frame can be determined to meet the preset post-processing conditions. If Wr is greater than w2 or Hr > h2, the resolution of the current decoded video frame is considered too large, and the second terminal's load cannot support high-resolution image post-processing, so image post-processing can be abandoned. In this case, the decoded video frame can be determined to not meet the preset post-processing conditions.

[0079] This solution accurately determines whether the resolution of the decoded video frame meets the preset post-processing conditions by comparing the resolution width and the second preset width of the decoded video frame, as well as the resolution height and the second preset height of the decoded video frame. It then adopts a more suitable image processing method, effectively balancing image processing efficiency and image processing quality, and improving the user's video call experience.

[0080] In one embodiment, if a decoded video frame does not meet the preset post-processing conditions, the decoded video frame can be rendered and played directly without processing, ensuring timely playback. In another embodiment, if a decoded video frame meets the preset post-processing conditions, it is post-processed using a preset image post-processing algorithm (e.g., super-resolution processing based on a preset image super-resolution algorithm, super-resolution processing based on a neural network, etc.), and then the post-processed decoded video frame is rendered and played. This solution improves video call quality and enhances the user's video call experience by post-processing the decoded video frame using a preset image post-processing algorithm when the preset post-processing conditions are met.

[0081] As described above, by determining the detection bandwidth for real-time video calls with the second terminal, the video playback frame rate of the second terminal, and the video packet loss rate of the first terminal, when the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table. The encoding resolution is determined based on the detection bandwidth and the target code table, and video encoding is performed on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information. In scenarios with average network conditions or significant network fluctuations, increasing the minimum and maximum bit rates of the preset levels in the preset code table makes it easier to trigger down-slicing of the image encoding resolution and less likely to trigger up-slicing. This makes it easier to trigger video super-resolution processing on the second terminal, improving call quality and enhancing user experience. Simultaneously, by increasing the minimum bit rate of the preset levels in the preset code table according to the preset minimum bit rate adjustment parameter, and increasing the maximum bit rate of the preset levels in the preset code table according to the preset maximum bit rate adjustment parameter, the preset code table can be quickly and accurately adjusted to obtain the target code table, ensuring call quality and improving user experience.

[0082] Figure 3 is a schematic diagram of the structure of an image encoding device for video calls provided in an embodiment of this application. Referring to Figure 3, the image encoding device for video calls can be applied to a first terminal, and the image encoding device for video calls includes a bandwidth detection module 31, an information acquisition module 32, a code table adjustment module 33, and a video encoding module 34.

[0083] Among them, the bandwidth detection module 31 is configured to determine the detection bandwidth for real-time video calls with the second terminal;

[0084] Information acquisition module 32 is configured to determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal;

[0085] The code table adjustment module 33 is configured to increase the minimum and maximum bit rates of the preset levels in the preset code table to obtain the target code table when the video playback frame rate and video packet loss rate meet the code table adjustment conditions.

[0086] The video encoding module 34 is configured to determine the encoding resolution based on the detection bandwidth and the target code table, and to perform video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information.

[0087] As described above, by determining the detection bandwidth for real-time video calls with the second terminal, the video playback frame rate of the second terminal, and the video packet loss rate of the first terminal, when the video playback frame rate and video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table. The encoding resolution is determined based on the detection bandwidth and the target code table, and the image to be encoded in the video call is encoded based on the encoding resolution to obtain encoded video information. In scenarios with average network conditions or large network fluctuations, the minimum and maximum bit rates of the preset levels in the preset code table are increased, making it easier to trigger the down-cutting of the image encoding resolution and more difficult to trigger the up-cutting of the image encoding resolution. This makes it easier to trigger video super-resolution processing on the second terminal, improving call quality and enhancing user experience.

[0088] In one possible embodiment, the video playback frame rate and the video packet loss rate satisfy the code table adjustment condition, which can be that the video playback frame rate is less than a preset frame rate threshold and the video packet loss rate is greater than a preset packet loss threshold.

[0089] In one possible embodiment, the code table adjustment module 33 increases the minimum and maximum bit rates of the preset levels in the preset code table, configured as follows:

[0090] Increase the minimum bitrate of the preset level in the preset code table according to the preset minimum bitrate adjustment parameter;

[0091] Increase the maximum bitrate of the preset level in the preset code table according to the preset maximum bitrate adjustment parameter.

[0092] In one possible embodiment, the video encoding module 34 performs video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information, configured as follows:

[0093] When the encoding resolution meets the processing computing power requirements, the first downsampling processing algorithm and the image preprocessing algorithm are used to perform first downsampling processing and image preprocessing on the image to be encoded in the video call, and then the image to be encoded after first downsampling processing and image preprocessing is used to perform video encoding to obtain encoded video information.

[0094] When the encoding resolution does not meet the processing computing power requirements, the image to be encoded in the video call is subjected to second downsampling processing based on the second downsampling processing algorithm, and the image to be encoded after the second downsampling processing is then video encoded to obtain encoded video information.

[0095] In one possible embodiment, the encoding resolution satisfies the processing power condition by having a resolution width less than or equal to a first preset width and a resolution height less than or equal to a first preset height.

[0096] In one possible embodiment, the image encoding device for video calls further includes a data transmission module, which is configured to send encoded video information to a second terminal for the second terminal to decode the encoded video information to obtain a decoded video frame, and, if the decoded video frame meets preset post-processing conditions, to perform image post-processing on the decoded video frame based on a preset image post-processing algorithm and play the image post-processed decoded video frame.

[0097] In one possible embodiment, the decoded video frame satisfies a preset post-processing condition, which may be that the resolution width of the decoded video frame is less than or equal to a second preset width, and the resolution height of the decoded video frame is less than or equal to a second preset height.

[0098] It is worth noting that in the above-described embodiments of the image encoding device for video calls, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0099] This application also provides an image encoding device for video calls, which can integrate the image encoding apparatus for video calls provided in this application. Figure 4 is a schematic diagram of the structure of an image encoding device for video calls provided in this application. Referring to Figure 4, the image encoding device for video calls includes: an input device 43, an output device 44, a memory 42, and one or more processors 41; the memory 42 is used to store one or more programs; when one or more programs are executed by one or more processors 41, the one or more processors 41 implement the image encoding method for video calls as provided in the above embodiments. The image encoding apparatus, device, and computer for video calls provided above can be used to execute the image encoding method for video calls provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0100] This application also provides a non-volatile storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the video call image encoding method provided in the above embodiments. Of course, the computer-executable instructions provided in this application are not limited to the video call image encoding method provided above; they can also perform related operations in the video call image encoding method provided in any embodiment of this application. The video call image encoding apparatus, device, and storage medium provided in the above embodiments can execute the video call image encoding method provided in any embodiment of this application. Technical details not described in detail in the above embodiments can be found in the video call image encoding method provided in any embodiment of this application.

[0101] Based on the above embodiments, this application also provides a computer program product. The technical solution of this application, in essence or in other words, the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions to cause a computer device, mobile terminal, or processor therein to execute all or part of the steps of the video call image encoding method provided in the various embodiments of this application.

Claims

1. An image encoding method for video calls, wherein, Applied to the first terminal, including: Determine the detection bandwidth for real-time video calls with the second terminal; Determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal; When the video playback frame rate and the video packet loss rate meet the code table adjustment conditions, the minimum and maximum bit rates of the preset levels in the preset code table are increased to obtain the target code table. The encoding resolution is determined based on the detection bandwidth and the target code table, and the image to be encoded in the video call is encoded based on the encoding resolution to obtain encoded video information.

2. The image encoding method for video calls according to claim 1, wherein, The video playback frame rate and the video packet loss rate satisfy the code table adjustment conditions, including: The video playback frame rate is less than a preset frame rate threshold, and the video packet loss rate is greater than a preset packet loss threshold.

3. The image encoding method for video calls according to claim 1, wherein, The method of increasing the minimum and maximum bit rates of preset levels in the preset code table includes: Increase the minimum bitrate of the preset level in the preset code table according to the preset minimum bitrate adjustment parameter; Increase the maximum bitrate of the preset level in the preset code table according to the preset maximum bitrate adjustment parameter.

4. The image encoding method for video calls according to claim 1, wherein, The process of encoding the image to be encoded in the video call based on the encoding resolution to obtain encoded video information includes: When the encoding resolution meets the processing computing power requirements, the first downsampling processing algorithm and the image preprocessing algorithm are used to perform first downsampling processing and image preprocessing on the image to be encoded in the video call, and the image to be encoded after the first downsampling processing and image preprocessing is then video encoded to obtain encoded video information. If the encoding resolution does not meet the processing computing power requirements, the image to be encoded in the video call is subjected to a second downsampling process based on the second downsampling processing algorithm, and the image to be encoded after the second downsampling process is then video encoded to obtain encoded video information.

5. The image encoding method for video calls according to claim 4, wherein, The encoding resolution satisfies the processing computing power requirements, including: The width of the encoding resolution is less than or equal to a first preset width, and the height of the encoding resolution is less than or equal to a first preset height.

6. The image encoding method for video calls according to claim 1, wherein, After performing video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information, the method further includes: The encoded video information is sent to the second terminal, which then decodes the encoded video information to obtain a decoded video frame. If the decoded video frame meets the preset post-processing conditions, the decoded video frame is post-processed based on a preset image post-processing algorithm, and the post-processed decoded video frame is played.

7. The image encoding method for video calls according to claim 6, wherein, The decoded video frames satisfy preset post-processing conditions, including: The resolution width of the decoded video frame is less than or equal to the second preset width, and the resolution height of the decoded video frame is less than or equal to the second preset height.

8. An image encoding device for video calls, applied to a first terminal, wherein, It includes a bandwidth detection module, an information acquisition module, a code table adjustment module, and a video encoding module, among which: The bandwidth detection module is configured to determine the detection bandwidth for real-time video calls with the second terminal; The information acquisition module is configured to determine the video playback frame rate of the second terminal and the video packet loss rate of the first terminal. The code table adjustment module is configured to increase the minimum and maximum bit rates of a preset level in the preset code table to obtain a target code table when the video playback frame rate and the video packet loss rate meet the code table adjustment conditions. The video encoding module is configured to determine the encoding resolution based on the detection bandwidth and the target code table, and to perform video encoding on the image to be encoded in the video call based on the encoding resolution to obtain encoded video information.

9. An image encoding device for video calls, wherein, include: Memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image encoding method for video calls as described in any one of claims 1-7.

10. A non-volatile storage medium for storing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the image encoding method for video calls as described in any one of claims 1-7.

11. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the image encoding method for video calls according to any one of claims 1-7.

Citation Information

Patent Citations

  • Encoder and control method

    CN106254875A

  • Method for dynamically adjusting definition mode in video call, and server

    CN108881780A

  • Video coding method and device, intelligent equipment and storage medium

    CN115623216A

  • Video coding method and device, storage medium and electronic device

    CN117082245A

  • Image coding method and device for video call, equipment, storage medium and product

    CN119583742A