Image processing method and device

Through the collaborative work of the decoding end and the encoding end, the anti-tearing or anti-blurring mode is judged and switched to solve the image tearing and blurring problems in wireless transmission, and the image quality and user experience of the video stream are improved.

CN115462086BActive Publication Date: 2025-10-14HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080100290.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-10-14
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

During wireless transmission, due to unstable channel bandwidth, the base layer or enhancement layer of the sub-image is lost, resulting in image tearing and blurring, affecting the user's video viewing experience.

Method used

Through the collaborative work of the decoding and encoding ends, it is determined whether image tearing or blurring problems will occur based on preset conditions, and the anti-tearing or anti-blurring mode is switched in time. The alternative image is sent and the reference image is stored. The encoding parameters are adjusted to reduce the amount of data and ensure image quality.

Benefits of technology

It effectively avoids image tearing and blurring, and improves the image quality of the video stream and the user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115462086B_ABST
    Figure CN115462086B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an image processing method and device, relates to the field of image processing, can avoid image tearing as much as possible, improves the image quality of the whole video stream, and improves the video watching experience of a user. The scheme is as follows: when a decoding end receives a first image in a video stream, whether the first image satisfies a first preset condition is judged according to a base layer of a first sub-image of the first image; if the first image satisfies the first preset condition, the decoding end sends a second image to display in a display period of the first image, the second image is an image in the video stream sent to display in the first display period, and the first display period is an image display period before the display period of the first image; each image includes a plurality of sub-images, and each sub-image includes a base layer. The embodiment of the present application is used for image processing in a video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing, and in particular to an image processing method and apparatus. Background Art

[0002] With the advancement of electronic technology, users have access to a growing variety of electronic devices. Some electronic devices have strong processing capabilities for video image rendering and encoding, while others offer superior video display quality. To provide a better video experience, the more powerful electronic device can serve as the source device, wirelessly transmitting the encoded video image to a destination device with better display quality. For example, a mobile phone or watch can wirelessly project a video stream onto a television for display.

[0003] During wireless transmission, to reduce latency, the system can employ sub-frame-level image processing. Specifically, the source device can divide a frame into multiple sub-images, encoding and transmitting each sub-image in turn. The destination device then decodes and displays each sub-image in turn. After the destination device completes decoding of each sub-image, it is displayed, minimizing system latency. Each sub-image can include a base layer and at least one enhancement layer. The base layer contains the basic image content, while the enhancement layer ensures higher image quality.

[0004] Because wireless transmission channels are easily affected by multiple factors, such as the operating environment and noise interference signals, the channel transmission bandwidth is unstable. Low channel bandwidth or channel interruption can easily lead to the loss of the base layer or enhancement layer of a sub-image, resulting in image quality issues. For example, if the base layer of a sub-image is lost due to low channel bandwidth or channel interruption, the sub-image without the base layer will not be displayed.

[0005] In the prior art, when a sub-image of a frame image is lost, the sub-image at the corresponding position of the sub-image in the image received before the frame image is used to fill it. Figure 1 As shown, the displayed image includes the contents of multiple sub-images that originally belonged to multiple frame images. Image tearing problems may occur at the junction of sub-images from different frame images (for example, the junction of sub-image (0) of image (n-1) and sub-image (1) of image (n)), thereby affecting the user's video viewing experience. Summary of the Invention

[0006] The embodiments of the present application provide an image processing method and device that can minimize image tearing, improve the image quality of the entire video stream, and enhance the user's video viewing experience.

[0007] To achieve the above objectives, the present invention adopts the following technical solutions:

[0008] In a first aspect, an embodiment of the present application provides an image processing method that can be applied to a decoding end. The method includes: when the decoding end receives a first image in a video stream, the decoding end determines whether the first image meets a first preset condition based on the base layer of the first sub-image of the first image. If the first image meets the first preset condition, the decoding end sends the second image for display within the display period of the first image. The second image is an image in the video stream that is displayed within the first display period, and the first display period is the image display period before the display period of the first image; each frame image in the video stream includes multiple sub-images, and each sub-image includes a base layer.

[0009] In this solution, the decoder can determine whether the first image is about to experience image tearing based on whether the first image meets a first preset condition. If the first image is about to experience image tearing, the decoder displays the second image during the first image's display period, so that the display terminal displays the second image instead of the first image that is about to experience image tearing. This minimizes image tearing in the first image, improves the image quality of the entire video stream, and enhances the user's viewing experience.

[0010] In a possible design, the first preset condition includes: a base layer of a first sub-image of the first image is lost or partially lost.

[0011] That is, if the decoding end does not receive the complete base layer of the first sub-image of the first image, it can be determined that the first preset condition is met, thereby determining that the first image will have an image tearing problem.

[0012] In another possible design, the decoding end stores the received first sub-image of the first image into a cache.

[0013] In this way, the decoding end can use the first sub-image of the received first image as a reference to decode other subsequent images received.

[0014] In another possible design, the method further includes: when a scene switch occurs, the decoder determines, based on the base layer of the first sub-image of the first image, whether the first image satisfies a second preset condition. If the first image satisfies the second preset condition, the decoder displays the second image within the display period of the first image.

[0015] The content of two adjacent frames before and after a scene switch is quite different. Compared with the image before the scene switch, the amount of data in the base layer of the image after the scene switch may increase, so the possibility of base layer loss is greater, and the probability of image tearing is also higher.

[0016] In this solution, at the scene switching, the decoding end can determine whether the first image will have the image tearing problem according to whether the first image satisfies the second preset condition. When the first image will have the image tearing problem, the decoding end sends the second image to be displayed in the display period of the first image, so that the display end displays the second image instead of the first image which will have the image tearing problem in the display period of the first image. Thus, the first image can be prevented from having the image tearing problem as much as possible, the image quality of the whole video stream is improved, and the user video watching experience is improved.

[0017] In another possible design, the second preset condition includes that a ratio of a basic layer encoded data amount of a first sub-image of the first image received in a first preset time length to the first channel bandwidth is greater than or equal to a first preset value, and a ratio of the basic layer encoded data amount of the first sub-image of the first image received in the first preset time length to the second channel bandwidth is greater than or equal to a second preset value. The first channel bandwidth is an average bandwidth in the first preset time length, the second channel bandwidth is an average bandwidth in a second preset time length, and the second preset time length is greater than the first preset time length, and the second preset time length is obtained by extending the first preset time length forward or backward along a time axis.

[0018] It should be understood that the first preset time length is usually very short, for example, can be 4 ms, and the average bandwidth in the first preset time length can be understood as an instantaneous bandwidth; the second preset time length can be understood as an average bandwidth in a certain time period. That is, the decoding end can predict whether the first image will have the image tearing according to the current receiving code rate, the instantaneous bandwidth, and the average bandwidth.

[0019] In another possible design, each sub-image further includes at least one enhancement layer, and the method further includes: when the scene switching occurs, determining whether the first image satisfies a third preset condition according to the enhancement layer of the first sub-image of the first image. If the first image satisfies the third preset condition, the decoding end enters a target mode, in which the second image is sent to be displayed, and the received sub-images of the first image are stored in the buffer.

[0020] Wherein, before and after the scene switching, the content difference between two adjacent images is large. Compared with the image before the scene switching, the data amount of the enhancement layer of the image after the scene switching can increase, so the possibility of the loss of the enhancement layer (including complete loss or partial loss) is large, and the probability of the image blur is also large.

[0021] In the scheme, the target mode can be understood as an anti-blurring mode. When the scene is switched, the decoding end can determine whether the first image will have an image blurring problem according to whether the first image satisfies a third preset condition. When the first image will have an image blurring problem, the decoding end can start the anti-blurring mode, so that the second image is continuously displayed in the anti-blurring mode, so that the second image is displayed and the first image that will soon have blurring is not displayed, so that the current frame image can be prevented from having image blurring as much as possible, the image quality of the entire video stream is improved, and the user video watching experience is improved.

[0022] In addition, the decoding end stores the sub-image of the received first image in the cache, which can facilitate decoding of other subsequent images received with the first image as an inter-frame coded reference frame by taking the sub-image of the first image as a reference.

[0023] In another possible design, the third preset condition includes that the number of enhancement layers of a first sub-image of the first image received by the decoding end is less than or equal to a third preset value, or a ratio of an encoded data amount of the enhancement layers of the first sub-image of the first image received within a third preset time period to a third channel bandwidth is greater than or equal to a fourth preset value, and a ratio of the encoded data amount of the enhancement layers of the first sub-image of the first image received within the third preset time period to a fourth channel bandwidth is greater than or equal to a fifth preset value. The third channel bandwidth is an average bandwidth within the third preset time period, the fourth channel bandwidth is an average bandwidth within a fourth preset time period, and the fourth preset time period is greater than the third preset time period, and the fourth preset time period is obtained by extending the third preset time period forward or backward along a time axis.

[0024] It should be understood that the average bandwidth within the third preset time period can be understood as an instantaneous bandwidth, and the fourth preset time period can be understood as an average bandwidth within a certain time period. That is, the decoding end can determine whether the third preset condition is satisfied according to the number of enhancement layers of the first sub-image of the first image received. Alternatively, the decoding end can determine whether the third preset condition is satisfied according to the current receiving code rate, the instantaneous bandwidth, and the average bandwidth.

[0025] In another possible design, when the decoding end enters the target mode, the method further includes that the decoding end sends first indication information to the encoding end to indicate the encoding end to enter the target mode.

[0026] That is, after the decoding end enters the anti-blurring mode, the encoding end can be notified to also enter the anti-blurring mode, so that corresponding processing operations are performed in the anti-blurring mode.

[0027] In another possible design, the method further includes: when the number of enhancement layers of each sub-image of the received first image is greater than or equal to a sixth preset value, the decoding end exits the target mode to display the most recently successfully decoded image. The decoding end sends second indication information to the encoding end to instruct the encoding end to exit the target mode.

[0028] In this scheme, if the decoding end determines that the number of enhancement layers of each sub-image of the first image is greater than or equal to the sixth preset value, it can be determined that a first image of better quality has been received, and thus the anti-blur mode can be exited, and the encoding end can be instructed to exit the anti-blur mode through the second indication information, so that the encoding end stops sending the first image.

[0029] In the second aspect, an embodiment of the present application provides an image processing method that can be applied to an encoding end. The method includes: in the process of sending the first image in the video stream to the decoding end, the encoding end enters the target mode after receiving the first indication information from the decoding end, and adjusts the encoding parameters in the target mode to reduce the amount of data after encoding other sub-images after the first sub-image in the first image. The encoding end sets the reference frame for inter-frame encoding of the image after the first image to the first image. After receiving the second indication information from the decoding end, the encoding end exits the target mode to stop sending the first image to the decoding end and sends the image in the latest successfully encoded video stream to the decoding end.

[0030] In this solution, after receiving the indication information of entering the anti-blur mode from the decoding end, the encoding end continues to send the first image in the anti-blur mode and adjusts the encoding parameters to reduce the amount of data after encoding the other sub-images after the first sub-image in the first image, or to increase the encoding compression rate of the subsequent images to be transmitted, so that the encoded data of the other sub-images after the first sub-image in the first image can be successfully sent to the decoding end under the current channel conditions. In addition, in the anti-blur mode, the encoding end sets the reference frame of the inter-frame encoding of the images after the first image to the first image, so that the other frame images after the first image in the anti-blur mode are inter-frame encoded with the first image as a reference, providing a better encoding reference for the inter-frame encoding of the subsequent image frames of the first image after the scene switch, thereby reducing the amount of data after encoding the subsequent image frames, so that the encoded data of the subsequent image frames can be successfully transmitted even when the current channel bandwidth is low, improving the problem of long-term image blur and enhancing the image display effect.

[0031] In a third aspect, the embodiments of the present application provide an image processing method, which can be applied to an encoding end. The method comprises: when a scene switching occurs, the encoding end determines whether a first image in a video stream meets a fourth preset condition according to a base layer of a first sub-image of the first image. If the first image meets the fourth preset condition, the encoding end sends third indication information to a decoding end to instruct the decoding end to display a second image in a display period of the first image. The second image is an image in the video stream displayed in the first display period, and the first display period is an image display period before the display period of the first image. Each frame of image in the video stream comprises a plurality of sub-images, and each sub-image comprises a base layer.

[0032] In this scheme, when the scene switching occurs, the encoding end can determine whether the first image will have an image tearing problem according to whether the first image meets the fourth preset condition. When the first image will have the image tearing problem, the encoding end timely informs the decoding end, so that the decoding end displays the second image in the display period of the first image, thereby causing the display end to display the second image instead of the first image which will have the image tearing problem in the display period of the first image. Thus, the first image can be prevented from having the image tearing problem as much as possible, the image quality of the entire video stream is improved, and the user video watching experience is improved.

[0033] In a possible design, the fourth preset condition comprises: a ratio of a data amount of the base layer of the first sub-image of the first image successfully sent by the encoding end in a fifth preset time length to a fifth channel bandwidth is greater than a seventh preset value, and a ratio of the data amount of the base layer of the first sub-image of the first image successfully sent in the fifth preset time length to a sixth channel bandwidth is greater than an eighth preset value. The fifth channel bandwidth is an average bandwidth in the fifth preset time length, the sixth channel bandwidth is an average bandwidth in a sixth preset time length, and the sixth preset time length is greater than the fifth preset time length, and the sixth preset time length is obtained by extending the fifth preset time length forward or backward along a time axis.

[0034] It should be understood that the average bandwidth in the fifth preset time length can be understood as an instantaneous bandwidth, and the sixth preset time length can be understood as an average bandwidth in a certain time period. That is, the encoding end can predict whether the first image will have the image tearing according to the current sending code rate, the instantaneous bandwidth, and the size of the average bandwidth.

[0035] In another possible design, each sub-image further includes at least one enhancement layer, and the method further includes: when the scene switching occurs, determining, by the encoding end, whether the first image satisfies a fifth preset condition according to the enhancement layer of the first sub-image of the first image. If the first image satisfies the fifth preset condition, the encoding end enters a target mode, and sends fourth indication information to the decoding end to indicate the decoding end to enter the target mode. The encoding end adjusts the encoding parameters in the target mode to reduce the data amount of the other sub-images after the first sub-image in the first image. The encoding end sets the reference frame for inter-frame encoding of the images after the first image as the first image.

[0036] When the scene switching occurs, the encoding end can determine whether the first image will have the image blur problem according to whether the first image satisfies the fifth preset condition. When the first image will have the image blur problem, the encoding end can instruct the decoding end to enter the anti-blur mode, and the encoding end itself also starts the anti-blur mode. The encoding end continuously sends the first image in the anti-blur mode, and adjusts the encoding parameters to reduce the data amount of the other sub-images after the first sub-image in the first image, or increases the encoding compression rate of the subsequent to-be-transmitted images to reduce the data amount of the to-be-transmitted images, so that the data of the other sub-images after the first sub-image in the first image can be successfully transmitted to the decoding end under the current channel condition.

[0037] In addition, in the anti-blur mode, the encoding end sets the reference frame for inter-frame encoding of the images after the first image as the first image, so that the other frame images after the first image in the anti-blur mode are inter-frame encoded with the first image as the reference, which provides a better encoding reference for inter-frame encoding of the subsequent image frames of the first image after the scene switching, thereby reducing the data amount of the subsequent image frames, enabling the data of the subsequent image frames to be successfully transmitted under the condition that the current channel bandwidth is low, improving the long-time image blur problem, and improving the image display effect.

[0038] In another possible design, the fifth preset condition includes that the number of the enhancement layers of the first sub-image of the first image successfully transmitted by the encoding end is less than or equal to a ninth preset value. Alternatively, a ratio of the data amount of the enhancement layers of the first sub-image of the first image successfully transmitted by the encoding end within a seventh preset time period to a seventh channel bandwidth is greater than or equal to a tenth preset value, and a ratio of the data amount of the enhancement layers of the first sub-image of the first image successfully transmitted by the encoding end within the seventh preset time period to an eighth channel bandwidth is greater than or equal to an eleventh preset value. The seventh channel bandwidth is an average bandwidth within the seventh preset time period, the eighth channel bandwidth is an average bandwidth within an eighth preset time period, and the eighth preset time period is greater than the seventh preset time period, and the eighth preset time period is obtained by extending the seventh preset time period forward or backward along a time axis.

[0039] It should be understood that the average bandwidth in the seventh preset time period can be understood as the instantaneous bandwidth, and the eighth preset time period can be understood as the average bandwidth in a certain time period. That is, the encoding end can determine whether the fifth preset condition is satisfied according to the number of enhancement layers of the first sub-image of the first image that is successfully transmitted. Alternatively, the encoding end can determine whether the fifth preset condition is satisfied according to the current transmission code rate, the instantaneous bandwidth, and the average bandwidth.

[0040] In another possible design, the method further includes: in a case where the number of enhancement layers of each sub-image of the first image that is successfully transmitted is greater than or equal to the twelfth preset value, the encoding end exits the target mode to stop transmitting the first image to the decoding end and transmitting the image in the video stream that is newly encoded successfully to the decoding end. The encoding end transmits fifth indication information to the decoding end to instruct the decoding end to exit the target mode.

[0041] In this scheme, in a case where the number of enhancement layers of each sub-image of the first image that is successfully transmitted is greater than or equal to the twelfth preset value, the encoding end determines that the decoding end has successfully received the first image with better quality, and thus can exit the anti-blurring mode and instruct the decoding end to also exit the anti-blurring mode.

[0042] In a fourth aspect, an embodiment of the present application provides an image processing method, which can be applied to a decoding end. The method includes: in a process of receiving a first image in a video stream from an encoding end, if third indication information from the encoding end is received, a second image is displayed in a display period of the first image. The second image is an image in the video stream that is displayed in a first display period, and the first display period is an image display period before the display period of the first image.

[0043] In this scheme, if the decoding end receives the relevant indication information from the encoding end, the second image is displayed in the display period of the first image, so that the display end displays the second image in the display period of the first image without displaying the first image.

[0044] In a possible design, the method further includes: if the decoding end receives fourth indication information from the encoding end, the decoding end enters a target mode, in which the second image is displayed and the sub-image of the first image that is received is stored in a cache. After receiving fifth indication information from the encoding end, the decoding end exits the target mode to display the image that is newly decoded successfully.

[0045] In the scheme, the target mode can be understood as an anti-blurring mode. After receiving the related indication information from the encoding end, the decoding end can enter the anti-blurring mode, so as to display the second image, so that the display end displays the second image instead of the first image. In addition, the decoding end can also store the received sub-image of the first image in the cache, so as to provide a decoding reference for subsequent images.

[0046] In another possible design, the scene switching occurs when: a ratio of an area of an intra-coded block of the first sub-image of the first image to a total area of the first sub-image of the first image is greater than or equal to a thirteenth preset value; and / or, a ratio of an amount of encoded data of the first sub-image of the first image to an amount of encoded data of a reference sub-image of a third image is greater than or equal to a fourteenth preset value. The third image is a previous frame image of the first image in the video stream, and the reference sub-image is a sub-image in the third image corresponding to a position of the first sub-image of the first image.

[0047] That is, the size of the area of the intra-coded block of the first sub-image, or the amount of encoded data of the first sub-image, and the like can be used to determine whether the scene switching occurs.

[0048] In a fifth aspect, an image processing apparatus is provided in the embodiments of the present application, which includes a transceiving module and a processing module. The processing module is configured to: when receiving a first image in a video stream through the transceiving module, determine whether the first image satisfies a first preset condition according to a base layer of a first sub-image of the first image. The processing module is further configured to: if the first image satisfies the first preset condition, display a second image in a display period of the first image through the transceiving module, the second image being an image in the video stream displayed in the first display period, and the first display period being an image display period before the display period of the first image; each frame of image in the video stream includes a plurality of sub-images, and each sub-image includes a base layer.

[0049] In the scheme, the image processing apparatus can determine whether the first image will have the image tearing problem according to whether the first image satisfies the first preset condition. When the first image will have the image tearing problem, the image processing apparatus displays the second image in the display period of the first image, so that the display end displays the second image instead of the first image which will have the image tearing problem. Therefore, the image tearing problem of the first image can be avoided as much as possible, the image quality of the entire video stream is improved, and the user video watching experience is improved. For example, the image processing apparatus is a decoding end.

[0050] In a possible design, the first preset condition includes that the base layer of the first sub-image of the first image is lost or partially lost.

[0051] In another possible design, the processing module is further configured to: when a scene switch occurs, determine whether the first image satisfies a second preset condition based on the base layer of the first sub-image of the first image. If the first image satisfies the second preset condition, the processing module is further configured to transmit the second image for display via the transceiver module during the display period of the first image.

[0052] In another possible design, the second preset condition includes: the ratio of the amount of data after base layer encoding of the first sub-image of the first image received within the first preset time length to the first channel bandwidth is greater than or equal to the first preset value, and the ratio of the amount of data after base layer encoding of the first sub-image of the first image received within the first preset time length to the second channel bandwidth is greater than or equal to the second preset value. Wherein, the first channel bandwidth is the average bandwidth within the first preset time length, the second channel bandwidth is the average bandwidth within the second preset time length, and the second preset time length is greater than the first preset time length, and the second preset time length is obtained by extending the first preset time length forward or backward along the time axis.

[0053] In another possible design, each sub-image also includes at least one enhancement layer. The processing module is further configured to, when a scene change occurs, determine whether the first image satisfies a third preset condition based on the enhancement layer of the first sub-image of the first image. If the first image satisfies the third preset condition, the processing module is further configured to enter target mode, display the second image via the transceiver module, and cache the sub-images of the first image received by the transceiver module.

[0054] In another possible design, the third preset condition includes: the number of enhancement layers of the first sub-image of the received first image is less than or equal to the third preset value. Alternatively, the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within the third preset time length to the third channel bandwidth is greater than or equal to the fourth preset value, and the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within the third preset time length to the fourth channel bandwidth is greater than or equal to the fifth preset value. The third channel bandwidth is the average bandwidth within the third preset time length, the fourth channel bandwidth is the average bandwidth within the fourth preset time length, and the fourth preset time length is greater than the third preset time length, and the fourth preset time length is obtained by extending the third preset time length forward or backward along the time axis.

[0055] In another possible design, when entering the target mode, the processing module is further used to: send first indication information to the encoding end through the transceiver module to instruct the encoding end to enter the target mode.

[0056] In another possible design, the processing module is further configured to: in a case where the number of enhancement layers of each sub-image of the first image received by the transceiving module is greater than or equal to a sixth preset value, exit the target mode, and send the latest successfully decoded image to the display through the transceiving module. The processing module is further configured to send second indication information to the encoding end through the transceiving module, to instruct the encoding end to exit the target mode.

[0057] In a sixth aspect, an embodiment of the present application provides an image processing apparatus, including a transceiving module and a processing module. The processing module is configured to: in a process of sending a first image in a video stream to a decoding end through the transceiving module, enter a target mode after receiving first indication information from the decoding end through the transceiving module, and adjust an encoding parameter in the target mode to reduce a data amount of sub-images after a first sub-image in the first image after encoding. The processing module is further configured to set a reference frame for inter-frame encoding of images after the first image as the first image. The processing module is further configured to exit the target mode after receiving second indication information from the decoding end through the transceiving module, to stop sending the first image to the decoding end through the transceiving module and sending an image in the video stream that is successfully encoded to the decoding end through the transceiving module.

[0058] In this solution, the image processing apparatus continuously sends the first image in the anti-fuzzing mode and adjusts the encoding parameter to reduce the data amount of the sub-images after the first sub-image in the first image after encoding, after receiving indication information from the decoding end to enter the anti-fuzzing mode, so that the data of the sub-images after the first sub-image in the first image after encoding can be successfully sent to the decoding end in the current channel condition. In the anti-fuzzing mode, the image processing apparatus sets the reference frame for inter-frame encoding of the images after the first image as the first image, so that the images after the first image in the anti-fuzzing mode are inter-frame encoded with the first image as the reference, which provides a better encoding reference for inter-frame encoding of the subsequent image frames after the first image in the scene switching, thereby reducing the data amount of the subsequent image frames after encoding, so that the data of the subsequent image frames after encoding can be successfully transmitted in the case of a low current channel bandwidth, the problem of long-time image fuzzing is improved, and the image display effect is improved. For example, the image processing apparatus is an encoding end.

[0059] In a seventh aspect, an embodiment of the present application provides an image processing apparatus, comprising a transceiving module and a processing module. The processing module is configured to: when a scene switch occurs, determining whether a first image in a video stream meets a fourth preset condition according to a base layer of a first sub-image of the first image. The processing module is further configured to: if the first image meets the fourth preset condition, sending third indication information to a decoding end through the transceiving module to instruct the decoding end to display a second image in a display period of the first image, the second image being an image in the video stream displayed in the first display period, and the first display period being an image display period before the display period of the first image; each image in the video stream comprises a plurality of sub-images, and each sub-image comprises a base layer.

[0060] In this scheme, when a scene switch occurs, the image processing apparatus can determine whether the first image will have a tearing problem according to whether the first image meets the fourth preset condition. When the first image will have a tearing problem, the image processing apparatus timely informs the decoding end, so that the decoding end displays the second image in the display period of the first image, thereby causing the display end to display the second image instead of the first image which will have a tearing problem in the display period of the first image. Thus, the first image can be prevented from having a tearing problem as much as possible, the image quality of the entire video stream is improved, and the user's video watching experience is improved. For example, the image processing apparatus is a coding end.

[0061] In a possible design, the fourth preset condition comprises: a ratio of an encoded data amount of the base layer of the first sub-image of the first image successfully transmitted in a fifth preset time period to a fifth channel bandwidth being greater than a seventh preset value, and a ratio of the encoded data amount of the base layer of the first sub-image of the first image successfully transmitted in the fifth preset time period to a sixth channel bandwidth being greater than an eighth preset value. The fifth channel bandwidth is an average bandwidth in the fifth preset time period, the sixth channel bandwidth is an average bandwidth in a sixth preset time period, and the sixth preset time period is greater than the fifth preset time period, and the sixth preset time period is obtained by extending the fifth preset time period forward or backward along a time axis.

[0062] In another possible design, each sub-image further comprises at least one enhancement layer. The processing module is further configured to: when a scene switch occurs, determining whether the first image meets a fifth preset condition according to the enhancement layer of the first sub-image of the first image. The processing module is further configured to: if the first image meets the fifth preset condition, entering a target mode, and sending fourth indication information to the decoding end through the transceiving module to instruct the decoding end to enter the target mode in the target mode. The processing module is further configured to: adjusting a coding parameter in the target mode to reduce an encoded data amount of sub-images after the first sub-image in the first image. The processing module is further configured to: setting a reference frame for inter-frame coding of images after the first image as the first image.

[0063] In another possible design, the fifth preset condition includes that the number of the enhancement layers of the first sub-image of the first image successfully transmitted is less than or equal to a ninth preset value. Alternatively, a ratio of the amount of the encoded data of the enhancement layers of the first sub-image of the first image successfully transmitted within a seventh preset time period to a seventh channel bandwidth is greater than or equal to a tenth preset value, and a ratio of the amount of the encoded data of the enhancement layers of the first sub-image of the first image successfully transmitted within the seventh preset time period to an eighth channel bandwidth is greater than or equal to an eleventh preset value. The seventh channel bandwidth is an average bandwidth within the seventh preset time period, the eighth channel bandwidth is an average bandwidth within an eighth preset time period, and the eighth preset time period is greater than the seventh preset time period, and the eighth preset time period is obtained by extending the seventh preset time period forward or backward along a time axis.

[0064] In another possible design, the processing module is further configured to: in a case where the number of the enhancement layers of each sub-image of the first image successfully transmitted by the transceiving module is greater than or equal to a twelfth preset value, exit the target mode to stop transmitting the first image to the decoding end by the transceiving module and transmitting the image in the video stream that is newly encoded successfully to the decoding end by the transceiving module. The processing module is further configured to: transmit fifth indication information to the decoding end by the transceiving module to instruct the decoding end to exit the target mode.

[0065] In an eighth aspect, an embodiment of the present application provides an image processing apparatus, including a processing module and a transceiving module. The processing module is configured to: in a process of receiving the first image in the video stream from the encoding end by the transceiving module, if third indication information from the encoding end is received by the transceiving module, display the second image by the transceiving module in a display period of the first image, the second image being an image in the video stream displayed in a first display period, and the first display period being an image display period before the display period of the first image.

[0066] In this scheme, if the image processing apparatus receives the relevant indication information from the encoding end, the second image is displayed in the display period of the first image, so that the display end displays the second image in the display period of the first image and does not display the first image. For example, the image processing apparatus is the decoding end.

[0067] In a possible design, the processing module is further configured to: if fourth indication information from the encoding end is received by the transceiving module, enter the target mode, in which the second image is displayed by the transceiving module, and the sub-image of the first image received by the transceiving module is stored in the cache. The processing module is further configured to: after fifth indication information from the encoding end is received by the transceiving module, exit the target mode to display the image that is newly decoded successfully by the transceiving module.

[0068] In a ninth aspect, an embodiment of the present application provides an image processing apparatus. The image processing apparatus can be at an encoding end or on the encoding end. Alternatively, the image processing apparatus can be at a decoding end or on the decoding end. The image processing apparatus comprises a processor and a transmission interface; the transmission interface and the processor are coupled; the transmission interface is configured to receive or send images in a video stream, and the processor is configured to invoke software instructions in a memory to perform an image processing method performed by the image processing apparatus in the first to fourth aspects or any possible design of the first to fourth aspects.

[0069] In a tenth aspect, an embodiment of the present application provides an image processing apparatus, comprising one or more processors; and a memory, the memory storing code. When the code is executed by the image processing apparatus, the image processing apparatus performs an image processing method performed by the image processing apparatus in the first to fourth aspects or any possible design of the first to fourth aspects.

[0070] In an eleventh aspect, an embodiment of the present application provides a computer readable storage medium, comprising computer instructions. When the computer instructions are run on an image processing apparatus, the image processing apparatus performs an image processing method in the first to fourth aspects or any possible design of the first to fourth aspects.

[0071] In a twelfth aspect, an embodiment of the present application provides a computer program product, when the computer program product is run on a computer, the computer performs an image processing method performed by an image processing apparatus in the first to fourth aspects or any possible design of the first to fourth aspects.

[0072] In a thirteenth aspect, an embodiment of the present application provides a chip system applied to an image processing apparatus. The chip system comprises one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected through lines; the interface circuits are configured to receive signals from a memory of the image processing apparatus and send signals to the processors, the signals comprising computer instructions stored in the memory; when the processors execute the computer instructions, the image processing apparatus performs an image processing method in the first to fourth aspects or any possible design of the first to fourth aspects.

[0073] In a fourteenth aspect, an embodiment of the present application provides an image processing system, comprising an encoding end and a decoding end, and the encoding end and the decoding end can be used to perform an image processing method in the first to fourth aspects or any possible design of the first to fourth aspects.

[0074] The beneficial effects of the other aspects correspond to the description of the beneficial effects of the method aspects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 A schematic diagram of an image tearing situation provided in an embodiment of the present application;

[0076] Figure 2 A schematic diagram of an image processing system provided in an embodiment of the present application;

[0077] Figure 3A A schematic diagram of the hardware structure of an image processing device provided in an embodiment of the present application Figure 1 ;

[0078] Figure 3B A schematic diagram of the hardware structure of an image processing device provided in an embodiment of the present application Figure 2 ;

[0079] Figure 3C Schematic diagram 3 of the hardware structure of an image processing device provided in an embodiment of the present application;

[0080] Figure 3D A schematic diagram of the hardware structure of an image processing device provided in an embodiment of the present application Figure 4 ;

[0081] Figure 4 A flowchart of an image processing method provided in an embodiment of the present application;

[0082] Figure 5 A schematic diagram of a hierarchical structure of a sub-image provided in an embodiment of the present application;

[0083] Figure 6 A schematic diagram of an image processing process provided in an embodiment of the present application;

[0084] Figure 7 A schematic diagram of another image processing process provided in an embodiment of the present application;

[0085] Figure 8 A flowchart of another image processing method provided in an embodiment of the present application;

[0086] Figure 9 A flowchart of another image processing method provided in an embodiment of the present application;

[0087] Figure 10 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0088] For ease of understanding, some examples of concepts related to the embodiments of this application are provided for reference as follows:

[0089] Video stream: continuous multi-frame video images.

[0090] Sub-images: Multiple small image blocks divided from a complete image frame. For example, a 1920*1080 pixel image can be divided into three 1920*360 pixel sub-images. In sub-frame image processing, a frame image can be divided into multiple sub-images.

[0091] Image base layer and enhancement layer: Each sub-image can be divided into multiple layers, including a base layer and at least one enhancement layer. The base layer can contain the basic content of the image, and the enhancement layer is used to ensure higher image quality.

[0092] Image tearing: When the displayed image includes the contents of multiple sub-images that originally belonged to multiple frame images, image tearing will occur at the junction of sub-images from different frame images.

[0093] Image blur: When the displayed image does not include an enhancement layer or includes a small number of enhancement layers, the image quality will be low, resulting in image blur.

[0094] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0095] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0096] The present application provides an image processing method that can be applied to Figure 2 The image processing system 20 shown. Figure 2 , the image processing system 20 may include a source device 21 and a destination device 22. The source device may encode continuous image frames and then transmit the encoded image frames to the destination device via wireless transmission. The destination device is configured to decode and display the received image frames. The continuous image frames referred to herein may also be referred to as image sequences, such as images in a video stream or associated multi-frame image sequences (e.g., multiple photos taken in continuous shooting mode).

[0097] For example, the image processing method provided in the embodiments of the present application can be applied to various wireless short-range projection application scenarios such as game projection, video projection (such as recorded video projection), or projection of associated multi-frame image sequences (such as PPT projection at work). Among them, the continuous game video images interacting between the encoding end and the decoding end in the game projection scenario can also be called a video stream. The following embodiments of this application will be explained by taking the encoding end sending a video stream to the decoding end as an example.

[0098] Exemplarily, the wireless transmission method may include wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), or infrared (IR communication) and other communication methods.

[0099] For example, the source device can be a mobile phone, wearable device (such as a watch or bracelet), tablet computer, in-vehicle device, laptop computer, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., which has strong processing capabilities such as graphics rendering and encoding. Alternatively, the source device can be a device with strong user interactivity and easy operation, such as a mobile phone or tablet computer. The destination device can be an electronic device with good display effects, such as a television, large-screen device, or augmented reality (AR) / virtual reality (VR) device.

[0100] Among them, the source-end device has functions such as image acquisition, encoding and sending, and can also be called an encoding-end device. Among them, image acquisition includes acquiring downloaded video images, recorded video images, video images generated by applications, or acquiring video images by other means, etc. Among them, for video images generated by applications (such as game video images) or other images, the source-end device also has functions such as graphics rendering. The destination-end device has image receiving, decoding and display functions. In some embodiments, the destination-end device is a physical device, including an interface module and a display module. In other embodiments, the destination-end device may include two independent physical devices, namely a decoding-end device and a display-end device. Among them, the decoding-end device is used to receive and decode images, and the display-end device is used to display images, and can also be used to perform related processing such as image enhancement on the images to be displayed.

[0101] In some embodiments, the decoding device and the display device may be integrated into one physical device, while the encoding device is a separate physical device. In other embodiments, the encoding device, the decoding device, and the display device are separate physical devices. In still other embodiments, the encoding device and the decoding device are located on the same physical device, while the display device is a separate physical device.

[0102] The above-mentioned encoding end device, decoding end device and display end device can be called image processing device or image processing device. Figure 3A A schematic diagram of the hardware structure of an image processing device 300 is provided. The image processing device 300 may include a processor 301 and a transmission interface 302, which is coupled to the processor 301. The transmission interface 302 is used to receive or send images in a video stream. The processor is configured to call software instructions stored in a memory to execute the image processing method provided in the embodiments of the present application.

[0103] The processor may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.

[0104] A transmission interface uses any device such as a transceiver to communicate with other devices or communication networks, such as radio access networks (RAN) and wireless local area networks (WLAN).

[0105] In some embodiments, the image processing device 300 can further include a memory 303 for storing software instructions as described above. The memory can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but not limited to. The memory can exist independently or be integrated with the processor.

[0106] In embodiments of the present application, when the image processing device is an encoding end device, referring to Figure 3B the processor can further include a video encoder for compressively encoding digital video. The encoding end device can support one or more video encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, or MPEG 4, etc. The processor can further include a graphics processing unit (GPU) for performing mathematical and geometric calculations for graphics rendering, and executing program instructions to generate or change display information, etc.

[0107] When the image processing device is a decoding end device, referring to Figure 3C the processor can further include a video decoder for decoding digital video images. The decoding end device can support one or more video decoders. In this way, the decoding end device can play videos in multiple encoding formats, such as: MPEG 1, MPEG 2, MPEG 3, or MPEG 4, etc.

[0108] When the image processing device is a display end device, referring to Figure 3DThe display terminal device further includes a display screen for displaying images. The display screen includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the display terminal device may include one or N display screens, where N is a positive integer greater than one.

[0109] In an embodiment of the present application, the encoding end device is capable of performing graphics rendering, encoding, and sending processing on images in a video stream; the decoding end device is capable of receiving and decoding images in a video stream; when the encoding end device or the decoding end device determines that image quality problems such as image tearing are about to occur based on the first sub-image of a frame image in the video stream, the display end device does not display the current image that is about to be torn, but displays the previous frame image; thereby minimizing image tearing in the current frame image, improving the image quality of the entire video stream, and enhancing the user's video viewing experience.

[0110] The following describes the image processing method provided by the embodiment of the present application from the perspective of the encoding end device, the decoding end device, and the display end device. It is understandable that any multiple of the encoding end device, the decoding end device, and the display end device can be integrated into the same physical device (for example, the decoding end device and the display end device can be integrated into the same physical device), or they can be independent physical devices, and the embodiment of the present application is not limited to this.

[0111] For the convenience of description, the encoding end device is referred to as the encoding end, the decoding end device is referred to as the decoding end, and the display end device is referred to as the display end.

[0112] In an image processing method provided in an embodiment of the present application, a decoding end determines whether image quality problems such as image tearing are about to occur. This method can include multiple image quality problem detection schemes, which are described below.

[0113] (1) Detect whether image tearing is about to occur based on the base layer of the first sub-image of the first image. Figure 4 , the process may include:

[0114] 401. When receiving a first image in a video stream, a decoding end determines whether the first image meets a first preset condition according to a base layer of a first sub-image of the first image.

[0115] The first image can be any frame of the video stream sent by the encoder to the decoder. As mentioned above, the encoder can send the video stream to the decoder through various wireless transmission methods (such as Wi-Fi or Bluetooth) in various application scenarios (such as wireless short-range screen projection scenarios such as game screen projection).

[0116] In the embodiment of the present application, when the sub-frame level image processing method is adopted, each frame image in the video stream may include multiple sub-images, which may also be called image slices or image blocks. Figure 5 Each sub-image can include a base layer and at least one enhancement layer. The base layer contains the basic content of the image, while the enhancement layer can be used to ensure higher image quality. When transmitting each sub-image in the video stream, the base layer is transmitted first, followed by the enhancement layer.

[0117] The first sub-image refers to the first sub-image in each frame to be processed (including encoding, transmission, and decoding, etc.). For example, the first sub-image may be the sub-image in the upper left corner of each frame, that is, the first sub-image of the first image may be the sub-image in the upper left corner of the first image.

[0118] For the first image, the decoding end first receives the first sub-image of the first image, and then receives the other sub-images of the first image. For the first sub-image, the decoding end first receives the base layer of the first sub-image, and then receives the enhancement layer of the first sub-image.

[0119] In step 401, when receiving a first image in a video stream, the decoder determines whether the first image satisfies a first preset condition based on the base layer of the first sub-image received first. The first preset condition is used to determine whether the first image will experience image tearing. That is, the decoder determines whether the first image will experience image tearing based on the base layer of the first sub-image received first.

[0120] 402. If the first image meets a first preset condition, the decoding end displays the second image within a display period of the first image.

[0121] 403. The display terminal displays the second image within the display period of the first image.

[0122] In steps 402-403, if the first image satisfies the first preset condition, the decoding end determines that the first image will have an image tearing problem, and therefore displays the second image instead of the first image during the display period of the first image, so that the display end displays the second image instead of the first image that will have an image tearing problem during the display period of the first image.

[0123] The second image is an image in the video stream displayed within the first display period, and the first display period is the image display period before the display period of the first image. For example, the first image is image frame i in the video stream, and the first display period is the display period corresponding to image frame i-1. If the display end displays image frame i-1 within the display period of image frame i-1, then the second image is image frame i-1, and the display end displays image frame i-1 within the display period of image frame i. For another example, if the display end displays image frame i-2 within the display period of image frame i-1, then the second image is image frame i-2, and the display end displays image frame i-2 within the display period of image frame i.

[0124] In this way, since the first image has image tearing problems, the second image may not have image tearing problems, and the probability of image tearing problems occurring continuously between adjacent image frames is small; therefore, when the first image is about to have image tearing problems, displaying the second image within the display period of the first image can reduce the probability of image tearing and avoid image tearing problems in the first image as much as possible, thereby improving the image quality of the entire video stream and enhancing the user's video viewing experience.

[0125] Moreover, for the first image, when the decoding end receives the first sub-image transmitted first of the first image, it can determine whether the first image is about to be torn according to the basic layer of the first sub-image, thereby performing anti-tearing processing in a timely and early manner when image tearing is about to occur, thereby minimizing the image display delay caused by the anti-tearing processing for the first image.

[0126] Compared with tearing detection based on subsequent images of the first sub-image of the first image, tearing detection based on the first sub-image can determine whether image tearing is about to occur after receiving the first sub-image, so that anti-tearing processing can be performed in time and the second image can be displayed, thereby reducing the image display delay during the first image display period. For example, in a game screen projection scenario, if tearing detection and anti-tearing processing are performed based on subsequent images of the first sub-image of the first image, the image display delay during the first image display period is large, and users can easily find that the game image is stuck on the display end. In addition, in order to minimize the image display delay, in the normal process when no image tearing is detected, the decoding end will send the first sub-image of the first image for display immediately after receiving it, so as to display the first sub-image as soon as possible.

[0127] For example, in some embodiments, the first preset condition may include: the base layer of the first sub-image of the first image received by the decoding end is lost or partially lost. When the channel bandwidth between the encoding end and the decoding end is low or a channel interruption occurs, the base layer of the first sub-image of the first image at the decoding end is easily lost or partially lost. Because the base layer includes the basic content of the image, when the base layer of the first sub-image of the first image received by the decoding end is lost or partially lost, that is, when the decoding end does not receive the complete base layer of the first sub-image, the first sub-image cannot be displayed, and the first image will experience image tearing.

[0128] Since the amount of data in the basic layer of the image is small and easy to be successfully received, the probability of basic layer loss or partial loss in multiple consecutive frames of images is small, the probability of image tearing in the second image displayed on the display end during the display period of the first image is small, and the probability of successful anti-tearing processing of the first image is high.

[0129] Moreover, when the decoding end receives the basic layer transmitted first by the first sub-image, it can determine whether the first image is about to be torn based on the basic layer, thereby performing anti-tearing processing in a timely and early manner when image tearing is about to occur, thereby minimizing the image display delay caused by the anti-tearing processing for the first image.

[0130] For example, see Figure 6 , the encoding end completes the encoding of a sub-image Slice within a unit time T. The sub-image is marked as FnSm, where n represents the number of frames and m represents the number of slices. For example, F2S1 represents the first sub-image of the second frame image. In the image processing method provided in the embodiment of the present application, for the first slice of each frame image (for example, F1S1, F2S1, F3S1), the decoding end performs image tearing detection according to the above-mentioned first preset condition. For example, if it is not detected that the first frame image is about to be torn according to F1S1, the first frame image is displayed in the display period corresponding to the first frame image according to the normal process. If it is detected that the second frame image is about to be torn according to F2S1, the anti-tearing mode can be started at arrow a, thereby repeatedly displaying the first frame image displayed in the previous image display period, that is, F1S1, F1S2 and F1S3, and no longer displaying F2S1, F2S2 and F2S3 of the second frame image. Then, the decoding end automatically exits the anti-tearing mode. Subsequently, if it is not detected according to F3S1 that the third frame image is about to be image-tearing, the third frame image is displayed in the display period corresponding to the third frame image according to the normal process. Figure 6 It can be seen that the display period of each frame of image is delayed after the receiving / decoding period, and the receiving / decoding period is delayed after the encoding period.

[0131] In addition, during the anti-tearing process, the decoding end can put the successfully received sub-images (including the base layer and the enhancement layer) into the cache to provide a decoding reference for subsequently received image frames.

[0132] In the image processing method provided in the embodiment of the present application, if the decoding end does not detect that the first image is about to experience image tearing according to the above-mentioned first preset condition, the first image is received and decoded according to the normal process in the prior art, and the first image is displayed within the display period of the first image.

[0133] It should be noted that the first image can be any image in the video stream, that is, each frame image in the video stream can use the method described in steps 401-403 above to perform image tearing detection and anti-tearing processing to improve the image quality of the entire video stream and enhance the user's video viewing experience.

[0134] In some embodiments of this application, see Figure 4 After step 401, if the first preset condition is not met, for example, if the base layer of the first sub-image of the first image is successfully received, the image processing method provided in this embodiment of the present application may further include the processes described in steps 404-406, or may further include the processes described in steps 407-415. That is, after performing anti-tearing detection based on the base layer of the first sub-image of each frame, when a scene change occurs, it is possible to predict whether image tearing will occur based on the base layer of the first sub-image of the current image, and / or predict whether image blur will occur based on the enhancement layer of the first sub-image of the current image, and perform corresponding preventive processing for these image quality issues.

[0135] In other embodiments of the present application, the solutions described in steps 404-406 and steps 407-415 are independent and parallel image processing solutions to the solution described in steps 401-403. The solutions described in steps 404-406 and steps 407-415 can be executed separately, without having to be executed again if the first preset condition is not met.

[0136] The image processing solutions described in steps 404-406 corresponding to (2) and the image processing solutions described in steps 407-415 corresponding to (3) are respectively described in detail below.

[0137] (2) When a scene switch occurs, the base layer of the first sub-image of the first image is used to predict whether image tearing will occur. Figure 4 , the process may include:

[0138] 404、In the case of scene switching, the decoding end determines whether the first image meets a second preset condition according to the base layer of the first sub-image of the first image.

[0139] The second preset condition is used to predict whether the first image will have a tearing problem, that is, the decoding end predicts whether the first image will have a tearing problem according to the base layer of the first sub-image to be received first.

[0140] Before and after the scene switching, the content of the two adjacent images is quite different. Compared with the image before the scene switching, the data amount of the base layer of the image after the scene switching can increase, and thus the possibility of base layer loss is higher, and the probability of image tearing is also higher.

[0141] In addition, the first image under the scene switching lacks a reference image frame, and thus the data amount of the base layer of the first image after encoding can significantly increase, thereby increasing the possibility of transmission loss and the probability of image tearing.

[0142] In addition, the image tearing phenomenon in the picture at the time of scene switching is more likely to attract the attention of the user, and the influence on the subjective feeling of the user is more obvious and strong.

[0143] Therefore, in the embodiments of the present application, the decoding end can predict whether the first image will have a tearing problem based on the base layer of the first sub-image of the first image at the time of scene switching, and perform a tearing prevention process in time and as early as possible when it is predicted that the first image will have a tearing problem.

[0144] In some embodiments, the decoding end can determine whether the scene switching occurs according to the area of the intra-coded block of the first sub-image, the data amount of the first sub-image after encoding, or other factors. For example, the scene switching occurs when the ratio of the area of the intra-coded block of the first sub-image of the first image to the total area of the first sub-image of the first image is greater than or equal to a thirteenth preset value, and / or the ratio of the data amount of the first sub-image of the first image after encoding to the data amount of the reference sub-image of the third image after encoding is greater than or equal to a fourteenth preset value. The third image is the previous image of the first image in the video stream, and the reference sub-image is the sub-image in the third image corresponding to the position of the first sub-image of the first image.

[0145] When the ratio of the area of the intra-coded block of the first sub-image of the first image to the total area of the first sub-image of the first image is greater than or equal to the thirteenth preset value, it can be indicated that the area of the intra-coded block of the first sub-image is large, the area of the inter-coded block of the first sub-image is small, the correlation of the first sub-image with the previous image of the first image is low, the content difference between the first image and the previous image is large, and the scene switching can occur currently.

[0146] When the ratio of the amount of encoded data of the first sub-image of the first image to the amount of encoded data of the reference sub-image of the third image is greater than or equal to the fourteenth preset value, it can indicate that the content of the first sub-image of the first image and the first sub-image of the previous frame image are significantly different, and the content of the first image and the previous frame image are significantly different, and a scene switch may have occurred.

[0147] Among them, there are many specific application scenarios for scene switching. For example, in an office scenario, the video stream is a PPT image stream. When the next PPT is played, the content of the two PPTs is quite different, and it can be considered that a scene switch has occurred. Another example, when the image in the video stream switches from a simple picture to a complex picture with rotation, it can be considered that a scene switch has occurred. Another example, when the scene corresponding to the image content changes from an indoor scene to an outdoor scene, it can be considered that a scene switch has occurred. Another example, when a work window is opened / closed, it can be considered that a scene switch has occurred.

[0148] 405. If the first image meets the second preset condition, the decoding end displays the second image within the display period of the first image.

[0149] 406. The display terminal displays the second image within the display period of the first image.

[0150] In steps 405-406, if the first image satisfies the second preset condition, the decoding end predicts that the first image will have an image tearing problem, and thus displays the second image instead of the first image during the display period of the first image, so that the display end displays the second image instead of the first image that will have an image tearing problem during the display period of the first image.

[0151] In some embodiments, the decoding end may predict whether image tearing will occur in the first image based on the current receiving bit rate, instantaneous bandwidth, average bandwidth, and the like.

[0152] For example, the second preset condition may include: the ratio of the amount of base layer encoded data of the first sub-image of the first image received by the decoding end within the first preset duration to the first channel bandwidth is greater than or equal to the first preset value, and the ratio of the amount of base layer encoded data of the first sub-image of the first image received within the first preset duration to the second channel bandwidth is greater than or equal to the second preset value. The first channel bandwidth is the average bandwidth within the first preset duration, the second channel bandwidth is the average bandwidth within the second preset duration, and the second preset duration is greater than the first preset duration.

[0153] The first preset time length is relatively short, and can be understood as a unit time length. For example, the first preset time length can be 0.001 s, 0.04 s, 0.004 s or 0.1 s, etc. The first channel bandwidth in the first preset time length is an average bandwidth in a short time window, and can be understood as an instantaneous bandwidth. The data amount of the base layer encoding of the first sub-image of the first image received in the first preset time length can be understood as a received code rate (or a received bitstream) of the base layer encoding of the first sub-image of the first image. The second channel bandwidth is an average bandwidth (or a statistical mean channel bandwidth) in a long time window corresponding to the second preset time length. For example, the second preset time length can be 1 s, 5 s or 10 s, etc. The specific time length of the first preset time length and the second preset time length is not limited in the embodiments of the present application. It should be understood that the time window corresponding to the first preset time length is located in the time window corresponding to the second preset time length. For example, the time window corresponding to the second preset time length can be obtained by extending the time window corresponding to the first preset time length forward or backward along the time axis. The first preset value and the second preset value can be a value less than 1 and close to 1, for example, 0.8, 0.9 or 0.95, etc., and the first preset value and the second preset value can be the same or different.

[0154] When the ratio of the received code rate of the base layer encoding of the first sub-image to the instantaneous bandwidth is greater than or equal to the first preset value (for example, 0.8), and the ratio of the received code rate of the base layer encoding of the first sub-image to the average bandwidth is greater than or equal to the second preset value (for example, 0.9), it can be indicated that the data amount of the base layer encoding of the first sub-image is large, the current channel bandwidth is close to the received code rate, and the current channel bandwidth can be insufficient to transmit the base layer of the subsequent sub-image of the first image. The decoding end can not receive the complete base layer of the subsequent sub-image of the first image, and thus the decoding end predicts that the first image will have image tearing.

[0155] In this way, similar to steps 401-403, in steps 404-406, since the first image has the image tearing problem and the second image can not have the image tearing problem, and the probability of continuous image tearing problems between images is small, when the first image will have the image tearing problem, displaying the second image in the display period of the first image can reduce the probability of image tearing, so as to avoid the image tearing problem of the first image as much as possible, improve the image quality of the entire video stream, and improve the user video watching experience.

[0156] In addition, for the first image, the decoding end can determine whether the first image will have the image tearing when receiving the first sub-image of the first image, that is, the first sub-image transmitted first, according to the base layer of the first sub-image, so that the anti-tearing processing can be performed in time and as early as possible when the image tearing occurs, and the image display delay caused by the anti-tearing processing for the first image can be reduced as much as possible.

[0157] During the anti-tearing process, the decoding end may put the successfully received sub-images (including the base layer and the enhancement layer) into a cache to provide a decoding reference for subsequently received image frames.

[0158] In addition, if the decoding end does not predict that image tearing will occur in the first image according to the above second preset condition, the first image is received and decoded according to the normal process in the prior art, and the first image is displayed within the display period of the first image.

[0159] It should be noted that in the scheme described in the above steps 404-406, since image tearing problems are likely to occur when scene switching occurs, tearing prediction and anti-tearing processing are only performed based on the first sub-image when scene switching occurs, and the delay caused by the anti-tearing processing during scene switching is not easy to be noticed by the user, and thus it is not easy to affect the user's viewing experience; in this way, it is possible to avoid image display delays, playback freezes or unsmoothness caused by image tearing prediction and anti-tearing processing for each frame of image in non-scene switching situations.

[0160] In the scheme described in steps 401-406 above, after the decoder determines that the first image is about to experience image tearing based on a first preset condition, or after the decoder predicts that the first image is about to experience image tearing based on a second preset condition, it can activate anti-tearing mode. In this mode, the second image is displayed during the display period of the first image, and then anti-tearing mode is automatically exited. In other words, anti-tearing mode is only effective for the frame that is about to experience image tearing and does not affect the next frame. Anti-tearing mode is intended to improve image tearing during drastic scene changes.

[0161] In the scheme described in the above steps 401-406, when the decoding end can determine that image tearing has occurred or is about to occur based on the first sub-image of a frame image in the video stream, it does not display the current image that is about to be torn, but displays the frame image that was displayed previously, thereby avoiding image tearing in the current frame image as much as possible, improving the image quality of the entire video stream, and enhancing the user's video viewing experience.

[0162] (3) When a scene change occurs, the enhancement layer of the first sub-image of the first image is used to predict whether image blurring will occur. Figure 4 , the process may include:

[0163] 407. When a scene switch occurs, the decoding end determines whether the first image satisfies a third preset condition according to the enhancement layer of the first sub-image of the first image.

[0164] For the description of scene switching, please refer to the relevant description in the above step 404 and will not be repeated here. The third preset condition is used to predict whether the first image will have an image blur problem, that is, the decoding end predicts whether the first image will have an image blur problem based on the enhancement layer of the first sub-image.

[0165] Before and after a scene change, the content of two adjacent frames differs significantly. Compared to the image before the scene change, the amount of data in the enhancement layer of the image after the scene change may increase, resulting in a greater possibility of enhancement layer loss (including complete loss or partial loss), and a higher probability of image blur.

[0166] Furthermore, since the first image in the scene switching scenario lacks a reference image frame, the amount of data after encoding the first image enhancement layer will increase significantly, thereby increasing the possibility of transmission loss and the probability of image blur.

[0167] In addition, the problem of image blur during scene switching is more likely to attract the user's attention. In particular, the phenomenon of image blur in the still picture after the scene switch is more easily perceived by the user, and the impact on the user's subjective experience is more obvious and strong. Moreover, the amount of data in the image enhancement layer is larger than that in the basic layer, and the similarity between adjacent multi-frame images after the scene switch is relatively large; if the enhancement layer of the first image is easily lost during the scene switch, the enhancement layers of the other frames after the first image are also easily lost due to the lack of reference image frames, which can easily lead to the blurring problem of multiple consecutive frames for a long time, affecting the user's video viewing experience.

[0168] Therefore, in an embodiment of the present application, the decoding end can predict whether the first image will have an image blur problem based on the enhancement layer of the first sub-image of the first image when the scene switches, and perform anti-blur processing in a timely and early manner when it is predicted that the image blur problem will occur.

[0169] 408. If the first image meets the third preset condition, the decoding end sends first instruction information to the encoding end to instruct the encoding end to enter the target mode.

[0170] If the first image satisfies the third preset condition, the decoding end predicts that the first image will have an image blur problem, and thus can instruct the encoding end to enter the target mode.

[0171] In some embodiments, the decoding end may predict whether the first image will be blurred based on the current receiving bit rate, instantaneous bandwidth, average bandwidth, and the like.

[0172] For example, the third preset condition may include: the number of enhanced layers of the first sub-image of the first image received by the decoding end is less than or equal to the third preset value. Or, the ratio of the amount of data after encoding the enhanced layer of the first sub-image of the first image received by the decoding end within the third preset time length to the third channel bandwidth is greater than or equal to the fourth preset value, and the ratio of the amount of data after encoding the enhanced layer of the first sub-image of the first image received by the decoding end within the third preset time length to the fourth channel bandwidth is greater than or equal to the fifth preset value. The third channel bandwidth is the average bandwidth within the third preset time length, the fourth channel bandwidth is the average bandwidth within the fourth preset time length, and the fourth preset time length is greater than the third preset time length.

[0173] Among them, if the number of enhancement layers of the first sub-image of the first image received by the decoding end is less than or equal to the third preset value, it may indicate that the decoding end has not received a sufficient number of enhancement layers, the quality of the first image is poor, and the probability of receiving a sufficient number of enhancement layers of subsequent sub-images of the first image is small, so the decoding end predicts that the first image will have image blurring.

[0174] The third preset duration is relatively short and can be understood as a unit duration. For example, the third preset duration can be 0.001s, 0.04s, 0.004s, or 0.1s. The third channel bandwidth within the third preset duration is the average bandwidth within a short time window and can be understood as the instantaneous bandwidth. The amount of data after enhancement layer encoding of the first sub-image of the first image received within the third preset duration can be understood as the received bit rate after enhancement layer encoding of the first sub-image of the first image. The fourth channel bandwidth is the average bandwidth within a longer time window corresponding to the fourth preset duration. For example, the fourth preset duration can be 1s, 5s, or 10s. It should be understood that the time window corresponding to the third preset duration is within the time window corresponding to the fourth preset duration. Exemplarily, the time window corresponding to the fourth preset duration can be derived by extending forward or backward along the time axis based on the time window corresponding to the third preset duration. This embodiment of the present application does not specifically limit the specific durations of the third and fourth preset durations. The fourth preset value and the fifth preset value may be values ​​less than 1 and close to 1, such as 0.8, 0.9 or 0.95, and the fourth preset value and the fifth preset value may be the same or different in size.

[0175] When the ratio of the received bit rate after encoding the enhancement layer of the first sub-image to the instantaneous bandwidth is greater than or equal to the third preset value (for example, 0.85), and the ratio of the received bit rate after encoding the enhancement layer of the first sub-image to the average bandwidth is greater than or equal to the fourth preset value (for example, 0.95), it can indicate that the amount of data after encoding the enhancement layer of the first sub-image is large, the current channel bandwidth is close to the received bit rate, and the current channel bandwidth may not be sufficient to transmit all the enhancement layers or a sufficient number of enhancement layers of the subsequent sub-images of the first image. The decoding end may not be able to receive all the enhancement layers or a sufficient number of enhancement layers of the subsequent sub-images of the first image. Therefore, the decoding end predicts that the first image will have an image blur problem, and the subsequent images are also likely to have an image blur problem.

[0176] For the first image, when the decoding end receives the first sub-image transmitted first of the first image, it can determine whether the first image is about to be blurred based on the enhancement layer of the first sub-image, so as to perform anti-blurring processing in a timely and early manner when image blurring is about to occur, thereby minimizing the image display delay caused by the anti-blurring processing for the first image.

[0177] The target mode may be referred to as an anti-blur mode. When the decoder predicts, based on the third preset condition, that the first image will be blurred, it may send first indication information to the encoder to instruct the encoder to enter the anti-blur mode, thereby performing corresponding processing operations in the anti-blur mode. For example, the first indication information may be an anti-blur mode flag. The anti-blur mode is intended to improve the problem of prolonged image blurring when a large scene change is followed by a relatively static scene.

[0178] 409. When the encoding end sends the first image in the video stream to the decoding end, after receiving the first indication information from the decoding end, the encoding end enters the target mode, adjusts the encoding parameters in the target mode to reduce the amount of data after encoding other sub-images after the first sub-image in the first image, and sets the reference frame of the inter-frame encoding of the image after the first image to the first image.

[0179] After receiving the first indication information from the decoding end, the encoder enters the anti-blur mode. In the anti-blur mode, the encoder adjusts encoding parameters, for example, increasing the quantization parameter QP of the discrete cosine transform, to achieve a higher encoding compression rate, thereby reducing the amount of encoded data of other sub-images after the first sub-image in the first image, so that the encoded data of other sub-images after the first sub-image in the first image can be successfully sent to the decoding end under the current channel condition.

[0180] In anti-blur mode, the encoder continuously sends the first image so that the decoder can receive a sufficient number of enhancement layers for each sub-image of the first image, that is, the decoder can receive an image of better quality. In addition, in anti-blur mode, the encoder can also set the reference frame for inter-frame coding of images after the first image to the first image, so that other frame images after the first image in anti-blur mode are inter-frame coded with the first image as a reference, so that when the decoder receives other frame images after the first image, it can decode with the received first image as a reference. For example, assume that the first image includes image frame a and image frame b. In anti-blur mode, the encoder encodes image frame a and image frame b, and both image frame a and image frame b are inter-frame coded with the first image as a reference. If the decoder receives image frame a or image frame b, it can decode with the received first image as a reference.

[0181] In addition, the encoding end sets the reference frame for inter-frame coding of images after the first image as the first image, which can provide a better coding reference for the inter-frame coding of subsequent image frames of the first image after the scene switching, thereby reducing the amount of data after encoding the subsequent image frames, so that the data after encoding the subsequent image frames can be successfully transmitted even when the current channel bandwidth is low, improving the problem of long-term image blur and enhancing the image display effect.

[0182] In anti-blur mode, the encoder sets the reference frame for inter-frame coding of images following the first image to the first image. This prevents the decoder from receiving frames after the first image that lack a decoding reference frame. For example, assuming that the first image is followed by image frames a and b, the encoder encodes image frames a and b in anti-blur mode. If image frame b is inter-coded with image frame a as a reference but not with the first image as a reference, and the decoder receives image frame b but not image frame a, then image frame b lacks a decoding reference frame.

[0183] 410. If the first image meets the third preset condition, the decoding end enters the target mode, displays the second image in the target mode, and stores the received sub-image of the first image in the cache.

[0184] When the decoder predicts that the first image will be blurred according to the third preset condition, the decoder can enter the anti-blur mode. In the anti-blur mode, the decoder can continue to receive the first image in response to the encoder continuing to send the first image.

[0185] Furthermore, in anti-blur mode, the decoding end can display the second image while continuously receiving the first image, so that the display end displays the second image instead of the first image that is about to become blurred, thereby minimizing image blur in the current frame image, improving the image quality of the entire video stream, and enhancing the user's video viewing experience. Since the duration of continuously receiving the first image is uncertain, the duration of the decoding end displaying the second image is also uncertain, and may be one image display period (the display duration corresponding to displaying one frame of image) or multiple image display periods. For example, one image display period may be 0.04s.

[0186] The decoding end can also store the received sub-image of the first image in a cache (i.e., the decoder's cache) so that the decoding end can use the sub-image of the first image as a reference to decode other subsequent images received with the first image as an inter-frame coded reference frame.

[0187] 411. The display terminal displays a second image.

[0188] After receiving the second image sent by the decoding end, the display end displays the second image. It should be noted that, compared with the duration of the second image sent by the decoding end, the duration of the second image displayed by the display end is also uncertain, and may be one image display period or multiple image display periods.

[0189] 412. When the number of enhancement layers of each sub-image of the received first image is greater than or equal to a sixth preset value, the decoding end sends second indication information to the encoding end to instruct the encoding end to exit the target mode.

[0190] When the decoding end is continuously receiving the first image, if it is determined that the number of enhancement layers of each sub-image of the first image is greater than or equal to the sixth preset value, it can be determined that a first image with good quality has been received, and the encoding end can be instructed to exit the anti-blur mode through the second indication information, so that the encoding end stops sending the first image.

[0191] 413. After receiving the second instruction information from the decoding end, the encoding end exits the target mode to stop sending the first image to the decoding end and sends the image in the latest successfully encoded video stream to the decoding end.

[0192] After receiving the second instruction information sent by the decoder, the encoder exits anti-blur mode according to the second instruction information, thereby stopping sending the first image and sending the most recently successfully encoded video frame to the decoder. For example, if the most recently successfully encoded image frame after the encoder exits anti-blur mode is image frame b, the encoder may send image frame b to the decoder. However, during the period described in step 409 above where the encoder continues to send the first image in anti-blur mode, image frame a that was encoded before image frame b will no longer be sent to the decoder.

[0193] 414. When the number of enhancement layers of each sub-image of the received first image is greater than or equal to a sixth preset value, the decoding end exits the target mode to display the latest successfully decoded image.

[0194] 415. The display terminal displays the latest image successfully decoded by the decoding terminal.

[0195] The most recently successfully decoded image is the image corresponding to the current display period. In steps 414-415, the decoding end determines that the number of enhancement layers of each sub-image of the received first image is greater than or equal to a sixth preset value. If a good-quality first image has been received, the decoding end may exit the anti-blur mode and display the most recently successfully decoded image (e.g., image frame b). The display end displays the most recently successfully decoded image.

[0196] For example, see Figure 7 , when the scene switches, the decoding end performs image blur detection based on the first Slice of the image. If, when the scene switches, the decoding end detects that the second frame image is about to be torn according to F2S1, the anti-blur mode can be started at arrow b, thereby repeatedly displaying the first frame image displayed in the previous image display period, i.e., F1S1, F1S2, and F1S3, and no longer displaying F2S1, F2S2, and F2S3 of the second frame image. The anti-blur mode is not exited until the decoding end determines that it has received a second frame image of better quality (i.e., the enhancement layer of each sub-image of the received second frame image is greater than or equal to the twelfth preset value). As Figure 7 As shown in , the transmission time of the second frame of image also occupies the transmission period of the third frame of image. After exiting the anti-blur mode, the decoding end receives and decodes the fourth frame of image and sends the fourth frame of image for display; Figure 7 As shown, after repeatedly displaying the first frame image, the display terminal no longer displays the second and third frame images, but directly displays the fourth frame image (ie, F4S1, F4S2 and F4S3) that has been successfully decoded the most recently.

[0197] Furthermore, if the decoding end does not predict that the first image will be blurred based on the third preset condition, the first image is received and decoded according to a normal process in the prior art, and the first image is displayed within the display period of the first image. It should be understood that in a normal process, the target image is displayed within the display period of the target image.

[0198] In the solution described in steps 407-415 above, after the decoder predicts that the first image will be blurred based on the third preset condition, it can activate anti-blur mode, thereby continuously displaying the second image in anti-blur mode until it receives a first image of satisfactory quality, only then exiting anti-blur mode. The first image can provide a good reference for encoding / decoding subsequent image frames, reducing the amount of data encoded in subsequent image frames, making it easier for the encoded data of subsequent image frames to be successfully transmitted to the decoder for decoding, thereby avoiding the problem of prolonged image blur.

[0199] In this way, when the decoding end predicts that image blurring is about to occur based on the first sub-image of a frame image in the video stream, it will not display the current image that is about to become blurred, but will continue to display the previous frame image, thereby minimizing image blurring in the current frame image, improving the image quality of the entire video stream, and enhancing the user's video viewing experience.

[0200] That is to say, when the scene switches, the decoding end can predict the reception / loss of subsequent sub-images of the first sub-image based on the relevant information of the basic layer or enhancement layer of the first sub-image of the first image and the bandwidth statistics of the channel, and adjust the encoding, transmission and display strategies according to the prediction results to improve image quality problems such as image blur and display complete and high-quality images as much as possible.

[0201] In addition, in the scheme described in the above steps 407-415, since image blurring problems are likely to occur when scene switching occurs, blur prediction and anti-blurring processing are performed based on the first sub-image only when scene switching occurs, and the delay caused by anti-blurring processing during scene switching is not easily perceived by the user, thereby not easily affecting the user's viewing experience; in this way, image display delays, playback freezes or unsmoothness caused by image blur prediction and anti-blurring processing for each frame of image in non-scene switching situations can be avoided.

[0202] In some embodiments of the present application, when a scene switch occurs, the decoding end predicts whether the first image will experience image tearing based on the second preset condition, and predicts whether the first image will experience image blurring based on the third preset condition in parallel. In this case, if the decoding end predicts that the first image will experience image tearing and image blurring, an anti-blurring method is used to process the first image to obtain a higher-quality first image, thereby providing a better reference for decoding subsequent images. If neither the anti-tearing mode nor the anti-blurring mode is activated, low-latency image display is performed according to the normal process in the prior art.

[0203] In the image processing method provided in the above-mentioned embodiment of the present application, the decoding end can determine that when image quality problems such as image tearing or image blurring occur or are about to occur based on the first sub-image of a frame image in the video stream, it can perform anti-tearing or anti-blurring processing in a timely and early manner, thereby avoiding image quality problems in the currently displayed image as much as possible, improving the image quality of the entire video stream, and enhancing the user's video viewing experience.

[0204] Other embodiments of the present application provide another image processing method, which can be used by the encoder to predict whether the image tearing problem will occur. For example, see Figure 8 , the image processing method may include the following steps 801-804.

[0205] It should be noted that the process of predicting whether image tearing will occur and performing anti-tearing processing on the encoder side described in steps 801-804 is similar to the process of predicting whether image tearing will occur and performing anti-tearing processing on the decoder side described in steps 404-406 above. The relevant descriptions of steps 404-406 above can be referred to. The following mainly provides supplementary explanations of the differences:

[0206] 801. When a scene switch occurs, the encoder determines, based on a base layer of a first sub-image of a first image in a video stream, whether the first image meets a fourth preset condition.

[0207] For an explanation of scene switching, please refer to the relevant description in step 404 above and will not be repeated here. The fourth preset condition is used to predict whether the first image will have image tearing. That is, the encoder predicts whether the first image will have image tearing based on the base layer of the first sub-image to be sent.

[0208] In some embodiments, the encoder may predict whether image tearing will occur in the first image based on the current transmission bit rate, instantaneous bandwidth, average bandwidth, and the like.

[0209] For example, the fourth preset condition may include: a ratio of the amount of base layer encoded data of the first sub-image of the first image successfully transmitted within a fifth preset duration to the fifth channel bandwidth is greater than a seventh preset value, and a ratio of the amount of base layer encoded data of the first sub-image of the first image successfully transmitted within the fifth preset duration to the sixth channel bandwidth is greater than an eighth preset value. The fifth channel bandwidth is the average bandwidth within the fifth preset duration, the sixth channel bandwidth is the average bandwidth within the sixth preset duration, and the sixth preset duration is greater than the fifth preset duration.

[0210] The fifth preset duration is relatively short and can be understood as a unit duration. The amount of data after base layer encoding of the first sub-image of the first image received within the fifth preset duration can be understood as the transmission bit rate (or transmission code stream) after base layer encoding of the first sub-image of the first image. The fifth channel bandwidth can be understood as the instantaneous bandwidth within a unit duration, and the sixth channel bandwidth is the average bandwidth within the longer time window corresponding to the sixth preset duration. For example, the fifth preset duration can be 4ms, and the sixth preset duration can be 1s. It should be understood that the time window corresponding to the fifth preset duration is within the time window corresponding to the sixth preset duration. Exemplarily, the time window corresponding to the sixth preset duration can be derived by extending forward or backward along the time axis based on the time window corresponding to the fifth preset duration. The seventh and eighth preset values ​​can be values ​​less than 1 and close to 1, such as 0.8, 0.9, or 0.95, and the seventh and eighth preset values ​​can be the same or different.

[0211] When the ratio of the transmitted bit rate of the base layer of the first sub-image after encoding to the instantaneous bandwidth is greater than or equal to the seventh preset value (for example, 0.8), and the ratio of the transmitted bit rate of the base layer of the first sub-image after encoding to the average bandwidth is greater than or equal to the eighth preset value (for example, 0.9), it can be indicated that the amount of data after the base layer of the first sub-image is large, the current channel bandwidth is close to the transmission bit rate, and the current channel bandwidth may not be sufficient to transmit the complete base layer of the subsequent sub-image of the first image. The encoding end may not be able to successfully send the complete base layer of the subsequent sub-image of the first image, and thus the encoding end predicts that image tearing will occur in the first image.

[0212] 802. If the first image satisfies a fourth preset condition, the encoder sends third instruction information to the decoder to instruct the decoder to display the second image within a display period of the first image.

[0213] For a description of the second image, see step 403 above and will not be repeated here. If the first image satisfies the fourth preset condition, the encoder predicts that the first image will experience image tearing, and may therefore send third indication information to the decoder to instruct the decoder to display the second image during the display period of the first image.

[0214] 803. When receiving the first image in the video stream from the encoding end, if the decoding end receives third instruction information from the encoding end, the decoding end displays the second image within the display period of the first image.

[0215] 804. The display terminal displays the second image within the display period of the first image.

[0216] In steps 803-804, if the decoding end receives the third indication information sent by the encoding segment, it can enter the anti-tearing mode to display the second image within the display period of the first image, so that the display end displays the second image within the display period of the first image, and then the decoding end automatically exits the anti-tearing mode. It should be understood that the anti-tearing mode is not a continuous process. After the second image is displayed within the display period of the first image, the current anti-tearing mode ends, and the second image will not be continuously displayed in the next display period. In theory, the encoding end will perform a tearing prediction for each frame of image sent. During the anti-tearing processing process, the decoding end can put the successfully received sub-images (including the base layer and the enhancement layer) into the cache to provide a reference for the decoding of subsequent image frames.

[0217] Since the first image has image tearing problems, while the second image may not have image tearing problems or may have image tearing problems, and the probability of image tearing problems occurring continuously between images is small; therefore, when the first image is about to have image tearing problems, displaying the second image during the display period of the first image instead of displaying the first image that is about to have image tearing problems can reduce the probability of image tearing, thereby avoiding image tearing problems in the first image as much as possible, improving the image quality of the entire video stream, and enhancing the user's video viewing experience.

[0218] Moreover, for the first image, when the encoding end sends the basic layer of the first sub-image that is first transmitted in the first image, it can determine whether the first image is about to be torn according to the basic layer of the first sub-image, thereby performing anti-tearing processing in a timely and early manner when image tearing is about to occur, and minimizing the image display delay caused by the anti-tearing processing for the first image.

[0219] Another image processing method provided in another embodiment of the present application is to predict whether the image blur problem will occur at the encoding end. Figure 9 , the image processing method may include the following steps 901-908.

[0220] It should be noted that the process of predicting whether an image blur problem will occur and performing anti-blur processing on the encoding side described in steps 901-908 is similar to the process of predicting whether an image blur problem will occur and performing anti-blur processing on the decoding side described in steps 407-415 above. The relevant descriptions of steps 407-415 above can be referred to. The following mainly provides supplementary explanations of the differences:

[0221] 901. When a scene switch occurs, the encoder determines, based on an enhancement layer of a first sub-image of the first image, whether the first image satisfies a fifth preset condition.

[0222] For the description of scene switching, please refer to the relevant description in the above step 404 and will not be repeated here. The fifth preset condition is used to predict whether the first image will have an image blur problem, that is, the encoder predicts whether the first image will have an image blur problem based on the enhancement layer of the first sub-image.

[0223] In some embodiments, the encoding end may predict whether the first image will be blurred based on the current transmission bit rate, instantaneous bandwidth, average bandwidth, and the like.

[0224] For example, the fifth preset condition may include: the number of enhanced layers of the first sub-image of the first image successfully sent is less than or equal to the ninth preset value. Alternatively, the ratio of the amount of data after encoding the enhanced layer of the first sub-image of the first image successfully sent within the seventh preset time length to the seventh channel bandwidth is greater than or equal to the tenth preset value, and the ratio of the amount of data after encoding the enhanced layer of the first sub-image of the first image successfully sent within the seventh preset time length to the eighth channel bandwidth is greater than or equal to the eleventh preset value. The seventh channel bandwidth is the average bandwidth within the seventh preset time length, the eighth channel bandwidth is the average bandwidth within the eighth preset time length, and the eighth preset time length is greater than the seventh preset time length.

[0225] Among them, if the number of enhancement layers of the first sub-image of the first image successfully sent by the encoding end is less than or equal to the ninth preset value, it may indicate that the encoding end has not successfully sent a sufficient number of enhancement layers, the quality of the first image received by the decoding end is poor, and the probability of successfully sending a sufficient number of enhancement layers of subsequent sub-images of the first image is small, so the encoding end predicts that the first image will have image blurring.

[0226] The seventh preset duration is relatively short and can be understood as a unit duration. The amount of data after enhancement layer encoding of the first sub-image of the first image successfully transmitted within the seventh preset duration can be understood as the transmission bit rate after enhancement layer encoding of the first sub-image of the first image. The seventh channel bandwidth within the first preset duration can be understood as the instantaneous bandwidth, and the eighth channel bandwidth is the average bandwidth within the longer time window corresponding to the eighth preset duration. For example, the seventh preset duration can be 4ms, and the eighth preset duration can be 1s. It should be understood that the time window corresponding to the seventh preset duration is within the time window corresponding to the eighth preset duration. Exemplarily, the time window corresponding to the eighth preset duration can be obtained by extending forward or backward along the time axis based on the time window corresponding to the seventh preset duration. The tenth and eleventh preset values ​​can be values ​​less than 1 and close to 1, and the tenth and eleventh preset values ​​can be the same or different.

[0227] When the ratio of the transmitted bit rate of the enhanced layer after encoding of the first sub-image to the instantaneous bandwidth is greater than or equal to the tenth preset value, and the ratio of the transmitted bit rate of the enhanced layer after encoding of the first sub-image to the average bandwidth is greater than or equal to the eleventh preset value, it can indicate that the amount of data after encoding the enhanced layer of the first sub-image is large, the current channel bandwidth is close to the transmitted bit rate, and the current channel bandwidth may not be sufficient to transmit all the enhanced layers or a sufficient number of enhanced layers of the subsequent sub-images of the first image. The encoding end may not be able to successfully send all the enhanced layers or a sufficient number of enhanced layers of the subsequent sub-images of the first image. Therefore, the encoding end predicts that the first image will have an image blur problem, and the subsequent images are also likely to have an image blur problem.

[0228] In addition, the above-mentioned first preset time length, third preset time length, fifth preset time length and seventh preset time length may be the same or different; the above-mentioned second preset time length, fourth preset time length, sixth preset time length and eighth preset time length may be the same or different; the embodiments of this application are not limited.

[0229] 902. If the first image meets the fifth preset condition, the encoding end enters the target mode, adjusts the encoding parameters in the target mode to reduce the amount of data after encoding other sub-images after the first sub-image in the first image, and sets the reference frame of the inter-frame encoding of the image after the first image to the first image.

[0230] The target mode may be referred to as an anti-blur mode. The description of step 902 may refer to the above step 409 and will not be repeated here.

[0231] 903. The encoder sends fourth instruction information to the decoder in the target mode to instruct the decoder to enter the target mode.

[0232] When the encoder predicts, based on the fifth preset condition, that the first image will be blurred, the encoder may send fourth indication information to the decoder to instruct the decoder to enter an anti-blur mode, thereby performing corresponding processing operations in the anti-blur mode. For example, the fourth indication information may be an anti-blur mode flag.

[0233] 904. If the decoding end receives the fourth instruction information from the encoding end, it enters the target mode, displays the second image in the target mode, and stores the received sub-image of the first image into the cache.

[0234] If the decoding end receives the fourth indication information sent by the encoding end, it enters the anti-blur mode. In the anti-blur mode, the decoding end continues to display the second image so that the display end continues to display the second image.

[0235] 905. When the number of enhancement layers of each sub-image of the first image successfully sent is greater than or equal to the twelfth preset value, the encoding end exits the target mode to stop sending the first image to the decoding end and sends the image in the latest successfully encoded video stream to the decoding end.

[0236] When the number of enhancement layers of each sub-image of the successfully sent first image is greater than or equal to the twelfth preset value, the decoding end receives the first image of better quality, so the encoding end can exit the anti-blur mode, thereby stopping sending the first image to the decoding end, and sending the latest successfully encoded image to the decoding end for decoding.

[0237] Correspondingly, the decoding end receives the latest successfully encoded image from the encoding end.

[0238] 906. When exiting the target mode, the encoder sends fifth instruction information to the decoder to instruct the decoder to exit the target mode.

[0239] When exiting the anti-ambiguity mode, the encoding end may further send fifth indication information to the decoding end to instruct the decoding end to also exit the anti-ambiguity mode.

[0240] 907. After receiving the fifth instruction information from the encoding end, the decoding end exits the target mode to display the latest successfully decoded image.

[0241] 908. The display terminal displays the latest successfully decoded image.

[0242] The latest successfully decoded image is the image corresponding to the current display period. In steps 907-908, after the decoding end exits the anti-blur mode, the latest successfully decoded image is sent for display, so that the display end displays the latest successfully decoded image.

[0243] In the solution described in steps 901-908 above, after the encoder predicts that the first image will be blurred based on the fifth preset condition, it can activate anti-blur mode, thereby continuously displaying the second image in anti-blur mode and only exiting anti-blur mode after receiving a first image of satisfactory quality. The first image can provide a good reference for encoding and decoding subsequent image frames, reducing the amount of data encoded in subsequent image frames, making it easier for the encoded data of subsequent image frames to be successfully transmitted to the decoder for decoding, thereby avoiding the problem of prolonged image blur.

[0244] Moreover, for the first image, when the encoding end sends the first sub-image that is first transmitted of the first image, it can determine whether the first image is about to be blurred based on the enhancement layer of the first sub-image, thereby performing anti-blurring processing in a timely and early manner when image blurring is about to occur, and minimizing the image display delay caused by the anti-blurring processing for the first image.

[0245] That is, in the scheme described in steps 901-908, when the scene switches, the encoding end can predict the reception / loss of subsequent sub-images of the first sub-image based on the relevant information of the basic layer or enhancement layer of the first sub-image of the first image and the bandwidth statistics of the channel, and adjust the encoding, transmission and display strategies according to the prediction results to improve image quality problems such as image blur and display a complete and high-quality image as much as possible.

[0246] The image processing method provided in the above embodiment can also be understood as a sub-image loss prediction method in a sub-frame level image processing system and a display method based on the prediction result.

[0247] When using sub-frame low-latency projection, if the image quality requirements and the nature of the image itself require a high-bitrate base layer, and at the same time encounter low channel bandwidth, image tearing will occur. In addition, if the channel bandwidth remains low, a very easily observable long-term low image quality problem will occur in static scenes after scene switching. In addition, due to the lack of reference for the current image when the scene switches, the bitrate of both the base layer and the enhancement layer will increase significantly, which greatly increases the possibility of transmission loss. At the same time, the image tearing problem that occurs during scene switching is easier to observe than the image tearing problem in continuous motion scenes, and the impact on subjective experience is also significantly stronger.

[0248] The image processing method provided in the embodiment of the present application uses the information of the first sub-image of a frame of image in the sub-frame level system and the bandwidth statistics of the channel to predict the reception / loss of subsequent sub-images of the frame of image, and adjusts the encoding, transmission and display strategies according to the prediction results to improve the problems of image tearing and long-term low-quality images. This method will temporarily increase the delay for a short period of time (after the scope of effectiveness of the invention solution ends, the delay will drop to a normal value). However, considering that the delay is less likely to be observed than tearing and long-term low image quality at high frame rates, this invention can significantly improve the subjective experience.

[0249] The image processing method provided in the embodiment of the present application can significantly reduce the occurrence of image tearing without increasing the number of images without tearing problems by performing image tearing detection / prediction on the first sub-image, and ultimately significantly improve the subjective effect of the displayed image. By performing image blur detection / prediction on the first sub-image, it can provide high-quality images when scenes are switched without increasing the number of images without blurring problems, and provide a better reference for subsequent inter-frame coding frames, which can effectively reduce the problem of long-term image blurring after scene switching when the bandwidth is low, and ultimately significantly improve the subjective effect of the displayed image. Therefore, this method can provide prediction results with significantly improved effects at a negligible computational cost; the resulting false detections are very limited, and the impact on the subjective quality of the final image is minimal.

[0250] In addition, in the embodiments of the present application, the detection and prevention of image quality problems such as image tearing and image blur can be implemented at the software layer, and thus can be achieved through software updates on existing devices without modifying the hardware, with a high degree of forward compatibility.

[0251] It is understandable that in order to implement the above functions, the image processing device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of this application.

[0252] In the embodiment of the present application, the image processing device can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. In actual implementation, other division methods may be used.

[0253] For example, in the case of dividing each functional module according to each function, Figure 10 A possible constituent schematic diagram of the image processing apparatus 1000 involved in the above embodiments is shown, as shown in the figure, the image processing apparatus 1000 can include a transceiver module 1001 and a processing module 1002. Figure 10

[0254] In some embodiments, the image processing apparatus is an encoding end device or located on the encoding end device. Wherein, the transceiver module 1001 can be used to support the image processing apparatus 1000 to perform the step 413 shown in the above embodiments; and / or other actions or functions performed by the encoding end device in the above method embodiments. The processing module 1002 can be used to support the image processing apparatus 1000 to perform the step 409 and the step 413 shown in the above embodiments; and / or other actions or functions performed by the encoding end device in the above method embodiments. Figure 4 Figure 4

[0255] In other embodiments, the image processing apparatus is a decoding end device or located on the decoding end device. Wherein, the transceiver module 1001 can be used to support the image processing apparatus 1000 to perform the step 402, the step 405 and the step 408, the step 410, the step 412 and the step 414 shown in the above embodiments; and / or other actions or functions performed by the decoding end device in the above method embodiments. The processing module 1002 can be used to support the image processing apparatus 1000 to perform the step 401, the step 404, the step 407 and the step 410 shown in the above embodiments; and / or other actions or functions performed by the decoding end device in the above method embodiments. Figure 4 Figure 4

[0256] Wherein, all the related contents of each step involved in the above method embodiments can be cited to the function description of the corresponding functional module, which will not be repeated here.

[0257] Or, in other embodiments, the image processing apparatus is an encoding end device or located on the encoding end device. Wherein, the transceiver module 1001 can be used to support the image processing apparatus 1000 to perform the step 802 shown in the above embodiments; Figure 8 Figure 9 The step 903, the step 905 and the step 906 shown in the above embodiments; and / or other actions or functions performed by the encoding end device in the above method embodiments. The processing module 1002 can be used to support the image processing apparatus 1000 to perform the step 801 shown in the above embodiments; Figure 8 Figure 9 ​​​​​​​Steps 901 and 902 shown; and / or other actions or functions performed by the encoding end device in the above method embodiments.

[0258] In other embodiments, the image processing device is a decoding end device or is located on a decoding end device. The transceiver module 1001 can be used to support the image processing device 1000 to perform the above embodiments. Figure 8 Step 803 shown; Figure 9 The processing module 1002 can be used to support the image processing apparatus 1000 in performing the above-mentioned steps 904 and 907; and / or other actions or functions performed by the decoding end device in the above-mentioned method embodiment. Figure 9 Step 904 in the above method embodiment; and / or other actions or functions performed by the decoding end device in the above method embodiment.

[0259] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here.

[0260] In the embodiment of the present application, the image processing device 1000 is presented in the form of dividing various functional modules in an integrated manner. The "module" here can refer to a specific ASIC, circuit, processor and memory that executes one or more software or firmware programs, integrated logic circuit, and / or other devices that can provide the above functions. In a simple embodiment, those skilled in the art can imagine that the image processing device 1000 can be used Figure 3A The form shown.

[0261] for example, Figure 3A The processor 301 in the image processing apparatus 1000 can call the computer instructions stored in the memory 303 to enable the image processing apparatus 1000 to execute the actions executed by the terminal device in the above method embodiment.

[0262] Specifically, Figure 10 The functions / implementation processes of the transceiver module 1001 and the processing module 1002 can be realized by Figure 3A The processor 301 in the embodiment calls the computer instructions stored in the memory 303 to implement. Or, Figure 10 The function / implementation process of the transceiver module 1001 can be achieved by Figure 3A The transmission interface 302 is implemented in Figure 10 The function / implementation process of the processing module 1002 can be achieved by Figure 3A The processor 301 in the memory 303 calls the computer instructions stored in the memory to implement it.

[0263] Optionally, embodiments of the present application further provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an image processing device, the image processing device executes the aforementioned related method steps to implement the image processing method in the aforementioned embodiment. For example, the image processing device may be the encoding device in the aforementioned method embodiment. Alternatively, the communication device may be the decoding device in the aforementioned method embodiment.

[0264] Optionally, embodiments of the present application further provide a computer program product that, when executed on a computer, causes the computer to execute the aforementioned steps to implement the image processing method performed by the image processing apparatus in the aforementioned embodiment. For example, the image processing apparatus may be the encoding device in the aforementioned method embodiment. Alternatively, the image processing apparatus may be the decoding device in the aforementioned method embodiment.

[0265] Optionally, embodiments of the present application further provide an image processing device, which may be a chip, component, module, or system-on-chip. The device may include a connected processor and memory; the memory is used to store computer instructions. When the device is running, the processor may execute the computer instructions stored in the memory, causing the chip to perform the image processing methods performed by the communication device in each of the above-mentioned method embodiments. For example, the image processing device may be the encoding device in the above-mentioned method embodiments. Alternatively, the image processing device may be the decoding device in the above-mentioned method embodiments.

[0266] Optionally, an embodiment of the present application further provides an image processing system, which includes an encoding end device and a decoding end device. The encoding end device and the decoding end device in the image processing system can respectively execute the image processing methods executed by the encoding end device and the decoding end device in the above embodiments. For example, the architectural diagram of the image processing system can be found in Figure 2 .

[0267] Among them, the image processing device, computer-readable storage medium, computer program product, chip or system-on-chip provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0268] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0269] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of the modules or units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0270] The units described as separate components can or can not be physically separate, and the components shown as units can be one physical unit or multiple physical units, that is, can be located in one place or distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0271] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0272] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage medium that can store program codes.

[0273] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image processing method, characterized in that: include: When receiving a first image in a video stream, determining whether the first image meets a first preset condition according to a base layer of a first sub-image of the first image; The first image includes a plurality of sub-images, and the first sub-image is the first sub-image received by the decoding end among the plurality of sub-images; The first preset condition includes: a base layer of a first sub-image of the first image is lost or partially lost; If the first image meets the first preset condition, the second image will be displayed within the display period of the first image, and the second image is the image in the video stream displayed within the first display period, and the first display period is the image display period before the display period of the first image; each frame image in the video stream includes multiple sub-images, and each of the sub-images includes a basic layer.

2. The method according to claim 1, characterized in that The method further comprises: When a scene switch occurs, determining whether the first image satisfies a second preset condition based on a base layer of a first sub-image of the first image; wherein the second preset condition includes: a ratio of an amount of base layer-encoded data of the first sub-image of the first image received within a first preset time duration to a first channel bandwidth is greater than or equal to a first preset value, and a ratio of an amount of base layer-encoded data of the first sub-image of the first image received within the first preset time duration to a second channel bandwidth is greater than or equal to a second preset value; The first channel bandwidth is the average bandwidth within the first preset time period, the second channel bandwidth is the average bandwidth within the second preset time period, and the second preset time period is greater than the first preset time period, and the second preset time period is obtained by extending the first preset time period forward or backward along the time axis; If the first image meets the second preset condition, the second image is displayed within the display period of the first image.

3. The method according to claim 1, characterized in that Each of the sub-images further comprises at least one enhancement layer, and the method further comprises: When a scene switch occurs, determining whether the first image satisfies a third preset condition based on the enhancement layer of the first sub-image of the first image; wherein the third preset condition includes: the number of enhancement layers of the first sub-image of the first image received is less than or equal to a third preset value, or the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within a third preset time duration to a third channel bandwidth is greater than or equal to a fourth preset value, and the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within the third preset time duration to a fourth channel bandwidth is greater than or equal to a fifth preset value; The third channel bandwidth is the average bandwidth within the third preset time period, the fourth channel bandwidth is the average bandwidth within the fourth preset time period, and the fourth preset time period is greater than the third preset time period, and the fourth preset time period is obtained by extending the third preset time period forward or backward along the time axis; If the first image meets the third preset condition, the target mode is entered, the second image is displayed in the target mode, and the received sub-image of the first image is stored in the cache.

4. The method according to claim 3, characterized in that When entering the target mode, the method further includes: Sending first indication information to the encoding end to instruct the encoding end to enter the target mode.

5. The method according to claim 4, characterized in that The method further comprises: When the number of enhancement layers of each sub-image of the received first image is greater than or equal to a sixth preset value, exiting the target mode to display the latest successfully decoded image; Sending second indication information to the encoding end to instruct the encoding end to exit the target mode.

6. An image processing method, characterized in that: include: In a process of sending a first image in a video stream to a decoding end, upon receiving first indication information from the decoding end, entering a target mode, and adjusting encoding parameters in the target mode to reduce an amount of data after encoding of sub-images following a first sub-image in the first image; The first image is used by the decoding end to determine whether the first image meets a first preset condition based on the base layer of the first sub-image; the first image includes multiple sub-images, and the first sub-image is the first sub-image received by the decoding end among the multiple sub-images; the first preset condition includes: the base layer of the first sub-image of the first image is lost or partially lost; if the first image meets the first preset condition, a second image is displayed within a display period of the first image, the second image being an image in the video stream displayed within a first display period, and the first display period is an image display period before the display period of the first image; each frame of the video stream includes multiple sub-images, and each of the sub-images includes a base layer; setting a reference frame for inter-frame coding of an image subsequent to the first image as the first image; After receiving the second instruction information from the decoding end, the target mode is exited to stop sending the first image to the decoding end, and the image in the video stream that is most recently successfully encoded is sent to the decoding end.

7. An image processing method, characterized in that: include: When a scene switch occurs, determining, according to a base layer of a first sub-image of a first image in a video stream, whether the first image satisfies a fourth preset condition; The first image includes a plurality of sub-images, and the first sub-image is the first sub-image received by the decoding end among the plurality of sub-images; The fourth preset condition includes: a ratio of an amount of base layer encoded data of a first sub-image of the first image successfully sent within a fifth preset time duration to a fifth channel bandwidth is greater than a seventh preset value, and a ratio of an amount of base layer encoded data of a first sub-image of the first image successfully sent within the fifth preset time duration to a sixth channel bandwidth is greater than an eighth preset value; The fifth channel bandwidth is the average bandwidth within the fifth preset time period, the sixth channel bandwidth is the average bandwidth within the sixth preset time period, and the sixth preset time period is greater than the fifth preset time period, and the sixth preset time period is obtained by extending the fifth preset time period forward or backward along the time axis; If the first image satisfies the fourth preset condition, a third indication message is sent to the decoding end to instruct the decoding end to display the second image within the display period of the first image, where the second image is an image in the video stream displayed within the first display period, and the first display period is an image display period before the display period of the first image; each frame image in the video stream includes multiple sub-images, and each of the sub-images includes a basic layer.

8. The method according to claim 7, characterized in that Each of the sub-images further comprises at least one enhancement layer, and the method further comprises: When a scene switch occurs, determining whether the first image satisfies a fifth preset condition based on the enhancement layer of the first sub-image of the first image; wherein the fifth preset condition includes: the number of enhancement layers of the first sub-image of the first image that are successfully sent is less than or equal to a ninth preset value, or the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image that is successfully sent within a seventh preset time duration to a seventh channel bandwidth is greater than or equal to a tenth preset value, and the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image that is successfully sent within the seventh preset time duration to an eighth channel bandwidth is greater than or equal to an eleventh preset value; The seventh channel bandwidth is the average bandwidth within the seventh preset time period, the eighth channel bandwidth is the average bandwidth within the eighth preset time period, the eighth preset time period is greater than the seventh preset time period, and the eighth preset time period is obtained by extending the seventh preset time period forward or backward along the time axis; If the first image satisfies the fifth preset condition, entering the target mode, and sending fourth indication information to the decoding end in the target mode to instruct the decoding end to enter the target mode; Adjusting encoding parameters in the target mode to reduce the amount of data after encoding other sub-images following the first sub-image in the first image; A reference frame for inter-frame coding of an image subsequent to the first image is set as the first image.

9. The method according to claim 8, characterized in that The method further comprises: If the number of enhancement layers of each sub-image of the successfully transmitted first image is greater than or equal to a twelfth preset value, exit the target mode to stop transmitting the first image to the decoding end and transmit the image in the most recently successfully encoded video stream to the decoding end; Sending fifth indication information to the decoding end to instruct the decoding end to exit the target mode.

10. An image processing method, characterized in that: include: During a process of receiving a first image in a video stream from an encoder, if third instruction information is received from the encoder, display a second image within a display period of the first image, where the second image is an image in the video stream displayed within the first display period, and the first display period is an image display period preceding the display period of the first image; The third indication information is indication information sent by the encoding end to the decoding end when it is determined that the first image meets a fourth preset condition; The first image includes a plurality of sub-images, and the first sub-image is the first sub-image received by the decoding end among the plurality of sub-images; The fourth preset condition includes: a ratio of an amount of base layer encoded data of a first sub-image of the first image successfully sent within a fifth preset time duration to a fifth channel bandwidth is greater than a seventh preset value, and a ratio of an amount of base layer encoded data of a first sub-image of the first image successfully sent within the fifth preset time duration to a sixth channel bandwidth is greater than an eighth preset value; Among them, the fifth channel bandwidth is the average bandwidth within the fifth preset time length, the sixth channel bandwidth is the average bandwidth within the sixth preset time length, and the sixth preset time length is greater than the fifth preset time length, and the sixth preset time length is obtained by extending the fifth preset time length forward or backward along the time axis.

11. The method according to claim 10, characterized in that The method further comprises: If fourth indication information is received from the encoding end, entering a target mode, displaying the second image in the target mode, and storing the received sub-image of the first image in a cache; After receiving the fifth indication information from the encoding end, the target mode is exited to display the latest successfully decoded image.

12. An image processing device, characterized in that: Includes a transceiver module and a processing module; In which, the processing module is used to: when receiving the first image in the video stream through the transceiver module, determine whether the first image meets the first preset condition based on the basic layer of the first sub-image of the first image; the first image includes multiple sub-images, and the first sub-image is the first sub-image received by the decoding end among the multiple sub-images; the first preset condition includes: the basic layer of the first sub-image of the first image is lost or partially lost; the processing module is also used to: if the first image meets the first preset condition, then send the second image for display through the transceiver module within the display period of the first image, the second image is the image in the video stream sent for display within the first display period, and the first display period is the image display period before the display period of the first image; each frame of the image in the video stream includes multiple sub-images, and each of the sub-images includes a basic layer.

13. The device according to claim 12, characterized in that The processing module is further configured to: when a scene switch occurs, determine, based on a base layer of a first sub-image of the first image, whether the first image satisfies a second preset condition; wherein the second preset condition includes: a ratio of an amount of base layer-encoded data of the first sub-image of the first image received within a first preset time duration to a first channel bandwidth is greater than or equal to a first preset value, and a ratio of an amount of base layer-encoded data of the first sub-image of the first image received within the first preset time duration to a second channel bandwidth is greater than or equal to a second preset value; The first channel bandwidth is the average bandwidth within the first preset time period, the second channel bandwidth is the average bandwidth within the second preset time period, and the second preset time period is greater than the first preset time period, and the second preset time period is obtained by extending the first preset time period forward or backward along the time axis; The processing module is further configured to: if the first image satisfies the second preset condition, transmit and display the second image via the transceiver module within a display period of the first image.

14. The device according to claim 12, characterized in that Each of said sub-images further comprises at least one enhancement layer; The processing module is further configured to: when a scene switch occurs, determine, based on the enhancement layer of the first sub-image of the first image, whether the first image satisfies a third preset condition; wherein the third preset condition includes: the number of enhancement layers of the first sub-image of the first image received is less than or equal to a third preset value, or the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within a third preset time duration to a third channel bandwidth is greater than or equal to a fourth preset value, and the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image received within the third preset time duration to a fourth channel bandwidth is greater than or equal to a fifth preset value; The third channel bandwidth is the average bandwidth within the third preset time period, the fourth channel bandwidth is the average bandwidth within the fourth preset time period, and the fourth preset time period is greater than the third preset time period, and the fourth preset time period is obtained by extending the third preset time period forward or backward along the time axis; The processing module is further configured to enter a target mode if the first image satisfies the third preset condition, display the second image via the transceiver module in the target mode, and store the sub-image of the first image received via the transceiver module into a cache.

15. The device according to claim 14, characterized in that When entering the target mode, the processing module is further configured to: First indication information is sent to the encoding end through the transceiver module to instruct the encoding end to enter the target mode.

16. The device according to claim 15, characterized in that The processing module is further configured to: when the number of enhancement layers of each sub-image of the first image received by the transceiver module is greater than or equal to a sixth preset value, exit the target mode to send and display the latest successfully decoded image through the transceiver module; The processing module is further configured to: send second indication information to the encoding end through the transceiver module to instruct the encoding end to exit the target mode.

17. An image processing device, characterized in that: Includes a transceiver module and a processing module; The processing module is configured to: during the process of sending the first image in the video stream to the decoding end through the transceiver module, enter a target mode after receiving first instruction information from the decoding end through the transceiver module, and adjust encoding parameters in the target mode to reduce the amount of data after encoding of other sub-images following the first sub-image in the first image; The first image is used by the decoding end to determine whether the first image meets a first preset condition based on the base layer of the first sub-image; the first image includes multiple sub-images, and the first sub-image is the first sub-image received by the decoding end among the multiple sub-images; the first preset condition includes: the base layer of the first sub-image of the first image is lost or partially lost; if the first image meets the first preset condition, a second image is displayed within a display period of the first image, the second image being an image in the video stream displayed within a first display period, and the first display period is an image display period before the display period of the first image; each frame of the video stream includes multiple sub-images, and each of the sub-images includes a base layer; The processing module is further configured to: set a reference frame for inter-frame coding of an image subsequent to the first image as the first image; The processing module is further configured to, after receiving second indication information from the decoding end through the transceiver module, exit the target mode to stop sending the first image to the decoding end through the transceiver module and to send the image in the most recently successfully encoded video stream to the decoding end through the transceiver module.

18. An image processing device, characterized in that: Includes a transceiver module and a processing module; The processing module is configured to: when a scene switch occurs, determine, based on the base layer of a first sub-image of a first image in a video stream, whether the first image satisfies a fourth preset condition; the first image includes multiple sub-images, and the first sub-image is the first sub-image received by the decoding end among the multiple sub-images; the fourth preset condition includes: a ratio of an amount of base layer-encoded data of the first sub-image of the first image successfully sent within a fifth preset duration to a fifth channel bandwidth is greater than a seventh preset value, and a ratio of an amount of base layer-encoded data of the first sub-image of the first image successfully sent within the fifth preset duration to a sixth channel bandwidth is greater than an eighth preset value; The fifth channel bandwidth is the average bandwidth within the fifth preset time period, the sixth channel bandwidth is the average bandwidth within the sixth preset time period, and the sixth preset time period is greater than the fifth preset time period, and the sixth preset time period is obtained by extending the fifth preset time period forward or backward along the time axis; The processing module is also used to: if the first image meets the fourth preset condition, send third indication information to the decoding end through the transceiver module to instruct the decoding end to display the second image within the display period of the first image, where the second image is an image in the video stream displayed within the first display period, and the first display period is an image display period before the display period of the first image; each frame image in the video stream includes multiple sub-images, and each of the sub-images includes a basic layer.

19. The device according to claim 18, characterized in that Each of said sub-images further comprises at least one enhancement layer, The processing module is further configured to: when a scene switch occurs, determine, based on the enhancement layer of the first sub-image of the first image, whether the first image satisfies a fifth preset condition; wherein the fifth preset condition includes: the number of enhancement layers of the first sub-image of the first image that are successfully sent is less than or equal to a ninth preset value, or the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image that is successfully sent within a seventh preset time duration to the seventh channel bandwidth is greater than or equal to a tenth preset value, and the ratio of the amount of data after encoding the enhancement layer of the first sub-image of the first image that is successfully sent within the seventh preset time duration to the eighth channel bandwidth is greater than or equal to an eleventh preset value; The seventh channel bandwidth is the average bandwidth within the seventh preset time period, the eighth channel bandwidth is the average bandwidth within the eighth preset time period, the eighth preset time period is greater than the seventh preset time period, and the eighth preset time period is obtained by extending the seventh preset time period forward or backward along the time axis; The processing module is further configured to: enter a target mode if the first image satisfies the fifth preset condition, and send fourth instruction information to the decoding end through the transceiver module in the target mode to instruct the decoding end to enter the target mode; The processing module is further configured to: adjust encoding parameters in the target mode to reduce the amount of data after encoding of other sub-images following the first sub-image in the first image; The processing module is further configured to set a reference frame for inter-frame coding of an image subsequent to the first image as the first image.

20. The device according to claim 19, characterized in that The processing module is further configured to: when the number of enhancement layers of each sub-image of the first image successfully transmitted through the transceiver module is greater than or equal to a twelfth preset value, exit the target mode to stop transmitting the first image to the decoding end through the transceiver module, and transmit the image in the video stream that is most recently successfully encoded to the decoding end through the transceiver module; The processing module is further configured to: send fifth indication information to the decoding end through the transceiver module to instruct the decoding end to exit the target mode.

21. An image processing device, characterized in that: Including processing module and transceiver module: The processing module is configured to: during a process of receiving a first image in a video stream from an encoding end through the transceiver module, if third indication information is received from the encoding end through the transceiver module, then display a second image through the transceiver module within a display period of the first image, where the second image is an image in the video stream displayed within a first display period, and the first display period is an image display period before a display period of the first image; The third indication information is indication information sent by the encoder to the decoder when determining that the first image satisfies a fourth preset condition; the first image includes multiple sub-images, and the first sub-image is the first sub-image received by the decoder among the multiple sub-images; the fourth preset condition includes: a ratio of an amount of base layer encoded data of the first sub-image of the first image successfully sent within a fifth preset time duration to a fifth channel bandwidth is greater than a seventh preset value, and a ratio of an amount of base layer encoded data of the first sub-image of the first image successfully sent within the fifth preset time duration to a sixth channel bandwidth is greater than an eighth preset value; Among them, the fifth channel bandwidth is the average bandwidth within the fifth preset time length, the sixth channel bandwidth is the average bandwidth within the sixth preset time length, and the sixth preset time length is greater than the fifth preset time length, and the sixth preset time length is obtained by extending the fifth preset time length forward or backward along the time axis.

22. The device according to claim 21, characterized in that The processing module is further configured to: if fourth indication information is received from the encoding end through the transceiver module, enter a target mode, display the second image through the transceiver module in the target mode, and store the sub-image of the first image received through the transceiver module into a cache; The processing module is further configured to: after receiving fifth indication information from the encoding end via the transceiver module, exit the target mode, so as to display the latest successfully decoded image via the transceiver module.

23. An image processing device, characterized in that: include: processor and transmission interface; The transmission interface is coupled to the processor; The transmission interface is used to receive or send images in a video stream, and the processor is configured to call software instructions in a memory to execute the image processing method according to any one of claims 1 to 11.

24. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on a computer or a processor, enable the computer or the processor to execute the image processing method according to any one of claims 1 to 11.

25. A computer program product, characterized in that When the computer program product is run on a computer or a processor, the computer or the processor is enabled to execute the image processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and device for decoding a bitstream

    US20130230108A1

  • REPLAYING OLD PACKETS FOR CONCEALING VIDEO DECODING ERRORS and VIDEO DECODING LATENCY ADJUSTMENT BASED ON WIRELESS LINK CONDITIONS

    US20160227257A1