Image processing method and device
By adjusting the code stream after image encoding according to network bandwidth in the cloud desktop system, the problem of lag in cloud desktop usage under insufficient bandwidth and high-delay networks is solved, achieving smoother image transmission and better user experience.
Patent Information
- Application Number
- CN202510064884.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-11-23
AI Technical Summary
Under networks with insufficient bandwidth and high latency, users will feel obvious lag when using cloud desktops, and they will even be unable to use cloud desktops normally.
By acquiring and encoding the current frame image, a code stream with a preset number of bytes is generated, the transmission time window of the image is calculated based on the predicted bandwidth, and when the transmission time window exceeds the maximum tolerance delay, the number of bytes of the code stream is adjusted to reduce the transmission delay.
In interactive scenarios, by adjusting the code stream after image encoding, the image transmission delay can be reduced, the lag caused by network congestion can be prevented in advance, and the user experience can be improved.
Smart Images

Figure CN120017828A_ABST
Abstract
Description
[0001] This invention is a divisional application with application number 2020113235034, application date November 23, 2020, and invention name “An image processing method and device”. Technical Field
[0002] The present disclosure relates to the field of image transmission, and in particular to an image processing method and device. Background Art
[0003] Cloud desktops have been widely used in all walks of life. The cloud desktop system runs the operating system desktop (i.e., cloud desktop) on a cloud server. Users only need a client and access to the network to access the server anytime and anywhere to operate their own private cloud desktop. However, in actual applications, the user experience of using the cloud desktop is closely related to the network. In a network with sufficient bandwidth and low latency, users can use the cloud desktop to achieve the same experience as using a local computer. In a network with insufficient bandwidth and high latency, users will experience obvious lag when using the cloud desktop, and may even be unable to use the cloud desktop normally. Summary of the invention
[0004] The disclosed embodiments provide an image processing method and device, which can solve the problem that users will experience obvious lag when using a cloud desktop, or even cannot use the cloud desktop normally, in a network with insufficient bandwidth and high latency. The technical solution is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:
[0006] Acquire a current frame image and encode the current frame image to generate a first code stream, where the number of bytes of the first code stream is a preset number of bytes;
[0007] Obtaining a predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth;
[0008] If it is determined that the receiving end device is currently in an interactive scenario, a first maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerable delay of each frame of the image, where the first preset maximum tolerable delay is the preset maximum tolerable delay in the interactive scenario;
[0009] If the predicted transmission time window is greater than the first maximum transmission time window, executing a first preset step, the first preset step comprising:
[0010] discarding the current frame image;
[0011] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0012] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0013] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0014] The image processing method provided by the embodiment of the present disclosure can adjust the number of bytes of the bitstream after image encoding according to the current predicted bandwidth of the network in an interactive scenario, and after adjusting the number of bytes of the bitstream after image encoding, obtain the predicted transmission time window of the image according to the number of bytes of the adjusted bitstream, and when the predicted transmission time window of the image is less than or equal to the maximum transmission time window, send the bitstream after image encoding to the receiving end device, and can reduce the transmission delay of the image by adjusting the number of bytes of the bitstream after image encoding. It can prevent the jamming phenomenon caused by network congestion in advance in the interactive scenario, and avoid the problem that in the interactive scenario, if the network bandwidth is insufficient and the delay is high, the user will feel obvious jamming when using the cloud desktop, or even cannot use the cloud desktop normally, thereby improving the user experience.
[0015] In one embodiment, the method further comprises:
[0016] If the predicted transmission time window is smaller than the first maximum transmission time window, executing a second preset step, wherein the second preset step includes:
[0017] Sending the first code stream to a receiving device;
[0018] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0019] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0020] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the first maximum transmission time window.
[0021] By executing the second preset step when the predicted transmission time window is less than the first maximum transmission time window, it is possible to gradually improve the bitstream after image encoding and improve picture clarity in an interactive scenario while the image can be transmitted smoothly. It is also possible to dynamically adjust the bitstream after image encoding to improve picture quality while meeting the requirements of image transmission delay, thereby achieving a better user experience.
[0022] In one embodiment, after obtaining the predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth, the method further includes:
[0023] If it is determined that the receiving end device is currently in a non-interactive scenario, a second maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time of each frame of the image, and a second preset maximum tolerable delay, where the second preset maximum tolerable delay is the preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay is greater than the first preset maximum tolerable delay;
[0024] If the predicted transmission time window is greater than the second maximum transmission time window, executing a third preset step, the third preset step comprising:
[0025] discarding the current frame image;
[0026] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0027] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0028] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0029] When it is determined that the receiving end device is currently in a non-interactive scenario, the second maximum transmission time window is obtained according to the average encoding time, average decoding time, average display time and the second preset maximum tolerance delay of each frame of the image, and when the predicted transmission time window is greater than the second maximum transmission time window, the third preset step is executed, so that the number of bytes of the encoded code stream of the image can be adjusted according to the current predicted bandwidth of the network in the non-interactive scenario, and after adjusting the number of bytes of the encoded code stream of the image, the predicted transmission time window of the image is obtained according to the number of bytes of the adjusted code stream, and when the predicted transmission time window of the image is less than or equal to the maximum transmission time window, the encoded code stream of the image is sent to the receiving end device, and the transmission delay of the image can be reduced by adjusting the number of bytes of the encoded code stream of the image, and the jamming phenomenon caused by network congestion can be prevented in advance in the non-interactive scenario, avoiding the problem that in the interactive scenario, if the network bandwidth is insufficient and the delay is high, the user will feel obvious jamming when using the cloud desktop, or even cannot use the cloud desktop normally, thereby improving the user experience.
[0030] In one embodiment, the method further comprises:
[0031] If the predicted transmission time window is smaller than the second maximum transmission time window, executing a fourth preset step, the fourth preset step comprising:
[0032] Sending the first code stream to a receiving device;
[0033] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0034] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0035] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the second maximum transmission time window.
[0036] By executing the fourth preset step when the predicted transmission time window is less than the second maximum transmission time window, it is possible to gradually improve the bitstream after image encoding and improve picture clarity in a non-interactive scenario while ensuring that the image can be transmitted smoothly. It is also possible to dynamically adjust the bitstream after image encoding to improve picture quality while meeting the requirements of image transmission delay, thereby achieving a better user experience.
[0037] In one embodiment, before acquiring the current frame image and encoding the current frame image, the method further includes:
[0038] Acquire at least one frame of image;
[0039] Encode each frame of the at least one frame of image and obtain the encoding time of each frame of image;
[0040] After encoding each frame of the at least one frame of the image, the at least one frame of the image is sent to a receiving end device and an average encoding time of each frame of the image is obtained according to the encoding time of each frame of the image.
[0041] By encoding each frame of at least one frame of image acquired and obtaining the encoding time of each frame of image before encoding the current frame of image, the encoding time of each frame of image can be accurately obtained, and then the first maximum transmission time window and the second maximum transmission time window can be obtained according to the encoding time of each frame of image.
[0042] Maximum transmission time window.
[0043] In one embodiment, before acquiring the current frame image and encoding the current frame image, the method further includes:
[0044] The average decoding time and the average display time of each frame of the image sent by the receiving end device are received, wherein the average decoding time and the average display time are obtained after the receiving end device decodes and displays each frame of the at least one frame of the image after receiving the at least one frame of the image.
[0045] By receiving the average decoding time and average display time of each frame image sent by the receiving end device before encoding the current frame image, the first maximum transmission time window and the second maximum transmission time window can be obtained according to the decoding time and the average display time of each frame image.
[0046] In one embodiment, obtaining the predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth includes:
[0047] Wp=(P / B)*1000, where Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
[0048] The predicted transmission time window can be accurately calculated using the above formula.
[0049] In one embodiment, obtaining the first maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerable delay of each frame of the image comprises:
[0050] W1=T1-E-D-S, where W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0051] The first maximum transmission time window can be accurately calculated using the above formula.
[0052] In one embodiment, obtaining the second maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the second preset maximum tolerable delay of each frame of the image comprises:
[0053] W2=T2-E-D-S, where W2 is the first maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0054] The second maximum transmission time window can be accurately calculated using the above formula.
[0055] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0056] A current frame image acquisition module, used for acquiring a current frame image and encoding the current frame image to generate a first code stream, wherein the number of bytes of the first code stream is a preset number of bytes;
[0057] A predicted transmission time window generating module, used for obtaining the predicted transmission time window of the current frame image according to the first bit stream and the current predicted bandwidth;
[0058] A first maximum transmission time window generating module is used to obtain a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time and a first preset maximum tolerable delay of each frame of an image if it is determined that the receiving end device is currently in an interactive scene, wherein the first preset maximum tolerable delay is a preset maximum tolerable delay in the interactive scene;
[0059] A first preset step execution module is configured to execute a first preset step if the predicted transmission time window is greater than the first maximum transmission time window, wherein the first preset step includes:
[0060] discarding the current frame image;
[0061] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0062] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0063] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0064] In one embodiment, the apparatus further comprises:
[0065] A second preset step execution module is configured to execute a second preset step if the predicted transmission time window is smaller than the first maximum transmission time window, wherein the second preset step includes:
[0066] Sending the first code stream to a receiving device;
[0067] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0068] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0069] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the first maximum transmission time window.
[0070] In one embodiment, the apparatus further comprises:
[0071] A second maximum transmission time window generating module is used to obtain a second maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame of image and a second preset maximum tolerable delay if it is determined that the receiving end device is currently in a non-interactive scenario, wherein the second preset maximum tolerable delay is a preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay is greater than the first preset maximum tolerable delay;
[0072] A third preset step execution module is configured to execute a third preset step if the predicted transmission time window is greater than the second maximum transmission time window, wherein the third preset step includes:
[0073] discarding the current frame image;
[0074] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0075] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0076] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0077] In one embodiment, the apparatus further comprises:
[0078] A fourth preset step execution module is used to execute a fourth preset step if the predicted transmission time window is less than the second maximum transmission time window, and the fourth preset step includes:
[0079] Sending the first code stream to a receiving device;
[0080] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0081] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0082] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the second maximum transmission time window.
[0083] In one embodiment, the apparatus comprises:
[0084] Average encoding time acquisition module, used for:
[0085] Acquire at least one frame of image;
[0086] Encode each frame of the at least one frame of image and obtain the encoding time of each frame of image;
[0087] After encoding each frame of the at least one frame of the image, the at least one frame of the image is sent to a receiving end device and an average encoding time of each frame of the image is obtained according to the encoding time of each frame of the image.
[0088] In one embodiment, the apparatus comprises:
[0089] An average decoding time receiving module is used to receive the average decoding time and average display time of each frame of the image sent by the receiving device, wherein the average decoding time and the average display time are obtained after the receiving device decodes and displays each frame of the at least one frame of the image after receiving the at least one frame of the image.
[0090] In one embodiment, the predicted transmission time window generation module is used to:
[0091] Wp=(P / B)*1000, where Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
[0092] In one embodiment, the first maximum transmission time window generating module is used to:
[0093] W1=T1-E-D-S, where W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0094] In one embodiment, the second maximum transmission time window generating module is used to:
[0095] W2=T2-E-D-S, where W2 is the first maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0096] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one computer instruction, and the instruction is loaded and executed by the processor to implement the steps performed in the image processing method described in any one of the first aspects.
[0097] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, wherein at least one computer instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the steps performed in the image processing method described in any one of the first aspects.
[0098] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0100] Figure 1 is a structural schematic diagram of an image processing system provided by an embodiment of the present disclosure;
[0101] Figure 2 This is a process of an image processing method provided by the embodiment of the present disclosure. Figure 1 ;
[0102] Figure 3 is a schematic diagram of a cloud desktop system provided by an embodiment of the present disclosure;
[0103] Figure 4 This is a process of an image processing method provided by the embodiment of the present disclosure. Figure 2 ;
[0104] Figure 5 The structure of an image processing device provided by the embodiment of the present disclosure is shown in FIG. Figure 1 ;
[0105] Figure 6 The structure of an image processing device provided by the embodiment of the present disclosure is shown in FIG. Figure 2 ;
[0106] Figure 7 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0107] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0108] Figure 1 Schematic diagram of the structure of an image processing system provided by an embodiment of the present disclosure. Figure 1 As shown, the system includes a sending device 101 and a receiving device 102. The sending device 101 and the receiving device 102 can be connected for communication, and the sending device 101 can transmit the acquired image to the receiving device 102 through the network. The image processing system can be applied to a cloud desktop system or to other image transmission scenarios, which is not limited in this embodiment.
[0109] When the image processing system is applied to a cloud desktop system, the source device in the server can be used as an image sending device, and the client device can be used as a receiving device. The source device obtains the current frame image (i.e., the current display screen of the cloud desktop) and encodes the current frame image to generate a first code stream, the number of bytes of the first code stream is a preset number of bytes; then the predicted transmission time window of the current frame image is obtained according to the first code stream and the current predicted bandwidth of the network; if it is determined that the client device is currently in an interactive scene, the first maximum transmission time window is obtained according to the average encoding time, average decoding time, average display time and the first preset maximum tolerable delay of each frame image (i.e., each frame of the display screen of the cloud desktop), and the first preset maximum tolerable delay is the preset maximum tolerable delay in the interactive scene; if the predicted transmission time window is greater than the first maximum transmission time window, the preset steps are executed, and the first preset steps include: 1. discarding the current frame image; 2. obtaining the next frame image (i.e., the next frame of the display screen of the cloud desktop) and encoding the next frame image to generate a second code stream, the number of bytes of the second code stream is the number of bytes after the number of bytes of the first code stream is reduced according to the preset algorithm; 3. taking the next frame image as the new current frame image and taking the second code stream as the new first code stream. When the predicted transmission time window is less than or equal to the first maximum transmission time window, the first code stream is sent to the client device.
[0110] The image processing system provided by the embodiment of the present disclosure can adjust the number of bytes of the image-encoded code stream according to the current predicted bandwidth of the network when the receiving device is in an interactive scenario, and send the image-encoded code stream to the receiving device when the predicted transmission time window of the image is less than or equal to the maximum transmission time window. This avoids the problem that, in an interactive scenario, if the network bandwidth is insufficient and the latency is high, the user will experience obvious lag when using the cloud desktop, or even be unable to use the cloud desktop normally, thereby improving the user experience.
[0111] Combine the following Figure 2 The embodiment further describes in detail how the image processing system provided by the embodiment of the present disclosure performs image processing. Figure 2 FIG. 1 is a flowchart of an image processing method provided by an embodiment of the present disclosure. Figure 2 As shown, the method includes:
[0112] S201, obtaining a current frame image and encoding the current frame image to generate a first code stream, wherein the number of bytes of the first code stream is a preset number of bytes.
[0113] In this embodiment, before obtaining the current frame image, at least one frame image is first obtained; then each frame image in the at least one frame image is encoded and the encoding time of each frame image is obtained; after encoding each frame image in the at least one frame image, the at least one frame image is sent to the receiving end device and the average encoding time of each frame image is obtained based on the encoding time of each frame image.
[0114] Furthermore, the average decoding time and the average display time of each frame of the image sent by the receiving device are received. The average decoding time and the average display time are obtained after the receiving device decodes and displays each frame of the at least one frame of the image after receiving the at least one frame of the image.
[0115] For example, in a cloud desktop system, before obtaining the current display screen, the source device obtains several frames of display screens transmitted within a preset time length (for example, 1s), and encodes each frame of the display screen in the several frames and obtains the encoding time of each frame of the display screen; after encoding each frame of the display screen in the several frames, the several frames of the display screen are sent to the client device and the average encoding time of each frame of the image is obtained according to the encoding time of each frame of the image. For example, when the number of frames per second (FPS, Frames Per Second) of the picture is 30, that is, the FPS cycle is 1s, and the number of frames per second is 30. The encoding time of the first frame of the display screen is E1, the encoding time of the second frame of the display screen is E2..., the encoding time of the 30th frame of the display screen is E30, and the average encoding time E = (E1+E2+...+E30) / 30.
[0116] After receiving the several frames of display pictures, the client device decodes and displays each frame of the display pictures, then obtains the decoding time and display time of each frame of the display pictures, and then obtains the average decoding time and average display time of each frame of the display pictures based on the decoding time and display time of each frame of the display pictures.
[0117] For example, after receiving 30 frames of display pictures transmitted by the source device within the FPS period, the client device decodes each of the 30 frames of display pictures, and the decoding time of the first frame of display pictures is obtained as D1, the decoding time of the second frame of display pictures is obtained as D2, ..., the decoding time of the 30th frame of display pictures is D30, and the average decoding time D = (D1 + D2 + ... + D30) / 30. The client device decodes each of the 30 frames of display pictures and displays them, and the display time of the first frame of display pictures is obtained as S1, the display time of the second frame of display pictures is obtained as S2, ..., the display time of the 30th frame of display pictures is S30, and the average decoding time S = (S1 + S2 + ... + S30) / 30. After obtaining the average decoding time and the average display time, the client device sends the average decoding time and the average display time to the source device.
[0118] Exemplarily, after obtaining the average encoding time and receiving the average decoding time and the average display time sent by the receiving device, the current frame image is encoded to obtain a first code stream, and the number of bytes of the first code stream is a preset number of bytes. For example, in this embodiment, after encoding the current frame image, the number of bytes of the first code stream obtained is 1000 bytes.
[0119] S202: Obtain a predicted transmission time window of the current frame image according to the first code stream and the current predicted bandwidth.
[0120] In this step, the predicted transmission time window of the current frame image can be calculated by formula (1).
[0121] Wp= (P / B) * 1000 (1).
[0122] Wherein, Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth (unit: Bps). In this implementation, while acquiring the current frame image, the current predicted bandwidth B of the network is acquired. Any bandwidth prediction method in the prior art can be used to acquire the current predicted bandwidth of the network, and this embodiment is not specifically limited here.
[0123] S203. If it is determined that the receiving device is currently in an interactive scenario, a maximum transmission time window is obtained according to the average encoding time, average decoding time, average display time and a first preset maximum tolerable delay of each frame of the image, where the first preset maximum tolerable delay is the preset maximum tolerable delay in the interactive scenario.
[0124] In one embodiment, if it is determined that the receiving device is currently in an interactive scene, the average encoding time of each frame of the image is obtained and the average decoding time and average display time of each frame of the image sent by the receiving device are received, and the first maximum transmission time window is calculated by formula (2).
[0125] W1=T1 – E – D – S (2).
[0126] Wherein, W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time. After the transmitting device obtains the average encoding time E of each frame of the image and receives the average decoding time D and the average display time S of each frame of the image sent by the receiving device, the first maximum transmission time window W1 can be calculated according to formula (2). T1 represents the maximum delay that can be tolerated when each frame of the image is transmitted in an interactive scenario, that is, in a non-interactive scenario, the time taken from the acquisition of each frame of the image, encoding, transmission, decoding, to the final display completion cannot exceed T1. T1 is closely related to the network bandwidth. If the network bandwidth is large, T1 can be large. If the network bandwidth is small, T1 can be small.
[0127] In another embodiment, if it is determined that the receiving device is not currently in an interactive scenario, the maximum transmission time window is obtained based on the average encoding time, average decoding time, average display time and a second preset maximum tolerable delay for each frame of the image, and the second preset maximum tolerable delay is the preset maximum tolerable delay in the interactive scenario, and the second preset maximum tolerable delay is smaller than the first preset maximum tolerable delay.
[0128] Exemplarily, it is determined that the receiving device is not currently in an interactive scenario, then after obtaining the average encoding time of each frame of the image and receiving the average decoding time and average display time of each frame of the image sent by the receiving device, the second maximum transmission time window is calculated by formula (3).
[0129] W2=T2 – E – D – S (3).
[0130] Wherein, W2 is the second maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time. After the transmitting end device obtains the average encoding time E of each frame of image and receives the average decoding time D and the average display time S of each frame of image sent by the receiving end device, the second maximum transmission time window W2 can be calculated according to formula (3). T2 represents the maximum delay that can be tolerated when each frame of image is transmitted in a non-interactive scenario, that is, in an interactive scenario, the time taken from the acquisition of each frame of image to the encoding, transmission, decoding, and final display completion cannot exceed T2. T2 is closely related to the network bandwidth. If the network bandwidth is large, the value of T2 can be large, and if the network bandwidth is small, the value of T2 can be small. Wherein, T2 is greater than T1. In some embodiments, the value range of T1 is between 40ms-80ms, for example, 60ms, and the value range of T2 is between 80ms-150ms, for example, 100ms.
[0131] It should be noted here that the interaction scenario refers to the scenario in which the user performs human-computer interaction on the receiving device. For example, the interaction scenario may include a keyboard interaction scenario, a mouse interaction scenario, a touch interaction scenario, a voice interaction scenario, a gesture interaction scenario, a body analysis interaction scenario, or a facial analysis interaction scenario.
[0132] When the receiving device has mouse operation, it is defined as a mouse interaction scenario; when the receiving device has keyboard operation, it is defined as a keyboard interaction scenario; when the receiving device has touch operation, it is defined as a touch interaction scenario; when the receiving device has voice interaction, it is defined as a voice interaction scenario; when the receiving device has gesture interaction, it is defined as a gesture interaction scenario; when the receiving device has identity analysis (such as face recognition, fingerprint recognition, etc.), it is defined as an identity analysis interaction scenario.
[0133] Scenarios other than interactive scenarios are defined as non-interactive scenarios. The reason for distinguishing these two scenarios is that in interactive scenarios, users are more sensitive to delay than in non-interactive scenarios, so T2 is greater than T1.
[0134] S204: If the predicted transmission time window is greater than the first maximum transmission time window, execute the first preset step.
[0135] In this embodiment, if the receiving end device is currently in an interactive scenario and the predicted transmission time window is greater than the first maximum transmission time window, the first preset step is performed:
[0136] 1. Discard the current frame image;
[0137] 2. Obtaining a next frame of image and encoding the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0138] 3. Use the next frame image as a new current frame image and use the second code stream as a new first code stream.
[0139] 4. Obtain a first predicted transmission time window according to the new first code stream and the current predicted bandwidth and determine whether the predicted transmission time window is greater than the first maximum transmission time window.
[0140] For example, when the number of bytes P of the first code stream of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the number of bytes of the first code stream and the current predicted bandwidth is 20ms, and the first maximum transmission time window W1 is 10ms, then the first preset step is performed: 1. The current frame image is discarded; 2. The next frame image is obtained and the next frame image is encoded to generate a second code stream, the number of bytes of the second code stream is 900 bytes, which is the number of bytes after the number of bytes of the first code stream 1000 is reduced by 10%; 3. The next frame image is used as the new current frame image and the second code stream is used as the new first code stream. 4. The predicted transmission time window Wp is obtained according to the number of bytes of the new first code stream 900 and the current predicted bandwidth, and it is determined whether the predicted transmission time window is greater than the first maximum transmission time window. If so, the first preset step is executed repeatedly until the predicted transmission time window is less than or equal to the first maximum transmission time window, and the first code stream is sent to the receiving end device so that the receiving end device decodes the first code stream to generate the current frame image.
[0141] In another implementation, if the receiving end device is currently in an interactive scenario and the predicted transmission time window is smaller than the first maximum transmission time window, a second preset step is performed, the second preset step comprising:
[0142] 1. Send the first code stream to a receiving device;
[0143] 2. Obtaining a next frame of image and encoding the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0144] 3. Use the next frame image as a new current frame image and use the third code stream as a new first code stream;
[0145] 4. A predicted transmission time window is obtained according to the new first bitstream and the current predicted bandwidth, and it is determined whether the predicted transmission time window is smaller than the first maximum transmission time window.
[0146] For example, when the number of bytes P of the first code stream of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the number of bytes of the first code stream and the current predicted bandwidth is 20ms, and the first maximum transmission time window W1 is 30ms, then the second preset step is performed: 1. The first code stream is sent to the receiving end device so that the receiving end device decodes the first code stream to generate the current frame image; 2. The next frame image is obtained and the next frame image is encoded to generate a third code stream, the number of bytes of the third code stream is 1100 bytes, which is the number of bytes after the number of bytes of the first code stream 1000 is increased by 10%; 3. The next frame image is used as the new current frame image and the third code stream is used as the new first code stream. 4. The predicted transmission time window Wp is obtained according to the number of bytes of the new first code stream 1100 and the current predicted bandwidth, and it is determined whether the predicted transmission time window is less than the first maximum transmission time window. If so, the second preset step is executed repeatedly until the predicted transmission time is equal to the maximum transmission time window.
[0147] Exemplarily, if the receiving end device is currently in a non-interactive scenario and the predicted transmission time window is greater than the second maximum transmission time window, the third preset step is executed.
[0148] In this embodiment, the third preset step includes:
[0149] 1. Discard the current frame image;
[0150] 2. Obtaining a next frame of image and encoding the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0151] 3. Use the next frame image as a new current frame image and use the second code stream as a new first code stream.
[0152] 4. Obtain a first predicted transmission time window according to the new first code stream and the current predicted bandwidth and determine whether the predicted transmission time window is greater than the second maximum transmission time window.
[0153] For example, when the number of bytes P of the first code stream of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the number of bytes of the first code stream and the current predicted bandwidth is 20ms, and the first maximum transmission time window W2 is 12ms, then the first preset step is performed: 1. The current frame image is discarded; 2. The next frame image is obtained and the next frame image is encoded to generate a second code stream, the number of bytes of the second code stream is 900 bytes, which is the number of bytes after the number of bytes of the first code stream 1000 is reduced by 10%; 3. The next frame image is used as the new current frame image and the second code stream is used as the new first code stream. 4. The predicted transmission time window Wp is obtained according to the number of bytes of the new first code stream 900 and the current predicted bandwidth, and it is determined whether the predicted transmission time window is greater than the second maximum transmission time window. If so, the third preset step is performed repeatedly until the predicted transmission time window is less than or equal to the first maximum transmission time window, and the first code stream is sent to the receiving end device so that the receiving end device decodes the first code stream to generate the current frame image.
[0154] In another implementation, if the receiving end device is currently in a non-interactive scenario and the predicted transmission time window is smaller than the second maximum transmission time window, a fourth preset step is performed, and the fourth preset step includes:
[0155] 1. Send the first code stream to a receiving device;
[0156] 2. Obtaining a next frame of image and encoding the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0157] 3. Use the next frame image as a new current frame image and use the third code stream as a new first code stream;
[0158] 4. Obtain a predicted transmission time window according to the new first code stream and the current predicted bandwidth and determine whether the predicted transmission time window is smaller than the second maximum transmission time window.
[0159] For example, when the number of bytes P of the first code stream of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the number of bytes of the first code stream and the current predicted bandwidth is 20ms, and the second maximum transmission time window W2 is 32ms, then the fourth preset step is performed: 1. The first code stream is sent to the receiving end device so that the receiving end device decodes the first code stream to generate the current frame image; 2. The next frame image is obtained and the next frame image is encoded to generate a third code stream, the number of bytes of the third code stream is 1100 bytes, which is the number of bytes after the number of bytes of the first code stream 1000 is increased by 10%; 3. The next frame image is used as the new current frame image and the third code stream is used as the new first code stream. 4. The predicted transmission time window Wp is obtained according to the number of bytes of the new first code stream 1100 and the current predicted bandwidth, and it is determined whether the predicted transmission time window is less than the second maximum transmission time window. If so, the fourth preset step is executed repeatedly until the predicted transmission time is equal to the second maximum transmission time window.
[0160] The image processing method provided by the embodiment of the present disclosure can adjust the number of bytes of the bitstream after image encoding according to the current predicted bandwidth of the network in an interactive scenario, and after adjusting the number of bytes of the bitstream after image encoding, obtain the predicted transmission time window of the image according to the number of bytes of the adjusted bitstream, and when the predicted transmission time window of the image is less than or equal to the maximum transmission time window, send the bitstream after image encoding to the receiving end device, and can reduce the transmission delay of the image by adjusting the number of bytes of the bitstream after image encoding. It can prevent the jamming phenomenon caused by network congestion in advance in the interactive scenario, and avoid the problem that in the interactive scenario, if the network bandwidth is insufficient and the delay is high, the user will feel obvious jamming when using the cloud desktop, or even cannot use the cloud desktop normally, thereby improving the user experience.
[0161] The image processing method provided by the embodiment of the present disclosure is further described in detail below.
[0162] This solution divides the application scenarios of cloud desktop into two scenarios: interactive scenarios and non-interactive scenarios.
[0163] When the client has mouse, keyboard, touch and other peripheral operations that cause the cloud desktop to change, it is defined as an interactive scene; otherwise, it is defined as a non-interactive scene. For these two different scenarios, the corresponding maximum tolerable delays T1 and T2 (unit: milliseconds) are defined respectively.
[0164] Definition of maximum tolerable delay: The time from the acquisition of the source image to the encoding, transmission, decoding, and final display completion cannot exceed the time.
[0165] For example, the maximum tolerable delay T1 is defined in an interactive scenario, and the maximum tolerable delay T2 is defined in a non-interactive scenario.
[0166] Wherein, T2 is greater than T1. In some embodiments, the value range of T1 is between 40ms-80ms, for example, 60ms, and the value range of the maximum tolerable delay T2 defined in the non-interactive scenario is between 80ms-150ms, for example, 100ms.
[0167] At the same time, the network bandwidth is estimated through the bandwidth prediction module to obtain the predicted bandwidth B (unit: Bps).
[0168] The statistics module is used to count the average encoding time E (unit: milliseconds) of all frames per second.
[0169] The average encoding time refers to the total time consumed for encoding N frames of images transmitted in a period T divided by N.
[0170] For example, in a cloud desktop, 30 frames are transmitted in each FPS cycle (usually 1s). The encoding time of the first frame is E1, the encoding time of the second frame is E2, and the encoding time of the 30th frame is E30. The average encoding time E = (E1+E2+...+E30) / 30.
[0171] According to the above method, the average decoding time D (unit: milliseconds) of all frames per second is counted.
[0172] The average encoding time refers to the total time consumed for decoding N frames of images transmitted in a period T (eg, each FPS period mentioned above) divided by N.
[0173] According to the above method, the average display time S (unit: milliseconds) of all frames within the period T is counted.
[0174] The average encoding time refers to the total time consumed for displaying N frames of images transmitted in a period T (eg, each FPS period mentioned above) divided by N.
[0175] For each frame, the size of the bitstream output after encoding by the encoder is defined as the output bitstream P (unit: byte).
[0176] Define (P / B)*1000 as the predicted transmission time window Wp, and Tn–E–DS as the maximum transmission time window Wr.
[0177] When (P / B)*1000>(Tn–E–D–S) (Tn is T1 or T2), the encoding end actively discards the current frame, dynamically reduces the encoder output bit rate (i.e., reduces the clarity), regenerates a new frame, and satisfies (P / B)*1000<=(Tn–E–DS) before transmission.
[0178] When (P / B)*1000<(Tn–E–DS), starting from the next frame, the encoder output bit rate P is dynamically increased (i.e., the clarity is improved), and (P / B)*1000<=(Tn–E–D–S) is satisfied before transmission.
[0179] Dynamic reduction or increase refers to adjusting the encoder's output bitstream parameters in proportion, such as 10% of the output bitstream size currently in use, and resetting the encoder.
[0180] When Wp is less than Wr, the picture can be transmitted smoothly, and under the premise that Wp is less than Wr, the encoder will gradually increase the output bit rate to improve the picture clarity. When Wp is greater than Wr, by adjusting the encoder output bit rate, packet loss and jamming caused by network congestion can be prevented in advance.
[0181] This solution distinguishes interactive and non-interactive scenarios and defines the maximum tolerable delay allowed in the two scenarios. Based on the predicted bandwidth, it dynamically calculates whether each frame of the encoded bitstream can reach the client within the maximum tolerable delay. If not, it actively drops frames and adjusts the encoder bitstream to meet the delay requirements. Conversely, while meeting the delay requirements, it dynamically adjusts the encoded bitstream to improve image quality and achieve a better user experience.
[0182] The invention focuses on setting different maximum tolerable delays in different scenarios. The examples given are interactive and non-interactive scenarios, which can be further refined into multiple types.
[0183] The application scenario of the present invention can be, for example, Figure 3 The cloud desktop system shown in the figure. The entire cloud desktop system consists of two parts: the source end and the client end. The source end generally refers to Figure 3 The client generally refers to a terminal device or a software system consisting of a decoder and a display module.
[0184] This solution divides the source screen of the cloud desktop into two scenarios: interactive scenario and non-interactive scenario.
[0185] The interaction scenario refers to the scenario in which the user performs human-computer interaction on the client.
[0186] For example, the interaction scenario may include a keyboard interaction scenario, a mouse interaction scenario, a touch interaction scenario, a voice interaction scenario, a gesture interaction scenario, a body analysis interaction scenario, or a facial analysis interaction scenario, etc.
[0187] When the client has mouse operation, it is defined as a mouse interaction scenario; when the client has keyboard operation, it is defined as a keyboard interaction scenario; when the client has touch operation, it is defined as a touch interaction scenario; when the client has voice interaction, it is defined as a voice interaction scenario; when the client has gesture interaction, it is defined as a gesture interaction scenario; when the client has identity analysis (such as face recognition, fingerprint recognition, etc.), it is defined as an identity analysis interaction scenario.
[0188] Scenarios other than interactive scenarios are defined as non-interactive scenarios. The reason for distinguishing between these two scenarios is that in interactive scenarios, users are more sensitive to delay than in non-interactive scenarios. When a user double-clicks to open a file, he hopes that the action will be executed immediately. In the cloud desktop system, the source screen generated by the double-click is transmitted to the client at the fastest speed after encoding, and decoded and displayed. In non-interactive scenarios, such as automatic loop playback of slides, even if the transmission delay is twice as long as in the interactive scenario, the user will not be aware of it, and it will not affect the user experience. In the case of fixed bandwidth, for the same source screen (before encoding), the smaller the bitstream after encoding, the faster the transmission, and the faster the user can see the changed screen. However, for the same screen with a fixed encoding algorithm, the smaller the bitstream, the worse the clarity. Therefore, there needs to be a balance between clarity and delay, that is, to transmit high-definition pictures as much as possible under the delay acceptable to users. Therefore, this solution divides the source screen into interactive and non-interactive scenarios, thereby defining different maximum tolerable delays T1 and T2. Then, based on different maximum tolerable delays, the encoder output bitstream is dynamically adjusted to obtain a picture of appropriate clarity, providing a better user experience.
[0189] The main workflow of this program is as follows Figure 4 As shown:
[0190] Step 1: The encoder, decoder and display module periodically calculate the average encoding time E, the average decoding time D and the average display time S respectively.
[0191] This periodicity is generally one second.
[0192] Step 2: At the source end, the bandwidth prediction module predicts the bandwidth in real time and outputs the predicted bandwidth value B;
[0193] Step 3: At the source end, each frame after collection is encoded by the encoder to obtain a bit stream size P;
[0194] Step 4: At the source end, use the formula (P / B)*1000>(Tn–E–D–S) to make a judgment, where Tn is selected as T1 or T2 according to different scenarios. If the formula calculates the result, execute step 5 or step 6 respectively;
[0195] Step 5: If the formula is not valid, it means that the current bitstream size meets the transmission delay in the current scenario, and is directly sent to the transmission module for transmission. At the same time, the encoder is notified to increase the output bitstream, that is, to improve the clarity.
[0196] The degree of improvement can be determined according to different encoding algorithms, for example, 10% each time;
[0197] Step 6: If the formula is established, it means that the current bitstream size does not meet the transmission delay in the current scenario, so the current bitstream is discarded and the encoder is notified to reduce the output bitstream, that is, reduce the definition, and then go back to step 3 and cycle again.
[0198] Usage scenario description (T2>T1):
[0199] The client is in a non-interactive scenario, and Tn is T2
[0200] At this time, the condition of (P / B)*1000≤(T2–E–D–S) is met, indicating that the current code stream size meets the transmission delay in the current scenario.
[0201] If the client state changes to interactive scene, Tn takes the value of T1
[0202] After the value of Tn changes, the condition of (P / B)*1000≤(T1–E–D–S) is not met, indicating that the current bitstream size does not meet the transmission delay in the current scenario. In this case, P should be appropriately reduced until the condition of (P / B)*1000≤(T2–E–D–S) is met.
[0203] The principle of the process of transitioning from an interactive state to a non-interactive state is the same.
[0204] Based on the above Figure 2 and Figure 4 The image processing method described in the corresponding embodiment is as follows: an embodiment of the device disclosed herein, which can be used to execute the embodiment of the method disclosed herein.
[0205] Figure 5 FIG. 1 is a schematic diagram of the structure of an image processing device provided by an embodiment of the present disclosure. Figure 5 As shown, the device 50 comprises:
[0206] The current frame image acquisition module 501 is used to acquire the current frame image and encode the current frame image to generate a first code stream, wherein the number of bytes of the first code stream is a preset number of bytes;
[0207] A predicted transmission time window generating module 502, configured to obtain a predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth;
[0208] A first maximum transmission time window generating module 503 is configured to obtain a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time and a first preset maximum tolerable delay of each frame of an image if it is determined that the receiving end device is currently in an interactive scenario, wherein the first preset maximum tolerable delay is a preset maximum tolerable delay in the interactive scenario;
[0209] The first preset step execution module 504 is configured to execute a first preset step if the predicted transmission time window is greater than the first maximum transmission time window, wherein the first preset step includes:
[0210] discarding the current frame image;
[0211] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0212] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0213] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0214] In one embodiment, Figure 6 As shown, the device 50 also includes:
[0215] The second preset step execution module 505 is configured to execute a second preset step if the predicted transmission time window is smaller than the first maximum transmission time window, wherein the second preset step includes:
[0216] Sending the first code stream to a receiving device;
[0217] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0218] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0219] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the first maximum transmission time window.
[0220] In one embodiment, the device 50 further comprises:
[0221] A second maximum transmission time window generating module 506 is configured to obtain a second maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame of an image and a second preset maximum tolerable delay if it is determined that the receiving end device is currently in a non-interactive scenario, wherein the second preset maximum tolerable delay is a preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay is greater than the first preset maximum tolerable delay;
[0222] The third preset step execution module 507 is configured to execute a third preset step if the predicted transmission time window is greater than the second maximum transmission time window. The third preset step includes:
[0223] discarding the current frame image;
[0224] Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0225] Using the next frame image as a new current frame image and using the second code stream as a new first code stream;
[0226] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0227] In one embodiment, the device 50 further comprises:
[0228] The fourth preset step execution module 508 is configured to execute the fourth preset step if the predicted transmission time window is smaller than the second maximum transmission time window. The fourth preset step includes:
[0229] Sending the first code stream to a receiving device;
[0230] Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0231] Using the next frame image as a new current frame image and using the third code stream as a new first code stream;
[0232] A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the second maximum transmission time window.
[0233] In one embodiment, the device 50 comprises:
[0234] The average encoding time acquisition module 509 is used to:
[0235] Acquire at least one frame of image;
[0236] Encode each frame of the at least one frame of image and obtain the encoding time of each frame of image;
[0237] After encoding each frame of the at least one frame of the image, the at least one frame of the image is sent to a receiving end device and an average encoding time of each frame of the image is obtained according to the encoding time of each frame of the image.
[0238] In one embodiment, the device 50 comprises:
[0239] The average decoding time receiving module 510 is used to receive the average decoding time and the average display time of each frame of the image sent by the receiving device, and the average decoding time and the average display time are obtained after the receiving device decodes and displays each frame of the at least one frame of the image after receiving the at least one frame of the image.
[0240] In one embodiment, the predicted transmission time window generation module 502 is used to:
[0241] Wp=(P / B)*1000, where Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
[0242] In one embodiment, the first maximum transmission time window generating module 503 is used to:
[0243] W1=T1-E-D-S, where W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0244] In one embodiment, the second maximum transmission time window generating module 506 is used to:
[0245] W2=T2-E-D-S, where W2 is the first maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0246] The image processing device provided in the embodiment of the present disclosure, its implementation process and technical effects can be seen in the above Figure 2 and Figure 4 The embodiments are not described in detail here.
[0247] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 7 As shown, the electronic device 70 includes a processor and a memory, wherein the memory stores at least one computer instruction, and the instruction is loaded and executed by the processor to implement Figure 2 and Figure 4 The steps performed in the image processing method described in the corresponding embodiment.
[0248] Based on the above Figure 2 and Figure 4 In accordance with the image processing method described in the embodiment, the embodiment of the present disclosure further provides a computer-readable storage medium. For example, the non-temporary computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, or an optical data storage device. The storage medium stores computer instructions for executing the above-mentioned Figure 2 and Figure 4 The image processing method described in the corresponding embodiment will not be repeated here.
[0249] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0250] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
Claims
1. An image processing method, characterized in that: include: Acquire a current frame image and encode the current frame image to generate a first code stream, where the number of bytes of the first code stream is a preset number of bytes; Obtaining a predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth; If it is determined that the receiving end device is currently in an interactive scenario, a first maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerable delay of each frame of the image, where the first preset maximum tolerable delay is the preset maximum tolerable delay in the interactive scenario; If the predicted transmission time window is smaller than the first maximum transmission time window, executing a second preset step, wherein the second preset step includes: Sending the first code stream to a receiving device; Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm; Using the next frame image as a new current frame image and using the third code stream as a new first code stream; Obtain a new predicted transmission time window according to the new first code stream and the current predicted bandwidth, and determine whether the new predicted transmission time window is smaller than the first maximum transmission time window; The step of obtaining the first maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerance delay of each frame of the image comprises: W1=T1-EDS, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
2. The method according to claim 1, characterized in that After obtaining the predicted transmission time window of the current frame image according to the first code stream and the current predicted bandwidth, the method further includes: If it is determined that the receiving end device is currently in a non-interactive scenario, a second maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time of each frame of the image, and a second preset maximum tolerable delay, where the second preset maximum tolerable delay is the preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay is greater than the first preset maximum tolerable delay; If the predicted transmission time window is greater than the second maximum transmission time window, executing a third preset step, the third preset step comprising: discarding the current frame image; Acquire a next frame of image and encode the next frame of image to generate a second code stream, wherein the number of bytes of the second code stream is the number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm; Using the next frame image as a new current frame image and using the second code stream as a new first code stream; Obtain a new predicted transmission time window according to the new first code stream and the current predicted bandwidth, and determine whether the new predicted transmission time window is greater than the second maximum transmission time window; The step of obtaining the second maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the second preset maximum tolerance delay of each frame of the image comprises: W2=T2-E-DS, where W2 is the first maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
3. The method according to claim 2, characterized in that The method further comprises: If the predicted transmission time window is smaller than the second maximum transmission time window, executing a fourth preset step, the fourth preset step comprising: Sending the first code stream to a receiving device; Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm; Using the next frame image as a new current frame image and using the third code stream as a new first code stream; A new predicted transmission time window is obtained according to the new first code stream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is smaller than the second maximum transmission time window.
4. The method according to claim 1, characterized in that Before acquiring the current frame image and encoding the current frame image, the method further includes: Acquire at least one frame of image; Encode each frame of the at least one frame of image and obtain the encoding time of each frame of image; After encoding each frame of the at least one frame of the image, the at least one frame of the image is sent to a receiving end device and an average encoding time of each frame of the image is obtained according to the encoding time of each frame of the image.
5. The method according to claim 4, characterized in that Before acquiring the current frame image and encoding the current frame image, the method further includes: The average decoding time and the average display time of each frame of the image sent by the receiving end device are received, wherein the average decoding time and the average display time are obtained after the receiving end device decodes and displays each frame of the at least one frame of the image after receiving the at least one frame of the image.
6. The method according to claim 1, characterized in that The step of obtaining the predicted transmission time window of the current frame image according to the first bitstream and the current predicted bandwidth includes: Wp=(P / B)*1000, where Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
7. An image processing device, characterized in that: include: A current frame image acquisition module, used for acquiring a current frame image and encoding the current frame image to generate a first code stream, wherein the number of bytes of the first code stream is a preset number of bytes; A predicted transmission time window generating module, used for obtaining the predicted transmission time window of the current frame image according to the first bit stream and the current predicted bandwidth; A first maximum transmission time window generating module is used to obtain a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time and a first preset maximum tolerable delay of each frame of an image if it is determined that the receiving end device is currently in an interactive scene, wherein the first preset maximum tolerable delay is a preset maximum tolerable delay in the interactive scene; A second preset step execution module is configured to execute a second preset step when the predicted transmission time window is less than the first maximum transmission time window, wherein the second preset step includes: Sending the first code stream to a receiving device; Acquire a next frame of image and encode the next frame of image to generate a third code stream, wherein the number of bytes of the third code stream is the number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm; Using the next frame image as a new current frame image and using the third code stream as a new first code stream; Obtain a new predicted transmission time window according to the new first code stream and the current predicted bandwidth, and determine whether the new predicted transmission time window is smaller than the first maximum transmission time window; The step of obtaining the first maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerance delay of each frame of the image comprises: W1=T1-EDS, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
Citation Information
Patent Citations
Video coding and network transmission method and video forwarding server
CN103475902A
Image data processing method and device
CN103929654A
Video encoding method and apparatus
CN106454355A
Image coding method and device, coding end equipment and storage medium
CN111954001A
Picture-encoding device and picture-transmission system using the same and quantization controlling method and mean through-put calculating method used for the same
JP1998028269A