An image processing method and apparatus
By dynamically adjusting the number of bytes in the encoded image stream and the transmission time window in the cloud desktop system, the problems of insufficient bandwidth and cloud desktop lag under high latency are solved, achieving smooth transmission and high-quality display in different scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN WANXIANG ELECTRONICS TECH CO LTD
- Filing Date
- 2020-11-23
- Publication Date
- 2026-04-10
AI Technical Summary
In networks with insufficient bandwidth and high latency, users will experience noticeable lag when using cloud desktops, or may even be unable to use cloud desktops normally.
By adjusting the number of bytes in the encoded image bitstream, the image transmission time window is dynamically adjusted in both interactive and non-interactive scenarios based on the current predicted network bandwidth. This avoids network congestion, ensures smooth image transmission, and improves user experience.
In both interactive and non-interactive scenarios, it reduces image transmission latency, avoids lag, and improves the user experience of cloud desktops.
Smart Images

Figure CN120017828B_ABST
Abstract
Description
[0001] The present application is a divisional application of application No. 2020113235034, titled "Image processing method and device", filed on November 23, 2020. TECHNICAL FIELD
[0002] The present disclosure relates to the field of image transmission, and in particular to an image processing method and device. BACKGROUND
[0003] Cloud desktops have been widely used in various industries. A cloud desktop system runs the desktop of an operating system (i.e., a cloud desktop) on a cloud server. A user only needs a client to access the network and can access the server anytime and anywhere to operate the private cloud desktop. SUMMARY
[0004] The present disclosure provides an image processing method and device, which can solve the problem that a user using a cloud desktop will experience obvious lag or even cannot normally use the cloud desktop in a network with insufficient bandwidth and high latency. The technical solution is as follows:
[0005] According to a first aspect of the present disclosure, an image processing method is provided, comprising:
[0006] obtaining a current frame image and encoding the current frame image to generate a first code stream, the number of bytes of the first code stream being a preset number of bytes;
[0007] obtaining a predicted transmission time window of the current frame image according to the first code stream and a current predicted bandwidth;
[0008] if it is determined that the receiving end device is currently in an interactive scenario, obtaining a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame image, and a first preset maximum tolerable delay, the first preset maximum tolerable delay being a preset maximum tolerable delay in the interactive scenario;
[0009] if the predicted transmission time window is greater than the first maximum transmission time window, performing a first preset step, the first preset step comprising:
[0010] discarding the current frame image;
[0011] acquire a next frame of image and encode the next frame of image to generate a second code stream, a byte number of the second code stream being a byte number after the byte number of the first code stream is reduced according to a preset first algorithm;
[0012] take the next frame of image as a new current frame of image and take the second code stream as a new first code stream;
[0013] obtain a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determine whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0014] The image processing method provided by the embodiments of the present disclosure can adjust the byte number of the code stream after image encoding according to the current predicted bandwidth of the network in an interactive scenario, and after adjusting the byte number of the code stream after image encoding, obtain a predicted transmission time window of the image according to the byte number of the adjusted code stream. When the predicted transmission time window of the image is less than or equal to the maximum transmission time window, the code stream after image encoding is sent to a receiving end device. The transmission delay of the image can be reduced by adjusting the byte number of the code stream after image encoding, which can prevent the phenomenon of lag caused by network congestion in advance in an interactive scenario, avoids the problem that in an interactive scenario, if the bandwidth of the network is insufficient and the delay is high, the user using the cloud desktop will feel obvious lag, or even cannot normally use the cloud desktop, and improves the user experience.
[0015] In one embodiment, the method further comprises:
[0016] if the predicted transmission time window is less than the first maximum transmission time window, a second preset step is performed, the second preset step comprising:
[0017] sending the first code stream to a receiving end device;
[0018] acquiring a next frame of image and encoding the next frame of image to generate a third code stream, a byte number of the third code stream being a byte number after the byte number of the first code stream is increased according to a preset second algorithm;
[0019] taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream;
[0020] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the first maximum transmission time window.
[0021] By performing the second preset step when the predicted transmission time window is less than the first maximum transmission time window, the code stream after image coding can be gradually improved, the picture definition can be improved, and the picture quality can be improved dynamically to meet the requirement of image transmission delay and achieve better user experience.
[0022] In one embodiment, after obtaining the predicted transmission time window of the current frame image according to the first code stream and the current predicted bandwidth, the method further comprises:
[0023] If it is determined that the receiving end device is currently in a non-interactive scenario, a second maximum transmission time window is obtained according to an average encoding time, an average decoding time, an average display time of each frame image, and a second preset maximum tolerable delay, the second preset maximum tolerable delay being a preset maximum tolerable delay in the non-interactive scenario, the second preset maximum tolerable delay being greater than the first preset maximum tolerable delay;
[0024] If the predicted transmission time window is greater than the second maximum transmission time window, a third preset step is performed, the third preset step comprising:
[0025] Discarding the current frame image;
[0026] Obtaining a next frame image and encoding the next frame image to generate a second code stream, the number of bytes of the second code stream being a number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0027] Taking the next frame image as a new current frame image and taking the second code stream as a new first code stream;
[0028] Obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0029] By determining that the receiving end device is currently in a non-interactive scene, then a second maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time and the second preset maximum tolerable delay of each frame of image, and when the predicted transmission time window is greater than the second maximum transmission time window, a third preset step is performed, which can adjust the byte quantity of the image coded stream according to the current predicted bandwidth of the network in the non-interactive scene, and after adjusting the byte quantity of the image coded stream, the predicted transmission time window of the image is obtained according to the byte quantity of the adjusted stream, and when the predicted transmission time window of the image is less than or equal to the maximum transmission time window, the image coded stream is sent to the receiving end device, which can reduce the transmission delay of the image by adjusting the byte quantity of the image coded stream, and can prevent the phenomenon of lag caused by network congestion in advance in the non-interactive scene, avoiding the problem that in the interactive scene, if the bandwidth of the network is insufficient and the delay is high, the user using the cloud desktop will feel obvious lag, and even cannot normally use the cloud desktop, thereby improving the user experience.
[0030] In one embodiment, the method further comprises:
[0031] If the predicted transmission time window is less than the second maximum transmission time window, a fourth preset step is performed, and the fourth preset step comprises:
[0032] sending the first stream to the receiving end device;
[0033] obtaining a next frame of image and encoding the next frame of image to generate a third stream, and the byte quantity of the third stream is the byte quantity increased from the byte quantity of the first stream according to a preset second algorithm;
[0034] taking the next frame of image as a new current frame of image and taking the third stream as a new first stream;
[0035] obtaining a new predicted transmission time window from the new first stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the second maximum transmission time window.
[0036] By performing the fourth preset step when the predicted transmission time window is less than the second maximum transmission time window, the image coded stream can be gradually improved to improve the picture definition under the premise that the image can be smoothly transmitted in the non-interactive scene, and the image coded stream can be dynamically adjusted to improve the picture quality under the requirement of image transmission delay, thereby achieving better user experience.
[0037] In one embodiment, before the obtaining a current frame of image and encoding the current frame of image, the method further comprises:
[0038] obtaining at least one frame of image;
[0039] encoding each of the at least one frame of image and obtaining an encoding time of each of the at least one frame of image;
[0040] after encoding each of the at least one frame of image, sending the at least one frame of image to a receiving end device and obtaining an average encoding time of each of the at least one frame of image according to the encoding time of each of the at least one frame of image.
[0041] By encoding each of the at least one frame of image obtained before encoding the current frame of image and obtaining the encoding time of each of the at least one frame of image, the encoding time of each of the at least one frame of image can be accurately obtained, and then the first maximum transmission time window and the second maximum transmission time window can be obtained according to the encoding time of each of the at least one frame of image.
[0042] In an embodiment, before the current frame of image is obtained and the current frame of image is encoded, the method further comprises:
[0043] receiving the average decoding time and the average display time of each of the at least one frame of image sent by the receiving end device, the average decoding time and the average display time being obtained by the receiving end device after receiving the at least one frame of image and decoding and displaying each of the at least one frame of image.
[0044] By receiving the average decoding time and the average display time of each of the at least one frame of image sent by the receiving end device before the current frame of image is encoded, the first maximum transmission time window and the second maximum transmission time window can be obtained according to the decoding time and the average display time of each of the at least one frame of image.
[0045] In an embodiment, the obtaining the predicted transmission time window of the current frame of image according to the first code stream and the current predicted bandwidth comprises:
[0046] Wp=(P / B) 1000, wherein Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
[0047] The predicted transmission time window can be accurately calculated by the above formula.
[0048] In an embodiment, the obtaining the first maximum transmission time window according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerant delay of each of the at least one frame of image comprises:
[0049] W1=T1–E–D–S, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerant delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0050] The first maximum transmission time window can be accurately calculated through the formula.
[0051] In one embodiment, the second maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time of each frame of image and the second preset maximum tolerable delay, and the second preset maximum tolerable delay is a preset maximum tolerable delay in the interactive scenario.
[0052] W2=T2-E-D-S, wherein W2 is the second maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0053] The second maximum transmission time window can be accurately calculated through the formula.
[0054] According to a second aspect of the embodiments of the present disclosure, an image processing device is provided, comprising:
[0055] A current frame image obtaining module is configured to obtain a current frame of image and encode the current frame of image to generate a first code stream, and the number of bytes of the first code stream is a preset number of bytes.
[0056] A predicted transmission time window generating module is configured to obtain a predicted transmission time window of the current frame of image according to the first code stream and a current predicted bandwidth.
[0057] A first maximum transmission time window generating module is configured to, if it is determined that a receiving end device is currently in an interactive scenario, obtain a first maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame of image and a first preset maximum tolerable delay, and the first preset maximum tolerable delay is a preset maximum tolerable delay in the interactive scenario.
[0058] A first preset step executing module is configured to, if the predicted transmission time window is greater than the first maximum transmission time window, execute a first preset step, and the first preset step comprises:
[0059] Discarding the current frame of image.
[0060] Obtaining a next frame of image and encoding the next frame of image to generate a second code stream, and the number of bytes of the second code stream is a number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm.
[0061] Taking the next frame of image as a new current frame of image and taking the second code stream as a new first code stream.
[0062] Obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0063] In one embodiment, the apparatus further comprises:
[0064] a second preset step execution module, configured to execute a second preset step if the predicted transmission time window is less than the first maximum transmission time window, the second preset step comprising:
[0065] sending the first code stream to a receiving end device;
[0066] obtaining a next frame of image and encoding the next frame of image to generate a third code stream, a byte quantity of the third code stream being a byte quantity increased from a byte quantity of the first code stream according to a preset second algorithm;
[0067] taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream;
[0068] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the first maximum transmission time window.
[0069] In one embodiment, the apparatus further comprises:
[0070] a second maximum transmission time window generation module, configured to, if it is determined that the receiving end device is currently in a non-interactive scenario, obtain a second maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame of image and a second preset maximum tolerable delay, the second preset maximum tolerable delay being a preset maximum tolerable delay in the non-interactive scenario, the second preset maximum tolerable delay being greater than the first preset maximum tolerable delay;
[0071] a third preset step execution module, configured to execute a third preset step if the predicted transmission time window is greater than the second maximum transmission time window, the third preset step comprising:
[0072] discarding the current frame of image;
[0073] obtaining a next frame of image and encoding the next frame of image to generate a second code stream, a byte quantity of the second code stream being a byte quantity reduced from a byte quantity of the first code stream according to a preset first algorithm;
[0074] taking the next frame of image as a new current frame of image and taking the second code stream as a new first code stream;
[0075] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0076] In one embodiment, the apparatus further comprises:
[0077] a fourth preset step execution module, configured to execute a fourth preset step if the predicted transmission time window is less than the second maximum transmission time window, the fourth preset step comprising:
[0078] sending the first code stream to a receiving end device;
[0079] obtaining a next frame of image and encoding the next frame of image to generate a third code stream, a byte number of the third code stream being a byte number obtained by increasing the byte number of the first code stream according to a preset second algorithm;
[0080] taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream;
[0081] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the second maximum transmission time window.
[0082] In one embodiment, the apparatus comprises:
[0083] an average encoding time obtaining module, configured to:
[0084] obtain at least one frame of image;
[0085] encode each frame of image in the at least one frame of image and obtain an encoding time of the each frame of image;
[0086] after encoding each frame of image in the at least one frame of image, send the at least one frame of image to a receiving end device and obtain an average encoding time of the each frame of image according to the encoding time of the each frame of image.
[0087] In one embodiment, the apparatus comprises:
[0088] an average decoding time receiving module, configured to receive an average decoding time and an average display time of the each frame of image sent by the receiving end device, the average decoding time and the average display time being obtained by the receiving end device after decoding and displaying each frame of image in the at least one frame of image after receiving the at least one frame of image.
[0089] In one embodiment, the predicted transmission time window generating module is configured to:
[0090] Wp = (P / B) 1000, wherein Wp is the predicted transmission time window, P is the byte number of the first code stream, and B is the current predicted bandwidth.
[0091] In one embodiment, the first maximum transmission time window generating module is configured to:
[0092] W1 = T1 - E - D - S, where W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0093] In one embodiment, the second maximum transmission time window generating module is configured to:
[0094] W2 = T2 - E - D - S, where W2 is the second maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0095] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which includes a processor and a memory, and the memory stores at least one computer instruction, which is loaded and executed by the processor to implement the steps performed in the image processing method according to any one of the first aspect.
[0096] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores at least one computer instruction, which is loaded and executed by a processor to implement the steps performed in the image processing method according to any one of the first aspect.
[0097] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0098] The accompanying drawings, which are incorporated into the specification and constitute part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0099] Figure 1 is a structural schematic diagram of an image processing system provided by the embodiments of the present disclosure;
[0100] Figure 2 is a flow of an image processing method provided by the embodiments of the present disclosure Figure 1 ;
[0101] Figure 3 is a schematic diagram of a cloud desktop system provided by the embodiments of the present disclosure;
[0102] Figure 4 is a flow of an image processing method provided by the embodiments of the present disclosure Figure 2 ;
[0103] Figure 5 is a structural schematic diagram of an image processing device provided by an embodiment of the present disclosure Figure 1
[0104] Figure 6 is a structural schematic diagram of an image processing device provided by an embodiment of the present disclosure Figure 2
[0105] Figure 7 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0106] The exemplary embodiments will be described in detail herein with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0107] Figure 1 is a structural schematic diagram of an image processing system provided by an embodiment of the present disclosure. As shown in Figure 1 , the system includes a sending end device 101 and a receiving end device 102. The sending end device 101 and the receiving end device 102 can be communicatively connected, and the sending end device 101 can transmit an acquired image to the receiving end device 102 through a network. The image processing system can be applied to a cloud desktop system, or can be applied to other image transmission scenarios, which are not limited herein.
[0108] When the image processing system is applied to a cloud desktop system, a source device in a server can be used as an image sending device, and a client device can be used as a receiving device. The source device obtains a current frame image (i.e., a current display screen of the cloud desktop) and encodes the current frame image to generate a first code stream, a byte quantity of the first code stream being a preset byte quantity. A predicted transmission time window of the current frame image is obtained according to the first code stream and a current predicted bandwidth of a network. If it is determined that the client device is currently in an interactive scenario, a first maximum transmission time window is obtained according to an average encoding time, an average decoding time, an average display time of each frame image (i.e., a display screen of each frame of the cloud desktop), and a first preset maximum tolerable delay, the first preset maximum tolerable delay being a preset maximum tolerable delay in the interactive scenario. If the predicted transmission time window is greater than the first maximum transmission time window, a preset step is performed, the first preset step including: 1. discarding the current frame image; 2. obtaining a next frame image (i.e., a display screen of a next frame of the cloud desktop) and encoding the next frame image to generate a second code stream, a byte quantity of the second code stream being a byte quantity obtained by reducing the byte quantity of the first code stream according to a preset algorithm; and 3. taking the next frame image as a new current frame image and taking the second code stream as a new first code stream. Until the predicted transmission time window is less than or equal to the first maximum transmission time window, the first code stream is sent to the client device.
[0109] The image processing system provided by the embodiments of the present disclosure can adjust a byte quantity of a code stream after image encoding according to a current predicted bandwidth of a network when a receiving device is in an interactive scenario, and can send the code stream after image encoding to the receiving device when a predicted transmission time window of the image is less than or equal to a maximum transmission time window, thereby avoiding the problem that a user using a cloud desktop will experience obvious lag or even cannot normally use the cloud desktop when the bandwidth of the network is insufficient and the delay is high in the interactive scenario, and improving user experience.
[0110] The image processing system provided by the embodiments of the present disclosure can adjust a byte quantity of a code stream after image encoding according to a current predicted bandwidth of a network when a receiving device is in an interactive scenario, and can send the code stream after image encoding to the receiving device when a predicted transmission time window of the image is less than or equal to a maximum transmission time window, thereby avoiding the problem that a user using a cloud desktop will experience obvious lag or even cannot normally use the cloud desktop when the bandwidth of the network is insufficient and the delay is high in the interactive scenario, and improving user experience. Figure 2 The image processing system provided by the embodiments of the present disclosure can adjust a byte quantity of a code stream after image encoding according to a current predicted bandwidth of a network when a receiving device is in an interactive scenario, and can send the code stream after image encoding to the receiving device when a predicted transmission time window of the image is less than or equal to a maximum transmission time window, thereby avoiding the problem that a user using a cloud desktop will experience obvious lag or even cannot normally use the cloud desktop when the bandwidth of the network is insufficient and the delay is high in the interactive scenario, and improving user experience. Figure 2 is a flowchart of an image processing method provided by the embodiments of the present disclosure. As shown in Figure 2 , the method includes the following steps.
[0111] S201, a current frame image is obtained and encoded to generate a first code stream, a byte quantity of the first code stream being a preset byte quantity.
[0112] In the embodiment, at least one frame of image is acquired before acquiring the current frame of image; each frame of image in the at least one frame of image is encoded to acquire the encoding time of each frame of image; after the encoding of each frame of image in the at least one frame of image, the at least one frame of image is sent to the receiving end device and the average encoding time of each frame of image is obtained according to the encoding time of each frame of image.
[0113] Further, the average decoding time and the average display time of each frame of image sent by the receiving end device are received, the average decoding time and the average display time being acquired after the receiving end device decodes and displays each frame of image in the at least one frame of image after receiving the at least one frame of image.
[0114] For example, in a cloud desktop system, a plurality of frames of display screen transmitted within a preset time length (for example, 1s) are acquired before acquiring the current display screen, each frame of display screen in the plurality of frames of display screen is encoded to acquire the encoding time of each frame of display screen; after the encoding of each frame of display screen in the plurality of frames of display screen, the plurality of frames of display screen are sent to the client device and the average encoding time of each frame of image is obtained according to the encoding time of each frame of image. For example, when the frame per second (FPS, Frames Per Second) is 30, that is, the FPS cycle is 1s and the frame per second is 30. The encoding time of the first frame of display screen is E1, the encoding time of the second frame of display screen is E2, and the encoding time of the thirtieth frame of display screen is E30. The average encoding time E is (E1+E2+...+E30) / 30.
[0115] After the client device receives the plurality of frames of display screen, each frame of display screen in the plurality of frames of display screen is decoded and displayed, the decoding time and the display time of each frame of display screen in the plurality of frames of display screen are acquired, and the average decoding time and the average display time of each frame of display screen in the plurality of frames of display screen are obtained according to the decoding time and the display time of each frame of display screen in the plurality of frames of display screen, respectively.
[0116] For example, after the client device receives 30 frames of display pictures transmitted in a FPS period of the source device, the client device decodes each of the 30 frames of display pictures, obtains a decoding time of the 1st frame of display picture as D1, a decoding time of the 2nd frame of display picture as D2,..., and a decoding time of the 30th frame of display picture as D30, and obtains an average decoding time D as (D1+D2+...+D30) / 30. After the client device decodes each of the 30 frames of display pictures, the client device displays the 30 frames of display pictures, obtains a display time of the 1st frame of display picture as S1, a display time of the 2nd frame of display picture as S2,..., and a display time of the 30th frame of display picture as S30, and obtains an average display time S as (S1+S2+...+S30) / 30. After the client device obtains the average decoding time and the average display time, the client device sends the average decoding time and the average display time to the source device.
[0117] For example, after obtaining the average encoding time and receiving the average decoding time and the average display time sent by the sink device, the current frame of image is encoded to obtain a first code stream, and a byte quantity of the first code stream is a preset byte quantity. For example, in this embodiment, after the current frame of image is encoded, a byte quantity of the first code stream is 1000 bytes.
[0118] S202, obtaining a predicted transmission time window of the current frame of image according to the first code stream and the current predicted bandwidth.
[0119] In this step, the predicted transmission time window of the current frame of image can be calculated by formula (1).
[0120] Wp=(P / B) 1000 (1).
[0121] Wherein, Wp is the predicted transmission time window, P is the byte quantity of the first code stream, and B is the current predicted bandwidth (unit: Bps). In this embodiment, the current predicted bandwidth B of the network is obtained while the current frame of image is obtained, and any method for predicting bandwidth in the prior art can be used to obtain the current predicted bandwidth of the network, which is not specifically limited herein.
[0122] S203, if it is determined that the sink device is currently in an interactive scenario, obtaining a maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame of image, and a first preset maximum tolerant delay, the first preset maximum tolerant delay being a preset maximum tolerant delay in the interactive scenario.
[0123] In one embodiment, if it is determined that the receiving end device is currently in the interactive scenario, the average encoding time of each frame of image is obtained, and the average decoding time and the average display time of each frame of image received from the receiving end device are received, and then the first maximum transmission time window is calculated according to formula (2).
[0124] W1=T1–E–D–S (2).
[0125] Wherein, W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time. After the average encoding time E of each frame of image is obtained and the average decoding time D and the average display time S of each frame of image received from the receiving end device are received, the first maximum transmission time window W1 can be calculated according to formula (2). T1 represents the maximum delay that can be tolerated when each frame of image is transmitted in the interactive scenario, that is, the time spent from the start of collection of each frame of image, through encoding, transmission, decoding, to the final display completion cannot exceed T1 in the non-interactive scenario. T1 is closely related to network bandwidth, if the network bandwidth is large, T1 can take a large value, if the network bandwidth is small, T1 can take a small value.
[0126] In another embodiment, if it is determined that the receiving end device is currently in the non-interactive scenario, the maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time of each frame of image and the second preset maximum tolerable delay, the second preset maximum tolerable delay is the preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay is greater than the first preset maximum tolerable delay.
[0127] Exemplarily, if it is determined that the receiving end device is currently in the non-interactive scenario, the average encoding time of each frame of image is obtained, and the average decoding time and the average display time of each frame of image received from the receiving end device are received, and then the second maximum transmission time window is calculated according to formula (3).
[0128] W2=T2–E–D–S (3).
[0129] Wherein, W2 is the second maximum transmission time window, T2 is the second preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time. The sender device acquires the average encoding time E of each frame of image, receives the average decoding time D and the average display time S of each frame of image sent by the receiver device, and then calculates the second maximum transmission time window W2 according to formula (3). T2 represents the maximum delay that can be tolerated when each frame of image is transmitted in a non-interactive scene, that is, in an interactive scene, the time spent from the start of each frame of image collection, encoding, transmission, decoding, to the final display completion cannot exceed T2. T2 is closely related to the network bandwidth, if the network bandwidth is large, T2 can take a larger value, if the network bandwidth is small, T2 can take a smaller value. Wherein, T2 is greater than T1. In some embodiments, the value range of T1 is 40ms-80ms, for example, 60ms, and the value range of T2 is 80ms-150ms, for example, 100ms.
[0130] It should be noted here that the interactive scene refers to a scene in which the user performs human-computer interaction on the receiver device. For example, the interactive scene can include a keyboard interaction scene, a mouse interaction scene, a touch interaction scene, a voice interaction scene, a gesture interaction scene, a body analysis interaction scene, or a face analysis interaction scene, etc.
[0131] When the receiver device has a mouse operation, it is defined as a mouse interaction scene; when the receiver device has a keyboard operation, it is defined as a keyboard interaction scene; when the receiver device has a touch operation, it is defined as a touch interaction scene; when the receiver device has voice interaction, it is defined as a voice interaction scene; when the receiver device has gesture interaction, it is defined as a gesture interaction scene; when the receiver device has identity analysis (such as face recognition, fingerprint recognition, etc.), it is defined as an identity analysis interaction scene.
[0132] Scenes other than the interactive scene are defined as non-interactive scenes. The reason for distinguishing between the two scenes is that in the interactive scene, the user's perception of delay is more sensitive than in the non-interactive scene, so T2 is greater than T1.
[0133] S204, if the predicted transmission time window is greater than the first maximum transmission time window, a first preset step is performed.
[0134] In this embodiment, if the receiver device is currently in an interactive scene and the predicted transmission time window is greater than the first maximum transmission time window, a first preset step is performed:
[0135] 1. Discard the current frame of image;
[0136] 2. obtaining a next frame image and encoding the next frame image to generate a second code stream, a byte number of the second code stream being a byte number obtained by reducing the byte number of the first code stream according to a preset first algorithm;
[0137] 3. taking the next frame image as a new current frame image and taking the second code stream as a new first code stream.
[0138] 4. determining whether a predicted transmission time window obtained according to the new first code stream and a current predicted bandwidth is greater than the first maximum transmission time window.
[0139] For example, when the byte number P of the first code stream of the current frame image is 1000 bytes, a predicted transmission time window Wp obtained according to the byte number of the first code stream and a current predicted bandwidth is 20 ms, and the first maximum transmission time window W1 is 10 ms, the first preset step is executed: 1. discarding the current frame image; 2. obtaining a next frame image and encoding the next frame image to generate a second code stream, a byte number of the second code stream being 900 bytes, which is a byte number obtained by reducing the byte number 1000 of the first code stream by 10%; 3. taking the next frame image as a new current frame image and taking the second code stream as a new first code stream. 4. determining whether a predicted transmission time window Wp obtained according to the byte number 900 of the new first code stream and a current predicted bandwidth is greater than the first maximum transmission time window. If yes, the first preset step is executed in a loop until the predicted transmission time window is less than or equal to the first maximum transmission time window, and the first code stream is sent to a receiving end device so as to generate the current frame image by decoding the first code stream by the receiving end device.
[0140] In another implementation, if the receiving end device is currently in an interactive scene and the predicted transmission time window is less than the first maximum transmission time window, a second preset step is executed, which includes:
[0141] 1. sending the first code stream to the receiving end device;
[0142] 2. obtaining a next frame image and encoding the next frame image to generate a third code stream, a byte number of the third code stream being a byte number obtained by increasing the byte number of the first code stream according to a preset second algorithm;
[0143] 3. taking the next frame image as a new current frame image and taking the third code stream as a new first code stream.
[0144] 4. determining whether a predicted transmission time window obtained according to the new first code stream and a current predicted bandwidth is less than the first maximum transmission time window.
[0145] For example, when the first code stream byte number P of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the first code stream byte number and the current predicted bandwidth is 20 ms, and the first maximum transmission time window W1 is 30 ms, the second preset step is executed: 1, the first code stream is sent to the receiving end device, so that the receiving end device decodes the first code stream to generate the current frame image; 2, the next frame image is obtained and encoded to generate the third code stream, the byte number of the third code stream is 1100 bytes, which is the byte number after increasing 10% of the byte number 1000 of the first code stream; 3, the next frame image is taken as a new current frame image and the third code stream is taken as a new first code stream. 4, the predicted transmission time window Wp obtained according to the byte number 1100 of the new first code stream and the current predicted bandwidth is determined whether it is less than the first maximum transmission time window. If yes, the second preset step is executed in a loop until the predicted transmission time is equal to the maximum transmission time window.
[0146] For example, if the receiving end device is currently in a non-interactive scene and the predicted transmission time window is greater than the second maximum transmission time window, a third preset step is executed.
[0147] In the embodiment, the third preset step includes:
[0148] 1, discarding the current frame image;
[0149] 2, obtaining the next frame image and encoding the next frame image to generate the second code stream, the byte number of the second code stream is the byte number after reducing the byte number of the first code stream according to a preset first algorithm;
[0150] 3, taking the next frame image as a new current frame image and taking the second code stream as a new first code stream.
[0151] 4, determining whether the first predicted transmission time window obtained according to the new first code stream and the current predicted bandwidth is greater than the second maximum transmission time window.
[0152] For example, when the first code stream byte number P of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the first code stream byte number and the current predicted bandwidth is 20 ms, and the second maximum transmission time window W2 is 12 ms, a third preset step is performed: 1, discarding the current frame image; 2, obtaining a next frame image and encoding the next frame image to generate a second code stream, the byte number of the second code stream being 900 bytes, which is the byte number after reducing the first code stream byte number 1000 by 10%; 3, taking the next frame image as a new current frame image and taking the second code stream as a new first code stream; 4, determining whether the predicted transmission time window Wp obtained according to the byte number 900 of the new first code stream and the current predicted bandwidth is greater than the second maximum transmission time window. If yes, the third preset step is repeatedly performed until the predicted transmission time window is less than or equal to the second maximum transmission time window, and the first code stream is sent to the receiving end device, so that the receiving end device decodes the first code stream to generate the current frame image.
[0153] In another implementation, if the receiving end device is currently in a non-interactive scene and the predicted transmission time window is less than the second maximum transmission time window, a fourth preset step is performed, which includes:
[0154] 1, sending the first code stream to the receiving end device;
[0155] 2, obtaining a next frame image and encoding the next frame image to generate a third code stream, the byte number of the third code stream being the byte number after increasing the first code stream byte number according to a preset second algorithm;
[0156] 3, taking the next frame image as a new current frame image and taking the third code stream as a new first code stream;
[0157] 4, determining whether the predicted transmission time window obtained according to the new first code stream and the current predicted bandwidth is less than the second maximum transmission time window.
[0158] For example, when the first code stream byte quantity P of the current frame image is 1000 bytes, the predicted transmission time window Wp obtained according to the first code stream byte quantity and the current predicted bandwidth is 20 ms, and the second maximum transmission time window W2 is 32 ms, the fourth preset step is executed: 1, the first code stream is sent to the receiving end device, so that the receiving end device decodes the first code stream to generate the current frame image; 2, the next frame image is obtained and encoded, and the third code stream is generated, the byte quantity of the third code stream is 1100 bytes, which is the byte quantity after increasing 10% of the byte quantity 1000 of the first code stream; 3, the next frame image is taken as a new current frame image and the third code stream is taken as a new first code stream. 4, the predicted transmission time window Wp obtained according to the byte quantity 1100 of the new first code stream and the current predicted bandwidth is determined whether it is less than the second maximum transmission time window. If yes, the fourth preset step is executed in a loop until the predicted transmission time is equal to the second maximum transmission time window.
[0159] The image processing method provided by the embodiments of the present disclosure can adjust the byte quantity of the code stream after image encoding according to the current predicted bandwidth of the network in the interactive scenario, and after adjusting the byte quantity of the code stream after image encoding, the predicted transmission time window of the image is obtained according to the byte quantity of the adjusted code stream. When the predicted transmission time window of the image is less than or equal to the maximum transmission time window, the code stream after image encoding is sent to the receiving end device. The transmission delay of the image can be reduced by adjusting the byte quantity of the code stream after image encoding. The phenomenon of lag caused by network congestion can be prevented in advance in the interactive scenario. The problem that users using cloud desktops will feel obvious lag or even cannot normally use the cloud desktops if the bandwidth of the network is insufficient and the delay is high in the interactive scenario is avoided, and the user experience is improved.
[0160] The image processing method provided by the embodiments of the present disclosure will be further described in detail below.
[0161] The application scenarios of the cloud desktop are divided into two scenarios: an interactive scenario and a non-interactive scenario.
[0162] When the client has peripherals such as a mouse, a keyboard, and a touch control device to operate and cause the cloud desktop to change, the scenario is defined as the interactive scenario. Conversely, the scenario is defined as the non-interactive scenario. The corresponding maximum tolerable delays T1 and T2 (unit: ms) are defined for the two different scenarios.
[0163] Maximum tolerable delay definition: the time that cannot be exceeded from the start of picture acquisition at the source end, after encoding, transmission, and decoding, to the final display completion.
[0164] For example, the maximum tolerable delay T1 is defined in the interactive scenario, and the maximum tolerable delay T2 is defined in the non-interactive scenario.
[0165] Wherein, T2 is greater than T1. In some embodiments, the value of T1 ranges between 40ms-80ms, for example, 60ms, and the value of the maximum tolerable delay T2 in non-interactive scenarios ranges between 80ms-150ms, for example, 100ms.
[0166] Meanwhile, the network bandwidth is estimated by the bandwidth prediction module to obtain a predicted bandwidth B (unit: Bps).
[0167] The average encoding time E (unit: ms) of all frames per second is counted by the counting module.
[0168] The average encoding time refers to the total time consumed by the encoding of N frames of images transmitted within a period T divided by N.
[0169] For example, in a cloud desktop, the number of frames transmitted per FPS period (usually 1s) is 30, the encoding time of the first frame is E1, the encoding time of the second frame is E2, and the encoding time of the 30th frame is E30. The average encoding time E is (E1+E2+...+E30) / 30.
[0170] According to the above method, the average decoding time D (unit: ms) of all frames per second is counted.
[0171] The average decoding time refers to the total time consumed by the decoding of N frames of images transmitted within a period T (for example, each FPS period described above) divided by N.
[0172] According to the above method, the average display time S (unit: ms) of all frames within a period T is counted.
[0173] The average display time refers to the total time consumed by the display of N frames of images transmitted within a period T (for example, each FPS period described above) divided by N.
[0174] For each frame, the size of the code stream output by the encoder is defined as the output code stream P (unit: bytes).
[0175] (P / B) 1000 is the predicted transmission time window Wp, and Tn-E-D-S is the maximum transmission time window Wr.
[0176] When (P / B) 1000>(Tn-E-D-S), (Tn takes the value of T1 or T2), the current frame is actively discarded at the encoding end, the output code stream of the encoder is dynamically reduced (i.e., the definition is reduced), a new frame is regenerated, and (P / B) 1000<=(Tn-E-D-S) is met, and transmission is performed again.
[0177] When (P / B) 1000<(Tn-E-D-S), then from the next frame, the dynamic promotion encoder output stream P (i.e., improve the clarity), and meet (P / B) 1000<=(Tn-E-D-S), and then transmission.
[0178] Dynamic reduction or promotion refers to adjusting the output stream parameters of the encoder by a certain percentage, such as 10% based on the size of the output stream currently in use, and resetting the encoder.
[0179] When Wp is less than Wr, the picture can be smoothly transmitted, and under the premise that Wp is less than Wr, the encoder gradually increases the output stream to improve the picture clarity. When Wp is greater than Wr, by adjusting the encoder output stream, the phenomenon of packet loss caused by network congestion can be prevented in advance.
[0180] The present scheme distinguishes between interactive and non-interactive scenarios and defines the maximum allowable tolerance delay in the two scenarios. According to the predicted bandwidth, it dynamically calculates whether the encoded frame stream can reach the client within the maximum tolerance delay. If not, it actively drops frames and adjusts the encoder stream to meet the delay requirement. Conversely, under the requirement of meeting the delay, the encoding stream is dynamically adjusted to improve the picture quality and achieve better user experience.
[0181] The invention point is to set different maximum tolerance delays for different scenarios. The examples given are interactive and non-interactive scenarios, and it can be further refined to multiple scenarios.
[0182] The use scenario of the present application can be, for example, a cloud desktop system as shown in Figure 3 The entire cloud desktop system consists of two parts: the source end and the client end. The source end generally refers to Figure 3 the acquisition module and the encoder. The client generally refers to a terminal device or a set of software systems composed of a decoder and a display module.
[0183] The present scheme divides the source picture of the cloud desktop into two scenarios: interactive scenario and non-interactive scenario.
[0184] Interactive scenario refers to a scenario where the user performs human-computer interaction on the client end.
[0185] For example, the interactive scenario can include keyboard interaction scenario, mouse interaction scenario, touch interaction scenario, voice interaction scenario, gesture interaction scenario, body analysis interaction scenario, or face analysis interaction scenario, etc.
[0186] When the client has mouse operation, it is defined as mouse interaction scene; when the client has keyboard operation, it is defined as keyboard interaction scene; when the client has touch operation, it is defined as touch interaction scene; when the client has voice interaction, it is defined as voice interaction scene; when the client has gesture interaction, it is defined as gesture interaction scene; when the client has identity analysis (such as face recognition, fingerprint recognition, etc.), it is defined as identity analysis interaction scene.
[0187] Scenes other than the interaction scene are defined as non-interactive scenes. The reason for distinguishing between the two scenes is that in the interaction scene, the user's perception of delay is more sensitive than in the non-interactive scene. When the user double-clicks to open a file, the user expects the action to be performed immediately. In response to the cloud desktop system, the source picture generated by the double-click is encoded and transmitted to the client at the fastest speed, and then decoded and displayed. In the non-interactive scene, such as automatic loop playing of slides, even if the transmission delay is twice as large as in the interactive scene, the user will not have much perception and will not affect the user experience. In the case of a fixed bandwidth, the same source picture (before encoding) naturally has a smaller code stream after encoding, and the transmission is faster, so the user can see the changed picture faster. But for a selected fixed encoding algorithm, the smaller the code stream is, the lower the clarity is. Therefore, a balance between clarity and delay is needed, that is, within the acceptable delay of the user, the picture with the highest clarity is transmitted as much as possible. Therefore, the source picture is divided into interactive and non-interactive scenes in the present scheme, and different maximum tolerable delays T1 and T2 are defined. Based on the different maximum tolerable delays, the encoder output code stream is dynamically adjusted to obtain a picture with appropriate clarity, so that the user experience is better.
[0188] The main workflow of the present scheme is shown in Figure 4
[0189] Step one, the encoder, decoder and display module periodically count the average encoding time E, average decoding time D and average display time S, respectively.
[0190] This periodicity is generally one second.
[0191] Step two, the bandwidth prediction module at the source end predicts the bandwidth in real time and outputs the predicted bandwidth value B.
[0192] Step three, at the source end, each frame after acquisition is encoded by the encoder to obtain the code stream size P.
[0193] Step four, at the source end, the formula (P / B) 1000>(Tn–E–D–S) is used for judgment, where Tn is selected as T1 or T2 according to the scene. If the calculation result of the formula is, step five or step six is executed, respectively.
[0194] Step five, the formula is not established, indicating that the current code stream size meets the transmission delay under the current scene, directly to the transmission module for sending, while informing the encoder to increase the output code stream, that is, to improve the clarity.
[0195] This promotion amplitude can be determined according to different encoding algorithms, such as 10% each time;
[0196] Step six, the formula is established, indicating that the current code stream size does not meet the transmission delay under the current scene, then discard the current code stream, inform the encoder to reduce the output code stream, that is, to reduce the clarity, and then reconvert to step three, and cycle again.
[0197] The use scene is explained (T2>T1):
[0198] The client is in a non-interactive scene, and Tn takes the value T2
[0199] At this time, the condition (P / B) 1000≤(T2–E–D–S) indicates that the current code stream size meets the transmission delay under the current scene.
[0200] If the state of the client changes to an interactive scene, Tn takes the value T1
[0201] After Tn takes the value changes, the condition (P / B) 1000≤(T1–E–D–S) indicates that the current code stream size does not meet the transmission delay under the current scene, at this time, P should be appropriately reduced until the condition (P / B) 1000≤(T1–E–D–S) is met.
[0202] The principle is consistent in the process of changing from an interactive state to a non-interactive state.
[0203] Based on the above Figure 2 and Figure 4 The image processing method described in the corresponding embodiment, the following is a device embodiment of the present disclosure, which can be used to execute the method embodiment of the present disclosure.
[0204] Figure 5 is a structural schematic diagram of an image processing device provided by an embodiment of the present disclosure. As Figure 5 shown, the device 50 includes:
[0205] A current frame image acquisition module 501 is configured to acquire a current frame image and encode the current frame image to generate a first code stream, and the number of bytes of the first code stream is a preset number of bytes.
[0206] A predicted transmission time window generation module 502 is configured to obtain a predicted transmission time window of the current frame image according to the first code stream and a current predicted bandwidth.
[0207] The first maximum transmission time window generation module 503 is used to obtain a first maximum transmission time window based on the average encoding time, average decoding time, average display time and first preset maximum tolerance delay of each frame of image if it is determined that the receiving device is currently in an interactive scenario. The first preset maximum tolerance delay is the preset maximum tolerance delay under the interactive scenario.
[0208] The first preset step execution module 504 is used to execute a first preset step if the predicted transmission time window is greater than the first maximum transmission time window. The first preset step includes:
[0209] Discard the current frame image;
[0210] The next frame image is acquired and encoded to generate a second bitstream. The number of bytes in the second bitstream is the number of bytes after the number of bytes in the first bitstream is reduced according to a preset first algorithm.
[0211] The next frame image is used as the new current frame image, and the second bitstream is used as the new first bitstream;
[0212] A new predicted transmission time window is obtained based on the new first bitstream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is greater than the first maximum transmission time window.
[0213] In one embodiment, such as Figure 6 As shown, the device 50 further includes:
[0214] The second preset step execution module 505 is used to execute a second preset step if the predicted transmission time window is less than the first maximum transmission time window. The second preset step includes:
[0215] Send the first bitstream to the receiving device;
[0216] The next frame image is acquired and encoded to generate a third bitstream. The number of bytes in the third bitstream is the number of bytes after the number of bytes in the first bitstream is increased according to a preset second algorithm.
[0217] The next frame image is used as the new current frame image, and the third bitstream is used as the new first bitstream;
[0218] A new predicted transmission time window is obtained based on the new first bitstream and the current predicted bandwidth, and it is determined whether the new predicted transmission time window is less than the first maximum transmission time window.
[0219] In one embodiment, the device 50 further includes:
[0220] The second maximum transmission time window generating module 506 is configured to, if it is determined that the receiving end device is currently in a non-interactive scenario, obtain a second maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame of image and a second preset maximum tolerable delay, the second preset maximum tolerable delay being a preset maximum tolerable delay in the non-interactive scenario, and the second preset maximum tolerable delay being greater than the first preset maximum tolerable delay;
[0221] The third preset step executing module 507 is configured to, if the predicted transmission time window is greater than the second maximum transmission time window, execute a third preset step, the third preset step including:
[0222] discarding the current frame of image;
[0223] obtaining a next frame of image and encoding the next frame of image to generate a second code stream, the number of bytes of the second code stream being a number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm;
[0224] taking the next frame of image as a new current frame of image and taking the second code stream as a new first code stream;
[0225] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is greater than the second maximum transmission time window.
[0226] In an embodiment, the apparatus 50 further includes:
[0227] The fourth preset step executing module 508 is configured to, if the predicted transmission time window is less than the second maximum transmission time window, execute a fourth preset step, the fourth preset step including:
[0228] sending the first code stream to a receiving end device;
[0229] obtaining a next frame of image and encoding the next frame of image to generate a third code stream, the number of bytes of the third code stream being a number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm;
[0230] taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream;
[0231] obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the second maximum transmission time window.
[0232] In an embodiment, the apparatus 50 includes:
[0233] The average encoding time obtaining module 509 is configured to:
[0234] obtain at least one frame of image;
[0235] encode each frame of image in the at least one frame of image and obtain an encoding time of each frame of image;
[0236] after encoding each frame of image in the at least one frame of image, send the at least one frame of image to a receiving end device and obtain an average encoding time of each frame of image according to the encoding time of each frame of image.
[0237] In an embodiment, the apparatus 50 comprises:
[0238] The average decoding time receiving module 510 is configured to receive an average decoding time and an average display time of each frame of image sent by the receiving end device, wherein the average decoding time and the average display time are obtained by the receiving end device after decoding and displaying each frame of image in the at least one frame of image.
[0239] In an embodiment, the predicted transmission time window generating module 502 is configured to:
[0240] Wp=(P / B) 1000, wherein Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
[0241] In an embodiment, the first maximum transmission time window generating module 503 is configured to:
[0242] W1=T1–E–D–S, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerant delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0243] In an embodiment, the second maximum transmission time window generating module 506 is configured to:
[0244] W2=T2–E–D–S, wherein W2 is the second maximum transmission time window, T2 is the second preset maximum tolerant delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
[0245] The image processing apparatus provided by the embodiments of the present disclosure can refer to the implementation process and technical effects of the above Figure 2 and Figure 4 embodiments, which will not be described here again.
[0246] Figure 7is a structural schematic diagram of an electronic device. As shown in the figure, the electronic device 70 includes a processor and a memory, the memory stores at least one computer instruction, the instruction is loaded and executed by the processor to realize Figure 7 and Figure 2 and Figure 4 the steps performed in the image processing method described in the corresponding embodiments.
[0247] Based on the above Figure 2 and Figure 4 the image processing method described in the corresponding embodiments, the embodiments of the present disclosure also provide a computer readable storage medium, for example, a non-transitory computer readable storage medium can be a read only memory (English: Read Only Memory, ROM), a random access memory (English: Random Access Memory, RAM), a CD-ROM, a magnetic tape, a floppy disk and an optical data storage device, etc. The storage medium stores computer instructions for executing the above Figure 2 and Figure 4 the image processing method described in the corresponding embodiments, which will not be described here.
[0248] Those of ordinary skill in the art can understand that all or part of the steps of the above embodiments can be completed by hardware, or by program instructions instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium, and the above mentioned storage medium can be a read only memory, a magnetic disk or an optical disk, etc.
[0249] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the present disclosure disclosed herein. The present application is intended to cover any variations, uses, or adaptive changes to the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional techniques in the art that are not disclosed by the present disclosure. The specification and examples are only considered as exemplary, and the true scope and spirit of the present disclosure are indicated by the following claims.
Claims
1. An image processing method, characterized by, The method comprises the following steps: acquiring a current frame image and encoding the current frame image to generate a first code stream, the number of bytes of the first code stream being a preset number of bytes; obtaining a predicted transmission time window of the current frame image according to the first code stream and a current predicted bandwidth; if it is determined that a receiving end device is currently in an interactive scenario, obtaining a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame image and a first preset maximum tolerant delay, the first preset maximum tolerant delay being a preset maximum tolerant delay in the interactive scenario; if the predicted transmission time window is smaller than the first maximum transmission time window, performing a second preset step, the second preset step comprising: sending the first code stream to the receiving end device; acquiring a next frame image and encoding the next frame image to generate a third code stream, the number of bytes of the third code stream being a number of bytes obtained by increasing the number of bytes of the first code stream according to a preset second algorithm; taking the next frame image as a new current frame image and taking the third code stream as a new first code stream; obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is smaller than the first maximum transmission time window; the obtaining of the first maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame image and the first preset maximum tolerant delay comprises: W1=T1-E-D-S, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerant delay, E is the average encoding time, D is the average decoding time and S is the average display time.
2. The method of claim 1, wherein, after the obtaining of the predicted transmission time window of the current frame image according to the first code stream and the current predicted bandwidth, the method further comprises: if it is determined that the receiving end device is currently in a non-interactive scenario, obtaining a second maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame image and a second preset maximum tolerant delay, the second preset maximum tolerant delay being a preset maximum tolerant delay in the non-interactive scenario, the second preset maximum tolerant delay being greater than the first preset maximum tolerant delay; if the predicted transmission time window is greater than the second maximum transmission time window, performing a third preset step, the third preset step comprising: discarding the current frame image; acquiring a next frame image and encoding the next frame image to generate a second code stream, the number of bytes of the second code stream being a number of bytes obtained by reducing the number of bytes of the first code stream according to a preset first algorithm; taking the next frame image as a new current frame image and taking the second code stream as a new first code stream; obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is greater than the second maximum transmission time window; the obtaining of the second maximum transmission time window according to the average encoding time, the average decoding time, the average display time of each frame image and the second preset maximum tolerant delay comprises: W2 = T2 - E - D - S, wherein, W2 is the second maximum transmission time window, T2 is the second preset maximum tolerance delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
3. The method of claim 2, wherein, The method further comprises: if the predicted transmission time window is less than the second maximum transmission time window, performing a fourth preset step, the fourth preset step comprising: sending the first code stream to a receiving end device; obtaining a next frame of image and encoding the next frame of image to generate a third code stream, the number of bytes of the third code stream being the number of bytes of the first code stream increased according to a preset second algorithm; taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream; obtaining a new predicted transmission time window according to the new first code stream and the current predicted bandwidth and determining whether the new predicted transmission time window is less than the second maximum transmission time window.
4. The method of claim 1, wherein, Before the obtaining a current frame of image and encoding the current frame of image, the method further comprises: obtaining at least one frame of image; encoding each frame of image in the at least one frame of image and obtaining an encoding time of each frame of image; after encoding each frame of image in the at least one frame of image, sending the at least one frame of image to a receiving end device and obtaining an average encoding time of each frame of image according to the encoding time of each frame of image.
5. The method of claim 4, wherein, Before the obtaining a current frame of image and encoding the current frame of image, the method further comprises: receiving an average decoding time and an average display time of each frame of image sent by the receiving end device, the average decoding time and the average display time being obtained by the receiving end device after receiving the at least one frame of image, decoding and displaying each frame of image in the at least one frame of image.
6. The method of claim 1, wherein, The obtaining a predicted transmission time window of a current frame of image according to a first code stream and a current predicted bandwidth comprises: Wp = (P / B) 1000 wherein Wp is the predicted transmission time window, P is the number of bytes of the first code stream, and B is the current predicted bandwidth.
7. An image processing apparatus characterized by comprising: comprising: a current frame of image obtaining module, configured to obtain a current frame of image and encode the current frame of image to generate a first code stream, the number of bytes of the first code stream being a preset number of bytes; a predicted transmission time window generating module, configured to obtain a predicted transmission time window of the current frame of image according to the first code stream and a current predicted bandwidth; a first maximum transmission time window generating module, configured to, if it is determined that a receiving end device is currently in an interactive scenario, obtain a first maximum transmission time window according to an average encoding time, an average decoding time, an average display time of each frame of image and a first preset maximum tolerance delay, the first preset maximum tolerance delay being a preset maximum tolerance delay in the interactive scenario; a second preset step performing module, configured to, if the predicted transmission time window is less than the first maximum transmission time window, perform a second preset step, the second preset step comprising: sending the first code stream to a receiving end device; obtaining a next frame of image and encoding the next frame of image to generate a third code stream, the number of bytes of the third code stream being the number of bytes of the first code stream increased according to a preset second algorithm; taking the next frame of image as a new current frame of image and taking the third code stream as a new first code stream; taking the next frame image as a new current frame image and taking the third code stream as a new first code stream; obtaining a new prediction transmission time window according to the new first code stream and the current prediction bandwidth, and determining whether the new prediction transmission time window is less than the first maximum transmission time window; the first maximum transmission time window is obtained according to the average encoding time, the average decoding time, the average display time and the first preset maximum tolerable delay of each frame image, and the first maximum transmission time window comprises: W1 = T1 - E - D - S, wherein W1 is the first maximum transmission time window, T1 is the first preset maximum tolerable delay, E is the average encoding time, D is the average decoding time, and S is the average display time.
Citation Information
Patent Citations
Picture-encoding device and picture-transmission system using the same and quantization controlling method and mean through-put calculating method used for the same
JP1998028269A
Method and device for transmitting data
US20080294789A1