Video encoding method based on cloud mobile phone, and server
By distinguishing the moving and non-moving areas of the image in the cloud mobile phone, adjusting the quantization parameters and performing sharpening processing, the problem of poor video encoding quality of cloud mobile phones is solved, and the display quality of terminal devices is improved.
Patent Information
- Application Number
- PCT/CN2024/142163
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-11
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
During the data transmission process between cloud mobile phones and terminal devices, due to the violent image movement, rich information, complex texture, and poor video encoding quality, the image quality displayed by terminal devices is poor, affecting the user experience.
By determining the moving area and non-motion area of the image in a cloud mobile phone, the quantization parameters of the non-motion area are reduced, the quantization parameters of the moving area are increased, and the image is sharpened to generate a video code stream to improve the image quality.
Without increasing the encoding rate, the image quality displayed by the terminal device is improved and the user experience is improved.
Smart Images

Figure CN2024142163_03072025_PF_FP_ABST
Abstract
Description
A video encoding method and server based on cloud phone
[0001] This application claims priority to the Chinese patent application with application number 202311824041.8 filed with the State Intellectual Property Office of China on December 26, 2023, and priority to the Chinese patent application with the invention name “A video transmission method, system and related equipment”; it also claims priority to the Chinese patent application with application number 202410592926.8 filed with the State Intellectual Property Office of China on May 11, 2024, and priority to the Chinese patent application with the invention name “A video encoding method and server based on cloud phone”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of video processing, and in particular to a video encoding method and server based on a cloud phone. Background Art
[0003] A cloud phone is a cloud service with an operating system running on a physical server and providing virtual phone functionality. During operation, the cloud phone generates multiple continuous images. These images are rendered, synthesized, captured, and encoded to create a video stream, which is then transmitted over the network to a connected terminal device, allowing the terminal device to display these multiple continuous images. Similar to cloud phones, cloud services that require the display of multiple continuous images on a terminal device also include cloud computers, cloud desktops, live game streaming, and video conferencing.
[0004] These multiple consecutive images exhibit characteristics such as intense motion, rich information, and complex textures. Given a constant data transmission bandwidth between the cloud phone and the terminal device, the more intense the motion, the richer the information, and the more complex the texture, the greater the amount of edge information and detail information contained in the image. Consequently, more content is lost after data compression during video encoding, and the quality of the resulting video stream deteriorates. Since the quality of the video stream determines the image quality displayed on the terminal device, a poor-quality video stream results in poorer image quality displayed on the terminal device, impacting the user experience. Summary of the Invention
[0005] The present application discloses a video encoding method and server based on a cloud phone. After determining the moving area and non-moving area in the image included in the video stream to be processed, the quantization parameter corresponding to the non-moving area is reduced to a first quantization parameter, and the quantization parameter corresponding to the moving area is increased to a second quantization parameter. Video encoding is performed according to the non-moving area and the first quantization parameter, and video encoding is performed according to the moving area and the second quantization parameter. This can improve the quality of video encoding, thereby improving the image quality of multiple continuous images displayed by the user device according to the video code stream generated by the video encoding.
[0006] In the first aspect, the present application provides a video encoding method based on a cloud phone. The method is applied to a cloud phone. The cloud phone runs in a server. The server includes a network card, so that the cloud phone can establish a network connection with a terminal device through the network card. The method includes: after the cloud phone generates a video stream to be processed including M consecutive images, it determines the motion area and non-motion area of the i-th image in the M consecutive images, where M is a positive integer greater than or equal to 2, and i is any positive integer between 1 and M-1. Afterwards, the QP value corresponding to the non-motion area of the i-th image is reduced to a first QP value, and the QP value corresponding to the motion area is increased to a second QP value. The cloud phone encodes the non-motion area of the image according to the first QP value and encodes the motion area of the image according to the second QP value, thereby generating a first video code stream and sending it to the terminal device through the network card.
[0007] In summary, since the quantization parameter has a significant impact on video encoding quality, by lowering the quantization parameter corresponding to the non-moving areas of the image to obtain the first quantization parameter, the video encoding quality corresponding to the first video stream can be improved. Therefore, the image quality displayed based on the first video stream sent by the cloud phone improves as the video encoding quality improves. In addition, by increasing the quantization parameter corresponding to the moving areas of the image, more encoding bitrate can be saved for the encoding process of the non-moving areas, thereby achieving improved video encoding quality without increasing the encoding bitrate.
[0008] Exemplarily, the specific process of the cloud phone reducing the QP value corresponding to the non-motion area to the first QP value is as follows: the cloud phone adds the first parameter to each QP value of the multiple QP values corresponding to the non-motion area to obtain multiple first QP values corresponding to the non-motion area, wherein the first parameter is a value less than zero, thereby achieving a reduction in the QP value corresponding to the non-motion area.
[0009] In this application, the first parameter is generally an integer, which can usually be -1, -2 or -3, and is not specifically limited here.
[0010] Exemplarily, the specific process of the cloud phone increasing the QP value corresponding to the motion area to the second QP value is as follows: the cloud phone adds the second parameter to each QP value of the multiple QP values corresponding to the motion area to obtain multiple second QP values corresponding to the motion area, wherein the second parameter is a value greater than zero, thereby achieving an increase in the QP value corresponding to the motion area.
[0011] In this application, the second parameter is generally a positive integer, which can usually be 1, 2 or 3, and is not specifically limited here.
[0012] Exemplarily, after determining the motion area and non-motion area of the i-th image in M consecutive images, the method also includes: the cloud phone performs a sharpening operation on the motion area of the i-th image in the M consecutive images to obtain the i-th sharpened image, wherein the i-th sharpened image includes a sharpened area and an unsharp area, the sharpened area corresponds to the motion area of the i-th image, and the unsharp area corresponds to the non-motion area of the i-th image, wherein the edge information of the sharpened area is stronger than the edge information of the motion area.
[0013] The sharpening operation on the moving area in the image makes the multiple pixel values corresponding to the sharpened area in the sharpened image different from the multiple pixel values corresponding to the moving area in the target image, which can enhance the edge information, detail information, etc. of the image corresponding to the sharpened area. In the subsequent process of video encoding based on the sharpened image, the encoding efficiency and video encoding quality can be improved, so that the terminal device can display a clearer and sharper image according to the generated first video code stream, thereby improving the image quality of the client display image perceived by the human eye.
[0014] Exemplarily, after performing a sharpening operation on the i-th image, the cloud phone performs video encoding on the sharpened area in the i-th sharpened image according to the second QP value.
[0015] By using the first quantization parameter to perform video encoding on the non-sharp areas of multiple images, the video encoding quality corresponding to the non-moving areas is improved, and the encoding bit rate is also improved. By using the second quantization parameter to perform video encoding on the sharp areas of multiple consecutive sharp images, the video encoding quality corresponding to the moving areas is reduced, and the encoding bit rate is also reduced. In summary, since the human eye is sensitive to changes in image quality in non-moving areas but not to changes in image quality in moving areas, although the video encoding quality corresponding to the moving areas is reduced, the overall image quality of the image displayed by the terminal device according to the first video stream perceived by the human eye is improved, thereby improving the user experience, and the encoding bit rate corresponding to the generated first video stream can be maintained at a low level and will not increase significantly due to the improvement in video encoding quality.
[0016] Exemplarily, the server is further provided with a graphics processing unit (GPU), and the cloud phone calls the GPU to implement the video encoding method provided in the first aspect. Implementing the video encoding method provided in the first aspect through the GPU can improve the execution efficiency of the method and save computing resources.
[0017] Exemplarily, the cloud phone is implemented by a virtual machine or container running on a server.
[0018] In a second aspect, the present application provides a server, characterized in that the server includes a cloud phone and a network card, the cloud phone runs in the server, and the cloud phone establishes a network connection with the terminal device through the network card. The cloud phone is used to generate a video stream to be processed including M consecutive images, where M is a positive integer greater than or equal to 2; to sequentially determine the motion area and non-motion area of the i-th image in the M consecutive images, where i is any positive integer between 1 and M-1, and the motion area and the non-motion area have corresponding quantization parameter QP values; to reduce the QP value corresponding to the non-motion area to a first QP value; to increase the QP value corresponding to the motion area to a second QP value; to encode the non-motion area of the i-th image according to the first QP value, and to encode the motion area of the i-th image according to the second QP value to generate a first video stream; to send the first video stream to the network card; and the network card is used to send the first video stream to the terminal device.
[0019] Exemplarily, the cloud phone is further used to reduce the QP value corresponding to the non-motion area according to the first parameter to obtain a first QP value corresponding to the non-motion area, where the first parameter is a value less than zero.
[0020] Exemplarily, the cloud phone is further used to increase the QP value corresponding to the motion area according to the second parameter to obtain a second QP value corresponding to the motion area, where the second parameter is a value greater than zero.
[0021] Exemplarily, the cloud phone is also used to perform a sharpening operation on the motion area of the i-th image to obtain the i-th sharpened image, wherein the i-th sharpened image includes a sharpened area and an unsharp area, wherein the sharpened area corresponds to the motion area of the i-th image, and the unsharp area corresponds to the non-motion area of the i-th image, and the edge information of the sharpened area is stronger than the edge information of the motion area.
[0022] Exemplarily, the cloud phone is specifically configured to encode the sharpened area of the i-th sharpened image according to the second QP value.
[0023] Exemplarily, the server also includes an image processing unit; the GPU is used to implement the cloud phone-based video encoding method provided in the first aspect through the call of the cloud phone.
[0024] Exemplarily, the server also includes at least one virtual machine, or at least one container; the at least one virtual machine and the at least one container are used to provide an operating environment for the cloud phone.
[0025] In a third aspect, the present application provides a video encoding device, which is applied to a server, the server including a network card, and establishing a network connection with a terminal device through the network card. The video encoding device includes an acquisition unit, an encoding unit, and a sending unit. The acquisition unit is used to acquire a video stream to be processed including M consecutive images, where M is a positive integer greater than or equal to 2; the encoding unit is used to sequentially determine the motion area and non-motion area of the i-th image in the M consecutive images, where i is any positive integer between 1 and M-1, and the motion area and the non-motion area both have corresponding quantization parameter QP values; reduce the QP value corresponding to the non-motion area to a first QP value; increase the QP value corresponding to the motion area to a second QP value; encode the non-motion area of the i-th image according to the first QP value, and encode the motion area of the i-th image according to the second QP value to generate a first video stream; and the sending unit is used to send the first video stream to the network card, so that the network card sends the first video stream to the terminal device.
[0026] Exemplarily, the encoding unit is further configured to reduce the QP value corresponding to the non-motion area according to a first parameter to obtain a first QP value corresponding to the non-motion area, where the first parameter is a value less than zero.
[0027] Exemplarily, the encoding unit is further configured to increase the QP value corresponding to the motion area according to a second parameter to obtain a second QP value corresponding to the motion area, where the second parameter is a value greater than zero.
[0028] Exemplarily, the encoding unit is further used to perform a sharpening operation on the motion area of the i-th image to obtain the i-th sharpened image, wherein the i-th sharpened image includes a sharpened area and an unsharp area, wherein the sharpened area corresponds to the motion area of the i-th image, and the unsharp area corresponds to the non-motion area of the i-th image, and the edge information of the sharpened area is stronger than the edge information of the motion area.
[0029] Exemplarily, the encoding unit is specifically configured to encode the sharpened area of the i-th sharpened image according to the second QP value.
[0030] In a fourth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method provided in the first aspect.
[0031] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes a storage device to execute the method provided in the first aspect.
[0032] In a sixth aspect, the present application provides a computer-readable storage medium, which includes computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method provided in the first aspect.
[0033] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments.
[0035] FIG1 is a schematic structural diagram of a video encoding system provided in an embodiment of the present application;
[0036] FIG2 is a flow chart of a video encoding method based on a cloud phone provided in an embodiment of the present application;
[0037] FIG3 is a schematic structural diagram of a video encoding device provided in an embodiment of the present application;
[0038] FIG4 is a schematic structural diagram of another video encoding device provided in an embodiment of the present application;
[0039] FIG5 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0040] FIG6 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0041] FIG7 is a schematic diagram of a structure in which one or more computing devices are connected via a network according to an embodiment of the present application;
[0042] FIG8 is a schematic diagram of another structure in which one or more computing devices are connected via a network, provided by an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0044] Some of the names involved in this application are defined below.
[0045] YUV: A color encoding method that encodes the brightness (Y component) and chrominance (U component and V component) of an image separately. The Y component is used to describe the brightness changes of the image, and the U component and V component are used to describe the color of the pixel. It is often used in the video encoding process of the encoder.
[0046] RGB: A color image encoding method in which the color of each pixel in the image is composed of the values of three channels: red (R), green (G), and blue (B).
[0047] H.264 / AVC and H.265 / HEVC are both video coding standards. H.265 / HEVC expands and improves the coding algorithm in H.264 / AVC and adds a series of coding features, which can achieve a compression rate twice that of H.264 / AVC while maintaining the same video coding quality.
[0048] The quantization parameter map (QP Map) is a parameter in the encoder that allows users to customize the video encoding quality of different image blocks within a frame. It is essentially a matrix, with each element in the matrix representing the QP of a 64x64 or 16x16 image block in the corresponding image. A larger QP value indicates worse video encoding quality for the corresponding image block. For example, for a frame with a resolution of 1280x720, the matrix dimensions of the QPMap corresponding to the image are 80x45. Users can determine the corresponding element in the QPMap for a specific image block in the encoder and set the desired QP value for that element to determine the video encoding quality of the image block.
[0049] Video coding is a technology that encodes each image block in multiple consecutive images and converts them into a data stream (video bitstream). Specifically, due to the high similarity between consecutive images, intra-frame coding and inter-frame coding are used to generate a residual matrix to remove redundant information from the consecutive images. The resulting residual matrix is more easily compressed than the original image data. The image data in the spatial domain is then transformed into coefficients in the frequency domain by performing a discrete cosine transform (DCT) or other form of transformation on the residual matrix. The frequency domain coefficients are then quantized based on a QP. Quantization is a technique for reducing data size by reducing image detail and precision, achieving image compression. Specifically, the quantization accuracy of the transform coefficients for the corresponding image block is determined based on the QP, which affects the quality and compression rate of the encoded video. Finally, encoding and packaging are performed to obtain the final video bitstream.
[0050] During the operation of real-time video cloud services such as cloud phones, each image in a to-be-processed video stream, including multiple consecutive images, undergoes rendering, synthesis, and capture operations to obtain corresponding YUV data. The encoder then performs motion detection based on the YUV data to determine the moving and non-moving areas of each image. Because the human eye is sensitive to image quality degradation in non-moving areas but insensitive to image quality degradation in moving areas, the encoder reduces the quantization parameters corresponding to the non-moving areas, thereby improving the video encoding quality of these areas and preserving more detailed information in these areas. This can improve the perceived image quality of the multiple consecutive images displayed on the terminal device. In this process, preserving more detailed information in the non-moving areas results in an increased encoding bitrate for these areas, resulting in more data being consumed by the video stream corresponding to these areas. Due to limited data transmission bandwidth between cloud-side devices and terminal devices, the data consumed by the video stream corresponding to the moving areas needs to be reduced. The encoder reduces the video encoding quality of these areas by increasing the quantization parameters corresponding to the moving areas. Although the image quality of the moving areas of the real-time image displayed on the terminal device is reduced, this does not affect the overall image quality of the multiple consecutive images displayed on the terminal device as perceived by the human eye. The above method improves the video encoding quality of the encoder, thereby improving the image quality of the image displayed on the terminal device perceived by the human eye, thereby improving the user experience.
[0051] The application scenarios involved in this application are described below. As shown in Figure 1, Figure 1 is a structural diagram of a video coding system 100 provided in an embodiment of the present application. The video coding system 100 includes a terminal device 110, a cloud data center 120, and a network 130. The terminal device 110 and the cloud data center 120 are connected via the network 130.
[0052] In a specific implementation, the terminal device 110 can be various types of user equipment (UE), such as a mobile phone, a tablet computer (pad), etc. It can also include wearable devices, integrated handheld devices, and other electronic devices with data processing and screen display capabilities. This application does not make specific limitations on this.
[0053] In a specific implementation, the cloud data center 120 may include at least one server 121 and a cloud platform 122 .
[0054] The server 121 may be a general physical server, for example, a physical server such as an X86 server or an ARM server, etc., which is not specifically limited in this application. The server may be connected to other servers or a cloud platform through an internal network.
[0055] Illustratively, the server 121 may include one or more virtual instances 123 , an operating system (OS) 124 , and hardware resources 125 .
[0056] The virtual instance 123 can be any of a container and a virtual machine (VM), and is not specifically limited here. The virtual instance can share the operating system and hardware resources of the server to run cloud phones, cloud computers, or similar real-time video cloud services. It can also be other types of cloud services, which are not specifically limited in this application.
[0057] The operating system 124 can be an operating system suitable for a container, a virtual machine, or a physical machine, such as an Android operating system, a Windows operating system, a Linux operating system, etc., and this application does not make any specific restrictions. It should be noted that the operating system can be an official complete operating system, or it can be an operating system in which individual driver modules of an official complete operating system are modified to adapt to the operation mode of the server, and this application does not make any specific restrictions.
[0058] Hardware resources 125 include computing resources, storage resources, and network resources, such as processors, memories, network cards, and the like. In the embodiment of the present application, in addition to a central processing unit (CPU), hardware resources 125 may also include a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and the like.
[0059] In addition, the hardware resources 125 may also include other hardware resources that may be required by the cloud phone, which is not specifically limited in this application.
[0060] The cloud platform 122 may be a general physical server or a virtual machine (VM) implemented based on a general physical server combined with NFV technology, which is not specifically limited in this application.
[0061] The cloud platform 122 is used to provide an access interface (such as an interface or API). Users can remotely access the access interface through a terminal device, register a cloud account and password on the cloud platform, and log in to the cloud platform. After that, the user can further pay to select and purchase a virtual machine with specific specifications (processor, memory, disk) on the cloud platform, and remotely log in using the virtual machine's remote login account and password to install and run the user's application in the virtual machine. It should be understood that the above examples are for illustrative purposes only and are not specifically limited in this application.
[0062] It should be understood that the video encoding system shown in FIG1 is only a possible example provided by an embodiment of the present application. The system may also include more devices, which is not specifically limited in the present application.
[0063] In the aforementioned system, the cloud phone running as a virtual instance needs to send the video generated during operation to the terminal device for display. Specifically, the video, consisting of multiple consecutive images, must be sent to the terminal device for display. Due to data transmission bandwidth limitations between the server and the terminal device, the server encodes the video using a video encoding algorithm representing the multiple consecutive images, achieving data compression and generating a video stream. This video stream is then sent to the terminal device via the network card in the hardware resources. The terminal device then performs operations such as decoding the video stream to restore the original image data and display the multiple consecutive images.
[0064] The data compression issues in the video encoding process described above result in the loss of detailed image information and other content. When an image includes more information, has more complex textures, or experiences more intense motion, the encoding device loses more detailed image information and other content during the video encoding process because the data transmission bandwidth between the server and the terminal device remains unchanged, resulting in poor video encoding quality in the generated video stream. Because the video stream also needs to be transmitted to the terminal device so that the terminal device can display multiple consecutive images, the video encoding quality corresponding to the video stream plays a decisive role in the image quality of the images displayed by the terminal device. In the case of poor video encoding quality, the image quality of the multiple consecutive images displayed by the terminal device is also poor.
[0065] To improve the image quality displayed on a terminal device, a cloud phone can employ methods to increase the encoding compression rate during the video encoding process for multiple consecutive images. For example, the configuration of sub-algorithms within the video encoding algorithm, such as the motion search algorithm, can be adjusted to improve the encoding compression rate. The motion search algorithm is used to determine the optimal motion vector, thereby enabling inter-frame prediction and motion compensation during the video encoding process. Adjusting the sub-algorithm configuration can adjust the motion search algorithm's motion search range, step size, and accuracy, thereby more accurately determining the optimal motion vector. Because the motion vector indicates the motion relationship between the current image frame and the reference image frame, accurately determining the optimal motion vector can reduce inter-frame differences, reduce the encoding of redundant information, improve encoding efficiency, and thus improve the encoding compression rate. Alternatively, the encoding algorithm can be replaced with the H.265 / HEVC encoding algorithm, which has a higher compression rate than the H.264 / AVC encoding algorithm. This method can reduce the encoding of redundant information while encoding more image details while maintaining the same size of the video stream generated by the video encoding algorithm, thereby improving the image quality displayed on the terminal device. However, this method will increase the complexity of the encoding device and consume a large amount of computing resources. In addition, it will also increase the encoding delay, affecting the smoothness of the image display in the terminal device, thereby affecting the user experience.
[0066] In addition to the above methods, the image quality displayed on the terminal device can also be improved by increasing the encoding bit rate. The encoding bit rate refers to the number of bits transmitted per second and determines the amount of data allocated to each frame by the encoding device. Increasing the encoding bit rate provides more data, encoding more image details, and thus improving image quality. However, increasing the encoding bit rate also increases the amount of data transmitted between the server and the terminal device, increasing operating costs.
[0067] In addition, cloud phones can also use a fixed-quality encoding mode to ensure the image quality displayed on the terminal device. Fixed-quality encoding mode is a bitrate control mode that maintains the video encoding quality of the video stream generated by the cloud phone at a fixed level, ensuring stable image quality regardless of factors such as image motion and information complexity. This stable image quality results in significant fluctuations in the encoding bitrate with changes in image motion and information complexity, which can affect stable data transmission, cause network congestion, and impact the user experience.
[0068] In order to solve the problems existing in the above-mentioned methods, the present application provides a new video encoding method based on cloud phones, which can improve the video encoding quality of the generated video stream without increasing the encoding bit rate, thereby improving the picture quality of the video stream obtained according to the video encoding displayed on the terminal device.
[0069] As shown in Figure 2, Figure 2 is a flowchart of a video encoding method based on a cloud phone provided in an embodiment of the present application. The method is applied to the video encoding system shown in Figure 1. The video encoding method provided by the present application includes the following multiple steps.
[0070] S210: The terminal device sends a video request to the cloud phone. Correspondingly, the cloud phone receives the video request sent by the terminal device.
[0071] The terminal device sends a video request to the cloud phone. The cloud phone generates a video stream to be processed including M consecutive images according to the received video request, where M is a positive integer greater than or equal to 2, for example, T1, T2...T M .
[0072] Exemplarily, when the cloud phone receives a cloud phone login request sent by a terminal device and passes the verification, it starts running and generates M consecutive images, where each image is the cloud phone interface corresponding to a certain time point.
[0073] Alternatively, the cloud phone can generate multiple continuous game images when it receives a multi-player cloud game application initiation request sent by the terminal device, or it can generate M continuous images when it receives a video conference initiation request sent by the terminal device, or a similar real-time communication application initiation request.
[0074] The video request sent by the terminal device to the cloud phone can be a cloud phone login request, a multi-person cloud game application initiation request, a video conference request, etc., and this application does not make any specific restrictions on this.
[0075] According to the above content, after receiving the video request sent by the terminal device, the cloud phone generates multiple continuous images. These images need to be sent to the terminal device for display, but cannot be sent directly in the form of images. Therefore, the cloud phone needs to encode the generated multiple continuous images and execute steps S220 to S260 to generate a video code stream that can be sent to the terminal device.
[0076] S220: Determine the motion region and the non-motion region of the i-th image in the M consecutive images.
[0077] The i-th image among M consecutive images includes a plurality of pixels, where i is any positive integer between 1 and M-1. The pixel value of each pixel in the i-th image can be represented using a YUV method, that is, the pixel value of each pixel includes three dimensions: brightness (Y), first chrominance (U), and second chrominance (V). Specifically, the brightness, first chrominance, and second chrominance of each pixel of the i-th image are encoded according to the YUV color encoding method to obtain the pixel values of each pixel of the i-th image. The i-th image can be represented in the form of the following matrix A.
[0078] Among them, P (Y,U,V) represents the pixel value of the pixel in the i-th image, m is the total number of rows of pixels in the i-th image, n is the total number of columns of pixels in the i-th image, and Y 11 Indicates the brightness component of the pixel corresponding to the first row and first column, U 11 Represents the first chromaticity component of the pixel corresponding to the first row and first column, V 11 Represents the second chromaticity component of the pixel point corresponding to the first row and first column. Other similar numerical values can be inferred based on the above content to obtain the corresponding represented components, which will not be repeated in this application.
[0079] The i-th image includes motion regions and non-motion regions. The motion region refers to the region in the i-th image where motion occurs, i.e., the pixel values of the motion region change over time. The non-motion region refers to the region in the i-th image that remains stationary, i.e., the pixel values of these regions remain essentially unchanged over time. The motion region and non-motion region in the i-th image can be determined by:
[0080] The cloud phone determines the motion area and non-motion area of the i-th image based on the pixel value of the i-th image and the motion detection algorithm.
[0081] Specifically, the cloud phone divides the current i-th image into multiple image blocks, each of which is usually 8x8 or 16x16 pixels in size, which is not specifically limited in this application. Afterwards, block matching is performed for each image block in the adjacent image frames of the i-th image to determine the area in the (i+1)-th image that is most similar to each image block, wherein block matching can be achieved by using mean square error or other similarity measurement methods to determine the difference in pixel values between regions, thereby determining the degree of matching between regions and determining the block matching result. After determining the matching image block of each image block, a motion estimation algorithm is used to calculate the motion vector of the pixel block or pixel point in each image block between the two images, wherein the motion vector is used to represent the displacement of the current pixel block or pixel point in the horizontal and vertical directions.
[0082] Then, the motion region and non-motion region in the i-th image can be determined based on the magnitude of the motion vector. When the motion vector is less than the vector threshold, the motion intensity of the corresponding region is small, and the corresponding region can be determined as a non-motion region based on the motion intensity. When the motion vector is greater than or equal to the vector threshold, the motion intensity of the corresponding region is large, and the corresponding region can be determined as a motion region based on the motion intensity. The vector threshold can be determined based on experience or user needs, and this application does not specifically limit this.
[0083] For example, if a pixel block includes multiple pixels and the motion vector corresponding to each pixel is both greater than and less than a vector threshold, a statistical analysis may be performed on the motion vector corresponding to each pixel. For example, the average value, variance, maximum value, and minimum value of the motion vectors corresponding to the multiple pixels may be calculated. When the average value of the motion vectors corresponding to the multiple pixels is calculated, the calculated average value is compared with the vector threshold. If the average value is greater than or equal to the vector threshold, the pixel block is determined to be a motion region. If the average value is less than the vector threshold, the pixel block is determined to be a non-motion region.
[0084] In addition to determining the motion area and non-motion area of the real-time picture based on the pixel value of the i-th image and the motion estimation algorithm, the processor can also determine it based on other types of technical algorithms and data. For example, the difference method is used to determine the motion area and non-motion area by calculating the difference between the pixel blocks between the i-th image and the adjacent image frame. The difference can be the difference between the pixel values, or other similarity measurement. When the difference exceeds the difference threshold, it indicates that the motion intensity of the image block is large, and the corresponding image block is determined to be the motion area; the optical flow estimation algorithm is used to determine the motion area by calculating the pixel displacement between the i-th image and the adjacent image frame. When the pixel displacement is greater than the displacement threshold, it indicates that the motion intensity of the image block is large, and the corresponding image block or pixel point is determined to be the motion area; or a motion model is used to determine the motion area of the image, etc. The above methods can be used alone or in combination with each other, and this application does not make specific limitations on this.
[0085] The above operations can be used to distinguish between the motion area and the non-motion area of the i-th image. Since the human eye is sensitive to the degradation of image quality in the non-motion area, but not to the degradation of image quality in the motion area, in order to improve the image quality of the i-th image displayed on the terminal device as perceived by the human eye, it is necessary to retain the detailed information of the non-motion area to a greater extent, etc., so as to improve the image quality of the non-motion area in the i-th image, while it is not necessary to maintain a high image quality for the motion area, and the effect of improving the overall image quality of the i-th image displayed on the terminal device as perceived by the human eye can also be achieved. Therefore, after determining the motion area and the non-motion area of the i-th image, the processor can simultaneously execute steps S230 to S250, and make different adjustments to the motion area and the quantization parameters corresponding to the motion area and the non-motion area, so as to improve the video encoding quality of the non-motion area and improve the overall image quality of the i-th image displayed on the terminal device as perceived by the human eye.
[0086] S230: Perform a sharpening operation on the motion region of the i-th image to obtain an i-th sharpened image.
[0087] The sharpened image is an image obtained by performing a sharpening operation on the motion region of the i-th image.
[0088] The sharpened image also includes a plurality of pixels. The pixel value of each pixel in the sharpened image can also be represented by YUV. Specifically, the sharpened image can be represented in the form of a matrix B.
[0089] Among them, Z (Y,U,V) represents the pixel value of the pixel in the sharpened image, l is the total number of rows of pixels in the sharpened image, l=m, k is the total number of columns of pixels in the sharpened image, k=n, and other values can be determined according to the components represented by the corresponding data in the matrix A representing the i-th image above, which will not be repeated here.
[0090] The sharpened image includes a sharpened area and an unsharp area, wherein the position of the sharpened area in the sharpened image is the same as the position of the motion area in the i-th image, and the position of the unsharp area in the sharpened image is the same as the position of the non-motion area in the i-th image. Specifically, when the i-th image is represented in the form of matrix A and the sharpened image is represented in the form of matrix B, assuming that the position of the sharpened area in the sharpened image is B[1:a, 1:b], wherein B represents matrix B, 1:a represents the 1st row and the a-th row in matrix B, 1:b represents the 1st column and the b-th column in matrix B, and B[1:a, 1:b] represents the area surrounded by the 1st row, the a-th row, the 1st column, and the b-th column in matrix B, then the position of the motion area in the i-th image is A[1:a, 1:b]. According to the above content, it can be seen that A[1:a, 1:b] represents the area surrounded by the 1st row, the a-th row, the 1st column, and the b-th column in matrix A.
[0091] After determining the sharpened area, the position of the unsharp area in the sharpened image can be determined as B[a:l, 1:b] and B[1:l, b:k]. Then, the position of the non-motion area in the i-th image is A[a:m, 1:b] and B[1:m, b:n]. The meaning of each value is the same as the position of the sharpened area described above, and will not be repeated here.
[0092] Because the pixels corresponding to the moving area are significantly displaced between the i-th image and the adjacent image frame, the motion compensation operation for the i-th image based on the motion vector between the i-th image and the adjacent image frame will be inaccurate during the video encoding process, and the accurate pixel values of the i-th image obtained by matching will also be inaccurate. As a result, the residual between the i-th image and the adjacent image frame cannot be effectively reduced, requiring more bits to represent the residual information, which increases the amount of data after encoding. In addition, it will also affect the encoding efficiency and video encoding quality. Therefore, the embodiments of the present application improve the encoding efficiency and video encoding quality by performing sharpening operations on the moving area and enhancing the edge information of the moving area.
[0093] Methods for sharpening the moving area to obtain a sharpened image may include unsharp masking (USM), Gaussian sharpening, sharpening filters, etc. The following will take the sharpening filter as an example for detailed description, wherein the sharpening filter may include Sobel, Prewitt, Laplacian, etc.
[0094] When the sharpening filter is the Sobel operator, the Sobel operator includes two 3*3 convolution kernels, and the first convolution kernel is:
[0095] The second convolution kernel is:
[0096] Since the Sobel operator is usually applied to single-channel grayscale images, and the motion area of the i-th image is represented in YUV format, the first chrominance component and the second chrominance component of the multiple pixel values corresponding to the motion area can be set to zero, retaining only the luminance component, and converting the motion area into a grayscale image. Afterwards, the above two convolution kernels are applied to each pixel point in the motion area, and the luminance component Y is convolved horizontally according to the first convolution kernel to obtain the horizontal edge intensity value G x , perform a vertical convolution operation on the brightness component Y according to the second convolution kernel to obtain the vertical edge intensity value G y For example, the pixel point in row i and column j in the grayscale image corresponding to the motion area is G x and G y The calculation formula is as follows: G x =Y(i+1,j-1)+2Y(i+1,j)+Y(i+1,j+1)- (Y(i-1,j-1)+2Y(i-1,j)+Y(i-1,j+1)) G y =Y(i-1,j+1)+2Y(i,j+1)+Y(i+1,j+1)- (Y(i-1,j-1)+2Y(i,j-1)+Y(i+1,j-1))
[0097] Afterwards, for each pixel, the total edge strength is calculated based on the edge strength values in the horizontal and vertical directions. The calculation formula for the total edge strength of each pixel in the motion area is as follows:
[0098] The calculation formula for the edge direction of each pixel is as follows:
[0099] According to the above steps, a processed grayscale image can be obtained. Subsequently, the processed grayscale image is combined with the first chroma component and the second chroma component in the multiple pixel values corresponding to the motion region to restore the processed image represented in YUV format. The processed image is the sharpened region in the sharpened image. In addition, after obtaining the processed grayscale image, the pre-processed grayscale image and the processed grayscale image can be superimposed. Thereafter, the first chroma component and the second chroma component in the multiple pixel values corresponding to the motion region are combined to restore the processed image represented in YUV format.
[0100] It can be understood that the above process is a possible example of a cloud phone performing a sharpening operation on a moving area. When performing a sharpening operation according to other operators or methods, the steps may be different, and this application does not make specific limitations on this.
[0101] To sum up, the cloud phone's sharpening operation on the moving area in the i-th image makes the multiple pixel values corresponding to the sharpened area in the sharpened image different from the multiple pixel values corresponding to the moving area in the i-th image. This can enhance the edge information of the image corresponding to the sharpened area, and can also enhance the detail information of the image. Therefore, in the subsequent process of video encoding based on the sharpened image, the encoding efficiency and video encoding quality can be improved, so that the terminal device can display a clearer and sharper image according to the generated video code stream, thereby improving the picture quality of the terminal device display screen perceived by the human eye and improving the user experience.
[0102] S240: Lower the QP value corresponding to the non-motion area of the i-th image to the first QP value.
[0103] The quantization parameter generally indicates the quantization step size and may also include parameters such as a quantization table, a quantization matrix, and a quantization factor, which are not specifically limited in this application. The quantization step size refers to the size of the interval that divides a continuous signal into discrete levels. The quantization step size is a positive integer that is used to divide a continuous input range and map it to discrete output levels. When the quantization step size is smaller, the interval between each quantization level is smaller, the quantization accuracy is higher, the image detail loss during the video encoding process is reduced, and the video encoding quality is higher. When the quantization step size is larger, the interval between each quantization level is larger, the quantization accuracy is lower, the image detail loss during the video encoding process is increased, and the video encoding quality is lower. Since video encoding quality affects the image quality displayed by the terminal device, the smaller the quantization parameter, the higher the video encoding quality, and the better the image quality displayed by the terminal device. In addition to the quantization step size, the elements in the quantization table or quantization matrix can also determine the quantization level, thereby affecting the video encoding quality. This application does not specifically limit this.
[0104] Since the human eye is more sensitive to the degradation of image quality in non-moving areas, in order to improve the image quality corresponding to the non-moving areas in the image displayed by the terminal device as perceived by the human eye, it is necessary to retain the detail information of the non-moving areas in the i-th image to a greater extent. The embodiment of the present application reduces the detail loss of the non-moving areas during the video encoding process by lowering the first quantization parameter corresponding to the non-moving areas, thereby improving the video encoding quality corresponding to the non-moving areas.
[0105] Each QP value in the QP Map corresponding to the i-th image corresponds to a 64*64 or 16x16 pixel image block in the i-th image. It can be determined that the image block includes multiple pixels. Therefore, the image block of the i-th image corresponding to each QP value may only include motion areas, or non-motion areas, or both motion areas and non-motion areas. Therefore, before reducing the quantization parameters corresponding to the non-motion areas, it is necessary to determine which quantization parameters can be used as the quantization parameters corresponding to the non-motion areas. The specific process is as follows:
[0106] When confirming that an image block corresponding to a certain quantization parameter includes only a non-motion area, the cloud phone determines that the quantization parameter is a quantization parameter corresponding to the non-motion area.
[0107] When the cloud phone confirms that an image block corresponding to a certain quantization parameter includes both a motion area and a non-motion area, it can make a determination by comparing the proportions of the non-motion area and the motion area in the image block. When the proportion of the non-motion area in the image block is greater than or equal to the proportion of the motion area in the image block, it is determined that the quantization parameter is the quantization parameter corresponding to the non-motion area. Otherwise, the quantization parameter is not the quantization parameter corresponding to the non-motion area. Alternatively, the cloud phone can also determine that as long as the image block corresponding to a certain quantization parameter includes a non-motion area, the quantization parameter can be used as the quantization parameter corresponding to the non-motion area. This application does not specifically limit the method for determining the quantization parameter corresponding to the non-motion area.
[0108] After obtaining the quantization parameter corresponding to the non-motion area in the i-th image by the above method, since the non-motion area of the i-th image corresponds to one or more QP values, in the case of corresponding multiple QP values, the QP values can be the same or different. The cloud phone adds a first parameter less than zero to each QP value in the multiple QP values corresponding to the non-motion area to obtain multiple first QP values corresponding to the non-motion area, where the first parameter is usually an integer. In the embodiment of the present application, the first parameter is in the range of [-3,0] and can be -1, -2 or -3. The first parameter can also take other possible values, which are not specifically limited in this application.
[0109] Specifically, when the motion intensity of the non-motion area corresponding to the original quantization parameter is very small, the value of the first parameter ΔQP1 can be determined to be -3, and the ΔQP1 with a value of -3 is added to the original quantization parameter to obtain the first QP value; when the motion intensity of the non-motion area corresponding to the original quantization parameter is slightly larger, the value of ΔQP1 can be determined to be -1, and the ΔQP1 with a value of -1 is added to the original quantization parameter to obtain the first QP value.
[0110] In the above process, the motion intensity can be determined based on the motion intensity of a non-motion region in the corresponding image block. In another possible embodiment, the motion intensity can be determined based on the motion intensity of the entire corresponding image block. The corresponding image block may include a motion region. As the proportion of the motion region in the image block gradually increases, the overall motion intensity of the image block increases, and the greater the motion intensity used in the above process of determining ΔQP1. This application does not specifically limit the method for determining the motion intensity used in the process of determining ΔQP1.
[0111] It can be understood that the above method for determining ΔQP1 is a possible example provided in the embodiment of the present application. The cloud phone can also determine ΔQP1 in other ways, which is not specifically limited in the present application.
[0112] Through the above method, the cloud phone reduces the original quantization parameter corresponding to the non-moving area by setting ΔQP1 of the original quantization parameter corresponding to the non-moving area to obtain a first quantization parameter. Since the smaller the quantization parameter, the higher the video encoding quality of the image block corresponding to the quantization parameter, the video encoding quality corresponding to the non-moving area is higher than the corresponding video encoding quality before the quantization parameter is reduced. The improvement in video encoding quality can further improve the image quality of the non-moving area in the real-time picture displayed by the terminal device.
[0113] After step S240, the quantization parameter corresponding to the non-motion area is reduced, which can retain more detailed information in the non-motion area. Therefore, after the video encoding operation, the video stream corresponding to the non-motion area needs to occupy more data to store more detailed information of the non-motion area, and the encoding bit rate corresponding to the non-motion area increases. However, due to the limitation of the data transmission bandwidth between the cloud phone and the terminal device, when the amount of data occupied by the video stream corresponding to the non-motion area increases, the amount of data occupied by the video stream corresponding to the motion area needs to be reduced. Therefore, the cloud phone needs to execute step S250.
[0114] S250: Raise the QP value corresponding to the motion region of the i-th image to a second QP value.
[0115] Similar to step S240, before increasing the QP value corresponding to the motion region, it is necessary to determine which QP values can be used as the QP value corresponding to the motion region. The specific process is as follows:
[0116] When confirming that an image block corresponding to a certain QP value includes only a motion area, the cloud phone determines that the QP value is a QP value corresponding to the motion area.
[0117] When the cloud phone confirms that the image block corresponding to a certain QP value includes both motion areas and non-motion areas, it can determine this by comparing the proportions of non-motion areas and motion areas in the image block. If the proportion of motion areas in the image block is greater than the proportion of non-motion areas in the image block, it is determined that the QP value can be used as the QP value corresponding to the motion area. Otherwise, the QP value is not the QP value corresponding to the motion area. Alternatively, the cloud phone can also determine that as long as the image block corresponding to a certain QP value includes a motion area, it can be used as the QP value corresponding to the motion area. This application does not specifically limit the method for determining the QP value corresponding to the motion area.
[0118] After obtaining multiple QP values corresponding to the motion area in the i-th image through the above method, since the motion area of the i-th image corresponds to one or more QP values, in the case of corresponding multiple QP values, the QP values can be the same or different. The cloud phone adds a second parameter greater than zero to each QP value in the multiple QP values corresponding to the motion area to obtain multiple second QP values corresponding to the motion area, where the second parameter is usually a positive integer. In the embodiment of the present application, the second parameter is in the range of [0, 3] and can be 1, 2 or 3. The second parameter can also take other possible values, which are not specifically limited in this application.
[0119] Specifically, when the motion intensity of the motion area corresponding to the original QP value is small, the value of the second parameter ΔQP2 can be determined to be 1, and the ΔQP2 with a value of 1 is added to the QP value to obtain the second QP value; when the motion intensity of the motion area corresponding to the QP value is large, the value of ΔQP2 can be determined to be 3, and the ΔQP2 with a value of 3 is added to the QP value to obtain the second QP value.
[0120] In the above process, the motion intensity can be determined based on the motion intensity of the motion region in the corresponding image block. In another possible embodiment, the motion intensity can be determined based on the motion intensity of the entire corresponding image block, where the corresponding image block includes a non-motion region. As the proportion of the non-motion region in the image block gradually increases, the overall motion intensity of the image block decreases, and the motion intensity used in the above process of determining ΔQP2 becomes smaller. This application does not specifically limit the method for determining the motion intensity used in the process of determining ΔQP2.
[0121] It can be understood that the above method for determining ΔQP2 is a possible example provided in the embodiment of the present application. The cloud phone can also determine ΔQP2 in other ways, which is not specifically limited in the present application.
[0122] During the above process, the cloud phone lowers the QP value corresponding to the non-moving area while increasing the QP value corresponding to the moving area, obtaining a second QP value corresponding to the moving area. This reduces the retention of detail information in the moving area during video encoding, reduces the video encoding quality of the moving area, and reduces the encoding bitrate corresponding to the moving area. Although the above operation of increasing the QP value of the moving area leads to a decrease in the image quality of the moving area in the image displayed by the terminal device to a certain extent, the human eye is not sensitive to the decrease in image quality of the moving area, and step S240 also improves the image quality of the non-moving area in the image displayed by the terminal device. Ultimately, the overall image quality of the image displayed by the terminal device as perceived by the human eye is improved.
[0123] S260: Encode the non-sharp area of the i-th image according to the first QP value, and encode the sharp area of the i-th image according to the second QP value to generate a first video code stream.
[0124] The specific process of the cloud phone performing video encoding on the unsharp area of the i-th image according to the first QP value is as follows: motion estimation is performed between the unsharp area in the i-th sharp image and the unsharp area in the reference sharp image frame, and the optimal motion vector is determined by matching the image blocks in the unsharp area in the i-th sharp image with the image blocks in the unsharp area in the reference sharp image frame; after determining the optimal motion vector, motion compensation is performed on the unsharp area in the i-th sharp image, that is, the pixel values corresponding to the unsharp area in the reference sharp image frame are moved according to the optimal motion vector to determine the optimal motion vector in the i-th sharp image. The pixel values of the unsharp area are used to reduce the residual between the unsharp area in the i-th sharpened image and the unsharp area in the reference sharpened image frame; a discrete cosine transform is performed on the unsharp area in the i-th sharpened image after motion compensation, and the corresponding pixel values are converted into frequency domain coefficients, which are used to indicate the changes in different frequencies in the image; thereafter, these coefficients are quantized according to a first QP value. When the first QP value is used as the quantization step size, each frequency domain coefficient can be divided by the first QP value to change the frequency domain coefficient to a smaller value; thereafter, an entropy coding technique (for example, Huffman coding or arithmetic coding) is used to encode the quantized coefficients.
[0125] The specific process of the cloud phone performing video encoding on the sharp area of the i-th image according to the second QP value is similar to the specific process of performing video encoding on the non-sharp area of the i-th image according to the first QP value, and will not be repeated here.
[0126] The cloud phone combines the compressed data obtained by performing video encoding on the non-sharp area according to the first QP value and the compressed data obtained by performing video encoding on the sharp area according to the second QP value to generate a first video code stream.
[0127] It can be understood that the specific process of the above-mentioned cloud phone performing video encoding based on quantization parameters and images can be a possible implementation method provided by the embodiment of the present application, and is not specifically limited here.
[0128] S270: Send the first video code stream to the terminal device. Correspondingly, the terminal device receives the first video code stream sent by the cloud phone.
[0129] The cloud phone sends the generated first video code stream to the terminal device. After receiving the first video code stream, the terminal device decodes the first video code stream through a decoder and parses the compressed video code stream data into YUV data representing the image. The YUV data is used for color adjustment, scaling, rotation and other processing. Afterwards, the terminal device renders the processed YUV data representing the image through a graphics library or graphics API, and uses a synthesizer or synthesis algorithm to synthesize the YUV data representing the image with the synthesized data. Finally, the terminal device displays the rendered and synthesized image on the terminal device's display.
[0130] To summarize, the video encoding method provided in the embodiment of the present application determines the motion area and non-motion area of the i-th image, performs a sharpening operation on the motion area to obtain the i-th sharpened image, enhances the edge information of the motion area, then reduces the quantization parameter corresponding to the non-motion area to obtain a first QP value, increases the quantization parameter corresponding to the motion area to obtain a second QP value, performs video encoding according to the sharpened area in the sharpened image and the second QP value, and performs video encoding according to the non-sharp area in the sharpened image and the first QP value to generate a video code stream.
[0131] The above method can reduce the loss of detail information in non-moving areas during the video encoding process, thereby improving the video encoding quality of non-moving areas. Although the encoding bit rate corresponding to the non-moving areas increases, the detail retention in the moving areas decreases, and the encoding bit rate corresponding to the moving areas is reduced, thereby maintaining a low encoding bit rate. Because the human eye is not sensitive to image quality degradation in moving areas but is more sensitive to changes in image quality in non-moving areas, even if the above method reduces the video encoding quality corresponding to the moving areas, the video encoding quality corresponding to the non-moving areas increases, thereby improving the image quality of the image displayed by the terminal device.
[0132] In a specific implementation, a terminal device acts as a video conference initiating device, and the cloud phone creates a conference group by running a video conferencing application. Another one or more terminal devices act as video conference participating devices and send a join request to the cloud phone to join the conference group using the generated conference group number or QR code. The join request can be used as a video request sent by the terminal device to the cloud phone to obtain the conference display video uploaded by the video conference initiating device to the cloud phone. The conference display video to be processed uploaded by the video conference initiating device to the cloud phone includes multiple continuous conference display images. The cloud phone determines the motion area and the non-motion area in the i-th conference display image according to steps S220 to S260 shown in Figure 2 and the conference display video to be processed, wherein the i-th conference display image is any one of the multiple continuous conference display images; the motion area is sharpened to obtain the i-th conference display sharpened image, wherein the sharpened area in the i-th conference display sharpened image corresponds to the motion area of the i-th conference display image, and the unsharp area in the i-th conference display sharpened image corresponds to the non-motion area of the i-th conference display image. While performing the sharpening operation, the first quantization parameter corresponding to the non-motion area is also reduced to obtain a first QP value, and the first quantization parameter corresponding to the motion area is reduced to obtain a second QP value; then, the unsharp area is video encoded according to the first QP value, and the sharpened area is video encoded according to the second QP value to generate a video code stream for the conference display.
[0133] The cloud phone sends the generated video code stream to one or more terminal devices that are participating devices in the video conference, so that the terminal devices display multiple continuous conference display images through decoding, rendering and synthesis operations.
[0134] It should be understood that the video encoding method provided in this application can be applied to not only the above-mentioned application scenarios, but also other possible scenarios such as live broadcast, and this application does not make any specific limitations on this.
[0135] As shown in Figure 3, Figure 3 is a structural schematic diagram of a video encoding device provided by an embodiment of the present application. The video encoding device 300 includes an acquisition unit 310, an encoding unit 320, and a sending unit 330, wherein the acquisition unit 310 is used to acquire a to-be-processed video stream comprising M consecutive images, where M is a positive integer greater than or equal to 2; the encoding unit 320 is used to sequentially determine a motion region and a non-motion region of the i-th image in the M consecutive images, where i is any positive integer between 1 and M-1, and both the motion region and the non-motion region have the same quantization parameter QP value; reduce the QP value corresponding to the non-motion region to a first QP value; increase the QP value corresponding to the motion region to a second QP value; encode the non-motion region of the i-th image according to the first QP value, and encode the motion region of the i-th image according to the second QP value to generate a first video stream; and the sending unit 330 is used to send the first video stream to the network card, so that the network card sends the first video stream to the terminal device.
[0136] The acquisition unit, encoding unit, and sending unit can all be implemented by software or hardware. For example, the implementation of the encoding unit is described below using the encoding unit as an example. Similarly, the implementation of the acquisition unit and sending unit can refer to the implementation of the encoding unit.
[0137] As an example of a software functional unit, a module may include a code running on a computing instance. The computing instance may include at least one of a virtual machine and a container. Furthermore, the computing instance may be one or more. For example, the coding unit may include code running on multiple virtual machines / containers. It should be noted that the multiple virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Generally, a region may include multiple AZs.
[0138] Similarly, multiple virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0139] As an example of a hardware functional unit, the coding unit may be integrated into a GPU, or may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0140] It should be noted that the encoding unit 320 can be used to execute steps S220 to S260 in Figure 2. In other embodiments, the specific steps executed by the acquisition unit 310, the encoding unit 320 and the sending unit 330 can be specified as needed, and this application does not make specific limitations.
[0141] It should be understood that the video encoding device 300 shown in FIG3 is only one possible implementation provided by the present application. The video encoding device 300 may also include more types and numbers of units, which is not specifically limited by the present application.
[0142] For example, as shown in FIG4 , FIG4 is a schematic structural diagram of another video encoding device provided by an embodiment of the present application. The acquisition unit 310 in the video encoding device 300 may include a rendering module 311, a synthesis module 312, and a capture module 313. The rendering module 311 is used to load the original data of the generated M continuous images and send rendering instructions to the synthesis module, wherein the rendering instructions include the position, size, rendering mode, layer overlay method, etc. of the rendering area, which is not specifically limited in this application; the synthesis module 312 is used to receive the rendering instructions and the synthesized image represented in RGB; the capture module 313 is used to obtain an image represented in YUV by converting the color space according to the received RGB data. In addition to being used to obtain an image represented in YUV, the acquisition unit can also be used to receive a video request sent by a terminal device, which is not specifically limited in this application.
[0143] For example, as shown in FIG4 , the encoding unit 320 in the video encoding apparatus 300 may include a motion detection module 321, an image quality enhancement module 322, and an encoding module 323. The motion detection module 321 is configured to determine a motion region and a non-motion region of the i-th image; the image quality enhancement module 322 is configured to sharpen the motion region of the i-th image to obtain an i-th sharpened image; reduce the quantization parameter corresponding to the non-motion region of the i-th image to a first QP value; and increase the quantization parameter corresponding to the motion region of the i-th image to a second QP value; and the encoding module 323 is configured to perform video encoding on the non-sharp region according to the first QP value and on the sharp region according to the second QP value to generate a first video stream.
[0144] It should be understood that the acquisition unit 310 and the encoding unit 320 shown in Figure 4 are only one possible implementation method provided by this application. The acquisition unit 310 and the encoding unit 320 can also include more types and numbers of modules, and this application does not make specific limitations on this.
[0145] As shown in Figure 5, Figure 5 is a schematic diagram of the structure of a computing device 500 provided in an embodiment of the present application. The computing device can be used as a server in the system shown in Figure 1 to implement the video encoding method shown in Figure 2. The computing device includes at least a bus 510, a processor 520, a processor 521, a memory 530, and a communication interface 540. The processor 520, the processor 521, the memory 530, and the communication interface 540 communicate with each other via the bus 510. It should be understood that this application does not limit the number of processors and memories in the computing device 500.
[0146] Bus 510 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG5 shows only one line, but this does not imply a single bus or type of bus. Bus 510 may include a path for transmitting information between various components of computing device 500 (e.g., processor 520, processor 521, memory 530, and communication interface 540).
[0147] The processor 520 may include any one or more processors such as a central processing unit (CPU), a microprocessor (MP), or a digital signal processor (DSP), and is configured to execute program codes.
[0148] The processor 521 may be a graphics processing unit (GPU), etc., and may integrate one or more functions of the acquisition unit 310, the encoding unit 320, and the sending unit 330, which is not specifically limited in this application.
[0149] The memory 530 may include a volatile memory, such as a random access memory (RAM). The memory 530 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The memory may also include a combination of the above types. The memory 530 stores executable program code. The memory stores executable program code corresponding to one or more units of the acquisition unit 310, the encoding unit 320, and the sending unit 330. The processor 520 executes the executable program code to respectively implement the functions of the aforementioned acquisition unit, encoding unit, and sending unit, thereby implementing the video encoding method shown in Figure 2. That is, the memory 530 stores instructions for executing the video encoding method. In addition, the memory 530 may also store more types and quantities of data, such as images represented in YUV format, etc., which is not specifically limited in this application.
[0150] The communication interface 540 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 500 and other devices or a communication network.
[0151] It should be noted that FIG5 is only a possible implementation of an embodiment of the present application. In actual applications, the computing device may also include more or fewer components, and the present application does not make any specific limitations on this.
[0152] As shown in Figure 6, Figure 6 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. The computing device cluster includes at least one computing device, which can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone. The memory 530 in one or more computing devices 500 in the computing device cluster can store the same instructions for executing the video encoding method provided in the present application.
[0153] In some possible implementations, the memory 530 of one or more computing devices 500 in the computing device cluster may also store some instructions for executing the video encoding method. In other words, the combination of one or more computing devices 500 can jointly execute the instructions for executing the video encoding method.
[0154] It should be noted that the memory 530 in different computing devices 500 in the computing device cluster can store different instructions, each for executing a portion of the encoder's functions. That is, the instructions stored in the memory 530 in different computing devices 500 can implement the functions of one or more of the motion detection module, the processing module, and the encoding module.
[0155] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN). FIG7 is a schematic diagram of a structure in which one or more computing devices are connected via a network, as provided in an embodiment of the present application.
[0156] As shown in Figure 7, computing device 500A includes a bus 510A, a processor 520A, a memory 530A, and a communication interface 540A, and computing device 500B includes a bus 510B, a processor 520B, a memory 530B, and a communication interface 540B. The two computing devices 500A and 500B are connected via a network. Specifically, they are connected to the network via the communication interface in each computing device. In this type of possible implementation, when the video encoding device 300 is stored in the form of software, the memory 530A in the computing device 500A stores instructions for executing the functions of the acquisition unit 310 and the sending unit 330. At the same time, the memory 530B in the computing device 500B stores instructions for executing the functions of the encoding unit 320. The computing devices 500A and 500B can implement the video encoding method provided in this application through data exchange.
[0157] As shown in Figure 8, Figure 8 is a schematic diagram of another structure of one or more computing devices connected via a network provided by an embodiment of the present application. Computing device 500C includes a bus 510C, a processor 520C, a processor 521C, a memory 530C, and a communication interface 540C. Computing device 500D includes a bus 510D, a processor 520D, a processor 521D, a memory 530D, and a communication interface 540D.
[0158] When video encoding apparatus 300 is represented in hardware, processor 521C in computing device 500C integrates the functions of acquisition unit 310 and transmission unit 330, while processor 521D in computing device 500D integrates the functions of encoding unit 320. Computing devices 500C and 500D exchange data to implement the video encoding method shown in FIG.
[0159] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a storage service device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the video encoding method shown in FIG2 .
[0160] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute a video encoding method shown in FIG2.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A video encoding method based on cloud phones, characterized in that, The method is applied to a cloud mobile phone, which runs on a server. The server is also provided with a network card. The cloud mobile phone establishes a network connection with a terminal device through the network card. The method includes: The cloud mobile phone generates a video stream to be processed including M consecutive images, where M is a positive integer greater than or equal to 2; The cloud mobile phone sequentially determines a motion area and a non-motion area of the i-th image among the M consecutive images, where i is any positive integer between 1 and M-1. Both the motion area and the non-motion area have corresponding quantization parameter QP values; The cloud mobile phone reduces the QP value corresponding to the non-motion area to a first QP value; The cloud mobile phone raises the QP value corresponding to the motion area to a second QP value; The cloud mobile phone encodes the non-motion area of the i-th image according to the first QP value and encodes the motion area of the i-th image according to the second QP value to generate a first video bitstream; The cloud mobile phone sends the first video bitstream to the terminal device through the network card.
2. The method according to claim 1, characterized in that, The cloud mobile phone reducing the QP value corresponding to the non-motion area to the first QP value includes: The cloud mobile phone reduces the QP value corresponding to the non-motion area according to a first parameter to obtain the first QP value corresponding to the non-motion area, and the first parameter is a value less than zero.
3. The method according to claim 2, wherein The cloud mobile phone raising the QP value corresponding to the motion area to the second QP value includes: The cloud mobile phone raises the QP value corresponding to the motion area according to a second parameter to obtain the second QP value corresponding to the motion area, and the second parameter is a value greater than zero.
4. The method according to any one of claims 1-3, characterized in that, After the cloud mobile phone sequentially determines the motion area and the non-motion area of the i-th image among the M consecutive images, the method further includes: The cloud mobile phone sequentially performs a sharpening operation on the motion area of the i-th image to obtain an i-th sharpened image, where the i-th sharpened image includes a sharpened area and an unsharpened area. The sharpened area corresponds to the motion area of the i-th image, and the unsharpened area corresponds to the non-motion area of the i-th image. The edge information of the sharpened area is stronger than the edge information of the motion area.
5. The method according to claim 4, wherein The cloud mobile phone encoding the motion area of the i-th image according to the second QP value includes; The cloud mobile phone encodes the sharpened area of the i-th sharpened image according to the second QP value.
6. The method according to any one of claims 1 to 5, characterized in that The server is also provided with a graphics processing unit GPU, and the cloud mobile phone calls the GPU to implement the video encoding method.
7. The method according to any one of claims 1 to 6, characterized in that The cloud mobile phone is implemented by a virtual machine or a container running in the server.
8. A server, characterized in that, Including a cloud mobile phone and a network card, the cloud mobile phone runs in a server, and the cloud mobile phone establishes a network connection with a terminal device through the network card; The cloud mobile phone is used to generate a video stream to be processed including M consecutive images, where M is a positive integer greater than or equal to 2; The cloud phone is used to sequentially determine the motion area and the non-motion area of the i-th image among the M consecutive images, where i is any positive integer between 1 and M-1, and both the motion area and the non-motion area have corresponding quantization parameter QP values; The cloud phone is used to reduce the QP value corresponding to the non-motion area to a first QP value; The cloud phone is used to increase the QP value corresponding to the motion area to a second QP value; The cloud phone is used to encode the non-motion area of the i-th image according to the first QP value and encode the motion area of the i-th image according to the second QP value to generate a first video bitstream; The cloud phone is used to send the first video bitstream to the network card; The network card is used to send the first video bitstream to the terminal device.
9. The method according to claim 8, wherein Specifically, the cloud phone is used for: Reducing the QP value corresponding to the non-motion area according to a first parameter to obtain the first QP value corresponding to the non-motion area, where the first parameter is a value less than zero.
10. The method according to claim 9, wherein, Specifically, the cloud phone is used for: Increasing the QP value corresponding to the motion area according to a second parameter to obtain the second QP value corresponding to the motion area, where the second parameter is a value greater than zero.
11. The server according to any one of claims 8-10, characterized in that, The cloud phone is further used for: Performing a sharpening operation on the motion area of the i-th image to obtain an i-th sharpened image, where the i-th sharpened image includes a sharpened area and an unsharpened area, where the sharpened area corresponds to the motion area of the i-th image, the unsharpened area corresponds to the non-motion area of the i-th image, and the edge information of the sharpened area is stronger than the edge information of the motion area.
12. The server according to claim 11, wherein Specifically, the cloud phone is used for: Encoding the sharpened area of the i-th sharpened image according to the second QP value.
13. The server according to any one of claims 8 to 12, characterized in that, The server further includes a graphics processing unit GPU; The GPU is used to implement the operations performed by the cloud phone described in claims 7-10 through the call of the cloud phone.
14. The method according to any one of claims 8 to 13, characterized in that The server further includes at least one virtual machine or at least one container; The at least one virtual machine and the at least one container are used to provide a running environment for the cloud phone.
15. A video encoding device, characterized in that, Applied to a server, the server includes a network card and establishes a network connection with a terminal device through the network card. The video encoding device includes: An acquisition unit, used to acquire a to-be-processed video stream of M consecutive images, where M is a positive integer greater than or equal to 2; An encoding unit, used to sequentially determine the motion area and the non-motion area of the i-th image among the M consecutive images, where i is any positive integer between 1 and M-1, and both the motion area and the non-motion area have the same quantization parameter QP value; Reducing the QP value corresponding to the non-motion area to a first QP value; Increasing the QP value corresponding to the motion area to a second QP value; Encoding the non-motion area of the i-th image according to the first QP value and encoding the motion area of the i-th image according to the second QP value to generate a first video bitstream; A sending unit, configured to send the first video stream to the network card, so that the network card sends the first video stream to the terminal device.
16. A cluster of computing devices, characterized in that, It includes at least one computing device, and each computing device includes a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions, characterized in that, When the instructions are run by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
18. A computer-readable storage medium, characterized in that, It includes computer program instructions, and when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Signal processing device, control programme, and integrated circuit
CN102771111A
Code rate control method and device for sport video
CN105898306A
Video data transmission method based on cloud mobile phone
CN112383775A
Video coding method and device, electronic equipment and storage medium
CN113676730A
Video coding method and device
CN115842915A