Image processing method and related equipment
By selecting the adapted target super-score model to perform super-resolution processing by receiving the video stream and its characteristic parameters, the impact of low image quality of the video stream to be super-score in video scenes on the super-score process is solved, and the quality of the video stream is improved.
Patent Information
- Application Number
- CN202410405804.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-07
- Filing Date
- 2024-04-03
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, it is difficult to effectively reduce the impact of the low image quality of the video stream to be supersegmented on the image supersegment process in video scenarios, resulting in low quality of the video stream after supersegment.
By receiving the video stream and its characteristic parameters, including noise parameters, selecting the adapted target super-score model for super-resolution processing. There is a mapping relationship between this target super-segment model and the noise parameters, which can reduce the impact of noise on the super-segment process of the image.
The impact of the low image quality of the video stream to be supersegmented on the image supersegment process is effectively reduced, and the quality of the final generated video stream is improved.
Smart Images

Figure CN119967236A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 7, 2023, with application number 202311473673.4 and application name “Image processing method, device, computing device cluster and storage medium”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The embodiments of the present application relate to the field of communication technology, and in particular, to an image processing method and related equipment. Background Art
[0003] At present, in order to improve the quality of images, image super-resolution processing technology can be used. Image super-resolution processing technology can also be called image super-resolution reconstruction (SR). Image super-resolution processing technology is a technology that uses image processing technology and machine learning technology to construct a high-resolution (HR) image from one or more low-resolution (LR) images of the same scene.
[0004] In order to achieve super-resolution processing of images, a lightweight or ultra-lightweight model can be used, that is, quantization, pruning, knowledge distillation and other methods are used to compress and accelerate the model, and the model is used to perform super-resolution processing on the image. In video scenes, the image quality of the received video stream may not be high due to the influence of lighting conditions, imaging sensors, compression and quantization, etc. The received video stream is the image that needs to be processed by image super-resolution, hereinafter referred to as the video stream to be super-resolutioned. The aforementioned methods using lightweight or ultra-lightweight models mainly focus on problems such as high model calculation complexity, and cannot reduce the impact of the low image quality of the video stream to be super-resolutioned in video scenes on the image super-resolution process, resulting in low quality of the video stream obtained after image super-resolution processing of the video stream.
[0005] It can be seen that there is an urgent need for a method that can reduce the impact of low image quality of the video stream to be super-resolved on the image super-resolution process. Summary of the invention
[0006] The embodiments of the present application provide an image processing method and related equipment, which can reduce the impact of low image quality of a video stream to be super-resolved on the image super-resolution process and improve the quality of a finally generated video stream.
[0007] In a first aspect, the present application provides an image processing method, which is applied to a computing device, and the method includes: receiving a video stream; receiving characteristic parameters of the video stream sent by a server, the characteristic parameters including a noise parameter, the noise parameter being used to indicate the noise level of the video stream, the noise parameter including one or more of a noise mean and a noise variance; performing super-resolution processing on the video stream through a target super-resolution model to obtain a target video stream, the resolution of the target video stream being higher than the resolution of the video stream, the target super-resolution model having a mapping relationship with the characteristic parameters, and the target super-resolution model being a neural network model.
[0008] In the first aspect, a computing device receives a video stream, which is a video stream to be super-resolved, and the computing device obtains characteristic parameters of the video stream, which include: a noise parameter, through which the noise level of the video stream can be determined, that is, when the scheme performs super-resolution processing on the video stream to be super-resolved, the noise level of the video stream to be super-resolved is taken into account. On this basis, there is a mapping relationship between the noise parameter and the target super-resolved model for super-resolution processing of the video stream, for example, the mapping relationship can be: when the noise parameter indicates that the noise level is too high, the target super-resolved model is a model with strong noise reduction capability. That is, in the first aspect, the target super-resolved model is adapted to the noise parameter of the video stream to be super-resolved, and is also adapted to the noise level indicated by the noise parameter, so the target super-resolved model can reduce the influence of noise on the image super-resolved process. Noise is an important reason for the low quality of the video stream to be super-resolved, so the influence of the low image quality of the video stream to be super-resolved on the image super-resolved process can be reduced, and the quality of the video stream finally generated can be improved.
[0009] Furthermore, in the first aspect, the characteristic parameters of the video stream are sent by the server. Compared with the noise parameters of the video stream analyzed by the computing device, directly receiving the parameters sent by the server can save the computing power of the computing device and reduce the computing pressure of the computing device. For example, in the first aspect, a terminal-cloud collaborative solution can be applied, that is, the noise parameters are obtained by the server with a larger computing power than the computing device. For example, the server can analyze the video stream by itself to obtain the noise parameters. A server can control one or more computing devices. In the scenario where the server controls multiple computing devices, the noise parameters obtained by the server can be sent to multiple computing devices at the same time. Then, only one server needs to obtain the noise parameters, and each computing device does not need to analyze the noise parameters. This method can save the computing power of the system where the server and the computing device are located.
[0010] In a specific design, the characteristic parameter can be obtained by the server through pre-analysis. In a more specific design, the server obtains a video stream; the server pre-analyzes the video stream to obtain the characteristic parameter of the video stream; the server sends the characteristic parameter to the computing device, so that the computing device performs super-resolution processing on the video stream according to the characteristic parameter.
[0011] In a possible implementation manner of the first aspect, the computing device is deployed with multiple super-resolution models, and a mapping relationship exists between the multiple super-resolution models and multiple values of the feature parameter. The target super-resolution model is determined from the multiple super-resolution models according to the feature parameter.
[0012] In this possible implementation, multiple super-resolution models are deployed in the computing device, and different super-resolution models are required in different scenarios. Therefore, compared with deploying only one super-resolution model, the deployment of multiple super-resolution models can be applied in more scenarios, and there is a mapping relationship between these multiple super-resolution models and feature parameters with different values. It can be seen that the above implementation provides a super-resolution model that can be applied to video streams with different feature parameters. On this basis, selecting a target super-resolution model according to the feature parameter can ensure that the target super-resolution model finally selected is adapted to the feature parameter, so that the video stream with the feature parameter can be better processed.
[0013] In a specific design, in addition to the ability to perform super-resolution processing on the video stream, the multiple super-resolution models also have the ability to adjust the characteristic parameters of the video stream, such as the ability to reduce noise. Specifically, when training multiple super-resolution models, training losses are set for the multiple super-resolution models respectively, wherein the training losses include: errors for adjusting the characteristic parameters of the multiple video streams; and training the multiple super-resolution models through the multiple video streams based on the training losses. By setting the above-mentioned training losses, the multiple super-resolution models can have the ability to adjust the characteristics of the video stream. For example, when it is necessary to train the super-resolution model, and the error for adjusting the noise needs to be considered during the training process, the super-resolution error and the noise reduction error can be set for the super-resolution model, and the noise reduction error is used to measure the ability of the super-resolution model to adjust the noise characteristics of the image.
[0014] In a specific design, a reparameterization technique can be used to process the multiple super-resolution models. Specifically, when training multiple super-resolution models, a multi-branch parallel structure is used to enable these models to obtain better feature expressions. When super-resolution processing of the video stream is required, the parallel structure is fused into a serial structure, thereby reducing the amount of calculation and the amount of parameters and improving the speed.
[0015] In a specific design, a computing device receives a first instruction through a control interface, wherein the first instruction is used to indicate whether to determine the target super-resolution model from the multiple super-resolution models based on the feature parameter.
[0016] Based on the above scheme, since it takes a certain amount of computing power to determine the target super-resolution model from multiple super-resolution models according to the characteristic parameters, the user's requirements for computing power and the quality of the noise parameters of the video stream may be different. For example, when the user's demand for improving the quality of the noise parameters of the video stream is higher than the demand for saving computing power, the interface can be set to an open state, thereby instructing the computing device to determine the target super-resolution model from multiple super-resolution models according to the characteristic parameters. When the user's demand for saving computing power is higher than the demand for improving the quality of the noise parameters of the video stream, the interface can be set to a closed state, thereby instructing the computing device not to determine the target super-resolution model from multiple super-resolution models according to the characteristic parameters. It can be seen that by configuring the control interface, the diverse needs of users can be met.
[0017] In a specific design, the control interface includes: an on / off interface.
[0018] In a possible implementation of the first aspect, a feature level of the video stream is determined based on the feature parameter, the feature level includes a noise level, and the noise level is used to indicate the noise level of the video stream; the target super-resolution model is determined from multiple super-resolution models based on the feature level, wherein there is a mapping relationship between the multiple super-resolution models and the feature level.
[0019] In this possible implementation, the feature parameters may not be able to directly correspond to the target super-resolution model. At this time, an intermediary, namely the feature level, is needed to first determine the feature level corresponding to the feature parameters, and then find the corresponding target super-resolution model based on the feature level, thereby improving the adaptability between the feature parameters and the target super-resolution model.
[0020] In a possible implementation manner of the first aspect, the feature parameter further includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur degree.
[0021] In this possible implementation, the image visual feature parameter is used as one of the feature parameters because the video stream is affected by various factors, resulting in large differences in brightness, contrast and blur. In the implementation, the differences in the image visual feature parameters are taken into account, and there is a mapping relationship between the image visual feature parameters and the target super-resolution model. Therefore, the target super-resolution model is suitable for processing the video stream with the image visual feature parameters, thereby reducing the influence of factors such as brightness, contrast and blur on the image super-resolution process.
[0022] In a specific design, the feature level also includes one of a brightness level, a contrast level, and a blur level.
[0023] A second aspect of the present application provides an image processing method, which is applied to a server, and the method includes: determining a video stream; pre-analyzing the video stream through a pre-analysis model to obtain characteristic parameters of the video stream, the pre-analysis including noise analysis, the characteristic parameters including noise parameters, the noise parameters being used to indicate the noise level of the video stream, the noise parameters including one or more of a noise mean and a noise variance; sending the video stream to a receiving end; sending the characteristic parameters to the receiving end so that the receiving end performs super-resolution processing on the video stream based on the characteristic parameters.
[0024] In the second aspect, the server analyzes the above-mentioned characteristic parameters because the video stream is affected by various factors, which may cause differences in these characteristic parameters. Therefore, in the above implementation, the server analyzes the above-mentioned characteristic parameters and sends these characteristic parameters to the receiving end, so that the receiving end can consider these characteristic parameters when performing super-resolution processing on the video stream, and reduce the influence of characteristic parameters such as noise parameters on the super-resolution processing process. And in the second aspect, a terminal-cloud collaboration solution can be implemented, that is, the server with a larger computing power than the computing device pre-analyzes the characteristic parameters, and the computing device with a smaller computing power than the server only needs to receive the characteristic parameters sent by the server, thereby saving the computing power of the computing device. For example, a server can control one or more computing devices. In the scenario where the server controls multiple computing devices, the noise parameters obtained by the server can be sent to multiple computing devices at the same time, then only one server is required to obtain the noise parameters, instead of each computing device analyzing the noise parameters. This method can save the computing power of the system where the server and the computing device are located.
[0025] In a possible implementation of the second aspect, the video stream includes multiple first image frames, and the video stream is subjected to frame extraction processing to obtain at least one target image frame among the multiple first image frames in the video stream; the at least one target image frame is pre-analyzed through a pre-analysis model to obtain characteristic parameters of the video stream.
[0026] In this possible implementation, at least one target image frame is extracted from multiple first image frames by extracting frames from the video stream, and the at least one target image frame is pre-analyzed. Compared with analyzing all image frames included in the video stream, this method can save computing power consumed in pre-analysis.
[0027] In a specific design, the server may pre-analyze the noise parameters of the video stream by using a noise estimation model. Optionally, the noise estimation model may be a Poisson-Gaussian noise model.
[0028] In a possible implementation manner of the second aspect, the feature parameter further includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur degree.
[0029] In this possible implementation, the image visual feature parameter is used as one of the feature parameters because the video stream may be affected by various factors, resulting in large differences in brightness, contrast and blur. Therefore, the difference in the image visual feature parameters is taken into account in the implementation, and the image visual feature parameters are sent to the receiving end, so that the receiving end can consider the image visual feature parameters when performing super-resolution processing on the video stream, thereby reducing the influence of parameters such as brightness, contrast and blur on the super-resolution processing process.
[0030] In a specific design, the server receives a second instruction through a control interface, wherein the second instruction is used to indicate whether to perform the pre-analysis on the video stream.
[0031] In this possible implementation, since the server needs to occupy part of the server's computing resources for pre-analysis of the video stream, the higher the load of the computing resources, the slower the video stream is pushed, and the user's requirements for the video stream push speed and the quality of the video stream may be different. For example, when the user has higher requirements for the image quality for push, he can set the control interface to the open state. At this time, the server will pre-analyze the video stream and send the characteristic parameters obtained by the pre-analysis to the receiving end, so that the receiving end will adaptively select the super-resolution processing method according to the characteristic parameters, thereby improving the effect of super-resolution processing, improving the quality of the video stream, and meeting the user's requirements for high image quality. On the contrary, when the user has a relatively low requirement for image quality and a relatively high requirement for the image push speed, he can choose to set the interface to the closed state. It can be seen that by configuring this control interface, the diverse needs of users can be met.
[0032] In a possible implementation of the second aspect, the pre-analysis model is trained through multiple second image frames, wherein the multiple second image frames have different feature parameters; multiple super-resolution models are trained through the multiple second image frames so that the multiple super-resolution models have the ability to perform super-resolution processing on video streams with different feature parameters; and the multiple super-resolution models are sent to the receiving end.
[0033] In this possible implementation, when training the pre-analysis model and multiple super-resolution models, a multi-model joint training method is adopted, that is, the pre-analysis model and multiple super-resolution models are simultaneously optimized through multiple second image frames, so that the multiple super-resolution models can have the ability to perform super-resolution processing on the multiple image frames, and the pre-analysis model has the ability to pre-analyze the multiple image frames, so that the multiple super-resolution models and pre-analysis models finally trained can jointly analyze images such as video streams.
[0034] In a third aspect of the present application, an image processing device is provided, which has the function of implementing the method of the first aspect or any possible implementation of the first aspect, or the method of the second aspect or any possible implementation of the second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as an interface module and a processing module.
[0035] In a fourth aspect, the present application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory, and the memory of at least one computing device stores computer execution instructions that can be run on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect, or the method of the second aspect or any possible implementation of the second aspect.
[0036] In a fifth aspect, the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a computing device cluster, the computing device cluster executes a method such as the first aspect or any possible implementation of the first aspect, or a method such as the second aspect or any possible implementation of the second aspect.
[0037] In a sixth aspect, the present application provides a computer program product storing one or more computer execution instructions. The computer program product includes computer execution instructions. When the computer execution instructions are executed by a computing device cluster, the computing device cluster implements the method of the above-mentioned first aspect or any possible implementation of the first aspect, or the above-mentioned second aspect or any possible implementation of the second aspect.
[0038] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting a computing device cluster to implement the functions involved in the first aspect or any possible implementation of the first aspect, or the second aspect or any possible implementation of the second aspect. In a possible design, the chip system may also include a memory, which is used to store necessary program instructions and data. The chip system may be composed of a chip, or may include a chip and other discrete devices.
[0039] Among them, the technical effects brought about by the third to seventh aspects or any possible implementation methods thereof can be referred to the technical effects brought about by the first aspect or the related possible implementation methods of the first aspect, or the second aspect or the related possible implementation methods of the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A schematic diagram of a system architecture provided for an embodiment of the present application;
[0041] Figure 2 A schematic diagram of an application architecture provided for an embodiment of the present application;
[0042] Figure 3 A flowchart of a method 100 provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of obtaining noise parameters provided in an embodiment of the present application;
[0044] Figure 5 A flowchart of a method 100 provided in an embodiment of the present application;
[0045] Figure 6 is a flow chart of Example 1;
[0046] Figure 7 is a flow chart of Example 2;
[0047] Figure 8 is a schematic diagram of a noise estimation process provided by an embodiment of the present application;
[0048] Fig. 9 A schematic diagram of the network structure of the super-resolution model provided in an embodiment of the present application;
[0049] Fig.10 A schematic diagram of the model structure of the training phase and the reasoning phase provided in the embodiment of the present application;
[0050] Fig.11 is a schematic diagram of an image processing device provided in an embodiment of the present application;
[0051] Fig.12 is a structural diagram of a computing device provided in an embodiment of the present application;
[0052] Fig.13 is a schematic diagram of a structure of a computing device cluster provided in an embodiment of the present application;
[0053] Fig.14 This is another structural diagram of the computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0055] The following are some terms related to the embodiments of the present application for explanation.
[0056] (1) Super-resolution processing: Super-resolution processing refers to emphasizing the overall or local characteristics of an image, making an originally unclear image clear or emphasizing certain interesting features, expanding the differences between the features of different objects in the image, and suppressing uninteresting features, so as to improve image quality, enrich information, enhance image interpretation and recognition effects, and meet the needs of certain special analysis. In this solution, a typical scenario of super-resolution processing is the scenario of video super-resolution processing. In this scenario, image super-resolution can be used to improve the quality of the video.
[0057] (2) Noise: Noise often appears as isolated pixels or blocks of pixels on an image that cause strong visual effects. Generally, noise signals are irrelevant to the object to be studied. They appear in the form of useless information and disrupt the observable information of the image. Noise can be divided into stationary noise and non-stationary noise according to its statistical characteristics. Stationary noise can be further divided into Gaussian noise, Poisson noise, etc. based on the probability density function after statistics.
[0058] (3) Gaussian noise: In digital image processing, Gaussian noise is a common type of image noise that simulates the statistical characteristics of many noise sources in the real world. The main characteristic of Gaussian noise is that the changes in its pixel values follow a normal distribution, also known as a Gaussian distribution, hence the name. Gaussian noise is usually caused by the cumulative effect of many random events, such as thermal noise in image sensors, random currents in electronic components, etc. These random events will introduce random interference into the image, causing the pixel values of the image to shift and fluctuate.
[0059] (4) Poisson noise: Poisson noise, or shot noise, is a type of noise that can be modeled by a Poisson process. In electronics, Poisson noise is caused by the properties of discrete charges. Poisson noise is also generated when optical devices count photons, which is related to the uncertainty of photons. Poisson noise is a type of noise that is related to light intensity. The greater the light intensity, the greater the fluctuation in the number of received photons, and therefore the more severe the Poisson noise.
[0060] (5) The terms "system" and "network" in the embodiments of the present application can be used interchangeably. "At least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the objects associated with each other are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, "at least one of A, B and C" includes A, B, C, AB, AC, BC or ABC. And, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects, and are not used to limit the order, timing, priority or importance of multiple objects.
[0061] The following is an example of the system architecture of the embodiment of the present application.
[0062] like Figure 1 As shown, Figure 1 A schematic diagram of a possible, non-limiting system architecture provided by this application. The solution provided by this application can be applied to Figure 1 System 1000 is shown.
[0063] The system 1000 includes an image acquisition terminal 101, a server 102, and an image push terminal 103. The image acquisition terminal 101 is the subject of image dissemination, and the image push terminal 103 is the object or audience of image dissemination. Taking the live broadcast scene as an example, the image acquisition terminal 101 can be the host's mobile phone or computer, the image push terminal 103 can be the audience's mobile phone or computer, and the server 102 can be a streaming media server, a cloud server, or a computing server.
[0064] Optionally, the image acquisition end 101 and the image push end 103 may be computing devices. Exemplarily, any computing device may be a terminal device, or may be a server, a container, or a virtual machine.
[0065] Optionally, still taking the live broadcast scene as an example, from the specific image transmission process, the image acquisition terminal 101 will collect the video stream of the merchant anchor and other users, and then process the video stream, and encapsulate the processed video stream and upload it to the server 102. The content delivery network (CDN) of the server 102 can distribute the video stream, and the server 102 can also perform content detection and transcoding on the video stream. The server 102 can then send the processed video stream to the image push terminal 103. The image push terminal 103 receives the video stream sent by the server 102, decodes the video stream, and then displays the video stream to the user. Taking the live broadcast scene as an example, the user can be a live broadcast audience.
[0066] In the above process, the quality of the video stream sent by the server 102 to the image push end 103 may not be high. At this time, the image push end 103 can use super-resolution processing technology to improve the quality of the video stream. At present, image super-resolution processing technology can be applied to multiple scenarios such as live broadcast, online viewing of recorded videos, etc. In these scenarios, image super-resolution processing technology is facing challenges. Taking the RTC live broadcast scenario as an example, in this scenario, due to the limited computing power of the image push end 103 and the need to achieve real-time image super-resolution processing, a model with fast reasoning speed and lightweight is required. The following takes image super-resolution processing as an example of super-resolution of images. The existing lightweight super-resolution scheme is mainly proposed for problems such as large model parameters, high computational complexity, and slow reasoning speed. The current mainstream lightweight super-resolution schemes include: 1) Using lightweight technology in network design, using deep separable convolution, heterogeneous convolution and other methods to improve speed and reduce parameter quantity. 2) Using quantization, pruning, knowledge distillation and other methods to compress and accelerate the model, reduce the algorithm complexity of the model and reduce the size of the model.
[0067] However, super-resolution processing of video streams faces other challenges. For example, due to the limitations of hardware conditions when collecting video streams and the influence of factors such as lighting conditions during collection, the image quality of the collected video streams to be super-resolved may not be high. For another example, super-resolution processing of video streams is completed at the image push end 103, and the computing power of the image push end 103 is limited, which means that the design of the super-resolution network must be relatively simple, and it is necessary to improve the image quality of the super-resolved image as much as possible under the condition of limited resources. For another example, a single super-resolved model is difficult to process complex and diverse images and has poor adaptability. The aforementioned lightweight super-resolution solution is mainly proposed to address the problems of large model parameters, excessive computational complexity, and slow inference speed. It cannot cope with the problem of low image quality of the video stream to be super-resolved in video scenes, which leads to low quality of the video stream obtained after image super-resolution processing of the video stream, limited computing power of the image push end, that is, the computing device, and poor adaptability of a single super-resolved model.
[0068] Based on the above problems, the present invention proposes an image processing method 100. When the method 100 is applied to Figure 1 When the system is shown, the following computing devices may be Figure 1 The image push terminal 103 shown in the figure, the following server can be Figure 1 The server 102 shown. In the method 100, the computing device obtains the characteristic parameters of the video stream from the server, and uses these characteristic parameters as the basis for selecting the target super-resolution model. That is, the method 100 selects the corresponding model for the adaptability of these large differences. Specifically, the computing device receives the video stream, which is the video stream to be super-resolved. The computing device obtains the characteristic parameters of the video stream, and the characteristic parameters include: a noise parameter, through which the noise level of the video stream can be determined, that is, when the method 100 performs image super-resolution processing on the video stream to be super-resolved, the noise level of the video stream to be super-resolved is considered. In the method 100, the target super-resolved model is adapted to the noise parameters of the video stream to be super-resolved, and is also adapted to the noise level indicated by the noise parameters. The target super-resolved model can reduce the influence of noise on the image super-resolved process. Noise is an important reason for the low quality of the video stream to be super-resolved, so the influence of the low image quality of the video stream to be super-resolved on the image super-resolved process can be reduced, and the quality of the video finally generated can be improved.
[0069] See also Figure 2 , Figure 2 A schematic diagram of an application architecture provided in an embodiment of the present application. Figure 2 As shown, the server is configured with an on / off interface, and the user can use the on / off interface to choose whether to pre-analyze the video stream. For example, when the user sets the interface to on, the server pre-analyzes the video stream, and the pre-analysis result is the noise parameter of the video stream. The server then sends the pre-analysis result to the computing device. The computing device is configured with an on / off interface, and the user can use the on / off interface to choose whether to perform model selection based on the pre-analysis result. For example, the model selection is that the computing device selects a model for super-resolution processing of the image from multiple super-resolution models based on the pre-analysis result sent by the server. Optionally, method 100 and the following method 200 can be applied to Figure 2 The application architecture shown.
[0070] It should be noted that Figure 2 This is only an example of this solution. In this solution, the server may also pre-analyze other features of the video stream. Accordingly, the result of the pre-analysis may also be other feature parameters besides the noise parameter, such as brightness, contrast, etc.
[0071] The following is an introduction to the application scenarios of method 100. Method 100 can be applied to the following scenarios: scenarios where the video stream resolution is too low. When the video stream resolution received by the computing device is low, for example, during network transmission or due to bandwidth limitations, the image may not be clear enough. Through super-resolution processing, low-resolution images can be upgraded to higher resolutions to improve the viewing experience; real-time communication (RTC) scenarios. In real-time communication applications, such as video conferencing and live broadcasting, users expect to obtain high-quality images. Client super-resolution can provide clearer images in real-time scenarios without waiting for server processing. Scenarios with unstable network environments. When the network is unstable, transmitting high-resolution images may cause delays or freezes. By performing super-resolution on the client, the amount of transmitted data can be reduced and fluency can be improved. In privacy protection scenarios, super-resolution processing can be completed locally on the computing device without transmitting the original high-resolution image to the server for image super-resolution, which helps protect user privacy.
[0072] The above describes the application architecture and scenarios of the image processing method provided in the embodiment of the present application. The following describes in detail the execution process of the image processing method provided in the embodiment of the present application. Figure 3 , Figure 3 A schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application. Figure 3 As shown, the method 100 provided in the embodiment of the present application includes the following steps S301-S303.
[0073] S301, a computing device obtains a video stream;
[0074] It should be noted that the computing device in this solution is the push end of the video stream. From the perspective of communication relationship, the computing device in this solution is the object of communication and the device that receives the transmitted video stream. Figure 1 Taking the system architecture shown as an example, the computing device in this solution may be, for example, the image push terminal 103 .
[0075] S302, the computing device receives characteristic parameters of the video stream sent by the server;
[0076] It should be noted that the characteristic parameter includes a noise parameter, and the noise parameter is used to indicate the noise level of the video stream. The noise parameter includes one or more of a noise mean and a noise variance.
[0077] It should be noted that the server is located between the computing device and the image acquisition terminal and is used to distribute and manage the video stream. Figure 1Taking the system architecture shown as an example, the server in this solution may be, for example, the server 102 , which receives the video stream from the image acquisition terminal 101 , and then transcodes the video stream and sends it to the image push terminal 103 .
[0078] It should be noted that the mean is the mean of the noise of one or more image frames in the video stream. The larger the mean, the stronger the noise intensity and the higher the noise level. The intensity of Gaussian noise is usually described by the variance. The larger the variance, the stronger the noise and the higher the noise level. For example, in N~(0,σ 2 ) represents the normal distribution, which is used to model Gaussian noise that is independent of the signal, σ 2 It represents the variance of Gaussian noise. The larger the σ value is, the larger the variance of Gaussian noise is, the stronger the noise intensity is, and the higher the noise level is.
[0079] It should be noted that the feature parameter refers to a parameter used to describe the attributes of an image. In a possible implementation, the feature parameter also includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur.
[0080] It should be noted that the brightness of a video stream is the brightness of the video stream, and contrast refers to the measurement of different brightness levels between the brightest white and the darkest black in the light and dark areas of the image frame in the video stream. The larger the difference range, the greater the contrast, and the smaller the difference range, the smaller the contrast. For example, for digital image transformation, let the original pixel grayscale be f(i,j) and the converted pixel grayscale be g(i,j), then the commonly used linear transformation is g(i,j)=af(i,j)+b. When the feature parameter is contrast, the feature parameter can be the coefficient a of the function, and when the feature parameter is brightness, the feature parameter can be the coefficient b of the function. The blur degree of a video stream is used to describe the clarity of an image. The blur degree can also be called blur or clarity. For example, when the feature parameter is blur degree, the feature parameter can be a mean blur parameter, a Gaussian blur parameter, a median modulus parameter, or a bilateral blur parameter.
[0081] It should be noted that in this solution, the feature parameters are not limited to include noise parameters. Optionally, the feature parameters may also only include image visual feature parameters, and the image visual feature parameters include one of brightness, contrast, and blur. In method 100, the image visual feature parameters are not limited to only one of brightness, contrast, and blur. When selecting the target super-resolution model, multiple image visual features of the video stream may also be comprehensively considered. In this case, the image visual feature parameters include multiple of brightness, contrast, and blur.
[0082] In this solution, the reason why the above-mentioned characteristic parameters are received, or in other words, these characteristic parameters are considered, is that the video stream will be affected by various factors and cause noise parameters, and there are large differences in brightness, contrast and blur. These parameters are important factors affecting the image quality of the video stream to be super-resolved. If these characteristic parameters are not considered, the quality of the final video stream will also be affected. In particular, when the indicators indicated by these characteristic parameters are not ideal, that is, when the image quality of the video stream to be analyzed is not high, it is particularly important to consider these characteristic parameters. For example, when the noise parameter indicates that the noise intensity of the video stream is too high, it is necessary to consider the noise parameter during super-resolution processing and try to reduce the influence of the noise intensity. Because the noise will be amplified during super-resolution processing, when the noise intensity is relatively high, due to this amplification effect, the final image quality will be relatively poor.
[0083] Furthermore, due to hardware limitations, the collected video stream itself contains noise, and in the video scene, in order to save the video stream, some high-resolution images will be downsampled on the server, so it contains blur caused by downsampling. In addition, compression quantization noise will be introduced when encoding the video stream. It can be seen that in the video stream scene, noise problems and blur problems are more prominent. Therefore, in the video scene, when performing super-resolution processing on the video stream, it is particularly necessary to take characteristic parameters such as noise parameters or blur degree into consideration.
[0084] Optionally, the characteristic parameter may be obtained by the server through pre-analysis. The process of the server pre-analyzing the video stream to obtain the characteristic parameter may refer to the following method 200, which will not be described in detail here.
[0085] Take this solution applied to live broadcast as an example. Figure 4 The server receives a video stream sent by a video acquisition end, that is, a host end. The video stream includes multiple image frames. The server can sample some images from the video stream, that is, extract frames from the video stream to obtain a target image frame video stream among multiple first image frames, and perform noise estimation and prediction on the image frame through a noise estimation model to obtain noise parameters. The server compresses or downsamples the video in the video stream, and sends the noise parameters and the compressed or downsampled video to a computing device, which then selects a super-resolution model based on the noise parameters.
[0086] S303: The computing device performs super-resolution processing on the video stream through a target super-resolution model to obtain a target video stream.
[0087] It should be noted that the resolution of the target video stream is higher than the resolution of the video stream, there is a mapping relationship between the target super-resolution model and the feature parameter, and the target super-resolution model is a neural network model.
[0088] Optionally, the mapping of feature parameters to the target super-resolution model may be a direct mapping or may be mapped after numerical calculation. For example, the mapping relationship between the feature parameters and the target super-resolution model may be preset, and the mapping relationship does not need to be calculated. For example, it is preset that when the noise mean is 100dB, the corresponding target super-resolution model is super-resolution model 1. For example, there may be a functional relationship between the feature parameters and the target super-resolution model, and the function needs to be calculated to determine the target super-resolution model corresponding to the feature parameters.
[0089] In a possible implementation, the computing device is deployed with multiple super-resolution models, and there is a mapping relationship between the multiple super-resolution models and multiple values of the feature parameter. The computing device determines the target super-resolution model from the multiple super-resolution models according to the feature parameter.
[0090] In the above implementation, multiple super-resolution models are deployed in the computing device, and different super-resolution models are required in different scenarios. Therefore, compared with deploying only one super-resolution model, the deployment of multiple super-resolution models can be applied in richer scenarios, and there is a mapping relationship between these multiple super-resolution models and feature parameters with different values. It can be seen that this solution provides a super-resolution model that can be applied to video streams with different feature parameters. On this basis, selecting the target super-resolution model according to the noise parameter can ensure that the target super-resolution model finally selected is adaptive, so that the super-resolution model with the feature parameter can be better processed.
[0091] Optionally, multiple super-resolution models are respectively applicable to super-resolution processing of video streams with different characteristics. For example, M image super-resolution models can be trained in advance, and the M image super-resolution models include: image super-resolution model No. 1. The M image super-resolution models are respectively applicable to super-resolution of video streams with M noise levels. For example: Image super-resolution model No. 1 is applicable to super-resolution of video streams with the highest noise level, that is, the processing effect of image super-resolution model No. 1 on this type of video stream is better than that of other super-resolution models in the M image super-resolution models.
[0092] Optionally, in addition to the ability to perform super-resolution processing on the video stream, the multiple super-resolution models also have the ability to adjust the characteristic parameters of the video stream, such as the ability to reduce noise. Specifically, when training multiple super-resolution models, training losses are set for the multiple super-resolution models respectively, wherein the training losses include: errors for adjusting the characteristic parameters of the multiple video streams; and training the multiple super-resolution models through the multiple video streams based on the training losses. By setting the above-mentioned training losses, the multiple super-resolution models can have the ability to adjust the characteristics of the image. For example, when the super-resolution model needs to be trained and the error for adjusting the noise needs to be considered during the training process, the super-resolution error and the super-resolution ability of the super-resolution model can be set for the super-resolution model. In addition, the noise reduction error needs to be set to measure the ability of the super-resolution model to adjust the noise characteristics of the image.
[0093] Optionally, a reparameterization technique can be used to process the multiple super-resolution models. Specifically, when training multiple super-resolution models, a multi-branch parallel structure is used to enable these models to obtain better feature expressions. When super-resolution processing of the video stream is required, the parallel structure is fused into a serial structure, thereby reducing the amount of calculation and the amount of parameters and improving the speed.
[0094] After obtaining multiple super-resolution models suitable for super-resolution processing of video streams with different features, and feature parameters, the computing device determines a target super-resolution model from the multiple super-resolution models based on the feature parameters. Specifically, this can be done in the following way: in one possible implementation, a feature level of the video stream is determined based on the feature parameters, and the feature level includes a noise level, which is used to indicate the noise level of the video stream; the target super-resolution model is determined from the multiple super-resolution models based on the feature level, wherein there is a mapping relationship between the multiple super-resolution models and the feature level.
[0095] It is understandable that the characteristic parameters may not be able to directly correspond to the target super-resolution model. At this time, it is necessary to use an intermediary, namely the noise level, to find the corresponding target super-resolution model. For example, when the pre-analysis model is a Poisson-Gaussian noise model, the characteristic parameters are α and σ, α is a parameter related to noise in the Poisson distribution, and σ represents the standard deviation of Gaussian noise. At this time, the corresponding target super-resolution model cannot be directly obtained according to α and σ. However, there is a corresponding relationship between α and σ and the noise level. For example, in the Gaussian distribution, the higher σ is, the higher the noise intensity is, and the corresponding noise level is also higher. Therefore, the noise level corresponding to α and σ can be determined first, and there is a corresponding relationship between the noise level and the target super-resolution model, and the corresponding target super-resolution model can be determined according to these parameters.
[0096] Optionally, the feature level also includes one of a brightness level, a contrast level, and a blur level. Optionally, in method 100, the feature level is not limited to including a noise level. Optionally, the feature level may also include only one of a brightness level, a contrast level, and a blur level. Optionally, in method 100, the feature level is not limited to only one of a brightness level, a contrast level, and a blur level. When selecting a target super-resolution model, multiple image visual features of a video stream may be comprehensively considered. The feature level may also include multiple of a brightness level, a contrast level, and a blur level.
[0097] Since it takes a certain amount of computing power to determine the target super-resolution model from multiple super-resolution models according to the feature parameters, the user's requirements for computing power and the quality of the video stream may be different. In order to cope with the diverse needs of users, a control interface can be set on the computing device for the user to independently choose whether to determine the target super-resolution model from multiple super-resolution models according to the feature parameters. Specifically, it can be done in the following way: In a possible implementation, the computing device receives a first instruction through the control interface, wherein the first instruction is used to indicate whether to determine the target super-resolution model from the multiple super-resolution models according to the feature parameters. For example, the control interface can be an on / off interface. When the user's demand for improving the quality of the video stream is higher than the demand for saving computing power, the interface can be set to the on state, thereby instructing the computing device to determine the target super-resolution model from multiple super-resolution models according to the feature parameters. When the user's demand for saving computing power is higher than the demand for improving the quality of the video stream, the interface can be set to the off state, thereby instructing the computing device not to determine the target super-resolution model from multiple super-resolution models according to the feature parameters.
[0098] Optionally, after obtaining the target video stream, the computing device displays the target video stream to push the video stream to the user.
[0099] In method 100, a computing device receives a video stream, which is a video stream to be super-resolved. The computing device obtains characteristic parameters of the video stream, which include: a noise parameter, through which the noise level of the video stream can be determined, that is, when the scheme performs image super-resolution processing on the video stream to be super-resolved, the noise level of the video stream to be super-resolved is taken into account. On this basis, there is a mapping relationship between the noise parameter and the target super-resolved model for super-resolution processing of the video stream. For example, the mapping relationship can be: when the noise parameter indicates that the noise level is too high, the target super-resolved model is a model with strong noise reduction capability. That is, in method 100, the target super-resolved model is adapted to the noise parameter of the video stream to be super-resolved, and is also adapted to the noise level indicated by the noise parameter. The target super-resolved model can reduce the influence of noise on the image super-resolved process. Noise is an important reason for the low quality of the video stream to be super-resolved, so the influence of the low image quality of the video stream to be super-resolved on the image super-resolved process can be reduced, and the quality of the final generated video can be improved.
[0100] Furthermore, in method 100, the computing device may be configured with multiple super-resolution models, and these multiple super-resolution models may be respectively adapted to super-resolution processing of video streams with different characteristics. When the characteristic parameters of the video stream are different, the selected target super-resolution model is also different, ensuring that the target super-resolution model is adapted to the characteristics of the video stream, thereby solving the problem of poor adaptability of a single super-resolution model and improving the effect of super-resolution analysis of the video stream using the target super-resolution model.
[0101] In order to solve the problem of limited computing power of the aforementioned computing device, the present solution provides an image processing method 200. In the method 200, a server with a larger computing power than a receiving end pre-analyzes the feature parameters, and a receiving end with a smaller computing power than the server receives the feature parameters sent by the server without analyzing the feature parameters. The receiving end is a computing device, thereby saving the computing power of the computing device.
[0102] It should be noted that method 200 can be applied to method 100, and the receiving end in method 200 is the computing device in method 100. When method 200 is used in method 100, the server analyzes and obtains the characteristic parameters. After obtaining the characteristic parameters, the server sends the characteristic parameters to the computing device in method 100, and the corresponding computing device receives the characteristic parameters sent by the server. Method 200 can also be applied to other scenarios that require the server to determine the characteristic parameters. When method 200 is applied to Figure 1 In the system architecture shown, the server in method 200 may be server 102, but this solution does not limit this. The server in method 200 may also be other servers, such as other cloud servers, media servers, or computing servers.
[0103] See also Figure 5 , Figure 5 A schematic diagram of a method 200 provided in an embodiment of the present application. Figure 5 As shown, the method 200 provided in the embodiment of the present application includes the following steps S501-S504.
[0104] S501, the server determines the video stream;
[0105] Exemplarily, determining a video stream means determining a video stream that needs to be pre-analyzed from one or more video streams. Optionally, the server may determine the video stream in one or more of the following ways: Optionally, the server may receive a video stream sent by an image acquisition terminal, which is the video stream that needs to be pre-analyzed. The server may also determine the video stream that needs to be pre-analyzed from video streams stored in a cache or a disk. Specifically, the server may receive an instruction sent by a receiving terminal, which instructs the sending of a certain video stream and instructs the server to pre-analyze the video stream, so that the server can determine the video stream that needs to be pre-analyzed.
[0106] S502: The server pre-analyzes the video stream using a pre-analysis model to obtain characteristic parameters of the video stream.
[0107] It should be noted that the pre-analysis includes noise analysis, the characteristic parameter includes a noise parameter, the noise parameter is used to indicate the noise level of the video stream, and the noise parameter includes one or more of a noise mean and a noise variance.
[0108] The server can pre-analyze the video stream in the following ways:
[0109] In one possible implementation, the video stream includes multiple first image frames, and the video stream is subjected to frame extraction processing to obtain at least one target image frame among the multiple first image frames in the video stream; the at least one target image frame is pre-analyzed through a pre-analysis model to obtain characteristic parameters of the video stream.
[0110] In the above method, by extracting frames from the video stream, at least one target image frame is extracted from multiple first image frames, and the at least one target image frame is pre-analyzed. Compared with analyzing all image frames contained in the video stream, this method can save computing power consumed in pre-analysis.
[0111] Optionally, the server may pre-analyze the noise parameters of the video stream by using a noise estimation model. Optionally, the noise estimation model may be a Poisson-Gaussian noise model.
[0112] In a possible implementation manner, the feature parameter further includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur degree.
[0113] In the above manner, the image visual feature parameter is used as one of the feature parameters because the video stream is affected by various factors, resulting in large differences in brightness, contrast and blur. Therefore, the difference in the image visual feature parameters is taken into account in the above implementation, and the image visual feature parameters are sent to the receiving end, so that the receiving end can consider the image visual feature parameters when performing super-resolution processing on the video stream, thereby reducing the influence of parameters such as brightness, contrast and blur on the image super-resolution processing process.
[0114] In addition, since the server needs to occupy a part of the server's computing resources for pre-analysis of the video stream, a control interface can be configured on the server. Specifically, in one implementation, the server can receive a user's instruction through the control interface, wherein the instruction is used to indicate whether to perform the pre-analysis on the video stream. For example, an on / off interface can be configured on the server. When the user has higher requirements for image quality, he can set the interface to the on state. At this time, the server will pre-analyze the video stream and send the feature parameters obtained by the pre-analysis to the receiving end, so that the receiving end can adaptively select the image super-resolution processing method according to the feature parameters, thereby improving the effect of image super-resolution processing, improving the quality of the video stream, and meeting the user's requirements for high image quality. On the contrary, when the user has a relatively low requirement for image quality and a relatively high requirement for the push speed of the image, he can choose to set the interface to the off state. It can be seen that by configuring the control interface, the diverse needs of users can be met.
[0115] S503: The server sends the video stream to the receiving end;
[0116] Optionally, the server may perform other processing on the video stream before sending the video stream to the receiving end. For example, the server may transcode the video stream and distribute the content. At this time, there may be differences between the video stream determined by the server in step S501 and the video stream sent in step S503, but the difference is mainly in the form of video format, so the two are collectively referred to as video streams here.
[0117] S504: The server sends the characteristic parameter to the receiving end;
[0118] It should be noted that the purpose of the server sending the characteristic parameter to the receiving end is to enable the receiving end to perform super-resolution processing on the video stream based on the characteristic parameter.
[0119] In addition to pre-analyzing the video, the server can also train the model. Optionally, the server can train the pre-analysis model and multiple target super-resolution models. Specifically, the following implementation methods can be used:
[0120] In one possible implementation, a pre-analysis model is trained through multiple video streams, wherein the multiple video streams have different characteristics; and multiple target super-resolution models are trained through the multiple video streams so that the multiple target super-resolution models have the ability to analyze video streams with different characteristics.
[0121] In one possible implementation, the pre-analysis model is trained using multiple second image frames, wherein the multiple second image frames have different feature parameters; multiple super-resolution models are trained using the multiple second image frames so that the multiple super-resolution models have the ability to perform super-resolution processing on video streams with different feature parameters; and the multiple super-resolution models are sent to the receiving end.
[0122] Based on the above method, when training the pre-analysis model and multiple super-resolution models, a multi-model joint training method is adopted, that is, the pre-analysis model and multiple super-resolution models are simultaneously optimized through multiple second image frames, so that the multiple super-resolution models can have the ability to perform super-resolution processing on the multiple image frames, and the pre-analysis model has the ability to pre-analyze the multiple image frames, so that the multiple super-resolution models and pre-analysis models finally trained can jointly analyze images such as video streams.
[0123] It should be noted that the device for training the model can be the server in method 200, or other devices with strong computing power.
[0124] In one possible implementation, the server sets training losses for the multiple target super-resolution models respectively, wherein the training losses include: errors for adjusting features of the multiple video streams, and the multiple target super-resolution models are trained through the multiple video streams based on the training losses.
[0125] It is understandable that by setting the above training loss, these multiple target super-resolution models can have the ability to adjust the characteristics of the image. For example, when the super-resolution model needs to be trained and the error of adjusting the noise needs to be considered during the training process, the super-resolution error and the error used to measure the super-resolution ability of the super-resolution model can be set for the super-resolution model. In addition, the noise reduction error needs to be set to measure the ability of the super-resolution model to adjust the noise characteristics of the image.
[0126] Optionally, a reparameterization technique can be used to process the multiple target super-resolution models. Specifically, when training multiple target super-resolution models, a multi-branch parallel structure is used to enable these models to obtain better feature expressions. When super-resolution processing of the video stream is required, the parallel structure is fused into a serial structure, thereby reducing the amount of calculation and the amount of parameters and improving the speed.
[0127] For ease of understanding, the above method 100 and method 200 will be introduced below with reference to specific examples.
[0128] Example 1
[0129] Embodiment 1 is an example of method 100. Embodiment 1 is an example in which the computing device in method 100 is a video playback end and the characteristic parameter is a noise parameter.
[0130] See also Figure 6 , Figure 6 A flow chart of Example 1 provided for the embodiments of the present application. Figure 6 The specific process includes:
[0131] S601, the server sends the noise parameter and image data to the video player via the video stream;
[0132] It should be noted that the video stream of step S601 can also be called a bit stream. The video stream of step S601 refers to the video stream obtained by the server after transcoding the original video stream sent by the video acquisition terminal after receiving it, and the video stream carries noise parameters and image data. The image data can be a video, and the video includes multiple image frames. That is, in Example 1, the feature parameters are sent together with the video stream. Since the server can provide service support for the video playback services of multiple video playback terminals, the server can send the video stream pair carrying noise parameters to multiple video playback terminals. The users of these video playback terminals are all the audiences of the video and need to obtain the video, so that these video playback terminals can share the noise parameters obtained by the server and guide the subsequent super-resolution processing of the video stream through the noise parameters.
[0133] S602: The video player selects a target super-resolution model from M super-resolution models according to a noise parameter;
[0134] It should be noted that M≥2, and M is a positive integer.
[0135] In step S602, the corresponding noise level can be determined according to the noise parameter. For example, the noise parameter is σ. The smaller the value of σ, the lower the Gaussian noise level. At this time, the super-resolution model that can analyze the video stream with a lower noise level is selected as the target super-resolution model. When the noise parameter σ is larger, it means that the noise level is higher. At this time, the super-resolution model that can analyze the video stream with a stronger noise level is selected as the target super-resolution model.
[0136] S603, the video player performs super-resolution processing on the video stream through the target super-resolution model;
[0137] It should be noted that after receiving the processed video stream sent by the server, the video player will decode the video stream. The video stream in step S603 refers to the video stream decoded by the video player.
[0138] S604: The video player displays the processed video to the user.
[0139] The above embodiment makes full use of the computing resources of the server, relieves the computing pressure of the video player, and can realize noise estimation and image super-resolution in real time. By embedding noise parameters in the video stream, the noise parameters can be efficiently sent to multiple video player terminals, avoiding repeated noise estimation of the video stream at different video player terminals. And the video player terminal can share the noise estimation result of the server, and select the super-resolution model according to the noise estimation result. The above embodiment overcomes the problem that a single super-resolution model is difficult to cope with real complex scene images. The video player terminal can adaptively select a suitable super-resolution model from multiple super-resolution models according to the noise parameters of the video stream to perform image super-resolution, so as to obtain a super-resolution result with higher subjective quality.
[0140] Example 2
[0141] Embodiment 2 is an example of method 200. Embodiment 2 is an example in which method 200 is applied to a server, and the receiving end is a video playback end. Embodiment 2 is an example in which the pre-analysis model in method 200 is used as a noise estimation model, and the server is also used for training the pre-analysis model and the super-resolution model.
[0142] See also Figure 7 , Figure 7 A flow chart of Example 2 provided in the embodiments of the present application. Figure 7 The specific process includes:
[0143] S701, the server obtains N video streams with different characteristics;
[0144] It should be noted that in order to train the model, the server needs to have strong computing power.
[0145] Since the noise parameters of the video stream are analyzed during the pre-analysis in Embodiment 2, the different features of the “N video streams with different features” refer to different noise features, such as different noise intensities.
[0146] S702: The server obtains a noise estimation model and K super-resolution models;
[0147] like Figure 8 As shown, the noise estimation model may include multiple residual modules, and these multiple residual modules constitute a noise estimation network. For example, the noise estimation model may be composed of four residual modules (resblock) cascaded together, through which the noise estimation model can estimate the noise parameters of the image frame Y, obtain the parameters a and e of the estimated noise distribution, optimize the noise estimation network through noise information supervision, and optimize the super-resolution network through reconstruction loss.
[0148] Here we take the Poisson-Gaussian noise model as an example. This model is often used to model complex noise in real scenes. The noise model is:
[0149] Y=αP+N,
[0150] Where P represents Poisson distribution, which is used to model sensor-related noise. The so-called sensor-related noise refers to the noise introduced by the image sensor when collecting signals. α represents the parameter related to Poisson distribution noise, N~(0,σ 2 ) represents the normal distribution, which is used to model Gaussian noise that is independent of the signal, and σ represents the standard deviation of the Gaussian noise.
[0151] Each of the K super-resolution models above consists of 5 convolution layers, where K is greater than or equal to 2 and is a positive integer. Each of the K super-resolution models above consists of 5 convolution layers, see Fig. 9 , Fig. 9 The neural network of the super-resolution model has a 5-layer convolution structure, including 4 layers of 3×3 convolution (conv) and 1 layer of 1×1 conv. On the one hand, in order to obtain more features of the video stream, multiple layers of convolution need to be set. On the other hand, the more convolution layers are set, the higher the computational complexity will be, and the computing power on the end side is limited, so the number of convolution layers should not be too high. Based on these two aspects, a 5-layer convolution structure was created. Fig. 9 The depth2space in is an image processing operation that rearranges pixels in the depth (channel) dimension into spatial dimensions. It can be used to convert high-dimensional tensors into low-dimensional tensors for subsequent processing or analysis.
[0152] In order to further improve the super-resolution effect of the model without introducing additional reasoning burden, re-parameterization technology can be used in the training and reasoning process of these K super-resolution models. Re-parameterization technology can enrich the representation ability of the model and enable the model to obtain some information that is difficult for simple models to obtain. For details, please refer to Fig.10 , Fig.10 There are two stages in the process: training and inference. In the training stage, a complex model can be obtained by replacing a single convolutional layer with a complex parallel structure, that is, Fig.10 The five-layer parallel structure in the convolution operation. The five-layer parallel structure includes: the Laplacian operator for edge detection, the Sobel operator for detecting the vertical gradient of the image, that is, the Sobel operator is used to find the vertical edge, and the Sobel operator is used to find the horizontal gradient of the image, that is, the Sobel operator is used to find the horizontal edge. Through these parallel convolution operations, the gradient and scale information of the image can be obtained, which helps the model to better learn the characteristics of the image. After the training is completed, the complex parallel structure is converted into the structure of the inference model through reparameterization, that is, Fig.10 In the inference stage, the structure of the inference model is used to perform super-resolution analysis on the video stream.
[0153] S703, the server simultaneously optimizes the noise estimation model and K super-resolution models through N video streams with different characteristics;
[0154] In the process of modeling the noise estimation model, the video stream is analyzed by the noise estimation model to obtain the parameters a and e of the noise distribution, i.e., the noise parameters. The noise information of these N video streams with different characteristics is known, and the noise estimation model can be supervised and optimized through this noise information.
[0155] Furthermore, in the process of modeling the noise estimation model, the super-resolution model can be optimized by reconstructing the loss. Specifically, when setting the training loss of the super-resolution model, the training loss includes not only the super-resolution error but also the noise error, so that the super-resolution model has a certain denoising ability. By constructing training data sets with different noise levels or sizes and performing joint training, super-resolution models with different noise removal capabilities can be obtained.
[0156] S704, the server configures the super-resolution model in the video playback terminal;
[0157] The server may send the super-resolution model trained in step S703 to the video player, so that the video player may configure the super-resolution model and then perform super-resolution analysis on the image through the super-resolution model.
[0158] S705: The server receives the original video stream sent by the video acquisition terminal;
[0159] It should be noted that the video acquisition end here can be a mobile phone, computer or other device used by the anchor, etc. The video acquisition end can collect or acquire the video recorded by the anchor, etc. The code stream here refers to the digital audio and video stream transmitted on the network after the analog audio and video signal is imaged, collected, and encoded. This specifically refers to the digital video stream transmitted on the network. After receiving the original video stream, the server can perform operations such as transcoding and data distribution on the original video stream.
[0160] S706: The server extracts some image frames from the original video stream, and analyzes the image frames using a noise estimation model to obtain noise parameters of the image frames;
[0161] The above-mentioned "extracting some image frames from the original video stream" means sampling the original video stream. Since the features of the image frames in the original video stream will show more similarities, and analyzing all the original video streams will consume a lot of computing power. Therefore, the noise parameters of the original video stream are analyzed here by sampling.
[0162] S707: The server sends the noise parameter to the video playback end.
[0163] Since the server can provide service support for video playback services of multiple video playback terminals, the server can send bitstream pairs carrying noise parameters to multiple video playback terminals. The users of these video playback terminals are all audiences of the video and need to obtain the video. Therefore, these video playback terminals can share the noise parameters obtained by the server and use the noise parameters to guide subsequent super-resolution processing of the video stream.
[0164] Through the training method in the above embodiment, the super-resolution model can have a certain noise reduction ability, and the super-resolution model can be adapted to the super-resolution analysis of video streams with different characteristics. In addition, the above embodiment adopts a joint training method. When analyzing the super-resolution model, the parameters obtained by the noise estimation model can be shared, thereby improving the generalization ability and training efficiency of the super-resolution model. At the same time, joint training can also improve the performance of the model through the mutual influence between the super-resolution task and the noise estimation task. In addition, in the above embodiment, an end-cloud collaborative solution is adopted, that is, the server pre-analyzes the video stream, and the video acquisition end performs super-resolution processing on the image. There is no need for the video acquisition end to pre-analyze the video stream, thereby reducing the computing pressure of the video acquisition end and saving the computing power of the video acquisition end.
[0165] The above is an introduction to the method flow provided by the present application. Based on the above method flow, the following is an introduction to the device provided by the present application.
[0166] See also Fig.11 , a schematic diagram of the structure of an image processing device provided by the present application, comprising:
[0167] See also Fig.11 , Fig.11 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application. Fig.11 In the example shown, the image processing device 1100 is used to implement the various steps performed by the distributed storage system in the above embodiments. The image processing device 1100 includes an interface module 1101 and a processing module 1102 .
[0168] As an example, the image processing device 1100 can implement the functions of the computing device in the above method 100, and thus can also achieve the beneficial effects possessed by the above method 100. Specifically, the interface module 1101 is used to: receive a video stream; receive characteristic parameters of the video stream sent by a server, the characteristic parameters including a noise parameter, the noise parameter being used to indicate the noise level of the video stream, the noise parameter including one or more of a noise mean and a noise variance; the processing module 1102 is used to: perform super-resolution processing on the video stream through a target super-resolution model to obtain a target video stream, the resolution of the target video stream being higher than the resolution of the video stream, the target super-resolution model having a mapping relationship with the characteristic parameter, and the target super-resolution model being a neural network model.
[0169] In a possible implementation, the computing device is deployed with multiple super-resolution models, and there is a mapping relationship between the multiple super-resolution models and multiple values of the feature parameter. The processing module 1102 is used to determine the target super-resolution model from the multiple super-resolution models according to the feature parameter.
[0170] In one possible implementation, the processing module 1102 is used to: determine a feature level of the video stream based on the feature parameters, the feature level including a noise level, and the noise level is used to indicate a noise level of the video stream; determine the target super-resolution model from a plurality of super-resolution models based on the feature level, wherein a mapping relationship exists between the plurality of super-resolution models and the feature level.
[0171] In a possible implementation manner, the feature parameter further includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur degree.
[0172] As another example, the image processing device 1100 can implement the function of the server in the above method 200, and thus can also achieve the beneficial effects of the above method 200. Specifically, the processing module 1102 is used to: determine a video stream; pre-analyze the video stream through a pre-analysis model to obtain characteristic parameters of the video stream, the pre-analysis includes noise analysis, the characteristic parameters include noise parameters, the noise parameters are used to indicate the noise level of the video stream, the noise parameters include one or more of a noise mean and a noise variance; the interface module 1101 is used to: send the video stream to a receiving end; send the characteristic parameters to the receiving end, so that the receiving end performs super-resolution processing on the video stream based on the characteristic parameters.
[0173] In a possible implementation, the video stream includes multiple first image frames, and the processing module 1102 is used to: perform frame extraction processing on the video stream to obtain at least one target image frame among the multiple first image frames in the video stream; and perform pre-analysis on the at least one target image frame through a pre-analysis model to obtain characteristic parameters of the video stream.
[0174] In a possible implementation manner, the feature parameter further includes an image visual feature parameter, and the image visual feature parameter includes one of brightness, contrast, and blur degree.
[0175] In one possible implementation, the processing module 1102 is used to: train the pre-analysis model through multiple second image frames, wherein the multiple second image frames have different feature parameters; train multiple super-resolution models through the multiple second image frames so that the multiple super-resolution models have the ability to perform super-resolution processing on video streams with different feature parameters; and the interface module 1101 is used to: send the multiple super-resolution models to the receiving end.
[0176] The interface module and the processing module may be implemented by software or hardware. For example, the implementation of the interface module is described below by taking the interface module as an example. Similarly, the implementation of the processing module may refer to the implementation of the interface module.
[0177] As an example of a software functional unit, the interface module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the interface module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region can include multiple AZs.
[0178] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0179] As an example of a hardware functional unit, the interface module may include at least one computing device, such as a server, etc. Alternatively, the interface module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0180] The multiple computing devices included in the interface module can be distributed in the same region or in different regions. The multiple computing devices included in the interface module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the interface module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0181] It should be noted that, in other embodiments, both the interface module and the processing module can be used to execute any step in the image processing method. The steps that the processing module is responsible for implementing can be specified as needed. The full functions of the image processing device can be realized by implementing different steps in the image processing method separately through the processing module.
[0182] See also Fig.12 , Fig.12 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. Fig.12 As shown, the computing device 120 includes: a bus 122, a processor 124, a memory 126 and a communication interface 128. The processor 124, the memory 126 and the communication interface 128 are coupled via the bus 122 (not marked in the figure). The memory 126 stores instructions. When the execution instructions in the memory 126 are executed, the computing device 120 executes the method executed by the computing device in the above method embodiment.
[0183] The computing device 120 may be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASIC), or one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA), or a combination of at least two of these integrated circuit forms. For another example, when a unit in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call a program. For another example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0184] The processor 124 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0185] The memory 126 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0186] The memory 126 stores executable program codes, and the processor 124 executes the executable program codes to respectively implement the functions of the aforementioned units or modules, thereby implementing the aforementioned data writing method. That is, the memory 126 stores instructions for executing the aforementioned data writing method.
[0187] The communication interface 128 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 120 and other devices or a communication network.
[0188] In addition to the data bus, the bus 122 may also include a power bus, a control bus, and a status signal bus. The bus may be a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0189] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0190] like Fig.13 As shown, the computing device cluster includes at least one computing device 120. The memory 126 in one or more computing devices 120 in the computing device cluster may store the same instructions for executing the image processing method.
[0191] In some possible implementations, the memory 126 of one or more computing devices 120 in the computing device cluster may also store partial instructions for executing the image processing method. In other words, the combination of one or more computing devices 120 may jointly execute instructions for executing the image processing method.
[0192] It should be noted that the memory 126 in different computing devices 120 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the white box testing apparatus. That is, the instructions stored in the memory 126 in different computing devices 120 can implement the functions of one or more modules in the interface module 1101 and the processing module 1102.
[0193] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.14 A possible implementation is shown. Fig.14As shown, two computing devices 120A and 120B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 126 in the computing device 120A stores instructions for executing the functions of the interface module 1101. At the same time, the memory 126 in the computing device 120B stores instructions for executing the functions of the processing module 1102.
[0194] It should be understood that Fig.14 The functionality of the computing device 120A shown in FIG. 1 may also be performed by multiple computing devices 120. Similarly, the functionality of the computing device 120B may also be performed by multiple computing devices 120.
[0195] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the image processing method.
[0196] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the image processing method.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image processing method, characterized in that: Applied to a computing device, the method comprises: Receive video stream; Receiving characteristic parameters of the video stream sent by a server, the characteristic parameters including noise parameters, the noise parameters are used to indicate the noise level of the video stream, and the noise parameters include one or more of a noise mean and a noise variance; The video stream is super-resolution processed by a target super-resolution model to obtain a target video stream, wherein the resolution of the target video stream is higher than the resolution of the video stream, and there is a mapping relationship between the target super-resolution model and the feature parameters, and the target super-resolution model is a neural network model.
2. The image processing method according to claim 1, characterized in that: The computing device is deployed with a plurality of super-resolution models, and a mapping relationship exists between the plurality of super-resolution models and a plurality of values of the feature parameter. Before the super-resolution processing is performed on the plurality of first image frames by the target super-resolution model, the method further includes: The target super-resolution model is determined from the multiple super-resolution models according to the feature parameters.
3. The image processing method according to claim 2, characterized in that: The step of determining the target super-resolution model from a plurality of super-resolution models according to the feature parameters comprises: Determining a feature level of the video stream according to the feature parameter, the feature level comprising a noise level, and the noise level is used to indicate a noise level of the video stream; The target super-resolution model is determined from a plurality of super-resolution models according to the feature level, wherein a mapping relationship exists between the plurality of super-resolution models and the feature level.
4. The image processing method according to any one of claims 1 to 3, characterized in that: The feature parameters also include image visual feature parameters, and the image visual feature parameters include one of brightness, contrast, and blur degree.
5. An image processing method, characterized in that: The method comprises: Determine the video stream; Pre-analyzing the video stream by using a pre-analysis model to obtain characteristic parameters of the video stream, wherein the pre-analysis includes noise analysis, and the characteristic parameters include noise parameters, and the noise parameters are used to indicate the noise level of the video stream, and the noise parameters include one or more of a noise mean and a noise variance; Sending the video stream to a receiving end; The characteristic parameters are sent to the receiving end, so that the receiving end performs super-resolution processing on the video stream based on the characteristic parameters.
6. The image processing method according to claim 5, characterized in that: The video stream includes a plurality of first image frames, and the pre-analysis of the video stream by the pre-analysis model to obtain characteristic parameters of the video stream includes: Performing frame extraction processing on the video stream to obtain at least one target image frame among the multiple first image frames; The at least one target image frame is pre-analyzed by using a pre-analysis model to obtain characteristic parameters of the video stream.
7. The image processing method according to claim 5 or 6, characterized in that: The feature parameters also include image visual feature parameters, and the image visual feature parameters include one of brightness, contrast, and blur degree.
8. The image processing method according to any one of claims 5 to 7, characterized in that: The method further comprises: Training the pre-analysis model through a plurality of second image frames, wherein the plurality of second image frames have different characteristic parameters; Training a plurality of super-resolution models through the plurality of second image frames, so that the plurality of super-resolution models have the ability to perform super-resolution processing on video streams with different feature parameters; The multiple super-resolution models are sent to the receiving end.
9. An image processing device, characterized in that: include: Interface modules for: Receive video stream; Receiving characteristic parameters of the video stream sent by a server, the characteristic parameters including noise parameters, the noise parameters are used to indicate the noise level of the video stream, and the noise parameters include one or more of a noise mean and a noise variance; A processing module is used to: perform super-resolution processing on the video stream through a target super-resolution model to obtain a target video stream, wherein the resolution of the target video stream is higher than the resolution of the video stream, and there is a mapping relationship between the target super-resolution model and the feature parameters, and the target super-resolution model is a neural network model.
10. The device according to claim 9, characterized in that The computing device is deployed with multiple super-resolution models, and there is a mapping relationship between the multiple super-resolution models and the multiple values of the feature parameters. The processing module is used to: determine the target super-resolution model from the multiple super-resolution models according to the feature parameters.
11. The device according to claim 10, characterized in that The processing module is used for: Determining a feature level of the video stream according to the feature parameter, the feature level comprising a noise level, and the noise level is used to indicate a noise level of the video stream; The target super-resolution model is determined from a plurality of super-resolution models according to the feature level, wherein a mapping relationship exists between the plurality of super-resolution models and the feature level.
12. The device according to any one of claims 9 to 11, characterized in that The feature parameters also include image visual feature parameters, and the image visual feature parameters include one of brightness, contrast, and blur degree.
13. An image processing device, characterized in that: include: Processing modules for: Determine the video stream; Pre-analyzing the video stream by using a pre-analysis model to obtain characteristic parameters of the video stream, wherein the pre-analysis includes noise analysis, and the characteristic parameters include noise parameters, and the noise parameters are used to indicate the noise level of the video stream, and the noise parameters include one or more of a noise mean and a noise variance; Interface modules for: Sending the video stream to a receiving end; The characteristic parameters are sent to the receiving end, so that the receiving end performs super-resolution processing on the video stream based on the characteristic parameters.
14. The device according to claim 13, characterized in that The video stream includes a plurality of first image frames, and the processing module is used to: perform frame extraction processing on the video stream to obtain at least one target image frame among the plurality of first image frames in the video stream; The at least one target image frame is pre-analyzed by using a pre-analysis model to obtain characteristic parameters of the video stream.
15. The device according to claim 13 or 14, characterized in that The feature parameters also include image visual feature parameters, and the image visual feature parameters include one of brightness, contrast, and blur degree.
16. The device according to any one of claims 13 to 15, characterized in that Processing modules for: Training the pre-analysis model through a plurality of second image frames, wherein the plurality of second image frames have different characteristic parameters; Training a plurality of super-resolution models through the plurality of second image frames, so that the plurality of super-resolution models have the ability to perform super-resolution processing on video streams with different feature parameters; The interface module is used to send the multiple super-resolution models to the receiving end.
17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 4.
18. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 5 to 8.
19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 4.
20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 5 to 8.
21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 4.
22. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method as claimed in any one of claims 5 to 8.
Citation Information
Cited By
Multi-user super-resolution processing method and device, equipment and medium
CN121437271A
Multi-user super-resolution processing method, apparatus, device and medium
CN121437271B