Resolution enhancement method, electronic equipment, storage medium and product
By using the super-segment model in the video stream to process the Y component and combine the UV component to generate a high-resolution video stream, the problem of image quality degradation caused by the reduction of video stream resolution when the network state is poor is solved, improving user experience and reducing computing costs.
Patent Information
- Application Number
- CN202510152844.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-17
AI Technical Summary
When the network state is poor, reducing the resolution of the video stream will lead to a degradation of image quality and poor user experience.
By acquiring the Y component, UV component and RGB image of the video stream, calculating the complexity of the RGB image and determining an appropriate super-segment model, processing the Y component to generate a high-resolution first image, and generating a target video stream in combination with the UV component.
When the resolution of the video stream is low, the resolution enhancement of the Y component can be effectively improved, the image quality and user experience can be improved, while reducing the calculation amount and reducing the calculation cost.
Smart Images

Figure CN120163711A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of LED display, and specifically, the present application relates to a resolution enhancement method, an electronic device, a storage medium, and a product. Background Art
[0002] With the vigorous development of multimedia technology, services based on video streams (such as online meetings, live broadcasts, etc.) have become an indispensable service in people's daily lives. At the same time, people's demand for video resolution has gradually developed from standard definition (SD), high definition (HD) to ultra-high definition (UHD), etc. In order to obtain higher picture quality and bring a better viewing experience to users, transmitting ultra-high definition videos usually requires a high-speed and stable downlink bandwidth. For example, the transmission bit rate of a 4K video (resolution 3840×2160) is usually between 20-50 Mbps, and the transmission bit rate of an 8K video (resolution 7680×4320) is higher than 100 Mbps. However, with the continuous increase in the user scale and video viewing duration in the Internet video industry, this demand for high-speed and stable bandwidth has grown explosively, posing a huge challenge to the existing network environment.
[0003] However, due to the real-time changes in the network environment, users often experience video stuttering due to insufficient bandwidth when watching ultra-high definition videos. In order to reduce the impact of stuttering, when transmitting a video stream, the video resolution and frame rate are often dynamically adjusted according to the current network condition. For example, when the network condition is poor, the resolution of the video stream will be reduced, but directly reducing the resolution will cause a significant decline in image quality and a poor user experience. Summary of the Invention
[0004] In view of the shortcomings of the existing methods, the present application provides a resolution enhancement method, an electronic device, a storage medium, and a product, which can solve the problem that when the existing network state is poor, reducing the resolution of the video stream results in poor image quality and a reduced user experience.
[0005] According to one aspect of the embodiments of the present application, the embodiments of the present application provide a resolution enhancement method, the method including:
[0006] If it is determined that the resolution of the video stream is less than a preset value, obtain the Y component, the UV component, and the RGB image corresponding to the video stream;
[0007] Obtain the complexity of the RGB image, and determine the super-resolution model corresponding to the Y component based on the complexity;
[0008] Process the Y component with the super-resolution model to generate a first image, and generate a target video stream based on the first image and the UV component, where the resolution of the first image is higher than that of the Y component.
[0009] In a possible implementation, the complexity of obtaining the RGB image includes:
[0010] Input the RGB image into an image complexity discriminator, and determine the complexity using the classification information output by the image complexity discriminator.
[0011] In a possible implementation, the image complexity discriminator includes multiple convolutional layers and activation functions. The obtaining of the classification information includes:
[0012] Use the convolutional layer to extract the features of the RGB image, generate a feature map, and obtain the predicted value corresponding to the feature map;
[0013] Process the predicted value using the activation function to obtain the classification information.
[0014] In a possible implementation, the processing of the Y component using the super-resolution model to generate the first image includes:
[0015] Slice the Y component into image blocks of a preset size, and obtain the complexity information corresponding to the image blocks. The complexity information includes a complexity score and an index;
[0016] Determine the processing branch corresponding to each image block based on the complexity information, and process the image block using the processing branch to obtain an inference result. The processing branch includes a convolutional branch and a self-attention branch;
[0017] Merge the inference results to obtain a merged result, and upsample the merged result to obtain the first image.
[0018] In a possible implementation, the complexity includes simple, medium, and complex. The proportionality coefficients of the super-resolution models corresponding to different complexities are different, and the proportionality coefficients are used to adjust the usage ratios of the convolutional branch and the self-attention branch.
[0019] In a possible implementation, the obtaining of the complexity information corresponding to the image block includes:
[0020] Input the RGB image including the image block into a preset convolutional layer to obtain a feature value;
[0021] Input the feature value into a mask prediction network to obtain a complexity probability distribution map. The complexity probability distribution map includes the complexity score corresponding to each image block and the index corresponding to the complexity score. The size of the complexity score corresponds to the complexity degree of the image block.
[0022] In a possible implementation, the generating of the target video stream based on the first image and the UV image includes:
[0023] Interpolate the UV components using interpolation to obtain a second image, where the resolution of the second image is the same as that of the first image;
[0024] Replace the Y component in the second image with the first image to obtain a target image, and generate a target video stream according to the target image.
[0025] According to one aspect of the embodiments of the present application, embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of any of the above methods.
[0026] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the steps of the above method are implemented.
[0027] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer program product, including a computer program, and when the computer program is executed, the steps of the above method are implemented.
[0028] The beneficial technical effects brought by the technical solutions provided by the embodiments of the present application include:
[0029] A resolution enhancement method provided by the present application has the beneficial effect that if it is determined that the resolution of a video stream is less than a preset value, the Y component, UV components, and RGB image corresponding to the video stream are obtained; the complexity of the RGB image is obtained, and a super-resolution model corresponding to the Y component is determined based on the complexity; the Y component is processed using the super-resolution model to generate a first image, and a target video stream is generated based on the first image and the UV components. The resolution of the first image is higher than that of the Y component. When the resolution of the video stream is less than the preset value, the present application selects a super-resolution model corresponding to the Y component according to the RGB image corresponding to the video stream, uses the super-resolution model to process the Y component to obtain a high-resolution first image, and obtains a high-resolution target video stream through the first image. The present application can effectively enhance the resolution of the video stream through the resolution enhancement of the Y component when the resolution of the video stream is low, so that the image quality can be improved and the user experience can be improved after the source resolution of the video stream is reduced, and the calculation amount can be reduced and the calculation cost can be reduced during resolution enhancement.
[0030] Additional aspects and advantages of the present application will be given in part in the following description, and these will become obvious from the following description, or can be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0032] Figure 1 It is a flowchart of the resolution enhancement method provided by the embodiment of the present application;
[0033] Figure 2 It is a flowchart of the resolution discrimination provided by the embodiment of the present application;
[0034] Figure 3 It is a working flowchart of the resolution enhancement provided by the embodiment of the present application;
[0035] Figure 4 It is a flowchart of the super-resolution model selection provided by the embodiment of the present application;
[0036] Figure 5 It is a flowchart of the super-resolution model inference provided by the embodiment of the present application;
[0037] Figure 6 It is a structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0038] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0039] Those skilled in the art of the present technology can understand that unless specifically stated, the "the" and "this" used here can also include the plural form. It should be further understood that the term "including" used in the specification of the present application means that there are the described features, integers, steps, operations, elements and / or components, but does not exclude other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here means at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0040] To make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0041] The embodiment of the present application provides a resolution enhancement method, and this method can be used in computers, servers, mobile phones, and other devices that can acquire video streams and perform resolution enhancement on the video streams.
[0042] Optionally, the resolution enhancement method can be used in video conferencing and can be executed by a server for video stream transmission or a conference terminal for playing the video stream in the video conferencing.
[0043] As Figures 1 - 5 shown, the resolution enhancement method of the present application includes:
[0044] S101: If it is determined that the resolution of the video stream is less than a preset value, obtain the Y component, UV component, and RGB image corresponding to the video stream.
[0045] Optionally, the video stream can be transmitted in a video conference. After receiving the video stream, perform a decoding operation on the video stream, and after decoding, detect the resolution of the decoded video stream.
[0046] Optionally, the format of the video stream can be h.265, h.264, av1, hevc, and other formats that can be used for video transmission.
[0047] Optionally, the preset value can be 3840*2160, or 1280x720, and other resolution values. The size of the preset value can be set according to user requirements. It is also possible to obtain the user's resolution setting information (such as 1080p, 720p, etc.) after playing the video stream, and determine the size of the preset value according to the resolution setting information and the aspect ratio of the screen for displaying the video stream.
[0048] Optionally, after it is determined that the resolution of the video stream is not less than the preset value, the image can be directly displayed according to the video stream.
[0049] In one embodiment, as Figure 2 shown, the Y component, UV component, and RGB image can be obtained through a super-resolution module. After obtaining the transmitted video stream, perform decoding processing on the video stream to obtain the resolution of the decoded video stream. Detect whether the resolution is less than the preset value. If so, transmit the video stream to the super-resolution module, and obtain the Y component, UV component, and RGB image corresponding to the video stream through the super-resolution module. If it is determined that the resolution is not less than the preset value, play the video according to the video stream.
[0050] Optionally, after the video stream is decoded, YUV data containing each frame of image will be obtained, and each frame of image in the YUV data is in YUV format. A hardware accelerator for separating components can be called to process multiple frames of images in the YUV data to obtain the Y component (i.e., the Y-channel image) and the UV component (i.e., the UV-channel image). Among them, the hardware accelerator can be called through the API of the hardware accelerator. In addition, the format of each frame of image in the YUV data can also be converted to obtain each frame of RGB image.
[0051] In one embodiment, as Figure 3 shown, the hardware accelerator can be an image accelerator. After decoding, the image accelerator is used to process the video stream, so as to separate the Y component, the UV component, and obtain the RGB image.
[0052] S102: Obtain the complexity of the RGB image, and determine the super-resolution model corresponding to the Y component based on the complexity.
[0053] Optionally, obtaining the complexity of the RGB image includes: inputting the RGB image into an image complexity discriminator, and determining the complexity using the classification information output by the image complexity discriminator. Wherein, the RGB image is an image obtained by converting the YUV image format corresponding to the Y component to be processed, and the image complexity discriminator can be set in the super-resolution module. After obtaining the RGB image, input the RGB image into the image complexity discriminator.
[0054] Optionally, the complexity can be obtained based on the idea of ClassSR. The image complexity discriminator can extract features in the RGB image through its own convolutional layer and global pooling, and determine the classification of the RGB image according to the features. Wherein, the information output by the image complexity discriminator can be the complexity value of the RGB image, and the classification of the RGB image is determined based on this value. The classification includes simple, medium, and difficult.
[0055] Optionally, the category of the RGB image can be recognized by the image complexity discriminator, and the complexity of the image can be determined according to the category. For example, if the category of the image is a simple type such as sky or cloud, it is determined as simple. If the category of the image is a complex type such as face or building, it can be determined as complex.
[0056] Optionally, the multiple frames of RGB images corresponding to the video stream can be divided into multiple batches according to the sorting of the RGB images. Each batch includes multiple frames of RGB images. The image complexity discriminator is used to determine the complexity of the RGB images in each batch. It is also possible to input the RGB images into the image complexity discriminator batch by batch, and the discriminator outputs the complexity of each frame of RGB image.
[0057] Optionally, the image complexity discriminator includes multiple convolutional layers and activation functions. The obtaining of the classification information includes: using the convolutional layer to extract the features of the RGB image to generate a feature map, and obtaining the predicted value corresponding to the feature map; using the activation function to process the predicted value to obtain the classification information.
[0058] In one embodiment, the number of convolutional layers can be 4, and the activation function can be the LeakyReLU activation function, the ReLU activation function, and other functions that can be used for the complexity classification of RGB images.
[0059] Optionally, the predicted value can be a three-dimensional vector, where each value in the three-dimensional vector represents the probability or confidence of an RGB image corresponding to a certain complexity level. Specifically, the three-dimensional vector can be [t0, t1, t2], where t0 represents the probability that the category of the RGB image is simple, t1 represents the probability that the complexity category of the RGB image is medium, and t2 represents the probability that the complexity category of the RGB image is difficult. The RGB image input to the image complexity discriminator is a three-channel image. Features are extracted through a convolutional layer, and a feature map with 32 channels is output. A linear layer in the image complexity discriminator is used to convert the feature map into a 3D vector, and a loss function is used to calculate the loss of the 3D vector to increase the difference in probabilities (predicted values) corresponding to different categories in the 3D vector.
[0060] In one embodiment, the calculation formula for loss calculation is:
[0061] sum-re i =|t i,0 -t i,1 |+|t i,0 -t i,2 |+|t i,1 -t i,2 |;
[0062]
[0063] Wherein, t i,0 represents the probability that the i-th RGB image belongs to the simple category, t i,1 represents the probability that the i-th RGB image belongs to the medium category, and t i,2 represents the probability that the i-th RGB image belongs to the difficult category. sum-re i represents the sum of the absolute values of the differences between different probabilities, n is the number of RGB images, the value of m1 is the number of categories minus one, and loss1 is the loss value corresponding to the i-th RGB image.
[0064] And for each category, the calculation formula for loss calculation includes:
[0065]
[0066]
[0067] Wherein, m is the number of categories, sumj represents the j-th category, and Loss2 is the loss value corresponding to the i-th category.
[0068] Optionally, the complexity includes simple, medium, and complex. The proportionality coefficients of the super-resolution models corresponding to different complexities are different. The proportionality coefficient is used to adjust the usage ratio of the convolutional branch and the self-attention branch. After determining the complexity of the RGB image, the super-resolution model corresponding to the Y component of the RGB image is determined according to the complexity.
[0069] In one embodiment, in the super-resolution model corresponding to simple complexity, the proportionality coefficient can be 0.25; in the super-resolution model corresponding to medium complexity, the proportionality coefficient can be 0.3; in the super-resolution model corresponding to complex complexity, the proportionality coefficient can be 0.5. This proportionality coefficient determines the proportion of windows in the sub-network of the super-resolution model that will be processed using the self-attention mechanism. If the proportionality coefficient is set to 0.5, then only half of the windows will use the self-attention branch, while the other half will use a simpler convolutional operation. By adjusting the proportionality coefficient, the computational complexity and memory consumption of the super-resolution model can be flexibly controlled. When the proportionality coefficient is low, the computational amount and memory consumption of the model will also be reduced accordingly, thereby improving the inference speed and efficiency. Specifically, as Figure 4 shown, the super-resolution model can be divided into a simple model, a medium model, and a complex model. The simple model corresponds to simple complexity, the medium model corresponds to medium complexity, and the complex model corresponds to complex complexity. After obtaining the Y component (Y-channel image) and the RGB image, the Y-channel image and the RGB image are input into an image complexity discriminator. The complexity classification is obtained through the image complexity discriminator, and the corresponding super-resolution model is determined based on the complexity classification. The high-resolution image corresponding to the Y-channel image is output using the corresponding super-resolution model.
[0070] S103: Process the Y component using the super-resolution model to generate a first image, and generate a target video stream based on the first image and the UV components.
[0071] Optionally, the resolution of the first image is higher than that of the Y component.
[0072] Optionally, processing the Y component using the super-resolution model to generate a first image includes: splitting the Y component into image blocks of a preset size, obtaining the complex information corresponding to the image blocks, where the complex information includes a complex score and an index; determining the processing branch corresponding to each image block based on the complex information, and processing the image blocks using the processing branch to obtain an inference result, where the processing branch includes a convolutional branch and a self-attention branch; merging the inference results to obtain a merged result, and upsampling the merged result to obtain the first image.
[0073] In one embodiment, the super-resolution model can process the Y component based on the CAMixerSR idea. When dividing the Y component into blocks, the division can be performed through a specified window size, and the size of the obtained image blocks can be 8*8 pixels. The size of the Y component image can be 80*80 and other numerical values, and the size of the image blocks can be set according to the size of the Y component image to ensure that the number of obtained image blocks is an integer.
[0074] Optionally, obtain the complex information corresponding to the image blocks, including: inputting the RGB image including the image blocks into a preset convolutional layer to obtain eigenvalues; inputting the eigenvalues into a mask prediction network to obtain a complex probability distribution map, where the complex probability distribution map includes the complex score corresponding to each image block and the index corresponding to the complex score, and the size of the complex score corresponds to the complexity of the image block. Determine whether the image block is a simple image block or a complex image block based on the complex score and / or index.
[0075] Optionally, an image with a complex score greater than a predetermined threshold can be determined as a complex image block, and an image block with a complex score less than or equal to the predetermined threshold can be determined as a simple image block.
[0076] Optionally, as Figure 5 shown, the processing branch corresponding to the image block can be determined according to the index of the image block. Specifically, when it is determined that the image block is a simple image block according to the index, a convolution branch can be used to process the image block, and when it is determined that the image block is a complex image block according to the index, a self-attention branch can be used to process the image block.
[0077] In one embodiment, the preset convolutional layer can be a 1x1 convolutional layer. Through this convolutional layer, the eigenvalue v of each image block is obtained, and the eigenvalue v is input into a mask prediction network (predictor) to obtain a complexity probability distribution map Mask (the index of the complex image block can be denoted as idx1, and the index of the simple image block can be denoted as idx2) and a feature offset value offsets. The image blocks are processed according to the index in the distribution map.
[0078] Optionally, generate a target video stream based on the first image and the UV image, including: interpolating the UV component using an interpolation method to obtain a second image, where the resolution of the second image is the same as that of the first image; replacing the Y component in the second image with the first image to obtain a target image, and generating a target video stream based on the target image.
[0079] Optionally, the interpolation method can be bilinear interpolation, bicubic interpolation, and other interpolation methods that can improve the image resolution.
[0080] In one embodiment, the preset resolution can be 3840x2160 resolution. After determining that the resolution of the video stream is less than this resolution, the Y component and the UV component are separated from the image of the video stream. First, the graphics processing unit interpolates the image (UV component) in NV12 format with a resolution of 1280x720 to 3840x2160 resolution through traditional interpolation methods such as bilinear interpolation to obtain a second image; the Y-channel image (Y component) with a resolution of 1280x720 is fed into the super-resolution model for inference to obtain a first image with a resolution of 3840x2160. The first image is a Y-channel image; finally, the Y channel of the interpolated first image is replaced with the Y channel of this Y-channel image to obtain a target image, and the target images are arranged according to the order of the images in the original video stream to obtain a target video stream.
[0081] Optionally, after using the first image to replace the Y component in the second image to obtain a target image, an upsampling operation can also be performed on the target image. Specifically, the PixeShuffle operation can be used for upsampling. The upsampling operation is used to deconstruct and rearrange the pixels corresponding to the Y component in the target image.
[0082] Optionally, the upsampling operation can be completed on the GPU to improve the computing efficiency.
[0083] The resolution enhancement method of this application has the following advantages:
[0084] 1. Dynamic model selection: Introduce a sub-image complexity classification mechanism, select different super-resolution models according to the classification of RGB images, and balance performance and image quality. Combine the sub-image classification of ClassSR and the content-aware routing of CAMixerSR to achieve more accurate sub-image classification and more efficient content-aware routing, thereby improving the accuracy and efficiency of the super-resolution algorithm.
[0085] 2. Visual enhancement in weak network environments: Improve the performance and effect of the super-resolution algorithm through a scaling factor, enhance the visual experience of users in video conferencing in weak network environments, and still provide good image clarity after the source resolution drops.
[0086] 3. Hardware and algorithm co-optimization: For performance optimization at the edge (heterogeneous computing), the super-resolution model usually has an upsampling operation at the end. This operation needs to deconstruct and rearrange the channel data of the tensor. These frequent data rearrangements involve a large amount of memory access. However, the memory bandwidth at the edge is usually low, and directly performing this operation on the NPU (neural network processing unit) will significantly reduce performance. While the GPU is naturally suitable for processing parallel computing, especially pixel-level operations. Therefore, performing this upsampling operation with the GPU will be much more efficient. After the GPU finishes processing, it can be directly displayed, reducing copy operations.
[0087] Based on the same inventive concept, an embodiment of the present application provides an electronic device, such as Figure 6 shown, Figure 6 The electronic device 2000 shown in the figure includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are communicatively connected, such as connected through a bus 2002.
[0088] The processor 2001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0089] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0090] The memory 2003 can be a ROM (Read-Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (random access memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read-Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0091] Optionally, the electronic device 2000 may further include a communication unit 2004. The communication unit 2004 can be used for receiving and sending signals. The communication unit 2004 can allow the electronic device 2000 to communicate with other devices wirelessly or wiredly to exchange data. It should be noted that in practical applications, the communication unit 2004 is not limited to one.
[0092] Optionally, the electronic device 2000 may further include an input unit 2005. The input unit 2005 can be used for receiving input digital, character, image, and / or sound information, or generating key signal inputs related to the user settings and function controls of the electronic device 2000. The input unit 2005 can include, but is not limited to, one or more of a touch screen, a physical keyboard, function keys (such as volume control buttons, power on / off buttons, etc.), a trackball, a mouse, a joystick, a shooting device, a pickup, etc.
[0093] Optionally, the electronic device 2000 may further include an output unit 2006. The output unit 2006 can be used for outputting or displaying the information processed by the processor 2001. The output unit 2006 can include, but is not limited to, one or more of a display device, a speaker, a vibration device, etc.
[0094] Although the figure shows the electronic device 2000 with various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0095] Optionally, the memory 2003 is used to store a computer program for executing the solution of this application, and is controlled by the processor 2001 to execute. The processor 2001 is used to execute the computer program stored in the memory 2003 to implement the steps of any method provided in the embodiments of this application.
[0096] Based on the same inventive concept, the embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided in this application / implements the steps of various alternative embodiments of the method provided in this application.
[0097] Based on the same inventive concept, the embodiments of this application provide a computer program product, which includes a computer program. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided in this application / implements the steps of various alternative embodiments of the method provided in this application.
[0098] Those skilled in the art of this technology can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in this application can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the related technologies that are the same as those disclosed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0099] In the description of this application, the directions or positional relationships indicated by the words "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are the exemplary directions or positional relationships based on the drawings, which are for the convenience of describing or simplifying the embodiments of this application, rather than indicating or implying that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.
[0100] The terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.
[0101] In the description of the present application, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0102] In the description of this specification, specific features, structures, materials, or characteristics may be combined in any one or more embodiments or examples in a suitable manner.
[0103] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.
Claims
1. A resolution enhancement method, characterized in that: The method comprises: If it is determined that the resolution of the video stream is less than the preset value, obtaining the Y component, UV component and RGB image corresponding to the video stream; Acquire the complexity of the RGB image, and determine the super-resolution model corresponding to the Y component based on the complexity; The Y component is processed by using the super-resolution model to generate a first image, and a target video stream is generated based on the first image and the UV component, wherein the resolution of the first image is higher than that of the Y component.
2. The resolution enhancement method according to claim 1, characterized in that: The obtaining the complexity of the RGB image includes: The RGB image is input into an image complexity discriminator, and the complexity is determined using classification information output by the image complexity discriminator.
3. The resolution enhancement method according to claim 2, characterized in that: The image complexity discriminator includes multiple convolutional layers and activation functions, and the acquisition of classification information includes: Extracting features of the RGB image using the convolutional layer, generating a feature map, and obtaining a prediction value corresponding to the feature map; The predicted value is processed by using the activation function to obtain the classification information.
4. The resolution enhancement method according to claim 1, characterized in that: The step of using the super-resolution model to process the Y component to generate a first image includes: The Y component is divided into image blocks of a preset size, and complex information corresponding to the image blocks is obtained, wherein the complex information includes a complex score and an index; Determine a processing branch corresponding to each image block based on the complex information, and use the processing branch to process the image block to obtain an inference result, wherein the processing branch includes a convolution branch and a self-attention branch; The inference results are combined to obtain a combined result, and the combined result is upsampled to obtain a first image.
5. The resolution enhancement method according to claim 4, characterized in that: The complexity includes simple, medium, and complex. The proportional coefficients of the super-resolution models corresponding to different complexities are different. The proportional coefficients are used to adjust the usage ratio of the convolution branch and the self-attention branch.
6. The resolution enhancement method according to claim 1, characterized in that: The obtaining of complex information corresponding to the image block includes: Inputting the RGB image including the image block into a preset convolution layer to obtain a feature value; The characteristic value is input into the mask prediction network to obtain a complex probability distribution map, which includes a complex score corresponding to each image block and an index corresponding to the complex score, and the size of the complex score corresponds to the complexity of the image block.
7. The resolution enhancement method according to claim 1, characterized in that: The generating a target video stream based on the first image and the UV image includes: interpolating the UV component by using an interpolation method to obtain a second image, where the resolution of the second image is the same as that of the first image; The first image is used to replace the Y component in the second image to obtain a target image, and a target video stream is generated according to the target image.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.