Video super-resolution method and device, storage medium and electronic device

Through the cloud-coordinated video super-resolution method, appropriate super-resolution strategies are selected at the video receiving end and the cloud based on the image texture complexity, which solves the problem of degraded video quality in real-time audio and video communications in weak network environments, and achieves high-quality audio and video call fluency and improved user experience.

CN116156243BActive Publication Date: 2025-10-10IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211574918.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-10-10
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Network devices that frequently move under mobile network coverage are easily affected by weak network scenarios caused by cellular signals or building obstructions. In this case, existing real-time audio and video communications reduce resolution to ensure smoothness, resulting in reduced video quality and a poor user viewing experience.

Method used

A cloud-based collaborative video super-resolution method is adopted. The video frame blocks are processed collaboratively by the video receiving end and the cloud. The decision of whether to perform super-resolution on the terminal or the cloud is made based on the complexity of the image texture. The appropriate strategy selection and splicing of video frame blocks are achieved by combining lightweight and large super-resolution models.

Benefits of technology

It improves video quality, ensures the smoothness and sustainability of audio and video calls, enhances the user viewing experience, and reduces the computing load on the video receiving end.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116156243B_ABST
    Figure CN116156243B_ABST
Patent Text Reader

Abstract

The application provides a video super-resolution method and device, a storage medium and an electronic device, and relates to the technical field of information. The video super-resolution method comprises the following steps: receiving a plurality of video frame blocks corresponding to a target video sent by a video sending end; if the plurality of video frame blocks fail to meet a video super-resolution condition of a video receiving end, obtaining image texture complexity of each of the plurality of video frame blocks; if there are P video frame blocks with image texture complexity greater than a preset complexity threshold in the plurality of video frame blocks, obtaining a super-resolution result of each of the P video frame blocks from a cloud; and determining a super-resolution result of the target video based on the super-resolution result of each of the P video frame blocks. Through the scheme in the application, the definition of the target video when displayed on the video receiving end can be improved, and the calculation amount of the video receiving end when super-resolved can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to a video super-resolution method, device, storage medium and electronic device. Background Art

[0002] Real-time audio and video aims to achieve multi-party communication within a few hundred milliseconds, placing high demands on the network. For network devices that frequently move within mobile network coverage, they are susceptible to weak network conditions caused by cellular signals or building obstructions. In these situations, real-time audio and video will reduce resolution to increase transmission smoothness, which can degrade audio and video quality and the user's viewing experience. Summary of the Invention

[0003] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a video super-resolution method, apparatus, storage medium and electronic device.

[0004] In a first aspect, an embodiment of the present application provides a video super-resolution method, which is applied to a video receiving end. The video super-resolution method includes: receiving multiple video frame blocks corresponding to a target video sent by a video sending end; if the multiple video frame blocks fail to meet the video super-resolution conditions of the video receiving end, obtaining the image texture complexity of each of the multiple video frame blocks; if there are P video frame blocks in the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, obtaining the super-resolution results of each of the P video frame blocks from the cloud, where P is a positive integer; and determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks.

[0005] In combination with the first aspect, in certain implementations of the first aspect, the video super-resolution method further includes: if there are Q video frame blocks whose image texture complexity is less than or equal to a preset complexity threshold among multiple video frame blocks, then super-resolution is performed on the Q video frame blocks to obtain super-resolution results of each of the Q video frame blocks, where Q is a positive integer; wherein, based on the super-resolution results of each of the P video frame blocks, the super-resolution result of the target video is determined, including: based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks, the super-resolution result of the target video is determined.

[0006] In combination with the first aspect, in certain implementations of the first aspect, the super-resolution result of the target video is determined based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks, including: if it is determined based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks that the super-resolution P video frame blocks and the super-resolution Q video frame blocks have overlapping areas, then pixel data of the overlapping areas in the video frame blocks corresponding to the overlapping areas are determined; distance variables of the overlapping areas in the video frame blocks corresponding to the overlapping areas are determined respectively; and the super-resolution result of the target video is determined based on the pixel data of the overlapping areas in the video frame blocks corresponding to the overlapping areas and the distance variables of the overlapping areas in the video frame blocks corresponding to the overlapping areas.

[0007] In combination with the first aspect, in certain implementations of the first aspect, the super-resolution result of the target video is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area, and the distance variable of the overlapping area in the video frame block corresponding to the overlapping area, including: determining the splicing weight of the video frame block corresponding to the overlapping area in P video frame blocks; determining the splicing weight of the video frame block corresponding to the overlapping area in Q video frame blocks; determining the super-resolution result of the target video based on the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in the P video frame blocks, and the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in the Q video frame blocks.

[0008] In combination with the first aspect, in certain implementations of the first aspect, the video super-resolution method further includes: if multiple video frame blocks can meet the video super-resolution conditions of the video receiving end, then based on the current resolutions of the multiple video frame blocks, the multiple video frame blocks are super-resolved at a preset ratio.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the video super-resolution condition at the video receiving end includes that current resolutions of the plurality of video frame blocks are greater than or equal to a preset resolution threshold.

[0010] In the second aspect, an embodiment of the present application provides a video super-resolution method, which is applied to the cloud. The video super-resolution method includes: receiving multiple video frame blocks corresponding to the target video sent by the video sending end; if the multiple video frame blocks meet the video super-resolution conditions of the cloud, obtaining the image texture complexity of the multiple video frame blocks from the video sending end; if there are P video frame blocks with image texture complexity greater than a preset complexity threshold among the multiple video frame blocks, super-resolving the P video frame blocks to obtain super-resolution results of the P video frame blocks, where P is a positive integer; and sending the super-resolution results of the P video frame blocks to the video receiving end.

[0011] On the third aspect, an embodiment of the present application provides a video super-resolution method, which is applied to a video sending end. The video super-resolution method includes: when the resolution of the target video to be transmitted is reduced, dividing the target video into blocks to obtain multiple video frame blocks corresponding to the target video; determining the image texture complexity corresponding to each of the multiple video frame blocks; sending the multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks to the cloud and / or the video receiving end, so that the cloud and / or the video receiving end can super-resolution the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks.

[0012] In a fourth aspect, an embodiment of the present application provides a video super-resolution device, which is applied to a video receiving end, and the video super-resolution device includes: a receiving module, which is used to receive multiple video frame blocks corresponding to a target video sent by a video sending end; a first acquisition module, which is used to obtain the image texture complexity of each of the multiple video frame blocks if the multiple video frame blocks fail to meet the video super-resolution conditions of the video receiving end; a second acquisition module, which is used to obtain the super-resolution results of each of the P video frame blocks from the cloud if there are P video frame blocks with image texture complexity greater than a preset complexity threshold among the multiple video frame blocks, where P is a positive integer; and a determination module, which is used to determine the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks.

[0013] In a fifth aspect, an embodiment of the present application provides a video super-resolution device, which is applied to the cloud. The video super-resolution device includes: a receiving module, which is used to receive multiple video frame blocks corresponding to the target video sent by the video sending end; an acquisition module, which is used to obtain the image texture complexity of the multiple video frame blocks from the video sending end if the multiple video frame blocks meet the video super-resolution conditions of the cloud; a super-resolution module, which is used to super-resolve the P video frame blocks if there are P video frame blocks with image texture complexity greater than a preset complexity threshold among the multiple video frame blocks, and obtain super-resolution results of the P video frame blocks, where P is a positive integer; and a sending module, which is used to send the super-resolution results of the P video frame blocks to the video receiving end.

[0014] In the sixth aspect, an embodiment of the present application provides a video super-resolution device, which is applied to a video sending end. The video super-resolution device includes: a division module, which is used to divide the target video into blocks when the resolution of the target video to be transmitted is reduced, so as to obtain multiple video frame blocks corresponding to the target video; a determination module, which is used to determine the image texture complexity corresponding to each of the multiple video frame blocks; and a sending module, which is used to send multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks to the cloud and / or the video receiving end, so that the cloud and / or the video receiving end can super-resolution the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks.

[0015] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program for executing the methods described in the first to third aspects.

[0016] In an eighth aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; and the processor for executing the methods described in the first to third aspects.

[0017] The video super-resolution method provided in the embodiments of the present application has the following beneficial effects.

[0018] This application performs super-resolution on the target video, improves the quality of the target video, ensures the smoothness and sustainability of the audio and video calls corresponding to the target video, and also improves the user's viewing experience. In addition, this application introduces the cloud to super-resolution the target video. When the video frame block corresponding to the target video does not meet the video super-resolution conditions of the video receiving end, the super-resolution results of the video frame block with image texture complexity greater than the preset complexity threshold are obtained from the cloud based on the image texture complexity of the video frame block. That is, the overall super-resolution efficiency of the target video is improved through cloud-to-cloud collaboration, reducing the computational workload of the video receiving end. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 The figure shows a scenario diagram applicable to the video super-resolution method provided in an embodiment of the present application.

[0021] Figure 2 FIG2 is a flow chart of a video super-resolution method provided by an exemplary embodiment of the present application.

[0022] Figure 3 FIG2 is a flow chart of a video super-resolution method provided by another exemplary embodiment of the present application.

[0023] Figure 4 The above is a flowchart of a video super-resolution method provided by another exemplary embodiment of the present application.

[0024] Figure 5 FIG2 is a schematic diagram of dividing video frame blocks provided by an exemplary embodiment of the present application.

[0025] Figure 6Fig. 1 shows a structural schematic diagram of a video super-resolution device according to an example embodiment of the present application.

[0026] Figure 7 Fig. 2 shows a structural schematic diagram of a video super-resolution device according to another example embodiment of the present application.

[0027] Figure 8 Fig. 3 shows a structural schematic diagram of a video super-resolution device according to yet another example embodiment of the present application.

[0028] Figure 9 Fig. 4 shows a structural schematic diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative work fall within the scope of protection of the present application.

[0030] Application Overview

[0031] The existing real-time audio and video communication can perceive the approximate network code rate in an extremely weak network environment through a code rate estimation algorithm, so as to reduce the resolution of real-time audio and video, for example, reducing the video resolution from 720P to 180P, so as to enhance the fluency of corresponding real-time audio and video transmission. This method sacrifices part of the video quality to ensure the fluency and sustainability of the entire audio and video call, but the decline in video quality will lead to poor visual experience of users in many scenarios.

[0032] Based on the above situation, the 180P can be up-sampled to 360P (2 times super-resolution) and 720P (4 times super-resolution) through a super-resolution model to improve the visual effect of real-time audio and video. The super-resolution algorithm based on deep learning is generally a relatively large model, and the calculation amount of this model is large whether it is running on a terminal or in the cloud. How to load these large amounts of calculation is a difficult problem.

[0033] This application proposes a cloud-to-cloud collaboration approach to load the computational load of the super-resolution model. First, due to the general-purpose computing chip of the central processing unit, the computational load of the video receiving end is not very strong, so a lightweight 2x super-resolution model can be used, as well as a lightweight super-resolution model that supports processing simple image textures. At the same time, the cloud supports large-scale super-resolution models that process complex image textures. The terminal divides the video frames of real-time audio and video into blocks. Each video frame block decides whether to super-resolution based on the resolution and image texture complexity, whether to perform 2x super-resolution or 4x super-resolution, terminal super-resolution or cloud super-resolution, and finally enables each video frame block to choose its own suitable strategy and implement block splicing of video frames at the terminal.

[0034] Example scenarios

[0035] Figure 1 The figure shows a scene diagram in which the video super-resolution method provided in the embodiment of the present application is applicable. Figure 1 As shown, the application scenario mentioned in the embodiment of the present application includes a video receiving end 11, a cloud 12, and a video sending end 13. In addition, the video receiving end 11, the cloud 12, and the video sending end 13 are connected in communication.

[0036] Specifically, when there is a video transmission request between the video sending end 13 and the video receiving end 11, if the video sending end 13 detects that the resolution of the target video to be transmitted is lower than the resolution at the previous moment, the video sending end 13 will cut the target video to be transmitted into blocks to obtain multiple video frame blocks corresponding to the target video to be transmitted, and perform image texture complexity detection on the multiple video frame blocks to obtain the image texture complexity corresponding to each video frame block.

[0037] Furthermore, the video transmitting end 13 sends multiple video frame blocks and their corresponding image texture complexities to the video receiving end 11 and the cloud 12. First, if the resolution of the video frame block meets the super-resolution condition of the video receiving end 11, the video receiving end directly super-resolutions the multiple video frame blocks, and splices and renders the super-resolution results; if the resolution of the video frame block does not meet the super-resolution condition of the video receiving end 11, then based on the image texture complexity, the cloud super-resolutions the video frame blocks whose image texture complexity is greater than a preset complexity threshold, and sends the super-resolution results to the video receiving end 11. At the same time, the video receiving end 11 super-resolutions the video frame blocks whose image texture complexity is less than or equal to the preset complexity threshold, obtains the super-resolution results, and splices and renders the super-resolution results sent by the cloud and the local super-resolution results.

[0038] Exemplary Methods

[0039] Figure 2FIG. 1 is a flow chart of a video super-resolution method provided by an exemplary embodiment of the present application. For example, the method is applied to a video receiving end. Figure 2 As shown, the video super-resolution method provided in the embodiment of the present application includes the following steps.

[0040] Step S210: receiving a plurality of video frame blocks corresponding to a target video sent by a video sending end.

[0041] For example, the video transmitter can be a mobile phone, tablet computer, desktop computer, or other terminal device with video transmission function. The target video refers to the video transmitted and communicated between the video receiver and the video transmitter, which can be a conference video or a WeChat video call.

[0042] Step S220 : If the plurality of video frame blocks fail to meet the video super-resolution condition of the video receiving end, the image texture complexity of each of the plurality of video frame blocks is obtained.

[0043] Exemplarily, the video super-resolution condition at the video receiving end includes that the current resolution of the plurality of video frame blocks is greater than or equal to a preset resolution threshold. For example, the preset resolution threshold is 360P. Exemplarily, the preset complexity threshold is 0.1. Furthermore, if the current resolution of the plurality of video frame blocks is 360P, and the image complexity of three of the plurality of video frame blocks is 0.3, then the super-resolution results of each of these three video frame blocks are obtained from the cloud.

[0044] In step S230 , if there are P video frame blocks among the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, super-resolution results of each of the P video frame blocks are obtained from the cloud, where P is a positive integer.

[0045] Step S240 : determining a super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks.

[0046] For example, if the image texture complexity of multiple video frame blocks is greater than a preset complexity threshold, the super-resolution results of each of the multiple video frame blocks are obtained from the cloud. If the image texture complexity of multiple video frame blocks is less than or equal to the preset complexity threshold, the video receiving end directly super-resolutions the multiple video frame blocks and determines the super-resolution results of each of the multiple video frame blocks.

[0047] Furthermore, for the above two situations, if there is no overlapping area between the multiple video frame blocks, the multiple super-resolved video frame blocks can be directly spliced ​​according to their positions based on the super-resolved results of the multiple video frame blocks. If there is overlapping area between the multiple video frame blocks, the pixel data of the overlapping area in the video frame block corresponding to the overlapping area is determined; the distance variable of the overlapping area in the video frame block corresponding to the overlapping area is determined respectively; and the super-resolved result of the target video is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and the distance variable.

[0048] Exemplarily, for two upper and lower video frame blocks with an overlapping area, the pixel value Value_a and the edge distance Dis_a of pixel point A in the overlapping area of ​​the upper video frame block are determined, and the pixel value Value_b and the edge distance Dis_b of pixel point B in the overlapping area of ​​the lower video frame block are determined. The edge distance is the distance between the pixel point and the bottom edge of the video frame block, with the distance to the edge being 1 and increasing in sequence. Exemplarily, if the upper and lower video frame blocks contain 8 pixels in the vertical overlapping area, then the pixel value Value_c of the overlapping area is = Dis_a / 8*a + Dis_b / 8*b.

[0049] This application performs super-resolution on the target video, improves the quality of the target video, ensures the smoothness and sustainability of the audio and video calls corresponding to the target video, and also improves the user's viewing experience. In addition, this application introduces the cloud to super-resolution the target video. When the video frame block corresponding to the target video does not meet the video super-resolution conditions of the video receiving end, the super-resolution results of the video frame block with image texture complexity greater than the preset complexity threshold are obtained from the cloud based on the image texture complexity of the video frame block. That is, the overall super-resolution efficiency of the target video is improved through cloud-to-cloud collaboration, reducing the computational workload of the video receiving end.

[0050] In one embodiment of the present application, the video super-resolution method also includes: if there are Q video frame blocks whose image texture complexity is less than or equal to a preset complexity threshold among multiple video frame blocks, then super-resolution is performed on the Q video frame blocks to obtain super-resolution results for each of the Q video frame blocks, where Q is a positive integer.

[0051] Specifically, when there are video frame blocks with image texture complexity greater than a preset complexity threshold and video frame blocks with image texture complexity less than or equal to the preset complexity threshold among multiple video frame blocks, the super-resolution results of the video frame blocks with image texture complexity greater than the preset complexity threshold are obtained from the cloud, and the video receiving end super-resolutions the video frame blocks with image texture complexity less than or equal to the preset complexity threshold to obtain the super-resolution results.

[0052] In this embodiment, determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks includes: determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks.

[0053] Specifically, if, based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks, it is determined that the super-resolved P video frame blocks and the super-resolved Q video frame blocks have overlapping areas, then the pixel data of the overlapping area in the video frame block corresponding to the overlapping area is determined; the distance variables of the overlapping area in the video frame block corresponding to the overlapping area are determined respectively; and the super-resolved result of the target video is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and the distance variables of the overlapping area in the video frame block corresponding to the overlapping area.

[0054] Exemplarily, the two video frame blocks with overlapping areas in the P super-resolved video frame blocks and the Q super-resolved video frame blocks are respectively recorded as video frame block p and video frame block q. Further, as described in step S240, the pixel value Value_a of pixel point A in the overlapping area of ​​video frame block p and the edge distance Dis_a of pixel point A are determined, and the pixel value Value_b of pixel point B in the overlapping area of ​​video frame block q and the edge distance Dis_b of pixel point B are determined. Exemplarily, if video frame block p and video frame block q contain g pixels in the vertical direction, then the pixel value Value_c of the overlapping area = Dis_a / g*a+Dis_b / g*b. Further, based on the pixel values ​​of the overlapping area and the super-resolved results of the non-overlapping area, the super-resolved result of the target video is determined.

[0055] In an embodiment of the present application, a fusion calculation is performed on the pixel values ​​of the overlapping areas in the P video frame blocks and the Q video frame blocks after super-resolution, which is beneficial to the smooth transition of the overlapping areas when the video frame blocks are spliced, thereby improving the viewing experience of the target video.

[0056] In one embodiment of the present application, a super-resolution result of a target video is determined based on pixel data of the overlapping area in the video frame block corresponding to the overlapping area, and a distance variable of the overlapping area in the video frame block corresponding to the overlapping area, including: determining the splicing weight of the video frame block corresponding to the overlapping area in P video frame blocks; determining the splicing weight of the video frame block corresponding to the overlapping area in Q video frame blocks; and determining the super-resolution result of the target video based on the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in the P video frame blocks, and the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in the Q video frame blocks.

[0057] For example, the splicing weight of the video frame blocks in the P video frame blocks is x, and the splicing weight of the video frame blocks in the Q video frame blocks is y. Continuing with the above example, the pixel value in the overlapping area is Value_c = Dis_a / g*a*x + Dis_b / g*b*y. For example, the value of x is greater than the value of y, for example, x is equal to 0.7, and y is equal to 0.3.

[0058] In the embodiment of the present application, because the super-resolution of video frame blocks at the cloud side is more complicated, the super-resolution of video frame blocks at the video receiving end is relatively simpler than that at the cloud side. When fusing the pixel values ​​of the pixel points in the overlapping area, the super-resolution results at the cloud side are given priority as much as possible, which can make the splicing boundaries of the overlapping area smoother.

[0059] In one embodiment of the present application, if the plurality of video frame blocks can meet the video super-resolution condition of the video receiving end, the plurality of video frame blocks are super-resolved at a preset ratio based on the current resolutions of the plurality of video frame blocks.

[0060] Exemplarily, the video super-resolution condition at the video receiving end includes that the current resolution of the plurality of video frame blocks is greater than or equal to a preset resolution threshold, for example, the preset resolution threshold is 360P.

[0061] For example, if the resolution of the video frame blocks received by the video receiving end is 720P, the video receiving end does not perform super-resolution processing on the video frame blocks, but directly splices and renders the video frame blocks. If the resolution of the video frame blocks received by the video receiving end is 360P, the video receiving end will perform double super-resolution on the video frame blocks to 720P, and then splice and render the super-resolved video frame blocks.

[0062] In an embodiment of the present application, the video receiving end independently performs lightweight super-resolution on some video frame blocks, which can speed up the display response speed of the target video and make splicing simpler and faster.

[0063] Figure 3 FIG. 1 is a flow chart of a video super-resolution method provided by another exemplary embodiment of the present application. For example, the method is applied to the cloud. Figure 3 As shown, the video super-resolution method provided in the embodiment of the present application includes the following steps.

[0064] Step S310: receiving a plurality of video frame blocks corresponding to a target video sent by a video sending end.

[0065] Step S320: If the plurality of video frame blocks meet the video super-resolution condition of the cloud, the image texture complexity of each of the plurality of video frame blocks is obtained from the video sending end.

[0066] Step S330 : If there are P video frame blocks among the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, super-resolution is performed on the P video frame blocks to obtain super-resolution results of the P video frame blocks, where P is a positive integer.

[0067] Step S340: sending the super-resolution results of the P video frame blocks to the video receiving end.

[0068] Exemplarily, the video super-resolution condition on the cloud includes that the current resolution of multiple video frame blocks is less than a preset resolution threshold. Exemplarily, the preset resolution threshold is 360P, and the preset complexity threshold is 0.1. If the resolution of the multiple video frame blocks received by the cloud is 180P, then the multiple video frame blocks meet the video super-resolution condition on the cloud, and if there are P video frame blocks among the multiple video frame blocks whose image complexity is greater than 0.1, the P video frame blocks can be super-resolved at a preset ratio with the help of the super computing power of the cloud. Exemplarily, the P video frame blocks are super-resolved by 2 times to become video frame blocks with a resolution of 360P, or by 4 times to become video frame blocks with a resolution of 720P.

[0069] In the embodiment of the present application, when the cloud's video super-resolution conditions are met, the cloud's super-computing power is used to carry the super-resolution calculation workload of the entire target video, so that the overall super-resolution increases the concurrent capacity through cloud collaboration. In addition, super-resolution of the target video can also improve the video quality of the target video.

[0070] Figure 4 The above is a flow chart of a video super-resolution method provided by another exemplary embodiment of the present application. For example, the method is applied to a video sending end. Figure 4 As shown, the video super-resolution method provided in the embodiment of the present application includes the following steps.

[0071] Step S410 : When the resolution of the target video to be transmitted is reduced, the target video is divided into blocks to obtain a plurality of video frame blocks corresponding to the target video.

[0072] Specifically, in the event of mobile network switching, building obstruction, or backbone network congestion, the target video to be transmitted at the video transmitter may become stuck. In these scenarios, to maintain the smooth transmission of the target video to be transmitted, a resolution adaptive mechanism is generally adopted. For example, when the video transmitter detects a decrease in network bandwidth through a bitrate estimation algorithm, it will correspondingly reduce the resolution of the target video to be transmitted, thereby ensuring that the bitrate can be effectively reduced to within the target predicted bandwidth in a timely manner to ensure the continuity of the target video to be transmitted.

[0073] Table 1 Correspondence between resolution and bit rate

[0074] Bitrate Resolution 300kbps 320*180 600kbps 640*360 1500kbps 1280*720

[0075] Once the resolution of the target video to be transmitted is reduced as described above, for example, the resolution of the target video to be transmitted is reduced from 1280*720 to 320*180, the video sending end divides the target video to be transmitted into blocks to obtain multiple video frame blocks corresponding to the target video.

[0076] Figure 5 The figure shows a schematic diagram of video frame block division provided by an exemplary embodiment of the present application. For example, a target video to be transmitted with a resolution of 320*180 can be divided into 64*64 video frame blocks. For example, the target video to be transmitted can be divided into 5 video frame blocks horizontally and 3 video frame blocks vertically, with each video frame block overlapping by 8 pixels.

[0077] Step S420: determining the image texture complexity corresponding to each of the plurality of video frame blocks.

[0078] Specifically, the plurality of video frame blocks are numbered. For example, the numbering method is Block ID. If there are 15 video frame blocks, the IDs are 0-14.

[0079] Furthermore, texture complexity is tested for each video frame block. For example, the Canny operator is used to detect image edges, ultimately generating an image grayscale edge map. Image texture complexity is determined based on the proportion of edge grayscale in the entire image.

[0080] Exemplarily, image texture complexity s=number of edge grayscale pixels / number of pixels in the entire image.

[0081] Step S430 : Sending multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks to the cloud and / or the video receiving end.

[0082] The purpose of step S430 is to enable the cloud and / or the video receiving end to perform super-resolution on the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks.

[0083] For example, the image texture complexity a is carried in the transmission protocol. Streaming media transmission generally uses the Real-time Transport Protocol (RTP). The image texture complexity of video frame blocks Block 0 to Block 14 can be carried with the help of the RTP extension protocol.

[0084] For example, considering 8-bit expansion, the texture complexity percentage uses a 128-point scale, which accounts for a total of 7 bits. For example, if a = 0.1, then the value is 13 (rounded up). The first bit is 1 for a valid value, and the first bit is 0 for an invalid value. The expansion value uses 8 bits, a total of 15 blocks, arranged in 15*8 bits at a time, and then aligned with 32 bits. Finally, 8 bits are padded, with the first bit being 0.

[0085] In an embodiment of the present application, the video transmitter uniformly performs fast segmentation and image texture complexity detection on the target video to be transmitted, reducing the amount of repeated calculations on the cloud and video receiver. By dividing the target video to be transmitted into blocks, a more refined segmentation of the target video to be transmitted is achieved, so that the video frame blocks are matched with a more optimal super-resolution method. In addition, performing image texture complexity detection on each video frame block can further match the video frame block with a more optimal super-resolution path.

[0086] Exemplary devices

[0087] Combined with the above Figures 2 to 5 , describes the method embodiment of the present application in detail, and the following is combined with Figures 6 to 8 , the device embodiment of the present application is described in detail. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment, so for parts not described in detail, reference can be made to the previous method embodiment.

[0088] Figure 6 FIG. 1 is a schematic diagram of the structure of a video super-resolution device provided by an exemplary embodiment of the present application. For example, the device is applied to a video receiving end, such as Figure 6 As shown, the video super-resolution device 60 provided in the embodiment of the present application includes:

[0089] The receiving module 610 is configured to receive a plurality of video frame blocks corresponding to a target video sent by a video transmitting end;

[0090] A first acquisition module 620 is configured to acquire image texture complexity of each of the plurality of video frame blocks if the plurality of video frame blocks fail to meet the video super-resolution condition of the video receiving end;

[0091] A second acquisition module 630 is configured to obtain super-resolution results of each of the P video frame blocks from the cloud if there are P video frame blocks with image texture complexity greater than a preset complexity threshold among the multiple video frame blocks, where P is a positive integer;

[0092] The determination module 640 is configured to determine a super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks.

[0093] In one embodiment of the present application, the determination module 640 is further used to, if there are Q video frame blocks whose image texture complexity is less than or equal to a preset complexity threshold among the multiple video frame blocks, perform super-resolution on the Q video frame blocks to obtain super-resolution results of each of the Q video frame blocks, where Q is a positive integer; wherein, based on the super-resolution results of each of the P video frame blocks, the super-resolution result of the target video is determined, including: determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks.

[0094] In one embodiment of the present application, the determination module 640 is further used to, if it is determined based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks that the super-resolved P video frame blocks and the super-resolved Q video frame blocks have overlapping areas, then determine the pixel data of the overlapping areas in the video frame blocks corresponding to the overlapping areas; respectively determine the distance variables of the overlapping areas in the video frame blocks corresponding to the overlapping areas; and determine the super-resolved result of the target video based on the pixel data of the overlapping areas in the video frame blocks corresponding to the overlapping areas and the distance variables of the overlapping areas in the video frame blocks corresponding to the overlapping areas.

[0095] In one embodiment of the present application, the determination module 640 is also used to determine the super-resolution result of the target video based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area, and the distance variable of the overlapping area in the video frame block corresponding to the overlapping area, including: determining the splicing weight of the video frame block corresponding to the overlapping area in P video frame blocks; determining the splicing weight of the video frame block corresponding to the overlapping area in Q video frame blocks; determining the super-resolution result of the target video based on the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in P video frame blocks, and the pixel data, distance variable and splicing weight of the video frame block corresponding to the overlapping area in Q video frame blocks.

[0096] In one embodiment of the present application, the determination module 640 is further configured to super-resolution the multiple video frame blocks at a preset ratio based on the current resolutions of the multiple video frame blocks if the multiple video frame blocks can meet the video super-resolution conditions of the video receiving end.

[0097] In one embodiment of the present application, the video super-resolution condition of the video receiving end includes that the current resolution of the plurality of video frame blocks is greater than or equal to a preset resolution threshold.

[0098] Figure 7 FIG. 1 is a schematic diagram of the structure of a video super-resolution device provided by another exemplary embodiment of the present application. For example, the device is applied to the cloud, such as Figure 7 As shown, the video super-resolution device 70 provided in the embodiment of the present application includes:

[0099] The receiving module 710 is configured to receive a plurality of video frame blocks corresponding to a target video sent by a video transmitting end;

[0100] An acquisition module 720 is configured to acquire the image texture complexity of each of the plurality of video frame blocks from the video transmitting end if the plurality of video frame blocks meet the video super-resolution condition of the cloud;

[0101] a super-resolution module 730 configured to, if there are P video frame blocks among the plurality of video frame blocks whose image texture complexity is greater than a preset complexity threshold, super-resolve the P video frame blocks to obtain super-resolved results for each of the P video frame blocks, where P is a positive integer;

[0102] The sending module 740 is configured to send the super-resolution results of the P video frame blocks to a video receiving end.

[0103] Figure 8 FIG. 1 is a schematic diagram of the structure of a video super-resolution device provided by another exemplary embodiment of the present application. For example, the device is applied to a video sending end, such as Figure 8 As shown, the video super-resolution device 80 provided in the embodiment of the present application includes:

[0104] A division module 810 is configured to divide the target video into blocks when the resolution of the target video to be transmitted is reduced, to obtain a plurality of video frame blocks corresponding to the target video;

[0105] A determination module 820 is configured to determine the image texture complexity corresponding to each of the plurality of video frame blocks;

[0106] The sending module 830 is used to send multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks to the cloud and / or the video receiving end, so that the cloud and / or the video receiving end can super-resolve the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks.

[0107] Below, reference Figure 9 To describe the electronic device according to the embodiment of the present application. Figure 9 Shown is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.

[0108] like Figure 9 As shown, the electronic device 90 includes one or more processors 901 and a memory 902 .

[0109] The processor 901 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.

[0110] The memory 902 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 901 may execute the program instructions to implement the methods of the various embodiments of the present application described above and / or other desired functions. Various contents such as multiple video frame blocks, image texture complexity thresholds, super-resolution results, and the current resolution of video frame blocks may also be stored in the computer-readable storage medium.

[0111] In one example, the electronic device 90 may further include an input device 903 and an output device 904 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0112] The input device 903 may include, for example, a keyboard, a mouse, and the like.

[0113] The output device 904 can output various information to the outside, including multiple video frame blocks, image texture complexity thresholds, super-resolution results, current resolution of video frame blocks, etc. The output device 904 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0114] Of course, to simplify, Figure 9 Only some of the components related to the present application in the electronic device 90 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 90 may further include any other appropriate components according to specific application scenarios.

[0115] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present application described above in this specification.

[0116] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. The embodiments of the present application are not limited by the

[0117] In addition, an embodiment of the present application can also be a computer readable storage medium, which stores computer program instructions, and when the computer program instructions are run on a processor, the processor executes the steps of the methods described above according to various embodiments of the present application.

[0118] The computer readable storage medium can be any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0119] The above describes the basic principles of the present application in combination with specific embodiments, but it should be noted that the advantages, advantages, effects and the like mentioned in the present application are only examples and are not limiting, and these advantages, advantages, effects and the like cannot be considered as the must-have of each embodiment of the present application. In addition, the above specific details are only for the purpose of example and understanding, and are not limiting, and the above details do not limit the present application to the must-use specific details to realize.

[0120] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0121] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0122] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0123] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A video super-resolution method, characterized in that: Applied to a video receiving end, the method includes: receiving a plurality of video frame blocks corresponding to a target video sent by a video transmitter; If the multiple video frame blocks fail to meet the video super-resolution condition of the video receiving end, obtaining the image texture complexity of each of the multiple video frame blocks; If there are P video frame blocks among the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, obtaining super-resolution results of each of the P video frame blocks from the cloud, where P is a positive integer; Determining a super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks; Among them, if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data and distance variable in the video frame block corresponding to the overlapping area, and the super-resolved result of the target video is determined based on the pixel value of the overlapping area and the super-resolved result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom of the video frame block.

2. The method according to claim 1, characterized in that Also includes: If there are Q video frame blocks whose image texture complexity is less than or equal to the preset complexity threshold among the multiple video frame blocks, super-resolution is performed on the Q video frame blocks to obtain super-resolution results of each of the Q video frame blocks, where Q is a positive integer; The step of determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks includes: Based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks, a super-resolution result of the target video is determined.

3. The method according to claim 2, characterized in that The determining the super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks includes: If it is determined based on the super-resolution results of each of the P video frame blocks and the super-resolution results of each of the Q video frame blocks that the super-resolved P video frame blocks and the super-resolved Q video frame blocks have an overlapping area, then determining pixel data of the video frame block corresponding to the overlapping area; respectively determining distance variables of the overlapping areas in the video frame blocks corresponding to the overlapping areas; The super-resolution result of the target video is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and the distance variable of the overlapping area in the video frame block corresponding to the overlapping area.

4. The method according to claim 3, characterized in that The determining the super-resolution result of the target video based on pixel data of the overlapping area in the video frame block corresponding to the overlapping area and a distance variable of the overlapping area in the video frame block corresponding to the overlapping area includes: Determining a splicing weight of the video frame blocks corresponding to the overlapping area in the P video frame blocks; Determining a splicing weight of the video frame blocks corresponding to the overlapping area in the Q video frame blocks; Based on the pixel data, distance variables and splicing weights of the video frame blocks corresponding to the overlapping areas in the P video frame blocks, and the pixel data, distance variables and splicing weights of the video frame blocks corresponding to the overlapping areas in the Q video frame blocks, the super-resolution result of the target video is determined.

5. The method according to any one of claims 1 to 4, characterized in that Also includes: If the multiple video frame blocks can meet the video super-resolution condition of the video receiving end, super-resolution of the multiple video frame blocks at a preset ratio is performed on the multiple video frame blocks based on current resolutions of the multiple video frame blocks.

6. The method according to any one of claims 1 to 4, characterized in that The video super-resolution condition of the video receiving end includes that the current resolution of the plurality of video frame blocks is greater than or equal to a preset resolution threshold.

7. A video super-resolution method, characterized in that: Applied to the cloud, the method includes: receiving a plurality of video frame blocks corresponding to a target video sent by a video transmitter; If the plurality of video frame blocks meet the video super-resolution condition of the cloud, obtaining the image texture complexity of each of the plurality of video frame blocks from the video sending end; If there are P video frame blocks among the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, super-resolving the P video frame blocks to obtain super-resolving results of each of the P video frame blocks, where P is a positive integer; The super-resolution results of each of the P video frame blocks are sent to a video receiving end. After the video receiving end receives the super-resolution results of each of the P video frame blocks, if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and a distance variable, and the super-resolution result of the target video is determined based on the pixel value of the overlapping area and the super-resolution result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom of the video frame block.

8. A video super-resolution method, characterized in that: Applied to a video sending end, the method includes: When the resolution of the target video to be transmitted is reduced, dividing the target video into blocks to obtain a plurality of video frame blocks corresponding to the target video; Determining the image texture complexity corresponding to each of the plurality of video frame blocks; The multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks are sent to the cloud and / or the video receiving end, so that the cloud and / or the video receiving end can super-resolve the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks. After the cloud and / or the video receiving end super-resolves the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks, if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and the distance variable, and the super-resolved result of the target video is determined based on the pixel value of the overlapping area and the super-resolved result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom of the video frame block.

9. A video super-resolution device, characterized in that: Applied to a video receiving end, the device comprises: A receiving module, configured to receive a plurality of video frame blocks corresponding to a target video sent by a video sending end; A first acquisition module is configured to acquire the image texture complexity of each of the plurality of video frame blocks if the plurality of video frame blocks fail to meet the video super-resolution condition of the video receiving end; A second acquisition module is configured to obtain, from the cloud, super-resolution results of each of the P video frame blocks if there are P video frame blocks among the multiple video frame blocks whose image texture complexity is greater than a preset complexity threshold, where P is a positive integer; A determination module is configured to determine a super-resolution result of the target video based on the super-resolution results of each of the P video frame blocks, wherein if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and a distance variable, and the super-resolution result of the target video is determined based on the pixel value of the overlapping area and the super-resolution result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom edge of the video frame block.

10. A video super-resolution device, characterized in that: Applied to the cloud, the device includes: A receiving module, configured to receive a plurality of video frame blocks corresponding to a target video sent by a video sending end; an acquisition module, configured to acquire, from the video transmitting end, the image texture complexity of each of the plurality of video frame blocks if the plurality of video frame blocks meet the video super-resolution condition of the cloud; a super-resolution module configured to, if there are P video frame blocks among the plurality of video frame blocks whose image texture complexity is greater than a preset complexity threshold, super-resolve the P video frame blocks to obtain super-resolved results for each of the P video frame blocks, where P is a positive integer; A sending module is used to send the super-resolution results of each of the P video frame blocks to a video receiving end, wherein after the video receiving end receives the super-resolution results of each of the P video frame blocks, if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data of the overlapping area in the video frame block corresponding to the overlapping area and a distance variable, and the super-resolution result of the target video is determined based on the pixel value of the overlapping area and the super-resolution result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom of the video frame block.

11. A video super-resolution device, characterized in that: Applied to a video sending end, the device includes: a dividing module, configured to divide the target video to be transmitted into blocks when the resolution of the target video is reduced, to obtain a plurality of video frame blocks corresponding to the target video; A determination module, configured to determine the image texture complexity corresponding to each of the plurality of video frame blocks; A sending module is used to send the multiple video frame blocks and the image texture complexity of each of the multiple video frame blocks to the cloud and / or the video receiving end, so that the cloud and / or the video receiving end can super-resolve the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks. After the cloud and / or the video receiving end super-resolves the multiple video frame blocks based on the image texture complexity of each of the multiple video frame blocks, if there is an overlapping area between the super-resolved video frame blocks, the pixel value of the overlapping area is determined based on the pixel data and distance variable in the video frame block corresponding to the overlapping area, and the super-resolved result of the target video is determined based on the pixel value of the overlapping area and the super-resolved result of the non-overlapping area; the distance variable includes the distance between the pixel point and the bottom of the video frame block.

12. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 8.

13. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Array image fusion method and device, medium and equipment

    CN113808059A

  • Cloud collaboration image super-resolution analysis method and system based on edge density

    CN114140326A

  • Super-resolution method and apparatus, terminal device, and storage medium

    WO2022160980A1