Video super-division method, electronic equipment and storage medium
By using dual cameras in tandem, the area to be super-resolution is determined and super-resolution processing is performed, which solves the problem of low image clarity in wide-angle camera videos and achieves higher image quality and more efficient processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-10
AI Technical Summary
In existing technologies, after super-resolution reconstruction of video footage captured by a wide-angle camera, the target clarity is not high and the processing efficiency is low, especially when super-resolution reconstruction is performed on the entire image.
The method of dual-camera linkage is adopted. The first camera is used to determine the shooting target and control the second camera to shoot. Based on the position of the shooting target in the second camera's image, the region to be super-resolution is determined and super-resolution processing is performed. Super-resolution processing is only performed on the region to be super-resolution, without processing the entire video frame.
It improves the image quality and clarity of the target, which is higher than that of a single-camera solution, and also improves the efficiency of super-resolution processing.
Smart Images

Figure CN121639474A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a video super-resolution method, an electronic device and a storage medium. BACKGROUND
[0002] Super-resolution reconstruction can improve image resolution, and can improve recognition ability and recognition accuracy in many computer vision tasks, such as image segmentation, target detection and the like. In the related art, the super-resolution reconstruction is performed on a video picture of a camera, so that the clarity of a target can be improved. However, the video picture is usually a wide-angle shot, and even if the super-resolution reconstruction is performed on the video picture, the clarity of the target is still not high. Moreover, the super-resolution reconstruction is performed on the entire image, and the processing efficiency is low. SUMMARY
[0003] Therefore, the present application provides a video super-resolution method, a program product, an electronic device and a storage medium.
[0004] The technical scheme of the present application is implemented as follows:
[0005] In one aspect, the present application provides a video super-resolution method, which comprises the following steps:
[0006] obtaining a first video picture shot by a first camera;
[0007] determining a shooting target based on the first video picture;
[0008] controlling a second camera to shoot the shooting target to obtain a second video picture containing the shooting target;
[0009] determining a super-resolution region to be super-resolved based on a position of the shooting target in the second video picture;
[0010] performing super-resolution processing on the super-resolution region to be super-resolved.
[0011] In the above scheme, after the shooting target is determined based on the first video picture, the following step is further included:
[0012] determining a rotation angle of the second camera according to a position of the shooting target in the first video picture;
[0013] The step of controlling the second camera to shoot the shooting target to obtain a second video picture containing the shooting target comprises the following step:
[0014] controlling the second camera to rotate to a specified angle based on the rotation angle, so that the shooting target is located in a central region of the second video picture.
[0015] In the above scheme, before the region to be super-resolution is determined based on the position of the shooting target in the second video frame, the method further comprises:
[0016] The position of the shooting target in the second video frame is determined.
[0017] In the above scheme, the position of the shooting target in the second video frame is determined by:
[0018] A target region in the first video frame is determined, which is equivalent to the second video frame, and the shooting target is in the target region;
[0019] The position of the shooting target in the second video frame is determined based on the position of the shooting target in the target region.
[0020] In the above scheme, the target region in the first video frame is determined by:
[0021] An initial region in the first video frame is determined, which is equivalent to the second video frame;
[0022] The initial region is expanded to obtain an expanded region, and the initial region is in the expanded region;
[0023] A target region in the expanded region is determined, which is most matched with the second video frame.
[0024] In the above scheme, the target region in the expanded region is determined by:
[0025] The target region in the expanded region is determined based on the second video frame by a matching algorithm.
[0026] In the above scheme, the target region in the expanded region is determined based on the second video frame by a matching algorithm, comprising:
[0027] The second video frame is used as a template to perform template matching in the expanded region;
[0028] The region in the expanded region with the highest matching degree with the template is determined as the target region.
[0029] In the above scheme, the target region in the expanded region is determined based on the second video frame by a matching algorithm, comprising:
[0030] Key points in the second video frame and the expanded region are determined.
[0031] The region in the extended region with the largest number of matched key points with the second video picture is determined as the target region through key point matching.
[0032] In the above scheme, the to-be-super-resolution region is determined based on the position of the shooting target in the second video picture, and the method comprises the steps of:
[0033] The position of the human head frame corresponding to the shooting target is determined based on the position of the shooting target in the second video picture, and the position of the human head frame is taken as the to-be-super-resolution region.
[0034] In the above scheme, the super-resolution processing of the to-be-super-resolution region comprises the steps of:
[0035] The position of the to-be-super-resolution region corresponding to the current video frame of the second video picture is determined.
[0036] The super-resolution region position of the previous video frame of the second video picture is obtained.
[0037] The super-resolution strategy is determined based on the distance between the to-be-super-resolution region position and the super-resolution region position.
[0038] The to-be-super-resolution region is processed based on the super-resolution strategy.
[0039] In the above scheme, the super-resolution strategy is determined based on the distance between the to-be-super-resolution region position and the super-resolution region position, and the method comprises the steps of:
[0040] If the distance is greater than a threshold value, the super-resolution reconstruction is performed based on the current video frame;
[0041] If the distance is less than or equal to a threshold value, the super-resolution reconstruction is performed based on the previous video frame and the current video frame.
[0042] In the above scheme, after the super-resolution processing of the to-be-super-resolution region, the method further comprises the steps of:
[0043] The super-resolution result of the to-be-super-resolution region is obtained.
[0044] The super-resolution result is backfilled to the second video picture.
[0045] In the above scheme, the super-resolution result is backfilled to the second video picture, and the method comprises the steps of:
[0046] The magnification of the super-resolution result is determined.
[0047] The magnification is applied to the second video picture.
[0048] The super-resolution result is backfilled to the second video picture after magnification.
[0049] In the above scheme, the method further comprises:
[0050] Based on the super-resolution result, the second video picture after zooming is cut out;
[0051] The target picture obtained by cutting out is taken as output.
[0052] In another aspect, the embodiments of the present application further provide a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the video super-resolution method.
[0053] In another aspect, the embodiments of the present application provide an electronic device, comprising a processor and a memory, which are connected to each other, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to invoke the program instructions to execute the steps of the video super-resolution method provided by the first aspect of the present application.
[0054] In another aspect, the embodiments of the present application provide a computer readable storage medium, comprising: the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of the video super-resolution method provided by the first aspect of the present application.
[0055] The embodiments of the present application obtain a first video picture captured by a first camera, determine a shooting target based on the first video picture, control a second camera to capture the shooting target to obtain a second video picture containing the shooting target. The region to be super-resolved is determined based on the position of the shooting target in the second video picture, and the region to be super-resolved is super-resolved. The embodiments of the present application are a target super-resolution processing scheme of dual-camera linkage, which determines the shooting target based on the first video picture captured by the first camera, captures the shooting target through the second camera, determines the region to be super-resolved according to the position of the shooting target in the second video picture of the second camera, and then performs super-resolution processing, so as to improve the picture quality and clarity of the target. Compared with the video super-resolution scheme using a single camera in the related art, the picture quality and clarity of the target after super-resolution processing are higher. Moreover, the embodiments of the present application only need to perform super-resolution processing on the region to be super-resolved, without performing super-resolution processing on the entire video frame, so that the super-resolution processing efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is an implementation flow diagram of a video super-resolution method provided by the embodiments of the present application;
[0057] Figure 2 is a scene diagram of a conference room provided by the embodiments of the present application;
[0058] Figure 3is a schematic diagram of a target region provided by an embodiment of the present application.
[0059] Figure 4 is a schematic diagram of a target region provided by an embodiment of the present application.
[0060] Figure 5 is a schematic diagram of a target region provided by an embodiment of the present application.
[0061] Figure 6 is a schematic diagram of a single-frame super-resolution scheme provided by an embodiment of the present application.
[0062] Figure 7 is a schematic diagram of a multi-frame super-resolution scheme provided by an embodiment of the present application.
[0063] Figure 8 is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0064] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0065] Super-resolution reconstruction of a video picture can improve the clarity of a target in the video picture. However, in related technologies, a wide-angle camera is usually used to capture a video in order to obtain a larger shooting range. The video picture is a global picture captured by the wide-angle camera. As a result, there are too many targets in the video picture. Even if the video picture is subjected to super-resolution reconstruction, the clarity and quality of the targets are still not high. Moreover, super-resolution reconstruction of the entire image has low processing efficiency.
[0066] To overcome the above-mentioned shortcomings of related technologies, an embodiment of the present application provides a video super-resolution method, which can improve the picture clarity of a target. In order to describe the technical solutions of the present application, specific embodiments will be described below.
[0067] Figure 1 is a schematic diagram of an implementation process of a video super-resolution method provided by an embodiment of the present application. The execution subject of the video super-resolution method is an electronic device. Referring to Figure 1 , the video super-resolution method comprises the following steps.
[0068] S101, a first video picture captured by a first camera is acquired.
[0069] In the embodiment, at least a first camera and a second camera are included, and the embodiment can be applied to a video conference system, and application scenarios include conference rooms, classrooms and other places where video pictures need to be shared.
[0070] The first camera can be a wide-angle camera, and the first camera captures a wide-angle picture. For example, as shown in a conference room, the first video picture captured by the first camera includes multiple targets. Figure 2
[0071] In S102, a target for shooting is determined based on the first video picture.
[0072] The target for shooting can be manually selected or automatically determined by an algorithm.
[0073] For example, target detection can be performed on the first video picture, target tracking is performed on the detected target, and the target for shooting is determined according to the target tracking result. For example, in a conference room, gestures / lip movements of a person are recognized by a target tracking algorithm, and the target for shooting can be determined by the gestures / lip movements.
[0074] For example, the target tracking algorithm can include but is not limited to a high-speed tracking with kernelized correlation filters (KCF) algorithm or an accurate scale estimation for robust visual tracking (DSST) algorithm.
[0075] In S103, the second camera is controlled to shoot the target for shooting, and a second video picture containing the target for shooting is obtained.
[0076] The second camera can be a long-focus lens, and the second camera can be arranged on a holder. The shooting angle of the second camera can be adjusted by rotating the holder, so that the second camera shoots the target for shooting and obtains a second video picture containing the target for shooting.
[0077] In S104, a super-resolution region to be super-resolved is determined based on the position of the target for shooting in the second video picture.
[0078] The position of the target for shooting in the second video picture is determined, the target for shooting is framed with a frame centered at the position of the target for shooting, and the region in the frame is the super-resolution region to be super-resolved.
[0079] In S105, the super-resolution region to be super-resolved is super-resolved.
[0080] Super-resolution processing refers to super-resolution reconstruction of an image in a region to be super-resolved. The super-resolution reconstruction is a process of improving a given low-resolution image to a high-resolution image by using a computer algorithm.
[0081] The embodiment can use any one of an interpolation-based super-resolution reconstruction algorithm, a degradation model-based super-resolution reconstruction algorithm, and a deep learning-based super-resolution reconstruction algorithm.
[0082] In the interpolation-based method, each pixel on the image is regarded as a point on the image plane, and the estimation of the super-resolution image can be regarded as a process of fitting unknown pixel information on the plane by using known pixel information, which is usually completed by a predefined transformation function or interpolation kernel. Common interpolation-based methods include nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, and the like.
[0083] The degradation model-based super-resolution reconstruction algorithm starts from the degradation model of the image, and assumes that the high-resolution image is obtained by appropriate motion transformation, blurring, and noise from the low-resolution image. This method extracts key information from the low-resolution image, and combines prior knowledge of the unknown super-resolution image to constrain the generation of the super-resolution image. Common methods include iterative back projection, convex set projection, and maximum a posteriori probability.
[0084] The deep learning-based super-resolution reconstruction algorithm learns a certain correspondence between the low-resolution image and the high-resolution image from a large amount of training data, and then predicts the high-resolution image corresponding to the low-resolution image according to the learned mapping relationship, so as to realize the super-resolution reconstruction process of the image. Common learning-based methods include manifold learning and sparse coding method.
[0085] After super-resolution processing is performed on the region to be super-resolved, the quality and clarity of the photographed target in the super-resolution image are greatly improved, and the user experience can be improved.
[0086] The embodiment of the application obtains a first video picture captured by a first camera, determines a shooting target based on the first video picture, controls a second camera to capture the shooting target to obtain a second video picture containing the shooting target. A region to be super-resolution processed is determined based on the position of the shooting target in the second video picture, and the region to be super-resolution processed is super-resolution processed. The embodiment of the application is a target super-resolution processing scheme of dual-camera linkage, the shooting target is determined by using the first video picture captured by the first camera, the shooting target is captured by the second camera, the region to be super-resolution processed is determined according to the position of the shooting target in the second video picture of the second camera, and then super-resolution processing is performed, so that the image quality and the definition of the target are improved. Compared with the video super-resolution scheme using a single camera in the related art, the image quality and the definition of the target after super-resolution processing are higher. Moreover, the embodiment of the application only needs to perform super-resolution processing on the region to be super-resolution processed, and does not need to perform super-resolution processing on the entire video frame, so that the super-resolution processing efficiency can be improved.
[0087] In an embodiment, after the shooting target is determined based on the first video picture, the method further includes:
[0088] A rotation angle of the second camera is determined according to the position of the shooting target in the first video picture;
[0089] The control of the second camera to capture the shooting target to obtain a second video picture containing the shooting target includes:
[0090] The second camera is controlled to rotate to a specified angle based on the rotation angle, so that the shooting target is located in a central region of the second video picture.
[0091] For example, the second camera is mounted on a gimbal, the gimbal angle coordinates of the second camera are determined according to the position of the shooting target in the first video picture by using the calibration result of the first camera and the second camera, and it is ensured that the shooting target is located in the central region of the second video picture after the gimbal rotates to the specified angle.
[0092] In an embodiment, before the region to be super-resolution processed is determined based on the position of the shooting target in the second video picture, the method further includes:
[0093] The position of the shooting target in the second video picture is determined.
[0094] For example, target detection and tracking are performed on the first video picture and the second video picture, the behavior (for example, hand gesture / lip movement) of the target is determined, and the target in the second video picture that has the same behavior as the shooting target is the shooting target, so that the position of the shooting target in the second video picture can be determined.
[0095] In an embodiment, the determination of the position of the shooting target in the second video picture includes:
[0096] determining a target region in the first video picture which is identical to the second video picture, the shooting target being in the target region;
[0097] determining the position of the shooting target in the second video picture based on the position of the shooting target in the target region.
[0098] For example, the target region can be determined by similarity matching of picture content.
[0099] As shown in Figure 3 , the upper half is the first video picture, and the lower half is the second video picture. The region in the dashed box in the first video picture is the target region. Figure 3 As shown in
[0100] , the position of the shooting target in the target region is known. Since the target region is identical to the second video picture, the position of the shooting target in the target region and the second video picture can be considered the same. Therefore, according to the position of the shooting target in the target region, the position of the shooting target in the second video picture can be determined. Figure 4 In an embodiment, the determining a target region in the first video picture which is identical to the second video picture comprises:
[0101] determining an initial region in the first video picture which is identical to the second video picture;
[0102] expanding based on the initial region to obtain an expanded region, the initial region being in the expanded region;
[0103] determining a target region in the expanded region which is most identical to the second video picture.
[0104] In an embodiment, the initial region contains the shooting target. In an embodiment, a region which is centered on the shooting target and identical to the second video picture can be the initial region.
[0105] The initial region is expanded, for example, the four edges of the initial region are respectively stretched outward by a certain distance.
[0106] The relationship between the expanded region and the initial region is shown in
[0107] , the initial region being in the expanded region. Figure 5 The left lower part is a schematic diagram of the expanded region, and the right lower part is a schematic diagram of the target region determined in the expanded region. In the expanded region, picture matching is performed with the second video picture. The region in the expanded region which is most identical to the second video picture is the target region. Figure 5
[0108] In an embodiment, the determining the target region in the extended region that is most matched with the second video frame comprises:
[0109] The target region is determined in the extended region based on the second video frame by a matching algorithm.
[0110] The matching algorithm comprises at least a key point matching algorithm and a template matching algorithm.
[0111] In an embodiment, the determining the target region in the extended region based on the second video frame by a matching algorithm comprises:
[0112] Template matching is performed in the extended region using the second video frame as a template.
[0113] The region in the extended region that has the highest matching degree with the template is determined as the target region.
[0114] For example, a mapping frame corresponding to the initial region (the mapping frame just contains the initial region) can be slid in the extended region, the pixel point similarity (matching degree) between the frame in the mapping frame and the template is calculated each time the mapping frame is slid (each time the mapping frame is slid by a fixed distance, and the mapping frame can be slid according to a preset track), until the mapping frame slides through the entire extended region, and the region in the mapping frame when the pixel point similarity is the highest is determined as the target region.
[0115] In an embodiment, the determining the target region in the extended region based on the second video frame by a matching algorithm comprises:
[0116] Key points in the second video frame and the extended region are determined.
[0117] The region in the extended region that has the most key points matched with the second video frame is determined as the target region by key point matching.
[0118] The size of the region is consistent with the initial region.
[0119] The key point matching algorithm refers to finding points (key points) with obvious features from two images, and then matching the key points in the two images to find one-to-one corresponding points. The more matching points, the higher the matching degree of the two images. According to this principle, the local frame in the extended region that is most matched with the second video frame can be determined, and the region where the local frame is located is the target region.
[0120] Common key point features include feature point detection and extraction (Oriented Fast and Rotated Brief, ORB), scale invariant feature transform matching (Scale Invariant Feature Transform, SIFT), etc.
[0121] ORB matching refers to detecting feature points in two images using the ORB algorithm and matching by calculating the descriptors of the feature points. SIFT can locate local features in an image, which are usually referred to as key points of the image. These key points are scale and rotation invariants.
[0122] In an embodiment, the to-be-super-resolution region is determined based on the position of the photographed target in the second video frame.
[0123] The position of the human head frame corresponding to the photographed target is determined based on the position of the photographed target in the second video frame, and the position of the human head frame is taken as the to-be-super-resolution region.
[0124] For example, if the type of the photographed target is "human", the position of the human head of the photographed target in the second video frame can be obtained, the human head of the photographed target is framed with a human head frame, and the region in the human head frame is the to-be-super-resolution region.
[0125] In an embodiment, the super-resolution processing of the to-be-super-resolution region includes:
[0126] The position of the to-be-super-resolution region corresponding to the current video frame of the second video frame is determined.
[0127] The position of the super-resolution region of the previous video frame of the second video frame is obtained.
[0128] A super-resolution strategy is determined based on the distance between the position of the to-be-super-resolution region and the position of the super-resolution region.
[0129] The to-be-super-resolution region is super-resolution processed based on the super-resolution strategy.
[0130] Here, the to-be-super-resolution region and the super-resolution region can both refer to the region in the human head frame. The distance between the positions of the human head frames in the two frames of pictures is calculated, and different super-resolution strategies are selected according to the distance. The super-resolution strategies include single-frame super-resolution and multi-frame super-resolution.
[0131] In an embodiment, the super-resolution strategy is determined based on the distance between the position of the to-be-super-resolution region and the position of the super-resolution region, including:
[0132] If the distance is greater than a threshold value, the super-resolution reconstruction is performed based on the current video frame.
[0133] If the distance is less than or equal to the threshold, then super-resolution reconstruction is performed based on the previous video frame and the current video frame.
[0134] This embodiment selects an appropriate super-resolution strategy based on the distance between the region to be super-resolution and the super-resolution region in two consecutive video frames. This ensures the effectiveness of the video super-resolution results, improves the clarity and detail of the video, and maintains the continuity and consistency of the video content.
[0135] In one embodiment, the super-resolution reconstruction based on the current video frame includes:
[0136] Copy the current video frame to obtain a copied video frame;
[0137] The copied video frame and the current video frame are fed together into the super-resolution model for super-resolution reconstruction.
[0138] Correspondingly, the super-resolution reconstruction based on the previous video frame and the current video frame includes:
[0139] The previous video frame and the current video frame are fed together into the super-resolution model for super-resolution reconstruction.
[0140] If the distance exceeds a threshold, a single-frame super-resolution scheme is used because the overlap between heads in two consecutive frames is insufficient; using two consecutive head images for super-resolution would yield poor results. Therefore, super-resolution reconstruction is performed based on the current video frame. For example, the current video frame is copied and used as the previous video frame, then fed together with the current video frame into the super-resolution model for processing. For example, ... Figure 6 As shown, Figure 6 It is a single-frame super-resolution scheme, and the area within the head frame is the area that needs to be super-resolution.
[0141] If the distance is less than or equal to the threshold, a multi-frame super-resolution scheme is used, where the current video frame and the previous video frame are fed together into the super-resolution model for processing. This multi-frame scheme ensures the stability and smoothness of the video super-resolution results. For example, ... Figure 7 As shown, Figure 7 It is a multi-frame super-resolution scheme, where the area to be super-resolution includes both the head position of the target in the previous video frame and the head position of the target in the current video frame.
[0142] The biggest difference between video and image super-resolution is that video super-resolution uses inter-frame information. This embodiment uses consecutive video frames for video super-resolution, improving the clarity and continuity of the video.
[0143] Continuous video frame super-resolution mainly involves the technology of using low-resolution video reference frames and multiple adjacent frames to reconstruct high-resolution video. This technology aims to improve the clarity and details of the video while maintaining the continuity and consistency of the video content. The methods of continuous video frame super-resolution mainly include motion estimation compensation-based techniques and deep learning-based methods.
[0144] Motion estimation compensation-based methods mainly rely on extracting inter-frame motion information and performing inter-frame warping operations to align them. Motion estimation calculates the motion between adjacent frames through the optical flow method, while motion compensation is used to align video frames based on this motion information.
[0145] Deep learning-based methods achieve the enhancement of inter-frame content continuity by learning the motion trajectory of consecutive frames as the prediction information for inter-frame super-resolution. This method uses deep neural network models to predict and restore high-frequency details, thereby improving the resolution and quality of the video.
[0146] In an embodiment, after the super-resolution processing of the to-be-super-resolved region, the method further comprises:
[0147] obtaining a super-resolution result of the to-be-super-resolved region;
[0148] backfilling the super-resolution result to the second video frame.
[0149] The super-resolution result of the to-be-super-resolved region is the obtained super-resolution region, and the super-resolution result is backfilled to the corresponding second video frame.
[0150] In an embodiment, the backfilling of the super-resolution result to the second video frame comprises:
[0151] determining a magnification factor of the super-resolution result;
[0152] applying the magnification factor to the second video frame;
[0153] backfilling the super-resolution result to the second video frame after magnification.
[0154] After super-resolution, the super-resolution region has a magnification factor relative to the original image size, the magnification factor is applied to the full image, and then the super-resolution result is backfilled to the original image.
[0155] In an embodiment, the method further comprises:
[0156] magnifying the second video frame after magnification with the super-resolution result as the center;
[0157] taking the target frame obtained by the cutout as the output.
[0158] Because the picture size of the video output is fixed, the enlarged video picture exceeds the original picture size, and therefore the second video picture after enlargement is cut out in this embodiment, and a suitable picture ratio is cut out as the output, with the super-resolution region as the center.
[0159] In an embodiment, a video super-resolution process is provided, comprising the following steps:
[0160] S201, obtaining a wide-angle picture of a main camera and a long-focus picture of a secondary camera (a gimbal camera) at the same time from a device end.
[0161] The device end can refer to a camera in a conference room, which has two cameras, i.e., a wide-angle camera and a long-focus camera. For example, pictures of the two cameras can be obtained at a specified frame rate (30 fps).
[0162] S202, tracking all targets on the wide-angle picture.
[0163] This is a multi-target tracking algorithm, and no limitation is imposed on which multi-target tracking algorithm is used.
[0164] S203, selecting (manually or automatically by an algorithm) a target on the wide-angle picture for tracking by the gimbal camera.
[0165] The target for close-up can be manually selected or automatically obtained by an algorithm, such as a gesture / lip movement recognition algorithm.
[0166] S204, mapping the target in the wide-angle picture to the picture of the long-focus camera.
[0167] The position of the target for close-up in the wide-angle picture is determined, and the angle coordinates of the gimbal are determined according to the calibration results of the main and secondary cameras, so that the target for close-up is in the center region of the picture of the secondary camera after the gimbal is turned to the specified angle.
[0168] S205, determining a region to be super-resolved according to the position of the target for tracking in the long-focus picture.
[0169] After the gimbal is turned to the specified angle, the region in the wide-angle picture that is equivalent to the long-focus picture is determined, and the region to be super-resolved is determined according to the method in the above embodiment.
[0170] S206, performing super-resolution processing on the region to be super-resolved.
[0171] The super-resolution model used here is a model that takes two consecutive frames as input, and can also be a model that takes one frame as input.
[0172] S207, filling the super-resolution result into the original picture according to a ratio.
[0173] After super-resolution, the super-resolution region has a magnification relative to the original size. The magnification is applied to the full image, and then the super-resolution result is backfilled to the original image.
[0174] S208, output the super-resolution result according to the predetermined strategy.
[0175] Centered on the super-resolution region, the appropriate picture ratio is deducted as the output.
[0176] The embodiment of the present application can realize fast secondary shooting tracking effect through the linkage of the dual cameras, and can improve the picture quality and clarity of the tracker.
[0177] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0178] It should be understood that when used in the present specification and the appended claims, the terms "include" and "contain" indicate the presence of described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.
[0179] It should be noted that the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0180] In addition, in the embodiments of the present application, "first", "second" and the like are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0181] In practical application, the video super-resolution method can be realized by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), or a programmable gate array (FPGA).
[0182] Based on the hardware implementation of the above program module, and in order to realize the method of the embodiment of the present application, the embodiment of the present application further provides an electronic device. Figure 8 The schematic diagram of the hardware composition structure of the electronic device of the embodiment of the present application is shown as Figure 8 The electronic device comprises:
[0183] a memory storing a computer program;
[0184] The communication interface is capable of information interaction with other devices such as network devices and the like;
[0185] The processor is connected with the communication interface to realize information interaction with other devices;
[0186] When the computer program is executed by the processor, the processor is configured to acquire a first video picture captured by a first camera; determine a shooting target based on the first video picture; control a second camera to capture the shooting target to obtain a second video picture containing the shooting target; determine a region to be super-resolution processed based on a position of the shooting target in the second video picture; and perform super-resolution processing on the region to be super-resolution processed.
[0187] In an embodiment, the processor is configured to:
[0188] Determine a rotation angle of the second camera according to a position of the shooting target in the first video picture;
[0189] In an embodiment, the processor is configured to:
[0190] Control the second camera to rotate to a specified angle based on the rotation angle, so that the shooting target is located in a central region of the second video picture.
[0191] In an embodiment, the processor is configured to determine the position of the shooting target in the second video picture.
[0192] In an embodiment, the processor is configured to:
[0193] Determine a target region in the first video picture that is equivalent to the second video picture, and the shooting target is located in the target region;
[0194] Determine the position of the shooting target in the second video picture based on a position of the shooting target in the target region.
[0195] In an embodiment, the processor is configured to:
[0196] Determine an initial region in the first video picture that is equivalent to the second video picture;
[0197] Expand the initial region to obtain an expanded region, and the initial region is located in the expanded region;
[0198] Determine a target region in the expanded region that is most equivalent to the second video picture.
[0199] In an embodiment, the processor is configured to:
[0200] determine the target region in the extended region based on the second video picture by a matching algorithm.
[0201] In an embodiment, the processor is configured to:
[0202] perform template matching in the extended region with the second video picture as a template;
[0203] determine the target region in the extended region as a region with the highest matching degree to the template.
[0204] In an embodiment, the processor is configured to:
[0205] determine key points in the second video picture and the extended region;
[0206] determine the target region in the extended region as a region with the largest number of key points matching the second video picture by key point matching.
[0207] In an embodiment, the processor is configured to:
[0208] determine a human head frame position corresponding to the shooting target based on a position of the shooting target in the second video picture, and take the human head frame position as the region to be super-resolved.
[0209] In an embodiment, the processor is configured to:
[0210] determine a region to be super-resolved position corresponding to a current video frame of the second video picture;
[0211] obtain a super-resolved region position of a previous video frame of the second video picture;
[0212] determine a super-resolution strategy based on a distance between the region to be super-resolved position and the super-resolved region position;
[0213] perform super-resolution processing on the region to be super-resolved based on the super-resolution strategy.
[0214] In an embodiment, the processor is configured to:
[0215] if the distance is greater than a threshold, perform super-resolution reconstruction based on the current video frame;
[0216] if the distance is less than or equal to the threshold, perform super-resolution reconstruction based on the previous video frame and the current video frame.
[0217] In an embodiment, the processor is configured to:
[0218] copy the current video frame to obtain a copied video frame;
[0219] The copied video frame and the current video frame are sent into the super-resolution model together for super-resolution reconstruction.
[0220] In an embodiment, the processor is configured to:
[0221] The last video frame and the current video frame are sent into the super-resolution model together for super-resolution reconstruction.
[0222] In an embodiment, the processor is configured to:
[0223] Obtain a super-resolution result of the super-resolution region;
[0224] Backfill the super-resolution result to the second video picture.
[0225] In an embodiment, the processor is configured to:
[0226] Determine a magnification of the super-resolution result;
[0227] Apply the magnification to the second video picture;
[0228] Backfill the super-resolution result to the second video picture after magnification.
[0229] In an embodiment, the processor is configured to:
[0230] Perform matting on the second video picture after magnification with the super-resolution result as the center.
[0231] Output the target picture obtained by matting.
[0232] Of course, in actual applications, various components in the electronic device are coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between the components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the bus system in Figure 8 .
[0233] The memory in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer programs used to operate on the electronic device.
[0234] It can be appreciated that the memory can be a volatile memory or a nonvolatile memory, and can also include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a ferromagnetic random access memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM). The magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example and not limitation, many forms of RAM can be used, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM). The memory described in the embodiments of the present application is intended to include but not limited to these and any other suitable types of memory.
[0235] The method disclosed in the embodiments of the present application can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip with a signal processing capability. In the implementation process, the steps of the above method can be completed by hardware integrated logic circuit or software form of instructions in the processor. The processor can be a general processor, DSP, or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiments of the present application, the hardware decoding processor can be directly embodied to execute the steps of the method, or the hardware and software modules in the decoding processor can be combined to execute the steps of the method. The software module can be located in a storage medium, which is located in a memory. The processor reads the program in the memory and combines the hardware to complete the steps of the method.
[0236] Alternatively, the processor executes the program to implement the corresponding processes realized by the electronic device in each method of the embodiments of the present application. For brevity, details are not repeated here.
[0237] In the exemplary embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program executable by the processor of the electronic device to complete the steps of the method disclosed in the embodiments of the present application.
[0238] In the exemplary embodiments, the embodiments of the present application also provide a storage medium, i.e. a computer storage medium, specifically a computer readable storage medium, for example, a first memory for storing a computer program, which can be executed by the processor of the electronic device to complete the steps of the method. The computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0239] In the several embodiments provided in the present application, it should be understood that the disclosed device, electronic device and method can be implemented by other manners. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interface, and indirect coupling or communication connection between the various components can be electrical, mechanical or other forms.
[0240] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the embodiment.
[0241] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0242] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction-related hardware, and the aforementioned program can be stored in a computer-readable storage medium, and the program executes the steps including the above-mentioned method embodiments when executed; and the aforementioned storage medium includes mobile storage devices, ROM, RAM, magnetic discs or optical discs and various storage medium that can store program codes.
[0243] Alternatively, the integrated units of the present application, if implemented in the form of software functional modules and sold or used as independent products, can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of software products, which are stored in a storage medium and include a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes mobile storage devices, ROM, RAM, magnetic discs or optical discs and various storage medium that can store program codes.
[0244] It should be noted that the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0245] In addition, in the present application, "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0246] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of video super-resolution, characterized in that, The method comprises: acquiring a first video picture captured by a first camera; determining a shooting target based on the first video picture; controlling a second camera to capture the shooting target to obtain a second video picture containing the shooting target; determining a super-resolution region to be super-resolved based on the position of the shooting target in the second video picture; performing super-resolution processing on the super-resolution region to be super-resolved.
2. The method of claim 1, wherein, After determining the shooting target based on the first video picture, the method further comprises: determining a rotation angle of the second camera according to the position of the shooting target in the first video picture; controlling the second camera to capture the shooting target to obtain a second video picture containing the shooting target, comprising: controlling the second camera to rotate to a specified angle based on the rotation angle, so that the shooting target is located in the center region of the second video picture.
3. The method of claim 1, wherein, Before determining the super-resolution region to be super-resolved based on the position of the shooting target in the second video picture, the method further comprises: determining the position of the shooting target in the second video picture.
4. The method of claim 3, wherein, The method of determining the position of the shooting target in the second video picture comprises: determining a target region equivalent to the second video picture in the first video picture, wherein the shooting target is located in the target region; determining the position of the shooting target in the second video picture based on the position of the shooting target in the target region.
5. The method of claim 4, wherein, The method of determining a target region equivalent to the second video picture in the first video picture comprises: determining an initial region equivalent to the second video picture in the first video picture; expanding based on the initial region to obtain an expanded region, wherein the initial region is located in the expanded region; determining a target region most matching the second video picture in the expanded region.
6. The method of claim 5, wherein, The method of determining a target region most matching the second video picture in the expanded region comprises: determining the target region in the expanded region based on the second video picture by a matching algorithm.
7. The method of claim 6, wherein, The method of determining the target region in the expanded region based on the second video picture by a matching algorithm comprises: performing template matching in the expanded region using the second video picture as a template; determining the region in the expanded region with the highest matching degree to the template as the target region.
8. The method of claim 6, wherein, The method of determining the target region in the expanded region based on the second video picture by a matching algorithm comprises: determining key points in the second video picture and the expanded region; determining the region in the expanded region with the largest number of matching key points to the second video picture as the target region by key point matching.
9. The method of claim 1, wherein, The method of determining a super-resolution region to be super-resolved based on the position of the shooting target in the second video picture comprises: determining a human head frame position corresponding to the shooting target based on the position of the shooting target in the second video picture, and taking the human head frame position as the super-resolution region to be super-resolved.
10. The method of claim 1, wherein, The method of performing super-resolution processing on the super-resolution region to be super-resolved comprises: determining a super-resolution region position corresponding to a current video frame of the second video picture; obtain a super-resolution region position of a previous video frame of the second video picture; determine a super-resolution strategy based on a distance between the to-be-super-resolved region position and the super-resolution region position; perform super-resolution processing on the to-be-super-resolved region based on the super-resolution strategy.
11. The method of claim 10, wherein, The determining of the super-resolution strategy based on the distance between the to-be-super-resolved region position and the super-resolution region position comprises: if the distance is greater than a threshold, performing super-resolution reconstruction based on the current video frame; if the distance is less than or equal to the threshold, performing super-resolution reconstruction based on the previous video frame and the current video frame.
12. The method of claim 11, wherein, The performing of the super-resolution reconstruction based on the current video frame comprises: copying the current video frame to obtain a copied video frame; sending the copied video frame and the current video frame together to a super-resolution model for super-resolution reconstruction. Correspondingly, the performing of the super-resolution reconstruction based on the previous video frame and the current video frame comprises: sending the previous video frame and the current video frame together to the super-resolution model for super-resolution reconstruction.
13. The method of claim 1, wherein, After the super-resolution processing on the to-be-super-resolved region, the method further comprises: obtaining a super-resolution result of the to-be-super-resolved region; backfilling the super-resolution result to the second video picture.
14. The method of claim 13, wherein, The backfilling of the super-resolution result to the second video picture comprises: determining a magnification of the super-resolution result; applying the magnification to the second video picture; backfilling the super-resolution result to the second video picture after magnification.
15. The method of claim 14, wherein, The method further comprises: performing matting on the second video picture after magnification with the super-resolution result as a center; taking a target picture obtained by the matting as an output.
16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the video super-resolution method as claimed in claims 1 to 15 when executing the computer program.
17. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program comprises program instructions which, when executed by a processor, cause the processor to execute the video super-resolution method as claimed in claims 1 to 15.