A sliding zoom method, device and storage medium based on depth camera

Through the depth camera, the main area and background area are divided, the main area is enlarged and background area is reduced, and the sliding zoom effect is solved in the existing technology is difficult to achieve, and the convenient sliding zoom effect is achieved.

CN114359005BActive Publication Date: 2025-06-06YUANLI TUXIN (CHONGQING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111341984.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-06-06
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

The prior art is difficult to achieve sliding zoom effect easily, especially when shooting videos, it often requires complex equipment and post-editing.

Method used

By using a depth camera to acquire scene maps and depth maps, and calibrate the depth maps using calibration data, dividing the main area and background area, the main area is enlarged and background area is reduced, thereby achieving the sliding zoom effect.

Benefits of technology

This method can easily and quickly realize the sliding zoom effect without the need for complex equipment and post-editing, improving the convenience and effect of shooting videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359005B_ABST
    Figure CN114359005B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a sliding zoom method, device and storage medium based on a depth camera, and the sliding zoom method includes: obtaining a scene map and a depth map; calibrating the depth map according to calibration data and the scene map to obtain a calibrated depth map; dividing the scene map into a main area and a background area according to the calibrated depth map; obtaining multiple groups of pairs of enlarged sub-images of the main area and reduced sub-images of the background area, and fusing each pair of the enlarged sub-images and the reduced sub-images to obtain multiple frames of display images. The technical solution of the present application only needs to collect a scene map with a subject and a depth map, and divide the foreground and background according to the depth map, enlarge the foreground, and reduce the background, so as to achieve a sliding zoom effect in which the subject moves forward and the background becomes farther away.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and specifically to an embodiment of the present application relating to a sliding zoom method, device, and storage medium based on a depth camera. Background Art

[0002] In order to create a stunning effect in a scene, a special camera movement technique is often used. By pushing and pulling the lens and zooming at the same time, the subject's position remains unchanged. Due to the perspective relationship, the subject's position and size remain unchanged, while the background keeps changing, creating a magical spatial effect. Figure 1 As shown, it should be noted that Figure 1 The faces are blurred in the figure. It is understandable that this processing does not affect the image’s function of explaining the technical solution.

[0003] Sliding zoom is also known as Hitchcock zoom, or Dolly Zoom in English. It is a shooting technique that changes the distance between the subject and the background to create an effect where the size of the subject itself does not change. Specifically, sliding zoom will create an effect where the size and position of the subject remains unchanged, but it feels like the subject is moving forward and the background is moving backward.

[0004] Currently, more and more shooting devices with stabilizers have built-in sliding zoom functions (for example, DJI's "Magic" zoom version and SMOOTH's mobile phone stabilizer). It is foreseeable that sliding zoom will become a standard feature of devices with gimbals in the future, but in the past two years, most devices still require post-editing to achieve the sliding zoom effect. In other words, the methods of implementing Dolly Zoom in related technologies include: using a slider when shooting a movie, or using a stabilizer or gimbal to shoot with a handheld device or a flying device. It is understandable that each of the above methods has inconvenient or unfriendly shooting methods, either with the help of a slider, a stabilizer or a gimbal, etc., which requires a high shooting cost for ordinary photographers, and the cost is also expensive, and it cannot be well popularized.

[0005] Therefore, how to more conveniently shoot a video with a sliding zoom effect has become a technical problem that needs to be solved urgently. Summary of the invention

[0006] The purpose of the embodiments of the present application is to provide a sliding zoom method, device and storage medium based on a depth camera. Through the technical solution of the present application, only one scene image and one depth image with a clear subject need to be collected to achieve a sliding zoom effect. For example, in some embodiments of the present application, the depth image and the scene image of the main camera are aligned using calibration data, and then algorithm processing is performed to divide the foreground and background according to the depth image, enlarge the foreground (corresponding to the subject area), and reduce the background (corresponding to the background area), so as to achieve a sliding zoom effect in which the subject moves forward and the background moves farther away.

[0007] In a first aspect, some embodiments of the present application provide a sliding zoom method based on a depth camera, the sliding zoom method comprising: acquiring a scene map and a depth map; calibrating the depth map according to calibration data and the scene map to obtain a calibrated depth map; dividing the scene map into a main area and a background area according to the calibrated depth map; acquiring pairs of enlarged sub-images of the main area and reduced sub-images of the background area, and fusing each pair of the enlarged sub-images and the reduced sub-images to obtain multiple frames of display images.

[0008] Some embodiments of the present application divide the scene graph into a main area and a background area according to the depth map, and gradually enlarge the main area in multiple frames (i.e., obtain an enlarged sub-image of the main area in each frame in the multiple frames), thereby gradually enlarging the object to be enlarged, and gradually reducing the background area in multiple frames (i.e., obtain a reduced sub-image of the background area in each frame in the multiple frames), thus realizing a sliding zoom method. Compared with the related art that must rely on a complex shooting process, the embodiments of the present application can realize a sliding zoom method by performing relevant processing based on the scene graph and the depth map of the photographed subject, and the image shooting process is simpler and the data processing process is also simpler.

[0009] In some embodiments, the scene map is obtained by photographing a main camera, and the depth map is obtained by photographing a depth camera; wherein, before calibrating the depth map according to the scene map, the sliding zoom method also includes: obtaining a conversion parameter calibration between the depth camera and the main camera to obtain the calibration data.

[0010] Some embodiments of the present application need to first correct the depth map based on the scene graph before segmenting the main area and the background area, so that the depth map corresponds to each pixel on the scene graph (that is, the pixel positions corresponding to the same photographed point in the space are the same on the depth map and the scene graph), so as to facilitate the subsequent segmentation of the main area and the background area on the scene graph according to the depth map.

[0011] In some embodiments, dividing the scene graph into a main body area and a background area according to the calibrated depth map includes: converting the calibrated depth map into a disparity map; obtaining a reference disparity value from the disparity map by obtaining a disparity value of a target focus point, and obtaining a reference depth value from the calibrated depth map; if it is determined that the reference depth value is less than a set distance threshold, all pixels that are greater than the reference disparity value by a first set threshold are confirmed as belonging to the pixel area where the main body is located, and the position corresponding to the pixel area where the main body is located on the scene graph is confirmed as the main body area, and the area on the scene graph other than the main body area is the background area; or, if it is confirmed that the reference depth value is greater than or equal to the set distance threshold, all pixels that are greater than the reference disparity value by a second set threshold are confirmed as belonging to the pixel area where the main body is located, and the position corresponding to the pixel area where the main body is located on the scene graph is confirmed as the main body area, and the area on the scene graph other than the main body area is the background area.

[0012] Some embodiments of the present application obtain a disparity map based on a depth map, and then segment the scene map into a main area and a background area based on the information on the disparity map and the depth map, convert the depth map into a disparity map, and complete the value range conversion. The depth distance with a relatively large value can be converted into a disparity value with a smaller value, which is more convenient to operate when segmenting the main area and the background area.

[0013] In some embodiments, before converting the calibrated depth map into a disparity map, the sliding zoom method also includes: obtaining a maximum depth value and a minimum depth value from the calibrated depth map or the depth map; confirming that the maximum depth value is greater than or equal to a set distance, and confirming that the difference between the maximum depth value and the minimum depth value is greater than or equal to a set distance difference.

[0014] Some embodiments of the present application need to confirm through the maximum depth value that the background of the subject being photographed is not too close, and confirm through the maximum depth value and the minimum depth value that the photographed image is not an image that cannot clearly distinguish the subject area and the background. Only when these conditions are met can the scene graph be converted into the subject area and the background area according to the depth map, avoiding the direct division of all scene graphs into subject areas and background areas, resulting in the failure of segmentation of some scene graphs.

[0015] In some embodiments, the obtaining of pairs of enlarged sub-images of the main area and reduced sub-images of the background area includes: obtaining a total magnification of the main area and a total reduction magnification of the background area; determining a target magnification of each frame according to the total magnification and the total number of frames of the composite video, and determining a target reduction magnification of each frame according to the total reduction and the total number of frames, wherein the total number of frames of the composite video is equal to the number of frames of the multiple display images; obtaining an enlarged sub-image of each frame according to the target magnification of each frame and the main area, and obtaining a reduced sub-image of each frame according to the target reduction magnification of each frame and the background area.

[0016] Some embodiments of the present application further distribute the solved total magnification ratio for the main area in each frame, and distribute the obtained total reduction ratio in each frame, so as to achieve a sliding zoom effect in which the main area is gradually enlarged and the background area is gradually reduced during video playback. Compared with the processing algorithms of the related technologies, the calculation amount is less and the calculation speed is faster.

[0017] In some embodiments, before obtaining a pair of enlarged sub-images of the subject area and reduced sub-images of the background area, the sliding zoom method further includes: obtaining pixels of interest from the scene graph; wherein, obtaining the enlarged sub-image of each frame according to the target magnification of each frame and the subject area includes: constructing an enlargement conversion matrix according to the target magnification of each frame and the pixels of interest; obtaining the enlarged sub-image of each frame according to the enlargement conversion matrix and the subject area; obtaining the reduced sub-image of each frame according to the target reduction ratio of each frame and the background area includes: constructing a reduction conversion matrix according to the target reduction ratio of each frame and the pixels of interest; obtaining the reduced sub-image of each frame according to the reduction conversion matrix and the background area.

[0018] Some embodiments of the present application can effectively improve the visual sensory effect by introducing a pixel of interest (in pixels) and, during post-processing (i.e., obtaining an enlarged sub-image and a reduced sub-image of each frame), enlarging the main area and reducing the background area with the pixel of interest as the center (i.e., through a constructed transformation matrix).

[0019] In some embodiments, obtaining the pixel of interest from the scene graph includes: calculating the center of gravity corresponding to the main body area according to the coordinate values ​​of each pixel included in the main body area; confirming that the target focus point is within the main body area, and obtaining the pixel of interest according to the following formula:

[0020] c=w1*p+w2*p1

[0021] Among them, p1 represents the coordinate value of the center of gravity, p represents the coordinate value of the target focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1.

[0022] In some embodiments of the present application, when the target focus point is located on the main area, the point of interest obtained by using the previous formula can be closer to the target focus point and closer to the center of gravity of the main area, which can ensure a better visual effect when the main area is enlarged.

[0023] In some embodiments, obtaining the pixel of interest from the scene graph includes: calculating the center of gravity of the main body area according to the coordinate values ​​of each pixel corresponding to the main body area; confirming that the target focus point is not in the main body area, taking the center of the scene graph as the new focus point p, and obtaining the initial pixel of interest according to the following formula: c=w2*p+w1*p1; if it is confirmed that the initial pixel of interest c is not inside the main body area, performing at least one iterative calculation of the following formula c=w2*c+w1*p1 until the set number of iterations is reached or it is confirmed that the initial pixel of interest c is inside the main body area, and taking the initial pixel of interest c as the pixel of interest; wherein p1 represents the coordinate value of the center of gravity, p represents the coordinate value of the new focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1.

[0024] In some embodiments of the present application, when the target focus point is not in the main area, the pixel of interest is obtained by using the center of the image and the center of gravity of the main area, and it is ensured as much as possible that the pixel of interest falls on the main area, which can ensure a better visual effect when the main area is enlarged.

[0025] In some embodiments, obtaining the total magnification of the main area and the total reduction ratio of the background area includes: obtaining the total reduction ratio according to the size of the full-size image and the size displayed on the specified interface, wherein the scene map and the depth map both belong to the full-size image.

[0026] Some embodiments of the present application determine the total reduction ratio by using the full-size image and the displayed image, which can effectively avoid the phenomenon of black edges appearing on the displayed image due to lack of image data during the process of reducing the background area.

[0027] In some embodiments, the total reduction ratio is obtained based on the size of the full-size image and the size of the specified interface display, including: taking the pixel of interest as the center, respectively calculating the distance s1_i from the pixel of interest to the four sides constituting the interface display image, and calculating the distance s2_i from the pixel of interest to the four sides constituting the full-size image, where the value of i is an integer greater than or equal to 1 and less than or equal to 4; calculating the ratio according to the following formula to obtain four ratios: r_i=s1_i / s2_i; and selecting the maximum ratio from the four ratios as the total reduction ratio.

[0028] Some embodiments of the present application determine the minimum value that can be reduced by calculating the distance between the pixel of interest and the four sides corresponding to the display size and the shooting size, and use the minimum value that can be reduced as the total reduction ratio of the background area. This can ensure that there is no lack of display data during the reduction of the background area, thereby obtaining a better display effect.

[0029] In some embodiments, the target magnification of each frame is determined according to the total magnification and the total number of frames of the synthesized video, including: if the total number of frames is N, the total magnification is represented by S1max, then the calculation formula of the target magnification of each frame is represented by:

[0030] S1i=1.0+(i-1)*(S1max-1.0) / (N-1)

[0031] Wherein, S1max is a number greater than 1, N is an integer greater than 1, i represents the number of any frame in N frames and i is an integer less than or equal to N.

[0032] Some embodiments of the present application can evenly distribute the total magnification of the main area in each frame through the formula of the target magnification, thereby ensuring the continuity of the main area magnification process between the previous and next frames during video display, thereby obtaining a better display effect.

[0033] In some embodiments, determining the target reduction ratio of each frame according to the total reduction ratio and the total number of frames of the multiple frames includes: if the total number of frames is N, the total reduction ratio is represented by S2min, then the calculation formula of the target reduction ratio of each frame is represented by:

[0034] S2i=1.0+(i-1)(S2min-1.0) / (N-1)

[0035] Wherein, S2min is a number less than 1 and greater than 0, N is an integer greater than 1, i represents the number of any frame in N frames, and i is an integer greater than or equal to 1 and less than or equal to N.

[0036] Some embodiments of the present application can evenly distribute the total reduction ratio of the background area in each frame through the formula of the target reduction ratio, thereby ensuring the continuity of the background area reduction process between the previous and next frames during video display, thereby obtaining a better display effect.

[0037] In some embodiments, the fusing of each pair of the enlarged sub-image and the reduced sub-image to obtain multiple frames of display images includes: confirming the overlapping area between the enlarged sub-image of any frame and the reduced sub-image included in any frame on the fused image; deleting the pixel values ​​located in the overlapping area on the reduced sub-image to obtain non-overlapping reduced sub-images; and fusing the enlarged sub-image of any frame and the non-overlapping reduced sub-image to obtain the display image of any frame.

[0038] In some embodiments of the present application, in order to avoid occlusion of the background area due to enlargement of the subject during image fusion, it is necessary to eliminate the overlapping area, that is, the background image needs to be processed before fusion. Because when the subject is enlarged and the background image is reduced, the subject part will cover a part of the background area, and this area needs to be eliminated, otherwise the fusion will be chaotic.

[0039] In some embodiments, a Poisson algorithm is used to fuse the enlarged sub-image of any frame image and the non-overlapping reduced sub-image.

[0040] In some embodiments, the fusing of the enlarged sub-image and the non-overlapping reduced sub-image of any frame image to obtain the display image of any frame image includes: fusing the enlarged sub-image and the non-overlapping reduced sub-image of any frame according to a Poisson algorithm to obtain a fused full-size image corresponding to any frame; and cropping the fused full-size image according to a size displayed on a specified interface to obtain a display image corresponding to any frame.

[0041] In some embodiments of the present application, the fused images are full-size images, so in order to display them on a display screen, each frame of the full-size image needs to be cropped to obtain a display image of each frame.

[0042] In some embodiments, after obtaining the multiple frames of display images, the sliding zoom method further includes: synthesizing the multiple frames of display images into a video and outputting the video.

[0043] In a second aspect, some embodiments of the present application provide a sliding zoom device based on a depth camera, the sliding zoom device comprising: a scene map and depth map acquisition module, configured to acquire a scene map and a depth map; a calibrated depth map acquisition module, configured to calibrate the depth map according to calibration data and the scene map to obtain a calibrated depth map; a main area and background area division module, configured to divide the scene map into a main area and a background area according to the calibrated depth map; a multi-frame display image acquisition module, configured to acquire pairs of enlarged sub-images of the main area and reduced sub-images of the background area, and fuse each pair of enlarged sub-images and reduced sub-images to obtain multi-frame display images.

[0044] In a third aspect, some embodiments of the present application provide a mobile terminal, comprising: a main camera configured to capture a scene graph; a depth camera configured to capture a depth map; a memory configured to store a computer program; a processor configured to execute the computer program based on the read scene graph and the depth image to implement the method described in any embodiment of the first aspect to obtain a plurality of frames of display images; and a display configured to display a video composed of the plurality of frames of display images.

[0045] In some embodiments of the present application, the mobile terminal also includes: an input unit configured to receive an input target focus point; wherein the processor is also configured to execute the computer program according to the scene graph, the depth map and the target focus point to obtain the multiple frames of display images.

[0046] In a fourth aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the method described in any embodiment of the first aspect above can be implemented.

[0047] In a fifth aspect, some embodiments of the present application provide an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement the method described in any embodiment of the first aspect when executing the program.

[0048] In a sixth aspect, some embodiments of the present application provide a computer program product, wherein the computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method described in any embodiment of the first aspect above can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0050] Figure 1 A demonstration diagram of the sliding zoom effect provided for the related technology;

[0051] Figure 2 A schematic diagram of the composition of a sliding zoom system based on a depth camera provided in an embodiment of the present application;

[0052] Figure 3 One of the flowcharts of the sliding zoom method based on the depth camera provided in the embodiment of the present application;

[0053] Figure 4 A flowchart for segmenting a subject area and a background area provided in an embodiment of the present application;

[0054] Figure 5 A flowchart for obtaining pixels of interest provided in an embodiment of the present application;

[0055] Figure 6-Figure 8 A schematic diagram of a full-size image taken by a mobile phone, an image of a size displayed on a display screen, and a depth image provided in an embodiment of the present application;

[0056] Fig. 9 The second flowchart of the sliding zoom method based on the depth camera provided in the embodiment of the present application;

[0057] Fig.10 A block diagram of a sliding zoom device based on a depth camera provided in an embodiment of the present application;

[0058] Fig.11 A block diagram of the mobile terminal provided in the embodiment of the present application;

[0059] Fig.12 A schematic diagram of the composition of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0061] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0062] At least in order to overcome the problems existing in the background technology, some embodiments of the present application provide a sliding zoom solution based on a depth map and a scene map, which divides the scene map into a main area and a background area according to the depth map, and gradually enlarges and displays the main area in multiple frames (i.e., obtains an enlarged sub-image of the main area in each frame in the multiple frames) and gradually reduces and displays the background area in multiple frames (i.e., obtains a reduced sub-image of the background area in each frame in the multiple frames) when the synthesized video is played, that is, a method for realizing sliding zoom. Compared with the related art that must rely on a complex shooting process, the embodiments of the present application can realize the sliding zoom method by performing relevant processing based on the scene map and the depth map of the photographed subject, and the process of shooting images is simpler and the data processing process is also simpler.

[0063] In some embodiments of the present application, the scene map and the depth map are images obtained by shooting the same scene at the same time or at different times (there is an object to be magnified on the image), wherein the scene map is an image of the scene obtained by an RGB camera, and the depth map is the depth information of the object in the scene, that is, the distance between the object in the scene and the depth camera or binocular camera can be obtained through the depth map. For example, some embodiments of the present application can be implemented using a mobile phone with a depth camera without the need for additional equipment. Specifically, a sliding zoom method can be implemented by shooting a group of images with a handheld mobile phone (that is, shooting a scene map and a depth map for the same scene at the same time). The shooting is simple and the requirements are not high. A shocking sliding zoom effect can be achieved, which can greatly enhance the user's multi-camera photo-taking experience.

[0064] Please see Figure 2 , Figure 2 A sliding zoom system based on a depth camera is provided in some embodiments of the present application, the system includes a terminal 10 and a server 20, wherein: Figure 2 The terminal 10 is configured to capture a scene image and a depth map, and send the captured scene image and depth map to the server 20. The server 20 is configured to obtain multiple frames of display images based on the scene image and the depth map, and then send the display images to the terminal 10. In this way, the terminal 10 can achieve a sliding zoom effect by playing multiple frames of display images.

[0065] The terminal 10 may include a mobile phone, a PAD, etc., and the specific type of the terminal is not limited in the embodiment of the present application. The terminal 10 and the server 20 may be interconnected via a wireless network or a wired network to achieve information transmission (e.g., depth map, scene map, each frame display image, etc.) between the two.

[0066] Figure 2 The terminal 10 includes a first camera 11 and a second camera 12. In some embodiments of the present application, the first camera 11 is a main RGB camera and the second camera 12 is a depth camera for capturing depth images. The first camera 11 is used to capture a scene image 21 of a target scene, and the second camera 12 is used to capture a depth image 22 of the target scene. In other embodiments of the present application, the first camera 11 and the second camera 12 are both RGB cameras, and the first camera 11 is used to capture a scene image 21 of a target scene. Fig.11 , and the depth map 22 is obtained by processing the images taken by the first camera 11 and the second camera 12 (i.e., the depth map is obtained by a binocular method). For how to obtain a depth map based on a binocular image, reference can be made to existing relevant literature, and no further details will be given here to avoid repetition.

[0067] That is to say, the depth map 22 required by the embodiment of the present application can be obtained by directly shooting with a depth camera, or it can be obtained by shooting an image with a binocular camera and then performing algorithm processing.

[0068] The depth camera or depth camera (i.e., RGBD camera) is also called a 3D camera because it can detect the depth of field distance of the shooting space, which is also the biggest difference from the ordinary camera (i.e., the first camera 11). The data obtained by the depth camera can determine the distance from the spatial point represented by each pixel in the depth map to the depth camera. The types of depth cameras include depth cameras based on structured light or depth cameras based on TOF solutions.

[0069] It is understandable that the pixel value of each pixel on the depth map 22 is the distance between the photographed point corresponding to the pixel and the depth camera, and the unit is a length unit such as meter, centimeter or millimeter. The first camera 11 and the second camera 12 have the same resolution, and the scene map 21 and the depth map 22 are obtained by using the first camera 11 and the second camera 12 to shoot the same scene at the same or adjacent time.

[0070] It should be noted that in some embodiments of the present application, the functions of the server 20 can also be performed by the terminal 10, that is, the terminal 10 will process the depth map and scene map obtained by shooting a certain scene to obtain the multi-frame display images required to achieve the sliding zoom effect.

[0071] For the convenience of description, the camera used to shoot scene Figure 21 will be referred to as the main camera in the following.

[0072] Combine the following Figure 3 Exemplary description by Figure 1 A sliding zoom method based on a depth camera is performed by the server 20 or the terminal 10.

[0073] like Figure 3 As shown, some embodiments of the present application provide a sliding zoom method based on a depth camera, the sliding zoom method comprising: S101, obtaining a scene map and a depth map. S102, calibrating the depth map according to calibration data and the scene map to obtain a calibrated depth map. S103, dividing the scene map into a main area and a background area according to the calibrated depth map. S104, obtaining multiple groups of pairs of enlarged sub-images of the main area and reduced sub-images of the background area, and fusing each pair of the enlarged sub-images and the reduced sub-images to obtain multiple frames of display images.

[0074] It should be noted that, in some embodiments, the scene graph of S101 is obtained by photographing the main camera, and the depth map of S101 is obtained by photographing the depth camera. Then, before executing S102, the sliding zoom method of some embodiments of the present application also includes: obtaining the conversion parameter calibration between the depth camera and the main camera to obtain the required calibration data. That is to say, some embodiments of the present application need to first correct the depth map based on the scene graph before segmenting the main area and the background area, so that the depth map corresponds to each pixel on the scene graph (that is, the same photographed point in space has the same pixel position on the depth map and the scene map), so as to facilitate the subsequent segmentation of the main area and the background area on the scene graph according to the depth map.

[0075] The following is an exemplary description of the relevant steps involved in the above process.

[0076] In order to ensure that the subject area obtained by cutting includes the complete subject object to be enlarged, in some embodiments of the present application, S103 includes: converting the calibrated depth map into a disparity map; obtaining a reference disparity value from the disparity map of the target focus point, and obtaining a reference depth value from the calibrated depth map; if it is determined that the reference depth value is less than a set distance threshold, all pixel points that are greater than the reference disparity value by a first set threshold are confirmed as belonging to the pixel area where the subject is located, and the position corresponding to the pixel area where the subject is located on the scene graph is confirmed as the subject area, and the area other than the subject area on the scene graph is the background area; or, if it is confirmed that the reference depth value is greater than or equal to the set distance threshold, all pixel points that are greater than the reference disparity value by a second set threshold are confirmed as belonging to the pixel area where the subject is located, and the position corresponding to the pixel area where the subject is located on the scene graph is confirmed as the subject area, and the area other than the subject area on the scene graph is the background area.

[0077] The unit of each pixel on the disparity map is pixel (pixel), and the pixel value of each pixel on the disparity map is the disparity value, and these disparity values ​​represent the coordinate difference of a point in space imaged on two parallel binocular images.

[0078] Some embodiments of the present application can focus the content of interest on the main area rather than the background area by setting a target focus point.

[0079] In order to pre-screen out scene images that cannot separate the background area and the main area (i.e., exclude images whose background is too close or whose main area and background cannot be clearly distinguished), in some embodiments of the present application, before converting the calibrated depth map into a disparity map in S103, the sliding zoom method also includes: obtaining a maximum depth value and a minimum depth value from the calibrated depth map or the depth map; confirming that the maximum depth value is greater than or equal to a set distance, and confirming that the difference between the maximum depth value and the minimum depth value is greater than or equal to a set distance difference.

[0080] Combine the following Figure 4 The above-mentioned process of S103 is described as an example.

[0081] like Figure 4 As shown, the process of dividing the scene graph into a main area and a background area according to the calibrated depth map exemplarily includes the following steps.

[0082] S1031, obtaining a calibrated depth map.

[0083] S1032, determine whether dmax<3000mm and dmax-dmin<1000mm are established. If so, execute S1038 to prompt the user that the photo does not meet the requirements and needs to be re-photographed. If not, execute S1033.

[0084] It should be noted that dmax is the maximum depth value obtained from the calibrated depth map, and dmin is the minimum depth value obtained from the calibrated depth map. As an example Figure 4 The set distance in is 3000mm, and the set distance difference is 1000mm. It is understandable that those skilled in the art can select appropriate set distance and set distance difference according to actual conditions, and the embodiments of the present application do not limit the specific values ​​of these parameters. Figure 4 The condition of dmax<3000mm can filter out scenes with backgrounds that are too close. Figure 4 The condition of dmax-dmin<1000mm can filter out scene images that cannot clearly distinguish the subject from the background. It is understandable that if these distance conditions are not met, the user can be prompted to reshoot the scene image and / or the depth image.

[0085] S1033, converting the calibrated depth map into a disparity map.

[0086] S1034, obtaining the disparity value d0 (i.e., reference disparity value) and the depth value d (i.e., reference depth value) of the target focus point (e.g., the diagonal point set in a preset manner). That is, the pixel value of the pixel position where the focus point is located is read from the disparity map to obtain the preset diagonal point disparity value d0, and the pixel value of the pixel position where the target focus point is located is read from the depth map to obtain the target focus point depth value d.

[0087] S1035, determine whether the target focus depth value d is less than 3000mm, if so, execute S1036, otherwise execute S1037. By executing this judgment, the subsequent steps can be divided into two situations to ensure the completeness of the main object included in the main area.

[0088] Figure 4 The set distance threshold is 3000mm, that is, Figure 4 It is necessary to confirm whether the reference depth value is less than the set distance threshold of 3000 mm. The embodiment of the present application does not limit the size of the set distance threshold. Those skilled in the art can select a suitable size of the set distance threshold according to actual needs and application scenarios.

[0089] S1036, the disparity map is divided into main pixel areas where -d0>=T1 (corresponding to the first set threshold, for example, the value of T1 is "-2pixel"), and the rest are background pixel areas. Figure 4In the above example, all pixels whose number of pixels is greater than the reference disparity value d0 by two pixels are confirmed as belonging to the pixel area where the subject is located, and the remaining area is confirmed as belonging to the background pixel area. Figure 4 The number of pixels is set to 2, but the embodiment of the present application does not limit the specific value of the set number of pixels. Those skilled in the art can select a suitable specific value of the set number of pixels based on factors such as application scenarios.

[0090] S1037, the disparity map -d0>=T2 (corresponding to the second set threshold, for example, the value of T2 is "2pixel") is divided into main pixel areas, and the rest are background pixel areas. Figure 4 In , all pixels that are two pixels larger than the reference disparity value d0 are identified as belonging to the pixel area where the subject is located, and the remaining area is the background area. Figure 4 The number of pixels is set to 2, but the embodiment of the present application does not limit the specific value of the set number of pixels. Those skilled in the art can select a suitable specific value of the set number of pixels based on factors such as application scenarios.

[0091] It should be noted that Figure 4 The significance of setting steps S1306 and S1037 is that this is because the subject area has depth, that is, there is a parallax interval, and the size of the parallax value decreases as the distance increases, so the segmented subject area needs to be larger than the parallax value of the target focus point (the parallax value of the target focus point is the parallax value in front of the subject, and the depth of the subject needs to be considered). This is because the subject is in front of the background and the distance is smaller than the background, so the parallax value of the subject area is greater than the reference parallax value.

[0092] It is not difficult to understand that by executing the above S103, the scene graph is divided into a main area and a background area. The implementation process of S104 is exemplarily described below.

[0093] In order to achieve a sliding zoom effect in which the main area is gradually enlarged and the background area is gradually reduced during video playback, in some embodiments of the present application, S104 includes: obtaining the total magnification of the main area obtained by executing S103, and obtaining the total reduction ratio of the background area obtained by executing S103; determining the target magnification of each frame according to the total magnification and the total number of frames included in the composite video, and determining the target reduction ratio of each frame according to the total reduction ratio and the total number of frames included in the composite video; obtaining the enlarged sub-image of each frame according to the obtained target magnification of each frame and the main area obtained by S103, and obtaining the reduced sub-image of each frame according to the obtained target reduction ratio of each frame and the background area obtained by S103. Afterwards, the enlarged sub-image and the reduced sub-image of each frame (or each pair) are fused to obtain the multi-frame display image. It can be understood that the total number of frames of the composite video is equal to the number of frames of the multi-frame display image.

[0094] That is to say, some embodiments of the present application distribute the solved total magnification ratio for the main area in each frame, and distribute the obtained total reduction ratio in each frame, thereby achieving a sliding zoom effect in which the main area is gradually enlarged and the background area is gradually reduced during video playback. Compared with the processing algorithms of the related technologies, the calculation amount is less and the calculation speed is faster.

[0095] It should be noted that in some embodiments of the present application, the total magnification ratio and the total reduction ratio can be any two pre-set values, for example, the total magnification ratio is 1.3 and the total reduction ratio is 0.8. In some embodiments of the present application, in order to avoid areas where image data is missing when reducing the image, a more accurate total reduction ratio can be determined based on the scene graph (i.e., the full-size image captured by the main camera) and the display size of the display screen.

[0096] For example, in order to obtain a more ideal total reduction ratio, in some embodiments of the present application, S104 involves obtaining the total magnification ratio of the main area and the total reduction ratio of the background area, including: obtaining the total reduction ratio according to the size of the full-size image and the size of the displayed image, wherein the scene map and the depth map are both the full-size image. The following is an example to illustrate the meaning of the full-size image and the displayed image. For example, the FOV of the mobile phone interface display image (i.e., the size of the displayed image) accounts for 75% of the real FOV (i.e., the full-size image size, determined by the shooting angle of the main camera). It can be understood that the display image size is generally smaller than the full-size image size. For example, in some embodiments of the present application, the total reduction ratio is obtained based on the size of the full-size image and the size displayed on the specified interface, including: calculating the distance s1_i from the pixel of interest to the four boundary lines of the displayed image, and calculating the distance s2_i from the pixel of interest to the four boundary lines of the full-size image, where the value of i is an integer greater than or equal to 1 and less than or equal to 4; calculating the ratio according to the following formula to obtain four ratios: r_i=s1_i / s2_i; and selecting the maximum ratio from the four ratios as the total reduction ratio.

[0097] That is to say, some embodiments of the present application determine the minimum value that can be reduced by calculating the distance between the pixel of interest and the four sides corresponding to the display size and the shooting size, and use the minimum value that can be reduced as the total reduction ratio of the background area. This can ensure that there is no lack of display data during the reduction of the background area, thereby obtaining a better display effect.

[0098] In order to quantize the target quantization magnification of each frame according to the total magnification, some embodiments of the present application may evenly distribute the total magnification of the main area in each frame. For example, the process of determining the target magnification of each frame according to the total magnification and the total number of frames of the synthesized video involved in S104 exemplarily includes: if the total number of frames set for the playback video is N, the total magnification is represented by S1max, and the calculation formula of the target magnification of each frame is represented by:

[0099] S1i=1.0+(i-1)*(S1max-1.0) / (N-1)

[0100] Wherein, S1max is a number greater than 1, N is an integer greater than 1, i represents the number of any frame in N frames and i is an integer less than or equal to N. In some embodiments of the present application, the total magnification of the main body area can be evenly distributed in each frame through the formula of the target magnification, so as to ensure the continuity of the magnification process of the main body area between the previous and next frames during video display, and obtain a better display effect.

[0101] In order to quantize the target quantization ratio of each frame according to the total reduction ratio, some embodiments of the present application may evenly distribute the total reduction ratio of the background area in each frame. For example, the process of determining the target reduction ratio of each frame according to the total reduction ratio and the total number of frames of the synthesized video in S104 exemplarily includes: if the total number of frames is N, and the total reduction ratio is represented by S2min, then the calculation formula of the target reduction ratio of each frame is represented by:

[0102] S2i=1.0+(i-1)(S2min-1.0) / (N-1)

[0103] Wherein, S2min is a number less than 1 and greater than 0, N is an integer greater than 1, i represents the number of any frame in N frames, and i is an integer greater than or equal to 1 and less than or equal to N. In some embodiments of the present application, the total reduction ratio of the background area can be evenly distributed in each frame through the formula of the target reduction ratio, so as to ensure the continuity of the background area reduction process between the previous and next frames when displaying the video, and obtain a better display effect.

[0104] In order to enhance the visual sensory effect and avoid the problem of poor display effect caused by magnification of non-subjects or magnification of subject edges, in some embodiments of the present application, before executing S104, the sliding zoom method also includes: obtaining pixels of interest from the scene graph, and the corresponding S104 involves obtaining the enlarged sub-image of each frame according to the target magnification of each frame and the subject area, including: constructing an enlargement conversion matrix according to the target magnification of each frame and the pixels of interest; obtaining the enlarged sub-image of each frame according to the enlargement conversion matrix and the subject area; the corresponding S104 involves obtaining the reduced sub-image of each frame according to the target reduction ratio of each frame and the background area, including: constructing a reduction conversion matrix according to the target reduction ratio of each frame and the pixels of interest; obtaining the reduced sub-image of each frame according to the reduction conversion matrix and the background area.

[0105] That is to say, some embodiments of the present application can effectively improve the visual sensory effect by introducing pixels of interest and enlarging the main area and reducing the background area with the pixels of interest as the center during post-processing (i.e., obtaining the enlarged sub-image and the reduced sub-image of each frame).

[0106] The following is an example of a method for obtaining pixels of interest.

[0107] In some embodiments of the present application, the above-mentioned obtaining of the pixel of interest from the scene graph includes: calculating the center of gravity corresponding to the main body area according to the coordinate values ​​of each pixel included in the main body area; confirming that the target focus point is within the main body area, and obtaining the pixel of interest according to the following formula:

[0108] c=w1*p+w2*p1

[0109] Among them, p1 represents the coordinate value of the center of gravity, p represents the coordinate value of the target focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1. In some embodiments of the present application, when the target focus point is located on the main area, the point of interest obtained by using the previous formula can be closer to the target focus point and closer to the center of gravity of the main area, which can ensure a better visual effect when the main area is enlarged.

[0110] In some other embodiments of the present application, the above-mentioned obtaining the pixel of interest from the scene graph includes: calculating the center of gravity of the main body area according to the coordinate values ​​of each pixel corresponding to the main body area; confirming that the target focus point is not in the main body area, taking the center of the scene graph as the new focus point p, and obtaining the initial pixel of interest according to the following formula:

[0111] c = w2*p+w1*p1;

[0112] If it is confirmed that the initial pixel of interest c is not inside the main area, then perform at least one more iteration of the following formula c=w2*c+w1*p1 until the set number of iterations is reached or it is confirmed that the initial point of interest c is inside the main area, then the initial point of interest c is used as the pixel of interest; wherein p1 represents the coordinate value with the center of gravity, p represents the coordinate value of the new focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1. In some embodiments of the present application, when the target focus point is not in the main area, the pixel of interest is obtained by the image center and the center of gravity of the main area, and the pixel of interest is ensured to fall on the main area as much as possible, so as to ensure a better visual effect when the main area is enlarged.

[0113] Combine the following Figure 5 The process of obtaining pixels of interest is exemplified.

[0114] like Figure 5 As shown, the process of obtaining the pixel of interest c includes:

[0115] S1041, calculating the center of gravity p1 of the main body area.

[0116] The center of gravity of the main area is calculated according to the coordinate values ​​of each pixel point included in the main area. For example, the average value of the coordinate values ​​of each pixel point included in the main area is calculated as the center of gravity p1 of the main area. It can be understood that the unit of the center of gravity p1 of the main area is pixel and represents the pixel position where the center of gravity is located. The coordinate values ​​of each pixel point in the main area are in the pixel coordinate system, and the origin of the pixel coordinate system is in the upper left corner of the scene graph.

[0117] S1042, determine whether the target focus point p is within the subject area, if so, execute S1034, otherwise point to S1044.

[0118] S1043, calculate the pixel of interest using the following formula: c=w1*p+w2*p1, pointing to S1036. It should be noted that for the specific meaning of this parameter, please refer to the above description.

[0119] S1034, the initial pixel of interest is calculated using the following formula: c=w2*c+w1*p1. It should be noted that the specific meaning of this parameter can be found in the above description.

[0120] S1035, determine whether the initial pixel of interest c is within the main area, if so, take the initial pixel of interest as the pixel of interest and output it, if not, replace the preset diagonal point (i.e., the pixel coordinate value of the preset diagonal point) with the value of the initial pixel of interest and repeatedly use the formula involved in S1034 to calculate the pixel of interest, and after the iteration is completed, use the final initial pixel of interest as the pixel of interest.

[0121] S1036, output the pixel of interest c.

[0122] In order to obtain the enlarged sub-image and the reduced sub-image of each frame according to the obtained pixel of interest, the target magnification and the target reduction ratio, in some embodiments of the present application, the process of obtaining the enlarged sub-image of the main area and the reduced sub-image of the background area included in each frame of the multiple frames of images involved in S104 exemplarily includes: constructing an image transformation matrix corresponding to each frame according to the obtained pixel of interest c(Cx, Cy), the target magnification and the target reduction ratio of each frame; obtaining the enlarged sub-image of each frame according to the transformation matrix of the main area and each frame image obtained in S103, and obtaining the reduced sub-image of each frame image according to the transformation matrix of the background area and each frame image obtained in S103.

[0123] That is to say, some embodiments of the present application construct an image transformation matrix for each frame through pixels of interest, target magnification and target reduction ratio, and obtain an enlarged sub-image of each frame through the image transformation matrix and the main area, and obtain a reduced sub-image of each frame through the image transformation matrix and the background area, so that when the video is displayed, each frame image is scaled with the pixel of interest as the center, thereby enhancing the user's visual experience.

[0124] In order to avoid occlusion of the background area due to enlargement of the subject during image fusion, it is necessary to eliminate the overlapping area. In some embodiments of the present application, the process of fusing each pair of enlarged sub-images and reduced sub-images involved in S104 to obtain the multiple frames of display images exemplarily includes: confirming the overlapping area of ​​the enlarged sub-image of any frame and the reduced sub-image included in any frame on the fused image; deleting the pixel values ​​located in the overlapping area on the reduced sub-image to obtain a non-overlapping reduced sub-image; and fusing the enlarged sub-image of any frame with the non-overlapping reduced sub-image to obtain the display image of any frame.

[0125] That is to say, in order to avoid occlusion of the background area caused by enlarging the subject, some embodiments of the present application need to eliminate the overlapping area during image fusion. That is, the background image needs to be processed before fusion. Because when the subject is enlarged and the background image is reduced, the subject part will cover a part of the background area. This area needs to be eliminated, otherwise the fusion will be chaotic.

[0126] For example, in some embodiments of the present application, Poisson fusion is used to fuse the enlarged sub-image and the non-overlapping reduced sub-image of any frame image.

[0127] It is understandable that the fused image is a full-size image, so in order to display it on the display screen, the full-size image of each frame needs to be cropped to obtain the display image of each frame. For example, in some embodiments of the present application, the process of fusing the enlarged sub-image of any frame image and the non-overlapping reduced sub-image to obtain the display image of any frame image involved in S104 exemplarily includes: fusing the enlarged sub-image of any frame and the non-overlapping reduced sub-image according to the Poisson algorithm to obtain a fused full-size image corresponding to any frame; cropping the fused full-size image according to the size displayed on the specified interface to obtain a display image corresponding to any frame.

[0128] It should be noted that, in order to demonstrate the sliding zoom effect, in some embodiments of the present application, after obtaining multiple frames of display images as described in S104, the sliding zoom method further includes: synthesizing the multiple frames of display images into a video and outputting the video.

[0129] The following takes a mobile phone with a depth camera and a main RGB camera module as an example to illustrate the process of implementing a sliding zoom method based on a depth camera by the mobile phone.

[0130] First, calibrate the depth camera and main camera on the mobile phone and calibrate the conversion parameters between the depth camera and the main camera.

[0131] Then, the depth camera and the main camera are used to collect scene images and depth images with obvious subjects, and the depth image and the main camera scene image are aligned using the calibration data.

[0132] Finally, algorithm processing is performed to separate the foreground and background according to the depth map, with the foreground enlarged and the background reduced, to achieve a sliding zoom effect in which the subject moves forward and the background becomes farther away.

[0133] It can be understood that in order to implement the sliding zoom method based on the depth camera of the present application, a mobile phone with a depth camera can be directly used to implement it, without the need for additional equipment. Specifically, a group of images (i.e., depth map and scene map) can be taken with the handheld mobile phone to form an image. The shooting is simple and the requirements are not high. The shocking sliding zoom effect can be achieved, which can greatly enhance the user's multi-camera photography experience.

[0134] Combine the following Figure 6-Figure 9 The sliding zoom method based on the depth camera implemented by a mobile phone is described as an example. Figure 6 and Figure 7 The faces are blurred in the figures. It is understandable that this processing does not affect the explanation of the technical solutions by these pictures.

[0135] The first step is interface design. The main camera of the image capture module on the mobile phone can choose a normal lens of about 80° and an RGBD lens.

[0136] Since the embodiments of the present application need to maintain an extra FOV (Field of view) for display, the FOV of the displayed image of the mobile phone interface accounts for 75% of the real FOV (i.e., the full-size image, i.e., the scene graph). The embodiments of the present application do not limit the specific ratio of the displayed image to the full-size image, for example, the FOV of the displayed image accounts for 80% of the real FOV.

[0137] like Figure 6 is a full-size image taken with the main camera, where the area enclosed by the frame is Figure 7 The display panel can display the display image, that is, the FOV of the image displayed on the mobile phone interface is smaller than the full-size image actually captured (the total reduction ratio of multiple frames is then determined based on the four boundary lines of the two images). This ensures that the image lacks display data during the reduction process, resulting in a poor sliding zoom effect. Figure 8 It is the depth map taken by RGBD lens.

[0138] It should be noted that Figure 6 , Figure 7 and Figure 8 These images are only used to illustrate the relationship between the various figures, and the actual size of the captured image cannot be determined based on the size of these images. Figure 8 The depth map and Figure 6 The scene graphs belong to images of the same size, Figure 6 The scene graph belongs to the full-size image.

[0139] The second step is data collection.

[0140] As attached Figure 7 As shown, when the user clicks the camera button on the phone, the platform will read the data required by the algorithm, which includes: calibration data, main image (full FOV) or scene image, depth image (full FOV), screen RECT (display FOV size), set target focus point (or preset focus point). It can be understood that the set target focus point can focus the content of interest on the subject (the subject area includes the subject, which is the magnified object) instead of focusing the content of interest on the background area or anywhere.

[0141] The third step is depth map alignment.

[0142] Since the RGBD camera and the main camera on the mobile phone do not overlap, the depth map obtained from the platform and the main image do not meet the pixel position alignment. Therefore, it is necessary to calibrate the main camera and the RGBD camera to obtain calibration data, and use the calibration data to align the depth map to the main image to obtain a calibrated depth map. For example, the specific calibration method can refer to the opencv function cv::rgbd::registerDepth and its principle. This method belongs to the prior art, so it will not be described in detail here.

[0143] The above process executes S201 and S202 as described in 9, wherein S201 includes: setting the screen display size, obtaining the main image, obtaining the depth map and obtaining the calibration data; S102 includes: aligning the depth map with the main image (or called the main image or scene image) according to the calibration data to obtain the calibrated depth map. Fig. 9 The example illustrates that a mobile phone executes a sliding zoom method based on a depth camera.

[0144] The fourth step is to segment the subject and correspond to Fig. 9 In S203, the main image is segmented according to the calibrated depth map to obtain a main area and a background area.

[0145] First, the input calibrated depth map is converted into a disparity map according to the conversion formula between the depth map and the disparity map x=b*f / d.

[0146] Among them, x represents the parallax, b represents the base distance (the distance between the optical centers of the two cameras), f represents the focal length (i.e. the focal length of the depth camera), and d represents the distance (i.e. the distance between each point in the captured scene and the camera).

[0147] The pixel value of each point on the depth map is the distance between the point in the scene and the image capture unit, and the unit is the length unit, while the unit of the disparity map is the pixel. It can be understood that since the pixels of the depth map and the main image are aligned before segmentation (i.e., the depth map is calibrated), the segmented disparity map can correspond to the segmented main image.

[0148] Then, according to the pixel position p of the target focus point, the reference disparity value d0 is obtained from the disparity map, and the pixel area with disparity -d0>=-2pixel on the disparity map is divided into the main area, and the rest is the background area. If the reference depth value of the target focus point d>3m (the 3m threshold is adjustable according to the actual module situation), it is considered that the focus is on the background. At this time, it is determined whether there is a subject on the depth map with d<3m. If there is no subject, the user is prompted to retake the image. If there is a subject, the disparity map is divided into the disparity map with disparity -d0>=2pixel as the main area, and the rest is the background area. The area corresponding to the main area on the scene map can be represented by number 1, and the area corresponding to the background area can be represented by number 2.

[0149] The fifth step is center calculation, which corresponds to Fig. 9 S204 obtains the pixel of interest according to the center of gravity of the subject area.

[0150] The significance of obtaining the region of interest c is that in post-processing, scaling with the pixel of interest as the center will provide better visual perception.

[0151] In some embodiments of the present application, the calculation method of the pixel of interest exemplarily includes: first calculating the centroid p1 of the main area, and then performing weighted calculation according to the position of the target focus point (i.e., the focus point pre-input in this example), if the focus point p is on the main area, then the pixel of interest c=w1*p+w2*p1, if the target focus point p is on the background area, then the focus point p is redefined to be at the center of the image, then the pixel of interest c=w2*p+w1*p1, for example, w1=0.7, w2=0.3, if the point of interest c is not inside the main body, then a secondary calculation c=w2*c+w1*p1 is performed. For the specific process, please refer to the relevant description and drawings above.

[0152] The sixth step is scale calculation, which includes Fig. 9 S205, that is, obtaining the total reduction ratio of the background area according to the position of the pixel of interest.

[0153] The total magnification of the main area can be any integer greater than 1. For example, in some embodiments, the total magnification of the main area can be set to 1.2 (adjustable). The embodiments of the present application do not limit the total magnification of the main area, but it is understood that if the magnification is too high, the entire display panel will not be able to display the completed main area.

[0154] The total reduction ratio of the background area is calculated according to the FOV. The ratio of the four sides is calculated based on the pixel of interest c as the center (first calculate the distance s1_i, s2_i from the point to the display size and the four sides of the full size, then calculate the ratio r_i = s1_i / s2_i), and finally select the maximum ratio as the total reduction ratio of the background area. The specific process can be referred to the above description.

[0155] It should be noted that in some embodiments of the present application, the total reduction ratio may also be set to a value, for example, the total reduction ratio is 0.85.

[0156] The seventh step is image scaling. This step corresponds to Fig. 9 S206 and S207, wherein S206 includes: calculating the target magnification ratio and the target reduction ratio of each frame in N frames, wherein N is the number of frames included in the played video, and N is an integer greater than 1. S207 includes: taking the pixel of interest as the center, obtaining an enlarged sub-image according to the calculated target magnification ratio and the main area of ​​each frame, and obtaining a reduced sub-image according to the calculated target reduction ratio and the background area of ​​each frame.

[0157] For example, set the number of images of the composite video (i.e., the played video) to N frames, and calculate the target zoom ratio of each frame of the image according to the scale value calculated in the sixth step (subject magnification S1max=1.2, background reduction ratio S2min=r_i<1.0). The target zoom ratio includes the target magnification ratio and the target reduction ratio. The calculation formula is: S1i=1.0+(i-1)*(S1max-1.0) / (N-1), S2i=1.0+(i-1)(S2min-1.0) / (N-1), i=1…N.

[0158] According to the pixel of interest c(Cx,Cy) found in the fifth step, the magnification transformation matrix of each frame is constructed as follows: The constructed reduction transformation matrix of each frame is The warp function is used to transform the main area and the background area to obtain the enlarged sub-image and the reduced sub-image included in each frame.

[0159] Step 8: Image fusion, corresponding to Fig. 9 S208 fuses the enlarged sub-image and the reduced sub-image of each frame to obtain the multi-frame display image.

[0160] The main image of each frame (ie, the enlarged sub-image) and the background image of each frame (ie, the reduced sub-image) are fused, and the fusion method may consider Poisson fusion.

[0161] It should be noted that the background image needs to be processed before fusion, because after the main area is enlarged and the background area is reduced, part of the main area will cover part of the background area. This area needs to be removed, otherwise the fusion will be chaotic.

[0162] The ninth step is video synthesis, corresponding to Fig. 9 Video synthesis of the S209.

[0163] The fused image in step 8 is cropped from the full-size image (i.e., on the scene graph) using the FOV size displayed on the interface, and synthesized into a video output.

[0164] Please refer to Fig.10 , Fig.10 The sliding zoom device based on the depth camera provided in the embodiment of the present application is shown. It should be understood that the device is similar to the above-mentioned Figure 3 Corresponding to the method embodiment, each step involved in the above method embodiment can be executed. The specific functions of the device can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in the memory in the form of software or firmware or solidified in the operating system of the device. The sliding zoom device based on the depth camera includes: a scene graph and depth map acquisition module 101, a calibrated depth map acquisition module 102, a main area and background area division module 103, and a multi-frame display image acquisition module 104.

[0165] The scene graph and depth map acquisition module 101 is configured to acquire a scene graph and a depth map, for example, the scene graph and the depth map include a magnified main object and a background.

[0166] The calibrated depth map acquisition module 102 is configured to calibrate the depth map according to the calibration data and the scene graph to obtain a calibrated depth map.

[0167] The main area and background area division module 103 is configured to divide the scene graph into a main area and a background area according to the calibrated depth map.

[0168] The multi-frame display image acquisition module 104 is configured to acquire multiple groups of pairs of enlarged sub-images of the subject area and reduced sub-images of the background area, and fuse each pair of enlarged sub-images and reduced sub-images to obtain multiple frames of display images.

[0169] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described device can refer to the aforementioned Figure 3The corresponding process in will not be elaborated here.

[0170] like Fig.11 As shown, some embodiments of the present application provide a mobile terminal 300, which includes: a main camera 310, configured to shoot a scene graph, that is, to shoot a scene graph containing a subject object; a depth camera 320, configured to shoot a depth map, that is, to shoot a depth map containing the same subject object; a memory 330, configured to store a computer program; a processor 340, configured to execute the computer program according to the read scene graph and the depth image to implement the following Figure 3 The method described above is used to obtain multiple frames of display images; the display 350 is configured to display a video composed of the multiple frames of display images.

[0171] In some embodiments of the present application, the mobile terminal also includes: an input unit configured to receive an input target focus point, wherein the processor is further configured to execute the computer program according to the scene graph, the depth map and the target focus point to obtain the multiple frames of display images.

[0172] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the device described above can refer to the aforementioned Figure 3 The corresponding process in will not be elaborated here.

[0173] Some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which can implement the above-mentioned Figure 3 The method described in the method embodiment.

[0174] like Fig.12 As shown, some embodiments of the present application provide an electronic device 500, including a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520. When the processor 520 reads the program through a bus 530 and executes the program, the following can be achieved: Figure 3 The method described in the related embodiments.

[0175] Fig.11 or Fig.12 A processor can process digital signals and can include various computing structures, such as a complex instruction set computer structure, a reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, the processor can be a microprocessor.

[0176] Fig.11 and Fig.12The memory of the present disclosure may be used to store instructions executed by the processor or data related to the execution of instructions. These instructions and / or data may include code to implement some or all functions of one or more modules described in the embodiments of the present disclosure. The processor of the present disclosure embodiment may be used to execute the instructions in the memory to implement Figure 3 The memory includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory or other memory known to those skilled in the art.

[0177] Some embodiments of the present application provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the method described in any embodiment of the first aspect above.

[0178] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0179] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0180] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0181] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0182] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0183] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A sliding zoom method based on a depth camera, It is characterized in that The sliding zoom method comprises: Get the scene graph and depth map; Calibrate the depth map according to the calibration data and the scene graph to obtain a calibrated depth map; Dividing the scene graph into a main area and a background area according to the calibrated depth map; Acquire multiple groups of pairs of enlarged sub-images of the subject area and reduced sub-images of the background area, and fuse each pair of enlarged sub-images and reduced sub-images to obtain multiple frames of display images; The step of obtaining a plurality of pairs of enlarged sub-images of the subject area and reduced sub-images of the background area includes: Acquire the total magnification of the subject area and the total reduction ratio of the background area; Determine a target magnification of each frame according to the total magnification and the total number of frames of the composite video, and determine a target reduction ratio of each frame according to the total reduction ratio and the total number of frames, wherein the total number of frames of the composite video is equal to the number of frames of the multiple display images; An enlarged sub-image of each frame is obtained according to the target enlargement ratio of each frame and the main area, and a reduced sub-image of each frame is obtained according to the target reduction ratio of each frame and the background area.

2. The sliding zoom method according to claim 1, It is characterized in that The scene image is obtained by photographing with a main camera, and the depth image is obtained by photographing with a depth camera; in, Before calibrating the depth map according to the scene map, the sliding zoom method further includes: Obtain a conversion parameter calibration between the depth camera and the main camera to obtain the calibration data.

3. The sliding zoom method according to any one of claims 1 to 2, It is characterized in that The step of dividing the scene graph into a main area and a background area according to the calibrated depth map comprises: Converting the calibrated depth map into a disparity map; Acquire a disparity value of a target focus point from the disparity map to obtain a reference disparity value, and acquire a depth value of the target focus point from the calibrated depth map to obtain a reference depth value; If it is determined that the reference depth value is less than the set distance threshold, all pixel points that are greater than the reference disparity value by a first set threshold are confirmed as belonging to the pixel area where the subject is located, and the position corresponding to the pixel area where the subject is located on the scene graph is confirmed as the subject area, and the area other than the subject area on the scene graph is the background area; or, if it is confirmed that the reference depth value is greater than or equal to the set distance threshold, all pixel points that are greater than the reference disparity value by a second set threshold are confirmed as belonging to the pixel area where the subject is located, and the position corresponding to the pixel area where the subject is located on the scene graph is confirmed as the subject area, and the area other than the subject area on the scene graph is the background area.

4. The sliding zoom method according to claim 3, It is characterized in that Before converting the calibrated depth map into a disparity map, the sliding zoom method further includes: Acquire a maximum depth value and a minimum depth value from the calibrated depth map or the depth map; It is confirmed that the maximum depth value is greater than or equal to a set distance, and it is confirmed that a difference between the maximum depth value and the minimum depth value is greater than or equal to a set distance difference.

5. The sliding zoom method according to claim 1, It is characterized in that Before acquiring a plurality of pairs of enlarged sub-images of the subject area and reduced sub-images of the background area, the sliding zoom method further includes: Obtaining pixels of interest from the scene graph; in, The step of obtaining the enlarged sub-image of each frame according to the target magnification of each frame and the main area comprises: Constructing a magnification conversion matrix according to the target magnification of each frame and the pixel of interest; Obtaining an enlarged sub-image of each frame according to the enlargement conversion matrix and the main area; The step of obtaining a reduced sub-image of each frame according to the target reduction ratio of each frame and the background area includes: Constructing a reduction transformation matrix according to the target reduction ratio of each frame and the pixel of interest; A reduced sub-image of each frame is obtained according to the reduced conversion matrix and the background area.

6. The sliding zoom method according to claim 5, It is characterized in that The obtaining of pixels of interest from the scene graph comprises: Calculating the center of gravity of the main body area according to the coordinate values ​​of each pixel point corresponding to the main body area; Confirm that the target focus point is within the subject area, and then obtain the pixel of interest according to the following formula: c=w1*p+w2*p1 Among them, p1 represents the center of gravity, p represents the coordinate value of the target focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1.

7. The sliding zoom method according to claim 5, It is characterized in that The acquiring of pixels of interest from the scene graph comprises: Calculating the center of gravity of the main body area according to the coordinate values ​​of each pixel point corresponding to the main body area; If it is confirmed that the target focus point is not in the subject area, the center of the scene graph is used as the new focus point p, and the initial pixel of interest is obtained according to the following formula: c = w2*p+w1*p1; If it is confirmed that the initial pixel of interest c is not inside the main body area, perform at least one more iteration of the following formula c=w2*c+w1*p1 until the set number of iterations is reached or it is confirmed that the initial pixel of interest c is inside the main body area, then use the initial pixel of interest c as the pixel of interest; Among them, p1 represents the coordinate value of the center of gravity, p represents the coordinate value of the new focus point on the scene graph, the value of the first weight coefficient w1 is greater than the value of the second weight coefficient w2, and the sum of the first weight coefficient and the second weight coefficient is 1.

8. The sliding zoom method according to any one of claims 1, 2, 4 to 7, It is characterized in that The obtaining of the total reduction ratio of the background area includes: The total reduction ratio is obtained according to the size of a full-size image and the size of a displayed image, wherein both the scene map and the depth map belong to the full-size image.

9. The sliding zoom method according to any one of claims 1, 2, 4 to 7, It is characterized in that The fusing each pair of the enlarged sub-image and the reduced sub-image to obtain multiple frames of display images includes: Confirming an overlapping area between an enlarged sub-image of any frame and a reduced sub-image included in any frame on the fused image; Deleting pixel values ​​located in the overlapped area on the reduced sub-images to obtain non-overlapping reduced sub-images; The enlarged sub-image of any frame and the non-overlapping reduced sub-image are merged to obtain a display image of any frame image.

10. A mobile terminal, It is characterized in that The mobile terminal comprises: A main camera, configured to capture a scene graph; a depth camera configured to capture a depth map; a memory configured to store a computer program; A processor, configured to execute the computer program according to the read scene graph and the depth image to implement the method according to any one of claims 1 to 9 to obtain multiple frames of display images; The display is configured to display a video composed of the multiple frames of presentation images.

11. The mobile terminal according to claim 10, It is characterized in that The mobile terminal further includes: An input unit configured to receive an input target focus point; in, The processor is further configured to execute the computer program according to the scene graph, the depth map and the target focus point to obtain the multiple frames of display images.

12. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method described in any one of claims 1 to 9 can be implemented.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, in, When the processor executes the program, the method described in any one of claims 1 to 9 can be implemented.

14. A computer program product, It is characterized in that The computer program product comprises a computer program, wherein the computer program can implement the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Video processing method, electronic equipment and storage medium

    CN111083380A

  • Image processing method and device and electronic equipment

    CN112532808A