Multi-lens video recording method and related devices

By using multi-camera fusion technology, the camera movement method is determined by using the subject and motion information of the image captured by the mobile phone, and high-quality video is generated. This solves the problem that high-quality video cannot be created online in the existing technology, and achieves a video effect with harmonious camera movement and prominent subject.

CN116405776BActive Publication Date: 2025-11-04HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310410670.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-27
Publication Date
2025-11-04
Estimated Expiration
2040-09-27

AI Technical Summary

Technical Problem

Existing mobile video editing software and multi-camera recording technology cannot meet users' needs for creating high-quality videos with good image quality and harmonious camera movements online using mobile phones.

Method used

The system uses a first camera to capture information about the main subject and its motion from multiple images, determines the target camera movement, and then fuses the images captured by the multiple cameras to generate a video with a prominent subject, a clean background, and smooth transitions.

Benefits of technology

It enables the creation of high-quality videos with good image quality and harmonious camera movements online using mobile phones, simplifies the video production process, and reduces the consumption of equipment, manpower and material resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405776B_ABST
    Figure CN116405776B_ABST
Patent Text Reader

Abstract

The application relates to the field of image processing, in particular to a multi-lens video recording method and related equipment. The method comprises the following steps: cooperative shooting of multiple cameras, determination of a target lens movement mode based on information of a main object in a shot image and movement information of the camera, image fusion based on the target lens movement mode, online lens movement fusion of multiple video stream images is realized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, the original application number is 202011037347.5, the original application date is September 27, 2020, and the entire contents of the original application are incorporated into the present application by reference. TECHNICAL FIELD

[0002] The present application relates to the field of video recording, in particular to a multi-lens video recording method and related equipment. BACKGROUND

[0003] Recording life, exchanging experiences, showing skills, and selling goods through videos are very typical social ways at present, and the video forms include but are not limited to short videos, micro videos, micro films, vlogs, embedded videos in web pages, GIFs, etc. In order to create high-quality videos with good imaging quality, prominent theme, clean background, and harmonious camera operation, first of all, expensive professional equipment such as single / micro single / handheld stabilizer / sound microphone needs to be matched, secondly, the script needs to be written and the story line needs to be clarified before shooting, in addition, professional techniques need to be used to improve the content performance during shooting, and finally, complicated editing work needs to be done after shooting, which consumes a lot of manpower, material resources, financial resources, and time. In recent years, the basic imaging quality of mobile phones has been greatly improved in terms of image quality, anti-shake, and focusing, and creating and sharing videos through mobile phones has become a new trend. Popular mobile video production software such as Big Piece / One Flash / Cat Pie belongs to offline post-editing mode, which mainly realizes functions such as filters, transitions, and background music through template means on pre-collected video sources.

[0004] On the other hand, with the improvement of the hardware and software capabilities of mobile phone chips and cameras, using multi-camera to improve the basic quality of videos has become an indispensable highlight at the launch of major mobile phone manufacturers. The Filmic pro application of iPhone 11Pro supports simultaneous recording of wide-angle / ultra-wide-angle / long-focus / front-facing lens, and previews in a multi-window form, and stores multiple video streams. Models such as Huawei P30Pro / Mate30 Pro have Dual View video recording function, which supports simultaneous recording of wide-angle / long-focus, and previews and stores videos in a two-grid form. It can be seen that neither the existing mobile video production software nor the existing mainstream multi-camera simultaneous video recording can meet the widespread demand of users to create high-quality videos with harmonious camera operation online using mobile phones. SUMMARY

[0005] The embodiment of the present application provides a multi-lens video recording method and related equipment, the subject object information of a plurality of first images collected by a first camera and the motion information of the first camera are used to determine a target panning mode, and then the plurality of first images and a plurality of second images collected by a second camera are fused based on the target panning mode to obtain a target video. The present application is not a simple transition or sorting output of image frames of each video stream, but edits, modifies and fuses each video image frame based on the subject information and the motion information, and then outputs a video with the characteristics of prominent subject, clean background and smooth transition, which can help users to create a high-quality video with good imaging quality and harmonious panning online for subsequent sharing.

[0006] In a first aspect, the embodiment of the present application provides a multi-lens video recording method, comprising:

[0007] A plurality of first images and a plurality of second images are obtained, the plurality of first images and the plurality of second images are obtained by a first camera and a second camera for a same shooting object, and the parameters of the first camera and the parameters of the second camera are different. Motion information of the first camera during the process of obtaining the plurality of first images is obtained. Subject object information in each of the plurality of first images is obtained. A target panning mode is determined according to the motion information of the first camera and the subject object information in each of the plurality of first images. The plurality of first images and the plurality of second images are fused according to the target panning mode to obtain a target video, and the panning effect of the target video is the same as the panning effect of a video obtained by using the target panning mode.

[0008] The parameters of the first camera and the parameters of the second camera are different, and specifically include:

[0009] The field of view (FOV) of the first camera is greater than the FOV of the second camera, or the focal length of the first camera is less than the focal length of the second camera.

[0010] Optionally, the first camera can be a wide-angle camera or a main camera, and the second camera can be a main camera or a long-focus camera.

[0011] Through the cooperative shooting of multiple cameras, the target panning mode is determined based on the information of the subject object in the obtained images and the motion information of the cameras, and then the images obtained by the multiple cameras are fused based on the target panning mode, so as to realize online panning fusion of multiple video streams, and not a simple transition or sorting of image frames of each video stream.

[0012] In combination with the first aspect and any of the possible implementation manners, the target panning mode is determined according to the motion information of the first camera and the subject object information in each of the plurality of first images, including:

[0013] It is determined whether the first camera is displaced during the shooting of the plurality of first images according to the motion information of the first camera; when it is determined that the first camera is not displaced, the target panning mode is determined according to the subject object information in each of the plurality of first images; when it is determined that the first camera is displaced, the target panning mode is determined according to the motion mode of the first camera and the subject object information in each of the plurality of first images.

[0014] Based on the determination of whether the first camera is displaced during the shooting of the plurality of first images, a suitable panning mode is selected, thereby providing a basis for obtaining a high-quality video with harmonious panning in the future.

[0015] In combination with the first aspect and any of the possible implementation manners, the target panning mode is determined according to the motion information of the first camera and the subject object information in each of the plurality of first images, including:

[0016] When the subject object information in each of the plurality of first images is used to indicate that the subject object is contained in each of the plurality of first images, and the first proportion of each of the plurality of first images is less than a first preset proportion, the target panning mode is determined as a push lens; wherein the first proportion of each of the plurality of first images is a ratio of an area of a region of interest (ROI) in which the subject object is located in the first image to an area of the first image, or; the first proportion of each of the plurality of first images is a ratio of a width of the ROI in which the subject object is located in the first image to a width of the first image, or the first proportion of each of the plurality of first images is a ratio of a height of the ROI in which the subject object is located in the first image to a height of the first image; when the subject object information in each of the plurality of first images is used to indicate that the subject object is not contained in a default ROI region in each of the plurality of first images, and an image in the default ROI in each of the plurality of first images is a part of the first image, the target panning mode is determined as a pull lens.

[0017] Optionally, the default ROI can be a central region in the first image, or other regions.

[0018] In the determination that the first camera is stationary, the target panning mode is further determined according to the subject object information of the first image collected by the first camera, thereby providing a basis for obtaining a high-quality video with a prominent subject and harmonious panning in the future.

[0019] In combination with the first aspect and any of the possible implementation manners described above, the target camera movement mode is determined according to the movement mode of the first camera and the subject object information of each of the plurality of first images, and the target camera movement mode includes:

[0020] When the movement mode of the first camera indicates that the first camera moves in the same direction during the shooting of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, the target camera movement mode is determined as a moving lens; when the movement mode of the first camera indicates that the first camera reciprocally moves in two opposite directions during the shooting of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, the target camera movement mode is determined as a panning lens; and when the movement mode of the first camera indicates that the first camera moves during the shooting of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images, the target camera movement mode is determined as a tracking lens.

[0021] When it is determined that the first camera is moving, the target camera movement mode is further determined according to the subject object information of the first images collected by the first camera and the movement mode of the first camera, thereby providing a basis for obtaining a high-quality video with a harmonious camera movement in the subsequent process.

[0022] In combination with the first aspect and any of the possible implementation manners described above, when the target camera movement mode is a zoom-in lens or a zoom-out lens, the plurality of first images and the plurality of second images are fused according to the target camera movement mode to obtain a target video, and the method includes:

[0023] obtaining a preset region of each of the plurality of first images, the preset region of each of the plurality of first images including the subject object, and when the target panning mode is a zoom-in mode, sizes of the preset regions in the plurality of first images gradually decrease in a time sequence; when the target panning mode is a zoom-out mode, the sizes of the preset regions in the plurality of first images gradually increase in the time sequence; obtaining a plurality of first image pairs from the plurality of first images and the plurality of second images, each of the plurality of first image pairs including a first image and a second image with a same time stamp; obtaining a plurality of first target images from the plurality of first image pairs, the plurality of first target images corresponding to the plurality of first image pairs one by one; wherein each of the plurality of first target images is the first image with the smallest angle of view range in the corresponding first image pair and includes an image in the preset region of the first image; cropping a part of each of the plurality of first target images that overlaps with an image in the preset region of the first image in the second image pair to which the first target image belongs to obtain a plurality of third images; processing each of the plurality of third images to obtain a plurality of processed third images, each of the plurality of processed third images having a preset resolution, and wherein the target video includes the plurality of processed third images.

[0024] It should be noted that the size of the preset region of the first image can be a pixel area of the preset region in the first image, or a proportion of the preset region in the first image, or a height and a width of the preset region.

[0025] It should be noted that the time stamp of the image in the present application is a time when the camera captures the image.

[0026] When the target panning mode is determined to be the zoom-in mode or the zoom-out mode, the target video with the prominent subject and the smooth transition can be obtained by fusing the first image and the second image, which can help the user to create a high-quality video with good imaging quality, prominent subject and harmonious panning for subsequent sharing by using the mobile phone online.

[0027] In combination with the first aspect and any possible implementation manner described above, when the target panning mode is the zoom-in mode, the zoom-out mode, the pan mode, the tilt mode or the track mode, the plurality of first images and the plurality of second images are fused according to the target panning mode to obtain the target video, including:

[0028] subject detection and extraction are performed on the plurality of first images, ROIs of each of the plurality of first images are obtained, and the ROIs of each of the plurality of first images are cropped to obtain a plurality of fourth images, wherein each of the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images correspond to each other, and a timestamp of each of the plurality of fourth images is the same as a timestamp of a first image corresponding to the fourth image; each of the plurality of fourth images is processed to obtain a plurality of fifth images, and a resolution of each of the plurality of fifth images is a preset resolution; the plurality of fourth images and the plurality of fifth images correspond to each other, and a timestamp of each of the plurality of fifth images is the same as a timestamp of a fourth image corresponding to the fifth image; a plurality of second image pairs are obtained from the plurality of second images and the plurality of fifth images, each of the plurality of second image pairs includes a fifth image and a second image with the same timestamp; a plurality of second target images are obtained from the plurality of second image pairs, and the plurality of second target images correspond to the plurality of image pairs one by one; wherein for each of the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image; wherein the target video includes the plurality of second target images.

[0029] When the target camera movement mode is determined to be follow, pan, or tilt, a target video with a prominent subject and smooth transition is obtained through fusion of the first image and the second image, which can help the user to create a high-quality video with good imaging quality, prominent subject, clean background, and harmonious camera movement for subsequent sharing by using a mobile phone.

[0030] With reference to the first aspect and any one of the possible implementation manners above, the method further includes:

[0031] displaying the image captured by the first camera, the image captured by the second camera, and the target video on the display interface;

[0032] The display interface includes a first region, a second region, and a third region, the first region displays the image captured by the first camera, the second region is used to display the image captured by the second camera, and the third region is used to display the target video.

[0033] By displaying the first image, the second image, and the target video on the display interface in real time, the user can observe the shooting conditions of the cameras and the visual effect of the target video in real time, and adjust the shooting angle or position of the first camera and the second camera in real time, so as to obtain a video with more ideal camera movement effect.

[0034] Further, the first region includes a fourth region, and the fourth region is configured to display a portion of the image captured by the first camera that overlaps the image captured by the second camera.

[0035] In combination with the first aspect and any possible implementation manner described above, the target video is obtained by fusing the plurality of first images and the plurality of second images according to the target lens movement manner.

[0036] The target video is obtained by fusing the plurality of first images, the plurality of second images, and the plurality of sixth images according to the target lens movement manner, wherein the sixth image is obtained by photographing the photographing object by a third camera, and the parameters of the third camera are different from the parameters of the first camera and the parameters of the second camera.

[0037] The FOV of the first camera is greater than the FOV of the second camera, the FOV of the third camera is greater than the FOV of the second camera, and less than the FOV of the first camera, or the focal length of the first camera is less than the focal length of the second camera, and the focal length of the third camera is greater than the focal length of the second camera. For example, the first camera is a wide-angle camera, the second camera is a long-focus camera, and the third camera is a main camera.

[0038] In combination with the first aspect and any possible implementation manner described above, when the target lens movement manner is a zoom-in or a zoom-out, the target video is obtained by fusing the plurality of first images, the plurality of second images, and the plurality of sixth images according to the target lens movement manner, including:

[0039] The preset region of each of the plurality of first images is obtained, the preset region of the plurality of first images includes the main object, and when the target lens movement manner is a zoom-in, the size of the preset region in the plurality of first images gradually decreases in time sequence, or when the target lens movement manner is a zoom-out, the size of the preset region in the plurality of first images gradually increases in time sequence. A plurality of third image pairs are obtained from the plurality of first images, the plurality of second images, and the plurality of sixth images, each of the plurality of third image pairs includes a first image, a second image, and a sixth image with the same time stamp. A plurality of third target images are obtained from the plurality of third image pairs, and the plurality of third target images correspond to the plurality of third image pairs in one-to-one correspondence. Each of the plurality of third target images is the image with the smallest view angle range in the corresponding third image pair and includes the image in the preset region of the first image. A portion of each of the plurality of third target images that overlaps the image in the preset region of the first image in the corresponding third image pair is cropped to obtain a plurality of seventh images. Each of the plurality of seventh images is processed to obtain a plurality of eighth images, and the resolution of each of the plurality of eighth images is a preset resolution. The target video includes the plurality of eighth images.

[0040] The third image is introduced on the basis of the first image and the second image for fusion, and a video with more ideal imaging quality can be obtained.

[0041] With reference to the first aspect and any possible implementation manner of the above, when the target panning mode is a follow shot, a move shot or a pan shot, the target video is obtained by fusing the plurality of first images, the plurality of second images and the plurality of sixth images according to the target panning mode, including:

[0042] The subject in the plurality of first images is detected and extracted, and the ROI of each first image in the plurality of first images is obtained; a plurality of fourth image pairs are obtained from the plurality of first images, the plurality of second images and the plurality of sixth images, each fourth image pair in the plurality of fourth image pairs including a first image, a second image and a sixth image with the same timestamp; a plurality of fourth target images are obtained from the plurality of fourth image pairs, the plurality of fourth target images corresponding to the plurality of fourth image pairs in one-to-one manner; wherein each fourth target image in the plurality of fourth target images is the one with the smallest perspective range in the corresponding fourth image pair and contains the image in the ROI of the first image; a part of each fourth target image in the plurality of fourth target images that overlaps with the image in the ROI of the first image in the fourth image pair to which the fourth target image belongs is cropped from the fourth target image to obtain a plurality of ninth images; each ninth image in the plurality of ninth images is processed to obtain a plurality of tenth images, and the resolution of each tenth image in the plurality of tenth images is a preset resolution, wherein the target video includes the plurality of tenth images.

[0043] The third image is introduced on the basis of the first image and the second image for fusion, and a video with more ideal imaging quality can be obtained.

[0044] With reference to the first aspect and any possible implementation manner of the above, the method further includes:

[0045] The image captured by the first camera, the image captured by the second camera and the image captured by the third camera and the target video are displayed on the display interface.

[0046] The display interface includes a first area, a second area and a third area, the first area includes a fifth area, and the fifth area includes a fourth area; the first area displays the image captured by the first camera, the second area is used to display the image captured by the second camera, the third area is used to display the target video, and the fourth area is used to display the part of the image captured by the first camera that overlaps with the image captured by the second camera; and the fifth area is used to display the part of the image captured by the second camera that overlaps with the image captured by the third camera.

[0047] By displaying the first image, the second image, the third image and the target video in real time on the display interface, the shooting conditions of the cameras and the visual effect of the target video can be observed in real time, and the user can adjust the shooting angle or position of the first camera, the second camera and the third camera in real time, so that the video with more ideal dolly effect can be obtained.

[0048] In a second aspect, the embodiments of the present application provide a master device, comprising:

[0049] The acquisition unit is configured to acquire a plurality of first images and a plurality of second images, the plurality of first images and the plurality of second images being obtained by the first camera and the second camera for a same shooting object, and the parameters of the first camera and the parameters of the second camera being different.

[0050] The acquisition unit is further configured to acquire subject object information in each of the plurality of first images.

[0051] The determination unit is configured to determine a target dolly mode according to the motion information of the first camera and the subject object information in each of the plurality of first images.

[0052] The fusion unit is configured to fuse the plurality of first images and the plurality of second images according to the target dolly mode to obtain a target video, and the dolly effect of the target video is the same as that of a video obtained by using the target dolly mode.

[0053] Optionally, the parameters of the first camera and the parameters of the second camera are different, and specifically include:

[0054] The FOV of the first camera is greater than the FOV of the second camera, or the focal length of the first camera is less than the focal length of the second camera.

[0055] In combination with the second aspect, the determination unit is specifically configured to:

[0056] determine whether the first camera has displacement during the process of obtaining the plurality of first images according to the motion information of the first camera; when it is determined that the first camera has no displacement, determine the target dolly mode according to the subject object information in each of the plurality of first images; and when it is determined that the first camera has displacement, determine the target dolly mode according to the motion mode of the first camera and the subject object information in each of the plurality of first images.

[0057] In combination with the second aspect and any one of the possible implementation manners, in the aspect of determining the target dolly mode according to the subject object information in each of the plurality of first images, the determination unit is specifically configured to:

[0058] when the subject object information of each of the plurality of first images indicates that the subject object is contained in each of the plurality of first images, and the first proportion of each of the plurality of first images is less than the first preset proportion, determining the target lens operation mode as a zoom-in lens operation mode; wherein the first proportion of each of the first images is a ratio of an area of a region of interest (ROI) in which the subject object is located in the first image to an area of the first image, or the first proportion of each of the first images is a ratio of a width of the ROI in which the subject object is located in the first image to a width of the first image, or the first proportion of each of the first images is a ratio of a height of the ROI in which the subject object is located in the first image to a height of the first image;

[0059] when the subject object information of each of the plurality of first images indicates that the subject object is not contained in a default ROI region in each of the plurality of first images, and the image in the default ROI in each of the plurality of first images is a part of the first image, determining the target lens operation mode as a zoom-out lens operation mode.

[0060] With reference to the second aspect and any one of the possible implementation manners above, in the aspect of determining the target lens operation mode according to the movement mode of the first camera and the subject object information of each of the plurality of first images, the determining unit is specifically configured to:

[0061] when the movement mode of the first camera indicates that the first camera moves in the same direction during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, determining the target lens operation mode as a shift lens operation mode;

[0062] when the movement mode of the first camera indicates that the first camera reciprocally moves in two opposite directions during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, determining the target lens operation mode as a tilt lens operation mode;

[0063] when the movement mode of the first camera indicates that the first camera moves during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images, determining the target lens operation mode as a follow lens operation mode.

[0064] With reference to the second aspect and any one of the possible implementation manners above, when the target lens operation mode is the zoom-in lens operation mode or the zoom-out lens operation mode, the fusing unit is specifically configured to:

[0065] obtaining a preset region of each of the plurality of first images, the preset region of each of the plurality of first images including the subject object, and when the target panning mode is a zoom-in mode, sizes of the preset regions in the plurality of first images gradually decrease in a time sequence; when the target panning mode is a zoom-out mode, the sizes of the preset regions in the plurality of first images gradually increase in the time sequence; obtaining a plurality of first image pairs from the plurality of first images and the plurality of second images, each of the plurality of first image pairs including a first image and a second image with a same time stamp; obtaining a plurality of first target images from the plurality of first image pairs, the plurality of first target images corresponding to the plurality of first image pairs in a one-to-one manner; wherein each of the plurality of first target images is the first image pair corresponding thereto with a smallest range of view angles, and contains an image in the preset region of the first image; cropping, from each of the plurality of first target images, a portion overlapping with an image in the preset region of the first image in the second image pair to which the first target image belongs, to obtain a plurality of third images; processing each of the plurality of third images to obtain a plurality of processed third images, each of the plurality of processed third images having a preset resolution, wherein the target video includes the plurality of processed third images.

[0066] With reference to the second aspect and any possible implementation manner described above, when the target panning mode is a shift mode, a tilt mode, or a track mode, the fusion unit is specifically configured to:

[0067] subject detection and extraction are performed on the plurality of first images, ROIs of each of the plurality of first images are obtained, and the ROIs of each of the plurality of first images are cropped to obtain a plurality of fourth images, wherein each of the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images correspond to each other, and a timestamp of each of the plurality of fourth images is the same as a timestamp of a first image corresponding to the fourth image; each of the plurality of fourth images is processed to obtain a plurality of fifth images, and a resolution of each of the plurality of fifth images is a preset resolution; the plurality of fourth images and the plurality of fifth images correspond to each other, and a timestamp of each of the plurality of fifth images is the same as a timestamp of a fourth image corresponding to the fifth image; a plurality of second image pairs are obtained from the plurality of second images and the plurality of fifth images, each of the plurality of second image pairs includes a fifth image and a second image with the same timestamp; a plurality of second target images are obtained from the plurality of second image pairs, and the plurality of second target images correspond to the plurality of image pairs one by one; wherein for each of the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image; and the target video includes the plurality of second target images.

[0068] With reference to the second aspect and any possible implementation manner of the above, the device further includes:

[0069] a display unit configured to display the image captured by the first camera, the image captured by the second camera, and the target video.

[0070] The display unit displays an interface including a first region, a second region, and a third region, the first region displays the image captured by the first camera, the second region is configured to display the image captured by the second camera, and the third region is configured to display the target video, and a field of view FOV of the first camera is greater than a FOV of the second camera.

[0071] Optionally, the first camera is a wide-angle camera or a main camera, and the second camera is a telephoto camera or a main camera.

[0072] Further, the first region includes a fourth region configured to display a portion of the image captured by the first camera that overlaps the image captured by the second camera.

[0073] With reference to the first aspect and any possible implementation manner of the above, the fusion unit is specifically configured to:

[0074] According to the target lens operation mode, the first images, the second images and the sixth images are fused to obtain a target video; the sixth images are obtained by a third camera aiming at the shooting object, and the parameters of the third camera and the first camera are different, and the parameters of the third camera and the second camera are different.

[0075] With reference to the second aspect and any possible implementation manner of the above, when the target lens operation mode is a zoom-in or a zoom-out, the fusion unit is specifically configured to:

[0076] obtain a preset region of each of the first images, the preset region of each of the first images including the main object, and when the target lens operation mode is a zoom-in, the size of the preset region of each of the first images gradually decreases in time sequence; when the target lens operation mode is a zoom-out, the size of the preset region of each of the first images gradually increases in time sequence; obtain a plurality of third image pairs from the first images, the second images and the sixth images, each of the third image pairs including a first image, a second image and a sixth image with the same time stamp; obtain a plurality of third target images from the plurality of third image pairs, the plurality of third target images corresponding to the plurality of third image pairs in one-to-one manner; each of the plurality of third target images is the one with the smallest view angle range in the corresponding third image pair and includes the image in the preset region of the first image; crop a part of each of the plurality of third target images that overlaps with the image in the preset region of the first image in the corresponding third image pair to obtain a plurality of seventh images; process each of the plurality of seventh images to obtain a plurality of eighth images, each of the plurality of eighth images having a preset resolution, and the target video includes the plurality of eighth images.

[0077] With reference to the second aspect and any possible implementation manner of the above, when the target lens operation mode is a zoom-in, a zoom-out, a pan, a shift or a tilt, the fusion unit is specifically configured to:

[0078] The subject detection and extraction are performed on the plurality of first images, and the ROI of each first image in the plurality of first images is obtained; a plurality of fourth image pairs are obtained from the plurality of first images, the plurality of second images, and the plurality of sixth images, each fourth image pair in the plurality of fourth image pairs includes a first image, a second image, and a sixth image with the same timestamp; a plurality of fourth target images are obtained from the plurality of fourth image pairs, the plurality of fourth target images correspond to the plurality of fourth image pairs in one-to-one correspondence; wherein each fourth target image in the plurality of fourth target images has the smallest view angle range in the corresponding fourth image pair, and contains the image in the ROI of the first image; a part of each fourth target image in the plurality of fourth target images that overlaps with the image in the ROI of the first image in the fourth image pair to which the fourth target image belongs is cropped from the fourth target image to obtain a plurality of ninth images; each ninth image in the plurality of ninth images is processed to obtain a plurality of tenth images, and the resolution of each tenth image in the plurality of tenth images is a preset resolution, wherein the target video includes the plurality of tenth images.

[0079] With reference to the second aspect and any one of the possible implementation manners above, the master device further includes:

[0080] The display unit is configured to display the image captured by the first camera, the image captured by the second camera, the image captured by the third camera, and the target video; wherein the FOV of the first camera is greater than the FOV of the second camera, the FOV of the third camera is greater than the FOV of the second camera, and less than the FOV of the first camera,

[0081] The interface displayed by the display unit includes a first area, a second area, and a third area, the first area includes a fifth area, the fifth area includes a fourth area, the first area is configured to display the image captured by the first camera, the second area is configured to display the image captured by the second camera, the third area is configured to display the target video, and the fourth area is configured to display the part of the image captured by the first camera that overlaps with the image captured by the second camera; the fifth area is configured to display the part of the image captured by the second camera that overlaps with the image captured by the third camera.

[0082] In a third aspect, an embodiment of the present application provides an electronic device, including a touch screen, a memory, and one or more processors; wherein one or more programs are stored in the memory; characterized in that, when the one or more processors execute the one or more programs, the electronic device implements part or all of the method of the first aspect.

[0083] In a fourth aspect, an embodiment of the present application provides a computer storage medium, characterized in that, including computer instructions, when the computer instructions run on an electronic device, the electronic device executes part or all of the method of the first aspect.

[0084] In a fifth aspect, an embodiment of the present application provides a computer program product, characterized by, when the computer program product runs on a computer, causing the computer to execute part or all of the method of the first aspect.

[0085] In the scheme of the embodiments of the present application, the information of the subject object in the image collected by the first camera and the motion information of the camera are used to determine the target panning mode, and then the images collected by the plurality of cameras (including the first camera and the second camera, etc.) are fused according to the target panning mode to obtain the target video with the visual effect of the video obtained by using the target panning mode. Instead of simply cutting or sorting the image frames of the video streams and then outputting, the embodiments of the present application edit, modify and fuse the image frames of the video streams according to the subject information and the motion information, and then output. The output has the characteristics of prominent subject, clean background and smooth transition, and can help the user to create a high-quality video with good imaging quality, prominent subject, clean background and harmonious panning online for subsequent sharing.

[0086] It should be understood that any one of the above possible implementation manners can be freely combined without violating the laws of nature, and the present application does not repeat here.

[0087] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in the present application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it can be understood that the description of a feature or a beneficial effect means that the specific technical feature, technical solution or beneficial effect is included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in the specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in the embodiments can be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects can be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0088] Figure 1a The application scenario schematic diagram provided for the embodiments of the present application is shown;

[0089] Figure 1b The application scenario schematic diagram provided for the embodiments of the present application is shown;

[0090] Figure 2a The structure schematic diagram of the master control device is shown;

[0091] Figure 2b The software structure block diagram of the master control device of the embodiments of the present application is shown;

[0092] Figure 3 A flowchart of a multi-camera video recording method is provided for the embodiments of the present application;

[0093] Figure 4 The ROI of the first image and the default ROI are shown;

[0094] Figure 5 The change of the subject object in the ROI of the first image and the moving direction of the lens are shown in the case of moving the lens;

[0095] Figure 6 The change of the subject object in the ROI of the first image and the moving direction of the lens are shown in the case of moving the lens;

[0096] Figure 7 The change of the subject object in the ROI of the first image and the moving direction of the lens are shown in the case of moving the lens;

[0097] Figure 8 The schematic diagrams of different scenes are shown;

[0098] Figure 9 The subject object in the ROI of the first image in the case of different scenes is shown;

[0099] Figure 10 The images contained in the second image pair are shown;

[0100] Figure 11a A display interface schematic diagram is provided for the embodiments of the present application;

[0101] Figure 11b Another display interface schematic diagram is provided for the embodiments of the present application;

[0102] Figure 12 A structure schematic diagram of a master control device is provided for the embodiments of the present application;

[0103] Figure 13 Another structure schematic diagram of a master control device is provided for the embodiments of the present application;

[0104] Figure 14 Another structure schematic diagram of a master control device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0105] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0106] First, the lens moving mode is introduced:

[0107] A push-in shot: A shot that moves slowly or rapidly forward while keeping the subject in the same position. The framing also changes from a long shot to a full shot, medium shot, close-up, or even an extreme close-up. This can be seen as the field of view (FOV) gradually decreasing, or the focal length gradually increasing. The main purpose of this shot is to emphasize the subject, gradually focusing the viewer's attention, enhancing the visual experience, and creating a sense of scrutiny.

[0108] Pulling out: The camera moves in the opposite direction to a push-in shot, moving away from the subject and into the distance. The framing widens, the subject shrinks, and the distance between the subject and the viewer gradually increases. The image expands from few elements to a whole, from a partial view to a complete composition. In terms of shot size, it progresses from a close-up or medium shot to a wide shot or long shot. This can be seen as the field of view (FOV) gradually increasing, or the focal length gradually decreasing. The main purpose of a pull-out shot is to establish the environment in which the subject is situated.

[0109] Panning shot: Used to film continuous action or large-scale scenes. The panning angle is small and the speed is constant. The camera does not move; instead, a movable base allows the lens to rotate up, down, left, right, and even around the camera, much like a person's gaze sweeping over the subject. A panning shot can represent a person's eye, observing everything around them. It plays a unique role in describing space and introducing the environment. Left and right panning is often used to introduce large scenes, while vertical panning is often used to showcase the grandeur and imposing presence of tall objects.

[0110] Tracking shot: A shot where the camera moves horizontally to expand the field of view. The camera moves left and right along the horizontal direction to shoot, similar to people walking and looking around in real life. Like panning shot, tracking shot can expand the two-dimensional spatial image of the screen, but because the camera is not fixed, it has more freedom than panning shot, and can break the limitations of the image and expand the space.

[0111] Tracking shot: A tracking shot is actually a variation of moving the camera, where the camera moves at a constant distance following the subject. Tracking shots always follow the moving subject, creating a strong sense of traversing space, and are suitable for continuously depicting a person's actions, expressions, or subtle changes.

[0112] Whirlwind shot: A rapid transition between two still images, where the image is blurred and becomes a stream of light.

[0113] Secondly, the application scenarios of this application will be introduced.

[0114] The master device 100 acquires a first video stream and a second video stream collected by a first camera and a second camera respectively; wherein the first video stream includes a plurality of first images, and the second video stream includes a plurality of second images; then the master device 100 fuses the first video stream and the second video stream according to the method of the present application, thereby obtaining a target video, which is a video stream with different camera effects or special effects; optionally, the master device 100 also acquires a third video stream through a third camera, and the third video stream includes a plurality of fifth images; then the master device 100 fuses the first video stream, the second video stream and the third video stream according to the method of the present application, thereby obtaining a target video; wherein the display interface of the master device 100 displays the video streams collected by the above cameras and the above target video. The above first video stream, second video stream and third video stream are respectively video streams obtained by the first camera, second camera and third camera for shooting the same shooting object.

[0115] The first camera includes a wide-angle camera, a main camera or a front camera; and the second camera includes a telephoto camera or a main camera, wherein the first camera and the second camera are not the same.

[0116] When the master device 100 only includes the first camera, the second camera or the third camera, or the master device 100 does not include the first camera, the second camera and the third camera, the master device 100 and one or more slave devices 200 form a network, such as Figure 1aAs shown, the master device 100 obtains the first video stream and the second video stream required by the master device 100 from the plurality of slave devices 200; optionally, the master device 100 also obtains a third video stream from the slave devices 200, for example, the master device 100 does not include the first camera and the second camera, the slave device A of the plurality of slave devices 200 includes the first camera, and the slave device B includes the second camera; optionally, the master device 100 does not include the third camera, and the slave device C includes the third camera; the master device 100 can obtain the first video stream collected by the first camera of the slave device A and the second video stream collected by the second camera of the slave device B; optionally, the master device 100 obtains the third video stream collected by the third camera of the slave device C. For another example, the master device 100 includes the first camera and the third camera, and the slave device 200 includes the second camera; the master device 100 obtains the first video stream through the first camera and the second video stream obtained by the slave device 200 through the second camera. The master device 100 fuses the first video stream and the second video stream according to the method of the present application, thereby obtaining a target video; for another example, the master device 100 includes the first camera and the third camera, and the slave device 200 includes the second camera; the master device 100 obtains the first video stream and the third video stream through the first camera and the third camera respectively, and obtains the second video stream obtained by the slave device 200 through the second camera. The master device 100 fuses the first video stream, the second video stream and the third video stream according to the method of the present application, thereby obtaining a target video; the target video is a video stream with different camera effects or special effects.

[0117] Optionally, after obtaining the target video, the master device 100 can display the target video on the display screen of the master device 100 or on the display device 300, as shown. Figure 1b

[0118] It should be noted that, in the process of obtaining the first video stream through the first camera, the device including the first camera can be moving or stationary. Optionally, in the process of obtaining the second video stream through the second camera, the device including the second camera can be moving or stationary; optionally, in the process of obtaining the third video stream through the third camera, the device including the second camera can be moving or stationary.

[0119] The related structure of the above electronic device (i.e., the master device 100 and the slave device 200) will be introduced below. Figure 2a The structural schematic diagram of the master device 100 is shown.

[0120] ​It should be understood that the host device 100 can have more or less components than those shown in the figure, can combine two or more components, or can have a different configuration of components. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0121] The host device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a camera 190, a key, and a display screen 150, etc. The sensor module 180 includes a gyroscope sensor 180A and an acceleration sensor 180B.

[0122] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the host device 100. In other embodiments of the present application, the host device 100 can include more or fewer components than shown in the figure, or combine certain components, or split certain components, or different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0123] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.

[0124] The controller can be the nerve center and command center of the host device 100. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching and executing instructions.

[0125] The processor 110 can also have a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using repeatedly. If the processor 110 needs to use the instructions or data again, it can be called directly from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0126] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0127] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the host device 100. In other embodiments of the present application, the host device 100 can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.

[0128] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger.

[0129] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the external memory, the display 150, the camera 190, and the wireless communication module 160, etc.

[0130] The wireless communication function of the host device 100 can be realized through the antenna 1, the antenna 2, the mobile communication module, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0131] The host device 100 implements display functions through a GPU, a display screen 150, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 150 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0132] The display screen 150 is used to display images, videos, etc. The display screen 150 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the host device 100 can include 1 or N display screens 150, and N is a positive integer greater than 1.

[0133] The host device 100 can implement a shooting function through an ISP, a camera 190, a video codec, a GPU, a display screen 150, and an application processor, etc.

[0134] The ISP is used to process data fed back by the camera 190. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electric signal, and the camera photosensitive element transmits the electric signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also algorithmically optimize the noise and brightness of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be disposed in the camera 190.

[0135] The camera 190 is used to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP to be converted into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard image signal in RGB, YUV, or the like. In embodiments of the present application, the camera 190 includes a camera for capturing images required for face recognition, such as an infrared camera or other camera. The camera for capturing images required for face recognition is generally located on the front of the electronic device, for example, above the touch screen, or can be located at other positions, which is not limited in embodiments of the present application. In some embodiments, the master device 100 can include other cameras. The electronic device can also include a dot matrix emitter (not shown in the figure) for emitting light. The camera captures the light reflected by the face to obtain a face image, and the processor processes and analyzes the face image by comparing it with the stored face image information to verify it.

[0136] Optionally, the camera 190 described above includes part or all of a wide-angle camera, a telephoto camera, a main camera, and a front camera.

[0137] The digital signal processor is used to process digital signals, which can process not only digital image signals but also other digital signals. For example, when the master device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0138] The video codec is used to compress or decompress digital videos. The master device 100 can support one or more video codecs. In this way, the master device 100 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0139] The NPU is a neural-network (NN) computing processor, which quickly processes input information by drawing on the structure of a biological neural network, for example, by drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the master device 100 can realize intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0140] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to extend the storage capacity of the host device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music and videos are saved in the external memory card.

[0141] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the host device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function (such as a face recognition function, a fingerprint recognition function, a mobile payment function, etc.), and the like. The data storage area can store data created during use of the host device 100 (such as face information template data, fingerprint information template, etc.). In addition, the internal memory 121 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0142] The host device 100 can implement an audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and an application processor, and the like. For example, music playing, recording, and the like.

[0143] The audio module 170 is used to convert digital audio information into an analog audio signal output and is also used to convert an analog audio input into a digital audio signal.

[0144] The speaker 170A, also known as a “loudspeaker”, is used to convert an audio electrical signal into a sound signal.

[0145] The receiver 170B, also known as a “earpiece”, is used to convert an audio electrical signal into a sound signal.

[0146] The microphone 170C, also known as a “microphone”, “sound transducer”, is used to convert a sound signal into an electrical signal.

[0147] The earphone interface 170D is used to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or can be a 3.5 mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0148] The gyroscope sensor 180A can be used to determine the motion posture of the host device 100. In some embodiments, the angular velocity of the host device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180A.

[0149] The acceleration sensor 180B can be used to determine the acceleration of the host device 100.

[0150] The touch sensor 180C, also referred to as a "touch panel". The touch sensor 180C can be disposed on the display screen 150, and the touch sensor 180C and the display screen 150 together form a touch screen, also referred to as a "touch screen". The touch sensor 180C is used to detect touch operations acting on or near it. The touch sensor 180C can pass the detected touch operation to the application processor to determine the touch event type. Visual output related to the touch operation can be provided through the display screen 150. In other embodiments, the touch sensor 180K can also be disposed on the surface of the host device 100, which is different from the position where the display screen 150 is located.

[0151] The keys 191 include a power-on key, a volume key, and the like. The keys 191 can be mechanical keys. They can also be touch keys. The host device 100 can receive key inputs and generate key signal inputs related to user settings and function control of the host device 100.

[0152] In the present application, the camera 190 captures a plurality of first images and a plurality of second images, the first images being, for example, images captured by a wide-angle camera or a main camera; the second images being, for example, images captured by a long-focus camera; wherein the object photographed by the first camera is the same as the object photographed by the second camera; motion information is obtained through the gyroscope sensor 180A and the acceleration sensor 180B in the sensor module 180, and then the plurality of first images and the plurality of second images and the motion information are transmitted to the processor 110; the processor 110 obtains the subject object information of each of the plurality of first images; which can be determined by using a subject recognition technology to determine the subject object in the first image, or according to user instructions, such as touch instructions detected by the touch sensor 180C or voice instructions of the user received by the microphone 170C to obtain the subject object information in the first image, and determine the target dolly mode according to the subject object information of each of the plurality of first images and the motion information, fuse the plurality of first images and the plurality of second images according to the target dolly mode, and obtain the target video; and display on the display screen 150; optionally, save the target video to an external memory card through the external storage interface 120.

[0153] The software system of the host device 100 can employ a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. Embodiments of the present application take an Android system with a layered architecture as an example to illustrate the software structure of the host device 100.

[0154] Figure 2b is a software structure block diagram of the host device 100 of embodiments of the present application.

[0155] A layered architecture divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, an Android system is divided into four layers, from top to bottom, an application layer, an application framework layer, an Android runtime and system library, and a kernel layer.

[0156] The application layer can include a series of application packages.

[0157] As shown in Figure 2b , the application packages can include camera, gallery, WLAN, Bluetooth, music, video, calendar, and other applications (also referred to as applications).

[0158] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.

[0159] As shown in Figure 2b , the application framework layer can include a window manager, a content provider, a view system, a resource manager, and the like.

[0160] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and take screenshots, etc.

[0161] The content provider is used to store and obtain data, and make the data accessible to applications. The data can include videos, images, audio, dialed and received calls, browsing history and bookmarks, phone books, and the like.

[0162] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, and the like. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.

[0163] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and the like.

[0164] The Android Runtime includes a core library and a virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0165] The core library includes two parts: one part is the function function that the java language needs to call, and the other part is the core library of Android.

[0166] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the management of the stack, the management of the thread, the management of the security and the exception, and the garbage collection and the like.

[0167] The system library can include a plurality of functional modules. For example: a surface manager, media libraries, a two-dimensional graphics engine (for example: SGL) and an image processing library and the like.

[0168] The surface manager is used for managing the display subsystem, and provides a plurality of applications with the fusion of 2D and 3D layers.

[0169] The media library supports a plurality of commonly used audio, video format playback and recording, and static image files and the like. The media library can support a plurality of audio and video coding formats, for example: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG and the like.

[0170] The two-dimensional graphics engine is a drawing engine for two-dimensional drawing.

[0171] The kernel layer is a layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver and a sensor driver.

[0172] In the embodiment of the application, the camera is controlled by the camera driver to obtain a plurality of first images and a plurality of second images. The image processing library obtains the subject object information of each first image in the plurality of first images and the motion information of the first camera, and determines a target dolly mode according to the subject object information of each first image in the plurality of first images and the motion information of the first camera, fuses the plurality of first images and the plurality of second images according to the target dolly mode, and obtains a target video; and displays the target video on the display screen through the display driver. For specific processes, please refer to the following related description.

[0173] The following describes how the host device 100 implements the multi-camera video recording process.

[0174] Referring to Figure 3 , Figure 3A flowchart of a multi-camera video recording method is provided for the embodiments of the present application. As shown in Figure 3 The method comprises the following steps.

[0175] S301, acquiring a plurality of first images captured by a first camera and a plurality of second images captured by a second camera.

[0176] The plurality of first images and the plurality of second images are respectively captured by the first camera and the second camera for the same shooting object, and the parameters of the first camera and the parameters of the second camera are different.

[0177] Specifically, the parameters of the first camera and the parameters of the second camera are different, and specifically include the following:

[0178] The field of view (FOV) of the first camera is greater than the FOV of the second camera, or the focal length of the first camera is less than the focal length of the second camera.

[0179] Optionally, the frame rate of the first camera when capturing the first image is the same as the frame rate of the second camera when capturing the second image. In other words, in the same time period, the number of frames of the first image captured by the first camera is the same as the number of frames of the second image captured by the second camera, and the first image and the second image captured at the same time have the same timestamp.

[0180] It should be pointed out that in the present application, the timestamp of an image is the time when the camera captures the image.

[0181] Optionally, the first camera and the second camera can be the same device or different devices.

[0182] In a feasible embodiment, the first camera is a wide-angle camera, a main camera or a front camera, and the second camera is a telephoto camera or a main camera.

[0183] In a feasible embodiment, the first image and the second image can be acquired from the same device or from different devices. For example, the main control device contains the first camera and the second camera, and the first image and the second image can be captured by the main control device. For another example, the main control device contains the first camera, and the slave device contains the second camera. The main control device acquires the plurality of second images captured by the slave device after capturing the first image by the first camera and the plurality of second images by the slave device. For another example, the slave device contains the first camera and the second camera, and the main control device does not contain the first camera and / or the second camera. The main control device acquires the plurality of first images and the plurality of second images captured by the slave device after the slave device captures the plurality of first images and the plurality of second images by the first camera and the second camera respectively.

[0184] In this way, the first image collected by the first camera and / or the second image collected by the second camera can be obtained from other devices under the premise that the host device does not contain the first camera and / or the second camera, so as to realize recording based on multi-camera video.

[0185] S302, obtain motion information of the first camera in the process of shooting a plurality of first images and subject object information of each of the plurality of first images.

[0186] Optionally, the subject object can be a person, an animal or an object such as a mobile phone, a cup, a bag, a car, etc.

[0187] In a feasible embodiment, the motion information of the first camera in the process of shooting a plurality of first images comprises:

[0188] Feature points of each of the plurality of first images are extracted to obtain the feature points of each of the plurality of first images; feature points of adjacent two images in the plurality of first images are matched to obtain a matching result; and then the motion information of the first camera is determined according to the feature matching result.

[0189] Or the motion information of the first camera is directly obtained from a sensor.

[0190] In an example, the motion information of the first camera in the process of shooting a plurality of first images determines whether the first camera is displaced in the process of shooting the plurality of first images, so as to determine whether the first camera is moving or stationary; further, in the case that the first camera is moving in the process of shooting the plurality of first images, the motion information further indicates that the first camera moves in the same direction or reciprocally moves in two opposite directions in the process of shooting the plurality of first images.

[0191] In an optional embodiment, the subject object information of each of the plurality of first images comprises:

[0192] Image saliency detection is performed on each of the plurality of first images to obtain a saliency region of each of the plurality of first images; a maximum rectangular region containing the saliency region in each of the plurality of first images is obtained; and deep learning based on a convolutional neural network is performed on the maximum rectangular region in each of the plurality of first images, such as adopting scene recognition technology, human body recognition technology or face recognition technology, to obtain the subject object information of each of the plurality of first images. The maximum rectangular region is the ROI.

[0193] When there are a plurality of subject objects in the first image, saliency detection is performed on the first image to obtain a plurality of saliency regions, and the ROI is a maximum rectangular region including the plurality of saliency regions.

[0194] It should be noted that the shape of the salient region is not limited to a rectangle, but can also be other regular shapes such as a square, a triangle, a circle, an ellipse, etc., and can also be an irregular shape.

[0195] In an optional embodiment, the subject object information of each of the plurality of first images is determined according to a user instruction.

[0196] Optionally, the user instruction can be a voice instruction, a touch instruction, or other forms of instructions.

[0197] For example, the voice instruction can be "determine a person as a subject object" or "the subject object is an animal"; when the master device receives the voice instruction, the person or the animal in the first image is determined as the subject object of the first image.

[0198] For another example, when the display interface of the master device displays the first image, the master device detects a touch instruction of the user, and when the touch instruction of the user is detected, the master device determines the object in the area where the touch instruction is detected as the subject object of the first image.

[0199] S303, determining a target panning mode according to the motion information of the first camera and the subject object information in each of the plurality of first images.

[0200] In an available embodiment, the target panning mode is determined according to the motion information of the first camera and the subject object information in each of the plurality of first images, including:

[0201] According to the motion information of the first camera, it is determined whether the first camera is displaced during the process of shooting the plurality of first images;

[0202] When it is determined that the first camera is not displaced, the target panning mode is determined according to the subject object information of each of the plurality of first images; when it is determined that the first camera is displaced, the target panning mode is determined according to the motion mode of the first camera and the subject object information of each of the plurality of first images.

[0203] According to the motion information of the first camera, it is determined whether the first camera is moving or stationary during the process of recording the video, and then two different ways are adopted to select a suitable panning mode for the two cases of the first camera being moving and stationary.

[0204] In an available embodiment, when it is determined that the first camera is stationary, the target panning mode is determined according to the subject object information of each of the plurality of first images, including:

[0205] When the subject object information of each of the multiple first images is used to indicate that each of the multiple first images contains a subject object, and the first ratio of each of the multiple first images is less than a first preset ratio, the target camera movement method is determined to be a push-in shot. Optionally, the first ratio of the first image is the ratio of the area of ​​the ROI where the subject object is located in the first image to the area of ​​the first image, or the first ratio of the first image is the ratio of the width of the ROI where the subject object is located in the first image to the width of the first image, or the first ratio of the first image is the ratio of the height of the ROI where the subject object is located in the first image to the height of the first image; preferably, the first ratio of the first image is the ratio of the height of the ROI where the subject object is located in the first image to the height of the first image.

[0206] When the subject object information of each of the multiple first images is used to indicate that there is no subject object in the default ROI region of each first image, and the image in the default ROI of each first image is a part of that first image, the target camera movement method is determined to be a pull-back shot.

[0207] It should be noted that when performing subject detection on the first image, a default ROI is first assigned, so that the main control device focuses on subject detection within the ROI.

[0208] Specifically, such as Figure 4 In the first image shown in Figures a and b, the small rectangle represents the Region of Interest (ROI). The area of ​​the ROI is smaller than the area of ​​the first image. Furthermore, the main subject within the first image is not clearly visible to the user. Therefore, a push-in shot is chosen to improve the user's focus and allow them to clearly see the main subject. Figure 4 As shown in Figure c, the default ROI does not contain the main object or does not contain a complete main object. In order to make the main object fully presented in the ROI, the target camera movement method is set to pull back.

[0209] In one feasible embodiment, when it is determined that the first camera is in motion, the target camera movement method is determined based on the movement mode of the first camera and the subject object information of each of the multiple first images, including:

[0210] When the movement mode of the first camera indicates that the first camera moves in the same direction during the capture of multiple first images, and the subject object information of each of the multiple first images is used to indicate that the subject object in the multiple first images has changed, the target camera movement mode is determined to be a camera pan.

[0211] When the movement mode of the first camera indicates that the first camera reciprocates in two opposite directions during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, the target panning mode is determined as a dolly-zoom.

[0212] When the movement mode of the first camera indicates that the first camera moves during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images, the target panning mode is determined as a follow.

[0213] As shown in FIG. a and FIG. b of Figure 5 , the first camera moves in one direction, and the subject object in the ROI in the first image changes, as shown in FIG. a of Figure 5 , the subject object is a male; as shown in FIG. b of Figure 5 , the subject object is a female, and the subject object changes,

[0214] Therefore, the target panning mode is determined as a pan.

[0215] As shown in FIG. a and FIG. b of Figure 6 , the first camera reciprocates in two directions, and the subject object in the ROI in the first image changes; for example, at a first time, the subject object in the ROI in the first image is a male, at a second time when the first camera moves to the right, the subject object in the ROI in the first image is a female; at the second time when the first camera moves to the left, the subject object in the ROI in the first image is a male; reciprocating in this way, therefore, the target panning mode is determined as a dolly-zoom.

[0216] As shown in FIG. a, FIG. b and FIG. c of Figure 7 , the first camera moves in one direction following the subject object, and the subject object in the ROI in the first image does not change, therefore, the target panning mode is determined as a follow.

[0217] In a feasible embodiment, the target panning mode can also be determined in the following manner:

[0218] The scene type of each of the plurality of first images is determined according to the subject object information of each of the plurality of first images, and the target panning mode is determined according to the scene type and the subject object information of each of the plurality of first images.

[0219] Specifically, the scene type of each of the plurality of first images is determined according to the subject object information of each of the plurality of first images, including:

[0220] When the number of subject objects in each of the plurality of first images is greater than or equal to the first preset number, and the scene understanding information is used to indicate that the surrounding environment needs to be presented, or when the motion amplitude of the subject objects in each of the plurality of first images is greater than the preset amplitude, and the scene understanding information is used to indicate that the surrounding environment needs to be presented, the shot type of the first image is a full shot;

[0221] When the number of subject objects in each of the plurality of first images is less than the first preset number, and the scene understanding information is used to indicate that the current scene has dialogue, action and emotional communication, the shot type of each of the first images is a medium shot;

[0222] When the number of subject objects in each of the plurality of first images is the second preset number, the motion amplitude of the subject objects in the first image is less than or equal to the preset amplitude, and the scene understanding information is used to indicate that the upper body fine action needs to be highlighted, the shot type of each of the first images is a close-up shot; the second preset number is less than the first preset number.

[0223] When the number of subject objects in each of the plurality of first images is the second preset number, the motion amplitude of the subject objects in the first image is less than or equal to the preset amplitude, and the scene understanding information is used to indicate that the local information needs to be highlighted, the shot type of each of the first images is a close-up shot.

[0224] For example, referring to Table 1 below, Table 1 is a correspondence table of subject object information and scene understanding information and shot type.

[0225] Subject object information and scene understanding information Scene type 1 subject object, no large movement, need to highlight local information such as facial expression Close-up 1 subject object, no large movement, need to highlight upper body movements Close-up 3 or fewer subject objects, dialogue, action and emotional exchange scene Medium shot 3 or more subject objects, or large movement, need to present the surrounding environment Wide shot

[0226] Table 1

[0227] As shown in Table 1, when the number of subject objects in the first image is 1, the subject object has no large motion, and the scene understanding information is used to indicate that the local information such as facial expression needs to be highlighted, the shot type of the first image is a close-up shot; when the number of subject objects in the first image is 1, the subject object has no large motion, and the scene understanding information is used to indicate that the upper body fine action needs to be highlighted, the shot type of the first image is a close-up shot; when the number of subject objects in the first image is less than or equal to 3, and the scene understanding information is used to indicate that the current scene is a scene with dialogue, action and emotional communication, the shot type of the first image is a medium shot; when the number of subject objects in the first image is greater than 3 and the scene understanding information is used to indicate that the surrounding environment needs to be presented, or when the subject objects in the first image have large motion and the scene understanding information is used to indicate that the surrounding environment needs to be presented, the shot type of the first image is a full shot.

[0228] It should be noted that the shot type refers to the difference in the range of the subject presented in the camera recorder due to the different distances between the camera and the subject. The shot type can be generally divided into five types, such asFigure 8 As shown, from near to far, they are close-up (pointing to the human shoulder), close-up (pointing to the human chest), medium shot (pointing to the human knee), full shot (the whole human body and the surrounding environment), long shot (the environment where the subject is located).

[0229] Optionally, in one possible embodiment, the scene type of each first image is determined according to the subject object information of each first image in the plurality of first images, comprising:

[0230] The second ratio of each first image is determined according to the subject object of each first image in the plurality of first images, and the second ratio is the ratio of the part of the subject object in the first image located in the ROI to the subject object in the first image;

[0231] When the second ratio of each first image is less than the second preset ratio, the scene type of each first image is close-up; when the second ratio of each first image is greater than or equal to the second preset ratio and less than the third preset ratio, the scene type of each first image is close-up; when the second ratio of each first image is greater than or equal to the third preset ratio and less than the fourth preset ratio, the scene type of each first image is medium shot; when the second ratio of each first image is equal to the fourth preset ratio, the third ratio of the first image is obtained, the third ratio is the ratio of the area occupied by the subject object in the first image to the ROI, and if the third ratio of the first image is greater than or equal to the fifth preset ratio, the scene type of the first image is full shot; if the third ratio is less than the fifth preset ratio, the scene type of the first image is long shot.

[0232] For example, as shown in FIG. 1, assuming that the second ratio is the ratio of the part above the human shoulder to the whole person, the third ratio is the ratio of the part above the human waist to the whole person, and the fourth ratio is 100%; as shown in FIG. 1a, the part of the subject object in the first image located in the ROI is the head, the ratio of the head to the whole person is less than the second ratio, so the scene type of FIG. 1a is close-up; as shown in FIG. 1b, the part of the subject object in the first image located in the ROI is the part above the chest, the ratio of the part above the chest to the whole person is greater than the second ratio but less than the third ratio, so the scene type of FIG. 1b is close-up; as shown in FIG. 1c, the part of the subject object in the first image located in the ROI is the part above the thigh, the ratio of the part above the thigh to the whole person is greater than the second ratio but less than the fourth ratio, so the scene type of FIG. 1a is medium shot; as shown in FIG. 1d, the part of the subject object in the first image located in the ROI is the whole person, so the scene type of FIG. 1d is full shot. Figure 9 Figure 9 Figure 9 Figure 9 Figure 9 Figure 9 Figure 9 Figure 9 ​​​​​​​As shown in FIG. d and FIG. e, the subject object in FIG. d and FIG. e is entirely located within the ROI, but the proportion of the subject object in the ROI in FIG. e is different from the proportion of the subject object in the ROI in FIG. d. The scene type of FIG. d in the image in FIG. d can be determined as a full view, and the scene type of FIG. e in the image in FIG. e can be determined as a long shot. Figure 9 Figure 9

[0233] It should be noted that the manner of determining the scene type of the image in the present application is not a limitation of the present application, and other manners can also be used.

[0234] Optionally, in an available embodiment, the target panning mode is determined according to the scene type and the subject object information of each of the plurality of first images, comprising:

[0235] When the scene types of the plurality of first images are the same, and the subject objects in the plurality of first images are the same, the target panning mode is determined as follow; when the scene types of the plurality of first images are different, and the subject objects in the plurality of first images are the same, the target panning mode is determined as ; when the scene types of the plurality of first images are the same, and the subject objects in the plurality of first images are different, the target panning mode is determined as move or shake; as shown in Table 2.

[0236] Subject object and scene information Camera movement Same subject object, different scene Push / pull Same subject object, same scene Follow Different subject objects, same scene Shift / pan

[0237] Table 2

[0238] It should be noted that the panning mode includes but is not limited to the above push / pull, follow, move / shake, and can also include a throw, and even include some special effects such as slow motion, Hitchcock zoom, cinema graph, etc. The correspondence between the subject object and the scene type information and the panning mode shown in Table 2 is only an example, and is not a limitation of the present application.

[0239] It should be noted that the scene types of the plurality of first images are different, which can be that the scene types of the plurality of first images are all different, or that the scene types of part of the plurality of first images are different; similarly, the subject objects of the plurality of first images are different, which can be that the subject objects of the plurality of first images are all different, or that the subject objects of part of the plurality of first images are different.

[0240] S304, according to the target panning mode, the plurality of first images and the plurality of second images are fused to obtain a target video.

[0241] The panning effect of the target video is the same as the panning effect of the video obtained by using the target panning mode.

[0242] ​​In one possible implementation, when the target panning mode is a push or pull lens, the target video is obtained by fusing the plurality of first images and the plurality of second images according to the target panning mode, including:

[0243] The preset region of each of the plurality of first images is obtained, the preset region of each of the plurality of first images includes the subject object, and when the target panning mode is a push lens, the size of the preset region in the plurality of first images gradually decreases in time sequence; when the target panning mode is a pull lens, the size of the preset region in the plurality of first images gradually increases in time sequence. A plurality of first image pairs is obtained from the plurality of first images and the plurality of second images, each of the plurality of first image pairs includes a first image and a second image with the same timestamp. A plurality of first target images is obtained from each of the plurality of second image pairs, the plurality of first target images correspond to the plurality of first image pairs in one-to-one manner. Each of the plurality of first target images is the first image pair corresponding thereto with the smallest range of view angles and includes the image in the preset region of the first image. A part of each of the plurality of first target images that overlaps with the image in the preset region of the first image in the second image pair to which the first target image belongs is cropped to obtain a plurality of third images. Each of the plurality of third images is processed to obtain a plurality of processed third images, and the resolution of each of the plurality of processed third images is a preset resolution. The target video includes the plurality of processed third images.

[0244] It should be noted that the size of the preset region of the first image can be the pixel area of the preset region in the first image, or the proportion of the preset region in the first image, or the height and width of the preset region.

[0245] Specifically, a preset region is set for each of the plurality of first images, wherein the preset region of each first image contains the subject object, and in order to reflect the visual effect of a push or pull lens, when the target lens operation mode is a push lens, the size of the preset region of the plurality of first images gradually decreases in the acquisition time sequence, that is, for any two adjacent first images in the plurality of first images in the acquisition time, the image in the preset region of the first image with the earlier timestamp contains the image in the preset region of the first image with the later acquisition timestamp; when the target lens operation mode is a pull lens, the size of the preset region of the plurality of first images gradually increases in the acquisition time sequence, that is, for any two adjacent first images in the plurality of first images in the acquisition time, the image in the preset region of the first image with the later timestamp contains the image in the preset region of the first image with the earlier acquisition timestamp, and then based on the preset region of the first image, a plurality of first target images are obtained from the plurality of first images and the second images corresponding to the plurality of first images, the first target image being the image in the preset region of the first image, but whether the image in the preset region is a part of the first image or a part of the second image is determined according to the following method:

[0246] When the image in the preset region of the first image is contained in the second image corresponding to the first image, the part overlapping the image in the preset region of the first image is cropped from the second image as the corresponding first target image; when the image in the preset region of the first image contains the second image corresponding to the first image, the image in the preset region is cropped from the first image as the corresponding first target image, that is, from the first image and the second image corresponding thereto, the image containing the image in the preset region of the first image and having the smallest view angle range is determined, and then the part overlapping the image in the preset region is cropped from the image having the smallest view angle range to obtain the first target image; wherein the second image corresponding to the first image is the image having the same timestamp as the first image;

[0247] After obtaining the plurality of first target images according to the above method, the part overlapping the image in the preset region of the corresponding first image is cropped from each of the plurality of first target images to obtain a plurality of third images; in order to display the third images according to the preset resolution, each of the plurality of third images is processed to obtain a plurality of processed third images, and the resolution of each of the plurality of processed third images is the preset resolution; the above target video includes the plurality of processed third images.

[0248] It should be noted that the preset region of each of the plurality of first images contains the subject object specifically refers to the part or all of the subject object contained in the preset region of the plurality of first images; further, when the target camera operation mode is a push lens, the preset region of the plurality of first images realizes the transition from the first part of the subject object to the second part of the subject object in time sequence, wherein the first part of the subject object contains the second part of the subject object, and the first part of the subject object is all or part of the subject object, and the second part of the subject object is part of the subject object; when the target camera operation mode is a pull lens, the preset region of the plurality of first images realizes the transition from the third part of the subject object to the fourth part of the subject object in time sequence, wherein the fourth part of the subject object contains the third part of the subject object, and the fourth part of the subject object is all or part of the subject object, and the third part of the subject object is part of the subject object.

[0249] According to the above timestamp, the third image is displayed, so as to realize the target video with the "push" or "pull" camera operation effect.

[0250] In a feasible embodiment, when the target camera operation mode is a shift lens, a follow lens or a pan lens, the plurality of first images and the plurality of second images are fused according to the target camera operation mode to obtain the target video, comprising:

[0251] The subject detection and extraction are performed on the plurality of first images, the ROI of each of the plurality of first images is obtained, and the ROI of each of the plurality of first images is cropped to obtain a plurality of fourth images, wherein each of the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images are one-to-one corresponding, and the timestamp of each of the plurality of fourth images is the same as the timestamp of the first image corresponding to the fourth image; each of the plurality of fourth images is processed to obtain a plurality of fifth images, and the resolution of each of the plurality of fifth images is a preset resolution; the plurality of fourth images and the plurality of fifth images are one-to-one corresponding, and the timestamp of each of the plurality of fifth images is the same as the timestamp of the fourth image corresponding to the fifth image; a plurality of second image pairs are obtained from the plurality of second images and the plurality of fifth images, and each of the plurality of second image pairs includes a fifth image and a second image with the same timestamp; a plurality of second target images are obtained from the plurality of second image pairs, and the plurality of second target images are one-to-one corresponding to the plurality of image pairs; wherein for each of the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image.

[0252] The target video includes a plurality of second target images.

[0253] It should be noted that each of the plurality of fourth images includes the subject object, and it can be understood that each of the plurality of fourth images includes part or all of the subject object.

[0254] It should be noted that the step of “performing subject detection and extraction on the plurality of first images to obtain the ROI of each of the plurality of first images” can not be performed, and the ROI obtained in step S302 can be directly used.

[0255] For example, assuming that the first camera collects 8 first images, the timestamps of the 8 first images are t1, t2, t3, t4, t5, t6, t7 and t8 respectively; the second camera collects 8 second images, the timestamps of the 8 second images are t1, t2, t3, t4, t5, t6, t7 and t8 respectively; performing subject detection and extraction on the ROIs of the 8 first images to obtain 8 fourth images, each of the 8 fourth images includes the subject object, and the timestamp of each of the 8 fourth images is the same as the timestamp of the first image to which the fourth image belongs, so the timestamps of the 8 fourth images are t1, t2, t3, t4, t5, t6, t7 and t8 respectively; then super-resolution is performed on each of the 8 fourth images to obtain 8 fifth images, the resolution of each of the 8 fifth images is the preset resolution, and the timestamps of the 8 fifth images are t1, t2, t3, t4, t5, t6, t7 and t8 respectively, as shown in FIG. 8, the second images and the fifth images with the same timestamp in the 8 second images and the 8 fifth images are divided into an image pair, and then 8 second image pairs are obtained, as shown by the dashed box in FIG. 9; obtaining 8 second target images from the 8 second image pairs; for each of the 8 second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the third target image corresponding to the second image pair is the fifth image, and the 8 second target images constitute a target video. Figure 10 Figure 10 As shown in FIG. 10, the contents of the second images and the fifth images in the second image pairs with timestamps t2, t4, t5 and t8 coincide, and the contents of the second images and the fifth images in the second image pairs with timestamps t1, t3, t6 and t7 do not coincide, so the target video includes the fifth images with timestamps t1, t3, t6 and t7 and the second images with timestamps t1, t3, t6 and t7. Figure 10 As shown in FIG. 10, the contents of the second images and the fifth images in the second image pairs with timestamps t2, t4, t5 and t8 coincide, and the contents of the second images and the fifth images in the second image pairs with timestamps t1, t3, t6 and t7 do not coincide, so the target video includes the fifth images with timestamps t1, t3, t6 and t7 and the second images with timestamps t1, t3, t6 and t7.​

[0256] Optionally, in an embodiment, the image captured by the first camera, the image captured by the second camera and the target video are displayed on the display interface. As shown in Figure 11a the display interface includes three areas: a first area, a second area and a third area. The first area is used to display the image captured by the first camera, the second area is used to display the image captured by the second camera, and the third area is used to display the target video.

[0257] Further, as shown in Figure 11a the first area includes a fourth area, which is used to display the part of the image captured by the first camera that overlaps the image captured by the second camera.

[0258] The first camera is a wide-angle camera or a main camera, the second camera is a telephoto camera or a main camera, and the FOV of the first camera is greater than the FOV of the second camera.

[0259] In an embodiment, the first images and the second images are fused according to the target lens movement mode to obtain the target video, including:

[0260] The first images, the second images and the sixth images are fused according to the target lens movement mode to obtain the target video.

[0261] The sixth image is captured by a third camera for a shooting object, the parameters of the third camera are different from the parameters of the first camera, and the parameters of the third camera are different from the parameters of the second camera.

[0262] Optionally, the FOV of the first camera is greater than the FOV of the second camera, the FOV of the third camera is greater than the FOV of the second camera and less than the FOV of the first camera, or the focal length of the first camera is less than the focal length of the second camera, and the focal length of the third camera is less than the focal length of the second camera.

[0263] For example, the first camera is a wide-angle camera, the second camera is a telephoto camera, and the third camera is a main camera.

[0264] Optionally, in an embodiment, when the target lens movement mode is a push lens or a pull lens, the first images, the second images and the sixth images are fused according to the target lens movement mode to obtain the target video, including:

[0265] The preset region of each of the plurality of first images is obtained, the preset region of the plurality of first images includes the main object, and when the target lensing mode is a zoom-in lensing mode, the size of the preset region in the plurality of first images gradually decreases in time sequence; when the target lensing mode is a zoom-out lensing mode, the size of the preset region in the plurality of first images gradually increases in time sequence; a plurality of third image pairs are obtained from the plurality of first images, the plurality of second images and the plurality of sixth images, each of the plurality of third image pairs includes a first image, a second image and a sixth image with the same time stamp; a plurality of third target images are obtained from the plurality of third image pairs, the plurality of third target images correspond to the plurality of third image pairs one by one; wherein each of the plurality of third target images is the one with the smallest view angle range in the corresponding third image pair and contains the image in the preset region of the first image; a part of each of the plurality of third target images is cropped, which overlaps with the image in the preset region of the first image in the third image pair to which the third target image belongs, to obtain a plurality of seventh images; each of the plurality of seventh images is processed to obtain a plurality of eighth images, and the resolution of each of the plurality of eighth images is a preset resolution, wherein the target video includes the plurality of eighth images.

[0266] Specifically, a preset region is set for each of the plurality of first images, wherein the preset region of each of the plurality of first images contains the main object, and in order to reflect the visual effect of zoom-in or zoom-out lensing, when the target lensing mode is a zoom-in lensing mode, the size of the preset region of the plurality of first images gradually decreases in the time sequence of acquisition, that is, for any two adjacent first images in the plurality of first images in the acquisition time, the image in the preset region of the first image with the earlier time stamp contains the image in the preset region of the first image with the later acquisition time stamp; when the target lensing mode is a zoom-out lensing mode, the size of the preset region of the plurality of first images gradually increases in the time sequence of acquisition, that is, for any two adjacent first images in the plurality of first images in the acquisition time, the image in the preset region of the first image with the later time stamp contains the image in the preset region of the first image with the earlier acquisition time stamp, and then a plurality of third target images are obtained from the plurality of first images, the second images corresponding to the plurality of first images and the sixth images corresponding to the plurality of first images based on the preset region of the first image, the third target image being the image in the preset region of the first image, but whether the image in the preset region is a part of the first image or a part of the second image or the sixth image is determined according to the following method:

[0267] When the image in the preset region of the first image is contained in the sixth image corresponding to the first image, the part of the sixth image that overlaps with the image in the preset region of the first image is cropped as the corresponding third target image; when the image in the preset region of the first image contains the second image and the sixth image corresponding to the first image, the image in the preset region of the first image is cropped from the first image as the corresponding third target image, that is, from the first image and the second image and the sixth image corresponding to the first image, the image in the preset region of the first image is determined in the image with the smallest angle of view range, and then the part of the image in the preset region of the first image is cropped from the image with the smallest angle of view range, so as to obtain the third target image; wherein the second image and the sixth image corresponding to the first image are respectively the second image and the sixth image with the same time stamp as the time stamp of the first image.

[0268] After a plurality of third target images are obtained according to the above method, the part of each third target image in the plurality of third target images that overlaps with the image in the preset region of the first image in the third image pair to which the third target image belongs is cropped to obtain a plurality of seventh images; in order to display the seventh images according to the preset resolution, each seventh image in the plurality of seventh images is processed to obtain a plurality of eighth images, and the resolution of each eighth image in the plurality of eighth images is the preset resolution; the above target video includes the plurality of eighth images.

[0269] It should be noted that the preset region of each first image in the plurality of first images contains the main object, which means that the preset region of the plurality of first images contains part or all of the main object; further, when the target camera movement mode is a push shot, the preset region of the plurality of first images realizes a transition from a first part of the main object to a second part of the main object in time sequence, wherein the first part of the main object contains the second part of the main object, and the first part of the main object is all or part of the main object, and the second part of the main object is part of the main object; when the target camera movement mode is a pull shot, the preset region of the plurality of first images realizes a transition from a third part of the main object to a fourth part of the main object in time sequence, wherein the fourth part of the main object contains the third part of the main object, and the fourth part of the main object is all or part of the main object, and the third part of the main object is part of the main object.

[0270] Optionally, in a feasible embodiment, when the target panning mode is a follow shot, a pan shot or a tilt shot, the target video is obtained by fusing the plurality of first images, the plurality of second images and the plurality of sixth images according to the target panning mode, comprising:

[0271] The subject detection and extraction are performed on the plurality of first images to obtain the ROI of each first image in the plurality of first images; a plurality of fourth image pairs are obtained from the plurality of first images, the plurality of second images and the plurality of sixth images, each fourth image pair in the plurality of fourth image pairs comprising a first image, a second image and a sixth image with the same timestamp; a plurality of fourth target images are obtained from the plurality of fourth image pairs, the plurality of fourth target images corresponding to the plurality of fourth image pairs one by one; wherein each fourth target image in the plurality of fourth target images is the one with the smallest angle of view range in the corresponding fourth image pair and contains the image within the ROI of the first image; a part of each fourth target image in the plurality of fourth target images that overlaps with the image within the ROI of the first image in the fourth image pair to which the fourth target image belongs is cropped from the fourth target image to obtain a plurality of ninth images; each ninth image in the plurality of ninth images is processed to obtain a plurality of tenth images, each tenth image in the plurality of tenth images having a preset resolution, wherein the target video comprises the plurality of tenth images.

[0272] Further, as shown in Figure 11b the display interface comprises a first area, a second area and a third area, the first area comprises a fifth area, the fifth area comprises a fourth area,

[0273] The first area displays the image captured by the first camera, the second area is used to display the image captured by the second camera, the third area is used to display the target video, the fourth area is used to display the part of the image captured by the first camera that overlaps with the image captured by the second camera, and the fifth area is used to display the part of the image captured by the second camera that overlaps with the image captured by the third camera.

[0274] In a feasible embodiment, after determining the subject object in the first image, the method of the present application comprises:

[0275] The master control device acquires the real-time position of the subject object by using a target tracking technology; and adjusts the field of view range of the first camera, the second camera and / or the third camera according to the real-time position of the subject object, so that the first camera, the second camera and / or the third camera capture images containing the subject object in real time.

[0276] For example, when it is determined according to the real-time position of the subject object that the subject object is about to leave the shooting range of the telephoto camera, the focal length of the telephoto lens is controlled to be smaller, so that the field of view range of the telephoto lens becomes larger.

[0277] Optionally, the wide-angle camera and the main camera are opened according to a certain opening rule, for example, when the main object is in the angle range of the main camera, the wide-angle camera is closed; when the main object is not in the angle range of the main camera, the main camera is closed and the wide-angle camera is opened.

[0278] In a feasible embodiment, the above-mentioned adjusting the framing range of the lens can also be sending a prompt message to the user to prompt the user to manually adjust the framing range of the lens, so as to ensure that the determined main object is always in the framing range of the lens.

[0279] Optionally, the above-mentioned prompt message can be a voice message, or can be a text message or other forms of messages displayed on the master control device.

[0280] As can be seen, in the embodiments of the present application, the information of the main object in the image collected by the first camera and the motion information of the camera are used to determine the target panning mode, and then the images collected by the multiple cameras (including the first camera and the second camera, etc.) are fused according to the target panning mode to obtain a target video with the visual effect of the video obtained by using the target panning mode. Instead of simply cutting or sorting the image frames of each video stream and outputting, the present application edits, modifies and fuses the image frames of each video stream according to the main object information and the motion information, and then outputs, which has the characteristics of highlighting the main object, clean background and smooth transition, and can help users to create high-quality videos with good imaging quality, highlighted main object, clean background and harmonious panning for subsequent sharing by using mobile phones online.

[0281] The specific structure of the above-mentioned master control device will be described in detail below.

[0282] Referring to Figure 12 , Figure 12 The structure schematic diagram of the master control device provided in the embodiments of the present application is shown in FIG. 12. As shown in FIG. 12, the master control device 1200 includes: Figure 12

[0283] The acquisition unit 1201 is configured to acquire multiple first images and multiple second images, the multiple first images and the multiple second images being respectively obtained by a first camera and a second camera for a same shooting object, parameters of the first camera and parameters of the second camera being different; and acquire motion information of the first camera in the process of obtaining the multiple first images.

[0284] The acquisition unit 1201 is further configured to acquire main object information in each of the multiple first images.

[0285] The determination unit 1202 is configured to determine a target panning mode according to the motion information of the first camera and the main object information in each of the multiple first images. ​

[0286] The fusing unit 1203 is configured to fuse the plurality of first images and the plurality of second images according to the target panning mode, to obtain a target video, and the panning effect of the target video is the same as that of a video obtained by using the target panning mode.

[0287] In an implementation, the determining unit 1202 is specifically configured to:

[0288] determine whether the first camera has displacement during shooting of the plurality of first images according to the motion information of the first camera; when it is determined that the first camera has no displacement, determine the target panning mode according to the subject object information of each of the plurality of first images; and when it is determined that the first camera has displacement, determine the target panning mode according to the motion mode of the first camera and the subject object information of each of the plurality of first images.

[0289] In an implementation, in the aspect of determining the target panning mode according to the subject object information of each of the plurality of first images, the determining unit 1202 is specifically configured to:

[0290] when the subject object information of each of the plurality of first images is used to indicate that the subject object is contained in each of the plurality of first images, and the first proportion of each of the plurality of first images is less than a first preset proportion, determine the target panning mode as a push lens; wherein the first proportion of each of the plurality of first images is a ratio of an area of a region of interest (ROI) in which the subject object is located in the first image to an area of the first image, or the first proportion of each of the plurality of first images is a ratio of a width of the ROI in which the subject object is located in the first image to a width of the first image, or the first proportion of each of the plurality of first images is a ratio of a height of the ROI in which the subject object is located in the first image to a height of the first image;

[0291] when the subject object information of each of the plurality of first images is used to indicate that the subject object is not contained in a default ROI region in each of the plurality of first images, and an image in the default ROI in each of the plurality of first images is a part of the first image, determine the target panning mode as a pull lens.

[0292] In an implementation, in the aspect of determining the target panning mode according to the motion mode of the first camera and the subject object information of each of the plurality of first images, the determining unit 1202 is specifically configured to:

[0293] when the motion mode of the first camera indicates that the first camera moves in the same direction during shooting of the plurality of first images, and the subject object information of each of the plurality of first images is used to indicate that the subject object changes in the plurality of first images, determine the target panning mode as a shift lens.

[0294] When the motion manner of the first camera indicates that the first camera reciprocates in two opposite directions during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, the target motion manner is determined as a pan shot.

[0295] When the motion manner of the first camera indicates that the first camera moves during the process of capturing the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images, the target motion manner is determined as a follow shot.

[0296] In a possible implementation, when the target motion manner is a push shot or a pull shot, the fusion unit 1203 is specifically configured to:

[0297] obtain a preset region of each of the plurality of first images, the preset region of each of the plurality of first images including the subject object, and when the target motion manner is a push shot, the size of the preset region in the plurality of first images gradually decreases in a time sequence; when the target motion manner is a pull shot, the size of the preset region in the plurality of first images gradually increases in the time sequence; obtain a plurality of first image pairs from the plurality of first images and the plurality of second images, each of the plurality of first image pairs including a first image and a second image with the same time stamp; obtain a plurality of first target images from the plurality of first image pairs, the plurality of first target images corresponding to the plurality of first image pairs in a one-to-one manner; wherein each of the plurality of first target images is the first image pair corresponding thereto with the smallest view angle range, and includes an image in the preset region of the first image; crop, from each of the plurality of first target images, a part overlapping with an image in the preset region of the first image in the second image pair to which the first target image belongs, to obtain a plurality of third images; and process each of the plurality of third images to obtain a plurality of processed third images, each of the plurality of processed third images having a preset resolution, wherein the target video includes the plurality of processed third images.

[0298] In a possible implementation, when the target motion manner is a pan shot, a tilt shot, or a follow shot, the fusion unit 1203 is specifically configured to:

[0299] The subject detection and extraction are performed on the plurality of first images, the ROI of each of the plurality of first images is obtained, and the ROI of each of the plurality of first images is cropped to obtain a plurality of fourth images, wherein each of the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images correspond to each other, and the timestamp of each of the plurality of fourth images is the same as the timestamp of the first image corresponding to the fourth image; each of the plurality of fourth images is processed to obtain a plurality of fifth images, and the resolution of each of the plurality of fifth images is a preset resolution; the plurality of fourth images and the plurality of fifth images correspond to each other, and the timestamp of each of the plurality of fifth images is the same as the timestamp of the fourth image corresponding to the fifth image; a plurality of second image pairs are obtained from the plurality of second images and the plurality of fifth images, each of the plurality of second image pairs includes a fifth image and a second image with the same timestamp; a plurality of second target images are obtained from the plurality of second image pairs, and the plurality of second target images correspond to the plurality of image pairs one by one; wherein for each of the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image; wherein the target video includes the plurality of second target images.

[0300] In one possible implementation, the host device 1200 further includes:

[0301] The display unit 1204 is configured to display the image captured by the first camera, the image captured by the second camera, and the target video.

[0302] The display unit 1204 displays an interface including a first region, a second region, and a third region, the first region displays the image captured by the first camera, the second region is configured to display the image captured by the second camera, and the third region is configured to display the target video, and the FOV of the first camera is greater than the FOV of the second camera.

[0303] Optionally, the first camera is a wide-angle camera or a main camera, and the second camera is a telephoto camera or a main camera.

[0304] Further, the first region includes a fourth region, and the fourth region is configured to display the part of the image captured by the first camera that overlaps the image captured by the second camera.

[0305] In one possible implementation, the fusion unit 1203 is specifically configured to:

[0306] According to the target lens operation mode, the plurality of first images, the plurality of second images and the plurality of sixth images are fused to obtain a target video; wherein the sixth image is obtained by a third camera aiming at a shooting object, the parameters of the third camera are different from the parameters of the first camera, and the parameters of the third camera are different from the parameters of the second camera.

[0307] In a feasible embodiment, when the target lens operation mode is a push lens or a pull lens, the fusion unit 1203 is specifically used for:

[0308] obtaining a preset region of each image in the plurality of first images, the preset region of the plurality of first images includes the main object, and when the target lens operation mode is a push lens, the size of the preset region in the plurality of first images gradually decreases in time sequence; when the target lens operation mode is a pull lens, the size of the preset region in the plurality of first images gradually increases in time sequence; obtaining a plurality of third image pairs from the plurality of first images, the plurality of second images and the plurality of sixth images, each third image pair in the plurality of third image pairs includes a first image, a second image and a sixth image with the same timestamp; obtaining a plurality of third target images from the plurality of third image pairs, the plurality of third target images correspond to the plurality of third image pairs one by one; wherein each third target image in the plurality of third target images is the smallest in the viewing angle range in the corresponding third image pair, and contains the image in the preset region of the first image; cropping the part of each third target image in the plurality of third target images that overlaps with the image in the preset region of the first image in the third image pair to which the third target image belongs to obtain a plurality of seventh images; processing each seventh image in the plurality of seventh images to obtain a plurality of eighth images, and the resolution of each eighth image in the plurality of eighth images is a preset resolution, wherein the target video includes the plurality of eighth images.

[0309] In a feasible embodiment, when the target lens operation mode is a follow lens, a shift lens or a pan lens, the fusion unit 1203 is specifically used for:

[0310] The subject detection and extraction are performed on the plurality of first images, and the ROI of each first image in the plurality of first images is obtained; a plurality of fourth image pairs are obtained from the plurality of first images, the plurality of second images, and the plurality of sixth images, each fourth image pair in the plurality of fourth image pairs includes a first image, a second image, and a sixth image with the same timestamp; a plurality of fourth target images are obtained from the plurality of fourth image pairs, the plurality of fourth target images correspond to the plurality of fourth image pairs in one-to-one correspondence; wherein each fourth target image in the plurality of fourth target images has the smallest view angle range in the corresponding fourth image pair, and contains the image in the ROI of the first image; a part of each fourth target image in the plurality of fourth target images that overlaps with the image in the ROI of the first image in the fourth image pair to which the fourth target image belongs is cropped from the fourth target image to obtain a plurality of ninth images; each ninth image in the plurality of ninth images is processed to obtain a plurality of tenth images, and the resolution of each tenth image in the plurality of tenth images is a preset resolution, wherein the target video includes the plurality of tenth images.

[0311] In one possible implementation, the master device 1200 further includes:

[0312] The display unit 1204 is configured to display the image captured by the first camera, the image captured by the second camera, the image captured by the third camera, and the target video; wherein the FOV of the first camera is greater than the FOV of the second camera, the FOV of the third camera is greater than the FOV of the second camera, and less than the FOV of the first camera,

[0313] The interface displayed by the display unit 1204 includes a first area, a second area, and a third area, the first area includes a fifth area, the fifth area includes a fourth area, the first area is configured to display the image captured by the first camera, the second area is configured to display the image captured by the second camera, the third area is configured to display the target video, the fourth area is configured to display the part of the image captured by the first camera that overlaps with the image captured by the second camera, and the fifth area is configured to display the part of the image captured by the second camera that overlaps with the image captured by the third camera.

[0314] It should be noted that the above-mentioned units (the acquisition unit 1201, the determination unit 1202, the fusion unit 1203, and the display unit 1204) are configured to perform the related steps of the above-mentioned method. For example, the acquisition unit 1201 is configured to perform the related content of steps S301 and S302, the determination unit is configured to perform the related content of S302, the fusion unit 1203 and the display unit 1204 are configured to perform the related content of S304.

[0315] In this embodiment, the host device 1200 is presented in the form of a unit. The "unit" here can refer to an application-specific integrated circuit (ASIC), a processor and a memory executing one or more software or firmware programs, an integrated logic circuit, and / or other devices that can provide the above functions. In addition, the above acquisition unit 1201, determination unit 1202, and fusion unit 1203 can be implemented by the processor 1401 of the host device shown in the figure. Figure 14

[0316] Referring to Figure 13 , Figure 13 Another structural schematic diagram of a host device provided by an embodiment of the present application is shown in the figure. As Figure 13 shown, the host device 1300 includes:

[0317] a camera 1301 configured to acquire a plurality of first images and a plurality of second images.

[0318] The camera 1301 includes a first camera and a second camera, and the plurality of first images and the plurality of second images are respectively obtained by the first camera and the second camera for the same shooting object, and the FOV of the first camera is greater than the FOV of the second camera.

[0319] a control unit 1302 configured to acquire motion information of the first camera during the process of obtaining the plurality of first images, acquire subject object information in each of the plurality of first images, and determine a target dolly mode according to the motion information of the first camera and the subject object information in each of the plurality of first images.

[0320] a fusion unit 1303 configured to fuse the plurality of first images and the plurality of second images according to the target dolly mode to obtain a target video, and the dolly effect of the target video is the same as that of a video obtained by using the target dolly mode.

[0321] a display and saving unit 1304 configured to display and save the first images, the second images, and the target video.

[0322] Optionally, the camera 1301 further includes a third camera, the FOV of the third camera is greater than the FOV of the second camera and less than the FOV of the first camera, and the third camera obtains a plurality of third images for the above shooting object.

[0323] the fusion unit 1303 is configured to fuse the plurality of first images, the plurality of second images, and the plurality of third images according to the target dolly mode to obtain the above target video.

[0324] ​The display and storage unit 1304 is further configured to display and store the third image.

[0325] It is to be noted that the specific implementation of the camera 1301, the control unit 1302, the fusion unit 1303 and the display and storage unit 1304 can refer to the descriptions of steps S301-S304, and will not be described herein.

[0326] As shown in FIG. 14, the master device 1400 can be implemented in the structure of the master device 100 in FIG. 1, and the master device 1400 includes at least one processor 1401, at least one memory 1402 and at least one communication interface 1403. The processor 1401, the memory 1402 and the communication interface 1403 are connected through a communication bus and complete communication with each other. Figure 14 Figure 12 The processor 1401 can be a general central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the above solutions.

[0327] The communication interface 1403 is configured to communicate with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0328] The memory 1402 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.

[0329] The memory 1402 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.

[0330] ​The memory 1402 is configured to store application code for implementing the above solutions, and the processor 1401 is configured to control the execution of the application code.

[0331] The code stored in the memory 1402 can implement the multi-lens video recording method provided above, such as:

[0332] The plurality of first images and the plurality of second images are obtained, the plurality of first images and the plurality of second images are obtained by the first camera and the second camera for the same shooting object, the parameters of the first camera and the parameters of the second camera are different; the motion information of the first camera in the process of obtaining the plurality of first images is obtained; the subject object information in each of the plurality of first images is obtained; the target lens operation mode is determined according to the motion information of the first camera and the subject object information in each of the plurality of first images; and the plurality of first images and the plurality of second images are fused according to the target lens operation mode to obtain a target video, the lens operation effect of the target video is the same as the lens operation effect of the video obtained by using the target lens operation mode.

[0333] Optionally, the above-mentioned host device 1400 further includes a display, configured to display the above-mentioned first image, second image and target video. Since the display is optional, it is not shown in the above-mentioned host device 1400. Figure 14

[0334] The embodiments of the present application also provide a computer storage medium, wherein the computer storage medium can store a program, and the program performs the part or all steps of any one of the multi-lens video recording methods described in the above method embodiments when executed.

[0335] It should be noted that, for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action order described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0336] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0337] ​In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely illustrative, and the division of units can be different, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0338] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0339] In addition, the functional units in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0340] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0341] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc.

[0342] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-lens video recording method, characterized by, The method comprises the following steps: obtaining a plurality of first images and a plurality of second images, wherein the plurality of first images and the plurality of second images are respectively obtained by a first camera and a second camera for a same shooting object, a field of view (FOV) of the first camera is larger than a FOV of the second camera, or a focal length of the first camera is smaller than a focal length of the second camera; obtaining motion information of the first camera during the process of obtaining the plurality of first images; obtaining subject object information in each of the plurality of first images; wherein the motion information of the first camera during the process of obtaining the plurality of first images is obtained by performing feature extraction on each of the plurality of first images to obtain feature points of each of the plurality of first images, matching the feature points of adjacent two images in the plurality of first images to obtain a feature matching result, and determining the motion information of the first camera according to the feature matching result; or the motion information of the first camera is directly obtained from a sensor; determining whether the first camera is displaced during the process of obtaining the plurality of first images according to the motion information of the first camera, determining a target dolly mode according to the subject object information of each of the plurality of first images when the first camera is not displaced, and determining the target dolly mode according to the motion information of the first camera and the subject object information of each of the plurality of first images when the first camera is displaced; wherein the target dolly mode is determined according to the subject object information of each of the plurality of first images, comprising: when the subject object information of each of the plurality of first images indicates that the subject object is contained in each of the plurality of first images, and a first proportion of each of the plurality of first images is less than a first preset proportion, the target dolly mode is determined as a push lens; wherein the first proportion of each of the plurality of first images is a ratio of an area of a region of interest (ROI) in which the subject object is located in the first image to an area of the first image, or the first proportion of each of the plurality of first images is a ratio of a width of the ROI in which the subject object is located in the first image to a width of the first image, or the first proportion of each of the plurality of first images is a ratio of a height of the ROI in which the subject object is located in the first image to a height of the first image; the target dolly mode is determined according to the motion information of the first camera and the subject object information of each of the plurality of first images, comprising: when the motion information of the first camera indicates that the first camera moves in a same direction during the process of obtaining the plurality of first images, and the subject object information of each of the plurality of first images is used to indicate that the subject object changes in the plurality of first images, the target dolly mode is determined as a dolly lens; fusing the plurality of first images and the plurality of second images according to the target dolly mode to obtain a target video.

2. The method of claim 1, wherein, The determining the target lens movement mode according to the subject object information of each of the plurality of first images further includes: When the subject object information of each of the plurality of first images indicates that no subject object is contained in a default ROI region in the each of the plurality of first images, and the image in the default ROI in the each of the plurality of first images is a part of the each of the plurality of first images, the target lens movement mode is determined as the zoom-in lens.

3. The method of claim 1, wherein, The determining the target lens movement mode according to the motion information of the first camera and the subject object information of each of the plurality of first images further includes: When the motion information of the first camera indicates that the first camera reciprocates in two opposite directions during the shooting of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images, the target lens movement mode is determined as the pan lens. When the motion information of the first camera indicates that the first camera moves during the shooting of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images, the target lens movement mode is determined as the follow lens.

4. The method of claim 2, wherein, When the target lens movement mode is the zoom-in lens or the zoom-out lens, the fusing the plurality of first images and the plurality of second images according to the target lens movement mode to obtain the target video includes: obtaining a preset region of each of the plurality of first images, the preset region of each of the plurality of first images containing the subject object, and when the target lens movement mode is the zoom-in lens, the size of the preset region in the plurality of first images gradually decreases in time sequence; when the target lens movement mode is the zoom-out lens, the size of the preset region in the plurality of first images gradually increases in time sequence; obtaining a plurality of first image pairs from the plurality of first images and the plurality of second images, each of the plurality of first image pairs containing a first image and a second image with the same timestamp; obtaining a plurality of first target images from the plurality of first image pairs, the plurality of first target images corresponding to the plurality of first image pairs one by one; wherein each of the plurality of first target images is the first image pair corresponding to the each of the plurality of first target images with the smallest view angle range and contains the image in the preset region of the first image; cropping a part of each of the plurality of first target images that overlaps with the preset region of the first image in the second image pair to which the each of the plurality of first target images belongs to obtain a plurality of third images; processing each of the plurality of third images to obtain a plurality of processed third images, the resolution of each of the plurality of processed third images being a preset resolution, wherein the target video includes the plurality of processed third images.

5. The method of claim 3, wherein, When the target panning mode is the lens moving, the lens shaking or the lens following, the fusing the plurality of first images and the plurality of second images according to the target panning mode to obtain the target video comprises: subject detection and extraction are performed on the plurality of first images, the ROI of each first image in the plurality of first images is obtained, and the ROI of each first image in the plurality of first images is cropped to obtain a plurality of fourth images, wherein each fourth image in the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images correspond to each other, and the timestamp of each fourth image in the plurality of fourth images is the same as the timestamp of the first image corresponding to the fourth image; each fourth image in the plurality of fourth images is processed to obtain a plurality of fifth images, and the resolution of each fifth image in the plurality of fifth images is a preset resolution; the plurality of fourth images and the plurality of fifth images correspond to each other, and the timestamp of each fifth image in the plurality of fifth images is the same as the timestamp of the fourth image corresponding to the fifth image a plurality of second image pairs are obtained from the plurality of second images and the plurality of fifth images, each second image pair in the plurality of second image pairs includes a fifth image and a second image with the same timestamp; a plurality of second target images are obtained from the plurality of second image pairs, and the plurality of second target images correspond to the plurality of image pairs one by one; wherein for each second image pair in the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image; wherein the target video includes the plurality of second target images.

6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: displaying the image captured by the first camera, the image captured by the second camera and the target video on a display interface; wherein the display interface includes a first area, a second area and a third area, the first area displays the image captured by the first camera, the second area is used to display the image captured by the second camera, and the third area is used to display the target video.

7. The method of claim 6, wherein, The first area includes a fourth area, and the fourth area is used to display the part of the image captured by the first camera that overlaps with the image captured by the second camera.

8. The method according to any one of claims 1 to 3, characterized in that, The fusing the plurality of first images and the plurality of second images according to the target panning mode to obtain the target video comprises: fusing the plurality of first images, the plurality of second images and a plurality of sixth images according to the target panning mode to obtain the target video; wherein the sixth image is obtained by a third camera aiming at the shooting object, the parameters of the third camera are different from the parameters of the first camera, and the parameters of the third camera are different from the parameters of the second camera.

9. The method of claim 8, wherein, When the target panning mode is the push or pull mode, the fusing the plurality of first images, the plurality of second images and the plurality of sixth images according to the target panning mode to obtain the target video comprises: acquiring a preset region of each of the plurality of first images, the preset region of each of the plurality of first images including the subject object, and when the target panning mode is the push mode, the size of the preset region of each of the plurality of first images gradually decreases in time sequence, and when the target panning mode is the pull mode, the size of the preset region of each of the plurality of first images gradually increases in time sequence; acquiring a plurality of third image pairs from the plurality of first images, the plurality of second images and the plurality of sixth images, each of the plurality of third image pairs including a first image, a second image and a sixth image with the same timestamp; acquiring a plurality of third target images from the plurality of third image pairs, the plurality of third target images corresponding to the plurality of third image pairs one by one, wherein each of the plurality of third target images is the one with the smallest view angle range in the corresponding third image pair and includes the image in the preset region of the first image; cropping a part of each of the plurality of third target images that overlaps with the image in the preset region of the first image in the third image pair to which the third target image belongs to obtain a plurality of seventh images; processing each of the plurality of seventh images to obtain a plurality of eighth images, each of the plurality of eighth images having a preset resolution, wherein the target video includes the plurality of eighth images.

10. The method of claim 8, wherein, When the target panning mode is the follow, shift or pan mode, the fusing the plurality of first images, the plurality of second images and the plurality of sixth images according to the target panning mode to obtain the target video comprises: performing subject detection and extraction on the plurality of first images to acquire a ROI of each of the plurality of first images; acquiring a plurality of fourth image pairs from the plurality of first images, the plurality of second images and the plurality of sixth images, each of the plurality of fourth image pairs including a first image, a second image and a sixth image with the same timestamp; acquiring a plurality of fourth target images from the plurality of fourth image pairs, the plurality of fourth target images corresponding to the plurality of fourth image pairs one by one, wherein each of the plurality of fourth target images is the one with the smallest view angle range in the corresponding fourth image pair and includes the image in the ROI of the first image; cropping a part of each of the plurality of fourth target images that overlaps with the image in the ROI of the first image in the fourth image pair to which the fourth target image belongs to obtain a plurality of ninth images; processing each of the plurality of ninth images to obtain a plurality of tenth images, each of the plurality of tenth images having a preset resolution, The target video includes the plurality of tenth images.

11. The method of claim 8, wherein, The method further includes: displaying the image captured by the first camera, the image captured by the second camera, the image captured by the third camera, and the target video on a display interface; The display interface includes a first region, a second region, and a third region, the first region includes a fifth region, and the fifth region includes a fourth region; The first region displays the image captured by the first camera, the second region is configured to display the image captured by the second camera, the third region is configured to display the target video, and the fourth region is configured to display the portion of the image captured by the first camera that overlaps the image captured by the second camera; and the fifth region is configured to display the portion of the image captured by the second camera that overlaps the image captured by the third camera.

12. A master device, comprising: includes: an acquisition unit configured to acquire a plurality of first images and a plurality of second images, the plurality of first images and the plurality of second images being captured by a first camera and a second camera for a same shooting object, a field of view (FOV) of the first camera being greater than a FOV of the second camera, or a focal length of the first camera being less than a focal length of the second camera; and acquire motion information of the first camera during capturing of the plurality of first images; The acquisition unit is further configured to acquire subject object information in each of the plurality of first images; wherein acquiring the motion information of the first camera during capturing of the plurality of first images includes: performing feature extraction on each of the plurality of first images to obtain feature points of each of the plurality of first images; matching the feature points of adjacent two images in the plurality of first images to obtain a feature matching result; and determining the motion information of the first camera according to the feature matching result; or directly acquiring the motion information of the first camera from a sensor; A determination unit is configured to determine, according to the motion information of the first camera, whether the first camera has moved during capturing of the plurality of first images, wherein when it is determined that the first camera has not moved, a target panning mode is determined according to the subject object information of each of the plurality of first images; and when it is determined that the first camera has moved, the target panning mode is determined according to the motion information of the first camera and the subject object information of each of the plurality of first images; wherein the determination of the target panning mode according to the subject object information of each of the plurality of first images includes: determining the target lens movement mode as a zoom-in lens when the subject object information of each of the plurality of first images indicates that the subject object is contained in each of the plurality of first images, and the first proportion of each of the plurality of first images is less than a first preset proportion, wherein the first proportion of each of the first images is a ratio of an area of a region of interest (ROI) in which the subject object is located in the first image to an area of the first image, or the first proportion of each of the first images is a ratio of a width of the ROI in which the subject object is located in the first image to a width of the first image, or the first proportion of each of the first images is a ratio of a height of the ROI in which the subject object is located in the first image to a height of the first image; determining the target lens movement mode according to the motion information of the first camera and the subject object information of each of the plurality of first images, comprising: determining the target lens movement mode as a zoom-out lens when the motion information of the first camera indicates that the first camera moves in the same direction during acquisition of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images; a fusion unit configured to fuse the plurality of first images and the plurality of second images according to the target lens movement mode to obtain a target video.

13. The apparatus of claim 12, wherein, In the aspect of determining the target lens movement mode according to the subject object information of each of the plurality of first images, the determining unit is further configured to: determining the target lens movement mode as a zoom-out lens when the subject object information of each of the plurality of first images indicates that the subject object is not contained in a default ROI region in each of the first images, and the image in the default ROI in each of the first images is part of the first image.

14. The apparatus of claim 12, wherein, In the aspect of determining the target lens movement mode according to the motion information of the first camera and the subject object information of each of the plurality of first images, the determining unit is further configured to: determining the target lens movement mode as a zoom-in lens when the motion information of the first camera indicates that the first camera reciprocally moves in two opposite directions during acquisition of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object changes in the plurality of first images; determining the target lens movement mode as a zoom-out lens when the motion information of the first camera indicates that the first camera moves during acquisition of the plurality of first images, and the subject object information of each of the plurality of first images indicates that the subject object does not change in the plurality of first images.

15. The apparatus of claim 13, wherein, when the target lens movement mode is the zoom-in lens or the zoom-out lens, the fusion unit is specifically configured to: obtaining a preset region of each of the plurality of first images, the preset region of each of the plurality of first images including the subject object, and when the target panning mode is the zoom-in, the size of the preset region of the plurality of first images gradually decreases in time sequence; when the target panning mode is the zoom-out, the size of the preset region of the plurality of first images gradually increases in time sequence; obtaining a plurality of first image pairs from the plurality of first images and the plurality of second images, each of the plurality of first image pairs including a first image and a second image with the same timestamp; obtaining a plurality of first target images from the plurality of first image pairs, the plurality of first target images corresponding to the plurality of first image pairs one by one; wherein each of the plurality of first target images is the first image with the smallest view angle range in the corresponding first image pair, and contains the image in the preset region of the first image; cropping a part of each of the plurality of first target images that overlaps with the image in the preset region of the first image in the second image pair to which the first target image belongs to obtain a plurality of third images; processing each of the plurality of third images to obtain a plurality of processed third images, the resolution of each of the plurality of processed third images being a preset resolution, wherein the target video includes the plurality of processed third images.

16. The apparatus of claim 14, wherein, When the target panning mode is the shift, the pan or the follow, the fusion unit is specifically configured to: perform subject detection and extraction on the plurality of first images to obtain an ROI of each of the plurality of first images, and crop the ROI of each of the plurality of first images to obtain a plurality of fourth images, wherein each of the plurality of fourth images includes a subject object; the plurality of first images and the plurality of fourth images correspond to each other, and the timestamp of each of the plurality of fourth images is the same as the timestamp of the corresponding first image; process each of the plurality of fourth images to obtain a plurality of fifth images, the resolution of each of the plurality of fifth images being a preset resolution; the plurality of fourth images and the plurality of fifth images correspond to each other, and the timestamp of each of the plurality of fifth images is the same as the timestamp of the corresponding fourth image obtaining a plurality of second image pairs from the plurality of second images and the plurality of fifth images, each of the plurality of second image pairs including a fifth image and a second image with the same timestamp; obtaining a plurality of second target images from the plurality of second image pairs, the plurality of second target images corresponding to the plurality of image pairs one by one; wherein for each second image pair in the plurality of second image pairs, when the content of the fifth image in the second image pair coincides with the content of the second image, the second target image corresponding to the second image pair is the second image; when the content of the fifth image in the second image pair does not coincide with the content of the second image, the second target image corresponding to the image pair is the fifth image; wherein the target video includes the plurality of second target images.

17. The apparatus of any of claims 12-16, wherein, The device further includes: a display unit configured to display the image captured by the first camera, the image captured by the second camera, and the target video; wherein the interface displayed by the display unit includes a first region, a second region, and a third region, the first region displays the image captured by the first camera, the second region is configured to display the image captured by the second camera, and the third region is configured to display the target video.

18. The apparatus of claim 17, wherein, The first region includes a fourth region configured to display the part of the image captured by the first camera that overlaps with the image captured by the second camera.

19. The apparatus of any one of claims 12-14, wherein, The fusion unit is specifically configured to: fuse the plurality of first images, the plurality of second images, and a plurality of sixth images according to the target lens operation mode to obtain the target video; wherein the sixth images are captured by a third camera for the shooting object, the parameters of the third camera are different from those of the first camera, and the parameters of the third camera are different from those of the second camera.

20. The apparatus of claim 19, wherein, When the target lens operation mode is the push lens or the pull lens, the fusion unit is specifically configured to: obtain a preset region of each image in the plurality of first images, the preset region of the plurality of first images includes the main object, and when the target lens operation mode is the push lens, the size of the preset region in the plurality of first images gradually decreases in time sequence; when the target lens operation mode is the pull lens, the size of the preset region in the plurality of first images gradually increases in time sequence; obtain a plurality of third image pairs from the plurality of first images, the plurality of second images, and the plurality of sixth images, each third image pair in the plurality of third image pairs includes a first image, a second image, and a sixth image with the same timestamp; obtain a plurality of third target images from the plurality of third image pairs, the plurality of third target images correspond to the plurality of third image pairs one by one; wherein each third target image in the plurality of third target images is the one with the smallest viewing angle range in the third image pair corresponding to the third target image and includes the image in the preset region of the first image; crop the part of each third target image in the plurality of third target images that overlaps with the image in the preset region of the first image in the third image pair to which the third target image belongs to obtain a plurality of seventh images; process each of the plurality of seventh images to obtain a plurality of eighth images, each of the plurality of eighth images having a preset resolution, wherein the target video comprises the plurality of eighth images.

21. The apparatus of claim 19, wherein, When the target lens movement mode is a follow shot, a pan shot, or a tilt shot, the fusion unit is specifically configured to: perform subject detection and extraction on the plurality of first images to obtain a ROI of each of the plurality of first images; obtain a plurality of fourth image pairs from the plurality of first images, the plurality of second images, and the plurality of sixth images, each of the plurality of fourth image pairs comprising a first image, a second image, and a sixth image having the same timestamp; obtain a plurality of fourth target images from the plurality of fourth image pairs, the plurality of fourth target images corresponding to the plurality of fourth image pairs in one-to-one correspondence; wherein each of the plurality of fourth target images has the smallest range of view angles in the corresponding fourth image pair and contains an image within the ROI of the first image; crop, from each of the plurality of fourth target images, a portion overlapping an image within the ROI of the first image in the fourth image pair to which the fourth target image belongs, to obtain a plurality of ninth images; process each of the plurality of ninth images to obtain a plurality of tenth images, each of the plurality of tenth images having a preset resolution, wherein the target video comprises the plurality of tenth images.

22. The apparatus of claim 19, wherein, The device further comprises: a display unit configured to display the image captured by the first camera, the image captured by the second camera, the image captured by the third camera, and the target video. The interface displayed by the display unit comprises a first area, a second area, and a third area, the first area comprises a fifth area, the fifth area comprises a fourth area, the first area displays the image captured by the first camera, the second area is configured to display the image captured by the second camera, the third area is configured to display the target video, and the fourth area is configured to display a portion of the image captured by the first camera overlapping the image captured by the second camera; the fifth area is configured to display a portion of the image captured by the second camera overlapping the image captured by the third camera.

23. An electronic device comprising a touch screen, a memory, one or more processors; wherein, The memory stores one or more programs; when the one or more processors execute the one or more programs, the electronic device implements the method of any one of claims 1 to 11.

24. A computer storage medium, comprising, The computer program product comprises computer instructions, when the computer instructions are run on an electronic device, the electronic device executes the method of any one of claims 1 to 11.

25. A computer program product, characterised in that, When the computer program product is run on a computer, the computer executes the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and apparatus for taking images using mobile communication terminal with plurality of camera lenses

    CN101090442A

  • Video processing method, electronic equipment and storage medium

    CN111083380A