Video shooting method and device and electronic equipment

By enabling the field-of-view angle expansion recording function in electronic devices and using IMU data to process images, the problem of a smaller field-of-view angle during video shooting is solved, video recording with a larger field-of-view angle is achieved, and the reliability and effect of video shooting are improved.

CN120676244APending Publication Date: 2025-09-19VIVO MOBILE COMM CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510746194.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

During video shooting, in order to ensure stability, electronic devices usually prevent shaking by cropping the video size, which results in a smaller field of view and reduces the reliability of video shooting.

Method used

By responding to the control input of the field of view expansion control, the field of view expansion recording function is started, and the first video is recorded so that its field of view is larger than the second field of view supported by the camera. The IMU data is used for image processing to achieve video recording with a larger field of view.

Benefits of technology

The reliability of video shooting by electronic devices is improved, more shooting scenes can be included, and the video recording effect is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676244A_ABST
    Figure CN120676244A_ABST
Patent Text Reader

Abstract

The invention discloses a video shooting method and device and electronic equipment, and belongs to the technical field of camera equipment. The method comprises the following steps: in response to control input of a field angle expansion control on a shooting preview interface, starting a field angle expansion video recording function; performing video recording in response to the video recording control input, and outputting a first video; wherein the first field angle of the first video is greater than a second field angle supported by a camera for recording the first video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of camera equipment, and specifically relates to a video shooting method, device and electronic equipment. Background Art

[0002] With the development of camera technology, users can use the cameras of electronic devices to shoot videos in different scenarios.

[0003] Generally speaking, when shooting videos, in order to ensure the stability of the video shooting, electronic devices usually perform video stabilization by cropping the video size. However, cropping the video size will cause the field of view of the video to become smaller. Therefore, the reliability of the video shot by the electronic device is low. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a video shooting method, device and electronic device, which can improve the reliability of electronic devices in shooting videos.

[0005] In a first aspect, an embodiment of the present application provides a video shooting method, the method comprising: starting a field of view angle expansion recording function in response to a control input of a field of view angle expansion control on a shooting preview interface; performing video recording and outputting a first video in response to a recording control input; wherein a first field of view angle of the first video is greater than a second field of view angle supported by a camera recording the first video.

[0006] In a second aspect, an embodiment of the present application provides a video shooting device, including: a processing module, which starts a field of view expansion recording function in response to a control input of a field of view expansion control on a shooting preview interface; and records a video in response to a recording control input, and outputs a first video; wherein the first field of view of the first video is greater than the second field of view supported by the camera recording the first video.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0011] In an embodiment of the present application, in response to a control input to a field-of-view expansion control on a shooting preview interface, a field-of-view expansion recording function is activated; in response to the recording control input, video recording is performed and a first video is output; wherein the first field-of-view of the first video is greater than the second field-of-view supported by the camera that recorded the first video. In this solution, since the field-of-view expansion recording function is activated, the field-of-view of the first video obtained by the electronic device during video recording can be greater than the field-of-view supported by the camera when recording the first video, that is, the first video can include more shooting scenes, thereby improving the reliability of video recording by the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic diagram of coordinate axes related to image mapping provided in some embodiments of the present application;

[0013] Figure 2 is a schematic diagram of coordinate axes related to image mapping provided in some embodiments of the present application;

[0014] Figure 3 is a flowchart of a video shooting method provided by some embodiments of the present application;

[0015] Figure 4A This is a schematic diagram of an example of inputting a video control provided by some embodiments of the present application;

[0016] Figure 4B This is a schematic diagram of an example of inputting a field of view expansion control provided by some embodiments of the present application;

[0017] Figure 4C is a schematic diagram of an example of inputting a recording control provided in some embodiments of the present application;

[0018] Figure 5A This is a schematic diagram of an example of inputting a video control provided by some embodiments of the present application;

[0019] Figure 5B This is a schematic diagram of an example of inputting a field of view expansion control provided by some embodiments of the present application;

[0020] Figure 5C is a schematic diagram of an example of inputting a recording control provided in some embodiments of the present application;

[0021] Figure 6Ais a schematic diagram of an example of a two-dimensional image provided by some embodiments of the present application;

[0022] Figure 6B is a schematic diagram of a three-dimensional sphere texture image provided by some embodiments of the present application;

[0023] Figure 6C is a schematic diagram of an example of a two-dimensional image provided by some embodiments of the present application;

[0024] Figure 6D is a schematic diagram of a three-dimensional sphere texture image provided by some embodiments of the present application;

[0025] Figure 7A is a schematic diagram of an example of a two-dimensional image provided by some embodiments of the present application;

[0026] Figure 7B is a schematic diagram of a three-dimensional sphere texture image provided by some embodiments of the present application;

[0027] Figure 7C is a schematic diagram of an example of a two-dimensional image provided by some embodiments of the present application;

[0028] Figure 7D is a schematic diagram of a three-dimensional sphere texture image provided by some embodiments of the present application;

[0029] Figure 8 is a schematic diagram of a fused three-dimensional spherical texture image provided by some embodiments of the present application;

[0030] Figure 9 is a schematic diagram of a fused three-dimensional spherical texture image provided by some embodiments of the present application;

[0031] Figure 10A is a schematic diagram of an example of a two-dimensional video image with an expanded field of view provided by some embodiments of the present application;

[0032] Figure 10B is a schematic diagram of an example of a two-dimensional video image with an expanded field of view provided by some embodiments of the present application;

[0033] Figure 11A is a schematic diagram of an example of a two-dimensional video image with an expanded field of view provided by some embodiments of the present application;

[0034] Figure 11B is a schematic diagram of an example of a two-dimensional video image with an expanded field of view provided by some embodiments of the present application;

[0035] Figure 12A This is a schematic diagram of an example of inputting an expansion ratio identifier provided in some embodiments of the present application;

[0036] Figure 12B is a schematic diagram of an example of inputting an expansion ratio control provided by some embodiments of the present application;

[0037] Figure 13A This is a schematic diagram of an example of inputting an expansion ratio identifier provided in some embodiments of the present application;

[0038] Figure 13B is a schematic diagram of an example of inputting an expansion ratio control provided by some embodiments of the present application;

[0039] Figure 14 is a schematic diagram of expanding the field of view provided by some embodiments of the present application;

[0040] Figure 15 is a schematic diagram of an example of a two-dimensional panoramic image provided by some embodiments of the present application;

[0041] Figure 16 is a schematic diagram of an example of a two-dimensional panoramic image provided by some embodiments of the present application;

[0042] Figure 17 is a schematic structural diagram of a video shooting device provided by some embodiments of the present application;

[0043] Figure 18 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application;

[0044] Figure 19 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION

[0045] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0046] The terms "first," "second," and the like in the specification of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification indicates at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0047] The terms "at least one" and "at least one of" in the specification of this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".

[0048] The terms used in the embodiments of this application are only used to explain the specific embodiments of this application and are not intended to limit this application. The following is a glossary of professional terms involved in the embodiments of this application:

[0049] Two-dimensional image: It is two-dimensional data composed of pixels on a plane, which only describes the planar features such as color and pattern of the object surface.

[0050] Three-dimensional texture image: It is volume data composed of voxels (three-dimensional pixel points) in three-dimensional space, which can simultaneously describe the detailed features of the surface and interior of an object.

[0051] The three-dimensional world coordinates are mapped to the two-dimensional plane coordinate system:

[0052] like Figure 1 As shown, P(X w ,Y w ,Z w ) is the world coordinate system O w -X w Y w Z w A little inside, O c -X c Y c Z c is the camera coordinate system, the uv plane is the camera imaging plane, and the imaging point corresponding to point P is p(x,y); the transformation from the world coordinate system to the camera coordinate system is a rigid body change, that is, the object will not deform, only rotate and translate, so we can get:

[0053] The coordinates of point P in the camera coordinate system for:

[0054] Where R is a 3×3 rotation matrix and T is a 3×1 translation matrix. Assume that the rotation angle of the camera coordinate relative to the world coordinate system is (θ x ,θ y ,θ z ), the translation offset is (t x ,t y ,t z), since the variation between frames is small, it can be simplified as R= Since the coordinate system is relative, in general, we will assume that the world coordinate system is the camera coordinate system when the first frame is input. Therefore, the subsequent rotation angles and translation offsets are the rotation and translation offsets of the Nth frame relative to the first frame, which can be obtained through the IMU.

[0055] like Figure 2 As shown, it can be seen that △ABO C ~△oCO C , △PBO C ~△pCO C ;

[0056] therefore,

[0057] therefore,

[0058] therefore,

[0059] Assuming that the center of the imaging plane is the origin, and since the distance from the camera coordinate system to the imaging plane is a fixed focal length f, we can obtain the formula for mapping the three-dimensional world coordinate system to the two-dimensional plane coordinate system based on the principle of similar triangles:

[0060]

[0061] Inertial Measurement Unit (IMU) data: This refers to data collected by IMU sensors that describes the motion and posture of electronic devices. IMU data includes at least one of the following: gyroscope (Gyro) data, Hall effect (Hall) data, and acceleration (ACC) data.

[0062] Gyro data: refers to the angular velocity data around the x-axis, y-axis, and z-axis of the electronic device measured by the gyroscope sensor of the electronic device.

[0063] Hall data refers to data related to the magnetic field measured by the Hall sensor of an electronic device. Specifically, the Hall sensor of an electronic device can measure the intensity information of the surrounding magnetic field on the x-axis, y-axis, and z-axis. This intensity information is the Hall data mentioned above.

[0064] ACC data: refers to acceleration data around the x-axis, y-axis, and z-axis of the electronic device measured by an ACC sensor of the electronic device. The ACC sensor may be a three-axis acceleration sensor.

[0065] The video shooting method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0066] The video shooting method in the embodiment of the present application can be applied to the scene of shooting video.

[0067] For example, it can be used to shoot a video of a performance by cartoon characters when going to an amusement park; or it can be used to shoot a video of the scenery on the top of a mountain when climbing a mountain.

[0068] The video shooting method provided in the embodiment of the present application may be executed by a video shooting device, which may be an electronic device, or a functional module or entity in the electronic device. The technical solution provided in the embodiment of the present application is described below using an electronic device as an example.

[0069] The embodiment of the present application provides a video shooting method, Figure 3 FIG. 1 shows a flow chart of a video shooting method provided by an embodiment of the present application, which can be applied to electronic devices. Figure 3 As shown, the video shooting method provided in the embodiment of the present application may include the following steps 201 and 202.

[0070] Step 201: The electronic device activates a field-of-view expansion recording function in response to a control input to a field-of-view expansion control on a shooting preview interface.

[0071] In some embodiments of the present application, the electronic device may receive a user's control input on a field of view expansion control on a shooting preview interface, and activate a field of view expansion recording function in response to the control input.

[0072] In some embodiments of the present application, the control input is used to enable a field of view expansion recording function of the electronic device.

[0073] In some embodiments of the present application, the control input includes, but is not limited to, touch input of the field of view expansion control by a user through a touch device such as a finger or a stylus, or voice commands input by the user, or specific gestures input by the user, or click input, or other feasible input. The specific input can be determined based on actual use needs and is not limited in the embodiments of the present application.

[0074] In some embodiments of the present application, the above-mentioned specific gesture can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double-press gesture, and a double-click gesture.

[0075] In some embodiments of the present application, the click input may be a single-click input, a double-click input, or any number of click inputs, and may also be a long-press input or a short-press input.

[0076] In some embodiments of the present application, the electronic device can receive user input on the camera application icon on the desktop, start the camera application, and display the above-mentioned shooting preview interface.

[0077] In some embodiments of the present application, the shooting preview interface may be a shooting preview interface in a video recording mode.

[0078] Specifically, after starting the camera application, the electronic device can receive the user's input on the recording control in the shooting preview interface, so that the electronic device displays the shooting preview interface in the recording mode, which includes the field of view angle expansion control. The electronic device can then receive the user's control input on the field of view angle expansion control to start the field of view angle expansion recording function.

[0079] In some embodiments of the present application, when the field of view expansion recording function is activated, the electronic device can adjust the display form of the field of view expansion control to remind the user that the field of view expansion recording function has been activated.

[0080] For example, the display mode of the field of view expansion control may include at least one of the following: thickening the border of the field of view expansion control, flashing the field of view expansion control, and magnifying the field of view expansion control. The specific mode may be determined based on actual usage needs and is not limited in this embodiment of the present application.

[0081] Step 202: The electronic device records a video in response to a video recording control input and outputs a first video.

[0082] In some embodiments of the present application, the video recording control input may be an input by the user to a recording control in a shooting preview interface.

[0083] In some embodiments of the present application, the video recording control input is used to record video.

[0084] In some embodiments of the present application, after starting the field of view angle expansion recording function, the electronic device can receive the user's input to the recording control in the shooting preview interface, perform video recording, and then end the video recording and output the first video when receiving the user's input to the recording control in the shooting preview interface again.

[0085] In an embodiment of the present application, the first field of view angle of the above-mentioned first video is greater than the second field of view angle supported by the camera recording the first video.

[0086] In some embodiments of the present application, the second field of view angle supported by the camera recording the first video may be the maximum field of view angle supported by the camera.

[0087] In some embodiments of the present application, the second field of view angle supported by the camera recording the first video can also be understood as: the field of view angle adopted by the electronic device when recording the first video.

[0088] Specifically, the field of view angle used by the electronic device when recording the first video can be any one of the following: the maximum field of view angle supported by the camera, the default field of view angle, and the field of view angle selected by the user when displaying the shooting preview interface in the video recording mode.

[0089] For example: the second field of view angle can be the maximum field of view angle supported by the camera, 140°, or it can be the field of view angle corresponding to the 1.5X zoom factor selected by the user when recording the first video; and the first field of view angle of the first video can be greater than the second field of view angle, for example, 160°.

[0090] In some examples, when displaying a shooting preview interface in a video recording mode, the electronic device may receive user input of a target identifier for at least one identifier displayed in the shooting preview interface, and record the video using the field of view angle indicated by the target identifier.

[0091] Exemplarily, each of the at least one identifier indicates a shooting field of view angle.

[0092] It is understandable that after the electronic device starts the field of view expansion recording function, the electronic device can process the recorded video so that the field of view of the video is greater than the field of view supported by the camera of the electronic device.

[0093] Example 1: Based on scenario 1, when a user goes to an amusement park and wants to shoot a video of a cartoon character performance, the user can click the camera application icon on the desktop to activate the camera application on the electronic device and shoot the video. Figure 4A As shown, the shooting preview interface 21 is displayed, and then the user can click the video control 22 in the shooting preview interface 21, as shown in FIG. Figure 4B As shown, the electronic device displays a shooting preview interface 23 in the recording mode, and the shooting preview interface 23 includes a field of view expansion control 24. When the user wants to shoot a video with a larger field of view and to include more scenes, the user can click the field of view expansion control 24 in the shooting preview interface 23, that is, the above-mentioned control input, so that the electronic device starts the field of view expansion recording function, such as Figure 4C As shown, when the electronic device starts the field of view angle expansion recording function, the display size of the field of view angle expansion control 24 becomes larger; then the user can click the recording control 25 in the shooting preview interface 23 to enable the electronic device to record the video and output the first video with a field of view angle greater than the field of view supported by the camera.

[0094] Example 2: In combination with scenario 2, when a user goes hiking, if he wants to take a video of the scenery on the top of the mountain, the user can click the camera application icon on the desktop to enable the electronic device to start the camera application and take a video of the scenery on the top of the mountain. Figure 5A As shown, the shooting preview interface 26 is displayed, and then the user can click the video control 27 in the shooting preview interface 26, as shown in FIG. Figure 5B As shown, the electronic device displays a shooting preview interface 28 in the recording mode, and the shooting preview interface 28 includes a field of view expansion control 29. When the user wants to shoot a video with a larger field of view and to include more scenes, the user can click the field of view expansion control 29 in the shooting preview interface 28, that is, the above-mentioned control input, so that the electronic device starts the field of view expansion recording function, such as Figure 5C As shown, when the electronic device starts the field of view angle expansion recording function, the display size of the field of view angle expansion control 29 becomes larger; then the user can click the recording control 30 in the shooting preview interface 28 to enable the electronic device to record the video and output the first video with a field of view angle greater than the field of view supported by the camera.

[0095] In the video shooting method provided in the embodiment of the present application, after the field of view angle expansion recording function is started, the field of view angle of the first video obtained by the electronic device for video recording can be greater than the field of view angle supported by the camera when recording the first video, that is, the first video can include more shooting scenes, thereby improving the reliability of the electronic device in shooting videos.

[0096] In some embodiments of the present application, the “electronic device performs video recording and outputs the first video” in the above step 202 can be specifically implemented through the following steps 202a to 202e.

[0097] Step 202a: During the video recording process, the electronic device collects N frames of images and IMU data corresponding to each frame of the N frames of images.

[0098] In some embodiments of the present application, each of the N frames of images captured by the electronic device is a two-dimensional image.

[0099] In some embodiments of the present application, the above-mentioned IMU data includes at least one of the following: Gyro data, Hall data, and ACC data.

[0100] Step 202b: The electronic device generates N three-dimensional spherical texture images of N frames of images based on each frame of image and the IMU data corresponding to each frame of image.

[0101] In an embodiment of the present application, the above-mentioned one frame of image corresponds to a three-dimensional spherical texture image.

[0102] In some embodiments of the present application, the above step 202b can be specifically implemented through the following steps 202b1 to 202b4.

[0103] Step 202b1: The electronic device calculates the rotation offset and the translation offset of each frame of image based on the IMU data corresponding to each frame of image.

[0104] In some embodiments of the present application, the rotation offset of each frame of image may be understood as the rotation offset of each frame of image relative to the first frame of image in N frames of image.

[0105] In some embodiments of the present application, the aforementioned translation offset of each frame of image may be understood as the translation offset of each frame of image relative to the first frame of image in N frames of image.

[0106] In some embodiments of the present application, the electronic device can determine the rotation offset of each frame of image relative to the first frame of image captured based on the Gyro data and Hall data corresponding to each frame of image to obtain the rotation offset of each frame of image.

[0107] Specifically, the electronic device can determine the offset of the electronic device movement of each frame of image relative to the first frame of image captured based on the Gyro data corresponding to each frame of image, and determine the offset of the lens compensation of each frame of image relative to the first frame of image captured based on the Hall data corresponding to each frame of image. The electronic device can then subtract the offset of the electronic device movement from the offset of the lens compensation to obtain the rotational offset of each frame of image relative to the first frame of image captured.

[0108] In some embodiments of the present application, the electronic device can determine the Gyro data acquisition range corresponding to each frame of image based on the image acquisition time of each frame of image, the time offset of the Gyro data relative to the image acquisition time of each frame of image, the exposure duration of each frame of image and the camera rolling shutter exposure duration of each frame of image. The electronic device can then collect the Gyro data corresponding to each frame of image within the Gyro data acquisition range corresponding to each frame of image.

[0109] In some embodiments of the present application, the time offset of the gyro data relative to the image acquisition time of each frame of the N frames of image is the same.

[0110] In some embodiments of the present application, the Gyro data acquisition range corresponding to the Nth frame image can be [T N -e N +T G ,T N +rs+T G ], N is a positive integer. Among them, T N is the image acquisition time of the Nth frame image, e Nis the exposure time of the Nth frame image, T G is the time offset of the Gyro data relative to the image acquisition time of the Nth frame image, and rs is the camera rolling shutter exposure duration of the Nth frame image.

[0111] For example, assuming that the image acquisition time of the Nth frame is 10:02:20, the exposure time of the Nth frame is e N The time offset T of the Gyro data relative to the image acquisition time of the Nth frame is 2 seconds. G The rolling shutter exposure time rs of the Nth frame is 1 second, and the Gyro data acquisition range corresponding to the Nth frame can be [10:02:19, 10:02:22]. The electronic device can acquire the Gyro data corresponding to the Nth frame between 10:02:19 and 10:02:22.

[0112] In some embodiments of the present application, each frame of image may correspond to n sets of Gyro data, and each set of Gyro data may be represented as in, Indicates the Gyro data of the i-th group of x-axis corresponding to a frame of image, Represents the Gyro data of the i-th group of y-axis corresponding to a frame of image, Indicates the Gyro data of the i-th group of z-axis corresponding to a frame of image, Indicates the acquisition time of the i-th group of Gyro data corresponding to a frame of image.

[0113] In some embodiments of the present application, the electronic device may determine the offset of the electronic device on the x-axis motion of the Nth frame image relative to the first frame image captured based on the Gyro data corresponding to the Nth frame image and formula (1), where formula (1) is as follows:

[0114]

[0115] Among them, s Nx Indicates the offset of the electronic device on the x-axis relative to the first frame of image captured. Indicates the Gyro data of the i-th group of x-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Gyro data.

[0116] In some embodiments of the present application, the electronic device may determine the offset of the electronic device on the y-axis motion of the Nth frame image relative to the first frame image captured based on the Gyro data corresponding to the Nth frame image and formula (2), where formula (2) is as follows:

[0117]

[0118] Among them, sNy Indicates the offset of the electronic device on the y-axis relative to the first frame of image captured. Indicates the Gyro data of the i-th group of y-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Gyro data.

[0119] In some embodiments of the present application, the electronic device may determine the offset of the electronic device in the z-axis motion of the Nth frame image relative to the first frame image captured based on the Gyro data corresponding to the Nth frame image and formula (3), where formula (3) is as follows:

[0120]

[0121] Among them, s Nz Indicates the offset of the electronic device moving along the z-axis relative to the first frame of image captured. Indicates the Gyro data of the i-th group of z-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Gyro data.

[0122] For example, the offset of the electronic device movement of the Nth frame image relative to the first frame image captured can be expressed as:

[0123]

[0124] In some embodiments of the present application, the electronic device can determine the Hall data acquisition range corresponding to each frame of image based on the image acquisition time of each frame of image, the time offset of the Hall data relative to the image acquisition time of each frame of image, the exposure duration of each frame of image and the camera rolling shutter exposure duration of each frame of image, and then the electronic device can collect the Hall data corresponding to each frame of image within the Hall data acquisition range corresponding to each frame of image.

[0125] In some embodiments of the present application, the time offset of the Hall data relative to the image acquisition time of each frame of the N frames of image is the same.

[0126] In some embodiments of the present application, the Hall data acquisition range corresponding to the Nth frame image may be [T N -e N +T H ,T N +rs+T H ]. Among them, T N is the image acquisition time of the Nth frame image, e N is the exposure time of the Nth frame image, T H is the time offset of the Hall data relative to the image acquisition time of the Nth frame image, and rs is the camera rolling shutter exposure duration of the Nth frame image.

[0127] For example, assuming that the image acquisition time of the Nth frame is 08:05:10, the exposure time of the Nth frame is e N The time offset T of the Hall data relative to the image acquisition time of the Nth frame is 1.5 seconds. H The camera rolling shutter exposure time rs of the Nth frame image is 0.5 seconds, and the Hall data acquisition range corresponding to the Nth frame image can be [08:05:09, 08:05:11]. The electronic device can acquire the Hall data corresponding to the Nth frame image between 08:05:09 and 08:05:11.

[0128] In some embodiments of the present application, each frame of image may correspond to k sets of Hall data, and each set of Hall data may be represented as in, Indicates the Hall data of the i-th group of x-axis corresponding to a frame of image, Indicates the Hall data of the i-th group of y-axis corresponding to a frame of image, Indicates the Hall data of the i-th group of z-axis corresponding to a frame of image, Indicates the acquisition time of the i-th group of Hall data corresponding to a frame of image.

[0129] In some embodiments of the present application, the electronic device may determine the offset of the lens compensation of the Nth frame image relative to the first frame image captured on the x-axis based on the Hall data corresponding to the Nth frame image and formula (4), where formula (4) is as follows:

[0130]

[0131] Among them, h Nx Indicates the offset of the lens compensation of the Nth frame image relative to the first frame image captured on the x-axis. Indicates the Hall data of the i-th group of x-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Hall data.

[0132] In some embodiments of the present application, the electronic device may determine the lens compensation offset of the Nth frame image relative to the first frame image captured on the y-axis based on the Hall data corresponding to the Nth frame image and formula (5), where formula (5) is as follows:

[0133]

[0134] Among them, h Ny Indicates the offset of the lens compensation of the Nth frame image relative to the first frame image captured on the y-axis. Indicates the Hall data of the i-th group of y-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Hall data.

[0135] In some embodiments of the present application, the electronic device may determine the lens compensation offset of the Nth frame image relative to the first frame image captured on the z-axis based on the Hall data corresponding to the Nth frame image and formula (6), where formula (6) is as follows:

[0136]

[0137] Among them, h Nz Indicates the offset of the lens compensation of the Nth frame image relative to the first frame image captured on the z-axis. Indicates the Hall data of the i-th group of z-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of Hall data.

[0138] For example, the offset of the lens compensation of the Nth frame image relative to the first captured frame image can be expressed as:

[0139]

[0140] In some embodiments of the present application, the electronic device may determine the translation offset of each frame image relative to the first captured frame image based on the ACC data.

[0141] In some embodiments of the present application, the electronic device can determine the ACC data acquisition range corresponding to each frame of image based on the image acquisition time of each frame of image, the time offset of the ACC data relative to the image acquisition time of each frame of image, the exposure duration of each frame of image, and the camera rolling shutter exposure duration of each frame of image. The electronic device can then collect the ACC data corresponding to each frame of image within the ACC data acquisition range corresponding to each frame of image.

[0142] In some embodiments of the present application, the time offset of the ACC data relative to the image acquisition time of each frame of the N frames of image is the same.

[0143] In some embodiments of the present application, the ACC data acquisition range corresponding to the Nth frame image can be [T N -e N +T A ,T N +rs+T A ]. Among them, T N is the image acquisition time of the Nth frame image, e N is the exposure time of the Nth frame image, T Ais the time offset of the ACC data relative to the image acquisition time of the Nth frame image, and rs is the camera rolling shutter exposure duration of the Nth frame image.

[0144] For example, assuming that the image acquisition time of the Nth frame is 06:20:10, the exposure time of the Nth frame is e N The time offset T of ACC data relative to the image acquisition time of the Nth frame is 0.5 seconds. A The exposure time of the camera rolling shutter of the N-th frame image is 0.5 seconds, and the exposure time rs of the camera rolling shutter of the N-th frame image is 1 second. Then the acquisition range of the ACC data corresponding to the N-th frame image can be [06:20:10, 06:20:11]. The electronic device can acquire the ACC data corresponding to the N-th frame image between 06:20:10 and 06:20:11.

[0145] In some embodiments of the present application, each frame of image may correspond to h groups of ACC data, and each group of ACC data may be represented as in, Indicates the ACC data of the i-th group of x-axis corresponding to a frame of image, Indicates the ACC data of the i-th group of y-axis corresponding to a frame of image, Indicates the ACC data of the i-th group of z-axis corresponding to a frame of image, Indicates the acquisition time of the i-th group of ACC data corresponding to a frame of image.

[0146] In some embodiments of the present application, the electronic device may determine the x-axis translation offset of the N-th frame image relative to the first frame image acquired based on the ACC data corresponding to the N-th frame image and formula (7), where formula (7) is as follows:

[0147]

[0148] Among them, a Nx Indicates the translation offset of the Nth frame image relative to the first frame image acquired on the x-axis. Indicates the ACC data of the i-th group of x-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of ACC data.

[0149] In some embodiments of the present application, the electronic device may determine the translation offset of the Nth frame image relative to the first frame image captured on the y-axis based on the ACC data corresponding to the Nth frame image and formula (8), where formula (8) is as follows:

[0150]

[0151] Among them, a Ny Indicates the translation offset of the Nth frame image relative to the first frame image acquired on the y-axis. Indicates the ACC data of the i-th group of y-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of ACC data.

[0152] In some embodiments of the present application, the electronic device may determine the translation offset of the Nth frame image relative to the first frame image captured on the z-axis based on the ACC data corresponding to the Nth frame image and formula (9), where formula (9) is as follows:

[0153]

[0154] Among them, a Nz Indicates the translation offset of the Nth frame image relative to the first frame image acquired on the z-axis. Indicates the ACC data of the i-th group of z-axis corresponding to the N-th frame image, Indicates the acquisition time of the i-th group of ACC data.

[0155] For example, the translation offset of the Nth frame image relative to the first captured frame image can be expressed as:

[0156]

[0157] Step 202b2: The electronic device maps each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system, to obtain an initial three-dimensional spherical texture image corresponding to each frame image.

[0158] In some embodiments of the present application, the electronic device can map each two-dimensional pixel in each frame image to a three-dimensional space based on the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel in each frame image on the optical axis of the camera coordinate system, to obtain the three-dimensional pixel points corresponding to each two-dimensional pixel in each frame image, so as to obtain the initial three-dimensional spherical texture image corresponding to each frame image.

[0159] In some embodiments of the present application, the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system can be expressed as Z c express.

[0160] In some embodiments of the present application, the electronic device may map the i-th two-dimensional pixel point in the N-th frame image to the three-dimensional space according to formula (10), and obtain the three-dimensional pixel point corresponding to the i-th two-dimensional pixel point in the N-th frame image. Formula (10) is as follows:

[0161]

[0162] Among them, (x i ,y i ) represents the i-th two-dimensional pixel in the N-th frame image, (X wi ,Y wi ,Z wi ) represents the three-dimensional pixel point corresponding to the i-th two-dimensional pixel point in the N-th frame image, f represents the focal length used by the camera recording the video, R represents the rotation matrix, T represents the translation matrix, Z c Represents the coordinate value of the i-th two-dimensional pixel point in the N-th frame image on the optical axis of the camera coordinate system.

[0163] For example, (θ x ,θ y ,θ z ) represents the rotation offset of the Nth frame image relative to the first frame image acquired.

[0164] For example, (t x ,t y ,t z ) is represented as the translation offset of the Nth frame image relative to the first frame image acquired.

[0165] For example, taking i=10 as an example, calculate the 3D pixel corresponding to the 10th 2D pixel in the Nth frame image. Assuming that the focal length f used by the camera recording the video is 50mm, the 10th 2D pixel is (100,150), Z c =500mm, the rotation offset is (0.1, 0.2, 0.3), and the translation offset is (10, 20, 30), then Substituting these values ​​into formula (10), we can obtain that the 3D pixel corresponding to the 10th 2D pixel is (120, 180, 60).

[0166] It should be noted that, since the formula for mapping a three-dimensional world coordinate system to a two-dimensional plane coordinate system is known, when the three-dimensional world coordinate system is known, the above formula (10) can be used to find the mapping of the two-dimensional plane coordinate system to the three-dimensional world coordinate system.

[0167] In this way, since each two-dimensional pixel point in each frame image can be mapped to a three-dimensional space through the above formula (10), the initial three-dimensional spherical texture image corresponding to each frame image can be obtained.

[0168] In some embodiments of the present application, after obtaining the three-dimensional pixel points corresponding to each two-dimensional pixel point in each frame image, the electronic device can use the color of each two-dimensional pixel point as the color of the three-dimensional pixel point corresponding to each two-dimensional pixel point to obtain the initial three-dimensional spherical texture image corresponding to each frame image.

[0169] For example, the color of the pixel (x, y) in the Nth frame image can be expressed as I(x, y, n), and the three-dimensional pixel point (X w ,Y w ,Z w ) can be expressed as c(X w ,Y w ,Z w )=I(x,y,n).

[0170] Step 202b3: The electronic device determines a rotation and translation error parameter value based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image.

[0171] In some embodiments of the present application, the above step 202b3 can be specifically implemented through the following steps 301 and 302.

[0172] Step 301: The electronic device predicts the image color value of each frame of image based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image.

[0173] In some embodiments of the present application, for a frame of image, after obtaining the initial three-dimensional spherical texture image corresponding to the frame of image, the electronic device can predict the image color value of the next frame of image of the frame by the image color value of the initial three-dimensional spherical texture image corresponding to the frame of image, thereby obtaining the predicted image color value of each frame of image, and then the electronic device can determine the rotation and translation error parameter value based on the predicted image color value of each frame of image and the actual image color value of each frame of image.

[0174] In some embodiments of the present application, for a first frame of image among N frames of image, the electronic device may use the actual image color value of the first frame of image as the predicted image color value of the first frame of image.

[0175] In some embodiments of the present application, the electronic device can predict the image color value of the next frame image of a frame image by the image color value of the initial three-dimensional spherical texture image corresponding to the frame image. Specifically, the electronic device can predict the pixel coordinates of the two-dimensional pixel points corresponding to each three-dimensional pixel point in the next frame image after mapping each three-dimensional pixel point in the initial three-dimensional spherical texture image corresponding to the frame image to the two-dimensional space based on the image color value of the initial three-dimensional spherical texture image corresponding to the frame image. Then, the electronic device can use the color of each three-dimensional pixel point as the color of the two-dimensional pixel point corresponding to the three-dimensional pixel point in the next frame image to predict the image color value of the next frame image of the frame image.

[0176] In some embodiments of the present application, the electronic device can subtract the rotation offset of a frame image relative to the first captured frame image from the rotation offset of the next frame image relative to the first captured frame image, that is, subtract the rotation offset of a frame image from the rotation offset of the next frame image to obtain the rotation offset between a frame image and the next frame image of the frame image.

[0177] In some embodiments of the present application, the electronic device can subtract the translation offset of a frame image relative to the first captured frame image from the translation offset of the next frame image relative to the first captured frame image, that is, subtract the translation offset of a frame image from the translation offset of the next frame image to obtain the translation offset between a frame image and the next frame image of the frame image.

[0178] In some embodiments of the present application, the electronic device may predict the pixel coordinates of the two-dimensional pixel corresponding to the i-th three-dimensional pixel in the N-th frame image after mapping the i-th three-dimensional pixel in the initial three-dimensional spherical texture image corresponding to the N-1-th frame image to the two-dimensional space by using the following formula (11). Formula (11) is as follows:

[0179]

[0180] in, represents the i-th three-dimensional pixel point in the initial three-dimensional spherical texture image corresponding to the N-1-th frame image, represents the 2D pixel corresponding to the ith 3D pixel in the Nth frame image, f represents the focal length used by the camera recording the video, R N represents the rotation matrix, T N represents the translation matrix, Z c Represents the coordinate value of the i-th two-dimensional pixel point in the N-th frame image on the optical axis of the camera coordinate system.

[0181] For example, (θ xN ,θ yN ,θ zN ) represents the rotation offset between the N-1th frame image and the Nth frame image; It is expressed as the translation offset between the N-1th frame image and the Nth frame image.

[0182] Among them, (γ0,γ1,γ2,γ3,γ4,γ5) is the first variable.

[0183] For example, taking i=5 as an example, after mapping the 3D pixel corresponding to the 5th 3D pixel in the initial 3D spherical texture image corresponding to the N-1th frame image to the 2D space, the pixel coordinates of the 2D pixel corresponding to the 5th 3D pixel in the Nth frame image are predicted. Assuming that the focal length f used by the camera recording the video is 30mm, the 5th 3D pixel is (200, 300, 400), Z c =500mm, the rotation offset between the N-1th frame image and the Nth frame image is (0.2, 0.3, 0.1), and the translation offset between the N-1th frame image and the Nth frame image is (10, 20, 30), then Substituting these values ​​into formula (11), we can obtain the pixel coordinates of the 2D pixel corresponding to the 5th 3D pixel in the Nth frame image as (18, 15.6).

[0184] In this way, since the above formula (11) can be used to predict the pixel coordinates of the two-dimensional pixel points corresponding to each three-dimensional pixel point in the N-1 frame image after mapping each three-dimensional pixel point in the initial three-dimensional spherical texture image corresponding to the N-1 frame image to the two-dimensional space, and the color of each three-dimensional pixel point is used as the color of the two-dimensional pixel point corresponding to the three-dimensional pixel point in the next frame image, the image color value of the N-1 frame image can be predicted, and the rotation and translation error parameter value can be determined based on the predicted image color value of each frame image and the actual image color value of each frame image.

[0185] Step 302: The electronic device determines a rotation and translation error parameter value based on the predicted image color value of each frame of image and the actual image color value of each frame of image.

[0186] In some embodiments of the present application, when the image color value of the next frame image of a frame image is predicted for the first time through the image color value of the initial three-dimensional spherical texture image corresponding to the frame image, the variable value of the first variable is the initial value, and the variable value of the first variable is adjusted once each prediction until the color difference between the image color value of each frame image predicted by the electronic device and the actual image color value of each frame image is less than a preset threshold value, and then the electronic device can use the variable value of the first variable adjusted for the last time as the rotation and translation error parameter value, and then the electronic device can obtain the three-dimensional spherical texture image corresponding to each frame image based on the rotation and translation error parameter value.

[0187] For example, the electronic device may calculate the color difference between the image color value of each frame of the image predicted by the electronic device and the actual image color value of each frame of the image by using the following formula (12), which is as follows:

[0188]

[0189] Among them, E represents the color difference, I i (x, y) represents the actual image color value of the i-th frame image of the electronic device, I i ′(x, y) represents the image color value of the i-th frame image predicted by the electronic device.

[0190] Step 202b4: The electronic device maps each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation and translation error parameter value, the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system, to obtain a three-dimensional spherical texture image corresponding to each frame image.

[0191] In some embodiments of the present application, the electronic device may map the i-th two-dimensional pixel point in the N-th frame image to the three-dimensional space according to formula (13), and obtain the three-dimensional pixel point corresponding to the i-th two-dimensional pixel point in the N-th frame image. Formula (13) is as follows:

[0192]

[0193] Among them, (x i ,y i ) represents the i-th two-dimensional pixel in the N-th frame image, (X wi ,Y wi ,Z wi ) represents the 3D pixel corresponding to the i-th 2D pixel in the N-th frame image, f represents the focal length used by the camera recording the video, R represents the rotation matrix, and T represents the translation matrix.

[0194] For example, (θ x ,θ y ,θ z ) represents the rotation offset of the Nth frame image relative to the first frame image acquired; (t x ,t y ,t z ) represents the translation offset of the Nth frame image relative to the first frame image acquired, and (γ0,γ1,γ2,γ3,γ4,γ5) represent the rotation and translation error parameters.

[0195] For example, taking i=3 as an example, calculate the 3D pixel corresponding to the third 2D pixel in the Nth frame image. Assume that the focal length f used by the camera recording the video is 20mm, the third 2D pixel is (80,60), Z c =300mm, the rotation offset is (0.1, 0.2, 0.3), and the translation offset is (5, 10, 15), then Substituting these values ​​into formula (13), we can obtain the 3D pixel point corresponding to the third 2D pixel point as (100, 120, 150).

[0196] In this way, since the electronic device can calculate the three-dimensional pixel point corresponding to each two-dimensional pixel point through formula (13), it can obtain an accurate three-dimensional spherical texture image corresponding to each frame image, thereby avoiding the occurrence of artifacts in the video after the field of view is expanded, and ensuring the continuity of the video image after the field of view is expanded.

[0197] Example 3, combined with Example 1, assume that the user shoots a video of a cartoon character's performance, which includes two frames of images, including Figure 6A For a frame of image 31 shown, the electronic device can obtain the following information based on the IMU data corresponding to the image 31: Figure 6B The three-dimensional sphere texture image corresponding to the image 31 shown, and the Figure 6C For a frame of image 32 shown in FIG. 3 , the electronic device can obtain the following information based on the IMU data corresponding to the image 32: Figure 6D The image 32 shown corresponds to a three-dimensional sphere texture image.

[0198] Example 4, combined with Example 2, assume that the user shoots a video of the scenery on the top of a mountain, including two frames of images, including Figure 7A The electronic device can obtain the IMU data corresponding to the image 41 according to the image 41. Figure 7B The three-dimensional sphere texture image corresponding to the image 41 shown, and the Figure 7C The electronic device can obtain the following information based on the IMU data corresponding to the frame image 42: Figure 7D The image 42 shown corresponds to a three-dimensional sphere texture image.

[0199] Step 202c: The electronic device selects a set of three-dimensional pixel points corresponding to each frame of image from the N three-dimensional spherical texture images based on the first field of view.

[0200] In the embodiment of the present application, each three-dimensional pixel point set includes at least one three-dimensional pixel.

[0201] In some embodiments of the present application, the above-mentioned first field of view angle can be a default field of view angle, a preset field of view angle, or a field of view angle selected by the user. It can be specifically determined according to actual usage requirements, and the embodiments of the present application are not limited to this.

[0202] In some embodiments of the present application, the electronic device may perform fusion processing on the N three-dimensional spherical texture images, and then select a three-dimensional pixel point set corresponding to each frame of the image from the fused three-dimensional spherical texture images.

[0203] Example 5, combined with Example 3, such as Figure 8 As shown, a fused three-dimensional spherical texture image is obtained by fusion processing of a three-dimensional spherical texture image corresponding to a frame image 31 and a three-dimensional spherical texture image corresponding to a frame image 32.

[0204] Example 6, combined with Example 4, such as Figure 9 As shown, a fused three-dimensional spherical texture image is obtained by fusion processing of a three-dimensional spherical texture image corresponding to a frame image 41 and a three-dimensional spherical texture image corresponding to a frame image 42.

[0205] In some embodiments of the present application, the above step 202c can be specifically implemented through the following steps 202c1 and 202c2.

[0206] Step 202c1: The electronic device determines the field of view projection range corresponding to each frame of the N three-dimensional spherical texture images according to the first field of view angle and the focal length of the camera used for recording the video.

[0207] In some embodiments of the present application, the electronic device can fuse the N three-dimensional spherical texture images mentioned above, and then determine the field of view angle projection range corresponding to each frame image from the fused three-dimensional spherical texture image based on the first field of view angle and the focal length used by the camera recording the video.

[0208] In some embodiments of the present application, the above-mentioned field of view angle projection range is an area on the fused three-dimensional spherical texture image.

[0209] In some embodiments of the present application, the above-mentioned field of view angle projection range can be understood as: when the camera of the electronic device actually adopts the above-mentioned first field of view angle and the above-mentioned focal length to shoot a video, the area that the camera of the electronic device can capture.

[0210] In some embodiments of the present application, the field of view angle projection range corresponding to each frame image can be expressed as Where P is the field of view expansion coefficient, and θ is the second field of view supported by the camera recording the video.

[0211] For example, assuming that the field of view supported by the camera recording the video is θ, and the first field of view is 1.1*θ, the field of view projection range corresponding to each frame of the image can be expressed as [-f*0.55*θ,f*0.55*θ].

[0212] Step 202c2: The electronic device determines the pixel coordinates of all three-dimensional pixels within the field of view projection range corresponding to each frame of image as a three-dimensional pixel set corresponding to each frame of image.

[0213] It is understandable that the electronic device can obtain all three-dimensional pixels within the field of view projection range corresponding to each frame of image, and then determine the pixel coordinates of all three-dimensional pixels as a three-dimensional pixel set corresponding to each frame of image.

[0214] Step 202d: The electronic device maps the three-dimensional pixel points in the three-dimensional pixel point set corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image.

[0215] In an embodiment of the present application, the field of view angle of the two-dimensional video image corresponding to each frame of image is a first field of view angle.

[0216] Example 7, combined with Example 3, such as Figure 10A As shown, image 51 is a two-dimensional video image corresponding to image 31, as shown in FIG. Figure 10B As shown, image 52 is a two-dimensional video image corresponding to image 32 , wherein the field of view angle of image 51 and image 52 is 1.1θ, ie, the first field of view angle mentioned above.

[0217] Example 8, combined with Example 4, such as Figure 11A As shown, image 61 is a two-dimensional video image corresponding to image 41, as shown in FIG. Figure 11B As shown, image 62 is a two-dimensional video image corresponding to image 42 , wherein the field of view angle of image 61 and image 62 is 1.3θ, which is the first field of view angle mentioned above.

[0218] In some embodiments of the present application, for a detailed description of an electronic device mapping three-dimensional pixel points in a three-dimensional pixel point set corresponding to each frame of an image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of an image, please refer to the related art, in which the electronic device performs two-dimensional mapping processing on the three-dimensional image to obtain a related description of the two-dimensional image, which will not be repeated here.

[0219] Step 202e: The electronic device performs image synthesis processing on the two-dimensional video images corresponding to the N frames of images according to the image acquisition sequence of each frame of image to obtain a first video.

[0220] In some embodiments of the present application, the electronic device may use video synthesis technology to perform image synthesis processing on two-dimensional video images corresponding to N frames of images in the order of image acquisition to obtain a first video.

[0221] In some embodiments of the present application, the electronic device may encode N frames of images and N three-dimensional spherical texture images into the first video.

[0222] In some embodiments of the present application, in addition to outputting the first video, the electronic device can also perform synthesis processing on N frames of images in the order of image acquisition to obtain a second video, and encode the three-dimensional spherical texture image corresponding to each frame of the N frames of images into the second video.

[0223] In some embodiments of the present application, the electronic device may store the first video in association with the second video.

[0224] This allows electronic devices to use IMU data to align the spatial positions of 2D images when acquiring a 3D image. This reduces computational costs and allows for faster results. Furthermore, the ability to expand the field of view of any video segment increases the flexibility of video processing.

[0225] In some embodiments of the present application, before the above step 202c, the video shooting method provided in the embodiment of the present application further includes the following steps 401 to 404.

[0226] Step 401: When video recording is finished, the electronic device performs synthesis processing on N frames of images according to the image acquisition sequence of the N frames of images to obtain a second video.

[0227] In an embodiment of the present application, the field of view of the second video is a second field of view.

[0228] In some embodiments of the present application, the electronic device may encode a three-dimensional spherical texture image corresponding to each frame of N frames of images into the second video.

[0229] Step 402: When displaying the video preview interface of the second video, the electronic device receives a first input from the user on an expansion ratio control in the video preview interface.

[0230] In some embodiments of the present application, the first input is used to obtain a video with an expanded field of view.

[0231] In some examples, the expansion ratio control can be a sliding control including a slide track and a slider; the electronic device can receive the user's sliding input for the expansion ratio control, control the slider to slide on the slide track, and then determine the field of view angle expansion coefficient according to the expansion ratio indicated by the slider.

[0232] In some examples, the expansion ratio control includes an input box, and the electronic device can receive user input into the input box and determine the field of view angle expansion coefficient based on the expansion ratio entered by the user in the input box.

[0233] For example, the expansion ratio may be 10%, 20%, 30%, etc.

[0234] In some embodiments of the present application, before receiving the first input, the electronic device may display an expansion ratio control in the video preview interface in response to the user's second input of the expansion ratio identifier in the video preview interface, and then determine the field of view angle expansion coefficient in response to the user's first input of the expansion ratio control.

[0235] Step 403: The electronic device determines a field of view angle expansion coefficient in response to the first input.

[0236] In some embodiments of the present application, the electronic device may add the above expansion ratio to 1 to obtain the above field of view angle expansion coefficient.

[0237] For example, when the expansion ratio is 10%, the expansion coefficient may be: 1+10%=1.1.

[0238] Example 9, combined with Example 3, after the user shoots a video of a performance of an animated character and obtains images 31 and 32, the electronic device can perform synthesis processing on images 31 and 32 in the order of image acquisition to obtain a second video; then Figure 12A As shown, when the video preview interface 71 of the second video is displayed, the user can click the expansion ratio mark 72 in the video preview interface 71, as shown in FIG. Figure 12B As shown, the electronic device displays the expansion ratio control 73 on the video preview interface 71; then the user can slide the expansion ratio control 73, that is, the first input mentioned above, to control the slider to slide on the slide rail so that the slider points to 10%, and then the electronic device can determine the field of view angle expansion coefficient according to the expansion ratio indicated by the slider.

[0239] Example 10, combined with Example 4, after the user shoots a video of the scenery on the top of a mountain and obtains image 41 and image 42, the electronic device can perform synthesis processing on image 41 and image 42 in the order of image acquisition to obtain a second video; then Figure 13A As shown, when the video preview interface 81 of the second video is displayed, the user can click the expansion ratio mark 82 in the video preview interface 81, as shown in FIG. Figure 13B As shown, the electronic device displays the expansion ratio control 83 on the video preview interface 81; then the user can slide the expansion ratio control 83, that is, the first input mentioned above, to control the slider to slide on the slide rail so that the slider points to 30%, and then the electronic device can determine the field of view angle expansion coefficient according to the expansion ratio indicated by the slider.

[0240] Step 404: The electronic device multiplies the field of view angle expansion coefficient by the second field of view angle to obtain the first field of view angle.

[0241] For example, when the second field of view angle supported by the camera recording the video is θ and the field of view angle expansion coefficient is 1.1, the first field of view angle may be 1.1*θ.

[0242] Example 11: Figure 14 As shown, the second field of view angle supported by the camera of the electronic device for recording video can be the solid line part, and the first field of view angle can be the dotted line part.

[0243] In this way, since the user can flexibly select the field of view expansion ratio, the electronic device can process videos at different field of view angles according to the field of view expansion ratio selected by the user, thereby improving the flexibility of video processing.

[0244] It is understood that after generating N three-dimensional spherical texture images for N frames of images, the electronic device can directly select a set of three-dimensional pixels corresponding to each frame of image from the N three-dimensional spherical texture images based on a default first field of view angle; map the three-dimensional pixels in the set of three-dimensional pixels corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image; and then perform image synthesis processing on the two-dimensional video images corresponding to the N frames of image according to the image acquisition order of each frame of image to output a first video. Alternatively, after video recording is completed, the electronic device can first perform image synthesis processing on the N frames of image according to the image acquisition order of the N frames of image to output a second video; then determine the first field of view angle based on the user input of the expansion ratio control in the video preview interface of the second video; and then select a set of three-dimensional pixels corresponding to each frame of image from the N three-dimensional spherical texture images based on the first field of view angle; and map the three-dimensional pixels in the set of three-dimensional pixels corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image; and then perform image synthesis processing on the two-dimensional video images corresponding to the N frames of image according to the image acquisition order of each frame of image to output the first video.

[0245] In some embodiments of the present application, the video shooting method provided in the embodiments of the present application further includes the following steps 501 and 502.

[0246] Step 501: When the field of view angle expansion coefficient is greater than or equal to the coefficient threshold, the electronic device selects a three-dimensional pixel point set corresponding to any frame image from N three-dimensional spherical texture images based on the first field of view angle.

[0247] In some embodiments of the present application, the electronic device can determine whether the field of view expansion coefficient is greater than or equal to the coefficient threshold, so that when the field of view expansion coefficient is greater than or equal to the coefficient threshold, based on the first field of view, a three-dimensional pixel point set corresponding to any frame image can be selected from N three-dimensional spherical texture images.

[0248] In some embodiments of the present application, the electronic device can select a three-dimensional pixel point set corresponding to any frame image from N three-dimensional spherical texture images based on the first field of view when the above-mentioned expansion ratio is greater than or equal to the expansion ratio threshold.

[0249] Step 502: The electronic device maps the three-dimensional pixel points in the three-dimensional pixel point set corresponding to any frame image to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any frame image.

[0250] Example 12, combined with Example 3, when the expansion ratio selected by the user is 100%, the electronic device can map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to the image 31 to the two-dimensional space, such as Figure 15 As shown, a two-dimensional panoramic image 91 corresponding to the image 31 is obtained.

[0251] Example 13, combined with Example 4, when the expansion ratio selected by the user is 100%, the electronic device can map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to the image 42 to the two-dimensional space, such as Figure 16 As shown, a two-dimensional panoramic image 92 corresponding to the image 42 is obtained.

[0252] In some embodiments of the present application, for a detailed description of an electronic device mapping three-dimensional pixel points in a three-dimensional pixel point set corresponding to any frame image to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any frame image, please refer to the related technology, where the electronic device performs two-dimensional mapping processing on the three-dimensional image to obtain the relevant description of the two-dimensional image, which will not be repeated here.

[0253] In this way, since when the field of view angle expansion coefficient selected by the user is greater than or equal to the coefficient threshold, the electronic device can select a three-dimensional pixel point set corresponding to any one of the N frames of collected images from N three-dimensional spherical texture images based on the larger field of view angle corresponding to the field of view angle expansion coefficient. Therefore, after mapping the three-dimensional pixel points in the three-dimensional pixel point set to two-dimensional space, a two-dimensional panoramic image corresponding to any one frame of image can be obtained, thereby improving the flexibility of video processing. At the same time, since the electronic device restores the collected two-dimensional images to three-dimensional space for splicing, the continuity of the above-mentioned two-dimensional panoramic image is guaranteed, and problems such as artifacts in the two-dimensional panoramic image are avoided.

[0254] It should be noted that the above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or, under the premise that there is no contradiction, can also be executed in combination with each other. The specific implementation can be determined according to actual usage requirements, and the embodiments of this application do not limit this.

[0255] It should be noted that the video shooting method provided in the embodiment of the present application can be executed by a video shooting device. In the embodiment of the present application, the video shooting device provided in the embodiment of the present application is described by taking the video shooting method performed by the video shooting device as an example.

[0256] Figure 17 FIG. 1 shows a possible structural diagram of a video shooting device involved in an embodiment of the present application. Figure 17 As shown, the video shooting device 70 may include: a processing module 71.

[0257] Among them, the processing module 71 is used to start the field of view expansion recording function in response to the control input of the field of view expansion control on the shooting preview interface; and in response to the recording control input, perform video recording and output the first video; wherein the first field of view of the first video is greater than the second field of view supported by the camera recording the first video.

[0258] An embodiment of the present application provides a video shooting device. After the field of view angle expansion recording function is activated, the field of view angle of the first video obtained by the electronic device when recording the video can be greater than the field of view angle supported by the camera when recording the first video, that is, the first video can include more shooting scenes, thereby improving the reliability of the electronic device in shooting videos.

[0259] In some possible implementations, the processing module 71 is specifically used to collect N frames of images and IMU data corresponding to each frame of the N frames during video recording; and based on each frame of image and the IMU data corresponding to each frame of image, generate N three-dimensional spherical texture images of the N frames of image, wherein one frame of image corresponds to one three-dimensional spherical texture image; and based on the first field of view, select a three-dimensional pixel point set corresponding to each frame of image from the N three-dimensional spherical texture images; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image, and the field of view of the two-dimensional video image corresponding to each frame of image is the first field of view angle; and perform image synthesis processing on the two-dimensional video images corresponding to the N frames of image according to the image acquisition order of each frame of image to obtain a first video.

[0260] In some possible implementations, the processing module 71 is specifically used to calculate the rotation offset and translation offset of each frame image based on the IMU data corresponding to each frame image; and map each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system to obtain an initial three-dimensional spherical texture image corresponding to each frame image; and determine the rotation and translation error parameter value based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame image; and map each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation and translation error parameter value, the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system to obtain a three-dimensional spherical texture image corresponding to each frame image.

[0261] In some possible implementations, the processing module 71 is specifically used to predict the image color value of each frame of image based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image; and determine the rotation and translation error parameter value based on the predicted image color value of each frame of image and the actual image color value of each frame of image.

[0262] In some possible implementations, the processing module 71 is specifically used to determine the field of view angle projection range corresponding to each frame image from N three-dimensional spherical texture images based on the first field of view angle and the focal length used by the camera recording the video; and determine the pixel coordinates of all three-dimensional pixel points within the field of view angle projection range corresponding to each frame image as the three-dimensional pixel point set corresponding to each frame image.

[0263] In some possible implementations, the video shooting device provided in the embodiments of the present application further includes: a receiving module;

[0264] The processing module 71 is further configured to, before selecting a set of three-dimensional pixels corresponding to each frame of the N three-dimensional spherical texture images based on the first field of view, perform synthesis processing on the N frames of images according to an image acquisition order of the N frames of images after video recording is completed to obtain a second video having a field of view of the second video at a second field of view angle.

[0265] a receiving module, configured to receive a first input from a user on an expansion ratio control in the video preview interface when the video preview interface of the second video is displayed;

[0266] The processing module 71 is further configured to determine a field of view angle expansion coefficient in response to the first input received by the receiving module; and multiply the field of view angle expansion coefficient by the second field of view angle to obtain a first field of view angle.

[0267] In some possible implementations, the processing module 71 is also used to select a three-dimensional pixel point set corresponding to any frame image from N three-dimensional spherical texture images based on the first field of view when the field of view angle expansion coefficient is greater than or equal to the coefficient threshold; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to any frame image to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any frame image.

[0268] The video capture device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0269] The video capture device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0270] The video shooting device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be described here.

[0271] Alternatively, as Figure 18 As shown, an embodiment of the present application also provides an electronic device 900, including a processor 901 and a memory 902, wherein the memory 902 stores a program or instruction that can be run on the processor 901, and when the program or instruction is executed by the processor 901, the various steps of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0272] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0273] Figure 19 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0274] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .

[0275] Those skilled in the art will understand that the electronic device 100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 19 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0276] Among them, the processor 110 is used to start the field of view expansion recording function in response to the control input of the field of view expansion control on the shooting preview interface; and in response to the recording control input, perform video recording and output the first video; wherein the first field of view of the first video is greater than the second field of view supported by the camera recording the first video.

[0277] An embodiment of the present application provides an electronic device. After the field of view angle expansion recording function is activated, the field of view angle of the first video obtained by the electronic device during video recording can be greater than the field of view angle supported by the camera when recording the first video. That is, the first video can include more shooting scenes, thereby improving the reliability of the electronic device in shooting videos.

[0278] In some embodiments of the present application, the processor 110 is used to start the field of view expansion recording function in response to the control input of the field of view expansion control on the shooting preview interface; and to record the video and output the first video in response to the recording control input; wherein the first field of view of the video is greater than the second field of view supported by the camera recording the video.

[0279] In some embodiments of the present application, the processor 110 is specifically used to collect N frames of images and IMU data corresponding to each frame of the N frames during video recording; and based on each frame of the image and the IMU data corresponding to each frame of the image, generate N three-dimensional spherical texture images of the N frames of image, wherein one frame of image corresponds to one three-dimensional spherical texture image; and based on the first field of view, select a three-dimensional pixel point set corresponding to each frame of image from the N three-dimensional spherical texture images; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image, and the field of view of the two-dimensional video image corresponding to each frame of image is the first field of view angle; and perform image synthesis processing on the two-dimensional video images corresponding to the N frames of image according to the image acquisition order of each frame of image to obtain a first video.

[0280] In some embodiments of the present application, the processor 110 is specifically used to calculate the rotation offset and translation offset of each frame image based on the IMU data corresponding to each frame image; and map each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system to obtain an initial three-dimensional spherical texture image corresponding to each frame image; and determine the rotation and translation error parameter value based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame image; and map each two-dimensional pixel point in each frame image to three-dimensional space based on the rotation and translation error parameter value, the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system to obtain a three-dimensional spherical texture image corresponding to each frame image.

[0281] In some embodiments of the present application, the processor 110 is specifically used to predict the image color value of each frame of image based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image; and determine the rotation and translation error parameter value based on the predicted image color value of each frame of image and the actual image color value of each frame of image.

[0282] In some embodiments of the present application, the processor 110 is specifically used to determine the field of view angle projection range corresponding to each frame image from N three-dimensional spherical texture images based on the first field of view angle and the focal length used by the camera recording the video; and determine the pixel point coordinates of all three-dimensional pixel points within the field of view angle projection range corresponding to each frame image as the three-dimensional pixel point set corresponding to each frame image.

[0283] In some embodiments of the present application, the processor 110 is further configured to, before selecting a set of three-dimensional pixel points corresponding to each frame of the N three-dimensional spherical texture images based on the first field of view, perform synthesis processing on the N frames of images according to an image acquisition order of the N frames of images upon completion of video recording to obtain a second video, where the field of view of the second video is the second field of view angle;

[0284] The user input unit 107 is configured to receive a first input from a user on an expansion ratio control in the video preview interface when the video preview interface of the second video is displayed;

[0285] The processor 110 is further configured to determine a field of view angle expansion coefficient in response to a first input received by the user input unit 107 ; and multiply the field of view angle expansion coefficient by the second field of view angle to obtain a first field of view angle.

[0286] In some embodiments of the present application, the processor 110 is further used to select a three-dimensional pixel point set corresponding to any frame image from N three-dimensional spherical texture images based on the first field of view when the field of view angle expansion coefficient is greater than or equal to the coefficient threshold; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to any frame image to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any frame image.

[0287] The electronic device provided in the embodiment of the present application can implement each process implemented in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0288] The beneficial effects of various implementations in this embodiment can be specifically referred to the beneficial effects of the corresponding implementations in the above method embodiment. To avoid repetition, they will not be described here.

[0289] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0290] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0291] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.

[0292] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0293] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0294] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0295] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0296] An embodiment of the present application provides a computer program / program product, which is stored in a storage medium. The program / program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0297] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0298] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0299] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A video shooting method, characterized in that: include: In response to a control input of a field of view expansion control on a shooting preview interface, starting a field of view expansion recording function; In response to a video recording control input, video recording is performed and a first video is output; wherein a first field of view angle of the first video is greater than a second field of view angle supported by a camera that records the first video.

2. The method according to claim 1, characterized in that The step of recording a video and outputting a first video includes: During the video recording process, collecting N frames of images and inertial measurement unit (IMU) data corresponding to each frame of the N frames of images; Based on each frame of image and the IMU data corresponding to each frame of image, generating N three-dimensional spherical texture images of the N frames of image, wherein one frame of image corresponds to one three-dimensional spherical texture image; Based on the first field of view angle, selecting a set of three-dimensional pixel points corresponding to each frame of the N three-dimensional spherical texture images; Mapping three-dimensional pixel points in a three-dimensional pixel point set corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image, wherein the field of view of the two-dimensional video image corresponding to each frame of image is the first field of view angle; According to the image acquisition order of each frame of image, the two-dimensional video images corresponding to the N frames of image are subjected to image synthesis processing to obtain the first video.

3. The method according to claim 2, characterized in that The method of generating N three-dimensional spherical texture images of the N frames of images based on each frame of image and the IMU data corresponding to each frame of image includes: According to the IMU data corresponding to each frame of image, the rotation offset and translation offset of each frame of image are calculated; Mapping each two-dimensional pixel point in each frame of image to a three-dimensional space based on a rotation offset of each frame of image, a translation offset of each frame of image, a focal length of a camera used to record the video, and a coordinate value of each two-dimensional pixel point in each frame of image on an optical axis of a camera coordinate system to obtain an initial three-dimensional spherical texture image corresponding to each frame of image; Determining a rotation and translation error parameter value based on an image color value of an initial three-dimensional spherical texture image corresponding to each frame of image; According to the rotation and translation error parameter value, the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system, each two-dimensional pixel point in each frame image is mapped to three-dimensional space to obtain a three-dimensional spherical texture image corresponding to each frame image.

4. The method according to claim 3, characterized in that The step of determining the rotation and translation error parameter value based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image includes: Predicting the image color value of each frame of image based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image; The rotation and translation error parameter value is determined based on the predicted image color value of each frame image and the actual image color value of each frame image.

5. The method according to claim 2, characterized in that The selecting, based on the first field of view angle, a set of three-dimensional pixel points corresponding to each frame of the N three-dimensional spherical texture images includes: Determining a field of view projection range corresponding to each frame of the N three-dimensional spherical texture images according to the first field of view angle and the focal length of the camera used to record the video; The pixel coordinates of all three-dimensional pixels within the projection range of the field of view corresponding to each frame of image are determined as a three-dimensional pixel set corresponding to each frame of image.

6. The method according to claim 2 or 5, characterized in that Before selecting a set of three-dimensional pixel points corresponding to each frame of image from the N three-dimensional spherical texture images based on the first field of view, the method further includes: When video recording is finished, performing synthesis processing on the N frames of images according to the image acquisition order of the N frames of images to obtain a second video, where the field of view of the second video is the second field of view angle; In a case where a video preview interface of the second video is displayed, receiving a first input from a user on an expansion ratio control in the video preview interface; In response to the first input, determining a field of view angle expansion coefficient; The first field of view angle is obtained by multiplying the field of view angle expansion coefficient by the second field of view angle.

7. The method according to claim 6, characterized in that The method further comprises: When the field of view angle expansion coefficient is greater than or equal to a coefficient threshold, selecting a set of three-dimensional pixel points corresponding to any frame image from the N three-dimensional spherical texture images based on the first field of view angle; The three-dimensional pixel points in the three-dimensional pixel point set corresponding to any one frame image are mapped to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any one frame image.

8. A video shooting device, characterized in that: include: a processing module, responsive to a control input to a field-of-view expansion control on a shooting preview interface, activating a field-of-view expansion recording function; And in response to the video recording control input, video recording is performed to output a first video; wherein, the first field of view angle of the first video is greater than the second field of view angle supported by the camera recording the first video.

9. The device according to claim 8, characterized in that The processing module is specifically used to collect N frames of images and IMU data corresponding to each frame of the N frames during video recording; and based on each frame of image and the IMU data corresponding to each frame of image, generate N three-dimensional spherical texture images of the N frames of image, wherein one frame of image corresponds to one three-dimensional spherical texture image; and based on the first field of view, select a three-dimensional pixel point set corresponding to each frame of image from the N three-dimensional spherical texture images; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to each frame of image to a two-dimensional space to obtain a two-dimensional video image corresponding to each frame of image, and the field of view of the two-dimensional video image corresponding to each frame of image is the first field of view angle; and perform image synthesis processing on the two-dimensional video images corresponding to the N frames of image according to the image acquisition order of each frame of image to obtain the first video.

10. The device according to claim 9, characterized in that The processing module is specifically used to calculate the rotation offset and translation offset of each frame of image based on the IMU data corresponding to each frame of image; and mapping each two-dimensional pixel point in each frame of image to a three-dimensional space based on a rotation offset of each frame of image, a translation offset of each frame of image, a focal length used by a camera recording the video, and a coordinate value of each two-dimensional pixel point in each frame of image on an optical axis of a camera coordinate system, to obtain an initial three-dimensional spherical texture image corresponding to each frame of image; and determining a rotation and translation error parameter value based on an image color value of an initial three-dimensional spherical texture image corresponding to each frame of image; And according to the rotation and translation error parameter value, the rotation offset of each frame image, the translation offset of each frame image, the focal length used by the camera recording the video, and the coordinate value of each two-dimensional pixel point in each frame image on the optical axis of the camera coordinate system, each two-dimensional pixel point in each frame image is mapped to a three-dimensional space to obtain a three-dimensional spherical texture image corresponding to each frame image.

11. The device according to claim 10, characterized in that The processing module is specifically used to predict the image color value of each frame of image based on the image color value of the initial three-dimensional spherical texture image corresponding to each frame of image; and determine the rotation and translation error parameter value based on the predicted image color value of each frame of image and the actual image color value of each frame of image.

12. The device according to claim 9, characterized in that The processing module is specifically configured to determine, from the N three-dimensional spherical texture images, a field of view angle projection range corresponding to each frame of image based on the first field of view angle and the focal length of the camera used to record the video; and to determine the pixel coordinates of all three-dimensional pixels within the field of view angle projection range corresponding to each frame of image as a three-dimensional pixel point set corresponding to each frame of image.

13. The device according to claim 9 or 12, characterized in that The device further includes: a receiving module; The processing module is further configured to, before selecting a set of three-dimensional pixel points corresponding to each frame of the N three-dimensional spherical texture images based on the first field of view, perform synthesis processing on the N frames of images according to an image acquisition order of the N frames of images upon completion of video recording to obtain a second video, where the field of view of the second video is the second field of view angle; The receiving module is configured to receive a first input from a user on an expansion ratio control in the video preview interface when the video preview interface of the second video is displayed; The processing module is further configured to determine a field of view angle expansion coefficient in response to the first input received by the receiving module; and multiply the field of view angle expansion coefficient by the second field of view angle to obtain the first field of view angle.

14. The device according to claim 13, characterized in that The processing module is also used to select a three-dimensional pixel point set corresponding to any frame image from the N three-dimensional spherical texture images based on the first field of view when the field of view expansion coefficient is greater than or equal to the coefficient threshold; and map the three-dimensional pixel points in the three-dimensional pixel point set corresponding to any frame image to a two-dimensional space to obtain a two-dimensional panoramic image corresponding to any frame image.

15. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the video shooting method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Roller shutter image motion compensation method and system based on XR positioning

    CN121837384A