Video generation method, device and panoramic camera
By acquiring and processing images by a panoramic camera, and using the attitude data and position information of the inertial measurement unit to generate flat videos, solving the problem of inefficiency in the prior art and achieving efficient flat video generation.
Patent Information
- Application Number
- CN202411679120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The existing flat video generation technology is inefficient.
The k-th plane image is acquired through the panoramic camera and sent to the designated terminal, the posture data is obtained using the inertia measurement unit, and the plane video is generated based on the position information of the image frame.
Improves the efficiency of flat video generation and can efficiently control viewing angles and video content.
Smart Images

Figure CN119183022B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video processing, and particularly relates to a video generation method, an apparatus, and a panoramic camera. Background Art
[0002] In recent years, with the continuous development of science and technology, various technologies have emerged in people's lives like bamboo shoots after a spring rain. For example, video generation technology, specifically, the technology of generating a planar video based on a panoramic video. However, there are still some drawbacks in the current planar video generation technology that need to be solved. For example, the current efficiency of planar video generation is relatively low. Summary of the Invention
[0003] Embodiments of this application provide a video generation method, an apparatus, a panoramic camera, and a computer-readable storage medium, which can solve the problem of relatively low efficiency in current planar video generation.
[0004] In a first aspect, embodiments of this application provide a video generation method. The video generation method is applied to a panoramic camera and includes: obtaining a k-th planar image, where the k-th planar image is a planar image corresponding to a target area in the k-th frame image of a first panoramic video, and k is a positive integer greater than or equal to 1; sending the k-th planar image to a specified terminal, where an inertial measurement unit is provided on the specified terminal, and the specified terminal can display the k-th planar image and obtain k-th pose data, where the k-th pose data is used to represent the pose data corresponding to the inertial measurement unit when the specified terminal displays the k-th planar image; obtaining the position information of the target area in the (k + 1)-th frame image according to the k-th pose data and the position information of the target area in the k-th frame image, and incorporating the position information of the target area in the (k + 1)-th frame image into a target data set, where the target data set further includes the position information of the target area in the k-th frame image; and generating a planar video according to the target data set.
[0005] In a possible implementation manner of the first aspect, the obtaining of the k-th planar image includes: if the panoramic camera is triggered to enter the instant capture and instant editing mode, then when the panoramic camera records the N-th frame image of the first panoramic video, the k-th planar image is obtained, where N is a positive integer greater than k.
[0006] In a possible implementation of the first aspect, at least two cameras are disposed on the panoramic camera. The two cameras include a first camera and a second camera. Before obtaining the k-th plane image, it includes: splicing a first image and a second image according to the calibration parameters of the two cameras to obtain a first target image, where the first target image is an image frame for constructing a first panoramic video. The first image is an image obtained based on the shooting of the first camera, and the second image is an image obtained based on the shooting of the second camera. The first target image includes the k-th frame image frame of the first panoramic video.
[0007] In a possible implementation of the first aspect, before generating a plane video according to the target data set, it includes: determining a target video segment in the first panoramic video; determining a first image to be re-spliced and a second image to be re-spliced based on the target video segment, where the first image to be re-spliced is the first image corresponding to the image frame in the target video segment, and the second image to be re-spliced is the second image corresponding to the image frame in the target video segment; determining the relative displacement between the same feature points between a first region and a second region, where the first region is a region in the first image to be re-spliced that has multiple same feature points with the second image to be re-spliced, and the second region is a region in the second image to be re-spliced that has multiple same feature points with the first image to be re-spliced; splicing the first image to be re-spliced and the second image to be re-spliced according to the relative displacement to obtain a second target image, where the second target image is an image frame for constructing a second panoramic video; correspondingly, generating a plane video according to the target data set includes: generating a plane video according to the target data set and the second panoramic video.
[0008] In a possible implementation of the first aspect, generating a plane video according to the target data set and the second panoramic video includes: unfolding each image frame in the second panoramic video into a target plane image based on the target data set, and generating a plane video based on all the target plane images.
[0009] In a possible implementation of the first aspect, the initial value of k is equal to a first specified positive integer, where the first specified positive integer is a positive integer greater than or equal to one. Obtaining the k-th plane image, sending the k-th plane image to a specified terminal, obtaining the position information of the target area in the (k + 1)-th frame image frame according to the k-th pose data and the position information of the target area in the k-th frame image frame, and incorporating the position information of the target area in the (k + 1)-th frame image frame into the target data set. The target data set also includes the position information of the target area in the k-th frame image frame, and it includes:
[0010] Step A1: Obtain the k-th planar image;
[0011] Step A2: Send the k-th planar image to a specified terminal;
[0012] Step A3: Obtain the position information of the target area in the (k + 1)-th image frame based on the k-th pose data and the position information of the target area in the k-th image frame, and incorporate the position information of the target area in the (k + 1)-th image frame into the target data set, where the target data set also includes the position information of the target area in the k-th image frame;
[0013] Step A4: Let k = k + 1;
[0014] Step A5: Determine whether k is less than a second specified positive integer. If k is less than the second specified positive integer, then return to execute Step A1, where the first specified positive integer is less than the second specified positive integer, and the difference between the second specified positive integer and the first specified positive integer is greater than or equal to 2.
[0015] In a possible implementation manner of the first aspect, the following relationship holds among the k-th pose data, the position information of the target area in the k-th image frame, and the position information of the target area in the (k + 1)-th image frame: , where the is used to represent the position information of the target area in the (k + 1)-th image frame, the is used to represent the k-th pose data, and the is used to represent the position information of the target area in the k-th image frame.
[0016] In a second aspect, an embodiment of the present application provides a video generation device, which is applied to a panoramic camera. The video generation device includes: an image acquisition unit for acquiring the k-th planar image, where the k-th planar image is a planar image corresponding to the target area in the k-th image frame of the first panoramic video, and k is a positive integer greater than or equal to 1; an image sending unit for sending the k-th planar image to a specified terminal, where an inertial measurement unit is provided on the specified terminal, and the specified terminal can display the k-th planar image and acquire the k-th pose data, and the k-th pose data is used to represent the pose data corresponding to the inertial measurement unit when the specified terminal displays the k-th planar image; an information acquisition unit for acquiring the position information of the target area in the (k + 1)-th image frame based on the k-th pose data and the position information of the target area in the k-th image frame, and incorporating the position information of the target area in the (k + 1)-th image frame into the target data set, where the target data set also includes the position information of the target area in the k-th image frame; and a generation unit for generating a planar video according to the target data set.
[0017] In a third aspect, an embodiment of the present application provides a panoramic camera, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above is implemented.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0019] It can be understood that the beneficial effects of the second to fourth aspects above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0020] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: In the present application, the panoramic camera can obtain the k-th plane image, where the k-th plane image is: the plane image corresponding to the target area in the k-th frame image of the first panoramic video, and k is a positive integer greater than or equal to 1; send the k-th plane image to a specified terminal, where an inertial measurement unit is set on the specified terminal, and the specified terminal can display the k-th plane image and obtain the k-th attitude data, and the k-th attitude data is used to represent: the attitude data corresponding to the inertial measurement unit when the specified terminal displays the k-th plane image; obtain the position information of the target area in the (k + 1)-th frame image according to the k-th attitude data and the position information of the target area in the k-th frame image, and classify the position information of the target area in the (k + 1)-th frame image into the target data set, and the target data set also includes the position information of the target area in the k-th frame image; generate a plane video according to the target data set. Since different attitude data correspond to different position information, different position information corresponds to different perspectives, and different perspectives correspond to different image contents, therefore, the present application can efficiently control the perspective by obtaining the attitude data corresponding to the inertial measurement unit, and then control the content of the generated plane video, that is, the present application can efficiently generate a plane video. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0022] Figure 1 It is a schematic flowchart of a video generation method provided by an embodiment of the present application;
[0023] Figure 2It is a schematic diagram of a video generation device provided by an embodiment of the present application. Detailed implementation manners
[0024] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0025] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.
[0028] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0029] The reference to "an embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0030] Embodiment 1.
[0031] Figure 1 The flowchart shows a video generation method provided by an embodiment of the present application, which is applied to a panoramic camera. The video generation method includes: step S101, step S102, step S103, and step S104. Details are as follows:
[0032] Step S101: Obtain the k-th plane image, where the k-th plane image is the plane image corresponding to the target area in the k-th frame image of the first panoramic video, and k is a positive integer greater than or equal to 1.
[0033] Among them, the image frames in the first panoramic video are (panoramic) spherical images, and the k-th plane image is the plane image unfolded from the target area in the k-th frame image of the first panoramic video.
[0034] As an example rather than a limitation, a virtual camera is set at the center of the sphere corresponding to the panoramic spherical image, and the virtual camera is used to determine the target area. For example, the field of view angle of the virtual camera forms a viewing range, and the area where the viewing range intersects the panoramic spherical image is the target area.
[0035] In some embodiments, step S101 includes: the panoramic camera unfolds the target area in the k-th frame image of the first panoramic video by using the equirectangular projection (ERP) method to obtain the k-th plane image, and the k-th plane image is the unfolding result corresponding to the target area.
[0036] Optionally, step S101 includes: if the panoramic camera is triggered to enter the instant capture and cut mode, then when the panoramic camera records the N-th frame image of the first panoramic video, it obtains the k-th plane image, where N is a positive integer greater than k.
[0037] In this embodiment, the instant capture and cut mode is used to represent the following mode: when the panoramic camera records the N-th frame image of the first panoramic video, it obtains the k-th plane image.
[0038] Specifically, when the panoramic camera records the N-th frame image of the first panoramic video, it at least executes obtaining the k-th plane image and sending the k-th plane image to a specified terminal, where N is a positive integer greater than k.
[0039] As an example rather than a limitation, N is equal to k + 1. Assuming k is equal to 2, then N is equal to 3, that is, when the panoramic camera records the third frame image of the first panoramic video, it obtains the second plane image.
[0040] That is, the panoramic camera in this embodiment can record the first panoramic video while performing relevant steps for video editing (editing the panoramic video to generate a planar video). Thus, the efficiency of data processing can be greatly improved.
[0041] Optionally, at least two cameras are provided on the panoramic camera. The two cameras include a first camera and a second camera. Before the step S101, it includes: the panoramic camera stitches a first image and a second image according to the calibration parameters of the two cameras to obtain a first target image. The first target image is an image frame for constituting the first panoramic video. The first image is an image obtained based on the shooting of the first camera, and the second image is an image obtained based on the shooting of the second camera. The first target image includes the k-th frame image of the first panoramic video.
[0042] Wherein, the calibration parameters are external parameter values obtained during the calibration of the panoramic camera. Among them, the external parameter values include: a rotation matrix and a translation matrix from a preset world coordinate system to a preset camera coordinate system.
[0043] As can be seen from the above, in this embodiment, the image frames of the first panoramic video are all stitched based on the calibration parameters of the two cameras. That is, the panoramic camera can quickly / efficiently implement image stitching based on the calibration parameters of the two cameras.
[0044] Step S102: Send the k-th planar image to a specified terminal. An inertial measurement unit is provided on the specified terminal. The specified terminal can display the k-th planar image and obtain k-th attitude data. The k-th attitude data is used to represent: the attitude data corresponding to the inertial measurement unit when the specified terminal displays the k-th planar image.
[0045] Specifically, the panoramic camera sends the k-th planar image to the specified terminal. The specified terminal can obtain the k-th planar image sent by the panoramic camera, then display the k-th planar image. The user using the specified terminal can view the k-th planar image through the specified terminal. During the period when the specified terminal displays the image, the user can change the attitude of the specified terminal. The specified terminal can obtain the attitude data corresponding to the inertial measurement unit when displaying the image, that is, can obtain the k-th attitude data. Then the specified terminal sends the k-th attitude data to the panoramic camera, so that the panoramic camera can generate a planar video based on the k-th attitude data. As can be seen from the above, the user can control the generation of the planar video through the specified terminal.
[0046] It should be noted that the specified terminal in this article does not include the panoramic camera.
[0047] As an example rather than a limitation, the specified terminal may be a virtual reality (VR) device, such as a head-mounted VR device. The VR device is capable of displaying the k-th plane image, and an inertial measurement unit is provided on the VR device. The inertial measurement unit may include a gyroscope, and the user can change the attitude of the VR device (inertial measurement unit) by turning their head.
[0048] Step S103: Obtain the position information of the target area in the (k + 1)-th image frame according to the k-th attitude data and the position information of the target area in the k-th image frame, and classify the position information of the target area in the (k + 1)-th image frame into the target data set, where the target data set further includes the position information of the target area in the k-th image frame.
[0049] Specifically, the panoramic camera obtains the k-th attitude data, and then obtains the position information of the target area in the (k + 1)-th image frame according to the k-th attitude data and the position information of the target area in the k-th image frame. Among them, the position information of the target area is information that can reflect the location where the target area is located.
[0050] As an example rather than a limitation, the position information of the target area may be at least one of the following: information representing the position of the midpoint of the target area, information representing the position of the boundary point of the target area, where the boundary point of the target area is a point on the boundary of the target area. For example, the boundary point in the upper left corner of the target area. In addition, the position information may be longitude and latitude. For example, the information representing the position of the midpoint of the target area includes: the longitude and latitude where the position of the midpoint of the target area is located. The position information can be represented in the form of a vector.
[0051] Optionally, the following relationship is satisfied among the k-th attitude data, the position information of the target area in the k-th image frame, and the position information of the target area in the (k + 1)-th image frame: where, the is used to represent the position information of the target area in the (k + 1)-th image frame, the is used to represent the k-th attitude data, and the is used to represent the position information of the target area in the k-th image frame.
[0052] Among them, is used to represent the conjugate quaternion of. The k-th attitude data can be represented by an attitude quaternion.
[0053] In this way, the position information of the target area in the (k + 1)-th image frame can be accurately and efficiently determined.
[0054] Optionally, the initial value of k is equal to a first specified positive integer, which is a positive integer greater than or equal to one. The three steps of step S101, step S102, and step S103 include: step A1, step A2, step A3, step A4, and step A5.
[0055] Step A1: Obtain the k-th plane image;
[0056] Step A2: Send the k-th plane image to a specified terminal;
[0057] Step A3: Obtain the position information of the target area in the (k + 1)-th image frame according to the k-th pose data and the position information of the target area in the k-th image frame, and classify the position information of the target area in the (k + 1)-th image frame into the target data set, which also includes the position information of the target area in the k-th image frame;
[0058] Step A4: Let k = k + 1;
[0059] Step A5: Determine whether k is less than a second specified positive integer. If k is less than the second specified positive integer, then return to execute step A1, where the first specified positive integer is less than the second specified positive integer, and the difference between the second specified positive integer and the first specified positive integer is greater than or equal to 2.
[0060] By way of example and not limitation, assume that the first specified positive integer is equal to 1, that is, the initial value of k is equal to 1, and the second specified positive integer is equal to 3. For the sake of description, it can be divided into multiple stages as follows:
[0061] The first stage: The panoramic camera obtains the first plane image, sends the first plane image to a specified terminal, and then the panoramic camera obtains the position information of the target area in the second image frame according to the first pose data and the position information of the target area in the first image frame, and classifies the position information of the target area in the second image frame into the target data set. At this time, the target data set includes: the position information of the target area in the first image frame and the position information of the target area in the second image frame; then let k = k + 1. Correspondingly, the current k is equal to 2. It is determined that the current k is less than the second specified positive integer, and then enter the second stage.
[0062] Second stage: The panoramic camera acquires a second planar image, sends the second planar image to a specified terminal, and then the panoramic camera acquires the position information of the target area in the third image frame according to the first attitude data and the position information of the target area in the second image frame, and incorporates the position information of the target area in the third image frame into the target data set. At this time, the target data set includes: the position information of the target area in the first image frame, the position information of the target area in the second image frame, and the position information of the target area in the third image frame; then let k = k + 1. Correspondingly, the current k is equal to 3. It is determined that the current k is not less than the second specified positive integer, and then step A1 is not returned for execution.
[0063] As can be seen from the above, in this embodiment, the target data set is updated in a cyclic manner, so that the updated target data set includes the position information of the target areas of multiple image frames in the first panoramic video, thereby providing sufficient data support for generating a planar video subsequently.
[0064] Step S104: Generate a planar video according to the target data set.
[0065] Specifically, the target data set at least includes: the position information of the target area in the k-th image frame and the position information of the target area in the (k + 1)-th image frame, and a planar video is generated based on the position information in the target data set.
[0066] Optionally, generating a planar video according to the target data set includes: generating a planar video according to the target data set and the first panoramic video.
[0067] Optionally, the image frames in the first panoramic video are stitched together from the first image and the second image according to the calibration parameters of the two cameras; before the step S104, the panoramic camera includes: determining a target video segment in the first panoramic video, and determining a first image to be re-stitched and a second image to be re-stitched based on the target video segment, where the first image to be re-stitched is the first image corresponding to the image frame in the target video segment, and the second image to be re-stitched is the second image corresponding to the image frame in the target video segment; determining the relative displacement between the same feature points between the first area and the second area, where the first area is: the area in the first image to be re-stitched that has multiple same feature points with the second image to be re-stitched, and the second area is: the area in the second image to be re-stitched that has multiple same feature points with the first image to be re-stitched; stitching the first image to be re-stitched and the second image to be re-stitched according to the relative displacement to obtain a second target image, where the second target image is the image frame used to form the second panoramic video; correspondingly, the step S104 includes: generating a planar video according to the target data set and the second panoramic video.
[0068] Among them, the target video segment in the first panoramic video can be understood as the video segment in the first panoramic video that needs to be used for editing to generate a planar video.
[0069] For ease of description, the first image corresponding to the m-th frame image in the target video segment is denoted as C m1 , and the second image corresponding to the m-th frame image in the target video segment is denoted as C m2 , where m is a positive integer greater than or equal to one, that is, the m-th frame image in the target video segment is based on C m1 and C m2 stitched together. For example, the first frame image in the target video segment is based on C 11 and C 12 stitched together, the second frame image in the target video segment is based on C 21 and C 22 stitched together, and so on. Correspondingly, the first image to be re-stitched includes: C 11 and C 21 , the second image to be re-stitched includes: C 12 and C 22 , the first region corresponding to C 11 is: the region in C 11 that has multiple identical feature points with C 12 , the second region corresponding to C 12 is: the region in C 12 that has multiple identical feature points with C 11 . Determine the first relative displacement, where the first relative displacement is the relative displacement between the identical feature points between the first region corresponding to C 11 and the second region corresponding to C 12 . Stitch C 11 and C 12 according to the first relative displacement to obtain D1, and D1 is used to represent: the stitching result obtained by stitching C 11 and C 12 according to the first relative displacement, and D1 belongs to the second target image; the first region corresponding to C 21 is: the region in C 21 that has multiple identical feature points with C 22 , the second region corresponding to C 22 is: the region in C 22 that has multiple identical feature points with C 21 . Determine the second relative displacement, where the second relative displacement is the relative displacement between the identical feature points between the first region corresponding to C 21 and the second region corresponding to C 22 . Stitch C 21 and C22 They are spliced to obtain D2, and D2 is used to represent: C according to the second relative displacement 21 and C 22 The splicing result obtained by splicing, D2 belongs to the second target image, that is, the image frames used to form the second panoramic video include D1 and D2. Correspondingly, generating a planar video according to the target data set and the second panoramic video includes: generating a planar video according to the target data set, D1, and D2.
[0070] In addition, the following example is given to illustrate the determination of the first relative displacement. For example, C 11 There is a feature point E1 in the corresponding first region, and C 12 There is a feature point E2 in the corresponding second region. The feature point E1 and the feature point E2 have the same feature, that is, the feature point E1 and the feature point E2 are the same feature point. Determine the first relative displacement, and the first relative displacement includes the relative displacement between the feature point E1 and the feature point E2.
[0071] In this embodiment, the panoramic camera adopts two different splicing methods. The first splicing method is to splice based on the calibration parameters of two cameras, and the second splicing method is to splice according to the relative displacement. The first splicing method can quickly realize image splicing, enabling users to view the first panoramic video in a timely manner. The second splicing method has higher image splicing accuracy, better image splicing effect, and only needs to operate on the target video segment, without the need to splice all the first images and all the second images corresponding to the first panoramic video, which can effectively improve the efficiency of the second splicing method.
[0072] In some embodiments, determining the target video segment in the first panoramic video includes: the panoramic camera obtains segment marking information, and determines the target video segment in the first panoramic video according to the segment marking information, where the segment marking information is used to represent: the starting image frame and the ending image frame of the target video segment.
[0073] In some embodiments, generating a planar video according to the target data set and the second panoramic video includes: expanding each image frame in the second panoramic video into a target planar image based on the target data set, and generating a planar video based on all the target planar images.
[0074] For example, in the target data set, there is position information corresponding to each frame image in the second panoramic video. For example, the target data set includes: the position information L1 of the target area in the first frame image of the first panoramic video and the position information L2 of the target area in the second frame image of the first panoramic video. The first frame image in the second panoramic video and the first frame image of the first panoramic video are stitched based on the same first image and second image. Therefore, L1 is the position information corresponding to the first frame image in the second panoramic video. Based on L1, the first frame image in the second panoramic video is unfolded into a first target plane image; the second frame image in the second panoramic video and the second frame image of the first panoramic video are stitched based on the same first image and second image. Therefore, L2 is the position information corresponding to the second frame image in the second panoramic video. Based on L2, the second frame image in the second panoramic video is unfolded into a second target plane image. A plane video is generated based on the first target plane image and the second target plane image.
[0075] In this application, the panoramic camera can obtain the k-th plane image. The k-th plane image is: the plane image corresponding to the target area in the k-th frame image of the first panoramic video, where k is a positive integer greater than or equal to 1; the k-th plane image is sent to a specified terminal. An inertial measurement unit is set on the specified terminal. The specified terminal can display the k-th plane image and obtain the k-th attitude data. The k-th attitude data is used to represent: the attitude data corresponding to the inertial measurement unit when the specified terminal displays the k-th plane image; the position information of the target area in the (k + 1)-th frame image is obtained according to the k-th attitude data and the position information of the target area in the k-th frame image, and the position information of the target area in the (k + 1)-th frame image is classified into the target data set. The target data set also includes the position information of the target area in the k-th frame image; a plane video is generated according to the target data set. Since different attitude data corresponds to different position information, different position information corresponds to different perspectives, and different perspectives correspond to different image contents, therefore, this application can efficiently control the perspective by obtaining the attitude data corresponding to the inertial measurement unit, and then control the content of the generated plane video, that is, this application can efficiently generate a plane video.
[0076] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0077] Embodiment 2.
[0078] Corresponding to the video generation method described in the above embodiment Figure 2The figure shows a schematic diagram of a video generation device provided by an embodiment of the present application. For ease of description, only parts related to the embodiment of the present application are shown.
[0079] The video generation device is applied to a panoramic camera. The video generation device includes: an image acquisition unit 201, an image transmission unit 202, an information acquisition unit 203, and a generation unit 204. Among them:
[0080] The image acquisition unit 201 is configured to acquire a k-th plane image, where the k-th plane image is a plane image corresponding to a target area in the k-th frame image of a first panoramic video, and k is a positive integer greater than or equal to 1.
[0081] Optionally, at least two cameras are provided on the panoramic camera. The two cameras include a first camera and a second camera. The video generation device further includes a first stitching unit, configured to: before the image acquisition unit 201 executes the acquisition of the k-th plane image, stitch a first image and a second image according to the calibration parameters of the two cameras to obtain a first target image, where the first target image is an image frame for constituting a first panoramic video, the first image is an image obtained based on the shooting of the first camera, the second image is an image obtained based on the shooting of the second camera, and the first target image includes the k-th frame image of the first panoramic video.
[0082] The image transmission unit 202 is configured to send the k-th plane image to a specified terminal. An inertial measurement unit is provided on the specified terminal. The specified terminal can display the k-th plane image and acquire k-th pose data, where the k-th pose data is used to represent the pose data corresponding to the inertial measurement unit when the specified terminal displays the k-th plane image.
[0083] The information acquisition unit 203 is configured to acquire the position information of the target area in the (k + 1)-th frame image according to the k-th pose data and the position information of the target area in the k-th frame image, and classify the position information of the target area in the (k + 1)-th frame image into a target data set, where the target data set further includes the position information of the target area in the k-th frame image.
[0084] The generation unit 204 is configured to generate a plane video according to the target data set.
[0085] Optionally, the video generating device further includes a second splicing unit, configured to: before the generating unit 204 generates a planar video according to the target data set, determine a target video segment in the first panoramic video; determine a first image to be re-spliced and a second image to be re-spliced based on the target video segment, where the first image to be re-spliced is a first image corresponding to an image frame in the target video segment, and the second image to be re-spliced is a second image corresponding to an image frame in the target video segment; determine the relative displacement between the same feature points between a first region and a second region, where the first region is a region in the first image to be re-spliced that has a plurality of the same feature points as the second image to be re-spliced, and the second region is a region in the second image to be re-spliced that has a plurality of the same feature points as the first image to be re-spliced; splice the first image to be re-spliced and the second image to be re-spliced according to the relative displacement to obtain a second target image, where the second target image is an image frame for constructing a second panoramic video; correspondingly, when the generating unit 204 generates a planar video according to the target data set, it is specifically configured to: generate a planar video according to the target data set and the second panoramic video.
[0086] It should be noted that for the technical details not described in detail in this embodiment, reference may be made to the video generating methods provided in the respective embodiments in the above-mentioned Embodiment 1.
[0087] Embodiment 3.
[0088] The panoramic camera in this embodiment includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned video generating method embodiments.
[0089] When the processor executes the computer program, it implements the steps in the above-mentioned video generating method embodiments, for example Figure 1 the steps S101 to S104 shown. Alternatively, when the processor executes the computer program, it implements the functions of the respective units in the above-mentioned device embodiments. For example, Figure 2 the functions of the units 201 to 204 shown.
[0090] Those skilled in the art can understand that this embodiment is only an example of a panoramic camera and does not constitute a limitation on the panoramic camera. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0091] The so-called processor may be a Central Processing Unit (CPU), and the processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0092] In some embodiments, the memory may be an internal storage unit of the panoramic camera, such as the hard disk or memory of the panoramic camera. In other embodiments, the memory may also be an external storage device of the panoramic camera, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the panoramic camera. Further, the memory may also include both the internal storage unit of the panoramic camera and the external storage device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory may also be used to temporarily store data that has been output or will be output.
[0093] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.
[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.
[0095] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0096] The embodiment of the present application provides a computer program product. When the computer program product runs on a panoramic camera, the panoramic camera can execute the steps in the above-mentioned method embodiments.
[0097] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the panoramic camera, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0098] In the above embodiments, the descriptions of the respective embodiments each have their own emphasis. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0100] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.
[0101] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0102] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A video generation method, characterized in that, The video generation method is applied to a panoramic camera, and the video generation method includes: Obtain a k-th plane image, where the k-th plane image is a plane image corresponding to a target area in the k-th frame image of a first panoramic video, and k is a positive integer greater than or equal to 1; Send the k-th plane image to a specified terminal, where an inertial measurement unit is set on the specified terminal, and the specified terminal can display the k-th plane image and obtain k-th pose data, and the k-th pose data is used to represent: the pose data corresponding to the inertial measurement unit when the specified terminal displays the k-th plane image; Obtain the position information of the target region in the (k + 1)-th image frame according to the k-th pose data and the position information of the target region in the k-th image frame, and incorporate the position information of the target region in the (k + 1)-th image frame into the target data set, where the target data set further includes the position information of the target region in the k-th image frame. The following relationship holds among the k-th pose data, the position information of the target region in the k-th image frame, and the position information of the target region in the (k + 1)-th image frame: wherein, the p k+1 is used to represent the position information of the target region in the (k + 1)-th image frame, the q k is used to represent the k-th pose data, the p k is used to represent the position information of the target region in the k-th image frame, and the is used to represent the conjugate quaternion of the q k ; Generate a plane video according to the target data set.
2. The video generation method according to claim 1, wherein The obtaining of the k-th plane image includes: If the panoramic camera is triggered to enter the instant capture and editing mode, then when the panoramic camera records the N-th frame image of the first panoramic video, obtain the k-th plane image, where N is a positive integer greater than k.
3. The video generation method according to claim 1, wherein At least two cameras are set on the panoramic camera, and the two cameras include a first camera and a second camera. Before obtaining the k-th plane image, it includes: Stitch a first image and a second image according to the calibration parameters of the two cameras to obtain a first target image, where the first target image is an image frame for constructing the first panoramic video, the first image is an image obtained based on the shooting of the first camera, the second image is an image obtained based on the shooting of the second camera, and the first target image includes the k-th frame image of the first panoramic video.
4. The video generation method according to claim 3, wherein Before generating the plane video according to the target data set, it includes: Determine a target video segment in the first panoramic video; Determine a first image to be re-stitched and a second image to be re-stitched based on the target video segment, where the first image to be re-stitched is the first image corresponding to the image frame in the target video segment, and the second image to be re-stitched is the second image corresponding to the image frame in the target video segment; determine the relative displacement between the same feature points between a first area and a second area, where the first area is an area in the first image to be re-stitched that has multiple same feature points with the second image to be re-stitched, and the second area is an area in the second image to be re-stitched that has multiple same feature points with the first image to be re-stitched; Stitch the first image to be re-stitched and the second image to be re-stitched according to the relative displacement to obtain a second target image, where the second target image is an image frame for constructing a second panoramic video; Correspondingly, the generating of the plane video according to the target data set includes: Generate a plane video according to the target data set and the second panoramic video.
5. The video generation method according to claim 4, characterized in that The generating of the plane video according to the target data set and the second panoramic video includes: Unfold each image frame in the second panoramic video into a target plane image based on the target data set, and generate a plane video based on all the target plane images.
6. The video generation method according to claim 1, wherein The initial value of k is equal to the first specified positive integer, which is a positive integer greater than or equal to one. Obtain the k-th planar image, send the k-th planar image to a specified terminal, obtain the position information of the target area in the (k + 1)-th image frame according to the k-th attitude data and the position information of the target area in the k-th image frame, and classify the position information of the target area in the (k + 1)-th image frame into the target data set, where the target data set further includes the position information of the target area in the k-th image frame, including: Step A1: Obtain the k-th planar image; Step A2: Send the k-th planar image to a specified terminal; Step A3: Obtain the position information of the target area in the (k + 1)-th image frame according to the k-th attitude data and the position information of the target area in the k-th image frame, and classify the position information of the target area in the (k + 1)-th image frame into the target data set, where the target data set further includes the position information of the target area in the k-th image frame; Step A4: Let k = k + 1; Step A5: Determine whether k is less than the second specified positive integer. If k is less than the second specified positive integer, then return to execute Step A1, where the first specified positive integer is less than the second specified positive integer, and the difference between the second specified positive integer and the first specified positive integer is greater than or equal to 2.
7. A video generation device, characterized in that, The video generation device is applied to a panoramic camera, and the video generation device includes: An image acquisition unit for acquiring the k-th planar image, where the k-th planar image is a planar image corresponding to the target area in the k-th image frame of the first panoramic video, where k is a positive integer greater than or equal to 1; An image sending unit for sending the k-th planar image to a specified terminal, where an inertial measurement unit is provided on the specified terminal, and the specified terminal can display the k-th planar image and acquire the k-th attitude data, where the k-th attitude data is used to represent the attitude data corresponding to the inertial measurement unit when the specified terminal displays the k-th planar image; An information acquisition unit, configured to obtain the position information of the target area in the (k + 1)-th image frame according to the k-th attitude data and the position information of the target area in the k-th image frame, and classify the position information of the target area in the (k + 1)-th image frame into a target data set, where the target data set further includes the position information of the target area in the k-th image frame, and the following relationship is satisfied among the k-th attitude data, the position information of the target area in the k-th image frame, and the position information of the target area in the (k + 1)-th image frame: wherein, the p k+1 is used to represent the position information of the target area in the (k + 1)-th image frame, the q k is used to represent the k-th attitude data, the p k is used to represent the position information of the target area in the k-th image frame, and the is used to represent the conjugate quaternion of the q k ; A generation unit for generating a planar video according to the target data set.
8. A panoramic camera, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video editing method and device and virtual reality technology equipment
CN118055298A
Video processing method and device, computer equipment and storage medium
CN118509713A