Film and television shooting method and film and television shooting device based on fpv
By acquiring and merging videos from multiple locations during FPV film shooting, the problems of inconsistent shooting content and shaky footage are solved, resulting in high-quality film and television works and improving user experience and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2026-03-17
AI Technical Summary
During FPV film shooting, the movement of the subject causes the shooting content to fail to meet the requirements of film and television creation, and the shaking of the screen during terminal shooting affects the visual experience, reducing the creative effect of film and television works.
By acquiring multiple FPV videos from different positions on the subject and performing video fusion processing, including gaze analysis, determination of effective gaze area, neural network model training, and video segment optimization, high-quality film and television works are generated.
It improves the visual effects of FPV film and television works, enhances the user experience, reduces labor costs, and increases the efficiency of film and television production.
Smart Images

Figure CN116614592B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and more specifically to an FPV-based film and television shooting method and an FPV-based film and television shooting device. Background Technology
[0002] With the continuous development of technology, the miniaturization and high-definition of cameras have greatly expanded their application scenarios, allowing people to meet a wider range of needs by wearing cameras on their bodies or setting them on terminal devices.
[0003] In the traditional field of photography, a script is usually prepared in advance, and then a professional actor performs the scenes. During the performance, a professional films the actors using a camera, thus creating a film based on the script. However, with the continuous development of information technology, people are exposed to more and more film and television works in their daily lives, leading to a certain degree of aesthetic fatigue. People are no longer satisfied with simple films and television shows shot by cameras, but expect to obtain a more immersive and realistic visual experience.
[0004] To address these technical challenges, researchers have considered mounting cameras on people or on devices to capture footage from a first-person perspective, aiming to provide more immersive cinematic experiences. However, in practice, several issues arise. First, the human camera's movement during filming inevitably leads to scenes that don't align with the intended cinematic experience, such as checking footing or surveying obstacles. Second, when filming from a device, random camera shake can significantly degrade the visual experience and the overall creative effect of the film. Summary of the Invention
[0005] In order to overcome the above-mentioned technical problems in the prior art, the present invention provides a film and television shooting method and device based on FPV. By acquiring multiple FPV-based shooting videos from different positions on the shooting subject and performing video fusion processing, invalid video segments are effectively filtered out, thereby achieving a better film and television work generation effect.
[0006] To achieve the above objectives, embodiments of the present invention provide a film and television shooting method based on FPV. The method includes: acquiring a first FPV-based shooting video and acquiring a second FPV-based shooting video, wherein the first shooting video and the second shooting video are shooting videos taken at different positions on a shooting subject; determining a main video and a secondary video in the first shooting video and the second shooting video; performing a first-view analysis on the main video to obtain a corresponding main first-view video segment; performing a first-view analysis on the secondary video to obtain a corresponding secondary first-view video segment; and performing a fusion processing on the main first-view video segment based on the secondary first-view video segment to generate a corresponding film and television work.
[0007] Preferably, determining the primary and secondary videos in the first and second captured videos includes: performing gaze analysis on the first and second captured videos to obtain corresponding first and second gazes; determining the effective gaze area corresponding to the current film and television creation; obtaining the first duty cycle of the first gaze within the effective gaze area and the second duty cycle of the second gaze within the effective gaze area; and determining the primary and secondary videos in the first and second captured videos based on the analysis of the first and second duty cycles.
[0008] Preferably, the method further includes: obtaining a preset neural network model before performing first-person perspective analysis on the main video; determining a film and television training dataset corresponding to the current film and television creation; training the preset neural network model based on the film and television training dataset to generate a film and television analysis model; and performing first-person perspective analysis on the main video and the secondary video respectively based on the film and television analysis model.
[0009] Preferably, the step of fusing the secondary first-view video segment based on the primary first-view video segment to generate a corresponding film or television work includes: determining an interrupted video segment based on the primary first-view video segment; extracting a replacement video segment corresponding to the interrupted video segment from the secondary first-view video segment; performing transition optimization processing on the connecting frames of the replacement video segment to obtain an optimized video segment; and performing fusion processing on the primary first-view video segment and the optimized video segment to generate a corresponding film or television work.
[0010] Preferably, the transition frame includes a start transition frame and an end transition frame. The step of performing transition optimization processing on the transition frames of the replacement video segment to obtain an optimized video segment includes: obtaining a preceding transition frame and a following transition frame corresponding to the replacement video segment from the main first-view video segment; performing line-of-sight analysis on the preceding transition frame to generate a preceding line of sight, and performing line-of-sight analysis on the following transition frame to generate a following line of sight; performing image distortion processing on a preset number of start transition frames of the replacement video segment based on the preceding line of sight to obtain a pre-processed video segment; and performing image distortion processing on a preset number of end transition frames of the pre-processed video segment based on the following line of sight to obtain an optimized video segment.
[0011] Accordingly, the present invention also provides a film and television shooting device based on FPV, the device comprising: a video acquisition unit, configured to acquire a first FPV-based shooting video and a second FPV-based shooting video, wherein the first shooting video and the second shooting video are shooting videos from different positions on the shooting subject; a primary and secondary video determination unit, configured to determine the primary video and the secondary video in the first shooting video and the second shooting video; a first analysis unit, configured to perform a first-view analysis on the primary video to obtain a corresponding primary first-view video segment; a second analysis unit, configured to perform a first-view analysis on the secondary video to obtain a corresponding secondary first-view video segment; and a work generation unit, configured to perform fusion processing on the primary first-view video segment based on the secondary first-view video segment to generate a corresponding film and television work.
[0012] Preferably, the primary and secondary video determination unit includes: a gaze analysis module, used to perform gaze analysis on the first and second captured videos to obtain corresponding first and second gazes; an effective area determination module, used to determine the effective gaze area corresponding to the current film and television creation; a duty cycle acquisition module, used to acquire the first duty cycle of the first gaze within the effective gaze area and the second duty cycle of the second gaze within the effective gaze area; and a primary and secondary video determination module, used to analyze and determine the primary and secondary videos in the first and second captured videos based on the first and second duty cycles.
[0013] Preferably, the device further includes a model training unit, which is specifically used for: obtaining a preset neural network model before performing a first-view analysis on the main video; determining a film and television training dataset corresponding to the current film and television creation; training the preset neural network model based on the film and television training dataset to generate a film and television analysis model; and performing first-view analysis on the main video and the secondary video based on the film and television analysis model.
[0014] Preferably, the work generation unit includes: an interrupted video segment determination module, used to determine an interrupted video segment based on the main first-person perspective video segment; a replacement video segment determination module, used to extract a replacement video segment corresponding to the interrupted video segment from the secondary first-person perspective video segment; an optimization module, used to perform transition optimization processing on the connecting frames of the replacement video segment to obtain an optimized video segment; and a fusion module, used to perform fusion processing on the main first-person perspective video segment and the optimized video segment to generate a corresponding film and television work.
[0015] Preferably, the connecting frames include a start connecting frame and an end connecting frame. The optimization module is specifically used to: obtain a preceding connecting frame and a following connecting frame corresponding to the replacement video segment from the main first-view video segment; perform line-of-sight analysis on the preceding connecting frame to generate a preceding line of sight, and perform line-of-sight analysis on the following connecting frame to generate a following line of sight; perform image distortion processing on a preset number of start connecting frames of the replacement video segment based on the preceding line of sight to obtain a pre-processed video segment; and perform image distortion processing on a preset number of end connecting frames of the pre-processed video segment based on the following line of sight to obtain an optimized video segment.
[0016] The present invention has at least the following technical effects through the technical solution provided by the present invention:
[0017] By acquiring FPV-based video footage from different positions of the subject, extracting effective first-person perspective video segments from the acquired footage, and merging the videos according to their primary and secondary relationships, an FPV film with optimal shooting effects is obtained, greatly improving the visual effects of FPV films and enhancing the user experience.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0020] Figure 1 This is a flowchart illustrating the specific implementation of the FPV-based film shooting method provided in this embodiment of the invention.
[0021] Figure 2 This is a flowchart illustrating the specific implementation of determining the main video and secondary video according to an embodiment of the present invention;
[0022] Figure 3 This is a flowchart illustrating the specific implementation of generating film and television works according to an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of the FPV-based film and television shooting device provided in an embodiment of the present invention. Detailed Implementation
[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0025] In this invention, the terms "system" and "network" are used interchangeably. "Multiple" refers to two or more; therefore, in this invention, "multiple" can also be understood as "at least two." "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that in the description of this invention, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.
[0026] Please see Figure 1 This invention provides a film shooting method based on FPV, the method comprising:
[0027] S10) Acquire a first FPV-based video and a second FPV-based video, wherein the first video and the second video are video taken at different positions on the subject.
[0028] S20) Determine the main video and secondary video in the first captured video and the second captured video;
[0029] S30) Perform first-view analysis on the main video to obtain the corresponding main first-view video segment;
[0030] S40) Perform first-view analysis on the secondary video to obtain the corresponding secondary first-view video segment;
[0031] S50) Based on the secondary first-person perspective video segment, the primary first-person perspective video segment is fused to generate the corresponding film and television work.
[0032] In one possible implementation, a first FPV (First Person View) video and a second FPV video are first acquired. The first and second FPV videos are taken from different positions on the subject. These different positions can be locations on the subject with different motion characteristics. Specifically, in this embodiment, the PFV video is captured using a human body. During the shooting process, multiple video shooting devices can be set on the human body. For example, the first shooting device is set on the head and moves with the head to obtain a video with a first-person view effect; the second shooting device is set on the chest, with its orientation consistent with the human body's orientation.
[0033] After obtaining the first and second captured videos, it is necessary to first determine the primary and secondary videos to identify the baseline videos for subsequent video fusion, thereby improving the accuracy of the fusion process. Please refer to [link / reference]. Figure 2 In this embodiment of the invention, determining the primary video and secondary video in the first and second captured videos includes:
[0034] S21) Perform line-of-sight analysis on the first captured video and the second captured video to obtain the corresponding first line of sight and second line of sight;
[0035] S22) Determine the effective visual area corresponding to the current film and television creation;
[0036] S23) Obtain the first duty cycle of the first line of sight within the effective line of sight area, and the second duty cycle of the second line of sight within the effective line of sight area, respectively;
[0037] S24) Based on the analysis of the first duty cycle and the second duty cycle, determine the main video and the secondary video in the first captured video and the second captured video.
[0038] In one possible implementation, gaze analysis is first performed on the first and second captured videos to obtain the corresponding first and second gaze lines. Specifically, the center point of each frame in the captured videos can be extracted, and the perspective direction of each frame can be analyzed to comprehensively determine the gaze line of each frame, i.e., determine the gaze line of the corresponding captured video. On the other hand, the effective gaze area corresponding to the current film and television creation is determined. Specifically, the effective gaze area of each scene in FPV can be determined based on the main scene content that needs to be shot in different scenarios of the current film and television creation. For example, when creating FPV video in the first scene, since the shooting theme and shooting direction of the scene are relatively clear, the effective video area in the scene can be the area corresponding to the smaller field of view in front of the shooting subject; when creating FPV video in the second scene, since there are more obstacles and interference in the scene, the effective field of view in the scene can be the area corresponding to the larger field of view around the shooting subject.
[0039] At this point, the first duty cycle of the first line of sight within the effective line of sight area and the second duty cycle of the second line of sight within the effective line of sight area are obtained respectively. Based on the first and second duty cycles, the primary and secondary videos in the first and second captured videos are determined. For example, in the first embodiment, it is necessary to shoot FPV film works with relatively stable shots. Therefore, by setting the effective line of sight area to a smaller area, the duty cycle of different lines of sight within the effective line of sight area is adjusted to increase the sensitivity to changes in the line of sight. At the same time, the primary and secondary rule is determined to determine the captured video with a larger duty cycle as the primary video and the captured video with a smaller duty cycle as the secondary video. In the second embodiment, it is necessary to shoot FPV film works with relatively active shots. Therefore, by expanding the area of the effective line of sight area, the sensitivity to changes in the line of sight is reduced. At the same time, the primary and secondary rule is determined to determine the captured video with a smaller duty cycle as the primary video and the captured video with a larger duty cycle as the secondary video.
[0040] In this embodiment of the invention, by distinguishing between primary and secondary video footage shot from different perspectives, specifically by distinguishing between primary and secondary video footage based on actual shot stabilization requirements, the video content that meets the shooting requirements can be preserved to the maximum extent during the subsequent video fusion process, thereby improving the visual effect of the final generated film and television work and enhancing the user experience.
[0041] After determining the main video and secondary video, a first-person perspective analysis is performed on the main video and secondary video respectively to identify the first-person perspective video segments in each video. This filters out video segments that do not meet the actual shooting requirements, ensuring that every segment in the final merged film and television work is a valid segment that is actually needed.
[0042] In existing technologies, the selection of video clips is mainly done manually. However, this method requires a huge amount of manpower, which increases the shooting cost and reduces the efficiency of film and television production. Therefore, in order to solve this technical problem, a method is used to train a machine model to recognize the shot video in order to obtain first-person perspective video clips that meet the requirements.
[0043] In this embodiment of the invention, the method further includes: obtaining a preset neural network model; determining a film and television training dataset corresponding to the current film and television creation before performing a first-view analysis on the main video; training the preset neural network model based on the film and television training dataset to generate a film and television analysis model; and performing a first-view analysis on the first shot video and the second shot video based on the film and television analysis model.
[0044] In one possible implementation, before performing first-person perspective analysis on the captured video, a corresponding neural network model can be trained based on the actual shooting needs of the film and television production, thereby assisting in the intelligent first-person perspective analysis of the subsequent captured video. Specifically, firstly, a preset neural network model is obtained, for example, a commonly used neural network model. Then, a film and television training dataset corresponding to the current film and television production is determined. For example, this film and television training dataset can be pre-captured images of the scenes and content required for the current film and television production. These pre-captured images are all on-site images based on FPV shooting. The aforementioned pre-captured images are used as the film and television training dataset, and then the preset neural network is trained using this film and television training dataset to generate a film and television analysis model. After generating the film and television analysis model, after subsequently obtaining the captured video, the film and television analysis model is used to perform first-person perspective analysis on the captured video to obtain the corresponding first-person perspective video segments.
[0045] For example, in one embodiment, a main video and a secondary video with a duration of 5 minutes are obtained. After performing a first-person perspective analysis on the main video, it is found that the video segments from 0-50s, 75-120s, 130-270s, and 280-300s are normal video segments, while the remaining video segments have abnormal line of sight or severe shaking. Therefore, the above video segments are taken as the main first-person perspective video segments. Based on the same principle, the corresponding secondary first-person perspective video segments are obtained.
[0046] At this point, to create a complete film or television work, specifically, the primary first-person perspective video segments are merged based on the secondary first-person perspective video segments to generate a complete film or television work.
[0047] Please see Figure 3 In this embodiment of the invention, the step of fusing the primary first-view video segment based on the secondary first-view video segment to generate a corresponding film or television work includes:
[0048] S51) Determine the interrupted video segment based on the main first-view video segment;
[0049] S52) Extract a replacement video segment corresponding to the interrupted video segment from the secondary first-view video segment;
[0050] S53) Perform transition optimization processing on the connecting frames of the replaced video segment to obtain the optimized video segment;
[0051] S54) Perform fusion processing on the main first-person perspective video segment and the optimized video segment to generate the corresponding film and television work.
[0052] In one possible implementation, an interrupted video segment is first determined based on the aforementioned primary first-view video segment. Then, a replacement video segment corresponding to the interrupted video segment is extracted from the secondary first-view video segment. Specifically, the time period corresponding to the interrupted video segment is mapped to the secondary first-view video segment to extract the replacement video segment corresponding to that time. In practical applications, if the replacement video segment is directly replaced by the interrupted video segment, since the first and second captured videos are shot from different positions of the subject (i.e., their perspectives are different), there will inevitably be a drastic change in the video image at the video transition point, which will have a significant visual impact on the user and reduce the user experience. To solve this technical problem, when replacing the replacement video segment with the interrupted video segment, a transition optimization process needs to be performed on the replacement video segment to quickly and smoothly transition the image to the other perspective, reducing the visual impact on the user.
[0053] In this embodiment of the invention, the connecting frame includes a start connecting frame and an end connecting frame. The step of performing transition optimization processing on the connecting frames of the replacement video segment to obtain an optimized video segment includes: obtaining a preceding connecting frame and a following connecting frame corresponding to the replacement video segment from the main first-view video segment; performing line-of-sight analysis on the preceding connecting frame to generate a preceding line of sight, and performing line-of-sight analysis on the following connecting frame to generate a following line of sight; performing image distortion processing on a preset number of start connecting frames of the replacement video segment based on the preceding line of sight to obtain a pre-processed video segment; and performing image distortion processing on a preset number of end connecting frames of the pre-processed video segment based on the following line of sight to obtain an optimized video segment.
[0054] In one possible implementation, when replacing a video segment, the preceding and following connecting frames corresponding to the video segment to be replaced are first obtained from the main first-view video segment. For example, when video replacement is required in the 45-60s video segment of the main first video segment, the 44s (i.e., the preceding connecting frame) and the 61s (i.e., the following connecting frame) are extracted. Of course, it should be noted that in order to better determine the implementation of changes, multiple consecutive frames can also be extracted as preceding or following connecting frames, which will not be elaborated on here.
[0055] After obtaining the aforementioned preceding and following connecting frames, line-of-sight analysis is performed on them to generate corresponding preceding and following lines of sight. Then, based on the preceding line of sight, image distortion processing is applied to a preset number (e.g., 10 frames) of frames before the replacement video segment to be inserted. For example, the distortion effect can be a gradual change from the preceding line of sight to the line of sight where the replacement video segment is located, thereby achieving the preceding processing of the replacement video segment. Then, based on the following line of sight, image distortion processing is further applied to a preset number of ending connecting frames of the preceding processed video segment to obtain the final optimized video segment.
[0056] At this point, the optimized video segment is inserted into the corresponding position in the first-person perspective video segment, thereby achieving the fusion processing of the first-person perspective video segment and the optimized video segment, and generating the final film and television work.
[0057] It should be noted that, in the embodiments of the present invention, multiple (more than two) shooting devices can also be used to shoot FPV videos at multiple locations on the shooting subject to provide a more reliable video source and further optimize the visual effects of the final generated film and television works. This should be something that those skilled in the art can easily conceive of based on the embodiments of the present invention, and should also fall within the protection scope of the present invention. It will not be elaborated further here.
[0058] In this embodiment of the invention, FPV video is automatically fused and generated by using shooting videos with multiple perspectives obtained from multiple positions of the shooting subject. This effectively filters and optimizes the problems of severe image shaking, frequent image switching, and invalid images that exist when using a single shooting source or a single shooting perspective. As a result, the film and television works produced can better meet the actual visual needs of users and improve the user experience.
[0059] The FPV-based film shooting device provided in the embodiments of the present invention will be described below with reference to the accompanying drawings.
[0060] Please see Figure 4Based on the same inventive concept, this invention provides an FPV-based film and television shooting device, comprising: a video acquisition unit for acquiring a first FPV-based shot video and a second FPV-based shot video, wherein the first shot video and the second shot video are shot videos from different positions on a shooting subject; a primary and secondary video determination unit for determining the primary video and the secondary video in the first shot video and the second shot video; a first analysis unit for performing a first-view analysis on the primary video to obtain a corresponding primary first-view video segment; a second analysis unit for performing a first-view analysis on the secondary video to obtain a corresponding secondary first-view video segment; and a work generation unit for performing a fusion processing on the primary first-view video segment based on the secondary first-view video segment to generate a corresponding film and television work.
[0061] In this embodiment of the invention, the primary and secondary video determination unit includes: a gaze analysis module, used to perform gaze analysis on the first and second captured videos to obtain corresponding first and second gazes; an effective area determination module, used to determine the effective gaze area corresponding to the current film and television creation; a duty cycle acquisition module, used to acquire the first duty cycle of the first gaze within the effective gaze area and the second duty cycle of the second gaze within the effective gaze area; and a primary and secondary video determination module, used to analyze and determine the primary and secondary videos in the first and second captured videos based on the first and second duty cycles.
[0062] In this embodiment of the invention, the device further includes a model training unit, which is specifically used for: obtaining a preset neural network model before performing a first-view analysis on the main video; determining a film and television training dataset corresponding to the current film and television creation; training the preset neural network model based on the film and television training dataset to generate a film and television analysis model; and performing a first-view analysis on the main video and the secondary video based on the film and television analysis model.
[0063] In this embodiment of the invention, the work generation unit includes: an interrupted video segment determination module, used to determine an interrupted video segment based on the main first-view video segment; a replacement video segment determination module, used to extract a replacement video segment corresponding to the interrupted video segment from the secondary first-view video segment; an optimization module, used to perform transition optimization processing on the connecting frames of the replacement video segment to obtain an optimized video segment; and a fusion module, used to perform fusion processing on the main first-view video segment and the optimized video segment to generate a corresponding film and television work.
[0064] In this embodiment of the invention, the connecting frame includes a start connecting frame and an end connecting frame. The optimization module is specifically used to: obtain a preceding connecting frame and a following connecting frame corresponding to the replacement video segment from the main first-view video segment; perform line-of-sight analysis on the preceding connecting frame to generate a preceding line of sight, and perform line-of-sight analysis on the following connecting frame to generate a following line of sight; perform image distortion processing on a preset number of start connecting frames of the replacement video segment based on the preceding line of sight to obtain a pre-processed video segment; and perform image distortion processing on a preset number of end connecting frames of the pre-processed video segment based on the following line of sight to obtain an optimized video segment.
[0065] The optional embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention.
[0066] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not describe the various possible combinations separately.
[0067] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0068] Furthermore, various different implementations of the present invention can be combined arbitrarily, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed in the present invention.
Claims
1. A first-person view (FPV) based film shooting method, characterized in that, The method comprises: acquiring a first shooting video based on a first-person view (FPV) and a second shooting video based on a first-person view (FPV), the first shooting video and the second shooting video being shooting videos of different positions on a shooting subject; determining a main video and a secondary video in the first shooting video and the second shooting video; performing first-view analysis on the main video to obtain a corresponding main first-view video segment; performing first-view analysis on the secondary video to obtain a corresponding secondary first-view video segment; based on the secondary first-view video segment, performing fusion processing on the main first-view video segment to generate a corresponding film and television work; The determination of the main video and the secondary video in the first shooting video and the second shooting video comprises: performing line-of-sight analysis on the first shooting video and the second shooting video to obtain a corresponding first line of sight and a second line of sight; determining an effective line-of-sight area corresponding to the current film and television creation; respectively acquiring a first duty cycle of the first line of sight within the effective line-of-sight area and a second duty cycle of the second line of sight within the effective line-of-sight area; based on the first duty cycle and the second duty cycle analysis, determining the main video and the secondary video in the first shooting video and the second shooting video; based on the main first-view video segment, determining an interrupted video segment; extracting a replacement video segment corresponding to the interrupted video segment from the secondary first-view video segment; performing transition optimization processing on the connection frames of the replacement video segment to obtain an optimized video segment; performing fusion processing on the main first-view video segment and the optimized video segment to generate a corresponding film and television work; The connection frames include start connection frames and end connection frames, and the transition optimization processing on the connection frames of the replacement video segment to obtain an optimized video segment comprises: acquiring a pre-connection frame and a post-connection frame corresponding to the replacement video segment from the main first-view video segment; performing line-of-sight analysis on the pre-connection frame to generate a pre-line-of-sight, and performing line-of-sight analysis on the post-connection frame to generate a post-line-of-sight; based on the pre-line-of-sight, performing image distortion processing on a preset number of start connection frames of the replacement video segment to obtain a pre-processed video segment; based on the post-line-of-sight, performing image distortion processing on a preset number of end connection frames of the pre-processed video segment to obtain an optimized video segment. The method further comprises:
2. The method of claim 1, wherein, before performing first-view analysis on the main video, acquiring a preset neural network model; determining a film and television training data set corresponding to the current film and television creation; based on the film and television training data set, training the preset neural network model to generate a film and television analysis model; based on the film and television analysis model, performing first-view analysis on the main video and the secondary video respectively. The device comprises:
3. A first-person view (FPV) based film shooting device, characterized in that, The video acquisition unit is configured to acquire a first shooting video based on a first-person view (FPV) and acquire a second shooting video based on the first-person view (FPV), the first shooting video and the second shooting video being shooting videos of different positions on a shooting subject; The main and secondary determination unit is configured to determine a main video and a secondary video from among the first shooting video and the second shooting video; The first analysis unit is configured to perform first-view analysis on the main video to obtain a corresponding main first-view video segment; The second analysis unit is configured to perform first-view analysis on the secondary video to obtain a corresponding secondary first-view video segment; The work generation unit is configured to perform fusion processing on the main first-view video segment based on the secondary first-view video segment to generate a corresponding film and television work. The main and secondary determination unit includes: The line-of-sight analysis module is configured to perform line-of-sight analysis on the first shooting video and the second shooting video to obtain a corresponding first line of sight and a second line of sight; The effective area determination module is configured to determine an effective line-of-sight area corresponding to a current film and television creation; The duty cycle acquisition module is configured to respectively acquire a first duty cycle of the first line of sight within the effective line-of-sight area and a second duty cycle of the second line of sight within the effective line-of-sight area; The main and secondary determination module is configured to determine a main video and a secondary video from among the first shooting video and the second shooting video based on the first duty cycle and the second duty cycle. The work generation unit includes: The interrupt video segment determination module is configured to determine an interrupt video segment based on the main first-view video segment; The replacement video segment determination module is configured to extract a replacement video segment corresponding to the interrupt video segment from the secondary first-view video segment; The optimization module is configured to perform transition optimization processing on a link frame of the replacement video segment to obtain an optimized video segment; The fusion module is configured to perform fusion processing on the main first-view video segment and the optimized video segment to generate a corresponding film and television work. The link frame includes a start link frame and an end link frame, and the optimization module is specifically configured to: Acquire a preceding link frame and a subsequent link frame corresponding to the replacement video segment from the main first-view video segment; Perform line-of-sight analysis on the preceding link frame to generate a preceding line of sight, and perform line-of-sight analysis on the subsequent link frame to generate a subsequent line of sight; Perform image distortion processing on a preset number of start link frames of the replacement video segment based on the preceding line of sight to obtain a preceding-processed video segment; Perform image distortion processing on a preset number of end link frames of the preceding-processed video segment based on the subsequent line of sight to obtain an optimized video segment.
4. The apparatus of claim 3, wherein, The device further includes a model training unit, which is specifically configured to: Acquire a preset neural network model before performing first-view analysis on the main video; Determine a film and television training data set corresponding to a current film and television creation; Train the preset neural network model based on the film and television training data set to generate a film and television analysis model; Perform first-view analysis on the main video and the secondary video based on the film and television analysis model, respectively.
Citation Information
Patent Citations
Wireless head tracker design method based on Kalman filtering and PPM coding
CN105974948A
Systems and methods for augmented stereoscopic display
CN109154499A