A video editing method and device, electronic equipment and storage medium
By performing attribute analysis on the scene in the video frame and utilizing the neural radiation field model, target video frames are generated, which solves the problem of the great limitations of video editing in the existing technology and realizes a variety of video editing effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-08
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video editing technologies have significant limitations and cannot achieve diverse video editing effects.
By performing attribute analysis on the scene in the sample video frames, spatial and semantic information is obtained. The target video frames are then generated using a neural radiation field model, enabling scene camera movement and shape editing.
It enables diverse video editing effects, including scene camera movement and shape editing, and generates highly realistic 3D models, reducing editing difficulty and time costs.
Smart Images

Figure CN119788920B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer, and particularly, to a video editing method and device, electronic equipment and storage medium. BACKGROUND
[0002] The existing video editing technology can generally realize video cutting, splicing, sound adjustment, speed change, and adding special effects (such as subtitles, stickers, filters, transitions, etc.). The editing method of the related art has great limitations. SUMMARY
[0003] Embodiments of the present disclosure provide a video editing method, device, electronic equipment and storage medium, which can realize diversified video editing.
[0004] In a first aspect, the embodiments of the present disclosure provide a video editing method, comprising:
[0005] performing attribute analysis on a scene in a sample video frame to obtain scene attribute information; wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame;
[0006] determining a target editing effect; wherein the target editing effect includes at least one of the following: editing scene panning and editing scene shape;
[0007] generating a target video frame based on the scene attribute information and the target editing effect through a neural radiance field model; wherein the neural radiance field model is constructed based on the sample video frame.
[0008] In a second aspect, the embodiments of the present disclosure further provide a video editing device, comprising:
[0009] an analysis module configured to perform attribute analysis on a scene in a sample video frame to obtain scene attribute information; wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame;
[0010] an effect determination module configured to determine a target editing effect; wherein the target editing effect includes at least one of the following: editing scene panning and editing scene shape;
[0011] an editing module configured to generate a target video frame based on the scene attribute information and the target editing effect through a neural radiance field model; wherein the neural radiance field model is constructed based on the sample video frame.
[0012] In a third aspect, the embodiments of the present disclosure further provide an electronic equipment, comprising:
[0013] one or more processors;
[0014] Storage device for storing one or more programs.
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the video editing method as described in any of the embodiments of this disclosure.
[0016] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video editing method as described in any of the embodiments of this disclosure.
[0017] The technical solution of this disclosure embodiment performs attribute analysis on the scene in the sample video frame to obtain scene attribute information; wherein, the scene attribute information includes at least the spatial information and semantic information corresponding to the scene in the sample video frame; determines the target editing effect; wherein, the target editing effect includes at least one of the following: editing scene camera movement and editing scene shape; and generates a target video frame based on the scene attribute information and the target editing effect through a neural radiation field model; wherein, the neural radiation field model is constructed based on the sample video frame.
[0018] Since the neural radiation field model is built upon sample video frames, it possesses the ability to represent scenes within those frames from any angle and distance. By analyzing the attributes of scenes within the sample video frames, at least spatial and semantic scene attribute information can be obtained, enabling scene understanding. Based on this scene attribute information, the neural radiation field model can edit the scenes in the sample video frames according to the desired editing effect, performing camera movements and / or shape editing to generate the target video frame, thus achieving diverse video editing capabilities. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a flowchart illustrating a video editing method provided in an embodiment of the present disclosure;
[0021] Figure 2 This is a schematic diagram of the scene coordinate reconstruction process in a video editing method provided in this embodiment of the present disclosure;
[0022] Figure 3 This is a flowchart illustrating a video editing method provided in an embodiment of the present disclosure;
[0023] Figure 4 A schematic diagram of changing a scene shooting dolly in a video editing method provided by an embodiment of the present disclosure;
[0024] Figure 5 A schematic diagram of changing a scene shape in a video editing method provided by an embodiment of the present disclosure;
[0025] Figure 6 A flow chart of a video editing method provided by an embodiment of the present disclosure;
[0026] Figure 7 A flow chart of a video editing method provided by an embodiment of the present disclosure;
[0027] Figure 8 A schematic diagram of a new target video frame in a video editing method provided by an embodiment of the present disclosure;
[0028] Figure 9 A structural schematic diagram of a video editing device provided by an embodiment of the present disclosure;
[0029] Figure 10 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided to make the present disclosure more thorough and complete. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.
[0031] It is understood that each step recited in the method embodiments of the present disclosure can be executed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions of other terms will be given in the description below.
[0033] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the sequence or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0035] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0036] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws, regulations and provisions.
[0037] Figure 1 A video editing method flowchart provided by the embodiments of the present disclosure, the embodiments of the present disclosure are applicable to the case of video editing, such as changing the camera operation in the scene shooting in the video, and / or changing the three-dimensional shape of the scene, etc. The method can be executed by a video editing device, which can be realized in the form of software and / or hardware, and the device can be configured in an electronic device, such as a computer, etc.
[0038] As shown in Figure 1 The video editing method provided by the embodiments of the present disclosure can include:
[0039] S110, attribute analysis is performed on the scene in the sample video frame to obtain scene attribute information.
[0040] In the embodiments of the present disclosure, the video frames of the video to be edited can be obtained by existing technologies, and the sample video frames can be determined based on the video frames. For example, a frame can be extracted from the video frames as a sample video frame every predetermined frame number or predetermined time.
[0041] In some optional implementations, before the attribute analysis is performed on the scene in the sample video frame, the method can further include: selecting the sample video frame based on at least one of the following: selecting the sample video frame according to the co-view relationship between the video frames; selecting the sample video frame according to the content of the video frames.
[0042] In these optional implementations, a pose solving method such as Structure from Motion (SfM) can be used to determine the camera positions of the video frames. When the change of the camera positions is less than a preset distance, it can be considered that the video frames have the common view relationship. Then, one of the video frames with the common view relationship can be selected as a sample video frame. In this way, there is sufficient camera motion, i.e., change of the view angle, between the selected sample video frames. By selecting the sample video frames according to the common view relationship between the video frames, it is beneficial to make the sample video frames carry more complete and comprehensive scene attribute information.
[0043] In this way, the sample video frames are selected according to the content of the video frames, which can include selecting the sample video frames according to at least one index such as the definition, brightness, and noise of the content of the video frames. For example, Fourier transform can be used to analyze whether the video frames are clear, and the video frames that are relatively blurred can be filtered out, and the remaining video frames can be used as the sample video frames. By selecting the sample video frames according to the content of the video frames, it is beneficial to improve the quality of the sample video frames and make the sample video frames carry more accurate scene attribute information.
[0044] In this way, the sample video frames are selected according to the content of the video frames, which can include selecting the sample video frames according to at least one index such as the definition, brightness, and noise of the content of the video frames. For example, Fourier transform can be used to analyze whether the video frames are clear, and the video frames that are relatively blurred can be filtered out, and the remaining video frames can be used as the sample video frames. By selecting the sample video frames according to the content of the video frames, it is beneficial to improve the quality of the sample video frames and make the sample video frames carry more accurate scene attribute information.
[0045] S120, determine a target editing effect.
[0046] In this way, the sample video frames are selected according to the content of the video frames, which can include selecting the sample video frames according to at least one index such as the definition, brightness, and noise of the content of the video frames. For example, Fourier transform can be used to analyze whether the video frames are clear, and the video frames that are relatively blurred can be filtered out, and the remaining video frames can be used as the sample video frames. By selecting the sample video frames according to the content of the video frames, it is beneficial to improve the quality of the sample video frames and make the sample video frames carry more accurate scene attribute information.
[0047] When the target editing effect comprises editing a scene shape, the target editing effect can comprise at least one of the following: magnification / reduction of at least a part of the scene, stretching / compression of at least a part of the scene in a predetermined direction, shaking of at least a part of the scene around a predetermined coordinate axis, bending, etc. When the target editing effect is magnification / reduction of at least a part of the scene, the at least a part of the scene after being mapped to a three-dimensional space can be magnified and / or the at least a part of the scene can be reduced, which can create a "big world / small world" scene for the edited video, in which the scene is different from the surrounding environment. When the target editing effect is stretching / compression of at least a part of the scene in a predetermined direction, the at least a part of the scene after being mapped to a three-dimensional space can be stretched and / or compressed. When the at least a part of the scene is continuously stretched and compressed after being mapped to a three-dimensional space, the edited video can have a scene fluctuation effect. When the scene is stretched, the edited video can have an effect of passing through a wormhole or moving at a very high speed, etc. When the target editing effect is shaking of at least a part of the scene around a predetermined coordinate axis or bending, the at least a part of the object after being mapped to a three-dimensional space can be shaken around the predetermined coordinate axis and / or bent, which can create a surreal scene for the edited video.
[0048] When the target editing effect comprises editing a scene shape, the target editing effect can comprise at least one of the following: magnification / reduction of at least a part of the scene, stretching / compression of at least a part of the scene in a predetermined direction, shaking of at least a part of the scene around a predetermined coordinate axis, bending, etc. When the target editing effect is magnification / reduction of at least a part of the scene, the at least a part of the scene after being mapped to a three-dimensional space can be magnified and / or the at least a part of the scene can be reduced, which can create a "big world / small world" scene for the edited video, in which the scene is different from the surrounding environment. When the target editing effect is stretching / compression of at least a part of the scene in a predetermined direction, the at least a part of the scene after being mapped to a three-dimensional space can be stretched and / or compressed. When the at least a part of the scene is continuously stretched and compressed after being mapped to a three-dimensional space, the edited video can have a scene fluctuation effect. When the scene is stretched, the edited video can have an effect of passing through a wormhole or moving at a very high speed, etc. When the target editing effect is shaking of at least a part of the scene around a predetermined coordinate axis or bending, the at least a part of the object after being mapped to a three-dimensional space can be shaken around the predetermined coordinate axis and / or bent, which can create a surreal scene for the edited video.
[0049] In the embodiments of the present disclosure, the video editing apparatus can provide at least one user interface to interact with the user during the video editing process. The user interface may, for example, comprise a selection interface of the target editing effect, in which a selection control of at least one editing effect can be presented, and the target editing effect can be determined from the editing effects in response to a triggering operation of the selection control by the user.
[0050] The step S110 and the step S120 have no strict time sequence relationship. For example, the video editing apparatus can first push at least one editing effect based on each item of scene attribute information obtained in the step S110 in a selection interface of a target editing effect based on each item of attribute information, and then determine the target editing effect from each pushed editing effect in response to a triggering operation of a selection control by a user. For another example, the video editing apparatus can analyze the related scene attribute information according to the target editing effect after determining the target editing effect from each editing effect in response to the triggering operation of the selection control by the user.
[0051] In the step S130, the target video frame is generated based on the scene attribute information and the target editing effect through a neural radiance field model.
[0052] The neural radiance field model is constructed based on sample video frames. The neural radiance field (NeRF) is a computer vision technology that extracts the collective shape and texture information of an object from multiple view images using deep learning technology, and then generates a continuous three-dimensional radiance field using the information, so that a highly realistic three-dimensional model / scene can be presented at any angle and distance. In the embodiment, the construction of the neural radiance field model can be realized only by using sample video frames in a video shot once, and the scene in the video is automatically implicitly reconstructed, i.e., the three-dimensional information of the scene is learned.
[0053] The scene understanding can be realized through the scene attribute information, and on this basis, the camera track and / or deformation processing operation of the scene can be generated according to the target editing effect. For example, when the target editing effect is to surround the target object, the target object in the scene can be determined through the scene attribute information, and on this basis, a new camera track can be generated.
[0054] On the premise of determining the camera track and / or deformation processing operation of the scene, the neural radiance field model can draw the deformed or non-deformed scene at each view angle in the track to obtain the target video frame. Based on the existing manner, a new video can be generated according to each target video frame, so that various video editing effects can be easily realized. There is no need to repeatedly shoot a video to obtain the desired camera track, and there is no need for professional personnel to edit the video using professional editing tools. A video with a film-level effect can be output at a low threshold and a low time cost.
[0055] In addition, the implicit reconstruction of the scene in the sample video frame through the NeRF not only can guarantee a highly realistic rendering effect, but also can achieve a reconstruction time of hours compared with the traditional technology of constructing a three-dimensional model based on a depth point cloud.
[0056] The technical solution of the embodiments of the present disclosure performs attribute analysis on a scene in a sample video frame to obtain scene attribute information, wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame; a target editing effect is determined, wherein the target editing effect includes at least one of the following: editing a scene camera angle and editing a scene shape; a target video frame is generated based on the scene attribute information and the target editing effect through a neural radiance field model, wherein the neural radiance field model is constructed based on the sample video frame.
[0057] Since the neural radiance field model is constructed based on the sample video frame, the neural radiance field model can have the ability to present the scene in the sample video frame at any angle and distance. By performing attribute analysis on the scene in the sample video frame, at least scene attribute information such as spatial information and semantic information corresponding to the scene can be obtained, which can realize scene understanding. Based on the scene attribute information, the neural radiance field model can edit the scene in the sample video frame according to the target editing effect, and generate a target video frame, thereby realizing diversified video editing.
[0058] The embodiments of the present disclosure can be combined with the various optional schemes of the video editing method provided in the above embodiments. The video editing method provided in the present embodiment is described in detail in the process of performing attribute analysis on the scene in the sample video frame. By performing at least one of the analysis operations of three-dimensional coordinate reconstruction, symmetry analysis, and semantic analysis, a foundation can be laid for scene understanding.
[0059] In the video editing method provided in the present embodiment, the attribute analysis on the scene in the sample video frame can include at least one of the following: reconstructing three-dimensional coordinates corresponding to the scene in the sample video frame; analyzing the symmetry of the scene in the sample video frame; and performing semantic analysis on the scene in the sample video frame.
[0060] Although the constructed neural radiance field model can perform three-dimensional reconstruction on the scene, it lacks a unified reference coordinate system when rendering the scene from different perspectives. Therefore, when performing attribute analysis on the scene in the sample video frame, the three-dimensional coordinates corresponding to the scene in the sample video frame can be reconstructed to ensure the uniformity of the positions of the scene rendered from different perspectives. For example, the world coordinates corresponding to the scene can be determined based on the camera positions and camera parameters of each sample video frame; for another example, the constructed neural radiance field model can be used to map the scene to a three-dimensional space, and a three-dimensional coordinate system can be determined according to straight lines in the scene in the three-dimensional space.
[0061] In the scene in the sample video frame, the symmetry is analyzed, for example, whether there is a symmetry axis in the sample video frame can be identified by existing image recognition methods (such as pattern matching method, optimization search method, etc.). The symmetry axis can be the symmetry axis of the whole scene, or the symmetry axis of the local scene. For example, see Figure 5 a in the a in the Figure 5 a in the a in the a can be considered as the center line of the road as the symmetry axis, and the whole scene is symmetrical on the left and right. By analyzing the symmetry of the scene in the sample video frame, a reference can be provided for subsequent editing of the scene dolly or editing of the scene shape. For example, when editing the scene dolly, at least part of the edited dolly path can be parallel to the symmetry axis, etc.; for example, when editing the scene shape, it can be stretched or compressed along the symmetry axis direction, etc.
[0062] In the scene in the sample video frame, the semantic analysis is performed, for example, the scene in the sample video frame can be semantically segmented by the existing image segmentation method. For example, in the case of an indoor scene, wall segmentation, floor segmentation, ceiling segmentation and furniture segmentation can be performed; in the case of an outdoor scene, building segmentation, road segmentation can be performed. By performing semantic analysis on the scene in the sample video frame, a foundation can be laid for the selection of objects in subsequent editing of the scene dolly or editing of the scene shape. For example, when editing the scene dolly, the target object to be circled during the dolly can be determined according to the semantic segmentation result; for example, when editing the scene shape, the target object to be edited can be determined, etc.
[0063] In these optional implementations, the scene attribute analysis process can include at least one of three-dimensional coordinate reconstruction, symmetry analysis and semantic analysis, which lays a foundation for scene understanding. In addition, other scene attribute analysis methods can also be applied, which will not be exhausted here.
[0064] For example, Figure 2 A flowchart of a scene coordinate reconstruction process in a video editing method provided by an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, in some optional implementations, the three-dimensional coordinates corresponding to the scene in the sample video frame can be reconstructed, which can include: Figure 2
[0065] S210, detecting a straight line in the scene in each sample video frame;
[0066] S220, determining the spatial position of each straight line by a neural radiance field model;
[0067] S230, clustering each straight line based on the spatial position of each straight line to obtain a target straight line;
[0068] S240, identifying a target plane in the scene, and determining a three-dimensional coordinate system of the scene according to the target plane and the target straight line.
[0069] In this embodiment, a detection method such as Hough transform can be used to detect the straight lines of the scene in each sample video frame, that is, to obtain the two-dimensional straight lines in the sample video frame.
[0070] The depth information of the two-dimensional straight lines can be rendered using a neural radiance field model. On this basis, since the camera position is known, the two-dimensional straight lines can be mapped to three-dimensional straight lines in space to obtain the spatial positions of the three-dimensional straight lines.
[0071] After the two-dimensional straight lines of the scene in each sample video frame are mapped to straight lines in space, due to errors, multiple straight lines are included near the true straight line. The straight lines can be clustered and analyzed according to their spatial positions to obtain target straight lines with higher confidence in space.
[0072] The results of the semantic analysis of the scene in the sample video frame can be used to identify the target plane corresponding to the scene, such as the ground or a table top. The target plane can be used to represent a horizontal plane. On this basis, the target straight lines can be analyzed. For example, straight lines that are perpendicular to each other are extracted, and two of the straight lines that are perpendicular to each other can be parallel to the target plane. At this time, the two straight lines that are parallel to the target plane can represent two axes on the horizontal plane, and the straight line that is perpendicular to the target plane can represent an axis that is perpendicular to the horizontal plane, thereby ensuring that the three-dimensional coordinate system established is consistent with the semantic information of the scene itself.
[0073] In these optional implementation manners, the constructed neural radiance field model can be used to map the scene to a three-dimensional space, and the straight lines in the scene can be combined with semantic analysis to determine a three-dimensional coordinate system, so as to ensure that the three-dimensional coordinate system established is consistent with the semantic information of the scene itself.
[0074] In some optional implementation manners, the analysis of the symmetry of the scene in the sample video frame can include: mapping each symmetry axis of the scene in the sample video frame to a three-dimensional space using the constructed neural radiance field model; and determining a target symmetry axis according to a symmetry axis in the three-dimensional space that is parallel to any axis in the three-dimensional coordinate system.
[0075] In these optional implementation manners, the symmetry axes can be determined based on the created three-dimensional coordinate system. By preferentially selecting a symmetry axis that is parallel to a coordinate axis in the reconstructed three-dimensional coordinate system as the basis for subsequent determination of editing and camera movement effects, the editing effects can be enriched.
[0076] The technical solutions of the embodiments of the present disclosure are described in detail. At least one of the three-dimensional coordinate reconstruction, the symmetry analysis and the semantic analysis can lay a foundation for scene understanding. The video editing method provided by the embodiments of the present disclosure belongs to the same disclosure concept as the video editing method provided by the above embodiments. The technical details not described in detail in the embodiments can be referred to the above embodiments, and the same technical features have the same beneficial effects in the embodiments and the above embodiments.
[0077] The embodiments of the present disclosure can be combined with the various optional solutions in the video editing method provided in the above embodiments. The video editing method provided by the embodiments limits the determination of the target editing effect. The editing scene dolly and / or editing scene shape can be determined as the target editing effect according to the scene attribute information, so as to provide more automated video editing capabilities.
[0078] For example, Figure 3 A flowchart of a video editing method provided by the embodiments of the present disclosure is shown. As Figure 3 shown, the video editing method provided by the embodiments can include:
[0079] S310, attribute analysis is performed on the scene in the sample video frame to obtain scene attribute information.
[0080] The scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame.
[0081] S320, determining a target editing effect according to the scene attribute information.
[0082] The target editing effect includes at least one of the following: editing scene dolly and editing scene shape.
[0083] The dolly track can be automatically generated according to the semantic information in the scene attribute information, and the target editing effect can include editing scene dolly.
[0084] For example, the target object in the scene can be panned. The determination process of the target object can include: segmenting different objects in each sample video frame through semantic analysis; selecting the target object in response to a target object selection operation of a user;
[0085] Or, the target object can be automatically determined according to the result of semantic analysis; for example, a target focus object category is defined in advance, and the object belonging to the category in the semantic segmentation result is taken as the target object; for another example, the object appearing in each sample video frame with a frequency higher than a preset frequency is taken as the target object.
[0086] For example, the camera orbiting track can be automatically generated according to the distribution of objects in the scene. Thus, the video-level camera movement special effect can be automatically generated with simple user interaction or without interaction.
[0087] In this case, the target editing effect can include editing scene camera movement, which can be automatically generated according to the symmetry plane in the scene attribute information. For example, the Hitchcock camera movement can be performed on the symmetry plane to create suspense, tension, or highlight strong emotional changes in the character's heart.
[0088] In this case, the target object can be morphed according to the semantic information in the scene attribute information, and the target editing effect can include editing scene shape. For example, the target object can be enlarged or shaken. The determination process of the target object can refer to the above description and will not be repeated here.
[0089] For example, Figure 4 FIG. 2 is a schematic diagram of changing scene camera movement in a video editing method according to an embodiment of the present disclosure. Referring to FIG. 2, Figure 4 ,the arrow curve in a of FIG. 2 is the original camera movement track, Figure 4 the arrow curve in b of FIG. 2 is the camera movement track after editing scene camera movement. The grand atmosphere can be created by orbiting the building. Figure 4 For example,
[0090] FIG. 3 is a schematic diagram of changing scene shape in a video editing method according to an embodiment of the present disclosure. Referring to FIG. 3, Figure 5 , a of FIG. 3 is the original street shape, Figure 5 b of FIG. 3 is the street shape after editing scene shape. The surreal atmosphere can be created by bending the street upwards. Figure 5 Figure 5 S330, generating the target video frame based on the scene attribute information and the target editing effect through the neural radiance field model.
[0091] In this case, the neural radiance field model is constructed based on the sample video frame.
[0092] The technical solution of the embodiment of the present disclosure limits the determination of the target editing effect. The editing scene camera movement and / or the editing scene shape can be determined as the target editing effect according to the scene attribute information, so that more automated video editing capabilities can be provided. The video editing method provided by the embodiment of the present disclosure belongs to the same disclosure concept as the video editing method provided by the above-mentioned embodiment. The technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiment, and the same technical features have the same beneficial effects in the present embodiment and the above-mentioned embodiment.
[0093] The technical solution of the embodiment of the present disclosure limits the determination of the target editing effect. The editing scene camera movement and / or the editing scene shape can be determined as the target editing effect according to the scene attribute information, so that more automated video editing capabilities can be provided. The video editing method provided by the embodiment of the present disclosure belongs to the same disclosure concept as the video editing method provided by the above-mentioned embodiment. The technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiment, and the same technical features have the same beneficial effects in the present embodiment and the above-mentioned embodiment.
[0094] The embodiments of the present disclosure can be combined with the various optional schemes of the video editing method provided in the above embodiments. The video editing method provided in the present embodiment can also be combined with the existing image generation model and the traditional rendering link, so as to present more rich video editing effects.
[0095] Exemplarily, Figure 6 A flowchart of a video editing method provided by an embodiment of the present disclosure is shown in FIG. 6. As shown in FIG. 6, in some implementations, the video editing process can first be processed in the video dimension based on the existing image generation model, and then the target video frame can be generated by the NeRF model, which can include the following steps. Figure 6
[0096] S610, attribute analysis is performed on the scene in the sample video frame to obtain scene attribute information.
[0097] The scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame.
[0098] S620, a target editing effect is determined.
[0099] The target editing effect includes at least one of the following: editing scene panning and editing scene shape.
[0100] S630, a target video frame is generated based on the scene attribute information and the target editing effect by a neural radiance field model.
[0101] The neural radiance field model is constructed based on the sample video frame.
[0102] S640, editing is performed on the key frame of the target video frame to obtain an updated key frame.
[0103] After the target video frame of each angle of the scene is rendered by the NeRF model, at least part of the target video frame can be used as a key frame. The key frame can be edited based on the existing image generation model, for example, the key frame can be modified, enhanced, etc. The editing of the key frame can include stylized processing of the key frame.
[0104] The present embodiment will be described by taking the stylized processing of the key frame by the image generation model as an example, and other editing methods can be similarly described, and will not be described exhaustively. After the stylized processing of the key frame, a video frame with a different style from the original sample video frame can be obtained. For example, a realistic sample video frame can be edited into a new video frame in the Lego style.
[0105] S650, optical flow information between the target video frames is determined by the neural radiance field model.
[0106] The step S640 and the step S650 do not have a strict time sequence relationship.
[0107] In this embodiment, since each pixel point in the scene rendered by the NeRF model has a unique spatial coordinate, the pixel points in the generated target video frame have a spatial correspondence relationship with the pixel points in the front and back target video frames, that is, correct optical flow. The optical flow information between the target video frames can be determined by the position changes of the pixel points in the current target video frame and the corresponding pixel points in the adjacent video frames. The optical flow information can be understood as information related to the motion of the pixel points, such as the motion direction and the displacement.
[0108] S660, interpolating the updated key frames based on the optical flow information to obtain intermediate video frames.
[0109] After determining the optical flow information between the target video frames, the optical flow information of the video frames between the adjacent key frames can be used as the optical flow information of the video frames between the adjacent updated key frames. Furthermore, the adjacent updated key frames can be interpolated based on the optical flow information of the video frames between the updated key frames to obtain intermediate video frames.
[0110] For example, the optical flow information of the video frames between the stylized front adjacent key frames can be used as the optical flow information of the intermediate video frames between the stylized rear adjacent key frames. The adjacent stylized key frames can be interpolated according to the optical flow information of the intermediate frames, so that the stylized intermediate video frames can be obtained.
[0111] S670, using the updated key frames and the intermediate video frames as new target video frames.
[0112] The new target video frames can be generated from the updated key frames and the intermediate video frames between the adjacent updated key frames. For example, the stylized target video frames can be generated from the stylized key frames and the stylized intermediate video frames between the stylized key frames.
[0113] The traditional video stylization processing can be considered as processing each frame by an image generation model. After stylization, the same pixel point in the scene may have different positions in each frame, which may cause the processed video to be jittery and flickering.
[0114] In this alternative embodiment, since each pixel point in the scene rendered by the NeRF model has a unique spatial coordinate, the pixel points in the generated image have a correspondence relationship with the pixel points in the front and back images, that is, correct optical flow. The NeRF model can use the correct optical flow between the target video frames to interpolate the intermediate frames between the target video frames, thereby reducing the jitter and flicker of the edited video.
[0115] Exemplary, Figure 7 A flowchart of a video editing method provided by an embodiment of the present disclosure is shown. As Figure 7 shown, in some implementations, the video editing process can first generate target video frames through a NeRF model, then process the target video frames based on an existing image generation model, and then feed back to the NeRF model, thereby realizing the processing of the video in the spatial dimension, which can include:
[0116] S710, attribute analysis is performed on the scene in the sample video frame to obtain scene attribute information.
[0117] The scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame.
[0118] S720, determine the target editing effect.
[0119] The target editing effect includes at least one of the following: editing scene panning and editing scene shape.
[0120] S730, generate target video frames based on scene attribute information and target editing effect through a neural radiance field model.
[0121] The neural radiance field model is constructed based on the sample video frame.
[0122] S740, edit the key frame of the target video frame to obtain an updated key frame.
[0123] After rendering the target video frames of the scene from various angles using the NeRF model, at least part of the target video frames can be used as key frames. The key frames can be edited based on the functions of the existing image generation model, such as modifying, enhancing, etc. The editing of the key frames can include style processing of the key frames. This embodiment will be described by taking the style processing of the key frames by the image generation model as an example, and other editing methods will not be described.
[0124] S750, adjust the parameters in the neural radiance field model using the updated key frame, so that the neural radiance field model with adjusted parameters generates new target video frames.
[0125] After obtaining the stylized key frame, the stylized key frame can be mapped back to the NeRF model to modify the color information of each pixel point of the scene after mapping to the three-dimensional space (i.e., modify the texture of the scene in the spatial dimension). Then, the updated NeRF model can be used for subsequent rendering to realize the stylization of the original video.
[0126] In these optional embodiments, by changing the texture of the target video frame and projecting back to the NeRF model, the NeRF model can render a target video frame with consistent texture, so that the original video after style processing can have consistency.
[0127] In some optional implementations, after generating the target video frame, the method further includes: generating a new target video frame containing the target three-dimensional model according to the depth information of the scene in the target video frame and the depth information of the target three-dimensional model.
[0128] Among them, the three-dimensional model can be understood as a three-dimensional special effect material that can be added to the scene in the sample video frame. Among them, before generating a new target video frame containing the target three-dimensional model, the three-dimensional model to be added can be obtained first. For example, the three-dimensional model to be added can be obtained from the preset three-dimensional model library according to the scene attribute information of the scene in the sample video frame. Specifically, a correspondence between the three-dimensional model and the scene semantics can be pre-set, and a three-dimensional model matching the semantic information of the scene in the sample video frame can be determined from the preset three-dimensional model library based on the correspondence as the three-dimensional model to be added.
[0129] Further, a new target video frame containing the target three-dimensional model can be generated according to the depth information of the scene in the target video frame and the depth information of the target three-dimensional model.
[0130] In the traditional rendering link, rendering can be performed after the three-dimensional model (such as a triangular facet model) is constructed. The spatial position of the three-dimensional model in the scene can be predefined or determined by a user setting received through a predefined user interface. The depth information corresponding to the scene of the target video frame and the depth information of the three-dimensional model at the spatial position can be obtained through the NeRF model. On this basis, because the geometric information of the three-dimensional model is known, the three-dimensional model and the scene in the new target video frame obtained after the three-dimensional model is rendered into the three-dimensional scene can have a reasonable occlusion relationship. The depth information of the three-dimensional model can be compared with the depth information corresponding to the scene of the target video frame, and the pixels with smaller depth can be determined from the three-dimensional model and the scene for rendering, that is, the three-dimensional model can be added to the scene of the target video frame and can have a reasonable occlusion relationship with the scene.
[0131] Exemplarily, Figure 8 A schematic diagram of a new target video frame of a video editing method provided by an embodiment of the present disclosure. Referring to Figure 8 When generating the target video frame, the football model can be added to the scene according to the depth information of the scene (the football gate) in the target video frame and the depth information of the football model, and the content of the scene can be enriched.
[0132] In these optional implementations, the target video frame can also be generated in combination with a traditional rendering link, to present more abundant video editing effects.
[0133] The technical solutions of the embodiments of the present disclosure can combine the video editing process with a generative artificial intelligence model and a traditional rendering link, to present more abundant video editing effects. The video editing method provided by the embodiments of the present disclosure belongs to the same disclosure concept as the video editing method provided by the above embodiments, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the same technical features have the same beneficial effects in the present embodiment and the above embodiments.
[0134] Figure 9 A structural schematic diagram of a video editing device provided by the embodiments of the present disclosure. The video editing device provided by the present embodiment is suitable for the case of video editing, such as the case of changing the camera movement when the scene in the video is shot, and / or the case of changing the three-dimensional shape of the scene, etc.
[0135] As shown in Figure 9 The video editing device provided by the embodiments of the present disclosure can include:
[0136] The analysis module 910 is configured to perform attribute analysis on the scene in the sample video frame to obtain scene attribute information, wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame;
[0137] The effect determination module 920 is configured to determine a target editing effect, wherein the target editing effect includes at least one of the following: editing scene camera movement and editing scene shape.
[0138] The editing module 930 is configured to generate a target video frame based on the scene attribute information and the target editing effect by using a neural radiance field model, wherein the neural radiance field model is constructed based on the sample video frame.
[0139] In some optional implementations, the analysis module can be configured to perform at least one of the following attribute analyses on the scene in the sample video frame:
[0140] Reconstructing the three-dimensional coordinates corresponding to the scene in the sample video frame;
[0141] Analyzing the symmetry of the scene in the sample video frame;
[0142] Performing semantic analysis on the scene in the sample video frame.
[0143] In some optional implementations, the analysis module can be configured to:
[0144] Detecting straight lines in the scene in each sample video frame.
[0145] determine spatial positions of the straight lines by a neural radiance field model;
[0146] cluster the straight lines based on the spatial positions of the straight lines to obtain target straight lines;
[0147] identify a target plane in the scene, and determine a three-dimensional coordinate system of the scene according to the target plane and the target straight lines.
[0148] In some optional implementation manners, the effect determination module can be configured to:
[0149] determine the target editing effect according to the scene attribute information.
[0150] In some optional implementation manners, the video editing apparatus can further include:
[0151] a sample video frame selection module configured to select sample video frames based on at least one of the following manners before performing attribute analysis on the scenes in the sample video frames:
[0152] select the sample video frames according to the co-view relationship between the video frames;
[0153] select the sample video frames according to the content of the video frames.
[0154] In some optional implementation manners, the video editing apparatus can further include:
[0155] a post-editing module configured to edit key frames of the target video frames to obtain updated key frames after the target video frames are generated;
[0156] determine optical flow information between the target video frames by a neural radiance field model;
[0157] interpolate the updated key frames based on the optical flow information to obtain intermediate video frames;
[0158] use the updated key frames and the intermediate video frames as new target video frames.
[0159] In some optional implementation manners, the post-editing module can be further configured to:
[0160] correspondingly, the model updating module can be configured to adjust parameters in the neural radiance field model by using the updated key frames, so that the neural radiance field model with the adjusted parameters generates new target video frames.
[0161] In some optional implementation manners, the post-editing module can be further configured to:
[0162] After the target video frame is generated, a new target video frame containing the target three-dimensional model is generated according to depth information of a scene in the target video frame and depth information of the target three-dimensional model.
[0163] The video editing apparatus provided by the embodiments of the present disclosure can perform the video editing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.
[0164] It is worth noting that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for easy mutual distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0165] Reference will be made to the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which Figure 10 a structural diagram of an electronic device (for example, a terminal device or a server in Figure 10 1000) suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (for example, vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 10 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0166] As shown in Figure 10 , the electronic device 1000 can include a processing apparatus (for example, a central processor, a graphics processor, and the like) 1001, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 1002 or loaded from a storage apparatus 1008 into a Random Access Memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processing apparatus 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0167] In general, the following devices can be connected to the I / O interface 1005: input devices 1006, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 1007, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 1008, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 1009. The communication devices 1009 can allow the electronic device 1000 to communicate wirelessly or via a wire with other devices to exchange data. Although Figure 10 The electronic device 1000 is shown with various devices, but it is understood that all of the illustrated devices are not required to implement or be present. More or fewer devices can alternatively be implemented or present.
[0168] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 1009, or installed from the storage devices 1008, or installed from the ROM 1002. When the computer program is executed by the processing devices 1001, the above-mentioned functions defined in the video editing method of embodiments of the present disclosure are performed.
[0169] The electronic device provided by the embodiments of the present disclosure and the video editing method provided by the above-mentioned embodiments belong to the same disclosure concept, and the technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiments, and the present embodiment has the same beneficial effects as the above-mentioned embodiments.
[0170] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the video editing method provided by the above-mentioned embodiments.
[0171] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with an instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0172] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0173] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.
[0174] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to:
[0175] Attribute analysis is performed on the scene in the sample video frame to obtain scene attribute information; wherein, the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame; a target editing effect is determined; wherein, the target editing effect includes at least one of the following: editing scene panning and editing scene shape; a target video frame is generated based on the scene attribute information and the target editing effect through a neural radiance field model; wherein, the neural radiance field model is constructed based on the sample video frame.
[0176] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0177] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0178] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the names of the units and modules do not constitute a limitation on the units and modules themselves.
[0179] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0180] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0181] According to one or more embodiments of the present disclosure, a video editing method is provided, the method comprising:
[0182] performing attribute analysis on a scene in a sample video frame to obtain scene attribute information; wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame;
[0183] determining a target editing effect; wherein the target editing effect includes at least one of the following: editing scene panning and editing scene shape;
[0184] generating a target video frame based on the scene attribute information and the target editing effect through a neural radiance field model; wherein the neural radiance field model is constructed based on the sample video frame.
[0185] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0186] In some optional implementation manners, the attribute analysis on the scene in the sample video frame comprises at least one of the following:
[0187] reconstructing three-dimensional coordinates corresponding to the scene in the sample video frame;
[0188] analyzing symmetry of the scene in the sample video frame;
[0189] performing semantic analysis on the scene in the sample video frame.
[0190] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0191] In some optional implementation manners, the reconstructing three-dimensional coordinates corresponding to the scene in the sample video frame comprises:
[0192] detecting straight lines in the scene in each of the sample video frames;
[0193] determining spatial positions of each of the straight lines by using the neural radiance field model;
[0194] clustering each of the straight lines based on the spatial positions of each of the straight lines to obtain a target straight line;
[0195] identifying a target plane corresponding to the scene in the sample video frame, and determining three-dimensional coordinates corresponding to the scene according to the target plane and the target straight line.
[0196] analyzing symmetry of the scene in the sample video frame, according to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0197] In some optional implementation manners, the determining a target editing effect comprises:
[0198] determining the target editing effect according to the scene attribute information.
[0199] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0200] In some optional implementation manners, after the generating the target video frame, further comprising:
[0201] selecting the sample video frame based on at least one of the following:
[0202] selecting the sample video frame according to a co-view relationship between video frames;
[0203] selecting the sample video frame according to content of the video frames.
[0204] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0205] In some optional implementations, after the target video frames are generated, the method further comprises:
[0206] editing the key frames of the target video frames to obtain updated key frames;
[0207] determining, by the neural radiance field model, optical flow information between the target video frames;
[0208] interpolating the updated key frames based on the optical flow information to obtain intermediate video frames;
[0209] using the updated key frames and the intermediate video frames as new target video frames.
[0210] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0211] In some optional implementations, after the target video frames are generated, the method further comprises:
[0212] editing the key frames of the target video frames to obtain updated key frames;
[0213] adjusting parameters in the neural radiance field model using the updated key frames, so that the neural radiance field model with the adjusted parameters generates new target video frames.
[0214] According to one or more embodiments of the present disclosure, a video editing method is provided, further comprising:
[0215] In some optional implementations, after the target video frames are generated, the method further comprises:
[0216] generating new target video frames containing the target three-dimensional model according to depth information of the scenes in the target video frames and depth information of the target three-dimensional model.
[0217] According to one or more embodiments of the present disclosure, a video editing device is provided, comprising:
[0218] an analysis module configured to perform attribute analysis on a scene in a sample video frame to obtain scene attribute information; wherein the scene attribute information at least includes spatial information and semantic information corresponding to the scene in the sample video frame;
[0219] an effect determination module configured to determine a target editing effect; wherein the target editing effect includes at least one of the following: editing scene panning and editing scene shape;
[0220] The editing module is configured to generate a target video frame based on the scene attribute information and the target editing effect by using a neural radiance field model, wherein the neural radiance field model is constructed based on the sample video frame.
[0221] The above description is merely that of preferred embodiments of the present disclosure and a description of the principles of the technology employed. It will be understood by those skilled in the art that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or equivalent features thereof without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0222] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination.
[0223] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of the example forms of implementing the claims.
Claims
1. A video editing method characterized by, Comprise: Attribute analysis is carried out on the scene in the sample video frame, and scene attribute information is obtained;Wherein, the scene attribute information at least includes the space information and semantic information corresponding to the scene in the sample video frame; Determine the target editing effect;Wherein, the target editing effect includes, at least one of the following: editing scene dolly and editing scene shape; Through the neural radiation field model, based on the scene attribute information and the target editing effect, generate target video frame;Wherein, the neural radiation field model is constructed based on the sample video frame; After the generation of the target video frame, further comprising: Editing the key frame of the target video frame to obtain the updated key frame; Through the neural radiation field model, determine the optical flow information between the target video frames; Based on the optical flow information, the updated key frame is interpolated to obtain the intermediate video frame; The updated key frame and the intermediate video frame are used as new target video frame.
2. The method of claim 1, wherein, The attribute analysis of the scene in the sample video frame includes at least one of the following: Reconstruct the three-dimensional coordinates corresponding to the scene in the sample video frame; Analysis of the symmetry of the scene in the sample video frame; Semantic analysis of the scene in the sample video frame.
3. The method of claim 2, wherein, The reconstruction of the three-dimensional coordinates corresponding to the scene in the sample video frame includes: Detect the straight line in the scene in each of the sample video frames; Through the neural radiation field model, determine the spatial position of each straight line; Based on the spatial position of each straight line, cluster each straight line to obtain a target straight line; Identify the target plane corresponding to the scene in the sample video frame, and determine the three-dimensional coordinates corresponding to the scene according to the target plane and the target straight line.
4. The method of claim 1, wherein, The determination of the target editing effect includes: According to the scene attribute information, determine the target editing effect.
5. The method of claim 1, wherein, Before the attribute analysis of the scene in the sample video frame, further comprising: Select the sample video frame based on at least one of the following ways: According to the mutual visibility relationship between each video frame, the sample video frame is selected; According to the content of the video frame, the sample video frame is selected.
6. The method of claim 1, wherein, After the generation of the target video frame, further comprising: Editing the key frame of the target video frame to obtain the updated key frame; Using the updated key frame, adjust the parameters in the neural radiation field model, so that the adjusted neural radiation field model generates new target video frame.
7. The method of claim 1, wherein, After the generation of the target video frame, further comprising: According to the depth information of the scene in the target video frame and the depth information of the target three-dimensional model, generate new target video frame containing the target three-dimensional model.
8. A video editing apparatus characterized by comprising: Comprise: Analysis module, for attribute analysis is carried out on the scene in the sample video frame, and scene attribute information is obtained;Wherein, the scene attribute information at least includes the space information and semantic information corresponding to the scene in the sample video frame; Effect determination module, for determining the target editing effect;Wherein, the target editing effect includes, at least one of the following: editing scene dolly and editing scene shape; An editing module is configured to generate target video frames based on the scene attribute information and the target editing effect by using a neural radiance field model, wherein the neural radiance field model is constructed based on the sample video frames; A post-editing module is configured to edit key frames of the target video frames to obtain updated key frames, determine optical flow information between the target video frames by using the neural radiance field model, interpolate the updated key frames based on the optical flow information to obtain intermediate video frames, and use the updated key frames and the intermediate video frames as new target video frames.
9. An electronic device, comprising: The electronic device includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the video editing method as claimed in any one of claims 1-7.
10. A storage medium containing computer executable instructions for performing the video editing method as claimed in any one of claims 1-7 when executed by a computer processor.
Citation Information
Patent Citations
Object processing method and system and processor
CN114821675A
Video processing method and related device
CN114979785A