Panoramic video editing method, electronic device, and storage medium
By determining the scene and preset viewpoint in panoramic video editing, performing structured analysis and target detection, and designing filtering rules and camera movement algorithms, the system automatically edits out the scenes that the user is most likely to select. This solves the problem that the editing results in existing technologies do not meet user expectations, and improves the richness of video clips and user satisfaction.
Patent Information
- Application Number
- PCT/CN2024/127867
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2024-10-28
- Publication Date
- 2026-02-19
AI Technical Summary
Existing technologies cannot edit panoramic videos to meet user needs, and the editing process is time-consuming. They also cannot allow users to input preferences before the final cut, resulting in editing results that do not meet user expectations.
By determining the scene and preset viewpoint corresponding to the panoramic video data, structured analysis and target detection are performed, filtering rules and camera movement algorithms are designed, the most likely scene to be selected by the user is automatically edited, and background music and camera movement are matched according to the scene to generate video clips.
It improves the visual richness and user satisfaction of video clips, lowers the barrier to entry for users to edit videos, and enables automatic editing of videos that meet user expectations based on their needs.
Smart Images

Figure CN2024127867_19022026_PF_FP_ABST
Abstract
Description
Panorama video editing method, electronic device and storage medium
[0001] Cross-reference to Related Applications
[0002] The present application is based on the PCT patent application No. PCT / CN2024 / 112534, filed on August 15, 2024, and claims priority to the PCT patent application No. PCT / CN2024 / 112534, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present application relates to the technical field of computers, and in particular to a panorama video editing method, an electronic device and a storage medium. BACKGROUND
[0004] In related technologies, panorama data is edited mainly by analyzing picture saliency targets, tracking saliency targets one by one, and editing videos of the saliency targets. Related technologies edit videos according to the properties of target detectors, for example, if a target detector detects that a target is a person, a building or an animal, then videos of the person, the building or the animal are edited. Related technologies provide users with videos of algorithm granularity without any picture understanding, without considering the actual needs of users, and thus cannot edit videos that users want.
[0005] SUMMARY
[0006] The present application provides a panorama video editing method, an electronic device and a storage medium.
[0007] The technical solution of the present application is implemented as follows:
[0008] The present application provides a panorama video editing method, which includes:
[0009] determining a scene corresponding to panorama video data;
[0010] determining at least one preset viewing angle corresponding to the scene;
[0011] editing the panorama video data based on the at least one preset viewing angle to obtain a video segment of each preset viewing angle.
[0012] In addition, according to at least one embodiment of the present application, the method further includes:
[0013] associating and displaying each video segment of each preset viewing angle obtained by editing with a corresponding preset viewing angle name.
[0014] In addition, according to at least one embodiment of the present application, before the panorama video data is edited based on the at least one preset viewing angle, the method further includes:
[0015] performing structured analysis on the panoramic video data to obtain structured data.
[0016] In addition, according to at least one embodiment of the present application, the clipping the panoramic video data based on the at least one preset view angle comprises:
[0017] determining a filtering rule corresponding to each preset view angle in the scene;
[0018] filtering the structured data based on the filtering rule to obtain clipping data corresponding to each preset view angle;
[0019] clipping the clipping data corresponding to each preset view angle into a film.
[0020] In addition, according to at least one embodiment of the present application, the performing structured analysis on the panoramic video data comprises:
[0021] performing target detection on targets in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target;
[0022] constructing the structured data based on the video sequence of each target.
[0023] In addition, according to at least one embodiment of the present application, the filtering rule is formed based on parameters of the target.
[0024] In addition, according to at least one embodiment of the present application, the parameters of the target in the filtering rule comprise one or more of the following:
[0025] an output category of the target;
[0026] a position of the target in a video frame;
[0027] an attribute contained by the target;
[0028] a video time length occupied by the target satisfying a set condition;
[0029] a forward direction of a camera.
[0030] In addition, according to at least one embodiment of the present application, the clipping the clipping data corresponding to each preset view angle into a film comprises:
[0031] determining a score of each target in the clipping data according to an output category of all targets in the clipping data and an attribute contained by all targets;
[0032] determining whether to clip a target picture corresponding to each target into the film according to the score of each target.
[0033] Further, according to at least one of the embodiments of the present application, the editing the clip data corresponding to each preset view angle into a video clip comprises:
[0034] determining background music according to the scene;
[0035] editing the background music into the video clip.
[0036] Further, according to at least one of the embodiments of the present application, the determining background music according to the scene comprises:
[0037] determining a background music segment according to the scene and a time length of the clip data;
[0038] when the time length of the clip data is greater than the time length of the background music segment, processing the clip data to match the time length of the background music segment.
[0039] Further, according to at least one of the embodiments of the present application, the editing the clip data corresponding to each preset view angle into a video clip comprises:
[0040] determining at least one camera movement according to the clip data and the scene;
[0041] displaying the at least one camera movement to a user;
[0042] editing the clip data into a video clip based on a camera movement selected by the user.
[0043] Further, according to at least one of the embodiments of the present application, the determining the scene corresponding to the panoramic video data comprises:
[0044] inputting the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, and obtaining a first type of scene;
[0045] frame extracting the panoramic video data, content detecting video frames obtained by the frame extraction, and determining a second type of scene under the first type of scene based on the content detection result;
[0046] Correspondingly, the determining the at least one preset view angle corresponding to the scene comprises:
[0047] determining the at least one preset view angle corresponding to the first type of scene and the second type of scene.
[0048] Further, according to at least one of the embodiments of the present application, after obtaining the video clip of each preset view angle, the method further comprises:
[0049] determining a sub-clip selected by a user in the plurality of video clips;
[0050] Combine the sub-clips selected by the user in multiple video clips into a combined video.
[0051] In addition, according to at least one embodiment of the present application, the determining the at least one preset visual angle corresponding to the scene comprises:
[0052] Finding the at least one preset visual angle corresponding to the scene from a preset database, different scenes corresponding to different preset visual angles.
[0053] The embodiment of the present application provides a computer program product, comprising a computer program, when the computer program is executed by a processor, the steps of the panoramic video editing method are realized.
[0054] At least one embodiment of the present application provides an electronic device, comprising:
[0055] The electronic device comprises a processor and a memory, the processor and the memory are connected to each other, wherein the memory is used to store a computer program, when the computer program is executed by the processor, the processor is configured to determine a scene corresponding to panoramic video data; determine at least one preset visual angle corresponding to the scene; edit the panoramic video data based on the at least one preset visual angle, to obtain a video clip of each preset visual angle.
[0056] In addition, according to at least one embodiment of the present application, the processor is configured to associate and display each video clip of the preset visual angle obtained by editing with the corresponding preset visual angle name.
[0057] In addition, according to at least one embodiment of the present application, the processor is configured to perform structured analysis on the panoramic video data to obtain structured data.
[0058] In addition, according to at least one embodiment of the present application, the processor is configured to determine a filtering rule corresponding to each preset visual angle under the scene; filter the structured data based on the filtering rule to obtain editing data corresponding to each preset visual angle; and edit the editing data corresponding to each preset visual angle into a video clip.
[0059] In addition, according to at least one embodiment of the present application, the processor is configured to perform target detection on the target in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target; and construct the structured data based on the video sequence of each target.
[0060] In addition, according to at least one embodiment of the present application, the filtering rule is formed based on the parameters of the target.
[0061] In addition, according to at least one embodiment of the present application, the parameters of the target in the filtering rule comprise one or more of the following:
[0062] an output category of the target;
[0063] a position of the target in the video frame;
[0064] an attribute contained by the target;
[0065] a video time length occupied by the target satisfying a set condition;
[0066] a forward direction of the camera.
[0067] In addition, according to at least one embodiment of the present application, the processor is configured to determine a score of each target in the clip data according to an output category of all targets and an attribute contained by all targets in the clip data; and determine whether to clip a corresponding target picture into a film according to the score of each target.
[0068] In addition, according to at least one embodiment of the present application, the processor is configured to determine background music according to the scene; and clip the background music into a film.
[0069] In addition, according to at least one embodiment of the present application, the processor is configured to determine a background music segment according to the scene and a time length of the clip data; and process the clip data to match a time length of the background music segment when the time length of the clip data is greater than the time length of the background music segment.
[0070] In addition, according to at least one embodiment of the present application, the processor is configured to determine at least one camera movement according to the clip data and the scene; display the at least one camera movement to a user; and clip the clip data into a film based on a camera movement selected by the user.
[0071] In addition, according to at least one embodiment of the present application, the processor is configured to input the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, and obtain a first type of scene; frame the panoramic video data, perform content detection on the video frames obtained by framing, determine a second type of scene in the first type of scene based on the content detection result, and determine at least one preset view angle corresponding to the first type of scene and the second type of scene.
[0072] In addition, according to at least one embodiment of the present application, the processor is configured to determine sub-clips selected by a user in a plurality of video clips; combine the sub-clips selected by the user in the plurality of video clips into a film to obtain a combined video.
[0073] In addition, according to at least one embodiment of the present application, the processor is configured to find at least one preset visual angle corresponding to the scene from a preset database, different scenes corresponding to different preset visual angles.
[0074] The embodiment of the present application provides a computer readable storage medium, comprising: the computer readable storage medium stores a computer program. The computer program is executed by a processor to realize the steps of the panoramic video clipping method provided in the first aspect of the embodiment of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0075] Fig. 1 is a schematic diagram of an implementation flow of a panoramic video clipping method provided by the embodiment of the present application;
[0076] Fig. 2 is a schematic diagram of an application program (APP) video clipping interface provided by the embodiment of the present application;
[0077] Fig. 3 is a schematic diagram of another APP video clipping interface provided by the embodiment of the present application;
[0078] Fig. 4 is a schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0079] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0080] The related art performs summary analysis on panoramic video data, selects representative segments through deduplication, then performs picture saliency target analysis, and tracks the saliency targets one by one after the analysis to obtain the video of the saliency target. The related art has the following defects:
[0081] The related art tracks the saliency targets one by one without any picture understanding, and many passers-by are selected, and the main body is selected incorrectly to cause the user to not obtain the desired shot. The related art needs to track the saliency targets one by one, and the analysis is time-consuming, and the target required by the user cannot be selected according to the scene recognition. The related art cannot allow the user to input preferences before the film is made, and the entire design is fully automatic, and the user cannot eliminate the target that he or she does not want before and after the film is made.
[0082] The related technology is to show the user a video of an algorithm granularity (target detector), for example, the target detector detects that the target is a person, a building or an animal, and then video clips are performed for the person, the building or the animal, and the video containing the person, the building or the animal is given to the user. The related technology has no picture understanding and does not consider the actual needs of the user, and the video clips are most likely not what the user wants, and the user needs to further clip by himself.
[0083] In view of the above-mentioned defects of the related technology, the embodiment of the present application provides a panoramic video clipping method which can automatically determine the picture most likely to be selected by the user in the current scene and automatically clip. In order to describe the technical solutions of the present application, the following will be described by specific embodiments.
[0084] FIG. 1 is a schematic diagram of the implementation process of a panoramic video clipping method provided by the embodiment of the present application, and the execution subject of the panoramic video clipping method is an electronic device. Referring to FIG. 1, the panoramic video clipping method comprises:
[0085] S101, determining a scene corresponding to panoramic video data.
[0086] The panoramic video is obtained by a panoramic camera, and the panoramic camera is composed of two or more fisheye lenses, and the obtained panoramic video is an original spherical video. The panoramic video data in this embodiment can be a collection of panoramic video data obtained by a plurality of panoramic cameras.
[0087] The scene in this embodiment includes scenes such as parent-child, pet, cycling, skiing, diving, sports and travel.
[0088] When the scene is determined, the scene can be determined by receiving the scene information input by the user. Or the scene is analyzed by recognizing the picture target in the panoramic video data.
[0089] For example, by detecting the picture saliency target in the panoramic video, it is detected that the video picture includes snow and a snowboard, and it can be determined that the scene is skiing. Or it is detected that the video picture includes a diving mask, fish or aquatic organisms, and it can be determined that the scene is diving.
[0090] S102, determining at least one preset visual angle corresponding to the scene.
[0091] In an embodiment, the at least one preset visual angle corresponding to the scene can be found from a preset database, and different scenes correspond to different preset visual angles.
[0092] Each preset visual angle corresponds to a lens in the scene, and different lenses correspond to different shooting visual angles. Different lenses correspond to different lens motion trajectories, speeds, time lengths and shooting objects.
[0093] In an embodiment, one preset view angle represents a type of shooting view angle in a corresponding scene.
[0094] Here, at least one preset view angle corresponding to each scene can be pre-stored in a preset database, and the at least one preset view angle associated with the scene information can be obtained by querying the preset database.
[0095] The at least one preset view angle corresponding to each scene is the view angle most likely to be selected / interested by the user in the scene. The user demand is pre-counted, and the at least one preset view angle corresponding to each scene is summarized.
[0096] For example, the at least one preset view angle corresponding to the skiing scene includes a skiing main character, a skiing highlight, a ski field view, a skiing friend, and a skiing forward direction. The at least one preset view angle corresponding to the diving scene includes a diving main character, an underwater creature, an underwater view, a diving highlight, and a diving partner. The at least one preset view angle corresponding to the parent-child scene includes a cute baby, a parent-child interaction, and a parent-child selfie.
[0097] The camera of the skiing main character is a selfie of the skiing scene, the camera of the skiing highlight is a skiing wrestling and a sports highlight, the camera of the ski field view is a landscape area and a special building of the skiing scene, the camera of the skiing friend is a person with a similar transportation tool / moving tool as the skiing main character, and the camera of the skiing forward direction is a skiing forward direction. The camera of the diving main character is a selfie of the diving scene, the camera of the underwater creature is an aquatic creature in the diving scene, the camera of the diving highlight is a highlight action of the diving main character, and the camera of the diving partner is a person diving in the diving scene. The camera of the cute baby is a child, the camera of the parent-child interaction is an adult and a child, and the camera of the parent-child selfie is a person holding a selfie stick.
[0098] In S103, the panoramic video data is clipped based on the at least one preset view angle to obtain a video clip of each preset view angle.
[0099] The panoramic video data includes a large number of video pictures of different targets and different view angles. According to the at least one preset view angle corresponding to the scene, the video pictures matching the preset view angle are selected from the panoramic video data to obtain a video clip of the preset view angle, thereby obtaining a video clip of each preset view angle in the at least one preset view angle.
[0100] A filtering algorithm can be designed for each preset view angle. The filtering algorithm is used to filter the video data corresponding to the preset view angle from the panoramic video data, and the unqualified targets and the video data not meeting the view angle are screened out. Then, the video clip of the preset view angle is obtained by clipping.
[0101] At least one preset view angle of the embodiment of the application is a view angle most likely selected by the user in the scene. The embodiment of the application is to edit the panoramic video data from the market and the user's perspective, and not to give the user algorithm granularity. The video clip of the preset view angle obtained by the embodiment of the application is the picture most likely selected by the user in the scene, and the picture content of the video clip can enable the user to understand what lens is played in the scene. For example, in the cycling scene, the content of each video clip can enable the user to understand what view angle corresponds to the cycling scene.
[0102] The embodiment of the application determines at least one preset view angle corresponding to the scene by determining the scene corresponding to the panoramic video data. The panoramic video data is edited based on the at least one preset view angle to obtain a video clip of each preset view angle. The corresponding scene is determined by picture understanding of the panoramic video data, and then at least one preset view angle is determined based on the scene. The embodiment of the application determines the preset view angle under the condition of picture understanding, so that each video clip obtained by editing is the picture most likely selected by the user in the current scene, the picture richness of the video clip is improved, the user satisfaction is improved, and the user editing threshold is reduced.
[0103] In an embodiment, the method further comprises:
[0104] The video clip of each preset view angle obtained by editing is associated with and displayed with the corresponding preset view angle name.
[0105] Many users edit videos in an application (APP). An APP interface is shown in FIG. 2. The video clip of each preset view angle obtained by editing is associated with and displayed with the corresponding view angle name. The user can slide left and right to switch the video clip of different preset view angles, or click the view angle name above to switch the video clip. When the video clip of a preset view angle is switched to, the view angle name of the preset view angle is highlighted and displayed in bold. As shown in FIG. 2, the content of the video clip is a skier, and in the view angle name bar above, “skiing selfie” is highlighted and displayed in bold. The view angle name displayed in the APP interface is from the user's scene or the user's language. The embodiment of the application associates and displays the video clip of each preset view angle with the corresponding view angle name, can intuitively enable the user to understand the general content of the video clip and what lens is played in the scene, and improves the user's experience.
[0106] In an embodiment, before the panoramic video data is edited based on the at least one preset view angle, the method further comprises:
[0107] The panoramic video data is analyzed to obtain structured data.
[0108] Video structured analysis is a process of analyzing and recognizing video content through intelligent analysis algorithms, extracting key information such as people, vehicles, behaviors, etc., and converting these information into text information that can be understood by computers and humans. This technology can convert raw, unstructured video data into structured data, facilitating subsequent search, query and application.
[0109] For example, a structured database of panoramic video can be directly generated through video image content analysis or feature recognition.
[0110] In an embodiment, the video data is edited based on the at least one preset perspective, comprising:
[0111] Determining the filtering rule corresponding to each preset perspective in the scene;
[0112] Filtering the structured data based on the filtering rule to obtain the editing data corresponding to each preset perspective;
[0113] Editing the editing data corresponding to each preset perspective into a film.
[0114] Here, a filtering rule can be designed for each preset perspective. The filtering rule is designed based on the structured data. The features of the structured data corresponding to the preset perspective constitute the corresponding filtering rule, which is beneficial for the filtering rule to filter the structured data, screen out unqualified targets and structured data that do not meet the perspective, and obtain the editing data corresponding to the preset perspective.
[0115] In an embodiment, the video data is structured analyzed, comprising:
[0116] Detecting the target in the video data based on a target detection algorithm to obtain the video sequence of each target;
[0117] The structured data is constituted based on the video sequence of each target.
[0118] Wherein, the target in the video data includes people, animals, buildings, sculptures, vehicles, sports equipment, plants, aquatic organisms, and natural landscape objects such as reefs.
[0119] The target detection algorithm is used to detect the target in the video data. The target detection is performed on each video frame. The detection results of the same target in multiple video frames constitute the video sequence of the target. The video sequences of all targets constitute the structured data.
[0120] Here, a target detection algorithm is used to detect a type of target, and different target detection algorithms are needed for different types of target detection. When using a target detection algorithm to detect a target, a target detection box is displayed on the target in the video frame.
[0121] The target detection algorithm can also be referred to as a detector. For example, a detector of the detector category "person" is needed to detect a target of the category "person" in panoramic video data.
[0122] The detector categories used by the embodiments of the present application include "person", "animal", "vehicle", "sports equipment", "building", "sculpture", etc. The detector to be used can be pre-set, and when performing target detection on panoramic video data, the pre-set detector is used to perform target detection on the panoramic video data.
[0123] When a salient target is detected in the current panoramic video frame being detected, a preset target tracking algorithm is used to track the salient target in the subsequent panoramic video frames of the frame. The target tracking algorithm can include but is not limited to a High-speed Tracking with Kernelized Correlation filters (KCF) algorithm or an Accurate Scale Estimation for Robust Visual Tracking (DSST) algorithm.
[0124] For each tracking segment, a tracking sequence containing a detection box is included. Specifically, each frame in the segment contains a detection box, and each detection box contains multiple attributes.
[0125] Taking a person class detection box as an example, the attributes of the person class detection box include actions (jumping, running, eating, playing skateboard, etc.), orientations (front, side, back), etc.
[0126] Among them, after detecting a target, a classifier can be used to evaluate the target. For example, an action classifier can be used to determine whether a person is performing a specific action. An orientation classifier can be used to determine the orientation of the human body, and different orientations correspond to different scores.
[0127] A large number of available tracking segments will be generated for a single video. To improve user editing efficiency, different Filters (filter or filtering rules) can be designed to filter different tracking segments, such as segments of selfie takers and travel companions. In addition, to improve the running efficiency of the Filter, the Filter is implemented based on structured data.
[0128] In an embodiment, the filtering rule is formed based on parameters of the target.
[0129] In an embodiment, the parameters of the target in the filtering rule include one or more of the following:
[0130] an output category of the target;
[0131] a position of the target in a video frame;
[0132] an attribute contained by the target;
[0133] a video duration occupied by the target satisfying a set condition;
[0134] a forward direction of a camera.
[0135] The position of the target in the video frame refers to a position of a detection box corresponding to the target in the video frame, and the attribute contained by the target includes a selfie-taker, etc.
[0136] Each preset visual angle corresponds to a Filter, and at least one Filter corresponding to each scene needs to be defined in advance. For example, one filtering rule is to extract clip data of children running or playing from structured data. For another example, one filtering rule is to extract clip data of a skier performing a highlight action of skiing from structured data.
[0137] The Filter is a filtering rule corresponding to a preset visual angle. According to the definition of the preset Filter, the structured data corresponding to the panoramic video data is filtered to obtain clip data corresponding to each preset visual angle.
[0138] In an embodiment, the clip data corresponding to each preset visual angle is clipped into a film, including:
[0139] determining a score of each target in the clip data according to an output category of all targets in the clip data and an attribute contained by all targets;
[0140] determining whether to clip a target picture corresponding to each target into the film according to the score of the target.
[0141] Here, the attribute contained by the target includes an identified action of the target, which can be an output result of an action classifier.
[0142] A score can be set for each category of the target, and the target is scored in combination with the action of the target. If the target is performing a highlight action or is doing something that meets the condition, the score of the target will be higher.
[0143] According to the score of each target, it is determined whether the corresponding target picture is cut into a film, for example, the targets with scores too small can be excluded according to the order from small to large, and the pictures of the targets with high scores are cut into a film.
[0144] When detecting the behavior of a person, human object interaction (HOI) detection can be performed, and whether the person and the object have an interaction relationship is determined according to whether the human detection frame and the other category object detection frame have an intersection, and the specific thing that the person is doing can be detected according to the HOI relationship. For example, the person and the bicycle have an HOI relationship, and it is considered that the person is riding the bicycle. The goal of human object interaction (HOI) detection is to locate the person, the object and their interaction behavior in the picture. When the human detection frame and the object detection frame have a certain IOU intersection, it is determined that there is an HOI relationship.
[0145] The area of the detection frame can be calculated, and the area can filter out targets that are too large or too small. If the number of targets detected by the detector exceeds N, the targets can be sorted in descending order of area, and only the first N targets are taken to participate in evaluation.
[0146] In an embodiment, the determining the scene corresponding to the panoramic video data comprises:
[0147] inputting the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, and obtaining a first type of scene;
[0148] frame extraction is performed on the panoramic video data, content detection is performed on the video frames obtained through the frame extraction, and a second type of scene in the first type of scene is determined based on the content detection result.
[0149] Correspondingly, the determining the at least one preset view angle corresponding to the scene comprises:
[0150] The at least one preset view angle is determined according to the first type of scene and the second type of scene.
[0151] The scene classification model is a pre-trained classification model, which is trained according to panoramic video data and labels (scenes) until the output of the scene classification model meets the user requirements. Here, the first type of scene output by the scene classification model is a large category, and the scene classification model cannot further subdivide the scene. The second type of scene is a subdivided scene in the first type of scene.
[0152] For example, the first type of scene includes a skiing scene, a diving scene, a forest scene, a cycling scene, a general sports scene, etc. Frame extraction can be performed at an Intra-coded Frame, which is an important frame in a video sequence that contains complete image information. Content detection is performed on the video frames obtained by frame extraction, and a second type of scene is determined based on the content detection result in the first type of scene.
[0153] For example, if the first type of scene output by the scene classification model is a forest scene, and it is detected that the extracted video frames include multiple targets (people) of the same type and the targets interact closely, it can be determined that the second type of scene is a parent-child outing scene or a family outing scene.
[0154] For example, if the first type of scene output by the scene classification model is a cycling scene, and it is detected that the extracted video frames include multiple cyclists, and the multiple cyclists are close in distance and follow each other for a long time, it can be determined that the second type of scene is a group cycling scene.
[0155] For example, if the first type of scene output by the scene classification model is a cycling scene, and it is detected that the extracted video frames include only one cyclist, and the cyclist is identified as a child, and an adult is close to the cyclist, it can be determined that the second type of scene is a parent-child scene.
[0156] According to the first type of scene and the second type of scene, at least one preset view angle corresponding to the first type of scene and the second type of scene is determined. The preset view angle determined according to the first type of scene and the second type of scene is more in line with user needs. For example, the first type of scene is skiing, and the second type of scene is a snow-covered forest scene, and at least one preset view angle corresponding to the first type of scene and the second type of scene includes a ski resort view and a skiing protagonist. For example, the first type of scene is diving, and the second type of scene is underwater reefs, and at least one preset view angle corresponding to the first type of scene and the second type of scene includes an underwater view and a diving protagonist.
[0157] In an embodiment, the editing of the clip data corresponding to each preset view angle into a film includes:
[0158] Determining background music according to the scene;
[0159] Editing the background music into the film.
[0160] The background music corresponding to each scene can be pre-set in the database, and the background music matching the scene can be selected to improve the quality of the film. The selected background music can also be spliced with the video segments, thereby improving the rhythm of the video film.
[0161] In an embodiment, the determining of the background music according to the scene includes:
[0162] determine a background music segment according to the scene and a time length of the clip data;
[0163] when the time length of the clip data is greater than the time length of the background music segment, processing the clip data to match the time length of the background music segment.
[0164] The background music segment needs to match the time length of the clip data. If the time length of the clip data is greater than the time length of the background music segment, the clip data can be processed according to the time length of the background music segment, for example, deleting unimportant video frames in the clip data, so that the time length of the clip data is equal to the time length of the background music segment, so that the background music is more matched with the video segment, and the quality of the film is improved.
[0165] In an embodiment, the clip data corresponding to each preset visual angle is edited into a film, comprising:
[0166] determine at least one camera movement according to the clip data and the scene;
[0167] display the at least one camera movement to the user;
[0168] edit the clip data into a film based on the camera movement selected by the user.
[0169] Wherein, the clip data is structured data after screening, a camera movement recommendation model can be trained, using a large number of existing scenes and structured data, taking the camera movement as a label, training the camera movement recommendation model until the camera movement output by the camera movement recommendation model can meet the user's demand. After training the camera movement recommendation model, taking the clip data and the scene as input, obtaining at least one camera movement output by the camera movement recommendation model.
[0170] For example, riding or driving scenes with a stable forward direction can recommend using a forward direction flip camera movement; travel, shopping, and passing through buildings scenes can recommend an overhead passing through buildings camera movement; walking around a target or passing through some targets can recommend a target curve variable speed camera movement; travel, shopping, and walking through buildings scenes with buildings on both sides can recommend a forward up-pan rotation camera movement; walking, riding, or driving scenes with a stable forward direction and buildings or sculptures on both sides can recommend a forward direction super camera movement; scenes with multiple people, at least one of whom is moving or making a spectacular action, can recommend a highlight cut target camera movement.
[0171] At least one of the camera movements can be displayed on the user's operation interface, for example, as shown in FIG. 3, in the video clip interface of the APP, at least one of the camera movements can be arranged at the bottom of the interface for the user to select, the user selects which camera movement, the video is automatically clipped using the camera movement, and the video is played in the video playing area of the interface. The camera movement selected by the user is highlighted in the APP interface, for example, the camera movement 2 in FIG. 3 is the camera movement selected by the user.
[0172] Some general camera movements such as the pan-tilt-zoom (PTZ) head, unmanned aerial vehicle (UAV) and the like can also be provided as preset camera movements for the user.
[0173] In an embodiment, after obtaining the video clip of each preset view angle, the method further comprises:
[0174] determining the sub-clip selected by the user in the multiple video clips;
[0175] combining the sub-clip selected by the user in the multiple video clips to obtain a combined video.
[0176] After obtaining the video clip of each preset view angle, the video clips can be mixed and cut, and all the video clips are cut by default.
[0177] In the embodiment, the user can select the required video clip, and can select the sub-clip in the video clip. The sub-clip refers to a part of the video clip, and the user can select multiple sub-clips in one video clip. All the sub-clips selected by the user are mixed and cut, and the combined video is obtained by automatically using the camera movement and the music. The embodiment can automatically generate the video, recommend the best music and camera movement through the algorithm, reduce the user's cutting threshold, and improve the picture richness of the video.
[0178] As shown in FIG. 2, the APP interface has virtual buttons of segmented export and immediate video generation at the bottom. The segmented export refers to exporting all the video clips to the user according to the preset view angle, and the immediate video generation refers to mixing and cutting all the video clips of the preset view angle to generate a video for the user.
[0179] The embodiment of the application provides a panoramic video clip process. The panoramic video frame is processed in a structured manner to generate multiple interest target sequences (Proposals) as raw materials for cutting. The Proposals are filtered according to the Filter, the highlight clips are extracted and de-duplicated, and finally the plane video is generated by projection, camera movement, transition and music.
[0180] Among them, the panoramic video content structured analysis includes in-camera analysis and download to mobile phone / computer analysis. When editing, the user can specify a certain perspective and picture target, and the best composition and camera operation are recommended by the algorithm. The embodiments of the application support automatic generation of a film, or the user specifies the target of interest from the filter to generate a film. The embodiments of the application can identify picture elements and scenes, help users deconstruct, filter out the most likely selected pictures in the current scene, and automatically operate the camera and edit, thereby reducing the user editing threshold and improving picture richness.
[0181] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0182] It should be understood that when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.
[0183] It should be noted that the technical solutions described in the embodiments of the application can be arbitrarily combined without conflict.
[0184] In addition, in the embodiments of the application, "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0185] The panoramic video editing device provided by the embodiments of the application comprises:
[0186] The first determination module is configured to determine a scene corresponding to the panoramic video data.
[0187] The second determination module is configured to determine at least one preset perspective corresponding to the scene.
[0188] The editing module is configured to edit the panoramic video data based on the at least one preset perspective, to obtain a video segment of each preset perspective.
[0189] In an embodiment, the device further comprises:
[0190] The display module is configured to display each video segment of the preset perspective obtained by editing in association with the corresponding preset perspective name.
[0191] In an embodiment, the device further comprises:
[0192] The structured module is configured to perform structured analysis on the panoramic video data to obtain structured data.
[0193] In an embodiment, the clipping module is configured to:
[0194] determine a filtering rule corresponding to each preset view angle under the scene;
[0195] filter the structured data based on the filtering rule to obtain clipping data corresponding to each preset view angle;
[0196] clip the clipping data corresponding to each preset view angle into a film.
[0197] In an embodiment, the structuring module is configured to:
[0198] perform target detection on targets in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target;
[0199] construct the structured data based on the video sequence of each target.
[0200] In an embodiment, the filtering rule is formed based on parameters of targets.
[0201] In an embodiment, the parameters of targets in the filtering rule include one or more of:
[0202] an output category of the target;
[0203] a position of the target in a video frame;
[0204] an attribute contained by the target;
[0205] a video time length occupied by the target that meets a set condition;
[0206] a forward direction of a camera.
[0207] In an embodiment, the clipping module is specifically configured to:
[0208] determine a score of each target in the clipping data according to output categories of all targets in the clipping data and attributes contained by all targets;
[0209] determine whether to clip a target picture corresponding to each target into a film according to the score of each target.
[0210] In an embodiment, the clipping module is specifically configured to:
[0211] determine background music according to the scene;
[0212] clip the background music into a film.
[0213] In an embodiment, the clipping module is specifically configured to:
[0214] determine a background music segment according to the scene and a time length of the clip data;
[0215] when the time length of the clip data is greater than the time length of the background music segment, process the clip data to match the time length of the background music segment.
[0216] In an embodiment, the clip module is specifically configured to:
[0217] determine at least one camera movement according to the clip data and the scene;
[0218] display the at least one camera movement to a user;
[0219] clip the clip data into a video segment based on a user-selected camera movement.
[0220] In an embodiment, the first determining module is specifically configured to: input the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, and obtain a first type of scene.
[0221] frame the panoramic video data, perform content detection on the video frames obtained by the framing, and determine a second type of scene under the first type of scene based on the content detection result.
[0222] determine at least one preset view angle corresponding to the first type of scene and the second type of scene.
[0223] In an embodiment, the apparatus further includes a combined video segment module configured to:
[0224] determine a sub-segment selected by a user from a plurality of video segments;
[0225] combine the sub-segment selected by the user from the plurality of video segments into a video segment to obtain a combined video.
[0226] In an embodiment, the second determining module is specifically configured to:
[0227] find at least one preset view angle corresponding to the scene from a preset database, wherein different scenes correspond to different preset view angles.
[0228] In actual application, the first determining module, the second determining module and the clip module can be implemented by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a field-programmable gate array (FPGA).
[0229] It should be noted that the panoramic video editing device provided in the above embodiment is only used as an example to illustrate the division of the above modules during panoramic video editing. In actual applications, the above processing can be completed by different modules according to needs, that is, the internal structure of the device is divided into different modules to complete all or part of the above-described processing. In addition, the panoramic video editing device and the panoramic video editing method provided in the above embodiment belong to the same concept, and the specific implementation process is described in the method embodiment, which will not be described here.
[0230] The panoramic video editing device can be in the form of an image file, which can be executed in the form of a container or a virtual machine to implement the panoramic video editing method described in the present application. Of course, it is not limited to the form of an image file, and any software form that can implement the panoramic video editing method described in the present application is within the protection scope of the present application.
[0231] Based on the hardware implementation of the above program modules, and in order to implement the method of the present application, the present application further provides an electronic device. FIG. 4 is a schematic diagram of the hardware composition structure of the electronic device according to an embodiment of the present application. As shown in FIG. 4, the electronic device includes:
[0232] a communication interface capable of information interaction with other devices such as network devices;
[0233] a processor connected with the communication interface to realize information interaction with other devices;
[0234] a memory storing a computer program, when the computer program is executed by the processor, the processor is configured to determine a scene corresponding to panoramic video data; determine at least one preset viewing angle corresponding to the scene; and edit the panoramic video data based on the at least one preset viewing angle to obtain a video segment of each preset viewing angle.
[0235] In an embodiment, the processor is configured to associate and display each video segment of the preset viewing angle obtained by editing with a corresponding preset viewing angle name.
[0236] In an embodiment, the processor is configured to perform structured analysis on the panoramic video data to obtain structured data.
[0237] In an embodiment, the processor is configured to determine a filtering rule corresponding to each preset viewing angle under the scene; filter the structured data based on the filtering rule to obtain editing data corresponding to each preset viewing angle; and edit the editing data corresponding to each preset viewing angle into a segment.
[0238] In an embodiment, the processor is configured to perform target detection on the targets in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target; and form the structured data based on the video sequence of each target.
[0239] In an embodiment, the filtering rule is formed based on parameters of the targets.
[0240] In an embodiment, the parameters of the targets in the filtering rule include one or more of the following:
[0241] an output category of the target;
[0242] a position of the target in a video frame;
[0243] an attribute contained by the target;
[0244] a video time length occupied by the target that meets a set condition;
[0245] a forward direction of a camera.
[0246] In an embodiment, the processor is configured to determine a score of each target in the clip data according to the output category of all targets and the attribute contained by all targets in the clip data; and determine whether to clip the corresponding target picture into a film according to the score of each target.
[0247] In an embodiment, the processor is configured to determine background music according to the scene; and clip the background music into a film.
[0248] In an embodiment, the processor is configured to determine a background music segment according to the scene and a time length of the clip data; and perform processing on the clip data to match the time length of the background music segment when the time length of the clip data is greater than the time length of the background music segment.
[0249] In an embodiment, the processor is configured to determine at least one camera movement according to the clip data and the scene; display the at least one camera movement to a user; and perform clip processing on the clip data based on a camera movement selected by the user.
[0250] In an embodiment, the processor is configured to input the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, to obtain a first type of scene; perform frame extraction on the panoramic video data, perform content detection on the video frames obtained by the frame extraction, and determine a second type of scene under the first type of scene based on the content detection result; and determine at least one preset view angle corresponding to the first type of scene and the second type of scene.
[0251] In an embodiment, the processor is configured to determine sub-clips selected by the user in the plurality of video clips; and combine the sub-clips selected by the user in the plurality of video clips into a video, to obtain a combined video.
[0252] In an embodiment, the processor is configured to find at least one preset view angle corresponding to the scene from a preset database, different scenes corresponding to different preset view angles.
[0253] Of course, in actual applications, various components in the electronic device are coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between the components. The bus system includes a data bus, a power supply bus, a control bus, and a status signal bus, in addition to the data bus. However, for the purpose of clear illustration, various buses are marked as a bus system in FIG. 4.
[0254] The memory in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of the data include any computer programs used for operation on the electronic device.
[0255] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM). The magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), and the like.
[0256] Memory), Dynamic Random Access Memory (DRAM)
[0257] The memories described in this application are intended to include, but are not limited to, these and any other suitable types of memories. These include: Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).
[0258] The methods disclosed in the embodiments of this application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory. The processor reads the program from the memory and, in conjunction with its hardware, completes the steps of the aforementioned method.
[0259] Optionally, when the processor executes the program, it implements the corresponding processes implemented by the electronic device in the various methods of the embodiments of this application. For the sake of brevity, these will not be described in detail here.
[0260] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a processor of an electronic device to perform the steps described in the method of this application.
[0261] In an example embodiment, the embodiments of the present application also provide a storage medium, specifically a computer storage medium, for example, a first memory for storing a computer program, which can be executed by a processor of an electronic device to complete the steps of the foregoing method. The computer readable storage medium can be a FRAM, a ROM, a PROM, an EPROM, an EEPROM, a Flash Memory, a magnetic surface memory, an optical disc, or a CD-ROM, etc.
[0262] In several embodiments provided in the present application, it should be understood that the disclosed apparatus, electronic device and method can be implemented by other manners. The apparatus embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be through some interfaces, indirect coupling or communication connection of the devices or units, which can be electrical, mechanical or other forms.
[0263] The units described above as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place or distributed on a plurality of network units; part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0264] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0265] Those of ordinary skill in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps including the above method embodiments when executed; and the foregoing storage medium includes a mobile storage device, a ROM, a RAM, a magnetic disc or an optical disc, and various media that can store program codes.
[0266] Alternatively, the above-mentioned integrated units of the present application, if realized in the form of software function modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: mobile storage devices, ROM, RAM, magnetic disks or optical disks, and various media that can store program codes.
[0267] It should be noted that the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0268] In addition, in the examples of the present application, "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0269] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A panoramic video clip method, comprising: determining a scene corresponding to panoramic video data; determining at least one preset view angle corresponding to the scene; clipping the panoramic video data based on the at least one preset view angle to obtain a video segment of each preset view angle.
2. The method of claim 1, wherein, The method further comprises: associating and displaying the video segment of each preset view angle obtained by clipping with a corresponding preset view angle name.
3. The method of claim 1, wherein, Before the clipping of the panoramic video data based on the at least one preset view angle, the method further comprises: performing structured analysis on the panoramic video data to obtain structured data.
4. The method of claim 3, wherein, The clipping of the panoramic video data based on the at least one preset view angle comprises: determining a filtering rule corresponding to each preset view angle under the scene; filtering the structured data based on the filtering rule to obtain clip data corresponding to each preset view angle; clipping the clip data corresponding to each preset view angle into a segment.
5. The method of claim 3, wherein, The structured analysis on the panoramic video data comprises: performing target detection on targets in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target; constructing the structured data based on the video sequence of each target.
6. The method of claim 4, wherein, The filtering rule is formed based on parameters of the target.
7. The method of claim 6, wherein, The parameters of the target in the filtering rule include one or more of the following: an output category of the target; a position of the target in a video frame; an attribute contained by the target; a video time length occupied by the target satisfying a set condition; a camera forward direction.
8. The method of claim 7, wherein, The clipping of the clip data corresponding to each preset view angle into a segment comprises: determining a score of each target in the clip data according to output categories of all targets in the clip data and attributes contained by all targets; determining whether to clip a corresponding target picture into a segment according to the score of each target.
9. The method of claim 1, wherein, The determination of the scene corresponding to the panoramic video data comprises: inputting the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model to obtain a first type of scene; performing frame extraction on the panoramic video data, performing content detection on video frames obtained by frame extraction, and determining a second type of scene under the first type of scene based on the content detection result; Correspondingly, the determination of the at least one preset view angle corresponding to the scene comprises: determining the at least one preset view angle corresponding to the first type of scene and the second type of scene.
10. The method of claim 4, wherein, The clipping of the clip data corresponding to each preset view angle into a segment comprises: determining background music according to the scene; clipping the background music into a segment.
11. The method of claim 10, wherein, The determination of the background music according to the scene comprises: determining a background music segment according to the scene and a time length of the clip data; when the time length of the clip data is greater than the time length of the background music segment, processing the clip data to match the time length of the background music segment.
12. The method of claim 4, wherein, The clipping of the clip data corresponding to each preset view angle into a segment comprises: determining at least one camera movement according to the clip data and the scene; displaying the at least one camera movement to a user; clipping the clip data into a segment based on a camera movement selected by the user.
13. The method of claim 1, wherein, After obtaining the video clip of each preset perspective, the method further comprises: determining a sub-clip selected by the user from the plurality of video clips; combining the sub-clip selected by the user from the plurality of video clips to obtain a combined video.
14. The method of claim 1, wherein, The determination of the at least one preset perspective corresponding to the scene comprises: finding the at least one preset perspective corresponding to the scene from a preset database, different scenes corresponding to different preset perspectives.
15. An electronic device comprising a memory and a processor, the memory storing a computer program, the processor being configured to determine a scene corresponding to panoramic video data when the computer program is executed by the processor. determining the at least one preset perspective corresponding to the scene; clipping the panoramic video data based on the at least one preset perspective to obtain a video clip of each preset perspective.
16. The electronic device of claim 15, wherein, The processor is configured to associate and display the video clip of each preset perspective obtained by clipping with the corresponding preset perspective name.
17. The electronic device of claim 15, wherein the processor is configured to perform structured analysis on the panoramic video data to obtain structured data.
18. The electronic device of claim 17, wherein, The processor is configured to determine a filtering rule corresponding to each preset perspective under the scene; filter the structured data based on the filtering rule to obtain clipping data corresponding to each preset perspective; clipping the clipping data corresponding to each preset perspective into a film.
19. The electronic device of claim 17, wherein, The processor is configured to perform target detection on the targets in the panoramic video data based on a target detection algorithm to obtain a video sequence of each target; The structured data is constituted based on the video sequence of each target.
20. The electronic device of claim 18, wherein, The filtering rule is formed based on the parameters of the target.
21. The electronic device of claim 20, wherein, The parameters of the target in the filtering rule comprise one or more of the following: output category of the target; position of the target in a video frame; attributes contained by the target; video time length occupied by the target satisfying a set condition; camera forward direction.
22. The electronic device of claim 21, wherein, The processor is configured to determine a score of each target in the clipping data according to the output category of all targets in the clipping data and the attributes contained by all targets; and determine whether to clip the corresponding target picture into a film according to the score of each target.
23. The electronic device of claim 18, wherein, The processor is configured to determine background music according to the scene; and clip the background music into a film.
24. The electronic device of claim 23, wherein, The processor is configured to determine a background music segment according to the scene and the length of the clipping data; and process the clipping data to match the length of the background music segment when the length of the clipping data is greater than the length of the background music segment.
25. The electronic device of claim 18, wherein, The processor is configured to determine at least one camera movement according to the clipping data and the scene; display the at least one camera movement to the user; and clip the clipping data into a film based on the camera movement selected by the user.
26. The electronic device of claim 15, wherein, The processor is configured to input the panoramic video data into a scene classification model to obtain a scene classification result output by the scene classification model, obtain a first type of scene; perform frame extraction on the panoramic video data, perform content detection on the video frames obtained by frame extraction, determine a second type of scene under the first type of scene based on the content detection result, and determine the at least one preset perspective corresponding to the first type of scene and the second type of scene.
27. The electronic device of claim 15, wherein, The processor is configured to determine sub-clips selected by a user in the plurality of video clips; and combine the sub-clips selected by the user in the plurality of video clips into a film to obtain a combined video.
28. The electronic device of claim 15, wherein, The processor is configured to search for at least one preset view angle corresponding to the scene from a preset database, different scenes corresponding to different preset view angles. 29.A computer readable storage medium, the computer readable storage medium storing a computer program, the computer program comprising program instructions, the program instructions causing a processor to execute the panoramic video clip method according to any one of claims 1 to 14 when executed by the processor.
Citation Information
Patent Citations
Method and system for panoramic video file editing as well as portable terminal
CN107547939A
Automatic editing method of panoramic video, panoramic camera, computer program product and readable storage medium
CN114598810A
Motion video generation method and device, terminal equipment and storage medium
CN116862946A
Video generation method and device, computer equipment and storage medium
CN117082304A
System of automatic creation of a scenario video clip with a predefined object
US20210021912A1