Viewing angle path determination method and apparatus for panoramic video, panoramic camera, and medium
By constructing a directed acyclic graph and determining the target view path based on the candidate view sequence, the problem of low efficiency in panoramic video editing is solved, achieving efficient one-shot editing and improving the user experience.
Patent Information
- Application Number
- PCT/CN2024/083108
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2026-01-29
AI Technical Summary
Existing technologies struggle to effectively determine the optimal viewpoint path for panoramic videos, resulting in low efficiency in one-shot editing.
By constructing a directed acyclic graph, the target view path is determined based on the candidate view sequence. Path search is performed using node weights and edge weights to select at least two target view sequences and construct the target view path to achieve one-shot editing.
It improves the automation and efficiency of panoramic video editing, ensures that the length of the edited video is consistent with the original video, and enhances the user experience.
Smart Images

Figure CN2024083108_29012026_PF_FP_ABST
Abstract
Description
Methods, devices, panoramic cameras, and media for determining the viewpoint path in panoramic video. Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, panoramic camera, and medium for determining the viewpoint path of panoramic video. Background Technology
[0002] Panoramic video is video captured by a panoramic camera in a 360-degree all-around view, allowing users to view dynamic video within the camera's shooting angle range. However, since a flat-screen display can only show one perspective of the panoramic video at any given time, determining the optimal viewing path from the panoramic video and then using this path for one-shot editing to obtain a flat-screen video is a pressing problem in this field.
[0003] Summary of the Invention
[0004] Therefore, it is necessary to provide a method, device, panoramic camera, and medium for determining the viewpoint path of panoramic video that can be edited in one continuous shot, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for determining the viewpoint path in panoramic video, including:
[0006] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0007] Secondly, this application provides a method for editing panoramic videos, including:
[0008] A target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0009] The panoramic video is edited based on the target view path to obtain a target video containing at least two target view sequences, the duration of which is the same as that of the panoramic video.
[0010] Thirdly, this application also provides a viewpoint path determination device for panoramic video, comprising:
[0011] The determination module is used to determine the target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0012] Fourthly, this application also provides a panoramic video editing device, comprising:
[0013] The determination module is used to determine a target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0014] The editing module is used to edit the panoramic video based on the target view path to obtain a target video containing at least two target view sequences, wherein the duration of the target video is the same as the duration of the panoramic video.
[0015] Fifthly, this application also provides a panoramic camera, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0016] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0017] Sixthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0018] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0019] In a seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0020] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0021] The aforementioned panoramic video viewpoint path determination method, apparatus, panoramic camera, and medium determine the target viewpoint path based on multiple candidate viewpoint sequences corresponding to the panoramic video. The target viewpoint path consists of at least two target viewpoint sequences, which are selected from the candidate viewpoint sequences. In this embodiment, the target viewpoint path is determined based on multiple candidate viewpoint sequences, and the obtained target viewpoint path includes at least two target viewpoint sequences. The panoramic video is then edited based on the target viewpoint path to obtain a one-shot edited video, laying the foundation for subsequent editing of the panoramic video based on the target viewpoint path.
[0022] The aforementioned panoramic video editing method determines a target viewpoint path based on multiple candidate viewpoint sequences corresponding to the panoramic video, and then edits the panoramic video based on the target viewpoint path to obtain a target video containing at least two target viewpoint sequences. The target viewpoint path consists of at least two target viewpoint sequences, which are selected from the candidate viewpoint sequences. The duration of the target video is the same as that of the panoramic video. In this embodiment, the target viewpoint path is determined based on multiple candidate viewpoint sequences, resulting in a target viewpoint path that includes at least two target viewpoint sequences. The duration of the target video is the same as that of the panoramic video, meaning the target video is a video obtained by editing the panoramic video in one continuous shot, rather than editing a highlight reel, thus improving the automation and efficiency of editing one-shot videos. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 illustrates a prior art method for determining the viewpoint path of panoramic video in one embodiment;
[0025] Figure 2 shows the application environment of the panoramic video view path determination method in one embodiment;
[0026] Figure 3 is a flowchart illustrating a method for determining the viewpoint path of panoramic video in one embodiment.
[0027] Figure 4 is a schematic diagram of target viewpoint identification in one embodiment;
[0028] Figure 5 shows a video frame obtained based on the target viewpoint in one embodiment;
[0029] Figure 6 is a schematic diagram of target viewpoint identifiers and candidate viewpoint identifiers in one embodiment;
[0030] Figure 7 is a flowchart illustrating a target view path determination method in one embodiment;
[0031] Figure 8 is a flowchart illustrating a method for constructing a directed acyclic graph in one embodiment;
[0032] Figure 9 is a schematic diagram of the construction of the view sequence pool in one embodiment;
[0033] Figure 10 is a flowchart illustrating a method for constructing a directed acyclic graph in another embodiment;
[0034] Figure 11 is a schematic diagram of a node array in one embodiment;
[0035] Figure 12 is a flowchart illustrating a node array diagram determination method in one embodiment;
[0036] Figure 13 is a flowchart illustrating a method for constructing a directed acyclic graph in another embodiment;
[0037] Figure 14 is a flowchart illustrating the target view path determination method in another embodiment;
[0038] Figure 15 is a schematic diagram of the node array in another embodiment;
[0039] Figure 16 is a flowchart illustrating the target view path determination method in another embodiment;
[0040] Figure 17 is a flowchart illustrating the second-view path determination method in another embodiment;
[0041] Figure 18 is a flowchart illustrating a method for determining the field of view in one embodiment;
[0042] Figure 19 is a flowchart illustrating a panoramic video editing method in one embodiment;
[0043] Figure 20 is a structural block diagram of a panoramic video view path determination device in one embodiment;
[0044] Figure 21 is an internal structure diagram of a panoramic camera in one embodiment. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0046] With the continuous development of internet technology, users can view panoramic videos from different perspectives and locations, improving their viewing experience.
[0047] Currently, the optimal path for obtaining panoramic video is mainly achieved through the following three methods. The first method involves unfolding the panoramic video to obtain a spherical image, taking a point every 10° of latitude and every 20° of longitude on the spherical image, and taking a segment every 5 seconds in time. Then, a greedy search algorithm is used to search for the local optimal node at each segment result to obtain the optimal viewpoint path. However, the viewpoint path obtained by this method is of poor quality.
[0048] The second method involves converting the panoramic video into a hexahedron. At each time point, the projected images of each face of the hexahedron are acquired. The optimal viewing path for the panoramic video is then obtained using models such as Long Short-Term Memory (LSTM) networks and these projected images. As shown in Figure 1, the panoramic video is converted into a hexahedron. Using the projected images on the six faces at each time point as nodes, the optimal viewing path is learned through an LTM network model. The computational cost of Figure 1 is related to the model size. Taking the common ResNet-50 model as an example, its computational cost is 4.145 GB, and the inference time on an A100 GPU is 2.5 ms, requiring execution once per keyframe.
[0049] The third approach models the viewpoint motion of each frame in a panoramic video as a reinforcement learning task. The frame sequence represents the observations in the reinforcement learning process, the optimal viewpoint in each frame is the target, and the movement of the viewpoint is the action. The optimal viewpoint path is then learned. This method, like the second approach, obtains the optimal viewpoint path through a model. However, this method suffers from high computational cost and is tightly coupled with the method used to acquire the viewpoint target, making it unsuitable for viewpoint targets acquired using other methods.
[0050] Therefore, this application proposes a method, apparatus, panoramic camera, and medium for determining the viewpoint path of panoramic video to solve the above-mentioned technical problems.
[0051] The panoramic video viewpoint path determination method, apparatus, panoramic camera, and medium provided in this application embodiment can be applied to the application environment shown in Figure 2. This application environment includes a panoramic camera, which can be a server, and its internal structure is shown in Figure 2. The panoramic camera includes a processor, memory, input / output interface (I / O), and communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the panoramic camera provides computing and control capabilities. The memory of the panoramic camera includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the panoramic camera stores viewpoint path determination related data. The I / O interface of the panoramic camera is used for information exchange between the processor and external devices. The communication interface of the panoramic camera is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a panoramic video viewpoint path determination method, apparatus, panoramic camera, and medium. A server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0052] In an exemplary embodiment, as shown in FIG3, a method for determining the viewpoint path of panoramic video is provided. Taking the application of this method to the panoramic camera in FIG2 as an example, the method includes the following S301, wherein:
[0053] S301, determine the target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0054] Panoramic video is created by using multiple cameras to capture the environment in a 360° omnidirectional manner, resulting in multiple video streams. These streams are then combined using technologies such as synchronization, stitching, and projection. Users can choose any angle within the 360° range to view the video. A monitor can only display a portion of the panoramic video at a time, and the content of that portion differs depending on the viewing angle. The viewing angle can be understood as the window through which the panoramic image is observed; the viewing angle determines the scene displayed in the panoramic image.
[0055] A viewpoint sequence refers to a sequence of viewpoints from multiple panoramic video frames arranged chronologically. To distinguish them, all viewpoint sequences corresponding to a panoramic video are called candidate viewpoint sequences, and the last selected viewpoint sequence is called the target viewpoint sequence. Viewpoint sequences can be obtained by tracking each target in the panoramic video to obtain the viewpoint sequence corresponding to each target. For example, if target 1 and target 2 are detected in the first frame of the panoramic video, and target 1 and target 2 are tracked respectively, and target 1 does not appear in frame 10, then the viewpoint sequence corresponding to target 1 is the tracking sequence consisting of frames 1-9. If target 2 does not appear in frame 20, then the viewpoint sequence corresponding to target 2 is the tracking sequence consisting of frames 1-19.
[0056] In this embodiment, a directed acyclic graph (DAG) can be constructed for multiple candidate viewpoint sequences corresponding to the panoramic video, and edge weights and node weights can be set in the DAG. Path search is performed based on the edge weights and node weights in the DAG to obtain the target viewpoint path. The process of performing path search is the process of filtering the target viewpoint sequence.
[0057] In one possible implementation, at least two target view sequences that can constitute the target view path can be determined from multiple candidate view sequences as needed. For example, at least two targets can be pre-selected, then these two targets can be tracked, and the target view sequence can be determined based on the tracking results. For instance, given targets A, B, and C, tracking A yields a view sequence about A, tracking B yields a view sequence about B, and tracking C yields a view sequence about C. If the view sequences of A, B, and C overlap in time, the target view can be selected based on the level of detail within that overlapping segment. If the union of the times covered by the view sequences of A, B, and C does not cover all time periods, the missing time periods can be filled using a default view sequence, such as using the camera's forward-looking view sequence as the default view sequence. For example, suppose the total duration of the panoramic video is 10 minutes. The A-view sequence takes up 0-2 minutes, the B-view sequence takes up 2-5 minutes, and the C-view sequence takes up 5-8 minutes. The remaining 2 minutes can be filled in using the default view sequence within that time period. Then the target view path is composed of the A-view sequence, the B-view sequence, the C-view sequence, and the default view sequence.
[0058] In the aforementioned method for determining the viewpoint path of panoramic video, a target viewpoint path is determined based on multiple candidate viewpoint sequences corresponding to the panoramic video. The target viewpoint path consists of at least two target viewpoint sequences, which are selected from the candidate viewpoint sequences. In this embodiment, the target viewpoint path is determined based on multiple candidate viewpoint sequences, resulting in a target viewpoint path that includes at least two target viewpoint sequences. The panoramic video is then edited based on the target viewpoint path to obtain a one-shot edited video, laying the foundation for subsequent editing of the panoramic video based on the target viewpoint path.
[0059] In one embodiment, the above-mentioned panoramic video viewpoint path determination method further includes: displaying a target viewpoint identifier on the panoramic video, wherein the target viewpoint identifier is used to identify the location of the target viewpoint.
[0060] The target viewpoint identifier is used to identify the location of the target viewpoint. Optionally, the target viewpoint identifier can be a rectangular identifier, a circular identifier, or an identifier with the same shape as the corresponding target. As shown in the panoramic video in Figure 4, the rectangular identifier in Figure 4 is the target viewpoint identifier on the panoramic video. Editing based on the target viewpoint where the rectangular identifier is located can produce the video frames shown in Figure 5.
[0061] In this embodiment of the application, a target viewpoint identifier is displayed on the panoramic video. The target viewpoint identifier can identify the location of the target viewpoint. Users can view the corresponding position of the video clipped based on the target viewpoint path in the panoramic video, thereby improving the user experience.
[0062] In one embodiment, the above-mentioned panoramic video view path determination method further includes: displaying candidate view identifiers and target view identifiers on the panoramic video, wherein the candidate view identifiers are used to represent the corresponding candidate view sequence, and the target view identifiers are used to represent the selected target view sequence.
[0063] Optionally, to distinguish between candidate viewpoint markers and target viewpoint markers, they can be displayed using different colors, different shapes, or different thicknesses. As shown in Figure 6, the solid-line rectangle markers represent target viewpoint markers, while the dashed-line rectangle markers represent candidate viewpoint markers.
[0064] Optionally, the candidate viewpoint identifier may include one or more.
[0065] In this embodiment, candidate viewpoint identifiers and target viewpoint identifiers are displayed on the panoramic video. The candidate viewpoint identifiers represent the corresponding candidate viewpoint sequence, and the target viewpoint identifiers represent the selected target viewpoint sequence, so as to facilitate clear viewing of the viewpoints of each target in the panoramic video and to track each target.
[0066] In one embodiment, in response to a user's instruction to change the target viewpoint in the panoramic video, the newly selected candidate viewpoint sequence is used as the target viewpoint sequence, and the target viewpoint path is updated.
[0067] In this embodiment, during the display of candidate and target viewpoints in the panoramic video, if the user is not satisfied with the selected target viewpoint, they can reselect it based on their instructions. For example, the panoramic video includes target viewpoint a, candidate viewpoint b1, candidate viewpoint b2, candidate viewpoint b3, and candidate viewpoint b4. During the playback of the panoramic video, the user can reselect candidate viewpoint b2, and the viewpoint sequence corresponding to candidate viewpoint b2 will be used as the target viewpoint sequence.
[0068] Furthermore, updating the target view path includes: determining the start and end positions of the reselected target view sequence, updating the view sequence from the start to the end position in the target view path, switching from the start position to the reselected target view sequence, and after the end position, switching back to the original target view path.
[0069] In this embodiment, each viewpoint sequence has a start position and an end position. The start position of the reselected target viewpoint sequence is the position corresponding to when the user selects the candidate viewpoint. The position corresponding to when the user selects the candidate viewpoint can also be the same as the start position of the candidate viewpoint sequence. The end position of the reselected target viewpoint sequence is the end position corresponding to the candidate viewpoint sequence itself.
[0070] For example, a panoramic video includes target viewpoints a1 and a2. The start and end positions of target viewpoint a1 are from frames 0 to 50, and the start and end positions of target viewpoint a2 are from frames 51 to 100. Candidate viewpoints b1, b2, b3, and b4 have start and end positions from frames 10 to 20, 10 to 40, 50 to 70, and 60 to 80, respectively. If the user selects candidate viewpoint b2 as the target viewpoint at frame 30, then from frames 30 to 40, the viewpoint sequence corresponding to candidate viewpoint b2 will be used as the target viewpoint sequence. From frames 40 to 50, the viewpoint sequence will revert to that corresponding to target viewpoint a1, and from frames 51 to 100, the viewpoint sequence will be that corresponding to target viewpoint a2.
[0071] In this embodiment, the start and end positions of the newly selected target view sequence are determined, the view sequence from the start to the end position in the target view path is updated, the view is switched from the start position to the newly selected target view sequence, and after the end position, it switches back to the original target view path. In this embodiment, if the user is not satisfied with the automatically generated target view path, they can manually adjust the target view, improving the flexibility of target view path determination.
[0072] Figure 7 is a flowchart illustrating a target viewpoint path determination method in one embodiment. As shown in Figure 7, this embodiment relates to a possible implementation of how to determine a target viewpoint path based on multiple candidate viewpoint sequences corresponding to a panoramic video. That is, the above-mentioned S301 may include:
[0073] S701, construct a directed acyclic graph based on multiple candidate view sequences corresponding to panoramic video, and determine the node weights and edge weights in the directed acyclic graph.
[0074] In this embodiment, multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences. The multiple nodes are sorted according to the time order of the multiple candidate view sequences to obtain a node array graph. The multiple nodes in the node array graph are connected according to preset rules to obtain a directed acyclic graph.
[0075] In one possible implementation, multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences. The nodes are then sorted according to the time order of the multiple candidate view sequences to obtain a node array graph. The nodes in the node array graph are connected according to a preset rule. Furthermore, a virtual start node and a virtual end node are set. The virtual start node is connected to the node containing the first keyframe, and the virtual end node is connected to the node containing the last keyframe to obtain the directed acyclic graph.
[0076] Furthermore, node weights and edge weights in a directed acyclic graph can be automatically assigned according to preset rules, and node weights and edge weights in the directed acyclic graph can also be determined based on user interaction.
[0077] S702 performs path search based on the node weights and edge weights of the directed acyclic graph to obtain the target view path.
[0078] In this embodiment, the total weight of each view path can be determined by the node weights and edge weights of each node and the edges between nodes in the directed acyclic graph, and the target view path can be obtained based on the total weight of each view path.
[0079] In one possible implementation, the node weights and edge weights between nodes in the directed acyclic graph can be used to determine the first total weight based on the node weights in each viewpoint path and the second total weight based on the edge weights in each viewpoint path. The viewpoint path that simultaneously satisfies the first total weight and the second total weight is taken as the target viewpoint path.
[0080] In this embodiment, a directed acyclic graph (DAG) is constructed based on multiple candidate viewpoint sequences corresponding to the panoramic video. The node weights and edge weights in the DAG are determined, and a path search is performed based on these weights to obtain the target viewpoint path. This embodiment innovatively constructs a DAG based on candidate viewpoint sequences and then determines the target viewpoint path based on the obtained node and edge weights. This not only yields good results but also achieves the globally optimal viewing viewpoint path for the panoramic video with very low time and space complexity, significantly reducing computational load. Furthermore, this application can use any method to obtain the candidate viewpoint sequences to acquire the DAG, thereby determining the target viewpoint path based on the DAG, resulting in broad applicability and enhanced scalability.
[0081] In one embodiment, determining the node weights and edge weights in a directed acyclic graph includes: automatically assigning node weights and edge weights to the directed acyclic graph according to preset rules.
[0082] In this embodiment, preset rules can be customized as needed to automatically assign node and edge weights to the directed acyclic graph. For example, a preset rule could be that the larger the area occupied by a target in a video frame of a panoramic video, the greater the node and edge weights corresponding to that target. Alternatively, a preset rule could be that targets appearing more frequently in the panoramic video correspond to greater node and edge weights.
[0083] In this embodiment, node weights and edge weights are automatically assigned to the directed acyclic graph (DAG) according to preset rules. The node weights and edge weights in the DAG can be determined according to preset rules; that is, users can customize the node weights and edge weights of the DAG as needed. This allows for the search and acquisition of different perspective paths based on different requirements, making it more widely applicable and highly scalable.
[0084] In one embodiment, determining the node weights and edge weights in a directed acyclic graph further includes: determining the node weights and edge weights in the directed acyclic graph based on user interactions.
[0085] In this embodiment, user input commands can be received, and in response to the user's operation commands, different weights can be assigned to the nodes and edges in the directed acyclic graph. That is, the user can directly input different weights for the nodes and edges in the directed acyclic graph.
[0086] In one possible implementation, the panoramic video can be analyzed to output the initial node weights and initial edge weights of each node and edge, providing the user with corresponding operation options and instructions. Based on the user's operation instructions, the node weights and edge weights are determined according to the initial node weights and initial edge weights.
[0087] In this embodiment, the node weights and edge weights in the directed acyclic graph are determined based on the interaction with the user. The user can customize the determination of node weights and edge weights, thereby improving the flexibility of determining node weights and edge weights.
[0088] In one embodiment, determining the node weights and edge weights in a directed acyclic graph based on user interaction includes: obtaining user interest points based on user interaction, and determining the node weights and edge weights in the directed acyclic graph based on user interest points.
[0089] In this embodiment, the user's interest points can be obtained based on the user's selected interest area, and the node weights and edge weights of the directed acyclic graph can be determined based on the user's interest points.
[0090] In another possible implementation, multiple candidate points of interest can be displayed, the user's interest in the multiple candidate points of interest can be obtained, and the node weights and edge weights of the directed acyclic graph can be determined based on the interest in each candidate point of interest.
[0091] Furthermore, based on user interaction, the user's points of interest are obtained, including: obtaining the region of interest selected by the user in the video frame image, and determining the user's points of interest based on the region of interest.
[0092] In this embodiment, the user selects a region of interest in a video frame image, and any point within that region is taken as the user's point of interest. Alternatively, the user selects a region of interest in a video frame image, and the center point of that region is taken as the user's point of interest.
[0093] In this embodiment, user interest points are obtained based on user interaction, and node weights and edge weights of the directed acyclic graph are determined based on user interest points, laying the foundation for subsequent determination of the target view path based on node weights and edge weights.
[0094] In one embodiment, determining the node weights and edge weights of a directed acyclic graph based on the user's points of interest includes: displaying multiple candidate points of interest, obtaining the user's interest level in the multiple candidate points of interest, and determining the node weights and edge weights of the directed acyclic graph based on the interest level of each candidate point of interest.
[0095] In this embodiment, multiple candidate points of interest are displayed. The interest level of each candidate point can be determined based on the historical click frequency of the target corresponding to that candidate point, or the user can preset the interest level for each candidate point. The higher the interest level, the greater the node weight and edge weight of the directed acyclic graph.
[0096] In this embodiment, multiple candidate points of interest are displayed, and the user's interest level among these candidate points is obtained. The node weights and edge weights of the directed acyclic graph (DAG) are then determined based on the interest level of each candidate point. This embodiment improves the flexibility of determining node and edge weights by basing the determination on the interest level of the candidate points.
[0097] In one embodiment, determining the node weights and edge weights in a directed acyclic graph based on user interaction includes: displaying the initial node weights and initial edge weights in the directed acyclic graph; receiving user operation instructions based on the initial node weights and initial edge weights; and, in response to the user operation instructions, determining the node weights and edge weights in the directed acyclic graph according to the initial node weights and initial edge weights.
[0098] In this embodiment, the target that needs to be focused on can be determined based on manual marking by the user, target tracking algorithm, or user-specified method. The panoramic camera assigns higher weights to the target that needs to be focused on and randomly assigns weights to other targets to obtain the corresponding initial node weights and initial edge weights, and outputs each initial node weight and initial edge weight.
[0099] At the same time, operation buttons are displayed on the panoramic camera's display interface. For each initial node weight and initial edge weight, the user can choose to directly use the "Keep" button to use the initial node weight and initial edge weight as the node weight and edge weight, or use the "Adjust" button to increase or decrease the initial node weight and initial edge weight to obtain the node weight and edge weight.
[0100] In this embodiment, the initial node weights and initial edge weights of the directed acyclic graph are displayed; user operation instructions based on the initial node weights and initial edge weights are received; in response to the user operation instructions, the node weights and edge weights of the directed acyclic graph are determined according to the initial node weights and initial edge weights. In this embodiment, the initial node weights and initial edge weights are adjusted based on the interaction with the user to obtain the node weights and edge weights of the directed acyclic graph, making the determination of node weights and edge weights more accurate, thereby making the target view path obtained based on node weights and edge weights more accurate.
[0101] Figure 8 illustrates a method for constructing a directed acyclic graph (DAG) in one embodiment. This application relates to how to construct a DAG based on multiple candidate view sequences corresponding to a panoramic video, as shown in Figure 8, which may include:
[0102] S801, construct a view sequence pool; the view sequence pool stores multiple candidate view sequences corresponding to panoramic videos.
[0103] Candidate view sequence can be determined in a variety of ways. In one embodiment, each target in the panoramic video is detected and tracked, and the view sequence of each tracked target is added to the view sequence pool.
[0104] As shown in Figure 9, in another possible implementation, the sphere of the panoramic video can be divided into sections with points every 10° of latitude and every 20° of longitude. Each section is then divided into blocks every 5 seconds. Candidate view sequences for each region are obtained from the starting block to the ending block, and these sequences are added to a view sequence pool. The candidate view sequences can also be determined in other ways; there are no limitations on how they are obtained here. For example, they can be randomly selected candidate view sequences.
[0105] S802, constructs a directed acyclic graph based on multiple candidate view sequences corresponding to panoramic video.
[0106] In this embodiment, multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences. These nodes are then sorted according to the chronological order of the candidate view sequences to obtain a node array graph. Finally, the nodes in the node array graph are connected according to a preset rule to obtain the directed acyclic graph. In one embodiment, each view sequence is treated as a node, and then the corresponding multiple nodes are sorted according to the chronological order of the candidate view sequences.
[0107] In one possible implementation, multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences. The nodes are then sorted according to the time order of the multiple candidate view sequences to obtain a node array graph. The nodes in the node array graph are connected according to a preset rule. Furthermore, a virtual start node and a virtual end node are set. The virtual start node is connected to the node containing the first keyframe, and the virtual end node is connected to the node containing the last keyframe to obtain the directed acyclic graph.
[0108] In this embodiment, a viewpoint sequence pool is constructed, and a directed acyclic graph (DAG) is built based on multiple candidate viewpoint sequences corresponding to the panoramic video. The viewpoint sequence pool stores multiple candidate viewpoint sequences corresponding to the panoramic video. In this embodiment, the DAG is constructed based on multiple candidate viewpoint sequences in the viewpoint sequence pool, which has low computational cost and can obtain the globally optimal viewing viewpoint path of the panoramic video with very low time and space complexity, laying the foundation for subsequent determination of the target viewpoint path based on the DAG.
[0109] In one embodiment, constructing a view sequence pool includes: determining the view sequence of each target in the panoramic video, and adding the view sequence of each target to the view sequence pool as a candidate view sequence.
[0110] Optionally, the targets in the panoramic video can be people, animals, buildings, trees, houses, sky, landscapes, etc.
[0111] In this embodiment, a starting frame can be determined from the panoramic video. A first target in the starting frame is detected to obtain a first target detection box. The first target detection box is tracked until the tracking signal is disconnected, resulting in a tracking sequence for the first target. This tracking sequence is added to the viewpoint sequence pool as a candidate viewpoint sequence. Next, it is determined whether a second target different from the first target appears in the next frame. If a second target appears, it is detected to obtain a second target detection box. This second target detection box is tracked to obtain a tracking sequence for the second target. This tracking sequence is added to the viewpoint sequence pool as a candidate viewpoint sequence. This process is repeated for each frame in the panoramic video until all targets in the panoramic video have been detected, resulting in viewpoint sequences for each target. These viewpoint sequences are then added to the viewpoint sequence pool to obtain the corresponding candidate viewpoint sequences.
[0112] In one embodiment, determining the viewpoint sequence of each target in a panoramic video includes: performing target detection on keyframes in the panoramic video, determining target detection boxes, tracking the target detection boxes to obtain a tracking sequence, and determining the viewpoint sequence of each target based on the tracking sequence of the target detection boxes.
[0113] In this embodiment, starting from a certain frame of the video as the starting frame, and then extracting one frame every N frames as a keyframe. At the keyframe, a multi-class detector is used to detect and obtain the target detection boxes of each target, and a multi-target tracking algorithm is used to track all detected target detection boxes until the tracking algorithm returns a disconnection signal, thus obtaining a tracking sequence. The tracking sequence of the target (i.e., the view sequence) is added to the view sequence pool.
[0114] In this embodiment, target detection is performed on keyframes in the panoramic video to determine target detection boxes. The target detection boxes are then tracked to obtain a tracking sequence. Based on the tracking sequence of the target detection boxes, the viewpoint sequence of each target is determined. This allows for the rapid acquisition of the viewpoint sequence of each target, thus accelerating the efficiency of determining the target viewpoint path.
[0115] Figure 10 is a flowchart illustrating a directed acyclic graph (DAG) construction method in another embodiment. As shown in Figure 10, this embodiment relates to a possible implementation of how to construct a DAG based on multiple view sequences corresponding to panoramic video, including the following steps:
[0116] S1001, determine multiple nodes in the directed acyclic graph based on multiple candidate view sequences.
[0117] In this embodiment, each candidate view sequence can be treated as a node in a directed acyclic graph, or each candidate view sequence can be divided into sub-view sequences at preset time intervals, and each sub-view sequence can be treated as a node.
[0118] S1002, based on the time order of each candidate view sequence, arrange the corresponding multiple nodes to obtain a node array diagram.
[0119] In this embodiment, multiple nodes are sorted according to the time order of each candidate view sequence to obtain a node array diagram. As shown in Figure 11, the view sequences corresponding to nodes 1, 4, 6, 10, and 12 have the same start time and are the earliest start time. The view sequences corresponding to nodes 2 and 7 have the same time and are later than the times of nodes 1, 4, 6, 10, and 12, and so on, to obtain the node array diagram shown in Figure 7.
[0120] S1003, Connect the nodes in the node array graph according to preset rules to obtain a directed acyclic graph.
[0121] Optionally, the preset rules can be that the time interval between each node meets the preset conditions, the spatial distance between each node meets the preset conditions, or both the time interval and the spatial distance between each node meet the preset conditions.
[0122] In this embodiment, nodes in the node array graph are connected according to their order and based on preset rules to obtain a directed acyclic graph.
[0123] In this embodiment, multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences. The nodes are then arranged according to the time order of each candidate view sequence to obtain a node array graph. The nodes in the node array graph are then connected according to preset rules to obtain a directed acyclic graph, which lays the foundation for determining the target view path based on the directed acyclic graph.
[0124] Figure 12 is a flowchart illustrating a node array diagram determination method in one embodiment. As shown in Figure 12, this embodiment relates to a possible implementation of how to arrange multiple nodes based on the temporal order of each candidate view sequence to obtain a node array diagram, including the following steps:
[0125] S1201, each candidate view sequence is segmented at preset intervals to obtain sub-view sequences.
[0126] The preset duration can be set freely. Optionally, the preset duration can be 1 second, 2 seconds, etc.
[0127] In this embodiment, each candidate viewpoint sequence T is segmented at preset time intervals of duration K to obtain T / K sub-viewpoint sequences. For example, if the preset duration is 2 seconds and one candidate viewpoint sequence has a duration of 10 seconds, segmenting the candidate viewpoint sequence according to K = 2 seconds can result in 5 sub-viewpoint sequences.
[0128] In one possible implementation, different preset durations can be set for different candidate view sequences. For example, there are a total of 10 candidate view sequences, of which 5 candidate view sequences have a preset duration of 2 seconds, 2 candidate view sequences have a preset duration of 1 second, and 3 candidate view sequences have a preset duration of 3 seconds.
[0129] S1202, treat each sub-view sequence as a node to obtain multiple nodes in the directed acyclic graph.
[0130] In this embodiment, each sub-view sequence is directly treated as a node to obtain multiple nodes in the directed acyclic graph.
[0131] S1203, based on the time order of each sub-view sequence, arrange the corresponding multiple nodes to obtain a node array diagram.
[0132] In this embodiment, the time corresponding to the starting frame in each sub-view sequence is determined, and the corresponding multiple nodes are sorted according to the time order of the starting frames in the sub-view sequence to obtain a node array diagram.
[0133] In this embodiment, each candidate view sequence is segmented at preset time intervals to obtain sub-view sequences. Each sub-view sequence is treated as a node, resulting in multiple nodes in a directed acyclic graph. These nodes are then arranged according to the temporal order of the sub-view sequences to obtain a node array graph. This embodiment further subdivides each candidate view sequence into sub-view sequences, and obtains a node array graph based on these sub-view sequences. This results in a more precise division of the node array graph, leading to a more accurate final path obtained from it.
[0134] Figure 13 is a flowchart illustrating a method for constructing a directed acyclic graph in another embodiment. As shown in Figure 13, this embodiment relates to a possible implementation of how to connect nodes in a node array graph based on preset rules to obtain a directed acyclic graph, including the following steps:
[0135] S1301, determine the time interval and spatial distance between every two nodes.
[0136] In this embodiment, the time interval and spatial distance between any two nodes can be determined based on the last video frame of the previous node and the first video frame of the next node.
[0137] S1302, if the time interval and spatial distance meet the preset conditions, then connect the corresponding two nodes to obtain a directed acyclic graph.
[0138] Optionally, the preset conditions can be a preset time interval or a preset time range; the spatial distance can be a preset spatial distance or a preset spatial range.
[0139] In this embodiment, if the time interval and the spatial distance are both within a preset range, then the two corresponding nodes are connected; or, if the time interval is within a preset range and the spatial distance is less than a preset spatial distance, then the two corresponding nodes are connected, resulting in a directed acyclic graph. The edge connecting the two corresponding nodes is a directed edge, which can only point from the end frame of the earlier node in the time sequence to the start frame of the later node in the time sequence.
[0140] Alternatively, other combinations may be used, which will not be elaborated in the embodiments of this application.
[0141] Furthermore, if the time interval and spatial distance meet the preset conditions, the corresponding two nodes will be connected, including: if the time interval is less than the preset time interval and the spatial distance is less than the preset spatial distance, the corresponding two nodes will be connected.
[0142] Optionally, the preset time interval and preset spatial distance can be directly set values, or values determined based on the time interval and spatial distance. For example, the variance of the time interval can be determined based on each time interval, and the variance can be used as the preset time interval.
[0143] In this embodiment, if the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance, then the two corresponding nodes are connected. As shown in Figure 11 above, if the time interval between node 1 and node 7 is less than a preset time interval and the spatial distance is less than a preset time interval, then node 1 and node 7 are connected.
[0144] In this embodiment, by determining the time interval and spatial distance between any two nodes, if the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance, the corresponding two nodes are connected to obtain a directed acyclic graph. Connecting nodes by time interval and spatial distance can avoid excessively long camera movement time when connecting two nodes, and also avoid camera movements with excessively large spatial distances.
[0145] In one embodiment, the method for constructing a directed acyclic graph further includes: determining a virtual start node and a virtual end node; connecting the node corresponding to the first keyframe to the virtual start node; and connecting the node corresponding to the last keyframe to the virtual end node to obtain a directed acyclic graph.
[0146] In this embodiment, a virtual start node and a virtual end node are determined. The virtual start node is placed before all other nodes, and the virtual end node is placed after all other nodes. The virtual start node is connected to the node including the first keyframe, and the virtual end node is connected to the node including the last keyframe. The remaining nodes, except for the virtual start node and the virtual end node, are connected according to the aforementioned preset rules to obtain a directed acyclic graph.
[0147] In this embodiment, a virtual start node and a virtual end node are determined. The node corresponding to the first keyframe is connected to the virtual start node, and the node corresponding to the last keyframe is connected to the virtual end node to obtain a directed acyclic graph. This ensures that the obtained target view path includes both the first and last keyframes, improving the accuracy of the target view path.
[0148] Figure 14 is a flowchart illustrating the target view path determination method in another embodiment. As shown in Figure 14, this embodiment relates to a possible implementation of how to perform path search based on the node weights and edge weights of a directed acyclic graph to obtain the target view path, including the following steps:
[0149] S1401, determine the total weight of each view path in the directed acyclic graph based on the node weights and edge weights of the directed acyclic graph.
[0150] The node weights can be determined based on the type of target being focused on, or they can be set to default values. For example, if the node weights are all set to 1 by default, or if the target's perspective path primarily focuses on the dancing target, then the weights of nodes containing the dancing target are set to be greater than the weights of other nodes.
[0151] Edge weights can be determined by the inverse ratio of the view distance between two nodes, 1 / d, where d is the distance between the view distance of the previous node's ending frame and the view distance of the next node's starting frame. Furthermore, since the virtual start node and virtual end node are virtual nodes, the weights of edges connected to them are set to 0.
[0152] In this embodiment, a dynamic programming algorithm is used to perform a global optimal path search in the constructed directed acyclic graph to obtain the total weight of the paths from each perspective. For example, when using the dynamic programming algorithm, the search proceeds from the bottom up, that is, it starts from the virtual termination node and calculates backwards, recording the total weight of the paths from each perspective.
[0153] S1402, determine the target view path based on the total weight of each view path.
[0154] In this embodiment, as shown in Figure 15, the view path with the largest total weight among all view paths can be directly used as the initial view path. Interpolation processing is then performed on the two adjacent nodes in the initial view path to obtain the target view path.
[0155] In one possible implementation, the view path with the largest total weight among all view paths can be used as the initial view path. Interpolation is then performed on adjacent nodes in the initial view path to obtain the second view path. The field of view angle corresponding to each video frame in the second view path is determined, and the target view path is determined based on the field of view angle corresponding to each video frame.
[0156] In this embodiment, the total weight of each viewpoint path on the directed acyclic graph (DAG) is determined based on the node weights and edge weights of the DAG, and the target viewpoint path is determined based on the total weight of each viewpoint path. This application allows for customizable node and edge weights to meet different needs, searches for different total weights of viewpoint paths, and determines the target viewpoint path based on the total weight of each viewpoint path, thus improving the personalization of the target viewpoint path.
[0157] Figure 16 is a flowchart illustrating the target view path determination method in another embodiment. As shown in Figure 16, this embodiment relates to a possible implementation of how to determine the target view path based on the total weight of each view path, including the following steps:
[0158] S1601, the view path corresponding to the largest total weight is used as the initial view path.
[0159] In this embodiment, the initial view path is the path after removing the virtual start node, virtual end node, edge connected to the virtual start node, and edge connected to the end node from the view path corresponding to the largest total weight. If multiple view paths with the largest total weight are included, then the view path corresponding to any one of the largest total weights is taken as the initial view path.
[0160] In one possible implementation, if a node is determined by segmentation through a view sequence, then after removing the virtual start node, virtual end node, edge connected to the virtual start node, and edge connected to the end node from the view path corresponding to the largest total weight, the segmented nodes are merged into the initial view path.
[0161] S1602, interpolate the two adjacent nodes in the initial view path to obtain the second view path.
[0162] In this embodiment, the end frame view of the preceding node and the start frame view of the following node in two adjacent nodes can be determined. An interpolation algorithm is used to interpolate the end frame view and the start frame view to obtain the view path between the end frame and the start frame. Based on the view path between the end frame and the start frame and each node, a second view path is obtained.
[0163] In one possible implementation, the keyframe view inside the node can also be obtained, the keyframe view inside the node can be interpolated to obtain the view path inside the node, and the second view path can be obtained based on the view path inside the node and the view path between the end frame and the start frame.
[0164] S1603, determine the field of view angle corresponding to each video frame in the second-view path.
[0165] In this embodiment, the field of view angle corresponding to each key frame in the second view path is determined, and the field of view angles corresponding to two adjacent key frames are interpolated to obtain the field of view angle of each video frame between two key frames; the field of view angles corresponding to each video frame in the second view path include the field of view angle corresponding to each key frame and the field of view angle of each video frame between two key frames.
[0166] S1604 determines the target view path based on the field of view corresponding to each video frame.
[0167] In this embodiment, the path obtained by interpolating the field of view angles corresponding to two adjacent keyframes in the second view path is used as the target view path.
[0168] In this embodiment, the viewpoint path corresponding to the largest total weight is used as the initial viewpoint path. Interpolation is performed on adjacent nodes in the initial viewpoint path to obtain the second viewpoint path. The field of view angle corresponding to each video frame in the second viewpoint path is determined, and the target viewpoint path is determined based on the field of view angle corresponding to each video frame. In this embodiment, the target viewpoint path is obtained by combining the field of view angle corresponding to each video frame in the second viewpoint path with the initial viewpoint path determined by the total weight of each viewpoint path, making the determination of the target viewpoint path more comprehensive.
[0169] Figure 17 is a flowchart illustrating a second-view path determination method in another embodiment. As shown in Figure 17, this embodiment relates to a possible implementation of how to interpolate two adjacent nodes in the initial view path to obtain a second-view path, including the following steps:
[0170] S1701, determine the ending frame view of the preceding node and the starting frame view of the following node in two adjacent nodes.
[0171] In this embodiment, the viewpoints of the ending frame and the beginning frame are obtained based on the target detection / tracking bounding boxes on the ending frame of the preceding node and the beginning frame of the following node. For example, the center coordinates of the tracking bounding box on the ending frame of the preceding node are used as the viewpoint of the ending frame, and the center coordinates of the target detection bounding box on the beginning frame of the following node are used as the viewpoint of the beginning frame.
[0172] S1702, interpolate the viewpoint of the end frame of the previous node and the viewpoint of the start frame of the next node to obtain the second viewpoint path.
[0173] In this embodiment, the edge between two adjacent nodes is considered as a visual motion path moving from the end frame view of the previous node to the start frame view of the next node. Therefore, by interpolating the end frame view of the previous node and the start frame view of the next node, the visual paths of the previous and next nodes can be obtained. Based on the visual paths of the previous and next nodes and each node, a second visual path is obtained. For example, the interpolation algorithm can employ nearest neighbor interpolation, spatial autocovariance optimal interpolation, etc.
[0174] In this embodiment, the ending frame view of the preceding node and the starting frame view of the following node are determined among two adjacent nodes. Interpolation processing is then performed on the ending frame view of the preceding node and the starting frame view of the following node to obtain the second view path. This application uses interpolation processing on the ending frame view and the starting frame view to obtain the second view path, which involves low computational complexity, a simple algorithm, and fast second view path determination.
[0175] Figure 18 is a flowchart illustrating a method for determining the field of view angle in one embodiment. As shown in Figure 18, this embodiment relates to a possible implementation of how to determine the field of view angle corresponding to each video frame in the second view path, including the following steps:
[0176] S1801, determine the field of view angle corresponding to each keyframe in the second-view path.
[0177] In this embodiment, the target bounding box in each keyframe of the second-view path is determined, and the field of view angle of each keyframe in the second-view path is determined based on the width of the target bounding box and the width of the keyframe. The target bounding box can be a target detection bounding box or a tracking bounding box obtained by tracking the target detection bounding box.
[0178] In one possible implementation, if the frame identifiers of each keyframe in the second-view path are discontinuous, a preset field of view angle is set for the keyframes between keyframes with discontinuous frame identifiers. That is, if there is no field of view angle between keyframes with discontinuous frame identifiers, a fixed field of view angle can be set. For example, the ending frame of node a is a keyframe a1, and the starting frame of node b is a keyframe b1. Nodes a and b are two connected nodes in the second-view path. However, there are keyframes c1 and c2 between keyframes a1 and b1. In this case, there is no target box for keyframes c1 and c2, and a fixed field of view angle can be set for keyframes c1 and c2.
[0179] S1802, interpolation is performed based on the field of view angles of two adjacent keyframes to obtain the field of view angle of each video frame between two adjacent keyframes.
[0180] In this embodiment, an interpolation algorithm is used to interpolate the field of view of two adjacent keyframes to obtain the field of view of each video frame between the two adjacent keyframes. The field of view of each video frame in the second view path includes the field of view of two adjacent keyframes and the field of view of each video frame between the two adjacent keyframes.
[0181] Furthermore, determining the field of view angle corresponding to each keyframe in the second-view path includes the following steps: determining the target box in each keyframe; and determining the field of view angle of the corresponding keyframe based on the width of the target box and the width of the keyframe.
[0182] In this embodiment, the target bounding box in the keyframe can be a target detection bounding box or a tracking bounding box obtained by tracking the target detection bounding box. The ratio of the width of the target bounding box to the width of the keyframe is determined, and this ratio is multiplied by a preset multiple of pi to obtain the field of view (FOV) of the corresponding keyframe. For example, the FOV of the keyframe is obtained using the following formula:
[0183] Where W1 is the width of the target bounding box and W2 is the width of the keyframe.
[0184] In this embodiment, a target bounding box is determined in each keyframe. Based on the width of the target bounding box and the width of the keyframe, the field of view of the corresponding keyframe is determined. Interpolation is performed based on the field of view of two adjacent keyframes to obtain the field of view of each video frame between two adjacent keyframes. This embodiment uses the width of the target bounding box and the width of the keyframe to determine the field of view of the keyframe, thereby obtaining the field of view of each video frame in the second-view path. This simple implementation makes the determination of the optimal path more efficient.
[0185] In one embodiment, a method for editing panoramic video is also provided, as shown in Figure 16, including the following steps:
[0186] S1901, determine the target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0187] For details on the implementation, please refer to S301 above.
[0188] S1902, the panoramic video is edited based on the target view path to obtain a target video containing at least two target view sequences, and the duration of the target video is the same as that of the panoramic video.
[0189] The target video has the same duration as the panoramic video, unlike related technologies that extract highlights from panoramic videos.
[0190] In this embodiment, a network model and target viewpoint paths can be used to edit the panoramic video to obtain a target video containing at least two target viewpoint sequences. Alternatively, a segmentation algorithm and target viewpoint paths can be used to determine at least two target viewpoint sequences from the panoramic video, and these at least two target viewpoint sequences can be stitched together to obtain the target video.
[0191] In this embodiment, a target viewpoint path is determined based on multiple candidate viewpoint sequences corresponding to the panoramic video. The panoramic video is then edited based on the target viewpoint path to obtain a target video containing at least two target viewpoint sequences. In this embodiment, the target viewpoint path is determined based on multiple candidate viewpoint sequences, and the obtained target viewpoint path includes at least two target viewpoint sequences. The duration of the target video is the same as the duration of the panoramic video; that is, the target video is a video obtained by editing the panoramic video in one continuous shot, rather than editing a highlight reel, thus improving the automation and efficiency of editing one-shot videos.
[0192] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0193] Based on the same inventive concept, this application also provides a panoramic video viewpoint path determination and editing device for implementing the panoramic video viewpoint path determination method, apparatus, panoramic camera and medium involved above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more viewpoint path determination apparatus embodiments provided below can be found in the limitations of the panoramic video viewpoint path determination method, apparatus, panoramic camera and medium above, and will not be repeated here.
[0194] In an exemplary embodiment, as shown in FIG20, a panoramic video viewpoint path determination device is provided, including: a determination module 11, wherein:
[0195] The determination module 11 is used to determine the target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0196] In one embodiment, the panoramic video view path determination device further includes:
[0197] The first display module is used to display target viewpoint markers on the panoramic video. The target viewpoint markers are used to indicate the location of the target viewpoint.
[0198] In one embodiment, the panoramic video view path determination device further includes:
[0199] The second display module is used to display candidate viewpoint identifiers and target viewpoint identifiers on the panoramic video. The candidate viewpoint identifiers represent the corresponding candidate viewpoint sequences, and the target viewpoint identifiers represent the selected target viewpoint sequences.
[0200] In one embodiment, the panoramic video view path determination device further includes:
[0201] The update module is used to respond to user commands to change the target viewpoint in the panoramic video by using the newly selected candidate viewpoint sequence as the target viewpoint sequence and updating the target viewpoint path.
[0202] In one embodiment, the update module includes:
[0203] The first determining unit is used to determine the start and end positions of the reselected target view sequence, update the view sequence from the start position to the end position in the target view path, switch from the start position to the reselected target view sequence, and switch back to the original target view path after the end position.
[0204] In one embodiment, determining the module includes:
[0205] The second determining unit is used to construct a directed acyclic graph based on multiple candidate view sequences corresponding to the panoramic video, and to determine the node weights and edge weights in the directed acyclic graph.
[0206] The third determining unit is used to perform path search based on the node weights and edge weights of the directed acyclic graph to obtain the target view path.
[0207] In one embodiment, the second determining unit is further configured to automatically assign node weights and edge weights to the directed acyclic graph according to preset rules.
[0208] In one embodiment, the second determining unit is further configured to determine the node weights and edge weights in the directed acyclic graph based on the interaction with the user.
[0209] In one embodiment, the second determining unit is further configured to obtain the user's points of interest based on the user's interaction, and determine the node weights and edge weights of the directed acyclic graph based on the user's points of interest.
[0210] In one embodiment, the second determining unit is further configured to acquire the region of interest selected by the user in the video frame image, and determine the user's point of interest based on the region of interest.
[0211] In one embodiment, the second determining unit is further configured to display multiple candidate points of interest and obtain the user's interest level in the multiple candidate points of interest;
[0212] The node weights and edge weights of the directed acyclic graph are determined based on the interest levels of each candidate interest point.
[0213] In one embodiment, the first determining unit is further configured to display the initial node weights and initial edge weights in the directed acyclic graph; receive operation instructions from the user based on the initial node weights and initial edge weights; and, in response to the user's operation instructions, determine the node weights and edge weights in the directed acyclic graph according to the initial node weights and initial edge weights.
[0214] In one embodiment, the determining module includes:
[0215] The first building unit is used to build a view sequence pool; the view sequence pool stores multiple candidate view sequences corresponding to the panoramic video.
[0216] The second building unit is used to construct a directed acyclic graph based on multiple candidate view sequences corresponding to the panoramic video.
[0217] In one embodiment, the first construction unit is further configured to determine the view sequence of each target in the panoramic video, and add the view sequence of each target to the view sequence pool as a candidate view sequence.
[0218] In one embodiment, the first building unit is further configured to perform target detection on keyframes in the panoramic video, determine target detection boxes, track the target detection boxes to obtain a tracking sequence, and determine the view sequence of each target based on the tracking sequence of the target detection boxes.
[0219] In one embodiment, the second construction unit is further configured to determine multiple nodes in a directed acyclic graph based on multiple candidate view sequences; arrange the corresponding multiple nodes according to the time order of each candidate view sequence to obtain a node array graph; and connect each node in the node array graph according to a preset rule to obtain a directed acyclic graph.
[0220] In one embodiment, the second building unit is further configured to divide each candidate view sequence into sub-view sequences at preset time intervals; treat each sub-view sequence as a node to obtain multiple nodes in a directed acyclic graph; and arrange the corresponding multiple nodes according to the time order of each sub-view sequence to obtain a node array graph.
[0221] In one embodiment, the second building unit is further configured to determine the time interval and spatial distance between every two nodes; if the time interval and spatial distance meet the preset conditions, the corresponding two nodes are connected to obtain a directed acyclic graph.
[0222] In one embodiment, the second building unit is further configured to connect the two corresponding nodes if the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance.
[0223] In one embodiment, the second building unit is further configured to determine a virtual start node and a virtual end node;
[0224] Connect the node corresponding to the first keyframe to the virtual start node; connect the node corresponding to the last keyframe to the virtual end node to obtain a directed acyclic graph.
[0225] In one embodiment, the third determining unit is further configured to determine the total weight of each view path on the directed acyclic graph based on the node weights and edge weights of the directed acyclic graph; and to determine the target view path based on the total weight of each view path.
[0226] In one embodiment, the third determining unit is further configured to take the view path corresponding to the largest total weight as the initial view path; perform interpolation processing on two adjacent nodes in the initial view path to obtain the second view path; determine the field of view angle corresponding to each video frame in the second view path; and determine the target view path based on the field of view angle corresponding to each video frame.
[0227] In one embodiment, the third determining unit is further configured to determine the ending frame view of the preceding node and the starting frame view of the following node among two adjacent nodes; and to perform interpolation processing on the ending frame view of the preceding node and the starting frame view of the following node to obtain the second view path.
[0228] In one embodiment, the third determining unit is further configured to determine the field of view angle corresponding to each key frame in the second view path; and to interpolate based on the field of view angles of two adjacent key frames to obtain the field of view angle of each video frame between two adjacent key frames.
[0229] In one embodiment, the third determining unit is further configured to determine the target bounding box in each keyframe; and to determine the field of view of the corresponding keyframe based on the width of the target bounding box and the width of the keyframe.
[0230] In one embodiment, a panoramic video editing device is also provided, comprising:
[0231] The determination module is used to determine the target view path based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0232] The editing module is used to edit the panoramic video based on the target view path to obtain a target video containing at least two target view sequences. The duration of the target video is the same as that of the panoramic video.
[0233] In one embodiment, the panoramic video editing device further includes:
[0234] The first display module is used to display target viewpoint markers on the panoramic video. The target viewpoint markers are used to indicate the location of the target viewpoint.
[0235] In one embodiment, the panoramic video editing device further includes:
[0236] The second display module is used to display candidate viewpoint identifiers and target viewpoint identifiers on the panoramic video. The candidate viewpoint identifiers represent the corresponding candidate viewpoint sequences, and the target viewpoint identifiers represent the selected target viewpoint sequences.
[0237] In one embodiment, the panoramic video editing device further includes:
[0238] The update module is used to respond to user commands to change the target viewpoint in the panoramic video by using the newly selected candidate viewpoint sequence as the target viewpoint sequence and updating the target viewpoint path.
[0239] In one embodiment, the update module includes:
[0240] The first determining unit is used to determine the start and end positions of the reselected target view sequence, update the view sequence from the start position to the end position in the target view path, switch from the start position to the reselected target view sequence, and switch back to the original target view path after the end position.
[0241] In one embodiment, determining the module includes:
[0242] The second determining unit is used to construct a directed acyclic graph based on multiple candidate view sequences corresponding to the panoramic video, and to determine the node weights and edge weights in the directed acyclic graph.
[0243] The third determining unit is used to perform path search based on the node weights and edge weights of the directed acyclic graph to obtain the target view path.
[0244] In one embodiment, the second determining unit is further configured to automatically assign node weights and edge weights to the directed acyclic graph according to preset rules.
[0245] In one embodiment, the second determining unit is further configured to determine the node weights and edge weights in the directed acyclic graph based on the interaction with the user.
[0246] In one embodiment, the second determining unit is further configured to obtain the user's points of interest based on the user's interaction, and determine the node weights and edge weights of the directed acyclic graph based on the user's points of interest.
[0247] In one embodiment, the second determining unit is further configured to acquire the region of interest selected by the user in the video frame image, and determine the user's point of interest based on the region of interest.
[0248] In one embodiment, the second determining unit is further configured to display multiple candidate points of interest and obtain the user's interest level in the multiple candidate points of interest;
[0249] The node weights and edge weights of the directed acyclic graph are determined based on the interest levels of each candidate interest point.
[0250] In one embodiment, the first determining unit is further configured to display the initial node weights and initial edge weights in the directed acyclic graph; receive operation instructions from the user based on the initial node weights and initial edge weights; and, in response to the user's operation instructions, determine the node weights and edge weights in the directed acyclic graph according to the initial node weights and initial edge weights.
[0251] In one embodiment, the determining module includes:
[0252] The first building unit is used to build a view sequence pool; the view sequence pool stores multiple candidate view sequences corresponding to the panoramic video.
[0253] The second building unit is used to construct a directed acyclic graph based on multiple candidate view sequences corresponding to the panoramic video.
[0254] In one embodiment, the first construction unit is further configured to determine the view sequence of each target in the panoramic video, and add the view sequence of each target to the view sequence pool as a candidate view sequence.
[0255] In one embodiment, the first building unit is further configured to perform target detection on keyframes in the panoramic video, determine target detection boxes, track the target detection boxes to obtain a tracking sequence, and determine the view sequence of each target based on the tracking sequence of the target detection boxes.
[0256] In one embodiment, the second construction unit is further configured to determine multiple nodes in a directed acyclic graph based on multiple candidate view sequences; arrange the corresponding multiple nodes according to the time order of each candidate view sequence to obtain a node array graph; and connect each node in the node array graph according to a preset rule to obtain a directed acyclic graph.
[0257] In one embodiment, the second building unit is further configured to divide each candidate view sequence into sub-view sequences at preset time intervals; treat each sub-view sequence as a node to obtain multiple nodes in a directed acyclic graph; and arrange the corresponding multiple nodes according to the time order of each sub-view sequence to obtain a node array graph.
[0258] In one embodiment, the second building unit is further configured to determine the time interval and spatial distance between every two nodes; if the time interval and spatial distance meet the preset conditions, the corresponding two nodes are connected to obtain a directed acyclic graph.
[0259] In one embodiment, the second building unit is further configured to connect the two corresponding nodes if the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance.
[0260] In one embodiment, the second building unit is further configured to determine a virtual start node and a virtual end node;
[0261] Connect the node corresponding to the first keyframe to the virtual start node; connect the node corresponding to the last keyframe to the virtual end node to obtain a directed acyclic graph.
[0262] In one embodiment, the third determining unit is further configured to determine the total weight of each view path on the directed acyclic graph based on the node weights and edge weights of the directed acyclic graph; and to determine the target view path based on the total weight of each view path.
[0263] In one embodiment, the third determining unit is further configured to take the view path corresponding to the largest total weight as the initial view path; perform interpolation processing on two adjacent nodes in the initial view path to obtain the second view path; determine the field of view angle corresponding to each video frame in the second view path; and determine the target view path based on the field of view angle corresponding to each video frame.
[0264] In one embodiment, the third determining unit is further configured to determine the ending frame view of the preceding node and the starting frame view of the following node among two adjacent nodes; and to perform interpolation processing on the ending frame view of the preceding node and the starting frame view of the following node to obtain the second view path.
[0265] In one embodiment, the third determining unit is further configured to determine the field of view angle corresponding to each key frame in the second view path; and to interpolate based on the field of view angles of two adjacent key frames to obtain the field of view angle of each video frame between two adjacent key frames.
[0266] In one embodiment, the third determining unit is further configured to determine the target bounding box in each keyframe; and to determine the field of view of the corresponding keyframe based on the width of the target bounding box and the width of the keyframe.
[0267] The various modules in the aforementioned panoramic video viewpoint path determination device and panoramic video editing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the panoramic camera in hardware form or independent of it, or stored in the memory of the panoramic camera in software form, so that the processor can call and execute the corresponding operations of each module.
[0268] In an exemplary embodiment, a panoramic camera is provided, which can be a terminal, and its internal structure can be as shown in Figure 21. The panoramic camera includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the panoramic camera provides computing and control capabilities. The memory of the panoramic camera includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the panoramic camera is used for exchanging information between the processor and external devices. The communication interface of the panoramic camera is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for determining the viewpoint path of panoramic video. The display unit of the panoramic camera is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the panoramic camera can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the housing of the panoramic camera, or external keyboards, touchpads, or mice, etc.
[0269] Those skilled in the art will understand that the structure shown in Figure 21 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the panoramic camera to which the present application is applied. A specific panoramic camera may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0270] In one embodiment, a panoramic camera is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0271] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0272] In one embodiment, a panoramic camera is also provided, including a display that, when executing a computer program, performs the following steps:
[0273] Display target viewpoint markers on the panoramic video; these markers indicate the location of the target viewpoint.
[0274] In one embodiment, the display performs the following steps when executing a computer program:
[0275] Candidate viewpoint identifiers and target viewpoint identifiers are displayed on the panoramic video. Candidate viewpoint identifiers represent the corresponding candidate viewpoint sequence, and target viewpoint identifiers represent the selected target viewpoint sequence.
[0276] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0277] In response to a user's instruction to change the target viewpoint in the panoramic video, the newly selected candidate viewpoint sequence is used as the target viewpoint sequence, and the target viewpoint path is updated.
[0278] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0279] Determine the start and end positions of the newly selected target view sequence, update the view sequence from the start to the end position in the target view path, switch from the start position to the newly selected target view sequence, and after the end position, switch back to the original target view path.
[0280] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0281] A directed acyclic graph is constructed based on multiple candidate view sequences corresponding to panoramic videos, and the node weights and edge weights in the directed acyclic graph are determined.
[0282] Path search is performed based on the node weights and edge weights of the directed acyclic graph to obtain the target view path.
[0283] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0284] Node weights and edge weights are automatically assigned to a directed acyclic graph according to preset rules.
[0285] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0286] The node weights and edge weights in a directed acyclic graph are determined based on user interactions.
[0287] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0288] Based on user interactions, user interests are obtained, and node weights and edge weights of the directed acyclic graph are determined based on these interests.
[0289] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0290] Obtain the region of interest selected by the user in the video frame image, and determine the user's points of interest based on the region of interest.
[0291] In one embodiment, the points of interest include a plurality of candidate points of interest, and when the display executes a computer program, it further performs the following steps: displaying the plurality of candidate points of interest;
[0292] When a processor executes a computer program, it also performs the following steps:
[0293] Obtain the user's level of interest in multiple candidate points of interest;
[0294] The node weights and edge weights of the directed acyclic graph are determined based on the interest levels of each candidate interest point.
[0295] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0296] Display the initial node weights and initial edge weights in the directed acyclic graph;
[0297] Receive user operation instructions based on initial node weights and initial edge weights;
[0298] In response to user commands, the node weights and edge weights in the directed acyclic graph are determined based on the initial node weights and initial edge weights.
[0299] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0300] Construct a view sequence pool; the view sequence pool stores multiple candidate view sequences corresponding to the panoramic video;
[0301] A directed acyclic graph is constructed based on multiple candidate view sequences corresponding to panoramic videos.
[0302] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0303] Determine the view sequence of each target in the panoramic video, and add the view sequence of each target to the view sequence pool as a candidate view sequence.
[0304] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0305] Target detection is performed on keyframes in panoramic video to determine target detection boxes, and the target detection boxes are tracked to obtain a tracking sequence;
[0306] Based on the tracking sequence of the target detection box, determine the view sequence of each target.
[0307] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0308] Multiple nodes in a directed acyclic graph are determined based on multiple candidate view sequences;
[0309] Based on the temporal order of each candidate view sequence, the corresponding multiple nodes are arranged to obtain a node array diagram;
[0310] The nodes in the node array graph are connected according to preset rules to obtain a directed acyclic graph.
[0311] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0312] Each candidate view sequence is segmented at preset intervals to obtain sub-view sequences;
[0313] Each sub-view sequence is treated as a node, resulting in multiple nodes in a directed acyclic graph;
[0314] The nodes are arranged according to the time order of each sub-view sequence to obtain a node array diagram.
[0315] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0316] Determine the time interval and spatial distance between every two nodes;
[0317] If the time interval and spatial distance meet the preset conditions, the corresponding two nodes will be connected to obtain a directed acyclic graph.
[0318] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0319] If the time interval is less than the preset time interval and the spatial distance is less than the preset spatial distance, then the two corresponding nodes will be connected.
[0320] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0321] Determine the virtual start node and virtual end node;
[0322] Connect the node corresponding to the first keyframe to the virtual start node;
[0323] Connect the node corresponding to the last keyframe to the virtual termination node to obtain a directed acyclic graph.
[0324] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0325] The total weight of each view path in the directed acyclic graph is determined based on the node weights and edge weights of the directed acyclic graph.
[0326] The target view path is determined based on the total weight of each view path.
[0327] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0328] Use the view path corresponding to the largest total weight as the initial view path;
[0329] Interpolation is performed on two adjacent nodes in the initial view path to obtain the second view path;
[0330] Determine the field of view corresponding to each video frame in the second-view path;
[0331] The target viewpoint path is determined based on the field of view corresponding to each video frame.
[0332] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0333] Determine the ending frame view of the preceding node and the starting frame view of the following node in two adjacent nodes.
[0334] Interpolate the viewpoint of the last frame of the previous node and the viewpoint of the first frame of the next node to obtain the second viewpoint path.
[0335] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0336] Determine the field of view corresponding to each keyframe in the second-view path;
[0337] Interpolation is performed based on the field of view angles of two adjacent keyframes to obtain the field of view angle of each video frame between the two adjacent keyframes.
[0338] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0339] Determine the bounding box in each keyframe;
[0340] Determine the field of view of the corresponding keyframe based on the width of the target bounding box and the width of the keyframe.
[0341] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0342] The target view path is determined based on multiple candidate view sequences corresponding to the panoramic video. The target view path consists of at least two target view sequences, which are selected from the candidate view sequences.
[0343] The panoramic video is edited based on the target view path to obtain a target video containing at least two target view sequences. The duration of the target video is the same as that of the panoramic video.
[0344] The implementation principles and technical effects of each step in this embodiment are similar to those of the panoramic video viewpoint path determination method and panoramic video editing method described above, and will not be repeated here.
[0345] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0346] The implementation principles and technical effects of each step in this embodiment when the computer program is executed by the processor are similar to those of the above-mentioned panoramic video viewpoint path determination method and panoramic video editing method, and will not be repeated here.
[0347] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0348] The implementation principles and technical effects of each step in this embodiment when the computer program is executed by the processor are similar to those of the above-mentioned panoramic video viewpoint path determination method and panoramic video editing method, and will not be repeated here.
[0349] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0350] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0351] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0352] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for determining the viewpoint path in panoramic video, characterized in that, The method comprises: determining a target view path based on a plurality of candidate view sequences corresponding to the panoramic video, the target view path being composed of at least two target view sequences, the target view sequences being selected from the candidate view sequences.
2. The method of claim 1, wherein, The method further comprises: displaying a target view identifier on the panoramic video, the target view identifier being used to identify the location of the target view.
3. The method of claim 1, wherein, The method further comprises: displaying a candidate view identifier and a target view identifier on the panoramic video, the candidate view identifier being used to represent the corresponding candidate view sequence, and the target view identifier being used to represent the selected target view sequence.
4. The method of claim 1, wherein, The method further comprises: updating the target view path by selecting a reselected candidate view sequence as the target view sequence in response to a user instruction to change the target view in the panoramic video.
5. The method of claim 4, wherein, The updating of the target view path comprises: determining the start position and the end position of the reselected target view sequence, updating the view sequence in the target view path from the start position to the end position, switching to the reselected target view sequence from the start position, and switching back to the original target view path after the end position.
6. The method of claim 1, wherein, The determining of the target view path based on the plurality of candidate view sequences corresponding to the panoramic video comprises: constructing a directed acyclic graph based on the plurality of candidate view sequences corresponding to the panoramic video, determining the node weight and the edge weight in the directed acyclic graph, and performing path search based on the node weight and the edge weight in the directed acyclic graph to obtain the target view path. The determining of the node weight and the edge weight in the directed acyclic graph comprises:
7. The method of claim 6, wherein, automatically assigning the node weight and the edge weight in the directed acyclic graph according to a preset rule. The determining of the node weight and the edge weight in the directed acyclic graph comprises:
8. The method of claim 6, wherein, determining the node weight and the edge weight in the directed acyclic graph based on the interaction with the user. The determining of the node weight and the edge weight in the directed acyclic graph based on the interaction with the user comprises:
9. The method of claim 8, wherein, obtaining the interest point of the user based on the interaction with the user, and determining the node weight and the edge weight in the directed acyclic graph based on the interest point of the user. The obtaining of the interest point of the user based on the interaction with the user comprises:
10. The method of claim 9, wherein, obtaining the interest region selected by the user in the video frame image, and determining the interest point of the user based on the interest region. The interest point comprises a plurality of candidate interest points, and the obtaining of the interest point of the user based on the interaction with the user and the determining of the node weight and the edge weight in the directed acyclic graph based on the interest point of the user comprises:
11. The method of claim 9, wherein, displaying the plurality of candidate interest points, obtaining the interest degree of the user for the plurality of candidate interest points, and determining the node weight and the edge weight in the directed acyclic graph according to the interest degree of each candidate interest point. The determining of the node weight and the edge weight in the directed acyclic graph comprises: displaying the initial node weight and the initial edge weight in the directed acyclic graph; 12. The method of claim 6, wherein, receiving an operation instruction of the user based on the initial node weight and the initial edge weight; determining the node weight and the edge weight in the directed acyclic graph according to the initial node weight and the initial edge weight in response to the operation instruction of the user. 13. The method of claim 6, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
14. The method of claim 13, wherein, The method comprises the following steps: The method comprises the following steps:
15. The method of claim 14, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
16. The method of claim 13, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
17. The method of claim 16, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
18. The method of claim 16, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
19. The method of claim 18, wherein, The method comprises the following steps: The method comprises the following steps:
20. The method of claim 18, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
21. The method of claim 6, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps:
22. The method of claim 21, wherein, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the interpolating two adjacent nodes in the initial view path to obtain a second view path; determining a field of view angle corresponding to each video frame in the second view path; determining the target view path based on the field of view angle corresponding to each video frame.
23. The method of claim 22, wherein, The interpolation of the two adjacent nodes in the initial view path to obtain a second view path comprises: determining the end frame view angle of the previous node and the start frame view angle of the next node in the two adjacent nodes; interpolating the end frame view angle of the previous node and the start frame view angle of the next node to obtain the second view path.
24. The method of claim 22, wherein, The determination of the field of view angle corresponding to each video frame in the second view path comprises: determining the field of view angle corresponding to each key frame in the second view path; interpolating the field of view angles of two adjacent key frames to obtain the field of view angle of each video frame between the two adjacent key frames.
25. The method of claim 24, wherein, The determination of the field of view angle corresponding to each key frame in the second view path comprises: determining a target frame in each key frame; determining the field of view angle of the corresponding key frame according to the width of the target frame and the width of the key frame.
26. A method of trimming a panoramic video, the method comprising: The method comprises: determining a target view path based on a plurality of candidate view sequences corresponding to a panoramic video, the target view path being composed of at least two target view sequences, the target view sequences being selected from the candidate view sequences; clipping the panoramic video based on the target view path to obtain a target video containing the at least two target view sequences, the target video having the same time length as the panoramic video.
27. The method of claim 26, wherein, The method further comprises: displaying a target view identifier on the panoramic video, the target view identifier being used to identify the location of the target view.
28. The method of claim 26, wherein, The method further comprises: displaying a candidate view identifier and a target view identifier on the panoramic video, the candidate view identifier being used to represent the corresponding candidate view sequence, and the target view identifier being used to represent the selected target view sequence.
29. The method of claim 26, wherein, The method further comprises: in response to a user instruction to change the target view in the panoramic video, updating the target view path by selecting a new candidate view sequence as the target view sequence.
30. The method of claim 29, wherein, The updating of the target view path comprises: determining the start position and the end position of the newly selected target view sequence, updating the view sequence in the target view path from the start position to the end position, switching to the newly selected target view sequence from the start position, and switching back to the original target view path after the end position.
31. The method of claim 26, wherein, The determination of the target view path based on a plurality of candidate view sequences corresponding to a panoramic video comprises: constructing a directed acyclic graph based on the plurality of candidate view sequences corresponding to the panoramic video, determining the node weight and the edge weight in the directed acyclic graph; performing path search based on the node weight and the edge weight of the directed acyclic graph to obtain the target view path.
32. The method of claim 31, wherein, The determination of the node weight and the edge weight in the directed acyclic graph comprises: automatically assigning the node weight and the edge weight of the directed acyclic graph according to a preset rule.
33. The method of claim 31, wherein, The determination of the node weight and the edge weight in the directed acyclic graph comprises: Determine node weights and edge weights in the directed acyclic graph based on the interaction with the user.
34. The method of claim 33, wherein, The determining node weights and edge weights in the directed acyclic graph based on the interaction with the user comprises: Obtaining interest points of the user based on the interaction with the user, and determining the node weights and edge weights of the directed acyclic graph based on the interest points of the user.
35. The method of claim 29, wherein, The obtaining interest points of the user based on the interaction with the user comprises: Obtaining an interest region in the video frame image that the user clicks, and determining the interest points of the user based on the interest region.
36. The method of claim 29, wherein, The interest points comprise a plurality of candidate interest points, and the obtaining interest points of the user based on the interaction with the user and determining the node weights and edge weights of the directed acyclic graph based on the interest points of the user comprise: Displaying a plurality of candidate interest points, and obtaining interest degrees of the user for the plurality of candidate interest points; Determining the node weights and edge weights of the directed acyclic graph according to the interest degrees of the respective candidate interest points.
37. The method of claim 36, wherein, The determining node weights and edge weights in the directed acyclic graph comprises: Displaying initial node weights and initial edge weights in the directed acyclic graph; Receiving operation instructions of the user based on the initial node weights and initial edge weights; In response to the operation instructions of the user, determining the node weights and edge weights in the directed acyclic graph according to the initial node weights and initial edge weights.
38. The method of claim 37, wherein, The constructing a directed acyclic graph based on a plurality of candidate view sequences corresponding to the panoramic video comprises: Constructing a view sequence pool; the view sequence pool stores a plurality of candidate view sequences corresponding to the panoramic video; Constructing a directed acyclic graph based on a plurality of candidate view sequences corresponding to the panoramic video.
39. The method of claim 38, wherein, The constructing a view sequence pool comprises: Determining view sequences of respective targets in the panoramic video, and adding the view sequences of the respective targets to the view sequence pool as the candidate view sequences.
40. The method of claim 39, wherein, The determining view sequences of respective targets in the panoramic video comprises: Performing target detection on key frames in the panoramic video to determine target detection boxes, and obtaining tracking sequences by tracking the target detection boxes; Determining the view sequences of the respective targets according to the tracking sequences of the target detection boxes.
41. The method of claim 38, wherein, The constructing a directed acyclic graph based on a plurality of candidate view sequences corresponding to the panoramic video comprises: Determining a plurality of nodes in the directed acyclic graph according to the plurality of candidate view sequences; Arranging a plurality of nodes corresponding to the respective candidate view sequences based on time sequences of the respective candidate view sequences to obtain a node array graph; Connecting the respective nodes in the node array graph based on a preset rule to obtain the directed acyclic graph.
42. The method of claim 41, wherein, The determining a plurality of nodes in the directed acyclic graph according to the plurality of candidate view sequences comprises: Splitting the respective candidate view sequences every preset time length to obtain sub-view sequences; Taking each sub-view sequence as a node to obtain the plurality of nodes in the directed acyclic graph; The arranging a plurality of nodes corresponding to the respective candidate view sequences based on time sequences of the respective candidate view sequences to obtain a node array graph comprises: arranging a plurality of nodes corresponding to the respective sub-view sequences based on time sequences of the respective sub-view sequences to obtain a node array graph.
43. The method of claim 41, wherein, The connecting the nodes in the node array graph based on the preset rule to obtain the directed acyclic graph comprises: determining a time interval and a spatial distance between each two nodes; if the time interval and the spatial distance satisfy a preset condition, connecting the corresponding two nodes to obtain the directed acyclic graph.
44. The method of claim 43, wherein, The connecting the nodes in the node array graph based on the preset rule to obtain the directed acyclic graph comprises: if the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance, connecting the corresponding two nodes.
45. The method of claim 43, wherein, The method further comprises: determining a virtual start node and a virtual end node; connecting a node corresponding to a first key frame to the virtual start node; connecting a node corresponding to a last key frame to the virtual end node to obtain the directed acyclic graph.
46. The method of claim 31, wherein, The path searching based on the node weight and the edge weight of the directed acyclic graph to obtain a target view path comprises: determining a total weight of each view path on the directed acyclic graph according to the node weight and the edge weight of the directed acyclic graph; determining the target view path according to the total weight of each view path.
47. The method of claim 46, wherein, The determining the target view path according to the total weight of each view path comprises: taking a view path corresponding to a maximum total weight as an initial view path; performing interpolation processing on adjacent two nodes in the initial view path to obtain a second view path; determining a field of view angle corresponding to each video frame in the second view path; determining the target view path based on the field of view angle corresponding to each video frame.
48. The method of claim 47, wherein, The performing interpolation processing on adjacent two nodes in the initial view path to obtain a second view path comprises: determining an end frame view angle of a previous node and a start frame view angle of a next node in the adjacent two nodes; performing interpolation processing on the end frame view angle of the previous node and the start frame view angle of the next node to obtain the second view path.
49. The method of claim 47, wherein, The determining a field of view angle corresponding to each video frame in the second view path comprises: determining a field of view angle corresponding to each key frame in the second view path; performing interpolation based on the field of view angles of adjacent two key frames to obtain a field of view angle of each video frame between the adjacent two key frames.
50. The method of claim 49, wherein, The determining a field of view angle corresponding to each key frame in the second view path comprises: determining a target box in each key frame; determining a field of view angle of a corresponding key frame according to a width of the target box and a width of the key frame.
51. An apparatus for determining a view path of a panoramic video, the apparatus comprising: The apparatus comprises: a determining module configured to determine a target view path based on a plurality of candidate view sequences corresponding to a panoramic video, the target view path being composed of at least two target view sequences, the target view sequences being selected from the candidate view sequences.
52. A panoramic camera comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program instructions, when executed by the processor, configure the processor to determine a target view path based on a plurality of candidate view sequences corresponding to a panoramic video, the target view path being composed of at least two target view sequences, the target view sequences being selected from the candidate view sequences.
53. The all-around camera of claim 52, wherein, The panoramic camera comprises a display configured to display a target view identifier on the panoramic video, the target view identifier being used to identify a location where a target view is located.
54. The all-around camera of claim 52, wherein, The panoramic camera comprises a display configured to display a candidate view identifier and a target view identifier on the panoramic video, the candidate view identifier being used to represent a corresponding candidate view sequence, and the target view identifier being used to represent a selected target view sequence.
55. The all-around camera of claim 52, wherein, The processor is configured to update the target view path by taking the reselected candidate view sequence as the target view sequence in response to an instruction of a user to change the target view in the panoramic video.
56. The all-around camera of claim 55, wherein, The processor is configured to determine a start position and an end position of the reselected target view sequence, update a view sequence in the target view path from the start position to the end position, switch to the reselected target view sequence from the start position, and switch back to the original target view path after the end position.
57. The all-around camera of claim 52, wherein, The processor is configured to construct a directed acyclic graph based on a plurality of candidate view sequences corresponding to the panoramic video, determine node weights and edge weights in the directed acyclic graph. The processor is configured to perform path searching based on the node weights and the edge weights of the directed acyclic graph to obtain a target view path.
58. The all-around camera of claim 57, wherein, The processor is configured to automatically assign the node weights and the edge weights of the directed acyclic graph according to a preset rule.
59. The all-around camera of claim 57, wherein, The processor is configured to determine the node weights and the edge weights in the directed acyclic graph based on interaction with a user.
60. The all-around camera of claim 59, wherein, The processor is configured to obtain a point of interest of a user based on interaction with the user, and determine the node weights and the edge weights of the directed acyclic graph based on the point of interest of the user.
61. The all-around camera of claim 60, wherein, The processor is configured to obtain an interest region pointed by the user in an image of a video frame, and determine the point of interest of the user based on the interest region.
62. The all-around camera of claim 60, wherein, The panoramic camera comprises a display, the point of interest comprises a plurality of candidate points of interest, the display is used to display the plurality of candidate points of interest, the processor is configured to obtain interest degrees of the plurality of candidate points of interest, and determine the node weights and the edge weights of the directed acyclic graph according to the interest degrees of the plurality of candidate points of interest.
63. The all-around camera of claim 57, wherein, The panoramic camera comprises a display, the display is used to display initial node weights and initial edge weights in the directed acyclic graph, the processor is configured to receive an operation instruction of a user based on the initial node weights and the initial edge weights, and determine the node weights and the edge weights in the directed acyclic graph according to the initial node weights and the initial edge weights in response to the operation instruction of the user.
64. The all-around camera of claim 57, wherein, The processor is configured to construct a view sequence pool, the view sequence pool stores a plurality of candidate view sequences corresponding to the panoramic video, and construct a directed acyclic graph based on the plurality of candidate view sequences corresponding to the panoramic video.
65. The all-around camera of claim 64, wherein, The processor is configured to determine view sequences of targets in the panoramic video, and add the view sequences of the targets to the view sequence pool as the candidate view sequences.
66. The all-around camera of claim 65, wherein, The processor is configured to perform target detection on key frames in the panoramic video, determine a target detection frame, and track the target detection frame to obtain a tracking sequence; According to the tracking sequence of the target detection frame, a view angle sequence of each target is determined.
67. The all-around camera of claim 64, wherein, The processor is configured to determine a plurality of nodes in the directed acyclic graph according to the plurality of candidate view angle sequences; arrange a plurality of nodes corresponding to the candidate view angle sequences in a time sequence to obtain a node array graph; According to a preset rule, each node in the node array graph is connected to obtain the directed acyclic graph.
68. The all-around camera of claim 67, wherein, The processor is configured to split each candidate view angle sequence into a sub-view angle sequence every preset time length to obtain a plurality of nodes in the directed acyclic graph; The node array graph is arranged according to the time sequence of each candidate view angle sequence.
69. The all-around camera of claim 67, wherein, The processor is configured to determine a time interval and a spatial distance between each two nodes; if the time interval and the spatial distance satisfy a preset condition, the corresponding two nodes are connected to obtain the directed acyclic graph.
70. The all-around camera of claim 69, wherein, If the time interval is less than a preset time interval and the spatial distance is less than a preset spatial distance, the corresponding two nodes are connected.
71. The all-around camera of claim 69, wherein, The processor is configured to determine a virtual starting node and a virtual ending node; connect a node corresponding to a first key frame to the virtual starting node; and connect a node corresponding to a last key frame to the virtual ending node to obtain the directed acyclic graph.
72. The all-around camera of claim 57, wherein, The processor is configured to determine a total weight of each view angle path on the directed acyclic graph according to a node weight and an edge weight of the directed acyclic graph; The target view angle path is determined according to the total weight of each view angle path.
73. The all-around camera of claim 72, wherein, The processor is configured to take a view angle path corresponding to a maximum total weight as an initial view angle path; perform interpolation processing on adjacent two nodes in the initial view angle path to obtain a second view angle path; determine a field of view angle corresponding to each video frame in the second view angle path; and determine the target view angle path based on the field of view angle corresponding to each video frame.
74. The all-around camera of claim 73, wherein, The processor is configured to determine an ending frame view angle of a previous node and a starting frame view angle of a next node in the adjacent two nodes; and perform interpolation processing on the ending frame view angle of the previous node and the starting frame view angle of the next node to obtain the second view angle path.
75. The all-around camera of claim 73, wherein, The processor is configured to determine a field of view angle corresponding to each key frame in the second view angle path; and perform interpolation on fields of view angle of adjacent two key frames to obtain a field of view angle of each video frame between the adjacent two key frames.
76. The all-around camera of claim 75, wherein, The processor is configured to determine a target frame in each key frame; and determine a field of view angle of a corresponding key frame according to a width of the target frame and a width of the key frame.
77. A computer readable storage medium having stored thereon a computer program, wherein The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 50.