Framing path correction method, electronic device, and storage medium

By identifying and correcting the initial framing path of panoramic video and generating the target framing path, the problems of cumbersome operation and low level of intelligence in existing technologies are solved, achieving a higher quality and simpler framing process.

WO2026050949A1PCT designated stage Publication Date: 2026-03-12ARASHI VISION INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing panoramic video framing methods are cumbersome to operate and have low levels of intelligence, resulting in framing position deviations and shaking, which affect the quality of the finished product and the user experience.

Method used

By obtaining the initial framing path, identifying the user's framing intention, and correcting the initial framing path based on the user's framing intention, the target framing path is generated by using object detection and saliency algorithms to identify points of interest.

Benefits of technology

It reduces framing position deviation and shake, improves image quality, reduces operational complexity, enhances the intelligence of framing path correction, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116964_12032026_PF_FP_ABST
    Figure CN2024116964_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video processing. Provided are a framing path correction method, an electronic device, and a storage medium. The framing path correction method comprises: acquiring an initial framing path obtained on the basis of user interaction; and identifying a user framing intent on the basis of the initial framing path, and correcting the initial framing path on the basis of the user framing intent, so as to obtain a target framing path.
Need to check novelty before this filing date? Find Prior Art

Description

Viewing path correction method, electronic device and storage medium TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video processing, in particular to a viewing path correction method, an electronic device and a storage medium. BACKGROUND

[0002] A panoramic video is a kind of video shot by a 3D panoramic camera in all directions of 360 degrees, and a user can watch it by adjusting different plane angles at will.

[0003] In related examples, a user can view a panoramic video by key frame dotting, body sensing, virtual joystick viewing or locking a viewing direction, but there are problems such as complicated operation and low intelligent degree, which affect the user experience.

[0004] SUMMARY

[0005] The present disclosure provides a viewing path correction method, an electronic device and a storage medium.

[0006] According to a first aspect, the present disclosure provides a viewing path correction method, comprising: obtaining an initial viewing path based on user interaction; and identifying a user viewing intention based on the initial viewing path, and correcting the initial viewing path based on the user viewing intention to obtain a target viewing path.

[0007] According to a second aspect, the present disclosure provides an electronic device, comprising: a processor configured to obtain an initial viewing path based on user interaction; identify a user viewing intention based on the initial viewing path, and correct the initial viewing path based on the user viewing intention to obtain a target viewing path.

[0008] According to a third aspect, the present disclosure provides a computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to implement the viewing path correction method described above.

[0009] According to embodiments of the present disclosure, the initial viewing path is identified based on the user viewing intention, and the initial viewing path is corrected based on the user viewing intention, which not only reduces the deviation or jitter of the viewing position, improves the image quality, but also reduces the complexity of the operation, improves the intelligent degree of the viewing path correction, and thus improves the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:

[0011] FIG. 1 schematically shows an application scenario diagram of the viewing path correction method according to an embodiment of the present disclosure;

[0012] FIG. 2 schematically shows a flowchart of a view path correction method according to an embodiment of the present disclosure;

[0013] FIG. 3 schematically shows a diagram of a view path correction method according to an embodiment of the present disclosure;

[0014] FIG. 4 schematically shows a diagram of a view path correction method according to another embodiment of the present disclosure;

[0015] FIG. 5 schematically shows a diagram of a method of calculating a view path similarity according to an embodiment of the present disclosure;

[0016] FIG. 6 schematically shows a diagram of a method of calculating a view path similarity according to another embodiment of the present disclosure;

[0017] FIG. 7 schematically shows a diagram of a view path correction according to an embodiment of the present disclosure;

[0018] FIG. 8 schematically shows a diagram of a view path correction according to another embodiment of the present disclosure.

[0019] FIG. 9 schematically shows a display page for soliciting user opinions according to an embodiment of the present disclosure;

[0020] FIG. 10 schematically shows a display page for comparing videos before and after correction according to an embodiment of the present disclosure; and

[0021] FIG. 11 schematically shows a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] To make the objects, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will be combined with the accompanying drawings for the embodiments of the present disclosure to clearly and completely describe the technical solutions of the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all. Based on the described embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present disclosure. It should be noted that throughout the drawings, the same elements are denoted by the same or similar reference numerals. In the following description, some specific embodiments are used only for the purpose of description, and should not be understood as any limitation on the present disclosure, but only as examples of the embodiments of the present disclosure. When it is possible to cause confusion to the understanding of the present disclosure, conventional structures or configurations will be omitted. It should be noted that the shapes and sizes of the components in the drawings do not reflect the true size and ratio, but only illustrate the content of the embodiments of the present disclosure.

[0023] Unless otherwise defined, technical terms or scientific terms used in the embodiments of the present disclosure should be understood as the common meanings to those skilled in the art. The terms “first”, “second”, and similar words of comparison used in the embodiments of the present disclosure do not represent any order, quantity, or importance, but are only used to distinguish different components.

[0024] In related videos, the main shooting modes for panoramic videos include key frame dotting, body sensing shooting, virtual joystick shooting, or locking a shooting direction, etc.

[0025] For the key frame dotting shooting mode, the user needs to manually dot key frames, preview, and constantly fine-tune according to the preview effect, which is tedious.

[0026] For the body sensing shooting and virtual joystick shooting modes, although the shooting operation can be quickly completed, the user's operation has a certain lag, which leads to interest point following lag, picture center deviation, and picture shaking, affecting the picture quality.

[0027] Therefore, the embodiments of the present disclosure identify the user's shooting path based on the initial shooting path, and correct the initial shooting path based on the user's shooting intention, which not only reduces the shooting position deviation or shaking, improves the picture quality, but also reduces the operation complexity, improves the intelligent degree of shooting path correction, and thus improves the user experience.

[0028] FIG. 1 schematically shows an application scenario diagram of a shooting path correction method according to an embodiment of the present disclosure.

[0029] As shown in FIG. 1, in the exemplary architecture 100, a first terminal device 101, a second terminal device 102, a server 103, and a network 104 can be included. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, and the server 103. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.

[0030] The user can use the first terminal device 101 and / or the second terminal device 102 to interact with the server 103 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101 and / or the second terminal device 102, such as applications with image acquisition functions, applications with image processing functions, etc.

[0031] The first terminal device 101 and / or the second terminal device 102 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc.

[0032] The server 103 can be a server providing various services, such as a background management server (for example only) providing support for content browsed by a user using the first terminal device 101 and / or the second terminal device 102. The background management server can perform analysis and the like on received user requests and the like, and feed back the processing results (such as a webpage, information, or data, or the like, obtained or generated according to a user request) to the terminal device.

[0033] It should be noted that the view path correction method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101 and / or the second terminal device 102.

[0034] Alternatively, the view path correction method provided by the embodiments of the present disclosure can also be executed by the server 103. The view path correction method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the first terminal device 101, the second terminal device 102, and / or the server 105.

[0035] For example, a user can use a terminal device to shoot a panoramic video, then identify the user's view intention based on the initial view path of the shot panoramic video, correct the initial view path based on the user's view intention, and obtain a target view path.

[0036] FIG. 2 schematically shows a flowchart of the view path correction method of the embodiments of the present disclosure.

[0037] As shown in FIG. 2, the view path correction method 200 can include operations S210-S220.

[0038] S210, obtaining an initial view path based on user interaction.

[0039] The view path is composed of determined viewing angles of each video frame. The initial view path is obtained through user interaction. The terminal determines the viewing angle of each video frame in the original video through user interaction, and then obtains the initial view path based on the viewing angle of each video frame. The role of the view path includes playing and displaying according to the view path or editing the original video according to the view path. The original video can be a panoramic video, a super wide-angle video, a single wide-angle video, a single fisheye video, or the like, which has a larger field of view. The original video can be shot and obtained by a corresponding camera of an electronic device, or can be shot by a panoramic camera or other camera device, obtained based on data transmission, or obtained by splicing multiple ordinary camera shots.

[0040] The terminal device interacting with the user can be a device for shooting a video, such as a camera, or a device for processing a video, such as a mobile phone. The interaction can be a framing operation on the original video, and the framing operation includes at least one of the following: a body-sensing framing operation, a virtual joystick framing operation, a predetermined direction framing operation, and a key frame dotting framing operation.

[0041] In S220, a user framing intention is identified based on the initial framing path, and the initial framing path is corrected based on the user framing intention to obtain a target framing path.

[0042] According to an embodiment of the present disclosure, the initial framing path indicates a time sequence of framing parameters of each video frame. The framing parameters can include a framing position, a framing direction, a field of view, and the like.

[0043] According to an embodiment of the present disclosure, the user framing intention can include, but is not limited to, a point of interest, a framing angle of view, a framing path, and the like. The point of interest can be a preset target object, a salient object, or a region where an object with a highlight performance is located, and the like, or can be a region where a collection of objects is located, for example, a key figure, a beautiful scenery, a group photo of multiple people, an interactive screen of a person and an animal, an interactive screen of a person and an object, a group photo screen of a person and a scenery, a forward direction, and the like. The point of interest can also be any point in the region described above, for example, a center point. The framing angle of view can be, for example, an asteroid angle of view.

[0044] According to an embodiment of the present disclosure, the point of interest can be identified based on the initial framing path by using a target detection algorithm or a saliency algorithm. The target detection algorithm includes, but is not limited to, Regin-CNN (Convolutional Neural Network), SPP-Net (Spatial Pyramid Pooling Networks), Fast Regin-CNN, R-FCN (Region-based Fully Convolutional Network), OverFeat, YOLOv1, YOLOv3, SSD (Single Shot MultiBox Detector), and RetinaNet, and the like. The saliency algorithm includes, but is not limited to, a detection method for a highlight video segment, and the like. It should be noted that the above algorithms are only exemplary, and do not necessarily mean that the embodiments of the present disclosure use these algorithms.

[0045] According to an embodiment of the present disclosure, the correction of the initial framing path based on the user framing intention can be a correction of the framing path of the video frame at the same video capture moment. The correction operation includes, but is not limited to, replacing or covering the initial framing path by using the framing path of the point of interest corresponding to the user framing intention, adjusting the framing parameters of the video frame whose target is offset from the center of the screen in the initial framing path, and the like.

[0046] For example, when the user takes a shot by key frame dotting, when the dotting density is low, the shot position is prone to deviation or jitter, thereby affecting the shot quality. When the dotting density is high, although the shot quality is improved, the number of dotting shot operations is also increased, and the operation complexity is increased.

[0047] Therefore, by identifying the user shot intention based on the initial shot path and correcting the initial shot path based on the user shot intention, not only the shot position deviation or jitter is reduced, the shot quality is improved, but also the operation complexity is reduced, the intelligent degree of shot path correction is improved, thereby improving the user experience.

[0048] According to an embodiment of the present disclosure, the operation S220 described above can include the following operations: determining a target interest point matched with the initial shot path based on the initial shot path, taking the target interest point as the user shot intention; and correcting the initial shot path based on the shot path of the target interest point to obtain a target shot path.

[0049] FIG. 3 schematically shows a schematic diagram of a shot path correction method according to an embodiment of the present disclosure.

[0050] As shown in FIG. 3, in this embodiment 300, a target detection algorithm or a saliency detection algorithm can be used to determine a target interest point 320 matched with the initial shot path 310, for example, target detection can be performed on the panoramic video to obtain at least one interest point, and the target interest point is taken as the user shot intention. Then, based on the at least one interest point, target tracking is performed on the panoramic video to generate a shot path 330 of the target interest point corresponding to the at least one interest point. The target tracking algorithm includes but is not limited to YOLO (You Only Look Once) v1-YOLOX, SSD (Single Shot MultiBox Detector) and RetinaNet. The SiamRPN (Siamese Region Proposal Network) can also be used to track the target interest point to obtain the shot path 330 of the target interest point. The shot path 330 of the target interest point can also be a shot path of a foreground direction centered on the target interest point. It should be noted that the above algorithms are only exemplary, and do not necessarily mean that the embodiments of the present disclosure use these algorithms.

[0051] According to an embodiment of the present disclosure, the at least one interest point and the panoramic video can also be input into a trained convolutional neural network to output an interest point shot path corresponding to the at least one interest point.

[0052] For example, the trained convolutional neural network can be a convolutional neural network constructed based on an attention mechanism.

[0053] Then, the initial framing path 310 is corrected based on the target framing path 330 of the target interest point. The target framing path 330 of the target interest point can replace the initial framing path 310 to obtain a target framing path 340.

[0054] According to an embodiment of the present disclosure, the initial framing path is corrected based on the target framing path of the interest point obtained by tracking the interest point. Since the target framing path of the interest point is more accurate and the target is more centered, the picture of the target framing path is more stable and has higher picture quality than the initial framing path.

[0055] In addition to taking the interest point as the framing intention, a plurality of candidate framing paths can be set in advance, and the matched candidate framing path can be taken as the framing intention. In an embodiment, the operation S220 can include the following operations: comparing the initial framing path with a plurality of candidate framing paths, determining the framing path matched with the initial framing path, taking the matched framing path as the user framing intention; and correcting the initial framing path based on the matched framing path to obtain a target framing path.

[0056] FIG. 4 schematically shows a schematic diagram of a framing path correction method according to another embodiment of the present disclosure.

[0057] As shown in FIG. 4, in the embodiment 400, the initial framing path 410 can be compared with the candidate framing path 420 to determine the framing path 430 matched with the initial framing path 410 from the candidate framing path 420. Then, the initial framing path 410 is corrected by using the framing path 430 to obtain a target framing path 440.

[0058] According to an embodiment of the present disclosure, the candidate framing path can include but is not limited to the framing path of the interest point and the historical framing path of the user. The historical framing path can be obtained based on at least one of the following framing operations: the historical key frame dotting operation of the user, the historical motion sensing framing operation, the historical virtual joystick framing operation, and the historical framing direction locking operation.

[0059] According to an embodiment of the present disclosure, the candidate framing path can be obtained by a camera, a mobile phone, or a cloud server interacting with the camera or the mobile phone. It can also be obtained by a gimbal interacting with a terminal device.

[0060] For example, the candidate framing path for the panoramic video can be received from the gimbal by detecting the existence of the gimbal interacting with the terminal device and establishing a wireless communication connection with the gimbal.

[0061] According to an embodiment of the present disclosure, the terminal device can be a mobile phone, and the terminal device can be fixed on a gimbal. Since wireless communication is established between the gimbal and the terminal device, during the process of shooting a panoramic video of a target by the user using the mobile phone, the posture of the gimbal can be adjusted according to a tracking mode selected by the user, for example, horizontal tracking or direction-locked tracking, to track the shooting target, so as to obtain a candidate framing path of the panoramic video.

[0062] According to an embodiment of the present disclosure, since the initial framing path and the candidate framing path are both time sequences of framing parameters of video frames. Therefore, when comparing the initial framing path with the plurality of candidate framing paths to determine the framing path matched with the initial framing path, the framing parameters in the initial framing path corresponding to the same time or time period can be compared with the framing parameters in the candidate framing path to calculate the framing parameter difference, and the candidate framing path with smaller framing parameter difference is determined as the framing path matched with the initial framing path.

[0063] According to an embodiment of the present disclosure, the framing parameter difference calculation method includes but is not limited to standard deviation, variance, and cosine distance.

[0064] For example, the candidate framing path with a framing parameter difference less than a difference threshold can be determined as the framing path matched with the initial framing path by setting the difference threshold.

[0065] According to an embodiment of the present disclosure, since the candidate framing path is no longer limited to the path of the interest point, the initial framing path can be corrected according to the historical operation preference of the user, and therefore, the adaptability to the personalized editing needs of the user in the framing path correction process is improved.

[0066] According to an embodiment of the present disclosure, the framing path similarity between the initial framing path and each candidate framing path can also be determined according to the framing parameters of the initial framing path and the framing parameters of the plurality of candidate framing paths; and the framing path matched with the initial framing path is determined from the plurality of candidate framing paths based on the framing path similarity.

[0067] According to an embodiment of the present disclosure, the similarity of the framing path can include the similarity between any dimension of the framing parameters. For example, the similarity can be the similarity between the framing position parameters, the similarity between the framing field of view parameters, the similarity between the framing direction parameters, or the similarity between the parameter matrix of any combination of the multi-dimensional parameters such as the framing position, the framing field of view, and the framing direction.

[0068] According to an embodiment of the present disclosure, the calculation method of the similarity includes but is not limited to the Euclidean distance algorithm, the Manhattan distance algorithm, and the cosine distance algorithm.

[0069] For example, the similarity between the initial framing path and each candidate framing path can be calculated based on the Euclidean distance algorithm. Then, the multiple candidate framing paths are sorted in descending order of similarity, and the candidate framing path ranked first is determined as the framing path matched with the initial framing path.

[0070] According to an embodiment of the present disclosure, the framing path matched with the initial framing path is determined based on the framing path similarity, the calculation logic is simple, the calculation amount is small, the requirement of hardware resources for the framing path comparison process can be reduced, and the data processing speed of the framing path correction process can be improved.

[0071] The process of calculating the framing path similarity based on the framing position parameter and the combination of the framing position parameter and the framing field angle parameter will be described in detail below with reference to FIGS. 5-6.

[0072] FIG. 5 schematically shows a method for calculating the framing path similarity according to an embodiment of the present disclosure.

[0073] As shown in FIG. 5, the method 500 for calculating the framing path similarity can include operations S510-S550.

[0074] S510, determine the similarity between the framing position parameter of each candidate framing path and the framing position parameter of the initial framing path on the same video frame, to obtain the video frame framing position similarity.

[0075] S520, determine whether the video frame framing position similarity is greater than a first preset threshold? If yes, perform operation S530, and if not, return to operation S510 for the next video frame.

[0076] S530, determine whether the current frame is the last frame? If yes, perform operation S540, and if not, return to operation S510 for the next video frame.

[0077] S540, count the number of video frames with the video frame framing position similarity greater than the first preset threshold in each candidate framing path, and return to operation S510 for the next video frame.

[0078] S550, obtain the framing path similarity between each candidate framing path and the initial framing path. The framing path similarity can be the number of video frames with the video frame framing position similarity greater than the first preset threshold in each candidate framing path.

[0079] According to an embodiment of the present disclosure, the framing position parameter can be the center point of the framing position,

[0080] For example, the parameters of the center point of the initial framing position of a video frame in the panoramic video can be represented in the spherical coordinate system as (r, θ0, ), the parameter of the position center point corresponding to the candidate framing path Path1 of the same video frame is (r, θ1, ). Then, the cosine value between (r, θ0, ) and (r, θ1, ) can be calculated to obtain the video frame framing position similarity between the initial framing path and the candidate framing path of the video frame. When the video frame framing position similarity is greater than a first predetermined threshold, for example, which can be 0.9, counting can be performed.

[0081] Next, the video frame framing position similarity calculation is performed again for the next video frame until the last frame is processed, and the number of video frames with video frame framing position similarity greater than the first preset threshold is counted for each candidate framing path. For example, the number of video frames with video frame framing position similarity greater than the first predetermined threshold for the candidate framing path Path1 is 32, and the number of video frames with video frame framing position similarity greater than the first predetermined threshold for the candidate framing path Path2 is 35. Then, the candidate framing path Path2 can be determined as the framing path matched with the initial framing path.

[0082] According to embodiments of the present disclosure, the accuracy of the video frame framing position directly affects the quality of the panoramic video picture, therefore, determining the framing path for correcting the initial framing path based on the video frame framing position similarity can further improve the framing accuracy, so that the framing target is more localized in the center position of the picture, thereby improving the picture quality.

[0083] During the shooting of the panoramic video, especially for the shooting of dynamic targets, the field of view angle is constantly changing to ensure that the shooting device can collect the target, therefore, in order to further improve the framing accuracy, the framing path similarity can be determined in combination with the framing field of view angle parameter and the framing position parameter.

[0084] FIG. 6 schematically shows a method for calculating framing path similarity according to another embodiment of the present disclosure.

[0085] As shown in FIG. 6, the framing similarity calculation method 600 can include operations S610-S670.

[0086] S610, determine the similarity between the framing position parameters of each candidate framing path on the same video frame and the framing position parameters of the initial framing path to obtain the video frame framing position similarity.

[0087] S620, determine whether the video frame framing position similarity is greater than a first preset threshold? If yes, perform operation S630; if not, return to operation S610 for the next video frame.

[0088] S630, determine the similarity between the field of view angle parameter of each candidate framing path and the field of view angle parameter of the initial framing path on the same video frame, to obtain a field of view angle similarity.

[0089] S640, determine whether the field of view angle similarity is less than a second preset threshold value? If yes, perform operation S650; if not, return to perform operation S610 for the next video frame.

[0090] S650, determine whether the current frame is the last frame? If yes, perform operation S660, if not, return to perform operation S610 for the next video frame.

[0091] S660, count the number of video frames in each candidate framing path whose video frame framing position similarity is greater than the first preset threshold value and whose field of view angle similarity is less than the second preset threshold value, and return to perform operation S610 for the next video frame.

[0092] S670, obtain the framing path similarity of each candidate framing path and the initial framing path. The framing path similarity can be the number of video frames in each candidate framing path whose video frame framing position similarity is greater than the first preset threshold value and whose field of view angle similarity is less than the second preset threshold value.

[0093] According to embodiments of the present disclosure, the calculation of the video frame framing position similarity is the same as described above and will not be repeated here. The field of view angle parameter can be an angle formed by two edges of the maximum range of the lens of the panoramic camera with the lens as the vertex, and the object image of the target being shot can pass through the lens.

[0094] For example, the field of view angle of the initial framing path of a video frame of a panoramic video can be represented as Fov1, and the field of view angle of the candidate framing path Path1 of the same video frame can be represented as Fov2. The field of view angle similarity can be defined as max(Fov1 / Fov2, Fov2 / Fov1). When the video frame framing position similarity is greater than the first predetermined threshold value and the field of view angle parameter is less than the second predetermined threshold value, for example, the second predetermined threshold value can be 1.2, counting can be performed.

[0095] Then, the video frame framing position similarity and the field of view angle similarity are calculated for the next video frame, until the last frame is processed, and the number of video frames in each candidate framing path whose video frame framing position similarity is greater than the first preset threshold value and whose field of view angle similarity is less than the second preset threshold value is counted.

[0096] For example, the number of video frames in which the candidate framing path Path 1 is similar to the framing position and the framing field of view is less than the second preset threshold is 16, and the number of video frames in which the candidate framing path Path 2 is similar to the framing position and the framing field of view is less than the second preset threshold is 11. The candidate framing path Path 1 can be determined as the framing path matched with the initial framing path.

[0097] According to an embodiment of the present disclosure, the framing path similarity is determined in combination with the framing field of view parameter and the framing position parameter, which can further improve the framing accuracy of the dynamic target and thus improve the picture quality of the dynamic target.

[0098] In an actual application scenario, the number of video frames included in the panoramic video is relatively large. Obviously, comparing the initial framing path and the candidate framing path of each video frame will reduce the processing speed in the framing path correction process. Therefore, the video frames to be processed can be determined from the multiple candidate framing paths and the initial framing path based on a frame extraction strategy, and the framing path similarity can be determined according to the framing parameters corresponding to the video frames to be processed.

[0099] According to an embodiment of the present disclosure, the frame extraction strategy includes but is not limited to: isochronous interval frame extraction, for example, extracting one frame every 0.5 ms; extracting key frames by detecting the key frames. For example, the key frames can be determined by detecting the scene changes in adjacent frames. The frame extraction strategy can be configured based on the actual needs of the application scenario, and the present disclosure does not make specific limitations thereto.

[0100] According to an embodiment of the present disclosure, the frame extraction strategy for determining the video frames to be processed from the multiple candidate framing paths and the initial framing path is the same, so as to calculate the framing path similarity for the framing parameters on the same video frame.

[0101] According to an embodiment of the present disclosure, the video frames to be processed are determined based on the frame extraction strategy, which can reduce the data processing amount in the framing path correction process and improve the efficiency of the framing path correction.

[0102] According to an embodiment of the present disclosure, the initial framing path can be replaced by the matched framing path to obtain a target framing path.

[0103] For example, the initial framing parameters corresponding to the target time period in the initial framing path can be replaced by the matched framing path to generate a target framing path. The target time period represents the same video capture period of the target framing path and the initial framing path of the target interest point.

[0104] According to an embodiment of the present disclosure, when the matched framing path is longer than the initial framing path, the matched framing path can be directly taken as the target framing path.

[0105] For example, the video capture period corresponding to the initial framing path can be t1-t 10 , the video capture period corresponding to the matched framing path can be t1-t 12 , and the target framing path can be the matched framing path.

[0106] According to an embodiment of the present disclosure, when the matched framing path is shorter than or equal to the initial framing path, the initial framing path of the same video capture period can be replaced by the matched framing path, and the initial framing path of other periods is retained.

[0107] For example, the video capture period corresponding to the initial framing path can be t1-t 10 , the video capture period corresponding to the matched framing path can be t1-t6. The initial framing path of the period t1-t6 can be replaced by the matched framing path, and the initial framing path of the period t7-t 10 is retained to obtain the target framing path.

[0108] According to an embodiment of the present disclosure, the matched framing path directly covers or replaces the initial framing path, which can reduce the complex operation in the framing path correction process and improve the correction efficiency compared with adjusting the framing parameters of each video frame.

[0109] In the framing scene of user key frame dotting, multiple targets can be framed in the initial framing path of the panoramic video, for example, the front direction is taken for 10 seconds, the transition is made for 10-12 seconds, the scene beside the road is taken for 12-20 seconds, the transition is made for 20-22 seconds, the target person is taken for 23-30 seconds, the transition is made for 30-32 seconds, and the asteroid framing view is switched for 32-30 seconds.

[0110] According to an embodiment of the present disclosure, in the framing scene, the panoramic video can be sliced into multiple video segments according to a predetermined time length, for example, 2s; the user framing intention corresponding to each video segment is identified based on the initial framing path corresponding to each video segment; the initial framing path corresponding to each video segment is corrected based on the user framing intention corresponding to each video segment to obtain the target framing path corresponding to each video segment.

[0111] FIG. 7 schematically shows a schematic diagram of framing path correction according to an embodiment of the present disclosure.

[0112] As shown in FIG. 7, in embodiment 700, the panoramic video 710 can be sliced according to a predetermined time length to obtain video segment P-1 to video segment P-n. For each video segment, the view path correction can be performed in parallel, for example, the user view intention I-1 is identified based on the initial view path of the video segment P-1, and the initial view path of the video segment P-1 is corrected based on the user view intention I-1 to obtain the view path Path-1. The processing procedures of other video segments are the same as that of the video segment P-1, which will not be described herein.

[0113] According to embodiments of the present disclosure, by slicing the panoramic video and performing view path correction on each video segment respectively, the view path after the user key frame dotting can be combined with the view operation of the user key frame dotting to further improve the accuracy of the view path after the user key frame dotting.

[0114] After the view path correction, since the target view path can include multiple interest points, and the target view path also includes part of the initial view path, the connection between adjacent interest points in the target view path can be discontinuous, which can easily cause the playback picture to shake.

[0115] Therefore, the view path correction method provided by the embodiments of the present disclosure can further include the following operation: when it is detected that the target view path includes multiple interest points, the view path between adjacent interest points in the multiple interest points is smoothed.

[0116] FIG. 8 schematically shows a view path correction diagram of another embodiment of the present disclosure.

[0117] As shown in FIG. 8, in embodiment 800, the initial view path 810 can include the initial view path of the interest point A, the initial view path of the interest point B, and the initial view path of the interest point C. The view path correction is performed by using the respective matching paths PA, PB and PC to obtain the corrected view path Path-A, Path-B and Path-C.

[0118] Then, the view path Path-A corresponding to the adjacent interest point A and the interest point B is smoothed, and the view path Path-B corresponding to the adjacent interest point B and the interest point C is smoothed, and finally the target view path 820 is obtained.

[0119] According to embodiments of the present disclosure, the smoothing processing includes but is not limited to the sliding window average method, the filtering method, etc.

[0120] According to the embodiment of the present disclosure, the target framing path is smoothed, the continuity between adjacent interest points is improved, the shaking probability of the playing picture is reduced, and the picture quality is further improved.

[0121] In actual application scenarios, the framing path correction method provided by the embodiment of the present disclosure can be combined with the current framing operation on the panoramic video.

[0122] For example, the initial framing path can be obtained in response to the framing operation on the panoramic video, the user framing intention is identified based on the initial framing path, and the initial framing path is corrected based on the user framing intention. The framing operation can include at least one of the following: a body sensing framing operation, a virtual joystick framing operation, a predetermined direction framing operation, and a key frame dotting framing operation.

[0123] For example, before the initial framing path is corrected based on the user framing intention to obtain the target framing path, a prompt message indicating whether the initial framing path needs to be corrected can be displayed; after the user's confirmation operation is obtained, the step of correcting the initial framing path based on the user framing intention is entered.

[0124] FIG. 9 schematically shows a display page for soliciting user opinions according to an embodiment of the present disclosure.

[0125] As shown in FIG. 9, in the display page 900, an inquiry text box of "whether to start AI correction" and buttons for user selection can be displayed. When the user selects "yes", the step of correcting the initial framing path based on the user framing intention is entered. When the user selects "no", the correction step is not performed and the current framing operation of the user is continued to be responded to.

[0126] In order to facilitate the user to intuitively compare the pictures before and after correction, after the initial framing path is corrected based on the user framing intention to obtain the target framing path, the first video corresponding to the initial framing path and the second video corresponding to the target framing path can be determined; the first video and the second video are compared and displayed.

[0127] FIG. 10 schematically shows a display page for comparing videos before and after correction according to an embodiment of the present disclosure.

[0128] As shown in FIG. 10, in the display page 1000, the initial video interface before correction and the video interface after correction can be displayed. The layout of the initial video interface before correction and the video interface after correction in the current display page can be configured according to actual application requirements, for example, they can be displayed above and below, left and right, or in different picture sizes, so that when the user clicks on a certain video interface, the display is enlarged to improve the flexibility of comparison and display and the flexibility of adaptation to the screen size of the display device.

[0129] FIG. 11 shows a schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown in FIG. 11, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0130] As shown in FIG. 11, the electronic device 1100 includes a processor 1101 that can perform various suitable actions and processes in accordance with computer programs stored in a read-only memory (ROM) 1102 or loaded into a random access memory (RAM) 1103 from a storage unit 1108. Various programs and data required by the device 1100 for operation can also be stored in the RAM 1103. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other by a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0131] Various components in the device 1100 are connected to the I / O interface 1105, including an input unit 1106, such as a keyboard, a mouse, etc.; a display 1107; a storage unit 1108, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0132] The processor 1101 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The processor 1101 performs various methods and processes described above, such as the view path correction method. For example, in some embodiments, the view path correction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the processor 1101, one or more steps of the view path correction method described above can be performed. Alternatively, in other embodiments, the processor 1101 can be configured to perform the view path correction method by any other appropriate means, such as by means of firmware.

[0133] For example, the processor 1101 is configured to obtain an initial view path based on a user interaction, identify a user view intention based on the initial view path, and correct the initial view path based on the user view intention to obtain a target view path.

[0134] For example, the processor 1101 is further configured to determine a target point of interest matching the initial view path based on the initial view path, take the target point of interest as the user view intention, and correct the initial view path based on a view path of the target point of interest to obtain the target view path.

[0135] For example, the processor 1101 is further configured to compare the initial view path with a plurality of candidate view paths, determine a view path matching the initial view path from the plurality of candidate view paths, take the matching view path as the user view intention, and correct the initial view path based on the matching view path to obtain the target view path.

[0136] According to embodiments of the present disclosure, the processor 1101 is further configured to determine a view path similarity between the initial view path and each candidate view path according to a view parameter of the initial view path and a view parameter of the plurality of candidate view paths, and determine a view path matching the initial view path from the plurality of candidate view paths based on the view path similarity.

[0137] According to an embodiment of the present disclosure, the framing parameters comprise: a framing position parameter of each video frame; the processor 1101 is further configured to determine a similarity between the framing position parameter of each candidate framing path and the framing position parameter of the initial framing path on the same video frame, to obtain a video frame framing position similarity; and calculate the framing path similarity with the initial framing path based on a number of video frames in each candidate framing path that meet a first predetermined condition, wherein the first predetermined condition comprises that the video frame framing position similarity is greater than a first preset threshold.

[0138] According to an embodiment of the present disclosure, the framing parameters further comprise: a framing field of view angle parameter of each video frame; the processor 1101 is further configured to determine a similarity between the framing field of view angle parameter of each candidate framing path and the framing field of view angle parameter of the initial framing path on the same video frame, to obtain a framing field of view angle similarity; and calculate the framing path similarity with the initial framing path based on a number of video frames in each candidate framing path that meet a second predetermined condition, wherein the second predetermined condition comprises that the video frame framing position similarity is greater than a first preset threshold and the framing field of view angle similarity is less than a second preset threshold.

[0139] According to an embodiment of the present disclosure, the processor 1101 is further configured to determine, based on the frame extraction strategy, to-be-processed video frames from the plurality of candidate framing paths, the initial framing path respectively; and determine the framing path similarity according to the framing parameters corresponding to the to-be-processed video frames.

[0140] According to an embodiment of the present disclosure, the processor 1101 is further configured to replace, by using the matched framing path, the initial framing parameters corresponding to the target time period in the initial framing path, to generate a target framing path, wherein the target time period represents a same video collection time period of the target interest point framing path and the initial framing path.

[0141] According to an embodiment of the present disclosure, the processor 1101 is further configured to, in response to detecting that the target framing path comprises a plurality of interest points, perform smoothing processing on the framing paths of adjacent interest points.

[0142] According to an embodiment of the present disclosure, the processor 1101 is further configured to slice the panoramic video according to a predetermined time length, to obtain a plurality of video segments; identify a user framing intention corresponding to each video segment based on the initial framing path corresponding to each video segment; correct the initial framing path corresponding to each video segment based on the user framing intention corresponding to each video segment, to obtain a target framing path corresponding to each video segment.

[0143] According to an embodiment of the present disclosure, the processor 1101 is further configured to perform target detection on the panoramic video, to obtain at least one interest point; perform target tracking on the panoramic video based on the at least one interest point, to generate an interest point framing path corresponding to the at least one interest point.

[0144] According to an embodiment of the disclosure, the processor 1101 is further configured to input the at least one interest point and the panoramic video into a trained convolutional neural network, and output an interest point framing path corresponding to the at least one interest point.

[0145] According to an embodiment of the disclosure, the processor 1101 is further configured to acquire an initial framing path in response to a framing operation for the panoramic video. The framing operation includes at least one of the following: a motion sensing framing operation, a virtual joystick framing operation, a predetermined direction framing operation, and a key frame dotting framing operation.

[0146] According to an embodiment of the disclosure, the processor 1101 is further configured to, in response to detecting that there is a gimbal interacting with the terminal device, establish a wireless communication connection with the gimbal, and receive a candidate framing path for the panoramic video from the gimbal.

[0147] According to an embodiment of the disclosure, the display 1107 is configured to display prompt information indicating whether the initial framing path needs to be corrected before the initial framing path is corrected based on the user framing intention to obtain a target framing path. The processor 1101 is further configured to, after obtaining the user confirmation operation, enter the step of correcting the initial framing path based on the user framing intention.

[0148] According to an embodiment of the disclosure, the processor 1101 is further configured to, after correcting the initial framing path based on the user framing intention to obtain the target framing path, determine a first video corresponding to the initial framing path and a second video corresponding to the target framing path. The display 1107 is further configured to compare and display the first video and the second video.

[0149] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0150] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0151] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0152] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0153] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0154] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between a client and a server is one of client-server. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0155] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.

[0156] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A view path correction method characterized by, The method comprises: obtaining an initial framing path based on user interaction; and identifying a user framing intention based on the initial framing path, and correcting the initial framing path based on the user framing intention to obtain a target framing path.

2. The method of claim 1, wherein, The identifying a user framing intention based on the initial framing path, and correcting the initial framing path based on the user framing intention to obtain a target framing path comprises: determining a target point of interest matching the initial framing path based on the initial framing path, and taking the target point of interest as the user framing intention; and correcting the initial framing path based on a framing path of the target point of interest to obtain the target framing path.

3. The method of claim 1, wherein, The identifying a user framing intention based on the initial framing path, and correcting the initial framing path based on the user framing intention to obtain a target framing path comprises: comparing the initial framing path with a plurality of candidate framing paths to determine a framing path matching the initial framing path, and taking the matching framing path as the user framing intention; and correcting the initial framing path based on the matching framing path to obtain the target framing path.

4. The method of claim 3, wherein, The comparing the initial framing path with a plurality of candidate framing paths to determine a framing path matching the initial framing path comprises: determining a framing path similarity between the initial framing path and each candidate framing path according to framing parameters of the initial framing path and framing parameters of the plurality of candidate framing paths; and determining a framing path matching the initial framing path from the plurality of candidate framing paths based on the framing path similarity.

5. The method of claim 4, wherein, The framing parameters comprise: a video frame framing position parameter; the determining a framing path similarity between the initial framing path and each candidate framing path according to framing parameters of the initial framing path and framing parameters of the plurality of candidate framing paths comprises: determining a video frame framing position similarity between the video frame framing position parameter of each candidate framing path and the video frame framing position parameter of the initial framing path; and calculating the framing path similarity with the initial framing path based on a number of video frames satisfying a first predetermined condition in each candidate framing path, wherein the first predetermined condition comprises: the video frame framing position similarity is greater than a first preset threshold.

6. The method of claim 5, wherein, The framing parameters further comprise: a video frame framing field angle parameter; the determining a framing path similarity between the initial framing path and each candidate framing path according to framing parameters of the initial framing path and framing parameters of the plurality of candidate framing paths further comprises: determining a video frame framing field angle similarity between the video frame framing field angle parameter of each candidate framing path and the video frame framing field angle parameter of the initial framing path; and calculating the framing path similarity with the initial framing path based on a number of video frames satisfying a second predetermined condition in each candidate framing path. The second predetermined condition comprises that the video frame framing position similarity is greater than the first preset threshold and the framing field of view similarity is less than a second preset threshold.

7. The method of claim 4, wherein, The method further comprises: The method further comprises: The method further comprises:

8. The method of claim 3, wherein, The method further comprises: The method further comprises:

9. The method of claim 8, wherein, The method further comprises: The method further comprises:

10. The method according to any one of claims 1 to 9, characterized in that, The method further comprises: The method further comprises: The method further comprises:

11. The method according to any one of claims 1 to 10, characterized in that, The method further comprises: The method further comprises: The method further comprises:

12. The method of claim 11, wherein, The method further comprises: The method further comprises:

13. The method of claims 1-12, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises:

14. The method of claims 1-13, wherein, The method further comprises: The method further comprises:

15. The method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises:

16. The method of claim 1, wherein, The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises: The method further comprises determine a first video corresponding to the initial framing path and a second video corresponding to the target framing path; compare and display the first video and the second video.

17. An electronic device comprising: a memory, a processor, and a display, when the computer program instructions are executed by the processor, the processor is configured to obtain an initial framing path based on user interaction, identify a user framing intention based on the initial framing path, and correct the initial framing path based on the user framing intention to obtain a target framing path.

18. The electronic device of claim 17, wherein: the processor is further configured to determine a target point of interest that matches the initial framing path based on the initial framing path, and take the target point of interest as the user framing intention; and correct the initial framing path based on a framing path of the target point of interest to obtain the target framing path.

19. The electronic device of claim 17, wherein: the processor is further configured to compare the initial framing path with a plurality of candidate framing paths, determine a framing path that matches the initial framing path, and take the matching framing path as the user framing intention; and correct the initial framing path based on the matching framing path to obtain the target framing path.

20. The electronic device of claim 19, wherein: the processor is further configured to determine a framing path similarity between the initial framing path and each candidate framing path based on framing parameters of the initial framing path and framing parameters of the plurality of candidate framing paths, and determine the framing path that matches the initial framing path from the plurality of candidate framing paths based on the framing path similarity.

21. The electronic device of claim 20, wherein: the framing parameters include a per-video-frame framing position parameter; the processor is further configured to determine a similarity between a framing position parameter of each candidate framing path and a framing position parameter of the initial framing path on a same video frame to obtain a video-frame framing position similarity, and calculate the framing path similarity with the initial framing path based on a number of video frames that satisfy a first predetermined condition in each candidate framing path, wherein the first predetermined condition includes that the video-frame framing position similarity is greater than a first preset threshold.

22. The electronic device of claim 21, wherein: the framing parameters further include a per-video-frame framing field of view parameter; the processor is further configured to determine a similarity between a framing field of view parameter of each candidate framing path and a framing field of view parameter of the initial framing path on a same video frame to obtain a framing field of view similarity, and calculate the framing path similarity with the initial framing path based on a number of video frames that satisfy a second predetermined condition in each candidate framing path, wherein the second predetermined condition includes that the video-frame framing position similarity is greater than the first preset threshold and the framing field of view similarity is less than a second preset threshold.

23. The electronic device of claim 21, wherein: The processor is further configured to determine, based on a frame extraction strategy, video frames to be processed from the multiple candidate framing paths and the initial framing path respectively, and determine the framing path similarity according to framing parameters corresponding to the video frames to be processed. 24.The electronic device of claim 19, wherein: The processor is further configured to replace the initial framing path with the matched framing path to obtain the target framing path. 25.The method of claim 24, wherein: The processor is further configured to perform smoothing processing on framing paths between adjacent interest points in the multiple interest points when it is detected that the target framing path includes the multiple interest points. 26.The electronic device of any one of claims 17-25, wherein: The processor is further configured to slice the panoramic video into multiple video segments according to a predetermined time length, identify a user framing intention corresponding to each video segment based on an initial framing path corresponding to each video segment, correct the initial framing path corresponding to each video segment based on the user framing intention corresponding to each video segment to obtain a target framing path corresponding to each video segment. 27.The electronic device of any one of claims 17-26, wherein: The processor is further configured to perform target detection on the panoramic video to obtain at least one interest point, perform target tracking on the panoramic video based on the at least one interest point to generate an interest point framing path corresponding to the at least one interest point, and take the interest point framing path as the candidate framing path. 28.The electronic device of claim 27, wherein: The processor is further configured to input the at least one interest point and the panoramic video into a trained convolutional neural network to output an interest point framing path corresponding to the at least one interest point. 29.The electronic device of any one of claims 17-28, wherein: The processor is further configured to obtain an initial framing path in response to a framing operation on the panoramic video. The framing operation includes at least one of: a body sensing framing operation, a virtual joystick framing operation, a predetermined direction framing operation, and a key frame dotting framing operation. 30.The electronic device of any one of claims 17-29, wherein: The processor is further configured to establish a wireless communication connection with a gimbal in response to detecting that the gimbal interacts with the terminal device, and receive a candidate framing path for the panoramic video from the gimbal.

31. The electronic device of claim 17, wherein: The electronic device further includes: a display configured to display prompt information indicating whether the initial framing path needs to be corrected before the initial framing path is corrected based on the user framing intention to obtain a target framing path; the processor is further configured to enter the step of correcting the initial framing path based on the user framing intention after obtaining a confirmation operation of the user.

32. The electronic device of claim 17, wherein: The electronic device further includes: The processor is further configured to determine a first video corresponding to the initial framing path and a second video corresponding to a target framing path after correcting the initial framing path based on the user framing intention to obtain the target framing path. The display is further configured to compare and display the first video and the second video.

33. A computer-readable storage medium, characterized in that, The non-transitory computer-readable medium has stored thereon executable instructions that, as a result of being executed by a processor, cause the processor to implement the method of any of claims 1-16.

Citation Information

Patent Citations

  • Target tracking and displaying method and device in panoramic video

    CN105843541A

  • Panoramic video processing method, device and equipment and storage medium

    CN111182218A

  • Panoramic video playing method and device, computer equipment and storage medium

    CN112954443A

  • View angle prediction method and device, equipment and storage medium

    CN115756158A

  • Information processing system, information processing apparatus, storage medium having stored therein information processing program, and information transmission / reception method

    US20140152764A1