Video generation method and device, computer equipment and storage medium
By determining the framing screen and adjusting the framing state when shooting large venues such as basketball courts, the problem of collecting the entire field motion range in the prior art is solved, and the efficiency and flexibility of video generation are improved.
Patent Information
- Application Number
- CN202311541121.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-16
AI Technical Summary
When shooting large venues such as basketball courts, the prior art is limited by the camera's viewing angle, making it difficult to collect the entire range of motion, resulting in cumbersome video processing and inefficient efficiency.
By determining the framing screen of the to-process video at the viewing angle, and adjusting the framing state according to the existence status of the transition event, collecting the target screen, and finally generating the target video.
It improves the efficiency of video generation, can flexibly change the framing state, adapt to the shooting needs of events such as passing and transition, and the generated target video reflects the content shot by multiple cameras more efficiently.
Smart Images

Figure CN120017970A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a video generation method, apparatus, computer equipment, storage medium and computer program product. Background Art
[0002] When shooting large venues such as basketball courts, the camera's field of view is limited. It is difficult to capture large-scale movements such as the entire court's range of motion with a single camera, and the target video cannot be obtained.
[0003] In traditional technology, manual camera movement is required to capture images of events such as passing and transitions, and the content shot by multiple cameras needs to be stitched into a target video in order to see the large-scale motion process through the target video. The video processing process of this shooting method is relatively cumbersome, and the processing efficiency needs to be improved. Summary of the invention
[0004] Based on this, it is necessary to provide a video generation method, apparatus, computer device, computer-readable storage medium and computer program product that can improve efficiency in response to the above technical problems.
[0005] In a first aspect, the present application provides a video generation method. The method comprises:
[0006] Determine the framing screen of the video to be processed under the framing perspective;
[0007] Determining a framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture;
[0008] Capturing the picture of the video to be processed according to the framing state to obtain a target picture;
[0009] A target video is generated according to the target picture.
[0010] In one of the embodiments, determining the framing picture of the video to be processed under the framing viewing angle includes:
[0011] Detecting a reference object in the video to be processed;
[0012] Determining a key framing angle for capturing an image of the reference object according to the position of the reference object;
[0013] According to the key framing angle, the video to be processed is captured to obtain a framing picture.
[0014] In one of the embodiments, the framing picture includes a sampling framing picture and an interpolation framing picture;
[0015] The step of determining a framing picture of the video to be processed under a framing viewing angle includes:
[0016] Based on each sampling moment, performing framing angle detection on the picture sampled from the video to be processed to obtain a sampled framing angle;
[0017] According to the sampling framing angle, the video to be processed is captured to obtain the sampling framing images at each sampling moment;
[0018] Performing interpolation calculation on the sampling framing angles at the sampling moments to obtain interpolated framing angles between the sampling moments;
[0019] At the non-sampling moments between the sampling moments, the video to be processed is captured according to the interpolation framing angle of view to obtain the interpolation framing pictures at the non-sampling moments.
[0020] In one of the embodiments, determining the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture includes:
[0021] If the framing state corresponding to the framing picture is a half-scene state and a transition initiation event of the framing picture is identified, adjusting the framing state to a transition state;
[0022] If the framing state corresponding to the framing picture is a transition state and a transition end event of the framing picture is identified, the framing state is adjusted to a half-field state.
[0023] In one of the embodiments, the framing perspective includes at least two key framing perspectives and a transition framing perspective, and the transition framing perspective is a framing perspective between the at least two key framing perspectives;
[0024] The step of determining a framing picture of the video to be processed under a framing viewing angle includes:
[0025] Under the transition framing angle, determining a transition framing screen of the video to be processed;
[0026] The determining the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture includes:
[0027] When the framing state of the transition framing picture is a transition state, determining whether a key framing angle matching the transition framing angle exists;
[0028] If yes, the framing state corresponding to the transition framing picture is adjusted to a half-scene state.
[0029] In one of the embodiments, the at least two key framing angles include a framing angle to be matched in a transition direction, and the framing angle to be matched is a framing angle according to the framing point to be matched;
[0030] The determining whether a framing angle matching the transition framing picture exists includes:
[0031] Determine the deviation distance between the viewing angle center of the transition viewing angle and the viewing angle to be matched;
[0032] If the deviation distance satisfies the matching condition, then the key framing angle matching the transition framing angle exists;
[0033] If the deviation distance does not satisfy the matching condition, then the key framing angle that matches the transition framing angle does not exist.
[0034] In one of the embodiments, the target picture includes a close-up picture;
[0035] The step of collecting the picture of the video to be processed according to the framing state to obtain the target picture includes:
[0036] When the framing state of the framing picture is a half-field state, detecting a close-up event based on a key framing angle of view;
[0037] If it is detected that the target object triggers the close-up event, adjusting the key framing angle according to the position of the target object to obtain a close-up framing angle;
[0038] The frame of the video to be processed is continuously captured according to the close-up framing angle to obtain the close-up frame, until the close-up frame satisfies the end condition of the close-up event, and the continuous capture of the frame of the video to be processed according to the close-up framing angle is stopped.
[0039] In one of the embodiments, the step of collecting the frame of the video to be processed according to the framing state to obtain the target frame includes:
[0040] When the framing state is a half-field state, detecting the moving direction and moving speed of the moving object group based on the framing picture;
[0041] According to the moving direction and the moving speed, the pan / tilt head is controlled to collect the picture of the video to be processed to obtain a half-field picture.
[0042] In one of the embodiments, after controlling the pan / tilt head to capture the image of the video to be processed according to the moving direction and the moving speed and obtaining the half-field image, the method further comprises:
[0043] If it is determined according to the half-field picture that the group of moving objects arrives at different half-field areas of the current scene, and the moving speed exceeds a transition threshold, the framing state is adjusted to a transition state.
[0044] In one embodiment, the target screen includes a transition screen; and determining the framing screen of the video to be processed under the framing viewing angle includes:
[0045] When the framing state is a transition state, collecting the picture of the video to be processed to obtain a transition framing picture;
[0046] The step of collecting the picture of the video to be processed according to the framing state to obtain the target picture includes:
[0047] Based on at least one of the picture acquisition objects in the moving object group and the target sphere in the transition framing picture, the picture acquisition is performed on the video to be processed to obtain a transition picture.
[0048] In one embodiment, the step of performing image acquisition on the video to be processed based on at least one image acquisition object of a group of moving objects and a target sphere in the transition framing image to obtain a transition image includes:
[0049] Determining a moving object group according to the density of moving objects in the transition framing picture;
[0050] According to the positions of the moving object group, the images of the video to be processed are captured to obtain transition images.
[0051] In one of the embodiments, the framing angle corresponding to the transition state further includes a transition framing angle;
[0052] The method of performing picture acquisition on the video to be processed based on at least one picture acquisition object among the moving object group and the target sphere in the transition framing picture to obtain the transition picture includes:
[0053] Determine the corresponding pan / tilt speed according to the movement speed of the image acquisition object;
[0054] Adjusting the transition framing angle of the transition framing picture according to the pan / tilt speed to obtain an adjusted transition framing angle;
[0055] The picture of the video to be processed under the adjusted transition viewing angle is collected to obtain a transition picture.
[0056] In one of the embodiments, the field of view of the video to be processed is greater than the field of view of the target video.
[0057] In a second aspect, the present application also provides a video generation device. The device comprises:
[0058] A framing module, used to determine the framing screen of the video to be processed under the framing perspective;
[0059] A detection module, used for determining a framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture;
[0060] A target picture acquisition module, used to acquire the picture of the video to be processed according to the framing state to obtain a target picture;
[0061] The video generation module is used to generate a target video according to the target picture.
[0062] In a third aspect, the present application also provides a handheld gimbal, comprising a motor and a processor, wherein the motor is used to control the rotation of the gimbal, and the processor implements the steps of video generation in any of the above embodiments when executing the computer program.
[0063] In a fourth aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of video generation in any of the above embodiments are implemented.
[0064] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of video generation in any of the above embodiments are implemented.
[0065] In a sixth aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of video generation in any of the above embodiments are implemented.
[0066] The above-mentioned video generation method, device, computer equipment, storage medium and computer program product determine the framing picture of the video to be processed under the framing angle of view, so that the original video to be processed is used as the source of the target picture to obtain the required framing angle of view; determine the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture, and collect the picture of the video to be processed according to the framing state to obtain the target picture, so that the subsequent target pictures of the framing picture can be in a different framing state from the framing picture, so that the framing state can be flexibly changed, so as to be able to collect pictures for events such as passing the ball and transitioning, and then generate a target video according to the target picture, and the content shot by multiple camera positions can be reflected through the target picture, so as to obtain the target video more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 A diagram showing an application environment of a video generation method in an embodiment;
[0068] Figure 2 is a flow chart of a video generation method in one embodiment;
[0069] Figure 3 A schematic diagram of key viewing points and key viewing angles in one embodiment;
[0070] Figure 4 is a schematic diagram of a video to be processed in one embodiment;
[0071] Figure 5 is a schematic diagram of a video to be processed in another embodiment;
[0072] Figure 6 A schematic diagram of determining a framing state corresponding to a framing picture in an embodiment;
[0073] Figure 7 is a structural block diagram of a video generating device in one embodiment;
[0074] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0076] The video generation method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 can be, but is not limited to, various cameras, video cameras, panoramic cameras, sports cameras, personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The terminal 102 can be fixed to the pan-tilt body by welding or the like, and can also be detachably connected or rotatably connected to the pan-tilt body.
[0077] In one embodiment, Figure 2 As shown, a video generation method is provided, which is applied to Figure 1 The terminal 102 in the example is used as an example to illustrate, and the following steps are included:
[0078] Step 202: determine the framing image of the video to be processed under the framing viewing angle.
[0079] The video to be processed is a video used for image extraction processing, and at least part of the images in the video to be processed can be captured through a framing angle. Optionally, the images of the video to be processed can be arranged in chronological order, or can be spliced from multiple groups of video clips. Specifically, the images in the venue can be captured by a video capture device, and then the images captured by the video capture device can be arranged in chronological order as the video to be processed; multiple groups of captured video clips can be spliced as needed to obtain the video to be processed.
[0080] The viewing angle is the viewing angle used to capture the picture from the video to be processed. The viewing angle can be determined based on two variables: the viewing point and the field of view (FOV) used for framing. Optionally, the viewing point and the field of view each include at least two dimensions, and the viewing point and the field of view in different time periods can be averaged in sequence to avoid jitter of the framing picture and sudden changes in the field of view. Optionally, the two dimensions (x, y) of the viewing point and the two dimensions (hfov, vfov) of the framing FOV can be respectively averaged in time sequence to obtain the viewing point and the field of view in each time period; wherein, in the two dimensions (x, y) of the viewing point, x is the position change of the image in the left and right direction, and y is the position change of the image in the up and down direction; in the two dimensions (hfov, vfov) of the framing FOV, hfov (Horizontal Field Of View) is the field of view range in the horizontal direction, and vfov (Vertical Field Of View) is the field of view range in the vertical direction. Optionally, sampling may be performed at time intervals, and interpolation processing of the viewing angle and the field of view angle may be performed based on the sampled images.
[0081] Optionally, the framing perspective may include at least one of a reference framing perspective and a key framing perspective, the reference framing perspective is a perspective of a certain reference angle, and the key framing perspective is a perspective for framing a key point of view, and the reference framing perspective and the key framing perspective may be the same perspective or different perspectives. The framing perspective also includes a transition framing perspective, which is a perspective in a transition state; optionally, the transition framing perspective may be a perspective between different reference framing perspectives, and the transition framing perspective may also be a perspective between different key framing perspectives.
[0082] The framing picture is a picture obtained by converting the picture of the video to be processed according to the framing angle. Optionally, the video to be processed is a video captured according to the angle of view of the video to be processed, and the framing angle and the angle of view of the video to be processed are different angles, so that the framing picture is different from the picture in the video to be processed.
[0083] In one embodiment, determining a framing picture of a video to be processed under a framing perspective includes: determining a plurality of sampling moments divided by time intervals in the video to be processed; determining sampling framing perspectives of the plurality of sampling moments; performing interpolation calculations based on the sampling framing perspectives at each sampling moment to obtain interpolated framing perspectives at each non-sampling moment; and performing screen capture on the video to be processed based on the sampling framing perspectives at each sampling moment and the interpolated framing perspectives at each non-sampling moment to obtain a framing picture.
[0084] Step 204: determining the framing state corresponding to the framing picture according to the transition event existence state of the framing picture.
[0085] The transition event existence state is the triggering condition of the transition event. The transition event is an event used to determine whether to transition. Optionally, the transition event includes a transition initiation event and a transition end event, the transition initiation event is used to characterize the start of the transition process, and the transition end event is used to characterize the end of the transition process. Optionally, the transition event of the framing state can be determined according to the framing state corresponding to the framing screen.
[0086] The framing state is a state determined based on the picture detection result of the framing picture. Optionally, the framing state is a state of performing picture acquisition in the current scene in the video to be processed. Optionally, the framing state includes a half-field state and a transition state; the half-field state is a state of performing picture acquisition in a certain half-field area of the current scene, and the transition state is a state of performing picture acquisition during the conversion process of different half-field areas of the current scene.
[0087] Optionally, when the framing state is the half-time state, it is determined whether the framing state is adjusted to the transition state according to the triggering of the transition initiation event. Optionally, when the framing state is the transition state, it is determined whether the framing state is adjusted to the half-time state according to the triggering of the transition end event.
[0088] In one embodiment, determining the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture includes: if the framing state corresponding to the framing picture is a half-time state and a transition initiation event of the framing picture is identified, adjusting the framing state to a transition state. In order to more clearly illustrate the identification process of the transition initiation event, multiple feasible embodiments are described.
[0089] In a feasible embodiment, identifying a transition initiation event of a framing screen includes: when the framing state corresponding to the framing screen is a half-time state, identifying an action of the framing screen; if the action meets the steal event, identifying a transition initiation event.
[0090] Among them, the action of the framing screen is identified, and it is determined whether the action meets the steal event; the judgment can be made based on the video model, and the video model specifically refers to a deep learning model that takes multiple frames of framing screen as input, such as the Expand 3D spatiotemporal domain network (Expand 3D, X3D), the Slowfast network (Slowfast) and other models. Optionally, the human body key points of the video screen of the target object can be detected respectively to obtain the skeleton information of multiple frames; the action recognition is performed based on the skeleton information of multiple frames in time sequence to obtain the action type; if the action type that meets the steal action is detected, the transition initiation event is identified. Taking the basketball scene as an example, the action of the target object is identified based on the sequence of multiple frames of framing screen. If it is identified that the action of the target object belongs to any jumping action with one foot suspended in the air, both feet suspended in the air, or both feet not suspended in the air, it is identified that the action of the framing screen meets the steal event.
[0091] In another feasible embodiment, identifying a transition initiation event of the framing screen includes: when the framing state corresponding to the framing screen is a half-field state, detecting the movement trajectory of the target sphere based on the framing screen; if the movement trajectory passes through different half-field areas of the current field, then identifying a transition initiation event.
[0092] In another feasible embodiment, a transition initiation event of a framing screen is identified, including: when the framing state corresponding to the framing screen is a half-field state, and the moving direction and movement speed of the moving object group are obtained, after capturing the target screen, if it is determined according to the target screen that the moving object group arrives at different half-field areas of the current scene, and the movement speed exceeds the transition threshold, then a transition initiation event is identified.
[0093] In another embodiment, the framing state corresponding to the framing picture is determined according to the existence state of the transition event of the framing picture, including: if the framing state corresponding to the framing picture is the transition state, and the transition end event of the framing picture is identified, the framing state is adjusted to the half-field state. In order to more clearly illustrate the identification process of the transition end event, two feasible embodiments are used for description.
[0094] In a feasible embodiment, identifying the transition end event of the framing screen includes: when the framing state of the framing screen is the half-time state, counting the number of moving objects in the framing screen of the key framing angle; if the number of moving objects detected in the framing screen of the key framing angle satisfies the number end transition condition, then identifying the transition end event of the framing screen. Wherein, satisfying the number end transition condition may mean that the number of moving objects in the framing screen is greater than the number end transition threshold, or may mean that the ratio between the number of moving objects in the framing screen and the total number of moving objects is greater than the ratio end transition threshold; wherein the moving object is at least one athlete.
[0095] In a feasible embodiment, identifying the transition end event of the framing picture includes: when the framing state corresponding to the transition picture is the transition state, determining whether a key framing angle matching the transition picture exists; if so, adjusting the framing state corresponding to the transition picture to the half-field state.
[0096] Step 206, capturing the frame of the video to be processed according to the framing state to obtain the target frame.
[0097] The target picture is obtained by capturing the picture of the video to be processed again after the framing picture is obtained. Optionally, after executing steps 202-206 for a certain time to obtain at least one frame of the target picture for generating the target video, the at least one frame of the target picture obtained this time can be used as the next execution of step 202, until the stop capture condition of the target picture is met, the target picture is stopped from being captured. Among them, the stop capture condition of the target picture can refer to that the number of target pictures is greater than the stop capture threshold, or that, in the video to be processed, the timestamp of capturing the picture of the video to be processed is greater than the stop capture time point. Exemplarily, if the pictures of the video to be processed at each moment have been captured, step 206 is stopped.
[0098] Optionally, the target picture includes at least one of a half-time picture and a transition picture. The half-time picture is a picture captured in a half-time state, and the transition picture is a picture captured in a transition state.
[0099] Optionally, the half-time picture includes at least one of a close-up picture and a key point half-time picture; the close-up picture is a close-up picture collected for the target object, and the key point half-time picture is a picture collected for the target object or the reference object at a key framing angle. Optionally, for the key framing angle, it can be a field of view under a wide angle, or a field of view in accordance with other schemes. Optionally, the target object includes picture collection objects such as a group of moving objects and a target sphere, and the viewing angle of the close-up picture and the key point half-time picture collection process can be controlled according to the position changes of the group of moving objects and the target sphere.
[0100] Optionally, the transition screen may include a screen captured according to a reference viewing angle, and also include a screen captured from a to-be-processed video based on at least one of a group of moving objects and a target sphere in a certain screen. Optionally, the reference field of view may be a field of view of a wide-angle lens, and the field of view of the wide-angle lens may be greater than or equal to 50 degrees and less than or equal to 70 degrees.
[0101] In one embodiment, the image of the video to be processed is captured according to the framing state to obtain the target image, including: when the framing state of the framing image is a half-field state, the image of the video to be processed is captured according to the wide-angle framing angle to obtain a half-field image. The half-field image captured in this way can be a close-up image or a key point half-field image.
[0102] In another embodiment, the target picture includes a close-up picture and a key point half-field picture. According to the framing state, the picture of the video to be processed is collected to obtain the target picture, including: when the framing state of the framing picture is the half-field state, if a close-up event is detected, the picture of the video to be processed is collected according to the close-up event to obtain the close-up picture; if no close-up event is detected, the picture of the video to be processed is collected according to the key framing point where the target object or the reference object is located to obtain the key point half-field picture.
[0103] In one embodiment, according to the framing state, the picture of the video to be processed is collected to obtain the target picture, including: when the framing state of the framing picture is a transition state, the picture of the video to be processed is collected according to the reference framing angle to obtain the transition picture. Thus, during the transition process, the picture is collected according to the reference framing angle.
[0104] In another embodiment, according to the framing state, the image of the video to be processed is collected to obtain the target image, including: according to the moving object or the target sphere in the transition framing image, the image of the video to be processed is collected to obtain the transition image. Thus, the transition control can be performed based on a certain moving object or a target sphere.
[0105] Optionally, the transition speed in the transition state may be a preset value, or may be estimated based on the moving object or the target sphere.
[0106] Step 208, generating a target video according to the target picture.
[0107] The target video is a video generated based on the target pictures obtained by performing steps 202-206 at least once. Optionally, multiple frames of target pictures can be spliced in chronological order to obtain the target video, or multiple frames of target pictures can be screened to combine target pictures belonging to a certain action type into the target video, or half-time pictures and transition pictures can be combined into the target video.
[0108] Optionally, a target video including only the transition picture can be generated based on the transition picture, a target video including only the half-court picture can be generated based on the half-court picture, and a target video including both the full-court picture and the half-court picture can be generated based on the half-court picture and the transition picture. Regardless of which of the three target videos is generated, there is no need to manually move the camera to capture pictures of events such as passing and transitions, and there is no need to splice the content shot by multiple camera positions into the target video.
[0109] Optionally, the field of view of the video to be processed is greater than the field of view of the target video.
[0110] The field of view of the video to be processed is the field of view of the video to be processed when it is displayed or data managed, and the field of view of the target video is the field of view of the target video when it is displayed. The above-mentioned display method can be playback or viewing according to timestamps. The above-mentioned data management can be a process of generating, exporting or saving as a video. Optionally, the video to be processed can be a panoramic video, a single wide-angle video, an ultra-wide-angle video, or a single fisheye video, and the target video is a flat video.
[0111] Optionally, the image of the video to be processed can be reprojected or cropped to obtain the target image. Optionally, when the video to be processed is a panoramic video, the image captured from the panoramic video can be reprojected to generate the target image; and the type of reprojection can be stereographic projection (Stereographic Projection) or azimuthal equidistant projection (Azimuthal equidistant). Optionally, when the video to be processed is a single wide-angle, ultra-wide-angle video, or a single fisheye video, the image of the video to be processed can be cropped or reprojected to obtain the target image.
[0112] When the field of view of the video to be processed is larger than that of the target video, the field of view covered by the video to be processed is larger than that of the target video, so that the video to be processed can capture more content, but the content covered will be distorted due to the field of view. The target video is generated based on the target picture obtained by capturing the picture of the video to be processed, and this distortion can be eliminated through image reprojection technology or image deformation technology to ensure that the video has high quality.
[0113] In one embodiment, Figure 3As shown, by performing backboard detection in the field, two backboards in opposite directions are detected. These two backboards are located in the dotted frame, and these two backboards are key viewing angles; based on these two key viewing angles, two key framing perspectives are determined. The key framing perspective is the perspective located in the solid frame. The solid frame represents the most typical framing frame. Most of the time, the players of both teams are attacking and defending in these two solid frames, which is called a half-court positional battle. The range between the two key viewing angles is the effective framing range. In a few cases, the process of players transitioning from one half court to another is generally called a transition. The left side of the left half court and the right side of the right half court are not effective framing ranges. That is to say, when exporting the edited plane video, we only focus on the content on the basketball court, so that the video of this part of the content is used as the framing video.
[0114] In an exemplary embodiment, Figure 4 The three types (a), (b), and (c) are explained based on the video to be processed. Figure 4 (a) is a situation where images are collected for a half-court positional battle, using key viewing angles and a wide-angle field of view as a reference field of view to collect images and obtain a half-court image; wherein the reference field of view angle may be 50-70 degrees. Figure 4 (b) in the figure is the case of collecting images for a half-court positional battle. When a close-up event such as a shot is detected in the half-court positional battle, the reference field of view is narrowed to the close-up field of view, and the images related to the close-up event are collected at the special effect field of view. For example, in the case of a close-up event such as a shot, the reference field of view is narrowed to within 30 degrees to obtain the close-up field of view. The field of view of the close-up field of view includes the person shooting, the basketball, and the basket. Figure 4 (c) in the figure is the moment of goal after shooting, which is the situation of collecting images of shooting, goal, etc. At this time, at the moment of goal after shooting, the FOV is narrowed to the basket area, and a close-up image of the target sphere is generated based on the relevant images when the goal is scored.
[0115] In another exemplary embodiment, Figure 5 In the transition state, the players of both sides move from the right half of the current field to the left half, and the viewing angle is determined based on the group of sports objects where most players are located, using a 70-degree wide-angle field of view.
[0116] In the above-mentioned video generation method, the framing picture of the video to be processed under the framing angle of view is determined, so that the original video to be processed is used as the source of the target picture to obtain the required framing angle of view; the framing state corresponding to the framing picture is determined according to the existence state of the transition event of the framing picture, and the picture of the video to be processed is collected according to the framing state to obtain the target picture, so that the subsequent target pictures of the framing picture can be in a different framing state from the framing picture, so that the framing state can be flexibly changed, so that the picture of events such as passing the ball and transition can be collected, and then the target video is generated according to the target picture. The target picture can reflect the content shot by multiple camera positions, so that the target video is obtained more efficiently.
[0117] In one embodiment, determining a framing picture of a video to be processed under a framing angle includes: detecting a reference object in the video to be processed; determining a key framing angle for capturing a picture of the reference object according to a position of the reference object; and capturing a picture of the video to be processed according to the key framing angle to obtain a framing picture.
[0118] The reference object is a reference object used to locate the key viewing point in the video to be processed. Optionally, the reference object can be a basket or a structure of the basket, or a certain identification symbol installed on the basket, or other objects set within the range of the basket.
[0119] The detection of the reference object can use common target detection methods, which can be detection methods based on manual features (such as template matching method, key point matching method, key feature method, etc.), or detection methods based on convolutional neural networks (CNN). In the detection method based on convolutional neural networks, the model used can be but not limited to You Only Look Once (YOLO), Single Shot MultiBox (SSD), Region-based Convolutional Neural Networks (R-CNN) or Mask Region-based Convolutional Neural Networks (Mask R-CNN) and other models, as long as the convolutional neural network can identify the features of the reference object.
[0120] In an optional embodiment, determining a key framing angle for capturing a picture of the reference object according to the position of the reference object includes: determining a key framing point in the picture according to the position of the reference object; and determining a key framing angle for capturing the picture toward the key framing point at a preset field of view angle. The preset field of view angle may be a reference field of view angle or a preset field of view angle specifically used for framing the picture at the key framing point.
[0121] In an optional embodiment, according to the key framing angle, the video to be processed is captured to obtain a framing picture, including: facing the key framing point, capturing the image of the video to be processed at a preset field of view angle to obtain a framing picture; or, during the conversion process between key framing points, continuously capturing the image of the video to be processed at a preset field of view angle to obtain a framing picture.
[0122] In this embodiment, a reference object in the video to be processed is detected, and a key framing angle for capturing images of the reference object is determined according to the position of the reference object; thereby, reference objects in a variety of venues can be obtained to avoid the limitation of a single venue, and key viewing angles can be determined to ensure that the application of this solution is highly efficient; and based on the key framing angle, images of the video to be processed are captured to obtain framing images, which makes the framing image capture process more flexible, and images can be captured in a direction toward key viewing angles, and the interval for capturing images can also be determined based on key viewing angles, so as to capture images flexibly and efficiently.
[0123] In one embodiment, the framing picture includes a sampled framing picture and an interpolated framing picture. The sampled framing picture is a picture that needs to be detected by framing angle of view to obtain a sampled framing angle of view, and is framed based on the sampled framing angle of view; the interpolated framing picture is a picture that is obtained by interpolating the sampled framing picture to obtain an interpolated framing angle of view, and then captured through the interpolated framing angle of view.
[0124] Determining the framing picture of the video to be processed under the framing perspective includes: based on each sampling moment, performing framing perspective detection on the picture sampled from the video to be processed to obtain a sampling framing perspective; according to the sampling framing perspective, performing picture acquisition on the video to be processed to obtain a sampling framing picture at each sampling moment; performing interpolation calculation on the sampling framing perspective at each sampling moment to obtain an interpolated framing perspective between each sampling moment; at a non-sampling moment between each sampling moment, performing picture acquisition on the video to be processed according to the interpolated framing perspective to obtain an interpolated framing picture at each non-sampling moment.
[0125] The sampling moment is used to determine the moment for sampling the video to be processed. When the sampling moment is determined, the non-sampling moment is also determined. Optionally, the sampling moment is also the moment for performing framing angle detection, and capturing the image of the video to be processed according to the sampling framing angle. Optionally, the sampling moment is a sampling moment calculated according to a preset interval. Optionally, each sampling moment can be the moment where a key frame is located. The picture obtained by sampling the video to be processed is the picture used for reference framing angle detection; optionally, the picture obtained by sampling the video to be processed is the picture at the sampling moment.
[0126] The sampling framing angle is the framing angle obtained by the framing angle detection, and each sampling moment may have its own sampling framing angle. The sampling framing picture is the framing picture collected according to the sampling framing angle. Optionally, the sampling framing picture can be a picture sampled based on the picture sampled from the video to be processed.
[0127] The interpolated framing angle is the angle obtained by interpolating the sampled framing angle, and each non-sampling moment between adjacent sampling moments may have its own non-sampling framing angle. The interpolated framing picture is the framing picture collected according to the interpolated framing angle.
[0128] Optionally, the sampling framing angle at each sampling moment is interpolated, including: performing linear interpolation calculation on the sampling framing angle at each sampling moment, or fitting the sampling framing angle at each sampling moment into a nonlinear curve to perform nonlinear interpolation calculation. Optionally, the interpolation calculation can take the sampling framing angle and the sampling moment as two-dimensional parameters of a point, so as to perform two-point linear calculation or multi-point linear calculation. Optionally, since each non-sampling moment is known, interpolation calculation can be performed based on each adjacent sampling moment and each sampling framing angle at each adjacent sampling moment to obtain the interpolated framing angle at each non-sampling moment.
[0129] In an optional implementation, based on each sampling moment, the frame obtained by sampling the video to be processed is detected for the framing angle, and before obtaining the sampling framing angle, it also includes: determining each sampling moment in the video to be processed according to the time interval since the end time of the video to be processed.
[0130] In one embodiment, based on each sampling moment, a framing angle detection is performed on the picture obtained by sampling the video to be processed to obtain a sampling framing angle, including: at each sampling moment, a viewing point and a field of view angle detection is performed on the picture obtained by sampling the video to be processed to obtain a sampling key viewing point and a sampling field of view angle for framing toward the sampling key viewing point.
[0131] Correspondingly, according to the sampling framing angle, the processed video is captured to obtain the sampling framing images at each sampling moment, including: toward each sampling key framing angle, the processed video is captured within the sampling field of view to obtain the sampling framing images at each sampling moment.
[0132] In one embodiment, after obtaining the sampled framing pictures at each sampling moment, the above-mentioned interpolation calculation is performed on the sampled framing angles at each sampling moment to obtain the interpolated framing angles between the sampling moments, including: for each sampling moment, performing interpolation position calculation on the sampling key viewing angles to obtain the interpolation position viewing angles between the sampling moments; for each sampling moment, performing interpolation calculation on the sampling field of view angles to obtain the interpolation field of view angles between the sampling moments.
[0133] Correspondingly, at the non-sampling moments between the sampling moments, the video to be processed is captured according to the interpolation framing angle to obtain the interpolation framing images of each non-sampling moment, including: at the non-sampling moments between the sampling moments, the image of the video to be processed within the interpolation field of view is captured toward the interpolation position framing angle to obtain the interpolation framing images of each non-sampling moment.
[0134] Take the sampling time as the sampling time point, the non-sampling time as the non-sampling time point, and use two-point linear interpolation to calculate the interpolation position for example. Assume that the non-sampling time point t+k is vacant, and its previous non-vacant sampling time point is t, and the position is p t , the next non-vacant time point is t+N, and the position is p t+N , then the position calculation formula at time point t+k is Among them, k and N are time intervals, and p represents the position.
[0135] Based on this, there is no need to perform viewing angle detection on each frame, but viewing angle detection is combined with key viewing angle detection, which saves time and allows for more efficient determination of the viewing image.
[0136] In one embodiment, Figure 6 As shown, the framing state corresponding to the framing picture is determined according to the existence state of the transition event of the framing picture, including: if the framing state corresponding to the framing picture is the half-field state and the transition initiation event of the framing picture is identified, the framing state is adjusted to the transition state; if the framing state corresponding to the framing picture is the transition state and the transition end event of the framing picture is identified, the framing state is adjusted to the half-field state.
[0137] In one embodiment, when the framing state is the half-time state, it is determined whether a transition initiation event is recognized; if so, the framing state is adjusted to the transition state; if not, the framing state continues to be the half-time state. Correspondingly, when the framing state is the transition state, it is determined whether a transition end event is recognized; if so, the framing state continues to be the transition state; if not, the framing state is adjusted to the half-time state.
[0138] Optionally, the video to be processed is analyzed from front to back in chronological order. The pictures are sampled at regular intervals for analysis. During the analysis, a perspective selection scheme based on a state machine is used, involving a half-time state and a transition state, a transition initiation event for jumping to a transition state, and a transition end event for jumping to a half-time state. In the half-time state, the half-time framing rules are executed to frame near the current key framing point; in the transition state, the framing rules of the transition process are executed to obtain a half-time framing picture.
[0139] Therefore, for the framing pictures at different times, the framing state can be changed from the half-time state to the transition state, and from the transition state to the half-time state. Through the alternating changes of these two framing states, different angles of shooting can be performed for events such as passing, transitions, or goals, so as to better ensure the acquisition efficiency of the target picture.
[0140] In one embodiment, a transition event in a transition state is described; the framing angle includes at least two key framing angles and a transition framing angle, and the transition framing angle is a framing angle between at least two key framing angles. Optionally, each key framing angle is a perspective for capturing images based on different key viewing angles, and the transition framing angle is a perspective for capturing images in the space between different key viewing angles. Optionally, the transition framing angle is a perspective between two adjacent key framing angles.
[0141] Determining the framing screen of the video to be processed under the framing perspective includes: determining the transition framing screen of the video to be processed under the transition framing perspective.
[0142] Correspondingly, the framing state corresponding to the framing picture is determined according to the existence state of the transition event of the framing picture, including: when the framing state of the transition framing picture is the transition state, judging whether there is a key framing angle matching the transition framing angle; if so, adjusting the framing state corresponding to the transition framing picture to the half-field state.
[0143] The transition framing picture is a framing picture corresponding to the transition state; the transition framing picture is a framing picture of the video to be processed under the transition framing perspective. Optionally, when the initialization state of the framing state is the transition state, the picture collected from the video to be processed is the transition framing picture; and the transition picture obtained by the last collection can be used as the transition framing picture.
[0144] In an optional implementation, determining whether a key framing angle that matches a transition framing angle exists includes: determining a transition framing angle of a transition framing picture when it is captured; calculating a degree of match between the transition framing angle and any key framing angle; if the degree of match is greater than a matching threshold, determining that a key framing angle that matches the transition framing angle exists; if the degree of match is less than the matching threshold, determining that a key framing angle that matches the transition framing angle does not exist.
[0145] The framing state corresponding to the transition framing picture is adjusted to the half-field state, so as to adjust the framing state and collect the picture of the video to be processed according to the half-field state to obtain the half-field picture, thereby realizing the switching of different framing states.
[0146] When the framing state is a transition state, the framing point of the transition framing perspective changes under different key framing perspectives, so as to capture the transition picture from the video to be processed; and the obtained transition picture can be used as a transition framing picture to determine whether the framing state is adjusted, so as to determine the way or means to capture the picture of the video to be processed after obtaining the transition framing picture, so as to capture the picture for events such as passing the ball and transition.
[0147] In one embodiment, at least two key framing angles include a framing angle to be matched in a transition direction, and the framing angle to be matched is a angle for capturing images according to the framing angle to be matched.
[0148] Judging whether a framing perspective that matches a transition framing perspective exists includes: determining a deviation distance between a perspective center of the transition framing perspective and a framing point to be matched; if the deviation distance satisfies a matching condition, a key framing perspective that matches the transition framing perspective exists; if the deviation distance does not satisfy the matching condition, a key framing perspective that matches the transition framing perspective does not exist.
[0149] The transition direction is the direction that matches the rotation direction of the gimbal in the transition state. If the gimbal rotates to the left, the transition direction is to the left side of the gimbal; if the gimbal rotates to the right, the transition direction is to the right side of the gimbal. Optionally, the above-mentioned gimbal can be a real gimbal or a virtual gimbal. The virtual gimbal is a position for capturing images, and the orientation and field of view of the position are determined by the viewing angle. Since the video to be processed exists, even if the virtual gimbal is not used for image capture, the terminal can determine the corresponding image based on the video to be processed and the position of the virtual gimbal. Therefore, the key viewing angle can be determined without capturing the image toward the key viewing angle through the virtual gimbal.
[0150] Optionally, the framing angle to be matched is a key framing angle located in the transition direction in the transition state. The framing angle to be matched is used to match the transition framing angle. The framing angle to be matched is obtained by framing based on the framing point to be matched.
[0151] Optionally, in the case where there are two key viewing angles, if the pan / tilt rotates to the left and the transition direction is to the left, the viewing angle to be matched is the key viewing angle on the left side of the pan / tilt; if the pan / tilt rotates to the right and the transition direction is to the right, the viewing angle to be matched is the key viewing angle on the right side of the pan / tilt. Optionally, in the case where there are multiple key viewing angles, if the pan / tilt rotates to the left and the transition direction is to the left, the viewing angle to be matched is located to the left of the key viewing angle most recently used for picture acquisition; if the pan / tilt rotates to the right and the transition direction is to the right, the viewing angle to be matched is located to the right of the key viewing angle most recently used for picture acquisition.
[0152] The perspective center of the transition framing perspective is the framing center during the transition process. Optionally, the perspective center of the transition framing perspective is the center position of the transition framing screen, which can be determined by the pixel coordinates in the transition framing screen, and the center position can also be determined based on the center position of the object in the transition framing screen; optionally, the perspective center of the transition framing perspective can be determined according to the position of the moving object group or the target sphere. Optionally, the perspective center of the transition framing perspective is a certain viewing point other than the key viewing point.
[0153] The deviation distance is used to characterize the distance difference between the viewing angle center of the transition framing picture and the viewing angle to be matched. Optionally, the distance difference can be the coordinate difference or coordinate ratio between the viewing angle center and the viewing angle to be matched, or it can be the distance difference between the viewing angle center and the viewing angle to be matched. The matching condition is a matching indicator defined for the deviation distance. When the deviation distance is a coordinate difference or a coordinate ratio, the matching condition is an indicator defined for the coordinate difference or the coordinate ratio; when the deviation distance is a distance difference, the matching condition is an indicator defined for the distance difference.
[0154] In an exemplary implementation, the deviation distance is a coordinate difference; when the deviation distance is less than a distance difference threshold, the deviation distance satisfies the matching condition; when the deviation distance is greater than the distance difference threshold, the deviation distance does not satisfy the matching condition.
[0155] In a specific embodiment, at least two key framing perspectives include a first key framing perspective and a second key framing perspective, wherein the first key framing perspective is a perspective for framing according to the first key viewing angle, and the second key framing perspective is a perspective for framing according to the second key viewing angle, and the second key viewing angle is a framing perspective to be matched in the transition direction, and the second key viewing angle is a framing perspective to be matched. At this time, it is a case of converting from the first key framing perspective to the second key framing perspective, and the framing perspective from the first key framing perspective to the second key framing perspective is a transition framing perspective.
[0156] At this time, judging whether there is a matching framing angle that matches the transition picture includes: determining the deviation distance between the perspective center of the transition picture and the second key viewing angle; if the deviation distance satisfies the matching condition, determining that the transition picture matches the framing picture of the transition end perspective; if the deviation distance does not satisfy the matching condition, determining that the transition picture does not match the framing picture of the transition end perspective.
[0157] Based on this, whether the transition is completed is judged by whether the deviation distance between the perspective center of the transition framing perspective and the to-be-matched viewing point meets the matching conditions; when the matching conditions are met, it can be determined that the camera has reached the half of the scene where the to-be-matched viewing point is located, and the data processed is relatively small, and the processing efficiency is relatively high.
[0158] In one embodiment, the target picture includes a close-up picture. According to the framing state, the picture of the video to be processed is collected to obtain the target picture, including: when the framing state of the framing picture is a half-field state, a close-up event is detected based on the key framing angle; if it is detected that the target object triggers the close-up event, the key framing angle is adjusted according to the position of the target object to obtain the close-up framing angle; the picture of the video to be processed is continuously collected according to the close-up framing angle to obtain the close-up picture, until the close-up picture meets the end condition of the close-up event, and the continuous collection of the picture of the video to be processed according to the close-up framing angle is stopped.
[0159] A close-up event is an event detected for a target object. Optionally, a close-up event is determined based on the detection result of the target object. Optionally, the target object can be determined to have triggered a close-up event based on the action type that the target object's action conforms to or the area where the target object's position is located. Optionally, if the target object is detected to have performed a key action that can be completed within half court, such as shooting or blocking, it is determined that the target object has triggered a close-up event. Optionally, the close-up event can be obtained through detection of the video to be processed, or it can be determined based on the framing screen. Optionally, the target object can be a target moving object or a target sphere.
[0160] The close-up framing angle is a perspective set for a close-up event. Optionally, different close-up events are detected for different target objects. The close-up framing angle is a perspective for capturing images at the position of the target object. Optionally, the close-up framing angle is a perspective obtained by reducing the field of view of the key framing angle.
[0161] The end condition of the close-up event is an indicator set for the detection result of the close-up picture. Optionally, if the close-up picture detected indicates that the current close-up event is over, the close-up picture satisfies the end condition of the close-up event; if the close-up picture detected indicates that the current close-up event is not completed, the close-up picture does not satisfy the end condition of the close-up event. Optionally, the end condition is used to determine whether to re-detect the close-up event.
[0162] In an optional implementation, adjusting the key framing angle according to the position of the target object to obtain the close-up framing angle includes: adjusting the viewing angle of the key framing angle to the position of the target object to obtain the close-up framing angle.
[0163] In another optional embodiment, the key framing angle is adjusted according to the position of the target object to obtain a close-up framing angle, including: during the time period when the shooting action is performed, adjusting the viewing angle of the key framing angle to the position of the target moving object to obtain a close-up framing angle of the target moving object; during the time period after the shooting action is performed, adjusting the viewing angle of the key framing angle to the position of the target sphere to obtain a close-up framing angle of the target sphere.
[0164] In one embodiment, the close-up event is a shooting event, and the close-up framing angle is first the close-up framing angle of the target moving object, and then the close-up framing angle of the target sphere. At this time, part of the close-up picture contains the target moving object, and the other part of the close-up picture contains the target sphere. Specifically, the picture of the video to be processed is continuously collected according to the close-up framing angle until the close-up picture meets the end condition of the close-up event, including: according to the close-up framing angle of the target moving object, the picture of the target moving object shooting is collected from the video to be processed; then according to the close-up framing angle of the target sphere, the picture of the target sphere is collected from the video to be processed until the target sphere drops to a preset height or a preset area in the picture of the target sphere.
[0165] Optionally, after stopping continuously capturing the images of the video to be processed according to the close-up framing angle, the method further includes: capturing the images of the video to be processed according to the key framing angle to obtain a key point half-field image.
[0166] Based on this, when the framing state of the framing screen is half-time, close-up events are detected based on the key framing angle to achieve real-time monitoring of close-up events; if a target object is detected to trigger a close-up event, the key framing angle is adjusted according to the position of the target object to obtain a close-up framing angle; the screen of the video to be processed is continuously collected according to the close-up framing angle to obtain a close-up screen, until the close-up screen meets the end condition of the close-up event, and the continuous collection of the screen of the video to be processed according to the close-up framing angle is stopped, so that the close-up screen related to the close-up event can be collected completely. In this way, close-up screens are collected for target spheres such as basketballs and footballs, and sports objects such as athletes, so that such screens do not need to be manually spliced, making the generation efficiency of target videos higher.
[0167] In one embodiment, the images of the video to be processed are captured according to the framing state to obtain a target image, including: when the framing state is a half-field state, detecting the moving direction and speed of the moving object group based on the framing image; and controlling the pan / tilt head to capture the images of the video to be processed according to the moving direction and speed to obtain a half-field image.
[0168] The motion object group is a group of objects that can control the target ball to move; the motion object group includes multiple motion objects. The motion object group can control the target ball to move by performing certain actions, or enable the target ball to trigger corresponding ball events.
[0169] In one embodiment, detecting the moving direction and speed of a moving object group based on a framing image includes: detecting the motion trajectory of each moving object in the moving object group based on multiple frames; and determining the moving direction and speed of the moving object group based on the motion trajectory of each moving object.
[0170] The motion trajectory of each moving object is a position sequence of each moving object arranged in time sequence; the position sequence can be a coordinate sequence or a picture area sequence, the picture area sequence is a picture area arranged in time sequence, and each picture area is used to represent the position of each moving object in the reference picture at a certain moment.
[0171] The moving direction and the moving speed are motion information generated based on the motion information of the moving objects in the moving object group; the moving direction and the moving speed can be determined based on the statistical results of the position change information of the moving objects within a certain time period, and the moving direction and the moving speed can also be determined based on the group position change information after the group position change information is determined based on the position change information of the moving objects.
[0172] In this implementation, the motion trajectory of each moving object in the moving object group is detected based on multiple frames of images, so that the multi-target detection method can be directly applied to the motion trajectory detection of each moving object, and the recognition speed and accuracy are guaranteed by a relatively mature detection method; on this basis, the group motion information of the moving object group is determined according to the motion trajectory of each moving object, and the group motion information is obtained by further performing data statistics through the motion trajectory of each moving object; and the group motion information can reflect the movement trend of the target sphere, and will not be affected by occlusion. The accuracy is affected, so as to better ensure the stability of image acquisition.
[0173] In another embodiment, the moving direction and speed of a group of moving objects are detected based on a framing picture, including: performing optical flow estimation on the group of moving objects based on multiple frames to obtain the optical flow displacement of each frame; accumulating the optical flow displacement of each frame to obtain a picture motion feature map corresponding to the multiple frames; the picture motion feature map includes multiple feature map pixels, and the direction of each feature map pixel is determined by the accumulation result of the optical flow displacement of each frame; according to the direction of each feature map pixel, the number of pixels of the feature map pixels displaced along the opposite transition direction are counted respectively to obtain the number of pixels in the opposite transition direction; determining a comparison result between the number of pixels in the opposite transition direction; when the comparison result meets the transition situation of the transition type, obtaining the moving direction and the displacement in the moving direction; determining the moving speed according to the ratio between the displacement in the moving direction and the optical flow estimation time period.
[0174] Optical flow displacement is the result of optical flow estimation based on multiple frames, and can be used to characterize the motion trend of a group of moving objects. The picture motion feature map is the cumulative result of the optical flow displacement of multiple frames in a certain interval. The feature map pixel is a pixel in the picture motion feature map, and each feature map pixel is used to characterize the optical flow estimation result in a certain interval. The direction of each feature map pixel is the optical flow direction; optionally, if the optical flow displacement is set along the transition direction, the direction of the feature map pixel is set along a certain transition direction. Opposite transition directions are at least one pair of opposite transition directions.
[0175] The comparison result between the numbers of pixels in opposite transition directions can be determined based on the ratio of the numbers of pixels in opposite transition directions, or based on the difference of the numbers of pixels in opposite transition directions, and the difference of the number of pixels has higher accuracy. For example, in a 1080p reference picture, if there are 1000 pixels to the right and 1 pixel to the left, the ratio will reach 1000. The difference in the number of pixels can represent how many pixels have changed, the ratio of the number of pixels cannot reflect the degree of such change, and the difference in the number of pixels can reflect the corresponding degree of change.
[0176] In this embodiment, the optical flow calculation method can still accurately determine the corresponding group transition information to ensure recognition accuracy without relying on the detection algorithm of the moving object group. Optionally, through the small number of trajectory conditions and the lack of reference objects, scenes such as the pan / tilt being blocked at a close distance and the backboard being blocked by a person are set. In these scenes, the group movement trend can be more accurately calculated through optical flow estimation.
[0177] In one embodiment, according to the moving direction and the moving speed, the pan-tilt head is controlled to capture the image of the video to be processed, including: controlling the pan-tilt head to rotate along the rotation direction corresponding to the moving direction at an angular velocity corresponding to the moving speed; during the rotation of the pan-tilt head, the image of the target sphere is captured.
[0178] In another embodiment, according to the moving direction and the moving speed, the pan / tilt is controlled to collect the picture of the video to be processed to obtain the half-field picture, including: adjusting the key framing angle according to the moving direction and the moving speed to obtain the half-field adjustment angle; controlling the pan / tilt to collect the picture of the video to be processed under the half-field adjustment angle to obtain the half-field picture. Thus, on the basis of collecting the picture according to the key framing angle, the adjustment is made according to the moving direction and the moving speed to obtain the half-field picture.
[0179] In this embodiment, when the framing state is the half-field state, the moving direction and moving speed of the moving object group are detected based on the framing screen; according to the moving direction and moving speed, the pan / tilt is controlled to collect the screen of the video to be processed to obtain the half-field screen. Thus, the half-field screen is synchronously adjusted based on the moving object group, thereby obtaining multiple types of half-field screens. On this basis, the close-up event triggered by the target object can be detected more accurately, and it is conducive to more timely detection of the transition initiation event, so that the half-field state can be switched to the transition state in a timely manner.
[0180] In one embodiment, according to the moving direction and the movement speed, the pan / tilt head is controlled to capture the picture of the video to be processed, and after obtaining the half-field picture, the method further includes: if it is determined according to the half-field picture that the moving object group arrives at different half-field areas of the current venue, and the movement speed exceeds the transition threshold, the framing state is adjusted to the transition state.
[0181] The current venue is a sports venue in the video to be processed. Optionally, the current venue may be a basketball court or a football court. Optionally, there are different half-court areas on both sides of the dividing line of the current venue. In the case of capturing a picture of a half-court area, the above half-court picture can be obtained. The transition threshold may be preset or calculated based on the movement speed of the group of moving objects in a certain historical time period.
[0182] In a feasible implementation, determining that the moving object group arrives at different half-court areas of the current field based on the half-court picture includes: determining whether the movement trajectory of each moving object passes through different half-court areas of the current field; if so, determining that the moving object group arrives at different half-court areas of the current field.
[0183] In another feasible implementation, determining that the moving object group arrives at different half-court areas of the current field based on the half-court picture includes: if the position of the moving object group passes through the boundary line of the current field, determining that the moving object group arrives at different half-court areas of the current field.
[0184] Determining the arrival of the moving object group at different half-court areas of the current venue based on the half-court picture is an established fact to judge whether there is a transition from the perspective of position; and the movement trend of the moving object group can be more accurately judged when the movement speed exceeds the transition threshold. Through the detection of these two angles, it is conducive to more accurate and timely detection of transition initiation events, so that the collection process of the half-court state and the transition state can be better connected.
[0185] In one embodiment, the target picture includes a transition picture.
[0186] Determining a framing picture of the video to be processed under a framing viewing angle includes: when the framing state is a transition state, collecting a picture of the video to be processed to obtain a transition framing picture.
[0187] Correspondingly, according to the framing state, the image of the video to be processed is collected to obtain the target image, including: based on at least one image collection object in the moving object group and the target sphere in the transition framing image, the image of the video to be processed is collected to obtain the transition image.
[0188] The picture acquisition object is a moving object in the transition framing picture, which is used to capture the subsequent pictures of the transition framing picture. Optionally, both the moving object group and the target sphere have a movement trend, and the picture acquisition process in the transition process can be more accurately judged by collecting based on the picture acquisition object.
[0189] In an optional implementation, capturing the images of the video to be processed to obtain the transition framing images includes: capturing the images of the video to be processed according to the current key framing angle during the time period when the half-time state is changed to the transition state to obtain the transition framing images. Exemplarily, for several frames of images when the half-time state is changed to the transition state, the current key framing angle can be temporarily retained for image capture.
[0190] In another optional implementation, the images of the video to be processed are captured to obtain a transition framing image, including: adjusting the pan / tilt head according to a preset speed and a preset direction to obtain a preset transition framing angle; and capturing the images of the video to be processed through the preset transition framing angle to obtain a transition framing image.
[0191] In another optional implementation, the image of the video to be processed is collected to obtain a transition framing picture, including: based on at least one of the image collection objects in the group of moving objects and the target sphere in the previous frame transition framing picture, adjusting the transition framing angle of the previous frame transition framing picture to the transition framing angle of the current frame transition framing picture; and collecting the image of the video to be processed according to the transition framing angle of the transition framing picture to obtain the current frame transition framing picture. Thus, in the process of image collection, the transition framing picture is also collected according to the image collection object, and the current frame transition framing picture is a transition picture relative to the previous frame transition framing picture, so that each frame transition picture can be collected based on the image collection object according to the time sequence, thereby improving the collection efficiency of the transition picture.
[0192] In this embodiment, when the framing state is a transition state, the image of the video to be processed is captured to obtain a transition framing image, so that the image capture process of the transition process forms a certain feedback mechanism. In this case, based on the moving object group and at least one image capture object in the target sphere in the transition framing image, the image of the video to be processed is captured to obtain a transition image, so that the process of the transition image is more accurate.
[0193] In one embodiment, based on the moving object group in the transition framing picture and at least one picture capture object in the target sphere, the picture of the video to be processed is captured to obtain a transition picture, including: determining the moving object group according to the density of the moving objects in the transition framing picture; and capturing the picture of the video to be processed according to the position of the moving object group to obtain a transition picture.
[0194] In one feasible implementation, the moving object group is determined according to the moving object density in the transition framing picture, including: in the transition framing picture, the number of moving objects in different areas is detected to obtain the number of moving objects in different areas; the number of moving objects in different areas is compared to obtain the area with the largest number of moving objects; wherein the number of moving objects in different areas is used to characterize the moving object density, and the area with the largest number of moving objects is the area where the moving object group is located.
[0195] In another feasible implementation, the moving object group is determined according to the moving object density in the transition framing picture, including: in the transition framing picture, detecting the number of moving objects in different areas to obtain the number of moving objects in different areas; obtaining the moving object density based on the ratio between the number of moving objects in different areas and the area of each area; and determining the area with the largest moving object density as the area where the moving object group is located.
[0196] In a feasible implementation, according to the position of the moving object group, the picture of the video to be processed is captured to obtain the transition picture, including: judging whether the position of the moving object group is detected to reach the boundary of the current field from the first half area; if so, capturing the picture in the second half area at a preset speed; wherein the first half area and the second half area are different areas, optionally, the first half area is captured through a key framing angle of a key viewing angle, and the second half area is captured through a key framing angle of another key viewing angle.
[0197] In another feasible implementation, according to the position of the moving object group, the picture of the video to be processed is captured to obtain a transition picture, including: determining the viewing angle center according to the position of the moving object group; according to the viewing angle center and the field of view of the reference framing angle, the picture of the video to be processed is captured to obtain a transition picture.
[0198] In this embodiment, the moving object group is determined according to the density of moving objects in the transition framing picture, and occlusion is not likely to occur, so the accuracy of its recognition is relatively high. Under this premise, the picture of the video to be processed is captured according to the position of the moving object group to obtain the transition picture, and the capture process of the transition picture can be controlled more accurately to improve the efficiency of transition picture capture.
[0199] In one embodiment, based on the group of moving objects in the transition framing screen and at least one of the target spheres, the video to be processed is captured to obtain the transition screen, including: detecting the position of the target sphere based on the transition framing screen; capturing the screen of the video to be processed according to the position of the sphere under the reference framing angle to obtain the transition screen. Thus, the screen capture process of the video to be processed can be controlled according to the movement of the target sphere to obtain the transition screen.
[0200] In one embodiment, the framing angle corresponding to the transition state further includes a transition framing angle;
[0201] Based on at least one of the image acquisition objects in the moving object group and the target sphere in the transition framing picture, the image is captured for the video to be processed to obtain a transition picture, including: determining the corresponding pan / tilt speed according to the movement speed of the image acquisition object; adjusting the transition framing angle of the transition framing picture according to the pan / tilt speed to obtain the adjusted transition framing angle; and capturing the image of the video to be processed under the adjusted transition framing angle to obtain the transition picture.
[0202] The gimbal speed is used to characterize the speed at which the virtual gimbal rotates. Optionally, the gimbal speed can be an angular velocity or a linear velocity. Optionally, a fitting function of the gimbal speed change can be generated according to the motion speed of the image acquisition object; the gimbal speed is determined by the fitting function; optionally, the motion speed of various image acquisition objects can be mapped to the gimbal speed.
[0203] The adjusted transition framing angle is the framing angle obtained by adjusting the transition framing angle. Optionally, the transition framing angle of the transition framing picture can be adjusted in terms of the angle of view center according to the pan / tilt speed to obtain the adjusted transition framing angle. The transition framing angle of the transition framing picture can be adjusted in terms of the angle of view center and the field of view angle according to the pan / tilt speed to obtain the adjusted transition framing angle. Exemplarily, in the case of fast transfer of the target sphere, the angle of view center can be determined according to the position of the target sphere, and the field of view angle can be increased to more accurately identify the target sphere.
[0204] In this embodiment, the corresponding pan / tilt speed is determined according to the movement speed of the picture acquisition object, so that the transition picture and the movement of the picture acquisition object can form a synchronization effect; the transition framing angle of the transition framing picture is adjusted according to the pan / tilt speed to obtain the adjusted transition framing angle, so that the adjustment amplitude of the transition framing angle each time is relatively small; in this case, the picture of the processed video under the adjusted transition framing angle is captured to obtain the transition picture, and a transition picture with a synchronization effect and a gradual property can be obtained.
[0205] In an exemplary embodiment, the present application provides an intelligent shooting and editing solution. After using a panoramic camera to shoot a video, the panoramic video is analyzed for video content, and the player group and basketball are identified for movement and / or events to determine events such as transitions and offensive events. A continuous camera movement trajectory is then generated based on the target / event of interest, and a planar video is automatically generated. In the identification of transition events, AI technology can be used to distinguish based on crowd density, orientation, and movement trends. A planar video that automatically follows key targets / events (basketball / crowd / transition / goal) is generated. For offensive events, it can be used as a category of close-up events to make the target video more diverse.
[0206] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0207] Based on the same inventive concept, the embodiment of the present application also provides a video generation device for implementing the video generation method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more video generation device embodiments provided below can refer to the limitations on the video generation method above, and will not be repeated here.
[0208] In one embodiment, Figure 7 As shown, a video generating device is provided, comprising:
[0209] A framing module 702 is used to determine a framing screen of the video to be processed under a framing viewing angle;
[0210] A detection module 704 is used to determine a framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture;
[0211] A target picture acquisition module 706 is used to acquire the picture of the video to be processed according to the framing state to obtain a target picture;
[0212] The video generation module 708 is used to generate a target video according to the target picture.
[0213] In one embodiment, the framing module 702 is used to:
[0214] Detecting a reference object in the video to be processed;
[0215] Determining a key framing angle for capturing an image of the reference object according to the position of the reference object;
[0216] According to the key framing angle, the video to be processed is captured to obtain a framing picture.
[0217] In one of the embodiments, the framing picture includes a sampling framing picture and an interpolation framing picture;
[0218] The framing module 702 is used to:
[0219] Based on each sampling moment, performing framing angle detection on the picture sampled from the video to be processed to obtain a sampled framing angle;
[0220] According to the sampling framing angle, the video to be processed is captured to obtain the sampling framing images at each sampling moment;
[0221] Performing interpolation calculation on the sampling framing angles at the sampling moments to obtain interpolated framing angles between the sampling moments;
[0222] At the non-sampling moments between the sampling moments, the video to be processed is captured according to the interpolation framing angle of view to obtain the interpolation framing pictures at the non-sampling moments.
[0223] In one embodiment, the detection module 704 is used to:
[0224] If the framing state corresponding to the framing picture is a half-scene state and a transition initiation event of the framing picture is identified, adjusting the framing state to a transition state;
[0225] If the framing state corresponding to the framing picture is a transition state and a transition end event of the framing picture is identified, the framing state is adjusted to a half-field state.
[0226] In one of the embodiments, the framing perspective includes at least two key framing perspectives and a transition framing perspective, and the transition framing perspective is a framing perspective between the at least two key framing perspectives;
[0227] The framing module 702 is used for:
[0228] Under the transition framing angle, determining a transition framing screen of the video to be processed;
[0229] Correspondingly, the detection module 704 is used to:
[0230] When the framing state of the transition framing picture is a transition state, determining whether a key framing angle matching the transition framing angle exists;
[0231] If yes, the framing state corresponding to the transition framing picture is adjusted to a half-scene state.
[0232] In one of the embodiments, the at least two key framing angles include a framing angle to be matched in a transition direction, and the framing angle to be matched is a framing angle according to the framing point to be matched;
[0233] The detection module 704 is used to:
[0234] Determine the deviation distance between the viewing angle center of the transition viewing angle and the viewing angle to be matched;
[0235] If the deviation distance satisfies the matching condition, then the key framing angle matching the transition framing angle exists;
[0236] If the deviation distance does not satisfy the matching condition, then the key framing angle that matches the transition framing angle does not exist.
[0237] In one of the embodiments, the target picture includes a close-up picture;
[0238] The target image acquisition module 706 is used to:
[0239] When the framing state of the framing picture is a half-field state, detecting a close-up event based on a key framing angle of view;
[0240] If it is detected that the target object triggers the close-up event, adjusting the key framing angle according to the position of the target object to obtain a close-up framing angle;
[0241] The frame of the video to be processed is continuously captured according to the close-up framing angle to obtain the close-up frame, until the close-up frame satisfies the end condition of the close-up event, and the continuous capture of the frame of the video to be processed according to the close-up framing angle is stopped.
[0242] In one embodiment, the target image acquisition module 706 is used to:
[0243] When the framing state is a half-field state, detecting the moving direction and moving speed of the moving object group based on the framing picture;
[0244] According to the moving direction and the moving speed, the pan / tilt head is controlled to collect the picture of the video to be processed to obtain a half-field picture.
[0245] In one embodiment, according to the moving direction and the moving speed, the pan / tilt head is controlled to collect the picture of the video to be processed, and after the half-field picture is obtained, the detection module 704 is used to:
[0246] If it is determined according to the half-field picture that the group of moving objects arrives at different half-field areas of the current scene, and the moving speed exceeds a transition threshold, the framing state is adjusted to a transition state.
[0247] In one embodiment, the target picture includes a transition picture; the framing module 702 is used to:
[0248] When the framing state is a transition state, collecting the picture of the video to be processed to obtain a transition framing picture;
[0249] The target image acquisition module 706 is used to:
[0250] Based on at least one of the picture acquisition objects in the moving object group and the target sphere in the transition framing picture, the picture acquisition is performed on the video to be processed to obtain a transition picture.
[0251] In one embodiment, the target image acquisition module 706 is used to:
[0252] Determining a moving object group according to the density of moving objects in the transition framing picture;
[0253] According to the positions of the moving object group, the images of the video to be processed are captured to obtain transition images.
[0254] In one of the embodiments, the framing angle corresponding to the transition state further includes a transition framing angle;
[0255] The target image acquisition module 706 is used to:
[0256] Determine the corresponding pan / tilt speed according to the movement speed of the image acquisition object;
[0257] Adjusting the transition framing angle of the transition framing picture according to the pan / tilt speed to obtain an adjusted transition framing angle;
[0258] The picture of the video to be processed under the adjusted transition viewing angle is collected to obtain a transition picture.
[0259] In one of the embodiments, the field of view of the video to be processed is greater than the field of view of the target video.
[0260] Each module in the above video generation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0261] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a video generation method is implemented. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.
[0262] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0263] In one embodiment, the present application also provides a handheld gimbal, including a motor and a processor, wherein the motor is used to control the rotation of the gimbal, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0264] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0265] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0266] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0267] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0268] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0269] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0270] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A video generation method, characterized in that: The method comprises: Determine the framing screen of the video to be processed under the framing perspective; Determining a framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture; Capturing the picture of the video to be processed according to the framing state to obtain a target picture; A target video is generated according to the target picture.
2. The method according to claim 1, characterized in that The step of determining a framing picture of the video to be processed under a framing viewing angle includes: Detecting a reference object in the video to be processed; Determining a key framing angle for capturing an image of the reference object according to the position of the reference object; According to the key framing angle, the video to be processed is captured to obtain a framing picture.
3. The method according to claim 1, characterized in that: The framing picture includes a sampling framing picture and an interpolation framing picture; The step of determining a framing picture of the video to be processed under a framing viewing angle includes: Based on each sampling moment, performing framing angle detection on the picture sampled from the video to be processed to obtain a sampled framing angle; According to the sampling framing angle, the video to be processed is captured to obtain the sampling framing images at each sampling moment; Performing interpolation calculation on the sampling framing angles at the sampling moments to obtain interpolated framing angles between the sampling moments; At the non-sampling moments between the sampling moments, the video to be processed is captured according to the interpolation framing angle of view to obtain the interpolation framing pictures at the non-sampling moments.
4. The method according to claim 1, characterized in that: The determining the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture includes: If the framing state corresponding to the framing picture is a half-scene state and a transition initiation event of the framing picture is identified, adjusting the framing state to a transition state; If the framing state corresponding to the framing picture is a transition state and a transition end event of the framing picture is identified, the framing state is adjusted to a half-field state.
5. The method according to claim 1, characterized in that The framing angles include at least two key framing angles and a transition framing angle, and the transition framing angle is a framing angle between the at least two key framing angles; The step of determining a framing picture of the video to be processed under a framing viewing angle includes: Under the transition framing angle, determining a transition framing screen of the video to be processed; The determining the framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture includes: When the framing state of the transition framing picture is a transition state, determining whether a key framing angle matching the transition framing angle exists; If yes, the framing state corresponding to the transition framing picture is adjusted to a half-scene state.
6. The method according to claim 5, characterized in that The at least two key framing angles include a framing angle to be matched in the transition direction, and the framing angle to be matched is a framing angle according to the framing point to be matched; The determining whether a framing angle matching the transition framing picture exists includes: Determine the deviation distance between the viewing angle center of the transition viewing angle and the viewing angle to be matched; If the deviation distance satisfies the matching condition, then the key framing angle matching the transition framing angle exists; If the deviation distance does not satisfy the matching condition, then the key framing angle that matches the transition framing angle does not exist.
7. The method according to claim 1, characterized in that The target picture includes a close-up picture; The step of collecting the picture of the video to be processed according to the framing state to obtain the target picture includes: When the framing state of the framing picture is a half-field state, detecting a close-up event based on a key framing angle of view; If it is detected that the target object triggers the close-up event, adjusting the key framing angle according to the position of the target object to obtain a close-up framing angle; The frame of the video to be processed is continuously captured according to the close-up framing angle to obtain the close-up frame, until the close-up frame satisfies the end condition of the close-up event, and the continuous capture of the frame of the video to be processed according to the close-up framing angle is stopped.
8. The method according to claim 1, characterized in that The step of collecting the picture of the video to be processed according to the framing state to obtain the target picture includes: When the framing state is a half-field state, detecting the moving direction and moving speed of the moving object group based on the framing picture; According to the moving direction and the moving speed, the pan / tilt head is controlled to collect the picture of the video to be processed to obtain a half-field picture.
9. The method according to claim 8, characterized in that After the pan / tilt head is controlled to capture the picture of the video to be processed according to the moving direction and the moving speed and a half-field picture is obtained, the method further comprises: If it is determined according to the half-field picture that the group of moving objects arrives at different half-field areas of the current scene, and the moving speed exceeds a transition threshold, the framing state is adjusted to a transition state.
10. The method according to claim 1, characterized in that The target screen includes a transition screen; the step of determining the framing screen of the video to be processed under the framing viewing angle includes: When the framing state is a transition state, collecting the picture of the video to be processed to obtain a transition framing picture; The step of collecting the picture of the video to be processed according to the framing state to obtain the target picture includes: Based on at least one of the picture acquisition objects in the moving object group and the target sphere in the transition framing picture, the picture acquisition is performed on the video to be processed to obtain a transition picture.
11. The method according to claim 10, characterized in that The method of performing picture acquisition on the video to be processed based on at least one picture acquisition object among the moving object group and the target sphere in the transition framing picture to obtain the transition picture includes: Determining a moving object group according to the density of moving objects in the transition framing picture; According to the positions of the moving object group, the images of the video to be processed are captured to obtain transition images.
12. The method according to claim 10, characterized in that The framing angle corresponding to the transition state also includes the transition framing angle; The method of performing picture acquisition on the video to be processed based on at least one picture acquisition object among the moving object group and the target sphere in the transition framing picture to obtain the transition picture includes: Determine the corresponding pan / tilt speed according to the movement speed of the image acquisition object; Adjusting the transition framing angle of the transition framing picture according to the pan / tilt speed to obtain an adjusted transition framing angle; The picture of the video to be processed under the adjusted transition viewing angle is collected to obtain a transition picture.
13. The method according to claim 1, characterized in that The field of view of the video to be processed is greater than the field of view of the target video.
14. A video generating device, characterized in that: The device comprises: A framing module, used to determine the framing screen of the video to be processed under the framing perspective; A detection module, used for determining a framing state corresponding to the framing picture according to the existence state of the transition event of the framing picture; A target picture acquisition module, used to acquire the picture of the video to be processed according to the framing state to obtain a target picture; The video generation module is used to generate a target video according to the target picture.
15. A handheld gimbal, characterized in that: It comprises a motor and a processor, wherein the motor is used to control the rotation of the pan / tilt head, and the processor is used to implement the steps of the method described in any one of claims 1 to 13.
16. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.