Video generation method, apparatus, device, and storage medium

By identifying the motion trajectory and posture changes of the target video frame, the system selects confidence video frames and generates video clips, solving the problem of low generation efficiency when the target object moves quickly, and achieving efficient and exciting scene editing.

CN119316626BActive Publication Date: 2026-01-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310848678.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-11
Publication Date
2026-01-27
Estimated Expiration
2043-07-11

AI Technical Summary

Technical Problem

Existing video generation methods struggle to efficiently select key segments when the target object moves quickly, resulting in low generation efficiency.

Method used

By identifying motion trajectories and posture changes in target video frames, confidence video frames are selected, and video clips are generated based on dynamic action categories.

Benefits of technology

It improves the efficiency and quality of video generation, enabling accurate editing of exciting moments during the movement of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316626B_ABST
    Figure CN119316626B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method, device and equipment and a storage medium. The method comprises: obtaining a target video, and determining a target video frame representing a motion trajectory of a target object in the target video; identifying motion information of the target object in the target video frame, and screening a confident video frame from the target video frame based on the identified motion information; wherein the motion information is used to represent at least a posture change of the target object; identifying a dynamic action category of the target object in the confident video frame, determining a to-be-edited video frame in the confident video frame based on the identified dynamic action category, and generating a video clip of the target object according to the to-be-edited video frame. The technical solution provided by the present disclosure can improve the efficiency of video generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, specifically to a video generation method, apparatus, device, and storage medium. Background Technology

[0002] In some practical applications, video clips related to a target object can be generated based on pre-recorded video content. For example, compelling segments of the target object can be selected from pre-recorded video content, and video clips of the target object can be generated based on the selected segments.

[0003] However, existing video generation methods often suffer from low generation efficiency. In certain scenarios, such as when the target object moves quickly, it is difficult to select accurate and exciting segments from pre-shot video content, resulting in low efficiency in the final video generation. Summary of the Invention

[0004] In view of this, one or more embodiments of the present disclosure provide a video generation method, apparatus, device and storage medium that can improve the efficiency of video generation.

[0005] This disclosure provides a video generation method, the method comprising: acquiring a target video and determining target video frames representing the motion trajectory of a target object in the target video; identifying motion information of the target object in the target video frames and, based on the identified motion information, filtering confidence video frames in the target video frames; wherein the motion information is at least used to represent the posture changes of the target object; identifying the dynamic action category of the target object in the confidence video frames and, based on the identified dynamic action category, determining video frames to be edited in the confidence video frames, and generating a video segment of the target object based on the video frames to be edited.

[0006] This disclosure also provides a video generation apparatus, the apparatus comprising: a trajectory recognition unit, configured to acquire a target video and determine target video frames representing the motion trajectory of a target object in the target video; a video frame filtering unit, configured to identify motion information of the target object in the target video frames and filter confidence video frames in the target video frames based on the identified motion information; wherein the motion information is at least used to represent the posture changes of the target object; and a video generation unit, configured to identify the dynamic action category of the target object in the confidence video frames, determine video frames to be edited in the confidence video frames based on the identified dynamic action category, and generate video segments of the target object based on the video frames to be edited.

[0007] This disclosure also provides an electronic device including a memory and a processor, the memory being used to store a computer program, which, when executed by the processor, implements the video generation method described above.

[0008] This disclosure also provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the video generation method described above.

[0009] This disclosure provides a technical solution through one or more embodiments, which identifies motion information of a target object from a target video frame representing the motion trajectory of that object. This motion information indicates the changes in the target object's posture during motion.

[0010] Generally speaking, changes in posture can indirectly reflect the intensity of a target object's movement, and the intensity of that movement also reflects the quality of the video footage. Therefore, by analyzing the aforementioned motion information, confident video frames can be selected from the target video frames. Within these confident video frames, the target object is more likely to produce exciting scenes.

[0011] Subsequently, by identifying the dynamic action categories of the confidence video frames, the current motion stage of the target object can be determined. Then, based on the identification results of the dynamic action categories, video clips that reflect the target user's motion process can be edited out.

[0012] As can be seen, through the technical solutions provided by one or more embodiments of this disclosure, confident video frames containing exciting scenes can be effectively determined from the target video frames based on motion information. By identifying the dynamic action categories in the confident video frames, the exciting scenes can be further located in detail within the confident video frames. Subsequently, video clips of the target object can be generated based on the located exciting scenes. Through the above method, not only can video clips of the target object's movement process be edited, but the motion state displayed in the video clips can also reflect the intensity of the movement process, thereby improving the efficiency and quality of video generation. Attached Figure Description

[0013] The features and advantages of the embodiments of this disclosure will be more clearly understood by referring to the accompanying drawings, which are illustrative and should not be construed as limiting the present disclosure in any way. In the drawings:

[0014] Figure 1 A schematic diagram of the video generation method steps in one embodiment of this disclosure is shown;

[0015] Figure 2 A schematic diagram illustrating the filtering of confidence video frames is shown in one embodiment of this disclosure;

[0016] Figure 3 A schematic diagram of the functional modules of a video generation apparatus in one embodiment of this disclosure is shown;

[0017] Figure 4 A schematic diagram of the structure of an electronic device according to one embodiment of the present disclosure is shown. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0019] This disclosure provides a video generation method according to one embodiment. Please refer to [link / reference]. Figure 1 The method may include the following steps.

[0020] S1: Acquire the target video and determine the target video frame representing the motion trajectory of the target object in the target video.

[0021] In this embodiment, the target object can be a user who needs to generate a video. It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, the user should be informed of the type, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and the user's authorization should be obtained.

[0022] For example, before implementing the technical solutions of the various embodiments of this disclosure, a prompt message can be sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as the electronic device, application program, server, or storage medium that performs the operations of the technical solutions of this disclosure, based on the prompt message.

[0023] As an optional but not limited implementation, the prompt message can be sent to the user in the form of a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control for the user to choose "agree" or "disagree".

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0026] It is understood that in the specific implementation of this application, data such as identity information and video data are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0027] In this embodiment, cameras deployed in offline scenarios can capture target videos containing the target object as it moves within those scenarios. It should be noted that in practical applications, there can be more than one target video. For example, multiple cameras can be deployed in the offline scenario, capturing target videos from different perspectives. Each target video can then be processed using the same or similar methods. Of course, the subsequent processing of the target videos, as well as the editing of their content, is performed only after authorization from the target object has been obtained. In practical applications, object detection algorithms can be used to analyze the target video frame by frame. After processing any video frame using the object detection algorithm, the location information of one or more bounding boxes corresponding to that video frame can be output. Each bounding box can represent a human body in the video frame; by using the location information of the bounding boxes, the human body in the video frame can be identified and located. Furthermore, after the object detection algorithm, each bounding box representing a human body will have its own bounding box identifier, which can be, for example, a globally unique number. After performing object detection processing on different video frames of the target video, the position information and bounding box identifiers of the bounding boxes mentioned above can be output for each video frame. Bounding boxes with the same identifier can represent the same human body. Thus, after frame-by-frame detection of the target video, by identifying the bounding boxes with the same identifier, the motion trajectory of the same human body can be tracked across different video frames.

[0028] In this embodiment, the target object can be any human body for which video generation is requested. In the target video, the target object may only appear in a subset of video frames and be identified by the object detection algorithm; these frames can then be used as target video frames representing the motion trajectory of the target object. Thus, by performing frame-by-frame object detection on the target video, the target video frames representing the motion trajectory of the target object can be determined within the target video.

[0029] S3: Identify the motion information of the target object in the target video frame, and based on the identified motion information, filter out confidence video frames in the target video frame; wherein the motion information is at least used to characterize the pose change of the target object.

[0030] In this embodiment, within the target video frames representing the motion trajectory of the target object, more refined feature recognition can be performed on the target object in each target video frame. Specifically, considering that the final edited video clips typically need to showcase the exciting moments of the target object during its movement, and these exciting moments often correspond to a variety of body movements of the target object, the motion information of the target object in the target video frames can be identified. Then, based on this motion information, the target video frames can be further filtered to select confidence video frames with a high probability of exhibiting exciting moments.

[0031] In this embodiment, the motion information of the target object can be represented in various ways. At a minimum, the motion information of the target object can characterize the changes in its posture. The posture of the target object can include the posture information of various designated parts of the target object. These designated parts can be, for example, human body parts that can reflect a state of motion, such as the target object's hands, legs, or waist.

[0032] In one implementation, the pose information of various designated parts of the target object can be identified using a trained pose recognition model. Specifically, the training samples for the pose recognition model during the training phase can be image samples containing designated parts of the human body. Furthermore, for each image sample, the pose information of the designated parts contained within the image sample can be pre-labeled. In practical applications, the pose information of the designated parts can be represented by pre-set state values. For example, for the hand, there are usually multiple poses such as raised, level, and lowered. Different state values ​​can be used to represent these different poses. For example, raised can correspond to state value 1, level can correspond to state value 2, and lowered can correspond to state value 3. Thus, the aforementioned state values ​​can serve as the pose information of the designated parts. In an image sample, there may be multiple designated parts, and each designated part can be characterized by a corresponding state value based on its current pose. The collection of pose information from all designated parts can then serve as the motion information of the human body in the image sample.

[0033] In practical applications, the order of various designated body parts can be predefined. Then, according to the order, the posture information of each designated body part is concatenated to form the motion information of the human body in the image sample. For example, if the order of the designated body parts is: hand, waist, and leg, then the motion information of the human body in the image sample can be represented by a vector (posture information of the hand, posture information of the waist, posture information of the leg). Suppose that in a certain image sample, the state value corresponding to the posture information of the hand is 1, the state value corresponding to the posture information of the waist is 1, and the state value corresponding to the posture information of the leg is 2, then the motion information of the human body in this image sample can be represented by the vector (1, 1, 2).

[0034] Using the methods described above, each image sample can possess pre-annotated motion information. The pose recognition model can be implemented using a neural network. Image samples are input as training data into the pose recognition model, and the output is supervised by the pre-annotated motion information, thereby generating corresponding error information. This error information can adjust the parameters of neurons in the pose recognition model, ensuring that the output remains consistent with the annotated motion information. Through training with a large number of image samples, a pose recognition model capable of accurately identifying human motion information in images can be obtained.

[0035] In this embodiment, the bounding box of the target object in the target video frame can be obtained through an object detection algorithm. Based on the bounding box, the region image of the target object can be extracted from the target video frame. Then, the region image can be input into the trained pose recognition model to obtain the pose information of each specified part of the target object.

[0036] In this embodiment, after identifying the pose information of each specified part from the target video frame, for any specified part, it can be determined whether the pose information of that specified part has changed compared to the previous target video frame. In practical applications, the pose information of the specified part can be represented in the form of a vector. Therefore, when determining whether the pose information of the specified part has changed, it is only necessary to subtract the state values ​​at corresponding positions in two adjacent vectors. For example, for the first target video frame, the pose information of each specified part of the target object can be represented as (1, 1, 2), and for the second target video frame, the pose information of each specified part of the target object can be represented as (3, 1, 2). By comparing the two vectors, the final judgment result can be obtained as follows: the pose of the target object's hand has changed, and the pose change value is 3-1=2.

[0037] In this way, by comparing the pose information of a specified part in adjacent target video frames, the pose change information of the specified part can be obtained. This pose change information can be represented by the pose change values ​​mentioned above. Finally, the combination of the pose change information of each specified part can serve as the motion information of the target object in the target video frame. For example, for the first and second target video frames mentioned above, the motion information of the target object can be represented as (2, 0, 0). It should be noted that each value in the motion information represents a pose change value. The larger the pose change value, the greater the pose change of the specified part. The pose change value here is not the same as the state value representing the pose information mentioned above.

[0038] In one implementation, the motion information of the target object can also be represented by velocity and acceleration. Specifically, the velocity of the target object can be determined jointly by adjacent first and second target video frames, wherein the first target video frame may be located before the second target video frame. Using an object detection algorithm, a first bounding box of the target object can be identified in the first target video frame, and a second bounding box of the target object can be identified in the second target video frame. Then, the positional difference between the second bounding box and the first bounding box can be determined. Specifically, when determining the positional difference, a first center position of the first bounding box and a second center position of the second bounding box can be determined first. Then, along the direction of the target object's motion, the positional difference between the first and second center positions is calculated, and this positional difference can be used as the positional difference between the second and first bounding boxes. The first and second center positions can be represented by pixel coordinates in the video frames. For example, the first center position can be represented as (100, 200), and the second center position can be represented as (200, 210). It can be found that the target object has been displaced in both the horizontal and vertical directions. However, by analyzing the horizontal and vertical coordinates, it can be determined that the target object is currently mainly displaced in the horizontal direction. At this time, the difference between the horizontal coordinates of the second center position and the first center position can be used as the position difference between the first center position and the second center position.

[0039] In this embodiment, the time difference between the first target video frame and the second target video frame is known. Therefore, the ratio of the position difference to the time difference can be used as the velocity of the target object. Generally speaking, the velocity of the target object determined based on two adjacent target video frames can be used as the velocity of the target object identified from the target video frame that is later in the temporal sequence (i.e., the second target video frame).

[0040] Of course, in some scenarios, if the target object has undergone a certain degree of displacement in both the horizontal and vertical directions, then the position difference can be calculated separately for the horizontal and vertical directions, and the respective velocities in the horizontal and vertical directions can be calculated based on the position difference between the horizontal and vertical directions.

[0041] It should be noted that the aforementioned positional difference can be represented by the number of pixels. For example, the difference of 100 in the horizontal coordinate between the first and second center positions can refer to 100 pixels. This allows the velocity of the target object to be calculated in the image coordinate system. In practical applications, the first and second center positions can be converted to coordinates in the world coordinate system through transformations between the image coordinate system, camera coordinate system, and world coordinate system. The calculated velocity in this way can represent the velocity of the target object in the world coordinate system. The specific coordinate system used for velocity calculation can be flexibly selected according to actual needs.

[0042] In one implementation, after determining the velocity of the target object, its acceleration can be further determined. Specifically, still taking the first target video frame and the second target video frame as examples, after determining the first velocity of the target object in the first target video frame and the second velocity of the target object in the second target video frame, the velocity difference between the second velocity and the first velocity can be calculated, and the acceleration of the target object can be determined based on the velocity difference. Specifically, the time difference between the first target video frame and the second target video frame is known, and the ratio of the aforementioned velocity difference to this time difference can be used as the acceleration of the target object. Similarly, the acceleration of the target object determined based on two adjacent target video frames can be used as the acceleration of the target object identified from the target video frame that is later in the temporal sequence (i.e., the second target video frame).

[0043] It should be noted that, with the change of application scenarios and the advancement of technology, the motion information of the target object identified from the target video frame can also take many other forms, which will not be elaborated here.

[0044] In this embodiment, motion information of the target object can be identified from each target video frame in the manner described above. This motion information can include multiple different types of motion parameters. For example, these motion parameters can be the aforementioned posture change values, velocity, acceleration, etc. Although the target video frames can show the motion trajectory of the target object, in some target video frames, the target object may be in a non-motion state or in an unattractive posture; these target video frames may not be used as material for subsequent video editing. Therefore, based on the motion information of the target video frames, the confidence level of each target video frame can be determined. This confidence level characterizes the probability that a target video frame contains exciting scenes. Thus, based on the confidence level, target video frames can be filtered to obtain video frames with a higher probability of containing exciting scenes.

[0045] In this embodiment, when determining the confidence level of each target video frame, the motion parameters can be weighted according to their preset weights to generate the confidence level of the target video frame. In practical applications, the preset weights of each motion parameter can be flexibly set according to the needs of video generation. Furthermore, different specified body parts may have different pose change values. During the confidence level calculation, the pose change values ​​of different specified body parts can be treated as independent motion parameters, and independent preset weights can be set. Of course, depending on the video generation requirements, the pose change values ​​of each specified body part can be averaged or weighted to obtain an overall comprehensive change value, and then this comprehensive change value can be weighted and summed with other motion parameters. This disclosure does not limit this approach.

[0046] It should be noted that different types of motion parameters may have different physical dimensions. To weight different types of motion parameters within the same numerical dimension, we can first normalize each type of motion parameter separately, thus processing the motion parameters into dimensionless data. This ensures consistency in data processing when weighting the dimensionless data subsequently.

[0047] In this embodiment, the level of confidence can represent the probability of having exciting scenes. After obtaining the confidence level of each target video frame, the target video frames with a confidence level greater than or equal to a specified threshold can be selected as the confident video frames.

[0048] Please see Figure 2 In a practical application example, the confidence level of each target video frame can be as follows: Figure 2 As shown. Figure 2 The dashed lines in the graph represent the specified threshold mentioned above. After filtering the target video frames using confidence scores, we can obtain... Figure 2 The confidence video frames for the two intervals shown.

[0049] In one implementation, when calculating the confidence score of a target video frame, in addition to considering motion information, the quality information of the region image where the target object is located can also be considered. Specifically, based on the detection results of the object detection algorithm, the region image where the target object is located can be extracted from the target video frame according to the bounding box of the target object. The quality information of the region image is generated by analyzing its contrast and brightness. Generally, this quality information can be a numerical value representing the quality; the larger the value, the better the quality. Subsequently, the weight values ​​of the motion information and the quality information can be obtained separately, and the motion information and the quality information can be weighted according to the obtained weight values ​​to obtain the confidence score of the target video frame. Then, target video frames with a confidence score greater than or equal to a specified threshold can be selected as confident video frames. The specific process can be found in the description of the foregoing implementation method, and will not be repeated here.

[0050] S5: Identify the dynamic action category of the target object in the confidence video frame, and based on the identified dynamic action category, determine the video frame to be edited in the confidence video frame, and generate a video segment of the target object according to the video frame to be edited.

[0051] In this embodiment, to automatically generate the final video based on the confidence video frames, the dynamic action category of the target object in the confidence video frames can be identified. In practical applications, dynamic action categories can include a series of preset categories such as waving, twisting, jumping, and sliding. After identifying the current dynamic action category of the target object from the confidence video frames, video frames corresponding to some dynamic action categories can be edited into the final video clip according to a preset video generation strategy.

[0052] For example, a preset video generation strategy can limit the editing of video clips to at least two different types of dynamic actions. After identifying the dynamic action category of the target object from the confidence video frames, the set of confidence video frames corresponding to that dynamic action category can be determined simultaneously. For example, the dynamic action category representing waving could correspond to confidence video frames from frame 10 to frame 50. Subsequently, when editing the video clip, for any dynamic action category, a specified number of confidence video frames can be selected from the set of confidence video frames corresponding to that dynamic action category. This specified number can be flexibly determined based on the length of the video clip. The specified number of confidence video frames selected from each dynamic action category can be used as the video frames to be edited. Of course, preprocessing such as deduplication and noise reduction can also be performed on the specified number of confidence video frames to obtain the video frames to be edited. After determining the video frames to be edited from the confidence video frames, the video clip of the target object can be generated based on the video frames to be edited. During the video clip generation process, the selected confidence video frames can be sorted, have effects added, or have filters changed, which will not be elaborated here.

[0053] Of course, in some application scenarios, the number of confidence video frames can be unlimited when filtering them. For example, after identifying the dynamic action category of the target object, confidence video frames belonging to that category can be filtered out, while those not belonging to that category can be discarded. Another example is that after identifying the dynamic action category of the target object, the correlation between each confidence video frame and that dynamic action category can be determined based on the content displayed in each frame. Subsequently, confidence video frames with a correlation higher than a certain threshold can be filtered out, while those with a correlation lower than that threshold can be discarded.

[0054] In this embodiment, each dynamic action category can correspond to its own action range, which can include multiple phased actions arranged in sequence. For example, for a dynamic action category representing waving, the phased actions in the action range can include the action of starting to wave, the action during the wave, and the action of ending the wave. As another example, for a dynamic action category representing jumping, the phased actions in the action range can include the action of taking off, the action of reaching the highest point of the jump, and the landing action. For different dynamic action categories, the phased actions included in the action range can be predefined.

[0055] In this embodiment, a large number of dynamic action categories' action ranges can be used as training samples to train the action prediction model. Specifically, for each action range of a dynamic action category, each stage of the action within that range can have an action label representing the dynamic action category, and also a stage label representing the stage of the action within the action range. For example, for the dynamic action category of jumping, each stage of the action range has an action label representing "jump," and furthermore, the take-off action, the action of reaching the highest point of the jump, and the landing action can each have their own stage label (e.g., corresponding to the stage labels "take-off," "highest point," and "landing," respectively). Thus, by inputting each stage of the action range into the action prediction model for training, predicted action labels and stage labels can be output. Subsequently, by comparing the predicted action labels and stage labels with pre-labeled action labels and stage labels, error information can be obtained, and this error information can be used to correct the action prediction model.

[0056] In practical applications, since action prediction models need to learn the associations between multiple stages of actions within an action interval, neural networks with temporal relationships can be used to communicate between the action prediction models. For example, action prediction models can be implemented using an LSTM (Long Short-Term Memory) neural network architecture or a Transformer architecture.

[0057] In this embodiment, after training the action prediction model, confidence video frames can be input into the model, which then outputs corresponding action labels and stage labels. This allows confidence video frames with the same action labels to be grouped into the same dynamic action category, and the confidence video frames within the same dynamic action category are sorted according to their stage labels. Thus, based on action labels, confidence video frames can be categorized to obtain a set of confidence video frames corresponding to each dynamic action category. The aforementioned action labels and stage labels can jointly represent the dynamic action category of the target object within the confidence video frames. Subsequently, during video generation, relatively exciting stage actions can be edited into video clips based on the stage labels of the confidence video frames.

[0058] This disclosure provides a technical solution through one or more embodiments, which identifies motion information of a target object from a target video frame representing the motion trajectory of that object. This motion information indicates the changes in the target object's posture during motion.

[0059] Generally speaking, changes in posture can indirectly reflect the intensity of a target object's movement, and the intensity of that movement also reflects the quality of the video footage. Therefore, by analyzing the aforementioned motion information, confident video frames can be selected from the target video frames. Within these confident video frames, the target object is more likely to produce exciting scenes.

[0060] Subsequently, by identifying the dynamic action categories of the confidence video frames, the current motion stage of the target object can be determined. Then, based on the identification results of the dynamic action categories, video clips that reflect the target user's motion process can be edited out.

[0061] As can be seen, through the technical solutions provided by one or more embodiments of this disclosure, confident video frames containing exciting scenes can be effectively determined from the target video frames based on motion information. By identifying the dynamic action categories in the confident video frames, the exciting scenes can be further located in detail within the confident video frames. Subsequently, video clips of the target object can be generated based on the located exciting scenes. Through the above method, not only can video clips of the target object's movement process be edited, but the motion state displayed in the video clips can also reflect the intensity of the movement process, thereby improving the efficiency and quality of video generation.

[0062] Please see Figure 3 One embodiment of this disclosure also provides a video generation apparatus, the apparatus comprising:

[0063] The trajectory recognition unit 100 is used to acquire the target video and determine the target video frame representing the motion trajectory of the target object in the target video;

[0064] The video frame filtering unit 200 is used to identify motion information of the target object in the target video frame, and to filter out confidence video frames in the target video frame based on the identified motion information; wherein the motion information is at least used to characterize the pose change of the target object;

[0065] The video generation unit 300 is configured to identify the dynamic action category of the target object in the confidence video frame, determine the video frame to be edited in the confidence video frame based on the identified dynamic action category, and generate a video segment of the target object based on the video frame to be edited.

[0066] The specific processing logic of each functional module can be found in the description of the aforementioned method implementation method, and will not be repeated here.

[0067] The various units described in the above embodiments can be implemented by a computer chip or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0068] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0069] Please see Figure 4 One embodiment of this disclosure also provides an electronic device, which includes a memory and a processor. The memory is used to store a computer program, which, when executed by the processor, implements the video generation method described above.

[0070] This disclosure also provides a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the video generation method described above.

[0071] The processor can be a central processing unit (CPU). It can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations thereof.

[0072] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the methods in the above-described embodiments.

[0073] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0074] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.

[0075] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, embodiments of apparatus, devices, and storage media are basically similar to method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0076] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

[0077] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A video generation method, characterized in that, The method includes: Acquire the target video, and determine the target video frame representing the motion trajectory of the target object in the target video; Extract the region image of the target object from the target video frame and generate the quality information of the region image; Identifying motion information of the target object in the target video frame, and based on the identified motion information, filtering out confidence video frames from the target video frame, including: The weight values ​​of the motion information and the quality information are obtained respectively, and the motion information and the quality information are weighted according to the obtained weight values ​​to obtain the confidence of the target video frame; Based on target video frames with a confidence level greater than or equal to a specified threshold, select confidence video frames are determined; wherein, the motion information is used at least to characterize the pose change of the target object; The dynamic action category of the target object is identified in the confidence video frame, and based on the identified dynamic action category, a video frame to be edited is determined in the confidence video frame, and a video segment of the target object is generated based on the video frame to be edited.

2. The method according to claim 1, characterized in that, Identifying the motion information of the target object in the target video frame includes: Based on the target video frame, identify the posture information of each specified part of the target object; For any specified part, determine whether the pose information of the specified part has changed compared to the previous target video frame, and generate the pose change information of the specified part based on the determination result; Based on the combination of posture change information of each of the specified parts, the motion information of the target object in the target video frame is determined.

3. The method according to claim 1 or 2, characterized in that, The motion information is also used to characterize the velocity of the target object; the velocity of the target object is determined in the following manner: For adjacent first target video frames and second target video frames, a first bounding box of the target object is identified in the first target video frame, and a second bounding box of the target object is identified in the second target video frame; wherein the first target video frame is located before the second target video frame; Determine the positional difference between the second bounding box and the first bounding box, and determine the velocity of the target object in the second target video frame based on the positional difference.

4. The method according to claim 1 or 2, characterized in that, The motion information is also used to characterize the acceleration of the target object; the acceleration of the target object is determined in the following manner: For adjacent first target video frames and second target video frames, a first velocity of the target object in the first target video frame and a second velocity of the target object in the second target video frame are determined; wherein the first target video frame is located before the second target video frame; The acceleration of the target object in the second target video frame is determined based on the velocity difference between the second velocity and the first velocity.

5. The method according to claim 1, characterized in that, The motion information includes multiple different types of motion parameters; based on the identified motion information, selecting confidence video frames from the target video frames includes: Based on the preset weights of each action parameter, the action parameters are weighted to generate the confidence score of the target video frame. Based on target video frames with a confidence level greater than or equal to a specified threshold, determine the selected confidence video frames.

6. The method according to claim 1, characterized in that, Different dynamic action categories correspond to their respective sets of confidence video frames; based on the identified dynamic action categories, the video frames to be edited are determined from the confidence video frames, including: For any dynamic action category, a specified number of confidence video frames are selected from the set of confidence video frames corresponding to the dynamic action category, and the video frames to be edited are determined based on the selected confidence video frames.

7. The method according to claim 1 or 6, characterized in that, Each dynamic action category corresponds to its own action range, which includes multiple phase actions arranged in chronological order; The categories of dynamic actions of the target object identified in the confidence video frame include: The confidence video frame is input into the trained action prediction model to output the action label and stage label of the confidence video frame; the action label and the stage label are used to characterize the dynamic action category of the target object in the confidence video frame.

8. The method according to claim 7, characterized in that, The method further includes: Confidence video frames with the same action label are grouped into the same dynamic action category, and the confidence video frames in the same dynamic action category are sorted according to their stage labels.

9. A video generation apparatus, characterized in that, The device includes: A trajectory recognition unit is used to acquire a target video and determine a target video frame representing the motion trajectory of a target object in the target video; extract an image of the region where the target object is located from the target video frame and generate quality information of the region image; A video frame filtering unit is configured to identify motion information of the target object in the target video frame, and filter out confidence video frames from the target video frame based on the identified motion information, including: obtaining weight values ​​of the motion information and the quality information respectively, and performing weighted processing on the motion information and the quality information according to the obtained weight values ​​to obtain the confidence level of the target video frame; determining the filtered confidence video frames based on target video frames with a confidence level greater than or equal to a specified threshold; wherein, the motion information is at least used to characterize the pose change of the target object; The video generation unit is configured to identify the dynamic action category of the target object in the confidence video frame, determine the video frame to be edited in the confidence video frame based on the identified dynamic action category, and generate a video segment of the target object based on the video frame to be edited.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store a computer program that, when executed by the processor, implements the method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video processing method and device, computer storage medium and intelligent interaction tablet

    CN115690635A

  • Video editing apparatus and method, movable platform, gimbal, and hardware device

    WO2022104637A1