Video production method and apparatus
By acquiring motion capture data and motion videos, identifying skeletal point data, generating virtual avatar videos and skeletal point videos, and merging them in an editor, the problem of low efficiency in fitness video production in existing technologies is solved, and the effect of automatically assisting users in learning standard movements is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING CALORIE INFORMATION TECH CO LTD
- Filing Date
- 2023-03-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing fitness video production is inefficient, requires a lot of human resources, and is difficult to effectively help users learn standard movements.
By acquiring motion capture data and motion videos, identifying skeletal point data, generating virtual avatar videos and skeletal point videos, and then merging them in an editor, a target video containing judgment information is automatically created.
It enables video creation without human intervention, improves video production efficiency, and can automatically assist users in learning standard actions.
Smart Images

Figure CN116320534B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of video processing technology, and in particular to a video production method. This specification also relates to a video production apparatus, a computing device, and a computer-readable storage medium. Background Technology
[0002] Fitness exercise is a type of physical activity that involves training with bare hands or various equipment, using specific movements and methods to develop muscles, increase physical strength, improve physique, and cultivate character. Fitness exercises are simple and easy to perform; appropriate exercise can effectively enhance physical fitness, improve health, develop overall muscle strength, increase power, and improve productivity. It can also improve body shape and posture, and cultivate positive emotions, making it widely popular.
[0003] With the development of the internet, it has become feasible to exercise without leaving home. However, the standards for movements have gradually become lower, which may not achieve the purpose of exercise. Currently, most exercise videos provided for user reference are produced by humans, which consumes a lot of human resources and is inefficient. Therefore, there is an urgent need for an effective solution to the above problems. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a video production method. This specification also relates to a video production apparatus, a computing device, and a computer-readable storage medium, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a video production method is provided, comprising:
[0006] Acquire motion capture data and motion videos, and identify skeletal point data from the motion videos;
[0007] A virtual avatar video is generated based on the motion capture data, and the skeletal point data is added to the motion video to obtain a skeletal point video;
[0008] Import the virtual avatar video and the skeletal point video into the editor;
[0009] The editor is used to fuse the virtual avatar video and the skeletal point video to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data.
[0010] According to a second aspect of the embodiments of this specification, a video production apparatus is provided, comprising:
[0011] The data acquisition module is configured to acquire motion capture data and motion videos, and to identify the skeletal point data from the motion videos;
[0012] The video generation module is configured to generate a virtual avatar video based on the motion capture data, and to add the skeletal point data to the motion video to obtain a skeletal point video.
[0013] The import module is configured to import the virtual avatar video and the skeletal point video into the editor;
[0014] The video processing module is configured to fuse the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data.
[0015] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0016] Memory and processor;
[0017] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions:
[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the video production method.
[0019] This specification provides a video production method and apparatus. The video production method includes: acquiring motion capture data and motion video, and identifying skeletal point data from the motion video; generating a virtual avatar video based on the motion capture data, and adding the skeletal point data to the motion video to obtain a skeletal point video; importing the virtual avatar video and the skeletal point video into an editor; and fusing the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data. By acquiring motion capture data and motion video, generating a virtual avatar video based on the motion capture data, and generating a skeletal point video based on the skeletal point data, and by fusing the virtual avatar video and the skeletal point video, a target video containing judgment information can be obtained. This achieves automatic video creation without human intervention, thereby improving video creation efficiency and enabling the target video to be used in downstream business to assist users in learning standard movements. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of a video production method provided in one embodiment of this specification;
[0021] Figure 2-1 This is a flowchart illustrating a video production method provided in one embodiment of this specification;
[0022] Figure 2-2 This is a recording diagram illustrating a video production method provided in one embodiment of this specification;
[0023] Figure 3 This is a schematic diagram of an editor illustrating a video production method provided in one embodiment of this specification;
[0024] Figure 4 This is a flowchart illustrating a video production method applied to a server, as provided in one embodiment of this specification.
[0025] Figure 5 This is a schematic diagram of the structure of a video production apparatus provided in one embodiment of this specification;
[0026] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0027] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0028] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0029] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0030] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0031] Image frame: The smallest unit that makes up a video.
[0032] Frame rate: The frequency (rate) at which a bitmap image appears continuously on a display, measured in frames.
[0033] This specification provides a video production method, and also relates to a video production apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.
[0034] See Figure 1 , Figure 1 This is a schematic diagram of a video production method provided in one embodiment of this specification. The embodiment of this specification provides a video production method including: acquiring motion capture data and motion video, identifying skeletal point data from the motion video, generating a virtual avatar video based on the motion capture data, adding the skeletal point data to the motion video to obtain a skeletal point video, importing the virtual avatar video and the skeletal point video into an editor, and fusing the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information.
[0035] The embodiments in this specification acquire motion capture data and motion videos, generate virtual avatar videos based on the motion capture data, and generate skeletal point videos based on skeletal point data. By fusing the virtual avatar videos and skeletal point videos, a target video containing judgment information can be obtained.
[0036] Figure 2-1 A flowchart of a video production method according to an embodiment of this specification is shown, which specifically includes the following steps:
[0037] Step S102: Acquire motion capture data and motion video, and identify the skeletal point data from the motion video.
[0038] Among them, motion capture data refers to motion capture data. Motion capture involves setting trackers on key parts of a moving object, and the data collected by the trackers can be motion capture data; motion video can be a video containing the action to be displayed, such as a video containing dance moves, a video containing martial arts moves, etc.; recognizing the motion video can be understood as using methods such as neural networks to identify the motion video; skeletal point data can be the position, depth, and other data of key points of the human skeleton in the video.
[0039] In practical applications, for recording fitness movements, dancers need to wear motion capture suits and perform dance or fitness movements in a motion capture room, where the entire perimeter can be a green screen. Motion capture data needs to be acquired from the dancers in their suits, and a camera is positioned in front of them to record video of their movements. This video is then analyzed to obtain skeletal point data. This specification does not limit the camera's resolution; for example, the camera resolution can be 640×480, 1920×1080, etc.
[0040] For example, motion capture equipment can be used to acquire motion capture data of dancers in the motion capture room in real time, and video of the dancers' movements can be obtained by cameras. By analyzing the video of the movements, the skeletal data of the dancers can be obtained.
[0041] It should be noted that music can be added during motion capture recording in the motion capture room. For example, the recording setup includes three modules: motion capture equipment, a camera, and music equipment. The motion capture equipment starts recording via a switch, freezing the dancer's movements in a frame before recording begins. Similarly, the camera starts recording via a switch, freezing the dancer's movements in a frame before recording begins. The music is switched off before recording starts. When preparing to start recording, the motion capture equipment and camera recording begin simultaneously, with the dancer remaining frozen in a frame. After a short period (e.g., 2 seconds), the music begins playing, and the dancer starts dancing. This short period is for the learner to prepare for the "3, 2, 1, go" countdown when the video motion begins. After the song recording is complete, the motion capture equipment, camera, and music equipment all stop.
[0042] It should also be noted that the camera can be positioned directly in front of the dancer, recording video of their movements from directly in front of them. Additionally, see [link to other documentation]. Figure 2-2 , Figure 2-2 This is a recording diagram illustrating a video production method provided in one embodiment of this specification. A dancer 204 performs dance movements in a motion capture room 202. A prompting device 206 (such as a screen or whiteboard) can be placed in front of the dancer 204 to indicate the dance movements to be performed. During recording, the dancer 204's mouth can open and close in sync with the lyrics, avoiding a blank facial expression and providing a better viewing experience. After recording, motion capture data and video footage captured by a camera are obtained. Further analysis and identification of the motion video yields skeletal point data. The start and end times of the motion video and skeletal point data are consistent.
[0043] Specifically, machine learning algorithms can be used to analyze motion videos to obtain skeletal point data of dancers.
[0044] In one possible implementation, the step of identifying skeletal point data from the motion video includes:
[0045] The image frame sequence of the action video is obtained, and the image frames in the image frame sequence are identified sequentially according to the human skeleton key point detection algorithm to obtain the skeleton point data of each image frame.
[0046] The image frame sequence can be a sequence of all image frames in the action video. For example, if the frame rate of the action video is 24 frames per second and the action video is 10 seconds long, then the image frame sequence is a sequence of 240 image frames in sequence. The human skeleton key point detection algorithm can be a skeleton key point detection algorithm based on machine learning. This algorithm can use conventional technical means in this field, and will not be described in detail in the embodiments of this specification.
[0047] In practical applications, to identify skeleton points in motion videos, it is necessary to analyze and identify each frame of the motion video to obtain the skeleton point data for each frame.
[0048] For example, if the frame rate of an action video is 24 frames per second and the video lasts for 10 seconds, then the image frame sequence is a sequence of 240 image frames in sequence. By using a human skeleton key point detection algorithm to detect skeleton points on the 240 image frames, the corresponding 240 sets of skeleton point data can be obtained.
[0049] The embodiments in this specification acquire motion capture data and motion videos, and identify skeletal point data from the motion videos to facilitate subsequent labeling of the videos generated from the motion capture data based on the skeletal point data.
[0050] Step S104: Generate a virtual avatar video based on the motion capture data, and add the skeletal point data to the motion video to obtain a skeletal point video.
[0051] Specifically, based on the acquisition of motion capture data and motion video, and the identification of skeletal point data in the motion video, a virtual avatar video can be generated according to the motion capture data, and the skeletal point data can be added to the motion video to obtain a skeletal point video.
[0052] Among them, virtual avatar videos can be videos of rendered virtual characters, such as digital human videos, animated videos, etc.; skeletal point videos can be action videos that include displays of skeletal points, such as skeletal points displayed on a character video in an action video, and skeletal points can include key points such as the head, hands, knees, and feet.
[0053] In practical applications, a video of a virtual character, i.e., a digital human video, can be rendered based on motion capture data. Skeletal points are then displayed in the motion video based on skeletal point data to generate a skeletal point video. Subsequently, the digital human video and the virtual character video can be compared to set the appropriate judgment points.
[0054] For example, motion capture data can be used to generate digital human videos in software, and then the corresponding skeletal point data can be added to the motion video to obtain a skeletal point video that displays the skeletal points.
[0055] This embodiment of the specification generates a virtual avatar video based on the motion capture data, and adds the skeletal point data to the motion video to obtain a skeletal point video. This facilitates comparison between the virtual avatar video and the skeletal point video, improving the accuracy of the judgment point marking.
[0056] In one possible implementation, generating the virtual avatar video based on the motion capture data includes:
[0057] An initial image video is generated based on the motion capture data;
[0058] Receive an effect addition instruction for the initial image video, and update the initial image video according to the effect addition instruction to obtain a virtual image video;
[0059] The virtual avatar video includes display effects corresponding to the effect addition instructions.
[0060] The initial image video mentioned above can be a digital human video generated directly from motion capture data without any added special effects; the effect addition command can be a command to add special effects to the initial image video, such as adding zoom, zoom, shake and other special effects; updating the initial image video according to the effect addition command can be understood as performing special effects processing on the initial image video.
[0061] In practical applications, the camera angles and special effects during virtual avatar video production can be modified and adjusted according to the song and dance moves. For example, the special effects in the climax scene may need to flash, and the camera may need to zoom in or move. Furthermore, two sets of virtual avatar videos can be output, with the same camera angle and timing. One video is a normal video (with special effects and a background), while the other video can have a white character, a completely black textured background, and no special effects. This latter video can be used for background removal and selection.
[0062] For example, firstly, an initial avatar video is generated based on the obtained motion capture data. Then, effects addition instructions are obtained for the initial avatar video. These instructions include: adding a shaking effect at the 10th second of the initial avatar video, adding a zoom-in effect at the 20th second, and adding a shrinking effect at the 21st second. The initial avatar video is then processed according to these effects addition instructions to ultimately obtain the virtual avatar video.
[0063] It should be noted that, to facilitate subsequent decision point marking, a video without special effects can also be generated. This improves efficiency when marking decision points in the video.
[0064] In summary, by adding display effects to the initial avatar video, the virtual avatar video can be made more visually appealing, and the resulting target video can be more personalized, thereby improving the viewing experience for users.
[0065] In one possible implementation, adding the skeletal point data to the motion video to obtain a skeletal point video includes:
[0066] Extract the position information of bone points from the bone point data of each image frame;
[0067] Based on the location information, the corresponding skeletal points are labeled in the corresponding image frames. Once the labeling is completed for each image frame, a skeletal point video is obtained.
[0068] The location information can be the coordinates of the skeleton point in the image, for example, the location information is (456, 789); marking the corresponding skeleton point in the corresponding image frame can be understood as marking the skeleton point in the corresponding image frame so that the skeleton point is displayed in the image frame.
[0069] In practical applications, it is necessary to annotate the corresponding image frames based on the skeletal point data of each image frame so that the skeletal points can be displayed in the skeletal point video.
[0070] For example, if the frame rate of the action video is 24 frames per second and the action video is 10 seconds long, then the image frame sequence is a sequence of 240 image frames in sequence. By using the human skeleton key point detection algorithm to detect skeleton points in the 240 image frames, the corresponding 240 sets of skeleton point data can be obtained. Based on the 240 sets of skeleton point data, the skeleton points are labeled in the 240 image frames.
[0071] This embodiment of the specification extracts the position information of the skeleton points in the skeleton point data of each image frame, and marks the corresponding skeleton points in the corresponding image frame according to the position information. When the annotation of each image frame is completed, a skeleton point video is obtained, so that the skeleton points in the skeleton point video are visible. Based on this, the target video is created, which can effectively improve the quality of the target video.
[0072] Step S106: Import the virtual avatar video and the skeletal point video into the editor.
[0073] Specifically, based on the above process of generating a virtual avatar video from the motion capture data and adding the skeletal point data to the motion video to obtain a skeletal point video, the virtual avatar video and the skeletal point video can be imported into the editor.
[0074] The editor can be a container for editing virtual avatar videos and skeletal point videos, such as video editing software.
[0075] In practical applications, after obtaining the virtual avatar video and the skeletal point video, both can be imported into editing software with the same interface for easy marking of judgment points. Marking judgment points can be understood as marking judgment points in the virtual avatar video while referring to the skeletal point video. These judgment points are used to subsequently judge the actions of the person following the movements in the video.
[0076] See Figure 3 , Figure 3 This is a schematic diagram of an editor for a video production method provided in one embodiment of this specification. The editor may include a video area, in which digital human video 302 and skeletal point video 304 can be imported. Specifically, the left side can be digital human video 302, and the right side can be skeletal point video 304. This embodiment of the specification does not limit this. Digital human video 302 and skeletal point video 304 will synchronously change as they are dragged and played along the timeline 306 below. Seconds and milliseconds are displayed below the video. The millisecond and second displays can be converted to each other, that is, the time unit can be set.
[0077] The marking of judgment point 314 can be edited on the digital human video 302. Specifically, you can click on the video with the mouse, and a draggable circular icon (20*20 pixels) will appear at the clicked location. After clicking and dragging the icon, an area can be edited to display the following information: id (identifier): the ID of this judgment point 314 on this spectrum; position information: the position information displayed in (x, y) format, where the unit of (x, y) can be pixels; judgment type: what kind of judgment is performed in this area, such as touch, stay, slide, etc. Other parameters can also be set: special parameters of the judgment, which need to be integrated after all judgment types are set. For example, the display size of judgment point 314, the stay time of judgment point 314, the sliding speed, etc. It also includes the time axis 306 area: the time axis 306 can be dragged left and right, which will change the display screen of digital human video 302 and skeletal point video 304. The minimum granularity is 50 milliseconds or frames, which is not limited in this embodiment. Frame rate information, such as 24 frames, can also be displayed.
[0078] Furthermore, songs can be imported. For example, to add a custom song, you can click "Import Song." Correspondingly, an audio timeline 308 is also included, and the song can be edited. The interface may also include a button area: a play / pause button 310 and an export to Excel button 312. The play and pause buttons 310 can play and pause the video from the current timeline position 306, while the export to Excel button 312 can export the edited decision points 314 in chronological order into an Excel spreadsheet, clearly displaying the order of the decision points 314.
[0079] In order to accurately add decision points to the virtual avatar video, it is necessary to ensure that the frame rate of the virtual avatar video and the skeletal point video are the same. The specific implementation method is as follows.
[0080] Before importing the virtual avatar video and the skeletal point video into the editor, the process also includes:
[0081] The frame rate of the virtual avatar video and the frame rate of the skeletal point video are adapted.
[0082] Accordingly, importing the virtual avatar video and the skeletal point video into the editor includes:
[0083] Import the adapted virtual avatar video and the skeletal point video into the editor.
[0084] The adaptation of the frame rate of the virtual avatar video and the frame rate of the skeletal point video can be understood as making the frame rate of the virtual avatar video the same as the frame rate of the skeletal point video.
[0085] In practical applications, to accurately add decision points to the virtual avatar video, it is necessary to ensure that the frame rates of the virtual avatar video and the skeletal point video are the same. This can be achieved by setting a fixed frame rate to adapt to the frame rates of the virtual avatar video and the skeletal point video, or by using the frame rate of the skeletal point video as a baseline. The specific implementation methods for these two frame rate adaptation methods are described below.
[0086] (1) In one possible implementation, adapting the frame rate of the virtual avatar video and the frame rate of the skeletal point video includes:
[0087] If the frame rate of the virtual avatar video is not the preset frame rate, the frame rate of the virtual avatar video is processed according to the preset frame rate.
[0088] If the frame rate of the skeletal point video is not a preset frame rate, the frame rate of the skeletal point video is processed according to the preset frame rate.
[0089] The preset frame rate can be understood as a pre-set frame rate, meaning that the frame rate of both the virtual character video and the skeletal point video must meet this frame rate.
[0090] If the frame rate of the virtual avatar video is not the preset frame rate, it is impossible to accurately add judgment points to the virtual avatar video. In order to accurately add judgment points to the virtual avatar video, it is necessary to ensure that the frame rates of the virtual avatar video and the skeletal point video are the same. Therefore, the frame rates of the virtual avatar video and the skeletal point video can be adjusted to the same frame rate.
[0091] For example, if the preset frame rate is 30 frames per second (fps), and the frame rate of the virtual avatar video is 24 fps, then the frame rate of the virtual avatar video does not meet the preset frame rate. Therefore, frame rate processing is performed on the virtual avatar video, adjusting its frame rate to 30 fps. Similarly, if the frame rate of the skeletal point video is 60 fps, then the frame rate of the skeletal point video does not meet the preset frame rate. Therefore, frame rate processing is performed on the skeletal point video, adjusting its frame rate to 30 fps.
[0092] In summary, by performing frame rate adaptation processing on the virtual avatar video and the skeletal point video respectively, the video can be aligned at a finer granular level when subsequent judgment information is labeled, thereby ensuring the accuracy of the added judgment information.
[0093] In one possible implementation, frame rate processing is performed on the virtual character video or the skeletal point video according to the preset frame rate, including:
[0094] If the frame rate of the virtual avatar video is less than the preset frame rate, the virtual avatar video is subjected to frame interpolation; or if the frame rate of the virtual avatar video is greater than the preset frame rate, the virtual avatar video is subjected to frame extraction.
[0095] If the frame rate of the skeletal point video is less than the preset frame rate, the skeletal point video is subjected to frame interpolation; if the frame rate of the skeletal point video is greater than the preset frame rate, the skeletal point video is subjected to frame extraction.
[0096] Frame interpolation can be understood as adding image frames to a video, while frame extraction can be understood as reducing the number of image frames in a video.
[0097] Continuing with the previous example, the preset frame rate is 30 frames per second (fps). If the virtual avatar video's frame rate is 24 fps, then the virtual avatar video's frame rate does not meet the preset frame rate. Therefore, frame rate processing is performed on the virtual avatar video, adjusting its frame rate to 30 fps. Specifically, this can be done by adding image frames evenly throughout the virtual avatar video, for example, adding a blank frame every four image frames. Similarly, if the skeletal point video has a frame rate of 60 fps, then the skeletal point video's frame rate does not meet the preset frame rate. Frame rate processing is also performed on the skeletal point video, adjusting its frame rate to 30 fps. Specifically, this can be done by removing image frames evenly throughout the skeletal point video, for example, removing one image frame every one image frame.
[0098] (2) In addition to the above frame rate adaptation method, the frame rate of the virtual image video can also be used to adapt the frame rate of the skeletal point video. The specific implementation method is as follows.
[0099] The adaptation of the frame rate of the virtual avatar video and the frame rate of the skeletal point video includes:
[0100] The frame rate of the skeletal point video is determined, and if the frame rate of the virtual avatar video and the frame rate of the skeletal point video do not match, the frame rate of the virtual avatar video is processed.
[0101] In practical applications, since the motion video is captured by a camera and has a fixed frame rate, the number of skeletal point data sets extracted from it is also fixed. Therefore, the frame rate of the virtual avatar video can be adapted to the frame rate of the skeletal point video, that is, to make the frame rate of the virtual avatar video and the frame rate of the skeletal point video the same.
[0102] In one possible implementation, frame rate processing of the virtual avatar includes:
[0103] If the frame rate of the virtual avatar video is less than the frame rate of the skeletal point video, the virtual avatar video is subjected to frame interpolation.
[0104] If the frame rate of the virtual avatar video is greater than the frame rate of the skeletal point video, then the virtual avatar video is subjected to frame extraction.
[0105] Using the previous example, the frame rate of the virtual avatar video is 24 frames per second, and the frame rate of the skeletal point video is 30 frames per second. Therefore, the frame rate of the virtual avatar video does not meet the frame rate of the skeletal point video. Frame rate processing is performed on the virtual avatar video, that is, the frame rate of the virtual avatar video is adjusted to 30 frames per second. Specifically, image frames can be added evenly in the virtual avatar video, such as adding an empty frame every 4 image frames.
[0106] For example, if the frame rate of the virtual avatar video is 40 frames per second and the frame rate of the skeletal point video is 30 frames per second, then the frame rate of the virtual avatar video does not meet the preset frame rate. The frame rate of the virtual avatar video needs to be processed, that is, the frame rate of the virtual avatar video is adjusted to 30 frames per second. Specifically, image frames can be removed evenly in the virtual avatar video, such as removing one image frame every three image frames.
[0107] Before importing the virtual avatar video and skeletal point video into the editor, the frame rate of the video is adjusted in this embodiment of the specification, which improves the accuracy of merging the virtual avatar video and skeletal point video.
[0108] Step S108: The virtual avatar video and the skeletal point video are fused using the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data.
[0109] Specifically, based on the above-mentioned import of the virtual avatar video and the skeletal point video into the editor, the virtual avatar video and the skeletal point video can be merged through the editor to obtain a target video containing judgment information.
[0110] The above fusion can be understood as mutual comparison; the judgment information can be understood as the information of the judgment point, such as ID, location information, judgment type, etc.
[0111] In one possible implementation, the step of fusing the virtual avatar video and the skeletal point video through the editor to obtain a target video containing determination information includes:
[0112] Receive a video progress adjustment instruction, and determine the target number of frames for the virtual character video and the skeletal point video based on the time identifier carried in the video adjustment instruction;
[0113] Display the image frames corresponding to the virtual avatar video and the image frames corresponding to the skeletal point video according to the target frame number.
[0114] Among them, the video progress adjustment command can be understood as the command generated by dragging the progress bar of the video; the time identifier can be the time identifier of the video progress bar, such as the 1000th millisecond of the video; the target frame number can be the frame number to be located, such as the 50th frame.
[0115] For example, if the time marker in the video progress adjustment instruction is 2000 milliseconds, then the 2000th millisecond is determined to be the 50th frame, and the 50th frame of the virtual avatar video and the 50th frame of the skeletal point video are displayed.
[0116] This embodiment of the specification receives a video progress adjustment command, and determines the target frame number of the virtual avatar video and the skeletal point video based on the time identifier carried in the video adjustment command, so as to accurately display the virtual avatar video and the skeletal point video accordingly.
[0117] In one possible implementation, the step of fusing the virtual avatar video and the skeletal point video through the editor to obtain a target video containing determination information includes:
[0118] Receive a decision point addition instruction submitted for a virtual avatar video in the editor, and generate a target video containing decision information based on the action decision data carried in the decision point addition instruction and the virtual avatar video.
[0119] The action determination data includes the skeletal point identifiers, action types, and action determination thresholds from the skeletal point data.
[0120] The above-mentioned decision point addition instruction can be used to add decision points to a virtual avatar video.
[0121] For example, if the time marker in the received video progress adjustment instruction is 2000 milliseconds, then the 2000th millisecond is determined to be the 50th frame. Therefore, the 50th frame of the virtual avatar video and the 50th frame of the skeletal point video are displayed. The instruction to add a judgment point for the 50th frame of the virtual avatar video is received, i.e., a circular icon is generated, which can be 20*20 pixels in size. The action judgment data corresponding to the icon includes: id (identifier): the ID of this judgment point in this image frame, and may also include the corresponding skeletal point identifier; position information: position information displayed in (x, y) format, where the unit of (x, y) can be pixels; judgment type (action type): what kind of judgment is performed in this area, such as: touch, stay, slide, etc. Other parameters can also be set: special parameters of the judgment, which need to be integrated after all judgment types are completed. For example, the judgment point display size is 50*50 pixels, and the judgment point stay time is 2 seconds, etc.
[0122] This embodiment of the specification receives a judgment point addition instruction submitted for a virtual avatar video in the editor, and generates a target video containing judgment information based on the action judgment data carried in the judgment point addition instruction and the virtual avatar video. Since the judgment points are obtained based on the skeletal point video, the judgment information has high accuracy.
[0123] In one possible implementation, the step of generating a target video containing judgment information based on the action judgment data carried in the judgment point addition instruction and the virtual avatar video includes:
[0124] The frame to be processed in the virtual character video and the position of the judgment point of the frame to be processed are determined based on the action judgment data.
[0125] Establish the correspondence between the determination point location and the corresponding skeletal point identifier, and establish the correspondence between the determination point location and the action type;
[0126] The action determination region of the frame to be processed is determined based on the action determination threshold and the determination point position, and a target video containing determination information is generated after all frames to be processed in the virtual avatar video have been processed.
[0127] Following the previous example, the motion determination data for frame 50 includes: id (identifier): the id of this determination point in this image frame, which may also include the corresponding skeleton point identifier; position information: position information displayed in (x, y) format, where the unit of (x, y) can be pixels; determination type (motion type): what kind of determination is performed in this area, such as: touch, stay, slide, etc. Other parameters can also be set: special parameters for determination, which need to be integrated after all determination types are completed. For example, the display size of the determination point is 50*50 pixels, and the stay time of the determination point is 2 seconds, etc. Establish the correspondence between id and the corresponding skeleton point identifier, and establish the correspondence between id and the corresponding determination type, with the determination area being 50*50 pixels. Process all image frames sequentially according to the above method to finally obtain the target video containing determination information.
[0128] After obtaining the target video containing the judgment information, the target video can be used to judge the user's video. The specific implementation method is as follows.
[0129] After obtaining the target video containing the determination information, the process further includes:
[0130] Play a target video to the user, wherein the target video contains an action determination area;
[0131] Collect user action videos associated with the target video, wherein the user actions in the user action videos are associated with standard actions in the target video;
[0132] If the target video is determined to play to a preset judgment interval, the skeletal point acquisition area corresponding to the action judgment area is determined in the user action video.
[0133] Based on the preset judgment interval and the skeletal point acquisition area, user skeletal point data is acquired in the user motion video;
[0134] Based on the standard skeletal point data associated with the action determination area, the user's skeletal point data is used to determine the action, and the action determination result is displayed.
[0135] In practical applications, a standard action video is played to the user on the terminal. This standard action video includes an action judgment area, allowing the user to subsequently perform corresponding actions based on the instructions in the standard action video. It should be noted that the action judgment method provided in this specification can play the standard action video through a video display device associated with the terminal. The video display device associated with the terminal can be understood as a video display device configured on the terminal, such as a mobile phone or laptop computer, and the video display device configured on the terminal can be understood as the screen on the mobile phone or the screen on the laptop computer, etc. Alternatively, it can be a video display device independent of the terminal but with a communication connection to it; for example, the terminal can be understood as a host computer or server, and the video display device configured on the terminal can be understood as a monitor, television, projector, or other device capable of displaying video that communicates with the host computer, or a monitor, television, projector, or other device capable of displaying video that communicates with the server, etc.
[0136] In this context, "terminal" can be understood as any device capable of implementing the action determination method, such as a user's mobile phone, computer, server, host, etc., without specific limitations in this specification. "Standard action video" can be understood as a video demonstrating a certain standard action, which the user can follow to perform the corresponding action. For example, in a teaching scenario, the standard action video could be a dance training video, a fitness tutorial video, a business etiquette training video, etc. Users can learn various skills based on this video; similarly, in a gaming scenario, the standard action video could be a dance video, a fitness video, etc., allowing users to imitate the standard actions in the video while playing the game.
[0137] The action judgment area can be understood as the region where action judgment is performed. Based on this area, users can understand the key points to focus on in the current or upcoming actions. In practical applications, if a standard action video only shows the standard movement, without highlighting the action judgment area for key parts of the standard movement during playback, users will be unable to learn or play in a targeted manner, resulting in low learning efficiency or a poor gaming experience. For example, in a push-up tutorial video, if the shoulder or arm movements are not emphasized, users will not be able to quickly grasp the key arm or shoulder movements while learning push-ups, leading to poor learning efficiency. Similarly, in a dance game video, without indicating the next movement, users may not be able to keep up with the rhythm of the dance, resulting in a poor gaming experience.
[0138] During the playback of a standard action video, this embodiment can synchronously capture user action video associated with the standard action video, wherein the user actions in the user action video are associated with the standard actions in the standard action video. It should be noted that the user action is an action performed by the user based on the standard actions in the standard action video. Furthermore, the action determination method provided in this specification can synchronously capture user action video through a video capture device associated with the terminal. The video capture device associated with the terminal can be understood as a video capture device configured on the terminal, such as a mobile phone, laptop, or other device, and the video capture device configured on the terminal can be understood as a camera on a mobile phone or a camera on a laptop, etc.; or a video capture device independent of the terminal but with a communication connection to it. For example, the terminal can be understood as a host, server, or other device, and the video capture device configured on the terminal can be understood as a camera communicating with a host or a camera communicating with a server, etc. This communication connection includes wireless and wired connections.
[0139] When the standard motion video is played to a preset judgment interval, the system can determine the skeletal point acquisition area corresponding to the motion judgment area in the user's motion video, and collect user skeletal point data in the user's motion video based on this skeletal point acquisition area. Then, based on the standard skeletal point data associated with the motion judgment area, the system performs motion judgment on the user's skeletal point data to obtain the judgment result, which is then displayed.
[0140] This embodiment of the specification determines the skeletal point acquisition area from the user's action video based on the preset judgment interval and the action judgment area in the standard action video when the standard action video is played to the preset judgment interval. Then, based on the standard skeletal point data associated with the action judgment area, it performs action judgment on the user's skeletal point data corresponding to the skeletal point acquisition area, thereby quickly and accurately identifying the difference between the user's current action and the standard action. The action judgment result is then displayed, allowing the user to adjust their actions accordingly, further improving the efficiency of skill learning or physical exercise.
[0141] This specification provides a video production method and apparatus. The video production method includes: acquiring motion capture data and motion video, and identifying skeletal point data from the motion video; generating a virtual avatar video based on the motion capture data, and adding the skeletal point data to the motion video to obtain a skeletal point video; importing the virtual avatar video and the skeletal point video into an editor; and fusing the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data. By acquiring motion capture data and motion video, generating a virtual avatar video based on the motion capture data, and generating a skeletal point video based on the skeletal point data, and by fusing the virtual avatar video and the skeletal point video, a target video containing judgment information can be obtained.
[0142] The following is in conjunction with the appendix Figure 4 Taking the application of the video production method provided in this specification on a server as an example, the video production method will be further explained. Figure 4 This specification illustrates a processing flowchart of a video production method applied to a server, according to an embodiment of the present invention, which specifically includes the following steps:
[0143] Step S402: Acquire motion capture data and motion video.
[0144] For example, motion capture equipment can be used to acquire motion capture data of dancers in the motion capture room in real time, and video of the dancers' movements can be obtained by cameras. By analyzing the video of the movements, the skeletal data of the dancers can be obtained.
[0145] Step S404: Obtain the image frame sequence of the action video, and identify the image frames in the image frame sequence in sequence according to the human skeleton key point detection algorithm to obtain the skeleton point data of each image frame.
[0146] For example, if the frame rate of an action video is 24 frames per second and the video lasts for 10 seconds, then the image frame sequence is a sequence of 240 image frames in sequence. By using a human skeleton key point detection algorithm to detect skeleton points on the 240 image frames, the corresponding 240 sets of skeleton point data can be obtained.
[0147] Step S406: Generate an initial image video based on the motion capture data; receive an effect addition instruction for the initial image video, and update the initial image video according to the effect addition instruction to obtain a virtual image video.
[0148] For example, firstly, an initial image video is generated based on the obtained motion capture data, and then an effect addition instruction is obtained for the initial image video. This effect addition instruction includes: adding a shaking effect at the 10th second of the initial image video.
[0149] Step S408: Extract the position information of the bone points in the bone point data of each image frame; mark the corresponding bone points in the corresponding image frame according to the position information, and obtain the bone point video after the marking of each image frame is completed.
[0150] For example, if the frame rate of the action video is 24 frames per second and the action video is 10 seconds long, then the image frame sequence is a sequence of 240 image frames in sequence. By using the human skeleton key point detection algorithm to detect skeleton points in the 240 image frames, the corresponding 240 sets of skeleton point data can be obtained. Based on the 240 sets of skeleton point data, the skeleton points are labeled in the 240 image frames.
[0151] Step S410: Adapt the frame rate of the virtual avatar video and the frame rate of the skeletal point video, and import the adapted virtual avatar video and the skeletal point video into the editor.
[0152] For example, if the preset frame rate is 30 frames per second (fps), and the frame rate of the virtual avatar video is 24 fps, then the frame rate of the virtual avatar video does not meet the preset frame rate. Therefore, frame rate processing is performed on the virtual avatar video, adjusting its frame rate to 30 fps. Similarly, if the frame rate of the skeletal point video is 24 fps, then the frame rate of the skeletal point video does not meet the preset frame rate. Therefore, frame rate processing is performed on the skeletal point video, adjusting its frame rate to 30 fps.
[0153] Step S412: Receive a judgment point addition instruction submitted for the virtual avatar video in the editor, and generate a target video containing judgment information based on the action judgment data carried in the judgment point addition instruction and the virtual avatar video.
[0154] For example, if the time marker in the received video progress adjustment instruction is 2000 milliseconds, then the 2000th millisecond is determined to be the 50th frame. Therefore, the 50th frame of the virtual avatar video and the 50th frame of the skeletal point video are displayed. The instruction to add a judgment point for the 50th frame of the virtual avatar video is received, i.e., a circular icon is generated, which can be 20*20 pixels in size. The action judgment data corresponding to the icon includes: id (identifier): the ID of this judgment point in this image frame, and may also include the corresponding skeletal point identifier; position information: position information displayed in (x, y) format, where the unit of (x, y) can be pixels; judgment type (action type): what kind of judgment is performed in this area, such as: touch, stay, slide, etc. Other parameters can also be set: special parameters of the judgment, which need to be integrated after all judgment types are completed. For example, the judgment point display size is 50*50 pixels, and the judgment point stay time is 2 seconds, etc.
[0155] Step S414: Collect user action video associated with the target video, collect user skeleton point data in the user action video, perform action determination on the user skeleton point data, and display the action determination result.
[0156] For example, if user action video associated with the target video is collected, and user skeleton point data is collected in the user action video, and the coordinates of the user's skeleton points are all within a range of 50*50 pixels centered at (x, y), then the user's action is determined to be: standard.
[0157] The embodiments in this specification fuse virtual avatar videos and skeletal point videos to obtain target videos containing judgment information. The judgment information contained in the target videos is obtained by fusing virtual avatar videos and skeletal point videos, thereby improving the accuracy of the judgment information in the target videos.
[0158] Corresponding to the above method embodiments, this specification also provides embodiments of a video production apparatus. Figure 5 A schematic diagram of a video production apparatus according to an embodiment of this specification is shown. Figure 5 As shown, the device includes:
[0159] The data acquisition module 502 is configured to acquire motion capture data and motion video, and to identify the motion video to obtain skeletal point data;
[0160] The video generation module 504 is configured to generate a virtual avatar video based on the motion capture data, and to add the skeletal point data to the motion video to obtain a skeletal point video.
[0161] Import module 506 is configured to import the virtual avatar video and the skeletal point video into the editor;
[0162] The video processing module 508 is configured to fuse the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data.
[0163] In one possible implementation, the data acquisition module 502 is further configured as follows:
[0164] The image frame sequence of the action video is obtained, and the image frames in the image frame sequence are identified sequentially according to the human skeleton key point detection algorithm to obtain the skeleton point data of each image frame.
[0165] In one possible implementation, the video generation module 504 is further configured as follows:
[0166] Extract the position information of bone points from the bone point data of each image frame;
[0167] Based on the location information, the corresponding skeletal points are labeled in the corresponding image frames. Once the labeling is completed for each image frame, a skeletal point video is obtained.
[0168] In one possible implementation, the video generation module 504 is further configured as follows:
[0169] An initial image video is generated based on the motion capture data;
[0170] Receive an effect addition instruction for the initial image video, and update the initial image video according to the effect addition instruction to obtain a virtual image video;
[0171] The virtual avatar video includes display effects corresponding to the effect addition instructions.
[0172] In one possible implementation, the video processing module 508 is further configured as follows:
[0173] Receive a decision point addition instruction submitted for a virtual avatar video in the editor, and generate a target video containing decision information based on the action decision data carried in the decision point addition instruction and the virtual avatar video.
[0174] The action determination data includes the skeletal point identifiers, action types, and action determination thresholds from the skeletal point data.
[0175] In one possible implementation, the video processing module 508 is further configured as follows:
[0176] The frame to be processed in the virtual character video and the position of the judgment point of the frame to be processed are determined based on the action judgment data.
[0177] Establish the correspondence between the determination point location and the corresponding skeletal point identifier, and establish the correspondence between the determination point location and the action type;
[0178] The action determination region of the frame to be processed is determined based on the action determination threshold and the determination point position, and a target video containing determination information is generated after all frames to be processed in the virtual avatar video have been processed.
[0179] In one possible implementation, the video processing module 508 is further configured as follows:
[0180] The frame rate of the virtual avatar video and the frame rate of the skeletal point video are adapted.
[0181] Accordingly, importing the virtual avatar video and the skeletal point video into the editor includes:
[0182] Import the adapted virtual avatar video and the skeletal point video into the editor.
[0183] In one possible implementation, the video processing module 508 is further configured as follows:
[0184] If the frame rate of the virtual avatar video is not the preset frame rate, the frame rate of the virtual avatar video is processed according to the preset frame rate.
[0185] If the frame rate of the skeletal point video is not a preset frame rate, the frame rate of the skeletal point video is processed according to the preset frame rate.
[0186] In one possible implementation, the video processing module 508 is further configured as follows:
[0187] If the frame rate of the virtual avatar video is less than the preset frame rate, the virtual avatar video is subjected to frame interpolation; or if the frame rate of the virtual avatar video is greater than the preset frame rate, the virtual avatar video is subjected to frame extraction.
[0188] If the frame rate of the skeletal point video is less than the preset frame rate, the skeletal point video is subjected to frame interpolation; if the frame rate of the skeletal point video is greater than the preset frame rate, the skeletal point video is subjected to frame extraction.
[0189] In one possible implementation, the video processing module 508 is further configured as follows:
[0190] The frame rate of the skeletal point video is determined, and if the frame rate of the virtual avatar video and the frame rate of the skeletal point video do not match, the frame rate of the virtual avatar video is processed.
[0191] In one possible implementation, the video processing module 508 is further configured as follows:
[0192] If the frame rate of the virtual avatar video is less than the frame rate of the skeletal point video, the virtual avatar video is subjected to frame interpolation.
[0193] If the frame rate of the virtual avatar video is greater than the frame rate of the skeletal point video, then the virtual avatar video is subjected to frame extraction.
[0194] In one possible implementation, the video processing module 508 is further configured as follows:
[0195] Receive a video progress adjustment instruction, and determine the target number of frames for the virtual character video and the skeletal point video based on the time identifier carried in the video adjustment instruction;
[0196] Display the image frames corresponding to the virtual avatar video and the image frames corresponding to the skeletal point video according to the target frame number.
[0197] In one possible implementation, the video processing module 508 is further configured as follows:
[0198] Play a target video to the user, wherein the target video contains an action determination area;
[0199] Collect user action videos associated with the target video, wherein the user actions in the user action videos are associated with standard actions in the target video;
[0200] If the target video is determined to play to a preset judgment interval, the skeletal point acquisition area corresponding to the action judgment area is determined in the user action video.
[0201] Based on the preset judgment interval and the skeletal point acquisition area, user skeletal point data is acquired in the user motion video;
[0202] Based on the standard skeletal point data associated with the action determination area, the user's skeletal point data is used to determine the action, and the action determination result is displayed.
[0203] This specification provides a video production method and apparatus. The video production apparatus includes: a data acquisition module configured to acquire motion capture data and motion video, and to identify skeletal point data from the motion video; a video generation module configured to generate a virtual avatar video based on the motion capture data, and to add the skeletal point data to the motion video to obtain a skeletal point video; an import module configured to import the virtual avatar video and the skeletal point video into an editor; and a video processing module configured to fuse the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information; wherein the judgment information corresponds to the skeletal point data. By acquiring motion capture data and motion video, generating a virtual avatar video based on the motion capture data, and generating a skeletal point video based on the skeletal point data, and by fusing the virtual avatar video and the skeletal point video, a target video containing judgment information can be obtained.
[0204] The above is an illustrative scheme of a video production apparatus according to this embodiment. It should be noted that the technical solution of this video production apparatus and the technical solution of the video production method described above belong to the same concept. For details not described in detail in the technical solution of the video production apparatus, please refer to the description of the technical solution of the video production method described above.
[0205] Figure 6 A structural block diagram of a computing device 600 according to an embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0206] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0207] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0208] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 600 can also be a mobile or stationary server.
[0209] The processor 620 is used to execute the following computer-executable instructions: when executed by the processor, these computer-executable instructions implement the steps of the above-described video production method.
[0210] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the video production method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the video production method described above.
[0211] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, are used to implement the steps of the video production method described above.
[0212] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium and the technical solution of the video production method described above belong to the same concept. Details not described in detail in the technical solution of the storage medium can be found in the description of the technical solution of the video production method described above.
[0213] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0214] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0215] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.
[0216] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0217] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A video production method, characterized in that, include: Acquire motion capture data and motion video, and identify skeletal point data from the motion video. The motion capture data is collected by setting trackers at key parts of the moving object. A virtual avatar video is generated based on the motion capture data, and the skeletal point data is added to the motion video to obtain a skeletal point video; Import the virtual avatar video and the skeletal point video into the editor; The editor merges the virtual avatar video and the skeletal point video to obtain a target video containing judgment information. This fusion process includes receiving a judgment point addition instruction submitted for the virtual avatar video in the editor; determining the frames to be processed in the virtual avatar video and the judgment point positions of the frames to be processed based on action judgment data; establishing a correspondence between the judgment point positions and their corresponding skeletal point identifiers, and establishing a correspondence between the judgment point positions and action types; determining the action judgment region of the frames to be processed based on an action judgment threshold and the judgment point positions; and generating a target video containing judgment information after all frames to be processed in the virtual avatar video have been processed. The judgment information corresponds to the skeletal point data, and the action judgment data includes the skeletal point identifiers, action types, and action judgment thresholds from the skeletal point data.
2. The method according to claim 1, characterized in that, The process of identifying the skeletal point data from the motion video includes: The image frame sequence of the action video is obtained, and the image frames in the image frame sequence are identified sequentially according to the human skeleton key point detection algorithm to obtain the skeleton point data of each image frame.
3. The method according to claim 2, characterized in that, The step of adding the skeletal point data to the motion video to obtain the skeletal point video includes: Extract the position information of bone points from the bone point data of each image frame; Based on the location information, the corresponding skeletal points are labeled in the corresponding image frames. Once the labeling is completed for each image frame, a skeletal point video is obtained.
4. The method according to claim 1, characterized in that, The step of generating a virtual avatar video based on the motion capture data includes: An initial image video is generated based on the motion capture data; Receive an effect addition instruction for the initial image video, and update the initial image video according to the effect addition instruction to obtain a virtual image video; The virtual avatar video includes display effects corresponding to the effect addition instructions.
5. The method according to claim 1, characterized in that, Before importing the virtual avatar video and the skeletal point video into the editor, the process also includes: The frame rate of the virtual avatar video and the frame rate of the skeletal point video are adapted. Accordingly, importing the virtual avatar video and the skeletal point video into the editor includes: Import the adapted virtual avatar video and the skeletal point video into the editor.
6. The method according to claim 5, characterized in that, The adaptation of the frame rate of the virtual avatar video and the frame rate of the skeletal point video includes: If the frame rate of the virtual avatar video is not the preset frame rate, the frame rate of the virtual avatar video is processed according to the preset frame rate. If the frame rate of the skeletal point video is not a preset frame rate, the frame rate of the skeletal point video is processed according to the preset frame rate.
7. The method according to claim 6, characterized in that, Frame rate processing is performed on the virtual character video or the skeletal point video according to the preset frame rate, including: If the frame rate of the virtual avatar video is less than the preset frame rate, the virtual avatar video is subjected to frame interpolation; or if the frame rate of the virtual avatar video is greater than the preset frame rate, the virtual avatar video is subjected to frame extraction. If the frame rate of the skeletal point video is less than the preset frame rate, the skeletal point video is subjected to frame interpolation; if the frame rate of the skeletal point video is greater than the preset frame rate, the skeletal point video is subjected to frame extraction.
8. The method according to claim 5, characterized in that, The adaptation of the frame rate of the virtual avatar video and the frame rate of the skeletal point video includes: The frame rate of the skeletal point video is determined, and if the frame rate of the virtual avatar video and the frame rate of the skeletal point video do not match, the frame rate of the virtual avatar video is processed.
9. The method according to claim 8, characterized in that, Frame rate processing of the virtual avatar includes: If the frame rate of the virtual avatar video is less than the frame rate of the skeletal point video, the virtual avatar video is subjected to frame interpolation. If the frame rate of the virtual avatar video is greater than the frame rate of the skeletal point video, then the virtual avatar video is subjected to frame extraction.
10. The method according to claim 1, characterized in that, The process of fusing the virtual avatar video and the skeletal point video using the editor to obtain a target video containing judgment information includes: Receive a video progress adjustment instruction, and determine the target number of frames for the virtual character video and the skeletal point video based on the time identifier carried in the video progress adjustment instruction; Display the image frames corresponding to the virtual avatar video and the image frames corresponding to the skeletal point video according to the target frame number.
11. The method according to claim 1, characterized in that, After obtaining the target video containing the determination information, the process further includes: Play a target video to the user, wherein the target video contains an action determination area; Collect user action videos associated with the target video, wherein the user actions in the user action videos are associated with standard actions in the target video; If the target video is determined to play to a preset judgment interval, the skeletal point acquisition area corresponding to the action judgment area is determined in the user action video. Based on the preset judgment interval and the skeletal point acquisition area, user skeletal point data is acquired in the user motion video; Based on the standard skeletal point data associated with the action determination area, the user's skeletal point data is used to determine the action, and the action determination result is displayed.
12. A video production apparatus, characterized in that, include: The data acquisition module is configured to acquire motion capture data and motion video, and to identify skeletal point data from the motion video. The motion capture data is collected by setting trackers at key parts of the moving object. The video generation module is configured to generate a virtual avatar video based on the motion capture data, and to add the skeletal point data to the motion video to obtain a skeletal point video. The import module is configured to import the virtual avatar video and the skeletal point video into the editor; The video processing module is configured to fuse the virtual avatar video and the skeletal point video through the editor to obtain a target video containing judgment information. The process of fusing the virtual avatar video and the skeletal point video through the editor to obtain the target video containing judgment information includes: receiving a judgment point addition instruction submitted for the virtual avatar video in the editor; determining the frames to be processed in the virtual avatar video and the judgment point positions of the frames to be processed based on action judgment data; establishing a correspondence between the judgment point positions and corresponding skeletal point identifiers, and establishing a correspondence between the judgment point positions and action types; determining the action judgment region of the frames to be processed based on an action judgment threshold and the judgment point positions; and generating a target video containing judgment information after all frames to be processed in the virtual avatar video have been processed. The judgment information corresponds to the skeletal point data, and the action judgment data includes the skeletal point identifiers, action types, and action judgment thresholds in the skeletal point data.
13. A computing device, characterized in that, It includes a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the video production method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer instructions, characterized in that, When executed by the processor, this instruction implements the steps of the video production method according to any one of claims 1 to 11.