Video generation method, apparatus, and electronic device

CN122601944APending Publication Date: 2026-08-18KEEPBEIJING CALORIE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610684410.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本申请的主要目的在于提供一种视频生成方法、装置以及电子设备,以解决相关技术中人工对视频进行处理的效率较低的问题

Benefits of technology

[0016]In this embodiment, the following steps are taken: First, acquire motion videos and motion data of the target user. Then, determine the target motion performed by the target user based on the motion videos. Next, determine the motion performance index values ​​of each initial video frame in the motion video based on the motion type and motion data of the target motion, resulting in M ​​motion performance index values, where M is a positive integer. Then, perform feature recognition operations on the images of each initial video frame to obtain the material feature information of each initial video frame. Based on the material feature information, determine the material performance index values ​​of each initial video frame, resulting in M ​​material performance index values. Finally, determine the value score of each initial video frame based on its motion performance index value and material performance index values. Based on the value scores, filter the initial video frames to obtain the target motion video. The process involves inputting target video frames into a large language model to obtain the target video of the target motion. The target video is obtained by processing the target video frames using the large language model. This is achieved by analyzing the motion video and motion data to determine the index value of each initial video frame in the motion video. Then, target video frames, i.e., keyframes, are selected from the motion video based on these index values. The large language model is then used to process these target video frames to obtain the final motion video. This process accurately selects keyframes from the motion video and generates the motion video based on them, thus achieving the technical effect of automatically generating motion videos. This improves the efficiency of motion video generation and solves the problem of low efficiency in manual video processing in related technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601944A_ABST
    Figure CN122601944A_ABST
Patent Text Reader

Abstract

This application discloses a video generation method, apparatus, and electronic device. Relating to the field of motion information processing, the method includes: acquiring motion video and motion data of a target user, and determining the target user's target motion; determining motion performance index values ​​for each initial video frame in the motion video based on the motion type and motion data of the target motion; performing feature recognition operations on the images of each initial video frame to obtain material feature information for each initial video frame, and determining the material performance index value for each initial video frame based on the material feature information; determining the value score for each initial video frame based on the motion performance index value and the material performance index value, and filtering the initial video frames based on the value score to obtain target video frames; and inputting the target video frames into a large language model to obtain the target video of the target motion. This application solves the problem of low efficiency in manual video processing in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of motion information processing, and more specifically, to a video generation method, apparatus, and electronic device. Background Technology

[0002] With the continuous development of the video industry, many users process their workout footage to create workout videos after completing their exercise, and then share these videos on streaming platforms.

[0003] Current methods for generating sports videos mainly rely on users manually selecting, editing, and splicing materials from sports-related applications. The entire process is complex and time-consuming, typically requiring several hours to complete a single video, which places high demands on users' operational proficiency and time commitment.

[0004] There is currently no effective solution to the problem of low efficiency in manual video processing in related technologies. Summary of the Invention

[0005] The main objective of this application is to provide a video generation method, apparatus, and electronic device to solve the problem of low efficiency in manual video processing in related technologies.

[0006] To achieve the above objectives, according to one aspect of this application, a video generation method is provided. The method includes: acquiring motion video and motion data of a target user, and determining the target motion performed by the target user based on the motion video; determining motion performance index values ​​for each initial video frame in the motion video based on the motion type and motion data of the target motion, obtaining M motion performance index values, where M is a positive integer; performing feature recognition operations on the images of each initial video frame to obtain material feature information for each initial video frame, and determining material performance index values ​​for each initial video frame based on the material feature information, obtaining M material performance index values; determining a value score for each initial video frame based on the motion performance index value and the material performance index value, and filtering the initial video frames based on the value scores to obtain target video frames; and inputting the target video frames into a large language model to obtain a target video of the target motion, wherein the target video is obtained by processing the target video frames using the large language model.

[0007] Optionally, determining the motion performance index values ​​of each initial video frame in the motion video based on the motion type and motion data of the target motion includes: for any initial video frame, obtaining preset weight information under the motion type, wherein the preset weight information includes multiple motion indicators and a preset weight value for each motion indicator; determining the initial motion indicators of the target motion based on the motion type to obtain N initial motion indicators, and calculating the index value of each initial motion indicator of the initial video frame based on the motion data under the initial video frame to obtain N target index values, wherein N is a positive integer; determining the importance score of each initial motion indicator based on the target index value of each initial motion indicator to obtain N importance scores; obtaining the first preset weight value of each initial motion indicator from the preset weight information to obtain N first preset weight values; and performing a weighted summation operation on the N importance scores based on the N first preset weight values ​​to obtain the motion performance index value of the initial video frame.

[0008] Optionally, feature recognition is performed on the images of each initial video frame to obtain the material feature information of each initial video frame, including: for any initial video frame, identifying the target information carried in the image of the initial video frame, wherein the target information includes at least one of the following: scene information, character information, and action information; obtaining the feature information of the target information, and determining the usability of the target information based on the feature information to obtain a usability index value; filtering out usable information from the target information based on the usability index value, and determining the feature information of the usable information as the material feature information.

[0009] Optionally, determining the material performance index value of each initial video frame based on the material feature information includes: for any initial video frame, determining the complexity of each sub-feature information in the material feature information of the initial video frame to obtain multiple complexities; determining the information type of the target information to which each sub-feature information belongs, and determining a second preset weight value for each sub-feature information based on the information type; and performing a weighted summation of the complexity of the sub-feature information based on the second preset weight value to obtain the material performance index value of the initial video frame.

[0010] Optionally, determining the value score of each initial video frame based on the motion performance index value and the material performance index value of each initial video frame includes: for any initial video frame, obtaining a first weight value for the motion performance index value and a second weight value for the material performance index value; performing a weighted summation of the motion performance index value and the material performance index value based on the first weight value and the second weight value to obtain the initial value score; determining the matching degree between the motion data and the image of the initial video frame, and determining the matching coefficient based on the matching degree; multiplying the initial value score by the matching coefficient to obtain the value score of the initial video frame.

[0011] Optionally, filtering initial video frames based on value scores to obtain target video frames includes: configuring the time length of a sliding window, and grouping the video frames of the motion video into P sliding window frame sets, where each sliding window frame set includes a number of initial video frames corresponding to the time length, and P is a positive integer; calculating the average value score of each initial video frame in each sliding window frame set to obtain P average values; determining the sliding window frame set with an average value greater than a preset threshold as the target frame set, and obtaining initial video frames from the motion video whose time difference with the target frame set is less than the time difference threshold to obtain candidate video frames; and determining the initial video frames and candidate video frames in the target frame set as target video frames.

[0012] Optionally, inputting the target video frame into the large language model to obtain the target video of the target motion includes: receiving music instruction information and scene transition information input by the target user, and generating prompt words based on the music instruction information and scene transition information; inputting the prompt words and the target video frame into the large language model to obtain the target video.

[0013] To achieve the above objectives, according to another aspect of this application, a video generation apparatus is provided. The apparatus includes: an acquisition unit, configured to acquire motion video and motion data of a target user, and determine a target motion performed by the target user based on the motion video; a first determination unit, configured to determine motion performance index values ​​for each initial video frame in the motion video based on the motion type and motion data of the target motion, obtaining M motion performance index values, where M is a positive integer; a second determination unit, configured to perform feature recognition operations on the images of each initial video frame, obtaining material feature information for each initial video frame, and determining material performance index values ​​for each initial video frame based on the material feature information, obtaining M material performance index values; a third determination unit, configured to determine a value score for each initial video frame based on the motion performance index value and material performance index value of each initial video frame, and filtering the initial video frames based on the value scores to obtain a target video frame; and a processing unit, configured to input the target video frame into a large language model to obtain a target video of the target motion, wherein the target video is obtained by processing the target video frame using the large language model.

[0014] To achieve the above objectives, according to another aspect of this application, an electronic device is provided, the electronic device including a memory storing an executable program; and a processor for running the program, wherein the program executes the video generation method described above when it runs.

[0015] To achieve the above objectives, according to another aspect of this application, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the video generation method described above.

[0016] In this embodiment, the following steps are taken: First, acquire motion videos and motion data of the target user. Then, determine the target motion performed by the target user based on the motion videos. Next, determine the motion performance index values ​​of each initial video frame in the motion video based on the motion type and motion data of the target motion, resulting in M ​​motion performance index values, where M is a positive integer. Then, perform feature recognition operations on the images of each initial video frame to obtain the material feature information of each initial video frame. Based on the material feature information, determine the material performance index values ​​of each initial video frame, resulting in M ​​material performance index values. Finally, determine the value score of each initial video frame based on its motion performance index value and material performance index values. Based on the value scores, filter the initial video frames to obtain the target motion video. The process involves inputting target video frames into a large language model to obtain the target video of the target motion. The target video is obtained by processing the target video frames using the large language model. This is achieved by analyzing the motion video and motion data to determine the index value of each initial video frame in the motion video. Then, target video frames, i.e., keyframes, are selected from the motion video based on these index values. The large language model is then used to process these target video frames to obtain the final motion video. This process accurately selects keyframes from the motion video and generates the motion video based on them, thus achieving the technical effect of automatically generating motion videos. This improves the efficiency of motion video generation and solves the problem of low efficiency in manual video processing in related technologies. Attached Figure Description

[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0018] Figure 1 A hardware structure block diagram of a computer terminal for implementing a video generation method is shown.

[0019] Figure 2 This is a flowchart of the video generation method provided in Embodiment 1 of this application;

[0020] Figure 3 This is a flowchart of an optional video generation method provided according to Embodiment 1 of this application;

[0021] Figure 4 This is a schematic diagram of a video generation apparatus according to Embodiment 2 of this application;

[0022] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] It should be noted that the video generation method, apparatus and electronic device defined in this disclosure can be used in the field of motion information processing, or in any field other than motion information processing. The application field of the video generation method, apparatus and electronic device defined in this disclosure is not limited.

[0027] It should be noted that all information, user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) used in this application are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with the relevant regulations and standards of the relevant regions, have taken necessary measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse use. If the user chooses to refuse, the process proceeds to the expert decision-making process. For example, this system has interfaces with relevant users or organizations. Before obtaining relevant information, a request to obtain the information needs to be sent to the aforementioned user or organization through the interface. After receiving consent from the aforementioned user or organization, the relevant information is obtained. Users can view the purpose of data use in real time through the authorization interface and have the right to withdraw authorization or delete data at any time. After authorization is withdrawn, the system will terminate the relevant data processing within 24 hours.

[0028] The embodiments or examples disclosed herein are not exhaustive, but merely illustrative of some embodiments or examples, and are not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment or example can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment or example can be arbitrarily interchanged. Furthermore, optional methods or examples in a particular embodiment or example can be arbitrarily combined; moreover, embodiments or examples can be arbitrarily combined. For example, some or all steps of different embodiments or examples can be arbitrarily combined, and a particular embodiment or example can be arbitrarily combined with optional methods or examples of other embodiments or examples.

[0029] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:

[0030] BPM: Beats Per Minute.

[0031] Example 1

[0032] According to an embodiment of this application, an embodiment of a video generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing a video generation method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, processing devices such as microprocessors or programmable logic devices), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface, a universal serial bus port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0034] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned video generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0037] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0038] Under the aforementioned operating environment, this application provides the following: Figure 2 The video generation method shown. Figure 2 This is a flowchart of the video generation method provided in Embodiment 1 of this application, as follows: Figure 2 As shown, the method includes:

[0039] Step S201: Obtain the target user's motion video and motion data, and determine the target motion performed by the target user based on the motion video.

[0040] It should be noted that the execution entity in this embodiment can be a processor or system adapted to a sports application. This system can obtain sports data and sports videos from the sports application when the target sports are authorized, and automatically generate a processed and edited sports video based on the sports data and sports videos, that is, the target video of the target sports.

[0041] It's important to clarify that the target users are those who use fitness apps to record videos and data related to a specific activity. Activity video refers to a continuous sequence of images captured or recorded by the user's device, containing visual information of the target user during exercise and providing the system with the raw input for the visual content. Activity data refers to time-series parameters related to exercise collected by wearable devices, smartphones, or other sensing terminals, which may include data such as heart rate, pace, cadence, altitude, and vertical displacement. The target activity refers to the specific type of exercise actually performed by the target user in the activity video, such as outdoor running.

[0042] For example, when generating a video through the system, it first needs to acquire the motion video recorded by the target user during the exercise, which includes the entire exercise process, as well as the motion data during the exercise. The system first obtains the motion video and motion data synchronously generated by the target user during the exercise through an authorized interface. The motion video is a continuous frame image sequence, and the motion data is a stream of physiological and trajectory parameters aligned with timestamps. The system performs time reference calibration on both, so that the video frames and motion data points have a corresponding time reference.

[0043] Furthermore, based on the environmental features, exercise equipment shapes, user limb movement patterns, and spatial distribution of movement trajectories presented in the exercise video, combined with the tag information in the exercise data, the system automatically determines the specific type of exercise performed by the target user, such as outdoor running, trail running, indoor high-intensity interval training, etc., thereby determining the target exercise for the video recording.

[0044] Step S202: Determine the motion performance index values ​​of each initial video frame in the motion video based on the motion type and motion data of the target motion, and obtain M motion performance index values, where M is a positive integer.

[0045] It should be noted that the motion type refers to the specific motion category of the target motion as determined by the system, while the motion performance index value refers to the quantitative value that represents the user's motion intensity and physiological information at the corresponding moment of each frame, calculated based on motion data and the target motion type.

[0046] For example, based on the motion type of the target motion and the time series of each parameter in the motion data, the corresponding motion performance index value can be calculated for each initial video frame, resulting in a total of M values, where M is the total number of video frames.

[0047] The sports performance index is driven by a multi-sports category weight matrix corresponding to the target sports type. This matrix defines the relative importance of sports data feature fields such as heart rate, pace, cadence, altitude, and vertical displacement for each sports type. First, the field values ​​of each sports data feature field in each frame of the sports data need to be normalized to the range [0, 1] to obtain the sports data vector of each frame. Then, the sports data vectors at the corresponding time points of each frame are weighted and summed according to the weight matrix to generate the sports performance index value of that frame. This value reflects the user's level of physical exertion at that time of the frame.

[0048] It should be noted that this calculation process is executed in parallel frame by frame, ensuring that each frame has independent criteria for evaluating motion performance. The weight matrix is ​​type-dependent; for example, in trail running, gradient and vertical displacement have higher weights than heart rate, while in high-intensity interval training, movement frequency and peak heart rate have dominant weights.

[0049] It's worth noting that exercise data can be obtained by reading relevant fields from the completed workout page. For example, a runner's workout page will show their running distance, average pace, heart rate, and calories burned. Users who upload data via smart bracelets or smartwatches can also have more in-depth data such as their oxygen uptake. If the user is running outdoors, their route can also be recorded, marked with information such as altitude, elevation gain, and geometric features of the trajectory (e.g., shuttle runs, loop runs, circular runs, straight-line climbs, etc.).

[0050] It should be noted that the sports data feature fields included in the sports data may include: Peak_HR_Ratio: current heart rate / user's historical maximum heart rate; Speed_Delta: pace improvement rate; TCEP (Training Cumulative Effort Point): training cumulative load point, identifying the user's ability to maintain a high heart rate zone for an extended period; Slope_Stability: the ability to maintain a stable pace when the incline changes; Vertical_Oscollation: vertical oscillation (identifying the feeling of airborne during running).

[0051] For example, the weight matrix can be shown in Table 1:

[0052] Table 1

[0053]

[0054] The weight allocation section only illustrates the characteristics of weight allocation; the specific weight values ​​are not shown here and can be adjusted and set according to actual needs.

[0055] Step S203: Perform feature recognition operation on the images of each initial video frame to obtain the material feature information of each initial video frame, and determine the material performance index value of each initial video frame based on the material feature information to obtain M material performance index values.

[0056] It should be noted that the material feature information refers to the structured visual features extracted by the system from the image, which may include descriptive parameters of dimensions such as scene, characters, and actions. The material performance index value refers to the quantitative value that represents the visual performance of each frame, calculated based on the material feature information.

[0057] For example, it is also necessary to perform multi-dimensional feature recognition on the image content of each initial video frame to extract material feature information including scene semantics, character posture, and action morphology. Scene semantics uses image segmentation technology to identify the background environment type, such as forest roads, city streets, or finish line arches, and assesses its visual complexity and spatial openness. Character posture is determined by key point detection to determine the torso angle, height in the air, and limb integrity, and to determine whether the composition of the image focuses on the moving subject. Action morphology is identified by spatiotemporal motion detection to identify typical motion patterns such as sprinting arm swings and cornering, and quantifies their motion amplitude and coordination.

[0058] After obtaining the material feature information of each initial video frame, the above feature information can be weighted and fused to form an independent material performance index value for each frame.

[0059] Step S204: Determine the value score of each initial video frame based on the motion performance index value and material performance index value of each initial video frame, and filter the initial video frames according to the value score to obtain the target video frames.

[0060] It should be noted that the value score refers to the comprehensive score of a single frame, which is obtained by combining the motion performance index value and the material performance index value and correcting for the consistency of the changes of the two.

[0061] For example, after obtaining the motion performance index value and the source material performance index value, the value score of each initial video frame can be determined based on these values. The value score can be determined by the product of the motion performance index value and the source material performance index value, and a correlation coefficient is introduced to correct for the consistency of their changes. When motion performance increases and visual performance increases synchronously, the correlation coefficient is positive, and the value score increases; when the two change in opposite directions or are unrelated, the correlation coefficient decreases, and the value score is suppressed.

[0062] The system sorts the value scores of all M initial video frames, selects frames with value scores higher than a preset threshold as candidates, and then clusters consecutive high-value frames through a sliding window to improve the temporal coherence of the generated segments. Finally, it selects the set of target video frames that constitute the target video skeleton.

[0063] Step S205: Input the target video frame into the large language model to obtain the target video of the target motion, wherein the target video is obtained by the large language model processing the target video frame.

[0064] It should be noted that a large language model can be an intelligent model with natural language understanding and generation capabilities, capable of processing sequential visual input and outputting structured video content.

[0065] For example, after selecting the target video frames, the final edited target video can be generated based on the target video frames. At this point, the selected target video frame sequence can be used as the sole input to the large language model. Based on the input target video frame sequence, the model automatically generates a complete target video containing the image sequence, synchronized audio, dynamic visual effects, and matching text descriptions, according to the target motion type, the time order between frames, and the motion and material performance information carried by each frame.

[0066] For example, the model divides narrative segments based on changes in exercise duration and intensity, such as initial calm, peak sprint, and fatigue persistence; it maps music rhythm based on pace fluctuations and cadence features; it generates matching narration and subtitle text based on scene semantics, and finally processes the target video frame sequence into the target video.

[0067] Therefore, steps S201-S205 achieve fully automated processing from raw footage to emotionally resonant final product by acquiring motion videos and data, quantifying frame-level motion performance and material performance, fusing and calculating value scores to select keyframes, and then generating the final video using a large language model. This method abandons the traditional model that relies on manual editing and static templates, achieving dynamic alignment between the physiological state of motion and the visual expression of the video, as well as automatic recognition of keyframes and video generation. It significantly improves video generation efficiency, realism, and personalization, is applicable to various types of motion, and has broad application value and technological advancement.

[0068] The video generation method provided in this application involves: acquiring motion videos and motion data of a target user; determining the target motion performed by the target user based on the motion video; determining motion performance index values ​​for each initial video frame in the motion video based on the motion type and motion data of the target motion, resulting in M ​​motion performance index values, where M is a positive integer; performing feature recognition operations on the images of each initial video frame to obtain material feature information for each initial video frame, and determining material performance index values ​​for each initial video frame based on the material feature information, resulting in M ​​material performance index values; determining a value score for each initial video frame based on the motion performance index value and the material performance index value, and filtering the initial video frames based on the value score. The process involves obtaining target video frames and inputting them into a large language model to generate the target video of the target motion. The target video is obtained by processing the target video frames using the large language model. This is achieved by analyzing the motion video and motion data to determine the index value of each initial video frame in the motion video. Based on these index values, target video frames, or keyframes, are selected from the motion video. The large language model is then used to process these target video frames to obtain the final motion video. This process accurately selects keyframes from the motion video and generates the motion video based on them, thus achieving the technical effect of automatically generating motion videos. This improves the efficiency of motion video generation and solves the problem of low efficiency in manual video processing in related technologies.

[0069] Optionally, in the video generation method provided in this application embodiment, determining the motion performance index value of each initial video frame in the motion video based on the motion type and motion data of the target motion includes: for any initial video frame, obtaining preset weight information under the motion type, wherein the preset weight information includes multiple motion indicators and a preset weight value for each motion indicator; determining the initial motion indicators of the target motion based on the motion type to obtain N initial motion indicators, and calculating the index value of each initial motion indicator of the initial video frame based on the motion data under the initial video frame to obtain N target index values, wherein N is a positive integer; determining the importance score of each initial motion indicator based on the target index value of each initial motion indicator to obtain N importance scores; obtaining the first preset weight value of each initial motion indicator from the preset weight information to obtain N first preset weight values; and performing a weighted summation operation on the N importance scores based on the N first preset weight values ​​to obtain the motion performance index value of the initial video frame.

[0070] It should be noted that the preset weight information refers to a preset structured parameter set, containing multiple motion indicators defined for a specific motion type and their corresponding preset weight values, used to guide the calculation of motion performance indicator values. The preset weight value refers to the numerical value pre-set for each motion indicator in the preset weight information, used to characterize the relative importance of that indicator in a specific motion type. Motion indicators refer to quantitative parameters directly related to motion performance, such as heart rate, pace, cadence, altitude, and vertical displacement. Initial motion indicators are the motion indicators needed to calculate motion performance indicator values ​​under the target motion type. Target indicator values ​​are the specific numerical values ​​of each initial motion indicator obtained through numerical calculation based on motion data at the corresponding moment of the motion video frame. The importance score is a score reflecting the relative intensity of the indicator in the motion state of that frame, obtained after normalizing and standardizing the target indicator value. The first preset weight value: refers to the preset weight value extracted from the preset weight information, corresponding one-to-one with each initial motion indicator.

[0071] For example, when determining the value of an athletic performance indicator, it is first necessary to obtain the preset weight information corresponding to that type. This information is a structured definition that includes multiple athletic indicators and their respective preset weight values, and each set of preset weight information is bound to a specific athletic type.

[0072] After obtaining the preset weight information, it is also necessary to determine the initial motion index of the target motion, that is, the initial motion index required for analyzing the motion performance index value of the target motion. After obtaining the initial motion index, the motion data corresponding to the current initial video frame is read, and for each initial motion index, its specific value at the time point of the frame is calculated based on the motion data in the frame, thereby obtaining the target index value of each initial motion index.

[0073] To eliminate differences in dimensions and biases in data distribution, each target indicator value can be normalized and mapped to the interval between zero and one to form a corresponding importance score, which reflects the relative intensity level of the indicator at the current moment.

[0074] Finally, the first preset weight value corresponding to each of the N initial motion indicators can be extracted from the preset weight information to obtain N first preset weight values. Then, the N initial motion indicators and the N first preset weight values ​​are used to perform a weighted summation operation to obtain the motion performance index value of the initial video frame.

[0075] For example, when the target sport is trail running, preset weight information is retrieved, which includes the sports indicators vertical displacement, slope, heart rate, and pace. The corresponding preset weight values ​​can be set to 0.5, 0.3, 0.15, and 0.05, respectively. The system reads the motion data at a specific moment in an initial video frame, measuring a vertical displacement of 2.8 meters per second, a slope of 0.12, a heart rate of 158 beats per minute, and a pace of 6 minutes and 10 seconds per kilometer. Using a preset data normalization method, the vertical displacement is normalized to 0.92, the slope to 0.88, the heart rate to 0.76, and the pace to 0.41, forming four target indicator values. Based on a preset non-linear mapping function, the target indicator values ​​are converted into corresponding importance scores: vertical displacement 0.90, slope 0.85, heart rate 0.72, and pace 0.38. The system takes the corresponding first preset weight values: 0.5, 0.3, 0.15, and 0.05 respectively, and performs a weighted summation operation: 0.90×0.50+0.85×0.30+0.72×0.15+0.38×0.05=0.783. Finally, the motion performance index value of the initial video frame is output as 0.783.

[0076] It should be noted that the above weight values ​​are only an optional setting example for illustration and are not actual weight values ​​used. This scheme does not restrict the setting of specific weight values, and you can set them as needed.

[0077] This embodiment achieves accuracy in calculating athletic performance indicators through a structured, type-bound preset weighting mechanism, ensuring inherent consistency and rationality in the evaluation logic across different sports types. By introducing importance scores to standardize the raw data and combining this with a first preset weight value for weighted summation, the problem of high-dimensional parameters dominating the calculation is effectively avoided, thus improving the accuracy of the final calculated athletic performance indicators.

[0078] Optionally, in the video generation method provided in this application embodiment, performing feature recognition operation on the images of each initial video frame to obtain the material feature information of each initial video frame includes: for any initial video frame, identifying target information carried in the image of the initial video frame, wherein the target information includes at least one of the following: scene information, character information, and action information; obtaining feature information of the target information, and determining the availability of the target information based on the feature information to obtain a availability index value; filtering out usable information from the target information based on the availability index value, and determining the feature information of the usable information as material feature information.

[0079] It should be noted that scene information refers to the environmental semantic content presented by the image background, such as forest roads, city streets, stadiums, and finish line arches, reflecting the spatial environment in which the movement occurs. Person information refers to attributes such as the target user's body shape, position, posture, and proportion within the frame, reflecting the subject's presentation in the image. Action information refers to the physical movement patterns performed by the target user in the image, such as sprinting with arm swings, cornering, and pedaling, reflecting the dynamic characteristics of the movement. Usability refers to the effectiveness and representativeness of the target information in the video's expression, reflecting whether it truly carries the narrative intent of the movement, rather than being distracting or invalid content. The usability index value is a numerical value calculated based on feature information, quantifying the degree of usability of the target information, used to determine whether the information qualifies for inclusion in the material's feature information.

[0080] For example, when acquiring material feature information, for any initial video frame, when acquiring the material feature information of that initial video frame, it is necessary to perform multi-dimensional target information recognition operation on the image of each initial video frame. Through image semantic segmentation, target detection and spatiotemporal action analysis, scene information, character information and action information are extracted simultaneously to obtain the target information carried in the image.

[0081] For each type of target information, it is also necessary to further extract its corresponding feature information, such as the background complexity, lighting quality and spatial openness of scene information, the torso tilt angle, jumping height and clothing contrast of character information, and the joint displacement amplitude, movement frequency and body coordination of action information, so as to obtain the feature information of the target information.

[0082] Subsequently, a usability assessment model needs to be constructed based on the feature information of each type of target information. This model comprehensively scores the feature information based on preset rules to determine whether it truly reflects the motion state, rather than being staged, obscured, or containing irrelevant interference, and outputs the corresponding usability index value. For example, if the torso proportion in the person information is below a threshold and the background is cluttered, its usability index value is significantly reduced; if the joint movement amplitude in the action information is below the baseline of the movement type, its usability index value is suppressed. The system filters each type of target information according to a preset usability threshold, retaining only the target information with a usability index value higher than the threshold, forming a set of usable information.

[0083] Finally, the system integrates the feature information of all available information to form the material feature information of the initial video frame. This information is a structured vector containing quantitative parameters of three dimensions: scene, character, and action, which are used to calculate the performance index values ​​of the subsequent material.

[0084] For example, by performing image analysis on an initial video frame, the system identifies the following target information: scene information: forest path; person information: user's torso leans forward at a 35-degree angle, face facing forward, clothing is dark red; action information: high-frequency arm swings, leg jump height of 28 centimeters. Based on this information, the following features can be extracted for the scene information: background texture complexity 0.75, lighting contrast 0.68; person information features: torso centering 0.82, clothing-background contrast 0.91; action information features: swing frequency 2.1 times per second, normalized jump height 0.86. The system determines the usability based on the following rules: background texture complexity above 0.7, clothing contrast above 0.8, and aerial height above 0.8. Therefore, the usability index values ​​for scene information are 0.72, character information is 0.87, and action information is 0.90, all of which are higher than the preset threshold of 0.6. The system retains all three items as usable information and finally combines the above feature information into the material feature information of this frame, including scene complexity of 0.75, clothing contrast of 0.91, torso tilt angle of 35 degrees, swing frequency of 2.1 times per second, and aerial height of 0.86.

[0085] Similarly, the material feature information of each initial video frame can be obtained through the above process, thereby obtaining the material feature information of each initial video frame.

[0086] It should be noted that the values ​​disclosed in the above cases are optional settings used for illustrative purposes and do not represent actual collected data or actual settings.

[0087] This embodiment effectively eliminates invalid, interfering, or non-realistic motion expressions in images by using a hierarchical identification, usability assessment, and filtering mechanism for target information. This ensures that the material feature information consists only of high-quality visual elements, achieving the technical effect of accurately acquiring material feature information.

[0088] Optionally, in the video generation method provided in this application embodiment, determining the material performance index value of each initial video frame based on the material feature information includes: for any initial video frame, determining the complexity of each sub-feature information in the material feature information of the initial video frame to obtain multiple complexities; determining the information type of the target information to which each sub-feature information belongs, and determining a second preset weight value for each sub-feature information based on the information type; and performing a weighted summation of the complexity of the sub-feature information based on the second preset weight value to obtain the material performance index value of the initial video frame.

[0089] It should be noted that sub-feature information refers to the smallest quantifiable visual parameter unit that constitutes the feature information of the material, such as the background complexity of the scene, the height of the character in the air, and the amplitude and frequency of the movement. It is the basic input unit for calculating the performance index value of the material. Complexity refers to the quantitative score of the information density, structural diversity, or expressive intensity of each sub-feature information in the image, reflecting its contribution to the expressiveness of the picture.

[0090] For example, when determining the performance index value of the material, for any initial video frame, it is necessary to determine the complexity of each sub-feature information under the initial video frame, and perform a weighted summation operation on the sub-feature information according to the complexity to obtain the complexity of the material feature information of the initial video frame, that is, the performance index value of the material.

[0091] First, it is necessary to structurally decompose the material feature information of each initial video frame and extract all the sub-feature information contained therein. Each sub-feature information corresponds to a specific visual description dimension, such as the background texture complexity, lighting uniformity, and spatial depth in scene information; the proportion of the torso in the frame, limb integrity, and color contrast between clothing and background in character information; and the range of joint movement, duration of movement, and rate of change of body coordination in action information.

[0092] For each sub-feature, its complexity is calculated using a pre-defined evaluation model. This complexity is a non-linear score, determined by a combination of the feature's distribution breadth, dynamic change intensity, and semantic saliency in the image. For example, the complexity of the action sub-feature information of a human posture with a high altitude and clear limb extension is higher than that of a low-amplitude swing.

[0093] Similarly, it is also necessary to identify the information type to which each sub-feature information belongs, such as scene, character, action, and determine the second preset weight value of each sub-feature information based on the information type.

[0094] Subsequently, the system performs a weighted summation operation on each sub-feature information, that is, multiplies the complexity by the second preset weight value to obtain the weighted contribution value, and finally sums the weighted contribution values ​​of all sub-feature information to obtain the material performance index value of the initial video frame.

[0095] For example, the system decomposes the feature information of an initial video frame to obtain sub-feature information: scene complexity 0.75 (belonging to scene information), clothing contrast 0.91 (belonging to character information), torso tilt angle 35 degrees (belonging to character information), swing frequency 2.1 times per second (belonging to action information), and height of flight 0.86 (belonging to action information). The system assigns a second preset weight value to each type of information: scene information is 0.15, character information is 0.2, and action information is 0.4. The system calculates the complexity of each sub-feature information: scene complexity 0.75 has a complexity score of 0.75, clothing contrast 0.91 has a complexity score of 0.91, a torso tilt angle of 35 degrees is a typical force exertion posture in running, so the complexity is 0.88, the swing frequency of 2.1 times per second is higher than the baseline value, so the complexity is 0.90, and the height of flight 0.86 has a complexity of 0.86. The system performs a weighted summation: 0.75×0.15+0.91×0.10+0.88×0.10+0.90×0.40+0.86×0.40=0.837, and finally outputs the material performance index value of this frame as 0.837.

[0096] It should be noted that the numerical values ​​and weight values ​​disclosed in the above examples are optional settings for illustrative purposes only and do not represent actual collected data or actual settings. This solution does not restrict the setting of specific weight values ​​and can be set as needed.

[0097] This embodiment achieves the technical effect of accurately determining the material performance index value of each initial video frame by introducing a second preset weight value that combines the complexity evaluation of sub-feature information with the binding of information type.

[0098] Optionally, in the video generation method provided in this application embodiment, determining the value score of each initial video frame based on the motion performance index value and the material performance index value of each initial video frame includes: for any initial video frame, obtaining a first weight value of the motion performance index value and a second weight value of the material performance index value; performing a weighted summation calculation on the motion performance index value and the material performance index value based on the first weight value and the second weight value to obtain an initial value score; determining the matching degree between the motion data and the image of the initial video frame, and determining a matching coefficient based on the matching degree; and multiplying the initial value score by the matching coefficient to obtain the value score of the initial video frame.

[0099] It should be noted that the first weight value refers to a preset fixed coefficient used to weight the motion performance index values. The second weight value refers to a preset fixed coefficient used to weight the material performance index values. The initial value score refers to the preliminary comprehensive score obtained by weighted summation, without considering the consistency correction between motion data and visual content, and serves as the raw output of the value score. The matching degree refers to the degree of synchronization between the trend of motion data change corresponding to the initial video frame and the trend of visual performance change in its image, under the premise of time alignment, and is used to determine whether physiological intensity and visual expression are truly coordinated.

[0100] For example, when calculating the value score of the initial video frame, the corresponding motion performance index value and material performance index value are first obtained, and the first weight value and the second weight value used for weighting are determined.

[0101] Furthermore, a weighted summation operation is performed based on the motion performance index value and the material performance index value of the initial video frame. The motion performance index value is multiplied by the first weight value, and the material performance index value is multiplied by the second weight value. The two products are then added together to obtain the initial value score of the frame.

[0102] Furthermore, since the initial value score only reflects the linear superposition of the physiological and visual dimensions and does not consider whether the two are synchronized in terms of temporal evolution, it is also necessary to analyze the consistency between the motion data and image content corresponding to the frame in the direction of change, i.e., the matching degree: if the motion data shows a significant increase (such as a sudden increase in heart rate or pace), and at the same time the amplitude of the person's movements in the image increases, the height of the jump increases, and the facial expression shows exhaustion or determination, then the matching degree is high; if the motion data is stable or decreases, but the visual content shows exaggerated posing or irrelevant movements, then the matching degree is low.

[0103] After obtaining the matching degree, the matching coefficient ρ can be determined based on the matching degree, and the final value score of the initial video frame can be determined based on the matching coefficient ρ and the initial value score, which can be calculated using the following formula:

[0104]

[0105] in, Value frame The initial video frame's value score, w1 is the first weight value, and S... P For sports performance indicators, w2 is the second weighted value, and S v ρ represents the performance index value of the material, and ρ is the matching coefficient.

[0106] For example, the system obtains a motion performance index value of 0.783 for an initial video frame, a material performance index value of 0.837, a preset first weight value of 0.6, and a second weight value of 0.4. The system performs a weighted sum: 0.783 × 0.60 + 0.837 × 0.40 = 0.804, resulting in an initial value score of 0.804. The system analyzes the motion data of this frame: the heart rate is increasing from 150 beats per minute to 162 beats per minute, and the pace is increasing from 6 minutes and 10 seconds per kilometer to 5 minutes and 50 seconds per kilometer; simultaneously, the angle of the person's torso leaning forward in the image increases from 30 degrees to 35 degrees, the swaying frequency increases from 1.9 times per second to 2.1 times per second, facial muscles are tense, and the reflection of sweat is enhanced. The system determines that the motion data and visual performance trends are consistent, with a matching degree of 0.94 and a matching coefficient of 0.94. The system multiplies the initial value score by the matching coefficient: 0.804 × (1 + 0.93) = 1.55, and finally outputs a value score of 1.55 for this frame.

[0107] It should be noted that the numerical values ​​and weight values ​​disclosed in the above examples are optional settings for illustrative purposes only and do not represent actual collected data or actual settings. This solution does not restrict the setting of specific weight values ​​and can be set as needed.

[0108] This embodiment introduces a matching coefficient to dynamically correct the initial value score, thereby verifying the causal consistency between physiological intent and visual expression. This improves the accuracy of the value score and effectively filters out false target video frames caused solely by visual impact or data anomalies. In turn, it achieves the technical effect of improving the accuracy of determining keyframes based on value scores.

[0109] Optionally, in the video generation method provided in this application embodiment, filtering initial video frames according to value scores to obtain target video frames includes: configuring the time length of a sliding window, and grouping the video frames of the motion video into P sliding window frame sets according to the sliding window, wherein each sliding window frame set includes a number of initial video frames corresponding to the time length, and P is a positive integer; calculating the average value score of each initial video frame in each sliding window frame set to obtain P average values; determining the sliding window frame set with an average value greater than a preset threshold as the target frame set, and obtaining initial video frames from the motion video whose time difference with the target frame set is less than the time difference threshold to obtain candidate video frames; and determining the initial video frames and candidate video frames in the target frame set as target video frames.

[0110] It's important to note that a sliding window refers to a dynamic time interval of fixed length that moves frame by frame along the video timeline to divide consecutive video frames into multiple overlapping subsequences. The time length refers to the span of time corresponding to the consecutive video frames covered by the sliding window. Sliding grouping refers to sequentially dividing the initial video frames in a moving video, using the sliding window as a unit, to form multiple sets of frames with temporal continuity. A sliding window frame set is an ordered set consisting of all the initial video frames covered by the sliding window at a given moment, with each set containing a number of frames corresponding to the time length.

[0111] For example, when determining the target video frame, since the target video frame is a combination of multiple frame sequences, the motion video can be grouped by frame according to the sliding window, and the sliding window frame set can be filtered according to the value score of each sliding window frame set, thereby obtaining the video frames that constitute the target video frame.

[0112] First, you need to configure the duration of the sliding window, for example, 3 to 5 seconds, and slide the window frame by frame along the video timeline with a fixed step size. All initial video frames are grouped by sliding to generate P sets of sliding window frames, each set containing a time span equal to the preset duration.

[0113] Furthermore, the value scores of all initial video frames within each sliding window frame set are arithmetically averaged to obtain P average values, which reflect the sustained level of highlight intensity within that local time period. The system compares all average values ​​with a preset threshold and retains only the sliding window frame set whose average value is higher than the threshold as the target frame set. This set represents the period in the video where there is significant and sustained high-value performance.

[0114] Furthermore, in order to ensure that the video segment composed of the target video frames has a natural start and end in time, the system further extracts all adjacent initial video frames from the motion video whose start and end times differ from the target frame set by less than a time difference threshold, as candidate video frames. The time difference threshold can be set to 1 to 2 seconds.

[0115] Finally, all initial video frames and candidate video frames in the target frame set are merged to form a complete target video frame set. This set is coherent in time and concentrated in value, preserving the core highlight segments while fully carrying the preparation and fallback processes before and after them, thus avoiding narrative breaks caused by segment cutting.

[0116] For example, if the sliding window duration is configured to four seconds and the video frame rate is thirty frames per second, then each sliding window frame set contains one hundred and twenty frames. The system slides the window one step per frame, generating a total of one hundred and fifty sliding window frame sets. The system calculates the average value score of the one hundred and twenty frames in each set. The average values ​​of sets 87 to 92 are 0.72, 0.78, 0.83, 0.85, 0.81, and 0.76 respectively, all higher than the preset threshold of 0.70. Therefore, sets 87 to 92 are determined as the target frame sets. The system sets a time difference threshold of two seconds, or sixty frames. The system extracts frames 81 to 86 and 93 to 98 from the moving video. Because their time difference with the target frame set is less than sixty frames, they are identified as candidate video frames. The system merges all frames 81 to 98 to form the target video frame set.

[0117] It should be noted that the numerical values ​​and weight values ​​disclosed in the above examples are optional settings for illustrative purposes only and do not represent actual collected data or actual settings. This solution does not restrict the setting of specific weight values ​​and can be set as needed.

[0118] This embodiment achieves dynamic identification of keyframes through a sliding window mechanism, effectively avoiding false screening caused by fluctuations in a single frame. Furthermore, by introducing candidate video frames, the smoothness of the video is improved, thereby achieving the technical effect of improving the effectiveness and accuracy of the target video frames.

[0119] Optionally, in the video generation method provided in this application embodiment, inputting the target video frame into a large language model to obtain the target video of the target motion includes: receiving music instruction information and scene transition information input by the target user, and generating prompt words based on the music instruction information and scene transition information; inputting the prompt words and the target video frame into the large language model to obtain the target video.

[0120] It should be noted that music instruction information refers to subjective descriptive instructions provided by the target user regarding the style, rhythm, or emotional atmosphere of the background music in the generated video. Scene transition information refers to semantic descriptions provided by the target user regarding the transition methods between different segments in the video.

[0121] For example, after obtaining the target video frame, since music is usually added when generating the video, upon receiving the music instruction information and scene transition information input by the target user, the system first uses a built-in nonlinear semantic mapping algorithm to extract objective features of the motion data, such as total duration, calorie consumption, motion type, and pace fluctuations. These features are then combined with the user-input music instruction information and scene transition information to construct structured prompts. These prompts can contain hierarchical semantic instructions. For example, the starting segment corresponds to "Style 1," generating visual and auditory guidance for environmental description; the peak segment corresponds to "Style 2," generating control instructions emphasizing impact, tension, and accelerated rhythm; and the fatigue / perseverance segment corresponds to "Style 3," generating expressive constraints highlighting resolute expressions, low-saturation tones, and slowly grading timbre.

[0122] It's important to note that the prompts need to embed a semantic mapping between motion parameters and user intent, such as "180 steps per minute corresponds to a music tempo of 120 BPM" or "Heart rate consistently above the threshold range drives the timbre overlay intensity to 80%." The system uses these prompts and the target video frame as joint input to a large language model. The model receives the target video frame as a visual sequence input and encodes the motion and visual features of each frame into a latent space representation through a visual-semantic alignment module. Simultaneously, it parses the semantic instructions in the prompts to construct narrative structure constraints. Based on the paragraph divisions, rhythm mappings, and emotional tones defined in the prompts, the model autonomously generates matching audio waveforms, image filter changes, special effects overlay sequences, and subtitle text, outputting the complete target video.

[0123] For example, the target user inputs music instructions as "the tempo gradually increases, ending with a drum burst," and scene transition information as "smooth transition." The system combines motion data from the target video frames: heart rate consistently exceeds the threshold in the last three minutes, pace increases significantly, a noticeable acceleration occurs in the last fifty frames, facial expressions change from tense to shouting, and the background gradually changes from a forest path to the final archway. The system generates prompts: "The initial section features low-pitched strings accompanied by ambient wind sounds, percussion is gradually added in the middle section, the drum density increases by 10% for every second the pace increases, a brass climax is added within the last three seconds, the scene slowly zooms in from the mountain path to a panoramic view of the archway, and the filter changes from cool gray to warm gold. Based on the above requirements, the system generates a complete video by combining the input video frames." The system inputs this prompt and the target video frames into a large language model. The model outputs a video containing the complete sequence of images, synchronized audio tracks, dynamic filter changes, and subtitles, forming the target video of the target motion.

[0124] This embodiment combines the user's music information request with video frames by using prompt words, thereby improving the accuracy of video generation.

[0125] Figure 3This is a flowchart of an optional video generation method provided according to Embodiment 1 of this application, such as... Figure 3 As shown, the process first requires determining the motion performance index and material performance index based on the motion video and motion data, and then associating the number of video frames with the motion data. Next, the value score of each single frame is calculated based on the motion performance index and material performance index, and the target video frame is selected from each single frame based on the value score. Finally, the target video frame is input into the large language model, which processes the target video frame to obtain the final output video.

[0126] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0127] Example 2

[0128] This application also provides a video generation apparatus. It should be noted that the video generation apparatus of this application can be used to execute the video generation method provided in the above embodiments. The video generation apparatus provided in this application will be described below.

[0129] According to an embodiment of this application, an apparatus for implementing the above-described video generation method is also provided. Figure 4 This is a schematic diagram of a video generation apparatus according to Embodiment 2 of this application, as shown below. Figure 4 As shown, the device includes:

[0130] To achieve the above objectives, according to another aspect of this application, a video generation apparatus is provided. The apparatus includes:

[0131] The acquisition unit 41 is used to acquire the motion video and motion data of the target user, and determine the target motion performed by the target user based on the motion video.

[0132] The first determining unit 42 is used to determine the motion performance index values ​​of each initial video frame in the motion video according to the motion type and motion data of the target motion, and obtain M motion performance index values, where M is a positive integer.

[0133] The second determining unit 43 is used to perform feature recognition operations on the images of each initial video frame to obtain the material feature information of each initial video frame, and to determine the material performance index value of each initial video frame based on the material feature information, thereby obtaining M material performance index values.

[0134] The third determining unit 44 is used to determine the value score of each initial video frame based on the motion performance index value and the material performance index value of each initial video frame, and to filter the initial video frames based on the value score to obtain the target video frame.

[0135] The processing unit 45 is used to input the target video frame into the large language model to obtain the target video of the target motion, wherein the target video is obtained by the large language model processing the target video frame.

[0136] The video generation apparatus provided in this application embodiment acquires motion video and motion data of a target user through an acquisition unit 41, and determines the target motion performed by the target user based on the motion video; a first determination unit 42 determines the motion performance index value of each initial video frame in the motion video based on the motion type and motion data of the target motion, obtaining M motion performance index values, where M is a positive integer; a second determination unit 43 performs feature recognition operation on the images of each initial video frame to obtain the material feature information of each initial video frame, and determines the material performance index value of each initial video frame based on the material feature information, obtaining M material performance index values; a third determination unit 44 determines the value score of each initial video frame based on the motion performance index value and the material performance index value of each initial video frame, and filters the initial video frames based on the value score to obtain the target video frame; a processing unit 45 inputs the target video frame into a large language model to obtain the target video of the target motion, wherein the target video is obtained by processing the target video frame by the large language model. By analyzing motion videos and motion data, the index value of each initial video frame in the motion video is determined. Then, target video frames, or keyframes, are selected from the motion video based on the index values. The target video frames are then processed using a large language model to obtain the final motion video. This achieves the goal of accurately selecting keyframes in the motion video and generating motion video based on the keyframes, thus realizing the technical effect of automatically generating motion videos. This improves the efficiency of motion video generation and solves the problem of low efficiency in manual video processing in related technologies.

[0137] Optionally, in the video generation apparatus provided in this application embodiment, the first determining unit 42 includes: a first acquiring module, configured to acquire preset weight information under a motion type for any initial video frame, wherein the preset weight information includes multiple motion indicators and a preset weight value for each motion indicator; a first determining module, configured to determine the initial motion indicators of the target motion according to the motion type, obtain N initial motion indicators, and calculate the indicator value of each initial motion indicator of the initial video frame according to the motion data under the initial video frame, obtain N target indicator values, wherein N is a positive integer; a second determining module, configured to determine the importance score of each initial motion indicator according to the target indicator value of each initial motion indicator, obtain N importance scores; a second acquiring module, configured to acquire the first preset weight value of each initial motion indicator from the preset weight information, obtain N first preset weight values; and a first calculating module, configured to perform a weighted summation operation on the N importance scores according to the N first preset weight values, obtain the motion performance indicator value of the initial video frame.

[0138] Optionally, in the video generation apparatus provided in this application embodiment, the second determining unit 43 includes: an identification module, used to identify target information carried in the image of any initial video frame, wherein the target information includes at least one of the following: scene information, character information, and action information; a third acquisition module, used to acquire feature information of the target information and determine the availability of the target information based on the feature information to obtain an availability index value; and a filtering module, used to filter out usable information from the target information based on the availability index value and determine the feature information of the usable information as material feature information.

[0139] Optionally, in the video generation apparatus provided in this application embodiment, the second determining unit 43 includes: a third determining module, used to determine the complexity of each sub-feature information in the material feature information of any initial video frame, and obtain multiple complexities; a fourth determining module, used to determine the information type of the target information to which each sub-feature information belongs, and determine a second preset weight value for each sub-feature information according to the information type; and a second calculation module, used to perform weighted summation of the complexity of the sub-feature information according to the second preset weight value, and obtain the material performance index value of the initial video frame.

[0140] Optionally, in the video generation apparatus provided in this application embodiment, the third determining unit 44 includes: a fourth obtaining module, used to obtain a first weight value of the motion performance index value and a second weight value of the material performance index value for any initial video frame; a third calculating module, used to perform a weighted summation calculation on the motion performance index value and the material performance index value according to the first weight value and the second weight value to obtain an initial value score; a fifth determining module, used to determine the matching degree between the motion data and the image of the initial video frame, and determine the matching coefficient according to the matching degree; and a fourth calculating module, used to multiply the initial value score by the matching coefficient to obtain the value score of the initial video frame.

[0141] Optionally, in the video generation apparatus provided in this application embodiment, the third determining unit 44 includes: a grouping module, used to configure the time length of the sliding window, and to group the video frames of the motion video according to the sliding window to obtain P sliding window frame sets, wherein each sliding window frame set includes a number of initial video frames corresponding to the time length, and P is a positive integer; a fifth calculation module, used to calculate the average value of each initial video frame in each sliding window frame set to obtain P average values; a sixth determining module, used to determine the sliding window frame set with an average value greater than a preset threshold as the target frame set, and to obtain initial video frames from the motion video whose time difference with the target frame set is less than the time difference threshold to obtain candidate video frames; and a seventh determining module, used to determine the initial video frames and candidate video frames in the target frame set as target video frames.

[0142] Optionally, in the video generation apparatus provided in this application embodiment, the processing unit 45 includes: a generation module, used to receive music instruction information and scene transition information input by the target user, and generate prompt words based on the music instruction information and scene transition information; and a processing module, used to input the prompt words and target video frames into a large language model to obtain the target video.

[0143] It should be noted that the aforementioned acquisition unit 41, first determination unit 42, second determination unit 43, third determination unit 44, and processing unit 45 correspond to steps S201 to S205 in Embodiment 1. The instances and application scenarios implemented by each of these units and their corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned modules or units can be hardware or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). These modules can also run as part of a device in the computer terminal 10 provided in Embodiment 1.

[0144] Example 3

[0145] Embodiments of this application may provide an electronic device. Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 (Only one is shown) processor 1002, memory 1004, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0146] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0147] Those skilled in the art will understand that Figure 5 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 5 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 5 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 5 The different configurations shown.

[0148] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0149] Example 4

[0150] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the video generation method provided in Embodiment 1.

[0151] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0152] Embodiments of this application also provide a computer program product, which, when executed on a data processing device, is adapted to perform the steps of a video generation method.

[0153] Embodiments of this application also provide a computer-readable storage medium, which includes a stored executable program, wherein the executable program controls the device where the computer-readable storage medium is located to execute the above-described video generation method when it runs.

[0154] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0155] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0156] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0160] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A video generation method, characterized in that, include: Acquire the target user's motion videos and motion data, and determine the target motion performed by the target user based on the motion videos; Based on the motion type of the target motion and the motion data, determine the motion performance index values ​​of each initial video frame in the motion video to obtain M motion performance index values, where M is a positive integer; Feature recognition is performed on the images of each initial video frame to obtain the material feature information of each initial video frame, and the material performance index value of each initial video frame is determined based on the material feature information to obtain M material performance index values; The value score of each initial video frame is determined based on the motion performance index value and the material performance index value of each initial video frame, and the initial video frames are filtered according to the value score to obtain the target video frames. The target video frame is input into a large language model to obtain a target video of the target motion, wherein the target video is obtained by the large language model processing the target video frame.

2. The method according to claim 1, characterized in that, Determining the motion performance index values ​​of each initial video frame in the motion video based on the motion type of the target motion and the motion data includes: For any initial video frame, obtain the preset weight information under the motion type, wherein the preset weight information includes multiple motion indicators and a preset weight value for each motion indicator; The initial motion index of the target motion is determined according to the motion type to obtain N initial motion indexes. The index value of each initial motion index of the initial video frame is calculated according to the motion data under the initial video frame to obtain N target index values, where N is a positive integer. Based on the target value of each initial motion indicator, an importance score is determined for each initial motion indicator, resulting in N importance scores; Obtain the first preset weight value for each initial motion index from the preset weight information to obtain N first preset weight values; The motion performance index value of the initial video frame is obtained by performing a weighted summation operation on the N importance scores based on the N first preset weight values.

3. The method according to claim 1, characterized in that, Perform feature recognition operations on the images of each initial video frame to obtain the material feature information of each initial video frame, including: For any initial video frame, identify the target information carried in the image of the initial video frame, wherein the target information includes at least one of the following: scene information, character information, and action information; The feature information of the target information is obtained, and the availability of the target information is determined based on the feature information to obtain an availability index value; Available information is filtered from the target information based on the availability index value, and the feature information of the available information is determined as the material feature information.

4. The method according to claim 1, characterized in that, The performance index values ​​for each initial video frame are determined based on the aforementioned material feature information, including: For any initial video frame, determine the complexity of each sub-feature information in the material feature information of the initial video frame to obtain multiple complexities; Determine the information type of the target information to which each sub-feature information belongs, and determine a second preset weight value for each sub-feature information based on the information type; The complexity of the sub-feature information is weighted and summed according to the second preset weight value to obtain the material performance index value of the initial video frame.

5. The method according to claim 1, characterized in that, The value score of each initial video frame is determined based on the motion performance index value and the material performance index value of each initial video frame, including: For any initial video frame, obtain the first weight value of the motion performance index value and obtain the second weight value of the material performance index value; The initial value score is obtained by weighting and summing the sports performance index value and the material performance index value according to the first weight value and the second weight value. Determine the matching degree between the motion data and the image of the initial video frame, and determine the matching coefficient based on the matching degree; The initial value score is multiplied by the matching coefficient to obtain the value score of the initial video frame.

6. The method according to claim 1, characterized in that, Based on the value scores, the initial video frames are filtered to obtain target video frames, including: Configure the time length of the sliding window, and group the video frames of the motion video according to the sliding window to obtain P sliding window frame sets, wherein each sliding window frame set includes a number of initial video frames corresponding to the time length, and P is a positive integer; Calculate the average value score of each initial video frame in each sliding window frame set to obtain P average values; The set of sliding window frames whose average value is greater than a preset threshold is determined as the target frame set, and initial video frames whose time difference with the target frame set is less than the time difference threshold are obtained from the motion video to obtain candidate video frames. The initial video frame and the candidate video frame in the target frame set are determined as the target video frame.

7. The method according to claim 1, characterized in that, Inputting the target video frame into the large language model yields the target video of the target motion, including: Receive music instruction information and scene transition information input by the target user, and generate prompt words based on the music instruction information and scene transition information; The prompt words and the target video frame are input into the large language model to obtain the target video.

8. A video generation apparatus, characterized in that, include: The acquisition unit is used to acquire motion videos and motion data of the target user, and determine the target motion performed by the target user based on the motion videos; The first determining unit is used to determine the motion performance index value of each initial video frame in the motion video according to the motion type of the target motion and the motion data, and obtain M motion performance index values, where M is a positive integer; The second determining unit is used to perform feature recognition operations on the images of each initial video frame to obtain the material feature information of each initial video frame, and to determine the material performance index value of each initial video frame based on the material feature information, thereby obtaining M material performance index values. The third determining unit is used to determine the value score of each initial video frame based on the motion performance index value and the material performance index value of each initial video frame, and to filter the initial video frames based on the value score to obtain target video frames. The processing unit is used to input the target video frame into a large language model to obtain a target video of the target motion, wherein the target video is obtained by the large language model processing the target video frame.

9. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the video generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the video generation method according to any one of claims 1 to 7.