Video description information generation method and device, equipment and storage medium

By obtaining the motion change information in the video, and using the large language model and the video description model for structured description, the problems of low efficiency and low accuracy of video description information generation are solved, and efficient and accurate video description is achieved.

CN120499467AActive Publication Date: 2025-08-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510947145.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-15
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In the prior art, video description information generation efficiency is low and the accuracy is not high, especially in sports scenes, description details are easily missing.

Method used

By obtaining the motion change information in the video to be analyzed, a large language model is used to perform semantic-level motion analysis, and a structured limb-level motion description is generated based on the video description model, including physical attribute calculations of time domain and frequency domain change information.

Benefits of technology

It improves the efficiency and accuracy of video description information generation, reduces the time-consuming generation, and the generated description is high in fine-grained, with high detail and high authenticity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499467A_ABST
    Figure CN120499467A_ABST
Patent Text Reader

Abstract

The invention relates to a video description information generation method and apparatus, a device and a storage medium. The method comprises the steps of determining motion change information of a to-be-analyzed object in a to-be-analyzed video; the motion change information is used for indicating physical attributes of object motion change in the to-be-analyzed video; inputting the to-be-analyzed video and first prompt information including the motion change information into the large language model to obtain motion event representation output by the large language model; the first prompt information is used for guiding the large language model to perform motion analysis of at least one semantic level on the to-be-analyzed video so as to generate structured motion representation; inputting the to-be-analyzed video and second prompt information including the motion event representation into a video description model to obtain video description information; the second prompt information is used for guiding the video description model to generate limb-level motion description for the to-be-analyzed video. According to the invention, the generation efficiency and accuracy of the video description information are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for generating video description information. Background Art

[0002] Video captioning is a key task in computer vision. It aims to extract features from a video, such as the subject, event, and scene, and convert them into a grammatically correct natural sentence or paragraph, known as the video description.

[0003] Video description information can be used in scenarios such as video generation and video search. For example, video generation typically requires a large number of video-description information sample pairs to train the video generation model. In related technologies, video description information can be generated through manual annotation or automatic annotation. The former is time-consuming, while the latter generally only provides a general summary of the video. Video descriptions of moving scenes, in particular, are often overly generalized and prone to missing details, reducing the efficiency and accuracy of video description generation. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, device, and storage medium for generating video description information to solve at least one problem in the related art. The technical solution of the present disclosure is as follows: According to one aspect of an embodiment of the present disclosure, a method for generating video description information is provided, including: Acquire a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; Acquiring motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object action change in the video to be analyzed; Inputting the video to be analyzed and the first prompt information into a large language model to obtain a motion event representation output by the large language model; the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis on the video to be analyzed at at least one semantic level to generate a structured motion representation; The video to be analyzed and the second prompt information are input into the video description model, and the video description information of the video to be analyzed is generated by the video description model; the second prompt information includes the motion event representation and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed.

[0005] In some embodiments, the motion event representation includes at least one motion representation unit matching the at least one semantic level, each semantic level corresponding to at least one motion representation unit; For each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the respective video frame range, and the text representation under the corresponding motion element; the motion element is used to represent the motion attributes related to the video motion description of the video to be analyzed.

[0006] In some embodiments, the first prompt information also includes auxiliary prompt information corresponding to the motion event representation, and the auxiliary prompt information includes at least one of hierarchical description information for representing each of the semantic levels, unit format information of the motion representation unit, and attribute definition information of the motion element.

[0007] In some embodiments, before inputting the video to be analyzed and the first prompt information into the large language model and obtaining the motion event representation output by the large language model, the method further includes: Obtaining a prompt template corresponding to the motion description information; The motion change information and the auxiliary prompt information are filled into the prompt template to generate the first prompt information.

[0008] In some embodiments, the at least one semantic level includes at least one of a first motion level for describing the overall action or overall movement of the object to be analyzed, a second motion level for describing the limb action of the object to be analyzed, and a third motion level for describing the action of the end part of the object to be analyzed.

[0009] In some embodiments, obtaining the motion change information of the object to be analyzed in the video to be analyzed includes: Extracting multiple video frames from the video to be analyzed; In the multiple video frames, skeleton key points of the object to be analyzed are detected to obtain posture estimation information of the object to be analyzed, where the posture estimation information is used to represent a position coordinate sequence of each skeleton key point of the object to be analyzed in the corresponding video frame; Based on the posture estimation information, motion change information of the object to be analyzed is determined.

[0010] In some embodiments, the motion change information includes time domain change information and frequency domain change information; and determining the motion change information of the object to be analyzed based on the posture estimation information includes: Determining temporal variation information of an object motion of the object to be analyzed based on the posture estimation information; the temporal variation information includes a first physical attribute for indicating a magnitude of a motion variation of a skeletal key point corresponding to the object motion; Based on the time domain change information, frequency domain change information of the object motion of the object to be analyzed is determined; the frequency domain change information includes a second physical attribute for indicating a change rhythm and / or movement intensity of the first physical attribute.

[0011] In some embodiments, determining the temporal variation information of the object motion of the object to be analyzed based on the posture estimation information includes: Based on the posture estimation information, obtaining motion speed information of each skeleton key point of the object to be analyzed; Based on the posture estimation information, calculating multiple joint angles of the object to be analyzed, as well as angular velocity information and average angular velocity information corresponding to each joint angle; The motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time domain variation information of the object motion of the object to be analyzed.

[0012] In some embodiments, determining the frequency domain change information of the object motion of the object to be analyzed based on the time domain change information includes: Converting the motion speed information and the angular velocity information into a motion spectrum through fast Fourier transform; performing feature extraction on the motion spectrum to obtain a first frequency domain feature representing a changing rhythm of the first physical attribute and a second frequency domain feature representing a motion intensity of the first physical attribute; The first frequency domain feature and the second frequency domain feature are determined as frequency domain change information of the object motion of the object to be analyzed.

[0013] According to another aspect of an embodiment of the present disclosure, a device for generating video description information is provided, including: A first acquisition module is configured to acquire a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; A first information determination module is configured to obtain motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object's motion change in the video to be analyzed; a semantic representation module configured to input the video to be analyzed and first prompt information into a large language model, and obtain a motion event representation output by the large language model; wherein the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis on the video to be analyzed at at least one semantic level to generate a structured motion representation; A generation module is configured to input the video to be analyzed and the second prompt information into a video description model, and generate video description information of the video to be analyzed through the video description model; the second prompt information includes the motion event representation and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed.

[0014] In some embodiments, the motion event representation includes at least one motion representation unit matching the at least one semantic level, each semantic level corresponding to at least one motion representation unit; For each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the respective video frame range, and the text representation under the corresponding motion element; the motion element is used to represent the motion attributes related to the video motion description of the video to be analyzed.

[0015] In some embodiments, the first prompt information also includes auxiliary prompt information corresponding to the motion event representation, and the auxiliary prompt information includes at least one of hierarchical description information for representing each of the semantic levels, unit format information of the motion representation unit, and attribute definition information of the motion element.

[0016] In some embodiments, the apparatus further comprises: A template acquisition module is configured to execute acquisition of a prompt template corresponding to the motion description information; The prompt information construction module is configured to fill the motion change information and the auxiliary prompt information into the prompt template to generate the first prompt information.

[0017] In some embodiments, the at least one semantic level includes at least one of a first motion level for describing the overall action or overall movement of the object to be analyzed, a second motion level for describing the limb action of the object to be analyzed, and a third motion level for describing the action of the end part of the object to be analyzed.

[0018] In some embodiments, the first information determination module includes: an extraction unit, configured to extract a plurality of video frames from the video to be analyzed; a detection unit configured to perform skeletal key point detection on the object to be analyzed in the plurality of video frames to obtain posture estimation information of the object to be analyzed, wherein the posture estimation information is used to represent a position coordinate sequence of each skeletal key point of the object to be analyzed in the corresponding video frames; The determining unit is configured to determine the motion change information of the object to be analyzed based on the posture estimation information.

[0019] In some embodiments, the motion change information includes time domain change information and frequency domain change information; and the first information determination module further includes: A time domain information determination unit is configured to determine time domain change information of the object motion of the object to be analyzed based on the posture estimation information; the time domain change information includes a first physical attribute for indicating a movement change size of a skeletal key point corresponding to the object motion; The frequency domain information determination unit is configured to determine the frequency domain change information of the object motion of the object to be analyzed based on the time domain change information; the frequency domain change information includes a second physical attribute for indicating the change rhythm and / or intensity of the movement of the first physical attribute.

[0020] In some embodiments, the time domain information determination submodule is further configured to execute: Based on the posture estimation information, obtaining motion speed information of each skeleton key point of the object to be analyzed; Based on the posture estimation information, calculating multiple joint angles of the object to be analyzed, as well as angular velocity information and average angular velocity information corresponding to each joint angle; The motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time domain variation information of the object motion of the object to be analyzed.

[0021] In some embodiments, the frequency domain information determination submodule is further configured to perform: Converting the motion speed information and the angular velocity information into a motion spectrum through fast Fourier transform; performing feature extraction on the motion spectrum to obtain a first frequency domain feature representing a changing rhythm of the first physical attribute and a second frequency domain feature representing a motion intensity of the first physical attribute; The first frequency domain feature and the second frequency domain feature are determined as frequency domain change information of the object motion of the object to be analyzed.

[0022] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method for generating video description information as described in any of the above embodiments.

[0023] According to another aspect of the present disclosure, an electronic device is provided, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for generating video description information as described in any one of the above embodiments.

[0024] According to another aspect of an embodiment of the present disclosure, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the method for generating video description information provided in any of the above embodiments is implemented.

[0025] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects: The disclosed embodiment obtains a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; obtains motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object action change in the video to be analyzed; the video to be analyzed and a first prompt information are input into a large language model to obtain a motion event representation output by the large language model; the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis of the video to be analyzed at at least one semantic level to generate a structured motion representation; the video to be analyzed and the second prompt information are input into the video description model, and the video description information of the video to be analyzed is generated by the video description model; the second prompt information includes the motion event representation and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed. In this way, by deeply integrating the calculation of motion physical properties with hierarchical language analysis, through the link of physical constraints, structured motion event representation and model prompt information, fine-grained video description information at the limb level is automatically generated, reducing the time consumption of video description information generation and improving the efficiency and accuracy of video description information generation.

[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of the present disclosure, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0028] Figure 1 The figure is a flowchart of a method for generating video description information according to an exemplary embodiment.

[0029] Figure 2 The figure is a schematic diagram showing a process of generating video description information according to an exemplary embodiment.

[0030] Figure 3 The flowchart of another method for generating video description information is shown according to an exemplary embodiment.

[0031] Figure 4 The present invention is a flowchart showing a step of determining time-domain variation information of an object motion of an object to be analyzed according to an exemplary embodiment.

[0032] Figure 5 The present invention is a flowchart showing a step of determining frequency domain change information of an object motion of an object to be analyzed according to an exemplary embodiment.

[0033] Figure 6 The figure is a block diagram of a device for generating video description information according to an exemplary embodiment.

[0034] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0035] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0036] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0038] Figure 1 The figure is a flowchart of a method for generating video description information according to an exemplary embodiment. Figure 2 It is a schematic diagram of a process for generating video description information according to an exemplary embodiment. The method for generating video description information disclosed herein can be applied to electronic devices, which can be terminals, servers, or similar computing devices. Take the electronic device as an example for explanation, wherein the server includes but is not limited to an independent server, or a server cluster or distributed system composed of multiple physical servers, or one or more cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, intermediate services, domain name services, security services, and big data and artificial intelligence platforms. As Figure 1 and Figure 2 As shown, the method includes the following steps.

[0039] In step S101 , a video to be analyzed is obtained, where the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed.

[0040] The video to be analyzed refers to the video for which video description information is to be generated. The video to be analyzed can be any video that reflects the motion or movement of the object to be analyzed, such as a filmed video, a clipped video, or an animation. The video to be analyzed includes the object to be analyzed and the object actions performed by the object to be analyzed. For example, the object to be analyzed may include, but is not limited to, people, virtual characters, animals, plants, objects, etc. The number of objects to be analyzed can be one or more. The object actions performed by the object to be analyzed may include, but are not limited to, at least one of dancing, playing ball, and surfing.

[0041] In the disclosed embodiments, an electronic device can obtain an initial video from various sources and segment the initial video into video segments, which are the videos to be analyzed. The video to be analyzed can then be uniformly sampled, extracting a preset number n1 of video frames to facilitate subsequent processing of these video frames. For example, n1 can be any value between 10 and 100 frames, such as 32 frames. Each extracted video frame can include at least one object to be analyzed and an object action performed by the object to be analyzed, such as a video frame of Xiaohong dancing.

[0042] In step S103, motion change information of the object to be analyzed in the video to be analyzed is obtained; the motion change information is used to indicate the physical properties of the object's action changes in the video to be analyzed.

[0043] The pose estimation information is used to reflect the pose of the object to be analyzed in each video frame. Exemplarily, the pose estimation information includes the position coordinates of each key point of the object to be analyzed in each video frame. The key points may be key locations associated with the object's actions.

[0044] In an embodiment of the present disclosure, an electronic device may obtain motion change information of an object to be analyzed in a video to be analyzed, calculated by another device. This motion change information indicates the physical properties of the object's motion changes in the video to be analyzed. For example, the physical properties may include, but are not limited to, at least one of speed and acceleration of the motion.

[0045] In some embodiments, obtaining motion change information of the object to be analyzed in the video to be analyzed includes: extracting multiple video frames from the video to be analyzed; In multiple video frames, skeleton key points of the object to be analyzed are detected to obtain posture estimation information of the object to be analyzed, and the posture estimation information is used to represent the position coordinate sequence of each skeleton key point of the object to be analyzed in the corresponding video frame; Based on the posture estimation information, the motion change information of the object to be analyzed is determined.

[0046] In an embodiment of the present disclosure, taking the object to be analyzed as a person as an example, the electronic device can uniformly sample the video to be analyzed and extract multiple video frames. Afterwards, a general human posture estimation model (such as the AlphaPose model) is called to perform skeletal key point (such as Halpe-26 points) detection processing on the object to be analyzed in each video frame extracted from the video to be analyzed, and obtain the posture estimation information output by the posture estimation model. The posture estimation information is used to characterize the position coordinate sequence of each skeletal key point of the object to be analyzed in the corresponding video frame. Next, the electronic device calculates the physical properties of the motion based on the posture estimation information and determines the motion change information of the object to be analyzed. The motion change information is used to indicate the physical properties of the object action change in the video to be analyzed.

[0047] In some embodiments, as Figure 3 As shown, the motion change information includes time domain change information and frequency domain change information; based on the posture estimation information, determining the motion change information of the object to be analyzed includes: In step S301, based on the posture estimation information, the temporal variation information of the object motion of the object to be analyzed is determined; the temporal variation information includes a first physical attribute for indicating the magnitude of the motion variation of the skeletal key points corresponding to the object motion; In step S303, frequency domain change information of the object motion of the object to be analyzed is determined based on the time domain change information; the frequency domain change information includes a second physical attribute indicating a change rhythm of the first physical attribute and / or a motion intensity.

[0048] In an embodiment of the present disclosure, motion change information is obtained from two levels, time domain and frequency domain, respectively. At this time, the motion change information includes time domain change information and frequency domain change information. Specifically, the electronic device first performs a time domain change analysis on the posture estimation information, and calculates the time domain change information of the object motion of the object to be analyzed. Among them, the time domain change information is used to reflect the change state of the time domain attribute. The time domain attribute can be a quantitative description of the motion state of each part of the object to be analyzed over time. The time domain change information includes a first physical attribute for indicating the size of the motion change of the skeletal key point corresponding to the object motion. Optionally, the first physical attribute includes but is not limited to at least one of the velocity, joint angle, angular velocity, etc. of each skeletal joint.

[0049] Afterwards, the electronic device performs frequency domain change analysis on the above-mentioned time domain change information to extract frequency domain change information of the object motion of the object to be analyzed. The frequency domain change information is used to reflect the change state of the frequency domain attribute. The frequency domain attribute can be at least one of the intensity, periodicity and complexity of the action. The frequency domain change information includes a second physical attribute for indicating the changing rhythm of the first physical attribute and / or the intensity of the movement. Optionally, the second physical attribute can be at least one of the frequency domain characteristics of the object motion of the object to be analyzed in different motion components, such as energy, main frequency, rhythm, etc.

[0050] In this way, by introducing two types of motion change information, time domain and frequency domain, the dual-domain physical representation provides detailed physical constraints for complex motion dynamics, thereby improving the generation details and precision of video description information, thereby improving the accuracy of video description information.

[0051] In some embodiments, as Figure 4 As shown, based on the posture estimation information, determining the time domain change information of the object motion of the object to be analyzed includes: In step S401, based on the posture estimation information, the motion speed information of each skeleton key point of the object to be analyzed is obtained; In step S403, based on the posture estimation information, multiple joint angles of the object to be analyzed, as well as angular velocity information and average angular velocity information corresponding to each joint angle are calculated; In step S405 , the motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time-domain variation information of the object motion of the object to be analyzed.

[0052] Generally speaking, the motion trajectory of an object to be analyzed can be viewed as a transition from an initial state (s0, p0) to a final state (s1, p1). Here, s0 and s1 represent different motion positions, and p0 and p1 represent different postures. This transition process can be understood as comprising the superposition of two sub-motions. Specifically, in the first sub-motion, the object to be analyzed moves from (s0, p0) to (s1, p0), maintaining the posture p0 unchanged while translating the position from s0 to s1. In the second sub-motion, the object to be analyzed transitions from (s0, p0) to (s0, p1), changing the posture from p0 to p1 while maintaining the fixed position s0. Based on the basic principles of vector decomposition and composition, the transition from the initial state (s0, p0) to the new final state (s1, p1) can be viewed as a combination of these two sub-motions.

[0053] Based on this, when determining the temporal variation of the object's motion, the motion of the object being analyzed is decomposed into two physical components: position translation (spatial displacement) and posture change (limb movement). These two physical components can be measured by velocity and angular velocity, respectively.

[0054] Specifically, taking the object to be analyzed as an example, including a person, the human skeleton key point K and the position coordinates of each skeleton key point in the corresponding video frame are obtained through posture estimation information. For example, the position coordinates of a certain skeleton key point a1 at the t-1th time point are (xt-1, yt-1), and the position coordinates at the tth time point are (xt, yt). Afterwards, the position change distance of the skeleton key point is determined based on the position coordinates in different video frames, and the ratio between the position change distance and the corresponding time interval is calculated to obtain the motion speed information of each skeleton key point. This motion speed information reflects the physical component of the motion in position translation. Optionally, the motion speed information here can include at least one of the motion speed of each skeleton key point and the average motion speed of each skeleton key point.

[0055] Optionally, multiple (e.g., 17) joint angles are selected from, for example, 26 human skeletal key points to capture meaningful aspects of body posture and dynamics. The joint angle is the angle corresponding to the joint formed by pairwise skeletal key points. In practical applications, the two limb segments that form an angle at each joint can be obtained, and the cosine theorem can be used to calculate the angle of each joint angle separately. After obtaining the angle of the joint angle, the angular velocity information and average angular velocity information of each joint angle changing over time can be calculated. Thereafter, the motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time domain variation information of the object motion of the object to be analyzed.

[0056] In some embodiments, as Figure 5 As shown, determining the frequency domain change information of the object motion of the object to be analyzed based on the time domain change information includes: In step S501, the motion speed information and the angular velocity information are converted into a motion spectrum by fast Fourier transform; In step S503, feature extraction is performed on the motion spectrum to obtain a first frequency domain feature representing the changing rhythm of the first physical attribute and a second frequency domain feature representing the intensity of the motion of the first physical attribute; In step S505 , the first frequency domain feature and the second frequency domain feature are determined as frequency domain change information of the object motion of the object to be analyzed.

[0057] In the embodiment of the present disclosure, in order to further capture the subtle differences in rhythm, we have integrated frequency domain analysis on the basis of time domain analysis. In the frequency domain analysis, the motion speed information and angular velocity information are first converted into a motion spectrum through fast Fourier transform; then the motion spectrum is subjected to feature extraction to obtain a first frequency domain feature and a second frequency domain feature. Among them, the first frequency domain feature is used to characterize the changing rhythm of the first physical attribute. The second frequency domain feature is used to characterize the intensity of the movement of the first physical attribute. Afterwards, the first frequency domain feature and the second frequency domain feature are determined as the frequency domain change information of the object motion of the object to be analyzed. Exemplarily, the first frequency domain feature may include the speed change rhythm and angular velocity change rhythm The second frequency domain feature can include at least one of motion energy, the proportion of high-frequency signals, and the standard deviation of the spectrum. In this way, by analyzing the rhythmic changes in motion intensity through fast Fourier transform, it is possible to effectively distinguish between intense and subtle motions, thereby improving the detail and accuracy of the video description.

[0058] In step S105, the video to be analyzed and the first prompt information are input into the large language model to obtain the motion event representation output by the large language model; the first prompt information includes motion change information and is at least used to guide the large language model to perform motion analysis of the video to be analyzed at least one semantic level to generate a structured motion representation.

[0059] The large language model refers to a pre-trained large language model, or a model built based on a pre-trained large language model, such as a large video language model. The large language model here has the function of generating corresponding content based on prompt instructions, and this disclosure does not specifically limit its specific structure.

[0060] Motion event representation may refer to representing the object action of an object to be analyzed in the form of a structured, parsing-oriented semantic framework.

[0061] In the disclosed embodiment, after acquiring motion change information, the electronic device adds the motion change information to a first prompt template to construct a first prompt message. The video to be analyzed and the first prompt message are then fed into a large language model. The first prompt message guides the large language model to perform motion analysis on the video to be analyzed at at least one semantic level, obtaining a structured motion event representation output by the large language model. In this way, the model automatically generates a structured motion event representation, ensuring the accuracy and comprehensiveness of subsequent fine-grained action decomposition.

[0062] In some embodiments, at least one semantic level includes at least one of a first motion level for describing the overall action or overall movement of the object to be analyzed, a second motion level for describing the limb action of the object to be analyzed, and a third motion level for describing the action of the end part of the object to be analyzed.

[0063] In the embodiment of the present disclosure, the object motion of the object to be analyzed is decomposed into three different but interrelated semantic levels, namely, the first motion level, the second motion level, and the third motion level.

[0064] The first motion level is used to describe the overall movement or action of the object to be analyzed, i.e., the individual level. For example, the object's actions at the first motion level can include "walking" or "jumping," which are used to describe the whole body displacement or overall movement.

[0065] The second level of motion, known as the limb-level, describes the subject's physical movements. Examples of such movements include "swinging the left arm" and "bending the right leg." These actions are broken down into the coordinated movements of the trunk and limbs, summarizing the coordinated movements of the main body parts.

[0066] The third level of motion, also known as the distal level, describes the movements of the extremities of the object being analyzed. Examples of this level include movements like "twisting the wrist" and "nodding the head." This level is further subdivided into movements of the hands, feet, and head, highlighting the subtle joints at the extremities.

[0067] In some embodiments, the motion event representation includes at least one motion representation unit matching at least one semantic level, each semantic level corresponding to at least one motion representation unit; For each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the respective video frame range, and the text representation under the corresponding motion element; the motion element is used to represent the motion attributes related to the video motion description of the video to be analyzed.

[0068] Optionally, based on linguistic knowledge, a motion event can be uniquely determined by motion elements such as "agent", "patient", "action type", "scene", "direction", and "amplitude". In order to ensure that the descriptions of "agent" and "patient" are not ambiguous, two motion elements, "modifier of agent" and "modifier of patient", can be added to assist in determining motion events of different objects.

[0069] On this basis, by defining a preset number n3 (e.g., 8) of motion elements closely associated with the motion descriptions in the video, and incorporating the motion change information obtained from physical calculations and the definitions of these motion elements into the prompt word, the model is guided to generate motion representation units that characterize each motion event in the video to be analyzed and match each semantic level. Each semantic level corresponds to at least one motion representation unit. For example, the first motion level includes one motion representation unit; the second and third motion levels can each include multiple motion representation units. Subsequently, at least one motion representation unit from each semantic level is combined in the order of the first, second, and third motion levels to form a complete motion event representation.

[0070] Among them, for each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the range of each video frame, and the text representation under the corresponding motion element. The motion element is used to characterize the motion attributes related to the video motion description of the video to be analyzed. Each motion representation unit can be a standardized motion representation. Exemplarily, the unit format of each motion representation unit may include: "[{begin_frame, end_frame}, (motion_subject, motion, motion_objects, motion_adverbial, motion_amplitude), (motion _subjects, modifiers_subject), (otion_objectly, modifers_objects)]", wherein the first unit {begin_frame, end_frame} represents the start frame and end frame of the motion. The second unit (motion_subject, motion, motion_objects, motion_adverbial, motion_amplitude) represents the subject of the action, the description of the action, the recipient of the action, the adverbial of the action, and the amplitude of the action. The third unit (motion_subjects, modifiers_subject) represents the modifiers of the action subject, and the fourth unit (otion_objectly, modifers_objects) represents the modifiers of the action receptor.

[0071] In some embodiments, the first prompt information also includes auxiliary prompt information corresponding to the motion event representation, and the auxiliary prompt information includes at least one of hierarchical description information for representing each semantic level, unit format information of the motion representation unit, and attribute definition information of the motion element.

[0072] Optionally, the hierarchical description information may include a hierarchical text description, such as "The specified movement is divided into the human body level (movement of the entire human body), the limb level (movement of the limbs), and the distal level (movement of the distal limbs, such as the palms, soles, fingers, and toes). Please output "human body level" in the second line, then output all human body level movement information, and output "limb level" in a new line, then output all limb level information, and then enter "distal level" in a new line, and output all distal level information." The unit format information of the motion representation unit is as above and will not be repeated here. The attribute definition information of the motion element may include the attribute definition text of each motion element that fully characterizes the motion event.

[0073] In some embodiments, before inputting the video to be analyzed and the first prompt information into the large language model and obtaining the motion event representation output by the large language model, the method further includes: Get the prompt template corresponding to the motion description information; The motion change information and the auxiliary prompt information are filled into the prompt template to generate the first prompt information.

[0074] Optionally, the electronic device can obtain a prompt template corresponding to the motion description information, and fill in the motion change information and auxiliary prompt information into the corresponding positions in the prompt template to generate a first prompt information, and guide the large language model through the first prompt information to generate a structured motion event representation that conforms to the unit format information of the motion representation unit.

[0075] In step S107, the video to be analyzed and the second prompt information are input into the video description model, and the video description information of the video to be analyzed is generated by the video description model; the second prompt information includes motion event representation and is at least used to guide the video description model to generate limb-level motion description for the video to be analyzed.

[0076] The video description model may be a pre-trained large language model, or a model constructed based on a pre-trained large language model, such as a video language large model, etc. The present disclosure does not specifically limit the specific structure of the video description model.

[0077] In an embodiment of the present disclosure, after the electronic device obtains the structured motion event representation, it adds the motion event representation to the second prompt template to construct the second prompt information, and inputs the video to be analyzed and the second prompt information into the video description model. The second prompt information is used to guide the video description model to generate a limb-level motion description for the video to be analyzed, so as to obtain the video description information output by the video description model. In this way, the motion event representation is used to guide the model to expand the motion to the limb level, and output video description information in natural language that conforms to grammatical, semantic, and physical property constraints. The video description information is a dense and fine-grained motion description, that is, the video description information is used to achieve a continuous and detailed text restoration of the complex action process in the video.

[0078] The above embodiment deeply integrates the calculation of motion physical properties with hierarchical language analysis, and automatically generates fine-grained video description information through the link of physical constraints, structured motion event representation and model prompt information. There is no need to rely on manual frame-by-frame generation of video description information, which reduces the time consumption of video description information generation and improves the efficiency and accuracy of video description information generation.

[0079] Furthermore, since manual frame-by-frame video description generation is no longer necessary, the scale and detail density of video-text pairing data are greatly expanded, providing a high-value data foundation for large-scale model training and downstream tasks such as video generation, action recognition, and intelligent editing. By deeply integrating motion physics with language structure, this method bridges the gap between video physical properties and structured semantic expression, effectively reducing hallucinations in large models and improving the authenticity, consistency, and interpretability of generated descriptions.

[0080] Figure 6 FIG. 1 is a block diagram of a device for generating video description information according to an exemplary embodiment. Figure 6 , the device comprises: A first acquisition module 601 is configured to acquire a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; The first information determination module 602 is configured to obtain motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object's motion change in the video to be analyzed; The semantic representation module 603 is configured to input the video to be analyzed and the first prompt information into the large language model to obtain the motion event representation output by the large language model; the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis on the video to be analyzed at least at one semantic level to generate a structured motion representation; The generation module 604 is configured to input the video to be analyzed and the second prompt information into a video description model, and generate video description information of the video to be analyzed through the video description model; the second prompt information includes the motion event representation, and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed.

[0081] In some embodiments, the motion event representation includes at least one motion representation unit matching the at least one semantic level, each semantic level corresponding to at least one motion representation unit; For each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the respective video frame range, and the text representation under the corresponding motion element; the motion element is used to represent the motion attributes related to the video motion description of the video to be analyzed.

[0082] In some embodiments, the first prompt information also includes auxiliary prompt information corresponding to the motion event representation, and the auxiliary prompt information includes at least one of hierarchical description information for representing each of the semantic levels, unit format information of the motion representation unit, and attribute definition information of the motion element.

[0083] In some embodiments, the apparatus further comprises: A template acquisition module is configured to execute acquisition of a prompt template corresponding to the motion description information; The prompt information construction module is configured to fill the motion change information and the auxiliary prompt information into the prompt template to generate the first prompt information.

[0084] In some embodiments, the at least one semantic level includes at least one of a first motion level for describing the overall action or overall movement of the object to be analyzed, a second motion level for describing the limb action of the object to be analyzed, and a third motion level for describing the action of the end part of the object to be analyzed.

[0085] In some embodiments, the first information determination module includes: an extraction unit, configured to extract a plurality of video frames from the video to be analyzed; a detection unit configured to perform skeletal key point detection on the object to be analyzed in the plurality of video frames to obtain posture estimation information of the object to be analyzed, wherein the posture estimation information is used to represent a position coordinate sequence of each skeletal key point of the object to be analyzed in the corresponding video frames; The determining unit is configured to determine the motion change information of the object to be analyzed based on the posture estimation information.

[0086] In some embodiments, the motion change information includes time domain change information and frequency domain change information; and the first information determination module further includes: A time domain information determination unit is configured to determine time domain change information of the object motion of the object to be analyzed based on the posture estimation information; the time domain change information includes a first physical attribute for indicating a movement change size of a skeletal key point corresponding to the object motion; The frequency domain information determination unit is configured to determine the frequency domain change information of the object motion of the object to be analyzed based on the time domain change information; the frequency domain change information includes a second physical attribute for indicating the change rhythm and / or intensity of the movement of the first physical attribute.

[0087] In some embodiments, the time domain information determination submodule is further configured to execute: Based on the posture estimation information, obtaining motion speed information of each skeleton key point of the object to be analyzed; Based on the posture estimation information, calculating multiple joint angles of the object to be analyzed, as well as angular velocity information and average angular velocity information corresponding to each joint angle; The motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time domain variation information of the object motion of the object to be analyzed.

[0088] In some embodiments, the frequency domain information determination submodule is further configured to perform: Converting the motion speed information and the angular velocity information into a motion spectrum through fast Fourier transform; performing feature extraction on the motion spectrum to obtain a first frequency domain feature representing a changing rhythm of the first physical attribute and a second frequency domain feature representing a motion intensity of the first physical attribute; The first frequency domain feature and the second frequency domain feature are determined as frequency domain change information of the object motion of the object to be analyzed.

[0089] It should be noted that, regarding the device in the above embodiment, the specific method and beneficial effects of each step have been described in detail in the embodiment of the aforementioned method, and will not be elaborated here.

[0090] Figure 7 FIG. 1 is a block diagram of an electronic device according to an exemplary embodiment. Figure 7 The electronic device includes a processor; a memory for storing instructions executable by the processor; wherein, when the processor is configured to execute the instructions stored in the memory, the steps of the method for generating video description information in any of the above embodiments are implemented.

[0091] The electronic device may be a terminal, a server or a similar computing device. For example, the electronic device is a server. Figure 7 This is a block diagram of an electronic device for generating video description information, according to an exemplary embodiment. The electronic device 1200 may vary significantly depending on configuration or performance. It may include one or more central processing units (CPUs) 1210 (processor 1210 may include, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA) processing device), a memory 1230 for storing data, and one or more storage media 1220 (e.g., one or more mass storage devices) for storing applications 1223 or data 1222. The memory 1230 and storage media 1220 may be either transient or persistent storage. The program stored in the storage medium 1220 may include one or more modules, each of which may include a series of instructions for operating on the electronic device. Furthermore, the CPU 1210 may be configured to communicate with the storage medium 1220 to execute the series of instructions stored in the storage medium 1220 on the electronic device 1200.

[0092] The electronic device 1200 may also include one or more power supplies 1260, one or more wired or wireless network interfaces 1250, one or more input and output interfaces 1240, and / or one or more operating systems 1221, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0093] The input / output interface 1240 can be used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the electronic device 1200. In one embodiment, the input / output interface 1240 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In an exemplary embodiment, the input / output interface 1240 may be a radio frequency (RF) module for wireless communication with the Internet.

[0094] It can be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components than shown, or with Figure 7 Different configurations shown.

[0095] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions. The instructions can be executed by a processor of the electronic device 1200 to perform the above method. Alternatively, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0096] In an exemplary embodiment, a computer storage medium is further provided. When instructions in the computer storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the method provided in any one of the above embodiments.

[0097] In an exemplary embodiment, a computer program product is also provided, comprising a computer program / instructions that, when executed by a processor, implements the method provided in any of the above-described embodiments. Optionally, the computer program is stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the method provided in any of the above-described embodiments.

[0098] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0099] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0100] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating video description information, characterized in that: include: Acquire a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; Acquiring motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object action change in the video to be analyzed; Inputting the video to be analyzed and the first prompt information into a large language model to obtain a motion event representation output by the large language model; the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis on the video to be analyzed at at least one semantic level to generate a structured motion representation; The video to be analyzed and the second prompt information are input into a video description model, and the video description information of the video to be analyzed is generated by the video description model; the second prompt information includes the motion event representation and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed.

2. The method according to claim 1, characterized in that The motion event representation includes at least one motion representation unit matching the at least one semantic level, each semantic level corresponding to at least one motion representation unit; For each semantic level, each motion representation unit includes the object action performed by the object to be analyzed within the respective video frame range, and the text representation under the corresponding motion element; the motion element is used to represent the motion attributes related to the video motion description of the video to be analyzed.

3. The method according to claim 2, characterized in that The first prompt information also includes auxiliary prompt information corresponding to the motion event representation, and the auxiliary prompt information includes at least one of hierarchical description information for representing each semantic level, unit format information of the motion representation unit, and attribute definition information of the motion element.

4. The method according to claim 3, characterized in that Before inputting the video to be analyzed and the first prompt information into the large language model and obtaining the motion event representation output by the large language model, the method further includes: Obtaining a prompt template corresponding to the motion description information; The motion change information and the auxiliary prompt information are filled into the prompt template to generate the first prompt information.

5. The method according to any one of claims 1 to 4, characterized in that: The at least one semantic level includes at least one of a first motion level for describing the overall action or overall movement of the object to be analyzed, a second motion level for describing the limb action of the object to be analyzed, and a third motion level for describing the action of the end part of the object to be analyzed.

6. The method according to claim 1, characterized in that The acquiring of the motion change information of the object to be analyzed in the video to be analyzed includes: Extracting multiple video frames from the video to be analyzed; In the multiple video frames, skeleton key points of the object to be analyzed are detected to obtain posture estimation information of the object to be analyzed, where the posture estimation information is used to represent a position coordinate sequence of each skeleton key point of the object to be analyzed in the corresponding video frame; Based on the posture estimation information, motion change information of the object to be analyzed is determined.

7. The method according to claim 6, characterized in that The motion change information includes time domain change information and frequency domain change information; and determining the motion change information of the object to be analyzed based on the posture estimation information includes: Determining temporal variation information of an object motion of the object to be analyzed based on the posture estimation information; the temporal variation information includes a first physical attribute for indicating a magnitude of a motion variation of a skeletal key point corresponding to the object motion; Based on the time domain change information, frequency domain change information of the object motion of the object to be analyzed is determined; the frequency domain change information includes a second physical attribute for indicating a change rhythm and / or movement intensity of the first physical attribute.

8. The method according to claim 7, characterized in that The determining, based on the posture estimation information, the time domain variation information of the object motion of the object to be analyzed includes: Based on the posture estimation information, obtaining motion speed information of each skeleton key point of the object to be analyzed; Based on the posture estimation information, calculating multiple joint angles of the object to be analyzed, as well as angular velocity information and average angular velocity information corresponding to each joint angle; The motion speed information, the angular velocity information corresponding to the joint angle, and the average angular velocity information are determined as the time domain variation information of the object motion of the object to be analyzed.

9. The method according to claim 8, characterized in that The determining, based on the time domain change information, the frequency domain change information of the object motion of the object to be analyzed includes: Converting the motion speed information and the angular velocity information into a motion spectrum through fast Fourier transform; performing feature extraction on the motion spectrum to obtain a first frequency domain feature representing a changing rhythm of the first physical attribute and a second frequency domain feature representing a motion intensity of the first physical attribute; The first frequency domain feature and the second frequency domain feature are determined as frequency domain change information of the object motion of the object to be analyzed.

10. A device for generating video description information, characterized in that: include: A first acquisition module is configured to acquire a video to be analyzed, wherein the video to be analyzed includes an object to be analyzed and an object action performed by the object to be analyzed; A first information determination module is configured to obtain motion change information of the object to be analyzed in the video to be analyzed; the motion change information is used to indicate the physical properties of the object's motion change in the video to be analyzed; a semantic representation module configured to input the video to be analyzed and first prompt information into a large language model, and obtain a motion event representation output by the large language model; wherein the first prompt information includes the motion change information and is at least used to guide the large language model to perform motion analysis on the video to be analyzed at at least one semantic level to generate a structured motion representation; A generation module is configured to input the video to be analyzed and the second prompt information into a video description model, and generate video description information of the video to be analyzed through the video description model; the second prompt information includes the motion event representation and is at least used to guide the video description model to generate a limb-level motion description for the video to be analyzed.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for generating video description information according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for generating video description information according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-feature fusion and time-space attention mechanism combination-based video description method

    CN108388900A

  • Athletic contest commentary generation method based on large language model

    CN119421013A

  • Curling match video description method based on large language model

    CN119603525A

  • Video analysis method and device, equipment and storage medium

    CN119919854A