Methods, devices, equipment, media, and products for generating digital human video ringback tones
By combining human motion information from benchmark audio and video libraries to generate digital human video ringback tones, the problem of insufficient correlation between music and visuals is solved, resulting in more attractive and personalized video ringback tone displays.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2024-08-08
- Publication Date
- 2026-05-26
AI Technical Summary
In existing video ringback tone displays, the connection between music and visuals is poor, resulting in a lack of novelty and engagement for users.
By acquiring the baseline video and the character's motion trajectory from the baseline audio, the target video is selected from the video library, the character's motion change information in the video frame corresponding to the keyword is determined, and the motion trajectory of the digital human is generated based on this information, and finally merged into a digital human video ringtone.
The generated digital human video ringback tones are more relevant and engaging with the music, providing a personalized and innovative calling experience and improving the utilization of the resource library videos and the user experience.
Smart Images

Figure CN119071389B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human technology, and in particular to methods, apparatus, devices, media, and products for generating digital human video ringback tones. Background Technology
[0002] Currently, most video ringback tones displayed in the market are either edited versions of existing videos or manually created animations. Edited video ringback tones, due to limitations in video material, often fail to fully match the rhythm and emotion of the music, resulting in a lack of coherence and appeal in the viewing experience. While animated video ringback tones can achieve a higher degree of creativity and personalization, their production costs are high, and they lack authenticity and emotional expression, making it difficult to resonate with users. In summary, existing video ringback tones display solutions suffer from poor correlation between music and visuals, resulting in a lack of novelty and engagement for the user experience. Summary of the Invention
[0003] This application provides a method, apparatus, device, medium, and product for generating digital human video ringback tones, in order to solve the technical problem that the existing video ringback tones display schemes have poor correlation between music and images, resulting in a lack of novelty and user experience.
[0004] To address the aforementioned technical problems, the embodiments of this application provide the following aspects:
[0005] In a first aspect, embodiments of this application provide a method for generating digital human video ringback tones, the method comprising:
[0006] Based on the reference audio, a reference video is determined, and the motion trajectory of the reference character in the reference video is obtained;
[0007] Target videos are selected from the video library based on the baseline character movement trajectory, wherein the background music of each video in the video library is the baseline audio.
[0008] Identify at least one keyword in the lyrics of the reference audio;
[0009] Based on the keywords, each video in the video library is filtered to obtain information on the changes in the movements of the characters in the video frames corresponding to the keywords.
[0010] Based on the aforementioned motion change information, the motion trajectory of the digital human is determined;
[0011] The reference audio, the target video, and the digital human are merged into a digital human video ringback tone.
[0012] Optionally, before determining the digital human's motion trajectory based on the motion change information, the method further includes:
[0013] Obtain the digital avatar information input by the user;
[0014] The digital human is generated based on the digital human image information.
[0015] Optionally, obtaining the reference character motion trajectory in the reference video includes:
[0016] Obtain the skeletal point coordinates of the person in the reference video, and set the center skeletal point coordinates in the skeletal point coordinates; wherein, the skeletal point coordinates include at least one of the following: head skeletal point coordinates, neck skeletal point coordinates, hand skeletal point coordinates, arm skeletal point coordinates, and leg skeletal point coordinates;
[0017] Determine the distances from the coordinates of the remaining bone points (excluding the coordinates of the central bone point) to the coordinates of the central bone point, and determine the limb movements of the character based on the change information of the distances.
[0018] Based on the described body movements, the baseline character movement trajectory is generated.
[0019] Optionally, filtering target videos from the video library based on the baseline character movement trajectory includes:
[0020] Determine the movement trajectory of the characters in each video in the video library;
[0021] The similarity of each character's movement trajectory to the baseline character's movement trajectory is compared to determine the similarity of each character's movement trajectory to the baseline character's movement trajectory.
[0022] Videos containing the character's movement trajectory with a similarity greater than or equal to a preset threshold are identified as candidate target videos;
[0023] The target video is determined from the candidate target videos according to the user's instructions.
[0024] Optionally, each video in the video library has the same video type as the baseline video, and the video type includes at least one of the following: dance type, performance type.
[0025] Optionally, the character's motion change information includes at least one of the following: the character's position change information and the speed information of the position change information. Each keyword corresponds to a set of video frames. Based on the motion change information, the motion trajectory of the digital human is determined as follows:
[0026] In a set of video frames corresponding to the keyword, determine the position change information of the person in each video frame and the speed information of the position change information. Based on the position change information of the person in each video frame and the speed information of the position change information, calculate the average position change information and the average speed information of the position change information.
[0027] The motion trajectory of the digital human is determined based on the average position change information and the average velocity information of the position change information.
[0028] Secondly, embodiments of this application provide a digital human video ringback tone generation apparatus, the apparatus comprising:
[0029] The acquisition module is used to determine the reference video based on the reference audio and acquire the reference human motion trajectory in the reference video;
[0030] An execution module is used to filter target videos from a video library based on the baseline character movement trajectory, wherein the background music of each video in the video library is the baseline audio.
[0031] Identify at least one keyword in the lyrics of the reference audio;
[0032] Based on the keywords, each video in the video library is filtered to obtain information on the changes in the movements of the characters in the video frames corresponding to the keywords.
[0033] Based on the aforementioned motion change information, the motion trajectory of the digital human is determined;
[0034] The reference audio, the target video, and the digital human are merged into a digital human video ringback tone.
[0035] Optionally, the execution module is further configured to obtain the digital human image information input by the user before determining the digital human's movement trajectory based on the movement change information;
[0036] The digital human is generated based on the digital human image information.
[0037] Optionally, the acquisition module is further configured to acquire the skeletal point coordinates of the person in the reference video, and set the center skeletal point coordinates in the skeletal point coordinates; wherein the skeletal point coordinates include at least one of the following: head skeletal point coordinates, neck skeletal point coordinates, hand skeletal point coordinates, arm skeletal point coordinates, and leg skeletal point coordinates.
[0038] Determine the distances from the coordinates of the remaining bone points (excluding the coordinates of the central bone point) to the coordinates of the central bone point, and determine the limb movements of the character based on the change information of the distances.
[0039] Based on the described body movements, the baseline character movement trajectory is generated.
[0040] Optionally, the execution module is further configured to determine the motion trajectory of the characters in each video in the video library;
[0041] The similarity of each character's movement trajectory to the baseline character's movement trajectory is compared to determine the similarity of each character's movement trajectory to the baseline character's movement trajectory.
[0042] Videos containing the character's movement trajectory with a similarity greater than or equal to a preset threshold are identified as candidate target videos;
[0043] The target video is determined from the candidate target videos according to the user's instructions.
[0044] Optionally, each video in the video library has the same video type as the baseline video, and the video type includes at least one of the following: dance type, performance type.
[0045] Optionally, the character's motion change information includes at least one of the following: the character's position change information and the speed information of the position change information. Each keyword corresponds to a set of video frames. The execution module is further configured to determine the character's position change information and the speed information of the position change information in each video frame of the set of video frames corresponding to the keyword, and to calculate the average position change information and the average speed information of the position change information based on the character's position change information and the speed information of the position change information in each video frame.
[0046] The motion trajectory of the digital human is determined based on the average position change information and the average velocity information of the position change information.
[0047] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of a method for generating a digital human video ringback tone as described in the first aspect.
[0048] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a method for generating a digital human video ringback tone as described in the first aspect.
[0049] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of a method for generating a digital human video ringback tone as described in the first aspect.
[0050] Therefore, by combining the baseline audio with the target video and through the actions of the digital human, the final digital human video ringback tone is more interesting and attractive, providing users with a more personalized and innovative calling experience; the target video is not limited to the baseline video, is more diverse, and is conducive to improving the rational use and playback of video resources in the resource library, and increasing the usage rate of low-popularity videos; the actions of the digital human are consistent with the content of the music lyrics, making it more interesting. Attached Figure Description
[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0052] Figure 1 A flowchart illustrating a method for generating digital human video ringback tones, provided as an embodiment of this application;
[0053] Figure 2 A schematic diagram of the motion lines of a person in a reference video provided for an embodiment of this application;
[0054] Figure 3 A flowchart illustrating a method for generating digital human video ringback tones, provided as an embodiment of this application;
[0055] Figure 4 A schematic diagram illustrating the display effect of a digital human video ringback tone provided in an embodiment of this application;
[0056] Figure 5 A structural block diagram of a digital human video ringback tone generation device provided in an embodiment of this application;
[0057] Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0059] Figure 1 This application illustrates a method for generating digital human video ringback tones according to an embodiment of the present application, such as... Figure 1 As shown, the method includes:
[0060] Step S101: Determine the reference video based on the reference audio, and obtain the reference character's motion trajectory in the reference video;
[0061] Step S102: Select target videos from the video library based on the baseline character movement trajectory;
[0062] The background music for each video in the video library is the base audio.
[0063] Step S103: Determine at least one keyword in the lyrics of the reference audio;
[0064] Step S104: Filter each video in the video library according to the keywords to obtain the motion change information of the characters in the video frames corresponding to the keywords;
[0065] Step S105: Determine the motion trajectory of the digital human based on the motion change information;
[0066] Step S106: Merge the base audio, target video, and digital human into a digital human video ringback tone.
[0067] In step S101, the reference audio and reference video must first be determined. The reference audio is typically selected based on specific characteristics, such as a popular song, a dialogue, or a musical clip. Selecting the reference video corresponding to the reference audio involves finding video material related to that audio, which could be a music video, movie clip, or advertising video. This pairing helps ensure that the audio and video match in content and emotion, thereby better achieving synchronization and fusion of audio and video in subsequent video processing. In the scenario shown in this application embodiment, the reference audio can be a song liked by the called party, and the corresponding reference video can be the official music video (MV) of that song.
[0068] In step S101, a reference person's motion trajectory in the reference video is also obtained. In one possible implementation, obtaining the reference person's motion trajectory in the reference video includes: obtaining the skeletal coordinates of the person in the reference video, and setting the coordinates of a central skeletal point in the skeletal coordinates; wherein the skeletal coordinates include at least one of the following: head skeletal coordinates, neck skeletal coordinates, hand skeletal coordinates, arm skeletal coordinates, and leg skeletal coordinates; determining the distances from the central skeletal coordinates to the coordinates of the remaining skeletal coordinates (excluding the central skeletal coordinates), and determining the person's limb movements based on the distance change information; and generating the reference person's motion trajectory based on the limb movements.
[0069] In specific application scenarios, the input can first be video data containing joint information of characters in a music video clip. The video data is obtained according to the object detection algorithm and output in the form of a five-dimensional matrix (N,C,T,V,M), where N is the amount of video data, that is, the frame number of the video. If it is the first frame, then N=1; C is the joint feature vector, including (x,y,acc), where x and y are the coordinates of the human joints, and acc is the coordinate confidence score; T is the number of keyframes extracted from the video; V represents the number of joints, assuming 25 joints; and M is the number of people in a video. Then, the input data is processed by BN (Batch Normalization) to facilitate subsequent data processing.
[0070] Next, based on the normalized data from the previous step, the temporal and spatial information of the video data is obtained, and the corresponding temporal and spatial edges of the trajectory are generated. GCN (Graph Convolutional Network) and TCN (Temporal Convolutional Network) are used alternately to normalize the input video joint data from the above steps and extract features, resulting in feature vectors containing temporal and spatial information. Spatial information refers to the position coordinates of the person and the coordinates of the skeleton points within a unit of time, while temporal information mainly refers to the changes in the person's position and the changes in the positions of the skeleton points along the timeline. Finally, the obtained feature vectors containing temporal and spatial information are used as the skeleton point representation for behavior recognition.
[0071] The video content is then calculated by analyzing the positional changes of individual skeletal points and the overall positional changes of all characters. Assuming the central skeletal point of the human body is the neck, the changes in the distances between different skeletal points and the central skeletal point of the same person are used to determine that person's actions. Then, the changes in the distances between different people's skeletal points and the central skeletal point of a particular person are used to determine the changes in the overall situation. This ultimately yields the actions of each person in the video. Finally, the overall video event can be obtained, and the collective action events in the video can be inferred based on the actions and positional changes of all characters. Assuming that some of the actions are dance, combining this with the positional changes of all characters in the video indicates a multi-person dance.
[0072] In one possible implementation, motion lines can be drawn based on the action. For example, temporal and spatial edges can be generated based on the time and space information obtained in the above steps. Connecting the obtained skeletal points can form multiple edges, including spatial edges that conform to the natural connections of joints. Simultaneously, combining the time information obtained in the above steps, there will also be temporal edges connecting the same joints in consecutive time steps; that is, the position information of the same joint will change at different time points. Connecting the coordinates of the same joint at different time points yields the lines of the same joint in consecutive time intervals, called temporal edges. Figure 2 As shown.
[0073] It should be noted that the generation process of temporal and spatial edges involves constructing multiple layers of spatiotemporal graph convolution based on the previous step. The generation process of temporal edges has been explained in the above steps, while spatial edges are the connections between joints within the same time frame. By lengthening the time axis and magnifying the space to the entire image, we can obtain the temporal and spatial edge information of joint changes of multiple people over a period of time. This allows information to be integrated along both spatial and temporal dimensions. The joints and spatial edges of a single human body, combined with the temporal dimension, generate temporal edges.
[0074] In step S102, target videos can be selected from the video library based on the baseline character movement trajectory, wherein the background music of each video in the video library is the baseline audio.
[0075] In one possible implementation, step S102, selecting target videos from the video library based on the baseline character movement trajectory, includes: determining the character movement trajectory in each video in the video library; comparing the similarity of each character movement trajectory with the baseline character movement trajectory to determine the similarity of each character movement trajectory relative to the baseline character movement trajectory; determining the video containing the character movement trajectory with a similarity greater than or equal to a preset threshold as a candidate target video; and determining the target video from the candidate target videos according to the user's instruction.
[0076] Each video in the video library has the same video type as the benchmark video, and the video type includes at least one of the following: dance or performance.
[0077] It should be noted that, optionally, the above filtering rules are progressive. First, videos with background music as the base audio can be filtered out. Then, videos of the same type as the base video are selected from these. Finally, videos with matching motion trajectories are selected as target videos according to the similarity matching method described above. Optionally, users can choose one of the three filtering rules or any combination of two filtering rules according to their actual needs. This application embodiment does not impose any restrictions on this.
[0078] Correspondingly, the temporal and spatial information of the candidate videos can be compared with the temporal and spatial information of the music videos obtained in the above steps using the Frechet distance. The system rules stipulate that the videos with the highest similarity ranking can be used as video clips matched to the music by the system. These clips can be randomly selected as reference resources for the subsequent fusion and innovation of digital human dance movements. The final video clips selected in this way provide new reference materials for subsequent digital human dance movements, inspire the fusion and innovation of digital human dance movements, bring users a new visual experience, and simultaneously increase the utilization rate and value of other dance videos in the resource library.
[0079] In steps S103 to S104, at least one keyword in the lyrics of the reference audio needs to be determined, and the video library is filtered according to the keyword to obtain the motion change information of the characters in the video frames corresponding to the keyword.
[0080] It should be noted that keywords can be characters, scenes, actions, emotions, etc. Assuming that the base audio is 3 minutes long and there is a keyword at the 30th second, then in the video with the base audio as background music, when the keyword is played (i.e., when the base audio is played to the 30th second), the corresponding video frame in the video is the video frame corresponding to the keyword.
[0081] In steps S105 to S106, the character's motion change information includes at least one of the following: the character's position change information and the speed information of the position change information. Each keyword corresponds to a set of video frames, such as... Figure 3 As shown, step S105, determining the digital human's motion trajectory based on motion change information, includes:
[0082] Step S301: Determine the position change information and speed information of the person in each video frame of the set of video frames corresponding to the keyword;
[0083] Step S302: Calculate the average position change information and the average velocity information of the position change information based on the position change information and velocity information of the person in each video frame.
[0084] Step S303: Determine the motion trajectory of the digital human based on the average position change information and the average velocity information of the position change information.
[0085] In specific application scenarios, all video clips of the selected music can be used, and the extracted lyrics can be used as a basis for subsequent judgments, referred to below as the "reference dataset". The lyrics content is broken down into characters, scenes, actions, and emotional words. The image frames of the characters appearing at the time points in the reference dataset mentioned above are found. This will yield multiple image frames of characters, scenes, actions, and emotional words appearing in the corresponding clips. This can be summarized in Table 1:
[0086] Table 1
[0087] Words (W) Refer to the dataset video ID Time point T Image frame F W_1 ID_1 T_1 F_1,F_2,... W_2 ID_1 T_2 F_7,F_8,... ... ... ... ... W_n ID_m T_t F_x1,F_x2,...,F_x
[0088] Table 1 shows the correspondence between keywords and keyframes. Keyframes can be screened for quality using algorithms such as image quality detection and subject detection to retain clear, high-quality image frames, which facilitates subsequent data analysis.
[0089] Based on the multiple frames of character images in the table above, several data points can be obtained: First, the positional information of the character's limbs before and after the word can be obtained from the image frame (head angle, gesture height, body range, leg movements, etc.). Combining this information, changes in positional information can also be derived, such as the character's head slowly shifting 30° from left to right in the segment before and after the word, or the leg quickly lifting and kicking out. In summary, the positional information and speed of change of key body parts corresponding to characters, scenes, actions, and emotional words can be obtained. Each part has different statistical standards and reference systems. Limb parts include, but are not limited to, the four categories listed in Table 2 below; all statistics are averages.
[0090] Table 2
[0091]
[0092]
[0093] In Table 2, the limb position information and change duration are average values. There are 7 instances of slow changes, thus yielding 7 sets of head start position information, end position information, change duration, and change position information. Assuming we calculate the head start position information S_L, where a 30° deviation to the left is recorded as -30 and a 30° deviation to the right as 30, the average value of S_L can be calculated using the following formula:
[0094]
[0095] In the calculation of the average value of the head start position information S_L, m is 7, |S_L i| represents the absolute value of the i-th head position information value. This is just an example; each part has different statistical standards and reference systems, which can be set according to actual needs. From this, we can obtain the position and speed of the changes in various parts of the human body corresponding to each word. Based on the table above, we obtain the speed distribution of each word. Combining the speed distributions of all words, we take the one with the highest speed distribution and the corresponding average limb position information and position change speed as a reference value for the rhythmic frequency of the digital human's gestures and limb movements.
[0096] Finally, the system can generate a target digital human video ringback tone by combining the user-selected music, filtered videos that match the audio, and the corresponding movements and rhythms of the digital human. The final result is a digital human video ringback tone that matches the video and music content and is displayed on the caller's mobile phone.
[0097] Additionally, it should be noted that before controlling the digital human's movement trajectory based on motion change information, the method also includes: obtaining the digital human's image information input by the user; and generating the digital human based on this image information. In other words, users can upload their personal photos to the system in advance, and the system will generate a corresponding digital human image based on the user's image. Alternatively, after matching the video content in step two, a unique digital human image can be created based on the video content, or users can set their own digital human image according to their preferences. For example, during a football match, the image could be changed to a team's jersey, or it could be changed to a digital human image from a TV series.
[0098] The method described in this application embodiment can find video content that matches the music preferred by the called party, and thereby generate digital human actions that match the rhythm of the music and the video content. Before the called party answers the call, the music, video content, and the animated digital human are displayed on the caller's mobile phone screen, realizing an end-to-end application of digital human video ringback tones. The ringback tone display effect is as follows: Figure 4 As shown.
[0099] In summary, the method described in this application calculates the similarity of time schedules and spatial edges generated by graph convolutional networks and temporal convolutional networks to find other video resources that match the rhythm of the music video. Then, it analyzes and calculates the position and speed of the characters' movements in the video resources to generate digital human movements and rhythm. By utilizing the similarity calculation of video content and the ability to generate digital human movements, digital human video ringtones can be generated, which can improve the user experience.
[0100] The method described in this application has the following technical effects: it finds other video resources that match the rhythm of the music video, making them more compatible with the music and the original music video content, giving users a new visual experience while maintaining a certain connection with the original music video, avoiding any sense of disharmony, increasing user interest and novelty, and improving user experience; it automatically generates digital human actions and rhythm changes based on the ringtone music and video content, which is richer and more harmonious than traditional music and video displays, resulting in a better user experience; the method of generating digital human video ringtones is related to the original video while introducing new video materials, and simultaneously generating digital human actions strongly related to the video content, creating a new application scenario that combines digital humans, bringing a brand-new visual experience to video ringtone users.
[0101] Figure 5 This application illustrates a digital human video ringback tone generation apparatus according to an embodiment of the present application, such as... Figure 5 As shown, the device 50 includes:
[0102] The acquisition module 501 is used to determine the reference video based on the reference audio and acquire the motion trajectory of the reference person in the reference video.
[0103] The execution module 502 is used to filter target videos from the video library based on the baseline human motion trajectory, wherein the background music of each video in the video library is the baseline audio.
[0104] Identify at least one keyword in the lyrics of the reference audio;
[0105] Based on keywords, each video in the video library is filtered to obtain information on the changes in the movements of people in the video frames corresponding to the keywords.
[0106] Determine the movement trajectory of the digital human based on the information about changes in movement;
[0107] The baseline audio, target video, and digital human are merged into a digital human video ringback tone.
[0108] In one possible implementation, the execution module 502 is further configured to obtain the digital human image information input by the user before determining the digital human's motion trajectory based on the motion change information;
[0109] A digital human is generated based on the digital human's image information.
[0110] In one possible implementation, the acquisition module 501 is further configured to acquire the skeletal point coordinates of a person in the reference video and set the center skeletal point coordinates in the skeletal point coordinates; wherein the skeletal point coordinates include at least one of the following: head skeletal point coordinates, neck skeletal point coordinates, hand skeletal point coordinates, arm skeletal point coordinates, and leg skeletal point coordinates.
[0111] In determining the coordinates of the skeleton points, the distances from the coordinates of the other skeleton points (excluding the central skeleton point) to the coordinates of the central skeleton point are used to determine the character's limb movements based on the changes in distance.
[0112] Generate a baseline character movement trajectory based on body movements.
[0113] In one possible implementation, the execution module 502 is further configured to determine the motion trajectories of characters in each video in the video library;
[0114] The similarity of each character's movement trajectory to the baseline character's movement trajectory is compared to determine the similarity of each character's movement trajectory to the baseline character's movement trajectory.
[0115] Videos containing the character's movement trajectory with a similarity greater than or equal to a preset threshold are identified as candidate target videos;
[0116] Based on the user's instructions, the target video is selected from the candidate target videos.
[0117] In one possible implementation, each video in the video library has the same video type as the base video, and the video type includes at least one of the following: dance type, performance type.
[0118] In one possible implementation, the character's motion change information includes at least one of the following: the character's position change information and the speed information of the position change information. Each keyword corresponds to a set of video frames. The execution module 502 is also used to determine the character's position change information and the speed information of the position change information in each video frame in the set of video frames corresponding to the keyword, and to calculate the average position change information and the average speed information of the position change information based on the character's position change information and the speed information of the position change information in each video frame.
[0119] The motion trajectory of the digital human is determined based on the average position change information and the average velocity information of the position change information.
[0120] In summary, it is possible to combine baseline audio with target video and, through the actions of digital humans, make the final digital human video ringback tones more interesting and attractive, providing users with a more personalized and innovative calling experience; the target video is not limited to the baseline video, is more diverse, and is conducive to improving the rational use and playback of video resources in the resource library, increasing the usage rate of low-popularity videos; the actions of digital humans are consistent with the content of music lyrics, making it more interesting.
[0121] This application also provides an electronic device 60, such as... Figure 6As shown, it includes: a processor 601, a memory 602, and a program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor, it implements the steps of the digital human video ringback tone generation method as shown in the above embodiment.
[0122] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of the digital human video ringback tone generation method shown in the above-described method embodiments, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0123] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the steps of the digital human video ringback tone generation method shown in the above method embodiments, and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0124] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0126] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for generating a digital person video ringer, characterized in that, The method includes: Based on the reference audio, a reference video is determined, and the motion trajectory of the reference character in the reference video is obtained; Target videos are selected from the video library based on the baseline character movement trajectory, wherein the background music of each video in the video library is the baseline audio. Identify at least one keyword in the lyrics of the reference audio; Based on the keywords, each video in the video library is filtered to obtain information on the changes in the movements of the characters in the video frames corresponding to the keywords. Based on the aforementioned motion change information, the motion trajectory of the digital human is determined; The reference audio, the target video, and the digital human are merged into a digital human video ringback tone.
2. The method of claim 1, wherein, Before determining the digital human's motion trajectory based on the motion change information, the method further includes: Obtain the digital avatar information input by the user; The digital human is generated based on the digital human image information.
3. The method of claim 1, wherein, Obtaining the baseline character's motion trajectory from the baseline video includes: Obtain the skeletal point coordinates of the person in the reference video, and set the center skeletal point coordinates in the skeletal point coordinates; wherein, the skeletal point coordinates include at least one of the following: head skeletal point coordinates, neck skeletal point coordinates, hand skeletal point coordinates, arm skeletal point coordinates, and leg skeletal point coordinates; Determine the distances from the coordinates of the remaining bone points (excluding the coordinates of the central bone point) to the coordinates of the central bone point, and determine the limb movements of the character based on the change information of the distances. Based on the described body movements, the baseline character movement trajectory is generated.
4. The method of claim 1, wherein, Filtering target videos from the video library based on the aforementioned baseline character movement trajectory includes: Determine the movement trajectory of the characters in each video in the video library; The similarity of each character's movement trajectory to the baseline character's movement trajectory is compared to determine the similarity of each character's movement trajectory to the baseline character's movement trajectory. Videos containing the character's movement trajectory with a similarity greater than or equal to a preset threshold are identified as candidate target videos; The target video is determined from the candidate target videos according to the user's instructions.
5. The method of claim 1, wherein, Each video in the video library has the same video type as the benchmark video, and the video type includes at least one of the following: dance type, performance type.
6. The method of claim 1, wherein, The motion change information of the character includes at least one of the following: the position change information of the character and the speed information of the position change information. Each keyword corresponds to a set of video frames. Based on the motion change information, the motion trajectory of the digital human is determined as follows: Determine the position change information of the person in each video frame and the speed information of the position change information in a set of video frames corresponding to the keyword; Based on the position change information of the person in each video frame and the velocity information of the position change information, calculate the average position change information and the average velocity information of the position change information; The motion trajectory of the digital human is determined based on the average position change information and the average velocity information of the position change information.
7. A generation device for digital human video ringback tones, characterized in that The device includes: The acquisition module is used to determine the reference video based on the reference audio and acquire the reference human motion trajectory in the reference video; An execution module is used to filter target videos from a video library based on the baseline character movement trajectory, wherein the background music of each video in the video library is the baseline audio. Identify at least one keyword in the lyrics of the reference audio; Based on the keywords, each video in the video library is filtered to obtain information on the changes in the movements of the characters in the video frames corresponding to the keywords. Based on the aforementioned motion change information, the motion trajectory of the digital human is determined; The reference audio, the target video, and the digital human are merged into a digital human video ringback tone.
8. An electronic device, comprising: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for generating a digital human video ringback tone as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the digital human video ringback tone generation method as described in any one of claims 1-6.
10. A computer program product, characterised in that, It includes computer instructions, which, when executed by a processor, implement the steps of the method for generating digital human video ringback tones as described in any one of claims 1-6.