Video display method and device, storage medium and electronic device

CN115599951BActive Publication Date: 2026-09-08GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110780926.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-09
Publication Date
2026-09-08
Estimated Expiration
2041-07-09

AI Technical Summary

Technical Problem

[0004]本发明实施例提供了一种视频显示方法和装置、存储介质及电子设备,以至少解决相关技术中为歌曲添加背景视频,或制作歌曲MV的效率较低的技术问题

Benefits of technology

[0009] In this embodiment, the process involves: obtaining a first tag for the target audio to be played; selecting a set of video frames matching the first tag from a video library; wherein each subset of video frames in the set includes an action frame sequence corresponding to a complete set of actions performed by a virtual object; selecting multiple subsets of video frames from the set and splicing them together to obtain a spliced ​​video frame subset; fusing the spliced ​​video frame subset with the target audio to obtain a fused video; playing the fused video; and playing the fused video. Based on the style type of the target video, multiple action frame sequences corresponding to a complete set of actions are selected from the video library, and the video frame subsets formed by the action frame sequences are spliced ​​together to obtain a spliced ​​video frame subset. This spliced ​​video frame subset can be the background video of a song, thereby improving the matching speed of the background video and the matching degree between the song and the background video, greatly improving the efficiency of adding background videos to songs or creating music videos, and solving the technical problem of low efficiency in adding background videos to songs or creating music videos in related technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599951B_ABST
    Figure CN115599951B_ABST
Patent Text Reader

Abstract

The application discloses a video display method and device, a storage medium and an electronic equipment. The method comprises the following steps: obtaining a first label of a target audio to be played; selecting a video frame set matched with the first label from a video library; wherein each video frame subset in the video frame set comprises a motion frame sequence corresponding to a complete set of motion combinations performed by a virtual object; selecting a plurality of video frame subsets from the video frame set for splicing to obtain a spliced video frame subset; fusing the spliced video frame subset with the target audio to obtain a fused video; and playing the fused video. The application solves the technical problem of low efficiency in adding a background video to a song or making a song MV in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computers, and more specifically, to a video display method and apparatus, a storage medium, and an electronic device. Background Technology

[0002] In related technologies, background videos are usually added to songs by manually selecting corresponding background videos or creating music videos (MVs). Existing technologies for creating full dance videos or montage MVs are slow, have limited background videos, and cannot quickly cover all songs. At the same time, the matching degree between background videos and songs is not good, which leads to low efficiency when adding background videos to songs or creating music videos.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This invention provides a video display method and apparatus, storage medium and electronic device to at least solve the technical problem of low efficiency in adding background video to songs or producing music videos in related technologies.

[0005] According to one aspect of the present invention, a video display method is provided, comprising: obtaining a first tag of a target audio to be played; selecting a set of video frames matching the first tag from a video library; wherein each subset of video frames in the set of video frames includes an action frame sequence corresponding to a complete set of actions performed by a virtual object; selecting multiple subsets of video frames from the set of video frames and splicing them together to obtain a spliced ​​video frame subset; fusing the spliced ​​video frame subset with the target audio to obtain a fused video; and playing the fused video.

[0006] According to another aspect of the present invention, a video display device is also provided, comprising: an acquisition unit for acquiring a first tag of a target audio to be played; a selection unit for selecting a set of video frames matching the first tag from a video library; wherein the video library stores multiple subsets of video frames, each subset of video frames including a sequence of action frames corresponding to a complete set of action combinations performed by a virtual object; a splicing unit for selecting multiple subsets of video frames from the set of video frames and splicing them together to obtain a spliced ​​video frame subset; a fusion unit for fusion of the spliced ​​video frame subset with the target audio to obtain a fused video; and a playback unit for playing the fused video.

[0007] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described video display method when it is run.

[0008] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the video display method described above through the computer program.

[0009] In this embodiment, the process involves: obtaining a first tag for the target audio to be played; selecting a set of video frames matching the first tag from a video library; wherein each subset of video frames in the set includes an action frame sequence corresponding to a complete set of actions performed by a virtual object; selecting multiple subsets of video frames from the set and splicing them together to obtain a spliced ​​video frame subset; fusing the spliced ​​video frame subset with the target audio to obtain a fused video; playing the fused video; and playing the fused video. Based on the style type of the target video, multiple action frame sequences corresponding to a complete set of actions are selected from the video library, and the video frame subsets formed by the action frame sequences are spliced ​​together to obtain a spliced ​​video frame subset. This spliced ​​video frame subset can be the background video of a song, thereby improving the matching speed of the background video and the matching degree between the song and the background video, greatly improving the efficiency of adding background videos to songs or creating music videos, and solving the technical problem of low efficiency in adding background videos to songs or creating music videos in related technologies. Attached Figure Description

[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0011] Figure 1 This is a schematic diagram of an application environment for an optional video display method according to an embodiment of this application;

[0012] Figure 2 This is a schematic diagram of an application environment for another optional video display method according to an embodiment of this application;

[0013] Figure 3 This is a flowchart of an optional video display method according to an embodiment of this application;

[0014] Figure 4 This is a schematic diagram of a virtual object for another optional video display method according to an embodiment of this application;

[0015] Figure 5 This is a schematic diagram of a combination of character movements according to another optional video display method based on an example of this application;

[0016] Figure 6This is a schematic diagram of a subset of video frames in an optional video display method according to an embodiment of this application;

[0017] Figure 7 This is a schematic diagram of the display interface of another optional video display method according to an embodiment of this application;

[0018] Figure 8 This is a schematic diagram of matching rules for another optional video display method according to an embodiment of this application;

[0019] Figure 9 This is a schematic diagram of the structure of an optional video display device according to an embodiment of this application;

[0020] Figure 10 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] According to one aspect of the embodiments of this application, a video display method is provided. Optionally, as an optional implementation, the video display method described above can be applied to, but is not limited to, [examples of other methods]. Figure 1The environment shown includes: a terminal device 102 for human-computer interaction, a network 104, and a server 106. The terminal device 102 may include, but is not limited to, in-vehicle electronic devices, handheld terminals, wearable devices, and portable devices. User 108 can interact with the terminal device 102, which runs a video display application client. The terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. The human-computer interaction screen 1022 displays a subset of spliced ​​video frames and a fused video. The processor 1024 obtains the first tag information of the target audio to be played; the memory 1026 stores the aforementioned subset of spliced ​​video frames and the fused video.

[0024] In addition, server 106 includes database 1062 and processing engine 1064. Database 1062 stores attribute information of the target audio, as well as a set of video frames and a fused video. Processing engine 1064 selects a set of video frames matching the first tag from the video library; selects multiple video suites from the video frame set for splicing to obtain a spliced ​​video frame subset; and fuses the spliced ​​video frame subset with the target audio to obtain a fused video.

[0025] The specific process is as follows: assuming... Figure 1 The terminal device 102 shown runs a video display application client. User 108 operates the human-computer interaction screen 1022 to manage and operate songs, such as in step S102, to obtain the attribute information of the target audio to be played. Then, step S104 is executed to send the first tag information of the target audio to the server 106 via network 104. After receiving the request, server 106 executes steps S106 to S110, selecting a set of video frames from the video library that matches the first tag; wherein the first tag is used to represent the style type of the target audio, and the video library stores multiple subsets of video frames, each subset of video frames including a sequence of action frames corresponding to a complete set of action combinations performed by a virtual object; selecting multiple subsets of video frames from the video frame set and splicing them together to obtain a spliced ​​video frame subset, wherein the sum of the playback durations of each of the multiple subsets of video frames is the same as the audio playback duration of the target audio in the first tag; fusing the spliced ​​video frame subset with the target audio to obtain a fused video; and as in step S112, notifying terminal device 102 through network 104 to return the fused video; and in step S114, playing the fused video on terminal device 102.

[0026] As another optional implementation, the video display method described above in this application can be applied to... Figure 2 The application environment shown. For example... Figure 2As shown, user 202 and user equipment 204 can interact with each other. User equipment 204 includes a memory 206 and a processor 208. In this embodiment, user equipment 204 can, but is not limited to, referencing and executing the operations performed by terminal device 102 to obtain the fused video.

[0027] Optionally, in this embodiment, the terminal device 102 and user device 204 may include, but are not limited to, at least one of the following: mobile phone (such as Android phone, iOS phone, etc.), laptop computer, tablet computer, PDA, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, etc. The target client may be a video client, instant messaging client, browser client, educational client, etc. The network 104 may include, but is not limited to, wired network and wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that enable wireless communication. The server 106 may be a single server, a server cluster composed of multiple servers, or a cloud server. The above is only an example, and no limitation is made in this embodiment.

[0028] Alternatively, as an optional implementation, such as Figure 3 As shown, the above video display method includes:

[0029] S302, Get the first tag of the target audio to be played;

[0030] S304, Select a set of video frames from the video library that matches the first tag mentioned above; wherein, each subset of video frames in the set of video frames includes a sequence of action frames corresponding to a complete set of action combinations performed by the virtual object;

[0031] S306, Select multiple subsets of video frames from the above video frame set and splice them together to obtain a spliced ​​video frame subset, wherein the sum of the playback durations of each of the multiple video frames is the same as the audio playback duration of the target audio in the first tag.

[0032] S308, the above-mentioned spliced ​​video frame subset is fused with the above-mentioned target audio to obtain a fused video;

[0033] S310, play the above-mentioned fused video.

[0034] In step S302, in practical applications, the target audio can be played through devices such as mobile phones, laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, and desktop computers. The first tag of the target audio can include, but is not limited to, the audio style type, the audio playback duration, and the audio tempo information. No limitations are made here.

[0035] In step S304, in practical applications, the first tag may include, but is not limited to, the style type of the target video, such as classical, jazz, pop, etc. Multiple subsets of video frames stored in the video library can be obtained through machine learning, such as... Figure 4 As shown, a human body model with multiple 3D positioning points is created; here, 360 3D positioning points can be selected. Then, multiple complete dance videos of different dance styles (such as...) are input and categorized. Figure 5 The dance videos shown in (a), (b), and (c) are input into a machine learning model to generate n subsets of dance movement video frames for each dance style, each ≥1s and ≤5s, structured around human body positioning points (e.g., ...). Figure 6 The video frame subsets shown represent a complete dance move (each subset contains a video frame subset). For example, if the video library stores N video frame subsets, then the number M of video frame subsets that match the target audio to be played is less than or equal to N. In other words, the video frame set corresponding to a song is a subset of the video frame set in the video library.

[0036] In step S306, in practical application, multiple subsets of the video frames are selected from the above video frame set and spliced ​​together to obtain a spliced ​​video frame subset. The sum of the playback durations of the multiple subsets of video frames is the same as the audio playback duration of the target audio in the first tag. That is, a background video (i.e., a set of video frames) with the same style type as the target audio is selected from the video library, and the background video must have the same playback duration as the target audio.

[0037] In step S306, in practical applications, such as Figure 7 As shown, the above-mentioned spliced ​​video frame subset is fused with the target audio to obtain a fused video, which is then played in the target application of the target device 702.

[0038] In this embodiment, the process involves: obtaining a first tag for the target audio to be played; selecting a set of video frames matching the first tag from a video library; wherein the video library stores multiple subsets of video frames, each subset including a sequence of action frames corresponding to a complete set of actions performed by a virtual object; selecting multiple subsets of video frames from the video frame set and splicing them together to obtain a spliced ​​video frame subset; fusing the spliced ​​video frame subset with the target audio to obtain a fused video; playing the fused video; and playing the fused video. Based on the style type of the target video, multiple sequences of action frames corresponding to a complete set of actions are selected from the video library, and the video frame subsets formed by these action frame sequences are spliced ​​together to obtain a spliced ​​video frame subset. This spliced ​​video frame subset can be the background video of a song, thereby improving the matching speed of the background video and the matching degree between the song and the background video, greatly improving the efficiency of adding background videos to songs or creating music videos, and solving the technical problem of low efficiency in adding background videos to songs or creating music videos in related technologies.

[0039] In one or more embodiments, step S304, selecting a set of video frames that match the first tag from the video library, includes: selecting a subset of video frames that are the same as the first style type indicated by the first tag from the video library to obtain the set of video frames.

[0040] In this embodiment of the application, the video library stores a subset of video frames of multiple style types, such as... Figure 6 As shown, each video frame subset contains a complete dance move. For example, if the target audio is a Japanese song, then the style of the target audio can be Japanese pop. Therefore, multiple video frame subsets with the style of Japanese otaku dance can be selected from the video library to obtain the above video frame set.

[0041] Through one or more embodiments provided in this application, by selecting a subset of video frames from the video library that are the same as the first style type indicated by the first tag, the above-mentioned video frame set can be obtained, which can accurately match the background video of the currently playing audio and improve the matching efficiency of the background video.

[0042] In one or more embodiments, step S306, selecting multiple subsets of video frames from the above-mentioned video frame set for splicing to obtain a spliced ​​video frame subset, includes:

[0043] Multiple subsets of the video frames are selected from the above video frame set and spliced ​​together to obtain a candidate spliced ​​video frame subset; wherein the playback duration of the candidate spliced ​​video frame subset is greater than or equal to the playback duration of the target audio; the audio beat interval duration of the target audio is determined; and the spliced ​​video frame subset is obtained based on the playback duration of the candidate spliced ​​video frame subset and the audio beat interval duration.

[0044] In one or more embodiments, obtaining the spliced ​​video frame subset based on the playback duration of the candidate spliced ​​video frame subset and the aforementioned audio beat interval duration includes:

[0045] Starting from the first beat of the target audio and ending beat, take each beat of the target audio as the current beat and perform the following adjustment steps to obtain the above-mentioned subset of spliced ​​video frames:

[0046] Compare the playback duration of the candidate video frame subset corresponding to the time period of the current beat audio with the duration of the audio beat interval; wherein, the candidate video frame subset is one or more video frame subsets in the candidate spliced ​​video frame subset;

[0047] When the playback duration of the candidate video frame subset corresponding to the current beat audio time period is equal to the audio beat interval duration, the playback speed of the candidate video frame subset remains unchanged.

[0048] When the playback duration of the candidate video frame subset corresponding to the current time period of the audio beat is greater than or less than the audio beat interval duration, the playback speed of the candidate video frame subset is adjusted so that the playback duration of the candidate video frame subset is the same as the audio beat interval duration.

[0049] In the embodiments of this application, such as Figure 8 As shown, the target video is an audio file with a duration of 60 seconds, a beat per minute (BPM) of 20, and a beat length of 3 seconds. The target video consists of three subsets of video frames: 802 (1 second), 804 (2 seconds), 806 (3 seconds), 808 (4 seconds), and 810 (5 seconds). Multiple video frames from these subsets are selected and spliced ​​together from the first beat to the last beat of the target audio to obtain a candidate spliced ​​video frame subset. Figure 8 (The rectangle in the image shows the time frame). The playback duration of the above candidate spliced ​​video frame subset is equal to the playback duration of the above target audio, both being 60 seconds.

[0050] When the playback duration of the candidate video frame subset (video frame subset 810) corresponding to the time period of the first beat audio is equal to the duration of the audio beat interval (both are 3s), the playback speed of the video frame subset 810 remains unchanged.

[0051] The third beat of the target audio is filled with a subset of video frames 810 with a playback duration of 5 seconds, but the audio beat interval is 3 seconds. In this case, video frame subset 810 can be played at 5 / 3 times the normal playback speed, meaning the playback speed is increased to make the playback duration of video frame subset 810 3 seconds. The sixteenth beat of the target audio is filled with a subset of video frames 804 with a playback duration of 2 seconds, but the audio beat interval is 3 seconds. In this case, video frame subset 804 can be played at 2 / 3 of the normal playback speed, meaning the playback speed is slowed down to make the playback duration of video frame subset 810 3 seconds.

[0052] By using one or more embodiments provided in this application, and by taking each beat of the target audio as the current beat audio sequentially from the first beat audio to the last beat audio, and adjusting the playback time of the video suite to be consistent with each beat audio of the target audio, a background video with consistent playback beat and time can be obtained, thereby improving the matching efficiency of the background video.

[0053] In one or more embodiments, the above-mentioned selection of multiple subsets of video frames from the above-mentioned video frame set for splicing to obtain a spliced ​​video frame subset includes:

[0054] The audio rhythms of multiple audio segments contained in the target audio are determined; multiple subsets of video frames matching the audio rhythms are obtained from the set of video frames; wherein audio segments with different music rhythms correspond to different subsets of video frames; the multiple subsets of video frames are spliced ​​together to obtain a spliced ​​subset of video frames.

[0055] In this embodiment of the invention, an audio file can be divided into multiple segments, each segment corresponding to a musical rhythm. Each musical rhythm contains a different number of notes, and therefore each musical rhythm can be divided into multiple levels of fast and slow tempos. The tempo of the musical rhythm can be obtained by: determining the position of each beat point in the audio to obtain each beat point, and inferring the tempo of each audio segment based on each beat point.

[0056] For example, the target audio includes four segments arranged in chronological order: G1, G2, G3, and G4. Each segment has a different audio rhythm, such as... Figure 8As shown, the target video is an audio file with a duration of 60 seconds, a BPM of 20, and a beat rate of 3 seconds per beat. Assuming the duration of segment G1 is 10 seconds, segment G2 is 15 seconds, segment G3 is 15 seconds, and segment G4 is 20 seconds, a subset of video frames can be selected within the first 10 seconds of the target audio, and another subset can be selected between the 10th and 25th seconds, until all segments of the target audio match a specific subset of video frames. Adjacent segments will have different subsets of video frames. Finally, these multiple matched subsets of video frames can be concatenated to obtain a concatenated video frame subset.

[0057] Through one or more embodiments provided in this application, by obtaining the rhythm of the target audio, it can be ensured that different rhythms match different subsets of video frames, allowing for flexible and diverse matching of the target audio with different background videos. In one or more embodiments, when the duration of the candidate video frame subset corresponding to the time period of the current beat audio is less than the duration of the audio beat interval, the video display method further includes: determining a plurality of video frame subsets whose playback durations are equal to the duration of the audio beat interval as a filling video frame subset set; and replacing the video frame subset corresponding to the time period of the current beat audio with the filling video frame subset set.

[0058] In the embodiments of this application, such as Figure 8 As shown, with an audio duration of 60s, BPM 20, and 3s / beat as the target video: a video frame subset 802 with a playback duration of 1s, a video frame subset 804 with a playback duration of 2s, a video frame subset 806 with a playback duration of 3s, a video frame subset 808 with a playback duration of 4s, and a video frame subset 810 with a playback duration of 5s are selected from the above video frame sets and spliced ​​together to obtain a candidate spliced ​​video frame subset. Figure 8 (A rectangle showing the time is displayed in the image). The playback duration of the above candidate spliced ​​video frame subsets is equal to the playback duration of the above target audio, both being 60s. The second beat of the target audio is filled with video frame subsets 802 with a playback duration of 1s and 804 with a playback duration of 2s as filling video frame subsets. The playback duration of video frame subsets 802 and 804 is 3s. Therefore, the filling video frame subsets including video frame subsets 802 and 804 can replace the video frame subsets in the second beat of the audio that have a playback duration of less than 3s.

[0059] Through one or more embodiments provided in this application, by determining a subset of video frames whose playback durations are equal to the duration of the audio beat interval, a set of filler video frames is obtained; by replacing the subset of video frames corresponding to the time period of the current beat audio with the set of filler video frames, the playback time of each beat audio of the target audio can be adjusted to be consistent, resulting in a background video with consistent playback beat and time, thereby improving the matching efficiency of the background video.

[0060] In one or more embodiments, the video display method further includes: the subsets of video frames corresponding to adjacent beat audios of the target audio are different.

[0061] In the embodiments of this application, such as Figure 8 As shown, the first beat audio corresponds to a video frame subset 802 with a playback duration of 3 seconds. The second beat audio corresponds to a filler playback video condition with a playback duration of 3 seconds, which includes video frame subsets 802 and 804. The third beat audio corresponds to a video frame subset 810 with a playback duration of 5 seconds. Here, the virtual objects in the video frame subsets can be the same or different. That is, virtual objects in different video frame subsets, such as cartoon characters, can be the same or different cartoon characters.

[0062] By using one or more embodiments provided in this application, and by setting different subsets of video frames corresponding to adjacent beat audio of the target audio, it is possible to ensure that the playback images or playback actions in the background video are different, thereby improving the flexibility of background video playback.

[0063] In one or more embodiments, after fusing the spliced ​​video frame subset with the target audio to obtain a fused video, the method further includes: obtaining multiple background videos corresponding to the target audio; wherein the background videos include the fused video; and determining the playback order of the background videos in the target application according to the playback priority of the multiple background videos.

[0064] In the embodiments of this application, a song can correspond to different background videos. The playback order of the background videos in the target application can be determined based on their playback priority. The playback priority can be set in the background or determined based on the click rate or the frequency of complete playback of the background video. Through the above technical means, the target audio can be more accurately configured with background videos that match the user's interests, thereby improving the matching degree of the background videos.

[0065] In one or more embodiments, determining the background video to be displayed in the current application based on the playback priority of the plurality of background videos includes:

[0066] S1, determine the playback priority of the background videos based on the frequency of each of the background videos being played completely; wherein the frequency is positively correlated with the playback priority of the background videos;

[0067] S2, determine the playback order of the background videos in the target application based on the playback priority of the multiple background videos.

[0068] In the embodiments of this application, a song can correspond to different background videos. The order in which the background videos are played in the target application can be determined according to the playback priority of the different background videos. The playback priority can be set in the background or determined according to the click rate or the frequency of complete playback of the background video. For example, background music with a higher frequency of complete playback can be played first.

[0069] In one or more embodiments, within a subset of video frames of different style types: the virtual objects are different; or, the display resources corresponding to the virtual objects are different.

[0070] In the embodiments of this application, for example, when the style type is classical music, the virtual object in the video suite can be a ballet dance animation; or, the skin of the virtual object is displayed as a ballet style; when the style type is traditional Chinese classical music, the virtual object in the video suite can be a traditional Chinese classical costume character; or, the skin of the virtual object is displayed as a Hanfu or cheongsam style.

[0071] Through one or more embodiments provided in this application, by setting different virtual objects in video suites with different style types; or by setting different display resources corresponding to the virtual objects, the background video of the target audio can be enriched, the audio playback viewing experience can be improved, and the matching degree of the background video can be enhanced.

[0072] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0073] The production speed of full dance tracks (background videos) / mixed MVs in related technologies is slow and concentrated, making it impossible to quickly cover all suitable songs; at the same time, the appearance and skill level of the dancers in the full dance tracks vary, resulting in poor viewing experience and matching.

[0074] To address the aforementioned technical problems, this application provides a video display method including the following: Here, the video suite can be the video frame set described in the above embodiments.

[0075] S1, as Figure 4 As shown, a human body model with multiple 3D positioning points is established; here, 360 3D positioning points can be selected.

[0076] S2 allows you to import multiple complete dance videos of different dance styles (e.g., ...). Figure 5 The dance videos shown in (a), (b), and (c) are input into a machine learning model to generate n dance movement video kits for each dance style, each ≥1s and ≤5s, with human body positioning points as the framework (e.g., Figure 6 The video suite shown (each video suite contains one complete dance move).

[0077] S3. Calculate the duration of the audio beat interval based on the number of beats per minute (BPM), and randomly fill in the dance move video kit. The matching logic is as follows:

[0078] 1) Calculate the audio beat interval duration = 60 ÷ BPM value (i.e., the number of seconds corresponding to one beat of the audio).

[0079] 2) Matching scheme: Randomly splice n dance move video kits to generate a long video with a duration equal to or greater than the audio duration;

[0080] A. If the duration of the video kit equals the duration of the audio beat interval, directly fill in a video kit with the same duration as that segment;

[0081] B. If the duration of the video suite is less than the duration of the audio beat interval, then randomly fill in class a / b;

[0082] a. Fill in a slow-motion version of the video;

[0083] b. Fill in multiple videos shorter than the duration of the segment. If the last video in the segment times out, fill in a speed-up version of that video.

[0084] C. If the duration of the video suite is greater than the duration of the audio beat interval, fill in the video with a speed-up version.

[0085] D. When the video duration is greater than or equal to the audio duration, stop splicing and cut off the video timeout portion based on the audio duration.

[0086] E. Regardless of whether the video lengths are the same, two adjacent dance move video sets are different.

[0087] Where: Speed / Slow Speed ​​Calculation Formula = Video Suite Duration ÷ Audio Beat Interval Duration;

[0088] In one embodiment, such as Figure 8As shown, with an audio duration of 60s, BPM 20, and 3s / beat as the target video: a video frame subset 802 with a playback duration of 1s, a video frame subset 804 with a playback duration of 2s, a video frame subset 806 with a playback duration of 3s, a video frame subset 808 with a playback duration of 4s, and a video frame subset 810 with a playback duration of 5s are selected from the above video frame sets and spliced ​​together to obtain a candidate spliced ​​video frame subset. Figure 8 The rectangles displaying time indicate that the playback duration of the candidate video frame subsets is equal to the playback duration of the target audio, both being 60 seconds. The third beat of the target audio is filled with a video frame subset 810 with a playback duration of 5 seconds, but the audio beat interval is 3 seconds. Therefore, video frame subset 810 can be played at 5 / 3 times the normal playback speed, meaning the playback speed is increased to make the playback duration of video frame subset 810 3 seconds. The sixteenth beat of the target audio is filled with a video frame subset 804 with a playback duration of 2 seconds, but the audio beat interval is 3 seconds. Therefore, video frame subset 804 can be played at 2 / 3 of the normal playback speed, meaning the playback speed is slowed down to make the playback duration of video frame subset 810 3 seconds.

[0089] S4. Based on song tags, categorize the dance styles that suit the songs, and design multiple looks for the virtual anime character for each dance style, for example:

[0090] Song Tags — Suitable Dance Styles — Character Design

[0091] 1) Popular Japanese culture – Japanese otaku dance – A. JK B. Lolita;

[0092] 2) Popular Korean culture – K-pop dance – A. Girl Crush B. Campus style C. Mature style;

[0093] 3) Classical, instrumental music – ballet – ballet styling;

[0094] 4) Jazz – European and American jazz dance – European and American modern fitness styling;

[0095] 5) Traditional Chinese Style – Classical Chinese Dance – A. Hanfu (traditional Han clothing) B. Qipao (cheongsam);

[0096] 6) Latin Dance – Latin Dance Styling;

[0097] 7) Popular dance music – modern dance, street dance – universally modern styling;

[0098] 8) Electronics – Modern dance, street dance – Y2K styling.

[0099] In one embodiment, the playback page of the merged video is sorted either by prioritizing the selection of merged videos from the video library for playback, or by sorting the background videos from high to low according to their completion rate (full playback rate).

[0100] In this embodiment, the following steps are taken: 1) Obtain the attribute information of the target audio to be played; 2) Select a set of video frames from a video library that matches the first tag in the attribute information; wherein the first tag represents the style type of the target audio, and the video library stores multiple video suites, each including a sequence of action frames corresponding to a complete set of actions performed by a virtual object; 3) Select multiple video suites from the video frame set and splice them together to obtain a spliced ​​video frame subset, wherein the sum of the playback durations of each of the multiple video suites is the same as the audio playback duration of the target audio in the attribute information; 4) Fuse the spliced ​​video frame subset with the target audio to obtain a fused video; 5) Play the fused video. Based on the style type of the target video, multiple action frame sequences corresponding to a complete set of actions are selected from the video library. The video suite composed of the above action frame sequences is then spliced ​​to obtain a spliced ​​video frame subset. The spliced ​​video frame subset can be the background video of a song, thereby improving the matching speed of the background video of the song and the matching degree between the song and the background video. This greatly improves the efficiency of adding background videos to songs or creating music videos, and solves the technical problem of low efficiency in adding background videos to songs or creating music videos in related technologies.

[0101] According to another aspect of the present invention, a video display apparatus for implementing the above-described video display method is also provided. For example... Figure 9 As shown, the device includes:

[0102] Get unit 902, get the first tag of the target audio to be played;

[0103] Selection unit 904 is used to select a set of video frames that match the first tag from the video library; wherein the video library stores multiple subsets of video frames, and each subset of video frames includes a sequence of action frames corresponding to a complete set of action combinations performed by a virtual object;

[0104] The splicing unit 906 is used to select multiple subsets of the video frames from the above video frame set and splice them to obtain a spliced ​​video frame subset.

[0105] The fusion unit 908 is used to fuse the above-mentioned spliced ​​video frame subset with the above-mentioned target audio to obtain a fused video;

[0106] Playback unit 910 is used to play the aforementioned fused video. In one or more embodiments of this application, the current target audio can be played through devices such as mobile phones, laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, and desktop computers. The attribute information of the target audio may include, but is not limited to, the audio style type, the audio playback duration, and the audio tempo information, etc., without any limitation.

[0107] In one or more embodiments of this application, the first tag may include, but is not limited to, the style type of the target video, such as classical, jazz, pop, etc. Multiple subsets of video frames stored in the video library can be obtained through machine learning, such as... Figure 4 As shown, a human body model with multiple 3D positioning points is created; here, 360 3D positioning points can be selected. Then, multiple complete dance videos of different dance styles (such as...) are input and categorized. Figure 5 The dance videos shown in (a), (b), and (c) are input into a machine learning model to generate n subsets of dance movement video frames for each dance style, each ≥1s and ≤5s, structured around human body positioning points (e.g., ...). Figure 6 The video frame subset shown (each video frame subset contains a complete dance move).

[0108] In one or more embodiments of this application, multiple subsets of video frames are selected from the above-mentioned video frame set and spliced ​​together to obtain a spliced ​​video frame subset. The sum of the playback durations of each of the multiple subsets of video frames is the same as the audio playback duration of the target audio in the above-mentioned attribute information. That is, a background video (i.e., a set of video frames) with the same style type as the target audio is selected from the above-mentioned video library, and the background video must have the same playback duration as the target audio.

[0109] In one or more embodiments of this application, such as Figure 7 As shown, the above-mentioned spliced ​​video frame subset is fused with the target audio to obtain a fused video, which is then played in the target application of the target device 702.

[0110] In this embodiment, the following steps are taken: 1) Obtain the attribute information of the target audio to be played; 2) Select a set of video frames from a video library that matches the first tag in the attribute information; wherein the first tag represents the style type of the target audio, and the video library stores multiple video suites, each including a sequence of action frames corresponding to a complete set of actions performed by a virtual object; 3) Select multiple video suites from the video frame set and splice them together to obtain a spliced ​​video frame subset, wherein the sum of the playback durations of each of the multiple video suites is the same as the audio playback duration of the target audio in the attribute information; 4) Fuse the spliced ​​video frame subset with the target audio to obtain a fused video; 5) Play the fused video. Based on the style type of the target video, multiple action frame sequences corresponding to a complete set of actions are selected from the video library. The video suite composed of the above action frame sequences is then spliced ​​to obtain a spliced ​​video frame subset. The spliced ​​video frame subset can be the background video of a song, thereby improving the matching speed of the background video of the song and the matching degree between the song and the background video. This greatly improves the efficiency of adding background videos to songs or creating music videos, and solves the technical problem of low efficiency in adding background videos to songs or creating music videos in related technologies.

[0111] According to another aspect of the present invention, an electronic device for implementing the above-described video display method is also provided. This electronic device may be... Figure 1 The terminal device or server shown. This embodiment uses the electronic device as a server as an example for illustration. Figure 10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps of any of the above method embodiments via the computer program.

[0112] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0113] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0114] S1, Get the first tag of the target audio to be played;

[0115] S2, Select a set of video frames that match the first tag from the video library; wherein, the video library stores multiple subsets of video frames, and each subset of video frames includes a sequence of action frames corresponding to a complete set of action combinations performed by a virtual object;

[0116] S3, Select multiple subsets of video frames from the set of video frames and splice them together to obtain a spliced ​​subset of video frames;

[0117] S4, the spliced ​​video frame subset is fused with the target audio to obtain a fused video;

[0118] S5, play the fused video.

[0119] Alternatively, as those skilled in the art will understand, Figure 10 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 10 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 10 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 10 The different configurations shown.

[0120] The memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video display method and apparatus in this embodiment of the invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, thereby realizing the aforementioned video display method. The memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include memory remotely located relative to the processor 1004, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1002 may be used, but is not limited to, for displaying data information such as fused video. As an example, such as... Figure 10 As shown, the memory 1002 may include, but is not limited to, the acquisition unit 902, selection unit 904, splicing unit 906, fusion unit 908, and playback unit 910 in the video display device. Furthermore, it may include, but is not limited to, other module units in the video display device, which will not be described in detail in this example.

[0121] Optionally, the transmission device 1006 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1006 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1006 is a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0122] In addition, the aforementioned electronic device also includes: a display 1008 for displaying the aforementioned fused video; and a connection bus 1010 for connecting the various module components in the aforementioned electronic device.

[0123] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0124] In one or more embodiments, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video display method described above. The computer program is configured to execute the steps of any of the method embodiments described above when running.

[0125] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:

[0126] S1, Get the first tag of the target audio to be played;

[0127] S2, Select a set of video frames that match the first tag from the video library; wherein, the video library stores multiple subsets of video frames, and each subset of video frames includes a sequence of action frames corresponding to a complete set of action combinations performed by a virtual object;

[0128] S3, Select multiple subsets of video frames from the set of video frames and splice them together to obtain a spliced ​​subset of video frames;

[0129] S4, the spliced ​​video frame subset is fused with the target audio to obtain a fused video;

[0130] S5, play the fused video.

[0131] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. This program can be stored in a computer-readable storage medium, which may include: a flash drive, read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. The sequence numbers of the above embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0132] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0133] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0134] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0136] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0137] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A video display method, characterized in that, include: Multiple styles of dance videos are input into a machine learning model to generate a subset of video frames corresponding to each of the multiple styles of dance videos. The human body model displayed in the subset of video frames includes multiple human body positioning points. Get the first tag of the target audio to be played; Select a set of video frames that match the first tag from the video library, and determine the target position of multiple human body positioning points in each video frame based on the action frame sequence corresponding to the action combination of virtual objects in each subset of the video frames in the video frame set. Selecting multiple subsets of video frames from the video frame set and splicing them together to obtain a spliced ​​video frame subset includes: selecting multiple subsets of video frames from the video frame set and splicing them together to obtain a candidate spliced ​​video frame subset; wherein the playback duration of the candidate spliced ​​video frame subset is greater than or equal to the playback duration of the target audio; determining the audio beat interval duration of the target audio; from the first beat audio to the last beat audio of the target audio, sequentially taking each beat audio of the target audio as the current beat audio, and performing the following adjustment steps to obtain the spliced ​​video frame subset: comparing the time period corresponding to the current beat audio. The playback duration of the candidate video frame subset is equal to the audio beat interval duration; wherein, the candidate video frame subset is one or more video frame subsets in the candidate spliced ​​video frame subset; when the playback duration of the candidate video frame subset corresponding to the time period of the current beat audio is equal to the audio beat interval duration, the playback speed of the candidate video frame subset remains unchanged; when the playback duration of the candidate video frame subset corresponding to the time period of the current beat audio is greater than or less than the audio beat interval duration, the playback speed of the candidate video frame subset is adjusted until the playback duration of the candidate video frame subset is the same as the audio beat interval duration; Determine the virtual skin of the virtual object that matches the first tag; The spliced ​​video frame subset is fused with the target audio to obtain a fused video, wherein in each video frame of the fused video, the human body positioning point on the human body model of the virtual object wearing the virtual skin is located at the corresponding target position; Play the fused video.

2. The method according to claim 1, characterized in that, The step of selecting a set of video frames from the video library that match the first tag includes: The video frame set is obtained by selecting a subset of video frames from the video library that are the same as the first style type indicated by the first tag.

3. The method according to claim 1, characterized in that, The step of selecting multiple subsets of video frames from the set of video frames and splicing them together to obtain a spliced ​​subset of video frames includes: Determine the audio rhythms of multiple audio segments contained in the target audio; Obtain multiple subsets of video frames from the set of video frames that match the audio rhythm; wherein audio segments with different music rhythms correspond to different subsets of video frames; Multiple subsets of video frames are spliced ​​together to obtain a spliced ​​subset of video frames.

4. The method according to claim 1, characterized in that, When the duration of the candidate spliced ​​video frame subset corresponding to the time period of the current beat audio is less than the duration of the audio beat interval, the method further includes: A subset of video frames whose playback durations sum to the duration of the audio beat interval is determined as the filler video frame set; Replace the subset of video frames corresponding to the time period in which the current beat audio is located with the set of filling video frames.

5. The method according to claim 1, characterized in that, The method further includes: the video frame subsets corresponding to adjacent beat audios of the target audio are different.

6. The method according to claim 1, characterized in that, The step of fusing the spliced ​​video frame subset with the target audio to obtain the fused video further includes: Obtain multiple background videos corresponding to the target audio; wherein, the multiple background videos include the fused video; The playback order of the background videos in the target application is determined based on the playback priority of the multiple background videos.

7. The method according to claim 6, characterized in that, The step of determining the background video to be displayed in the target application based on the playback priority of the multiple background videos includes: The playback priority of the background videos is determined based on the frequency with which the multiple background videos are played completely; wherein the frequency is positively correlated with the playback priority of the background videos. The playback order of the background videos in the target application is determined based on the playback priority of the multiple background videos.

8. The method according to any one of claims 1 to 7, characterized in that, In video frame subsets with different style types: the virtual objects are different; or, the display resources corresponding to the virtual objects are different.

9. A video display device, characterized in that, include: The acquisition unit inputs dance videos of various styles into a machine learning model to generate a subset of video frames corresponding to each of the dance videos of various styles, wherein the human body model displayed in the subset of video frames includes multiple human body positioning points; and acquires the first tag of the target audio to be played. The selection unit is used to select a set of video frames that match the first tag from the video library, and determine the target position of multiple human body positioning points in each video frame according to the action frame sequence corresponding to the action combination of virtual objects in each subset of the video frames in the set of video frames. The splicing unit is configured to select multiple subsets of video frames from the video frame set for splicing to obtain a spliced ​​video frame subset, including: selecting multiple subsets of video frames from the video frame set for splicing to obtain a candidate spliced ​​video frame subset; wherein the playback duration of the candidate spliced ​​video frame subset is greater than or equal to the playback duration of the target audio; determining the audio beat interval duration of the target audio; and, from the first beat audio to the last beat audio of the target audio, sequentially taking each beat audio of the target audio as the current beat audio, performing the following adjustment steps to obtain the spliced ​​video frame subset: comparing the time of the current beat audio... The playback duration of the candidate video frame subset corresponding to the segment is equal to the audio beat interval duration; wherein, the candidate video frame subset is one or more video frame subsets in the candidate spliced ​​video frame subset; when the playback duration of the candidate video frame subset corresponding to the time segment of the current beat audio is equal to the audio beat interval duration, the playback speed of the candidate video frame subset remains unchanged; when the playback duration of the candidate video frame subset corresponding to the time segment of the current beat audio is greater than or less than the audio beat interval duration, the playback speed of the candidate video frame subset is adjusted until the playback duration of the candidate video frame subset is the same as the audio beat interval duration; A fusion unit is used to determine the virtual skin of the virtual object that matches the first tag; The spliced ​​video frame subset is fused with the target audio to obtain a fused video, wherein in each video frame of the fused video, the human body positioning point on the human body model of the virtual object wearing the virtual skin is located at the corresponding target position; The playback unit is used to play the fused video.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method according to any one of claims 1 to 8.

11. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 8 through the computer program.

Citation Information

Patent Citations

  • Music short video generation method and device, electronic equipment and storage medium

    CN111935537A