Multimedia playing method and device, electronic equipment and computer readable medium

By identifying separators and matching mute video keyframes in multimedia playback, the resource waste problem caused by digital human voice-over insertion is solved, and efficient multimedia playback is achieved.

CN120264074APending Publication Date: 2025-07-04JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503531.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing voice-over audio insertion method causes digital people to stay still or make audio into video, wasting server resources, and inefficient multimedia playback.

Method used

By identifying the separator in the multimedia data, determine the segmented multimedia data, and obtain the characteristics of the last frame at the end of the currently played audio video, match the keyframe of the mute video, play the mute video from the keyframe, and play the voice-over audio at the same time until the voice-over audio is played, and continue to play the next segmented multimedia data.

Benefits of technology

It saves server resources, improves the efficiency and continuity of multimedia playback, and reduces resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264074A_ABST
    Figure CN120264074A_ABST
Patent Text Reader

Abstract

The invention discloses a multimedia playing method and device, electronic equipment and a computer readable medium, and relates to the technical field of computers, and the method comprises the steps: obtaining multimedia data according to a multimedia playing request, and recognizing separators to determine each piece of segmented multimedia data in the multimedia data; and if the currently played segment multimedia data is an audio video and the next segment multimedia data is an out-of-picture audio, when the currently played segment multimedia data is completely played, playing the next segment multimedia data, and if the currently played segment multimedia data is an audio video and the next segment multimedia data is an out-of-picture audio. Obtaining the feature of the last frame of the currently played segmented multimedia data, and matching the feature with the feature of each frame in the corresponding mute video to obtain a key frame; and starting from the key frame, playing the mute video, playing the corresponding out-of-picture audio at the same time, and continuing to play the next segment of multimedia data in response to the completion of the playing of the corresponding out-of-picture audio. Server resources are saved, and multimedia playing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a multimedia playback method, apparatus, electronic device, and computer-readable medium. Background Art

[0002] Currently, during the playback of a digital human script, it is often necessary to insert and play off-screen voice audio that is not the host's voice for warming up the scene. When inserting the off-screen voice audio, the digital human needs to remain silent. Currently, the implementation method of inserting off-screen voice audio is either that the digital human remains static, or the audio is also made into a video to ensure that the digital human is dynamic, which wastes server resources and results in low multimedia playback efficiency. Summary of the Invention

[0003] In view of this, embodiments of this application provide a multimedia playback method, apparatus, electronic device, and computer-readable medium, which can solve the problem that the existing implementation methods of inserting off-screen voice audio are either that the digital human remains static, or the audio is also made into a video to ensure that the digital human is dynamic, wasting server resources and resulting in low multimedia playback efficiency.

[0004] To achieve the above object, according to one aspect of the embodiments of this application, a multimedia playback method is provided, including:

[0005] Obtain multimedia data according to a received multimedia playback request, and identify delimiters in the multimedia data;

[0006] Based on the delimiters, determine each segmented multimedia data in the multimedia data;

[0007] Play each segmented multimedia data according to the time sequence. If the currently played segmented multimedia data is an audio-visual video and the next segmented multimedia data is off-screen voice audio, when the currently played segmented multimedia data is played to completion, obtain the feature of the last frame of the currently played segmented multimedia data;

[0008] Obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the features of each frame in the silent video to obtain the key frame that is matched;

[0009] Start playing the silent video from the key frame, and at the same time play the corresponding off-screen voice audio. In response to the completion of the playback of the corresponding off-screen voice audio, continue to play the next segmented multimedia data.

[0010] Optionally, the method further includes:

[0011] If the corresponding off-screen voice audio has not been played to completion when the silent video is played to completion, obtain the time point at which the target state of the virtual image in the silent video is recorded;

[0012] Loop and play the silent video from a time point until the corresponding voice-over audio finishes playing, and then stop looping and playing the silent video.

[0013] Optionally, loop and play the corresponding silent video from a time point, including:

[0014] Intercept the video of the silent video from the time point to the end of the silent video as the target silent video;

[0015] Loop and play the target silent video.

[0016] Optionally, match the features of the last frame with the features of each frame in the silent video to obtain the matched key frames, including:

[0017] Perform cosine similarity matching on the mouth state features of the last frame and the mouth state features of each frame in the silent video to obtain respective cosine similarities;

[0018] Determine the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame.

[0019] Optionally, identify the delimiters in the multimedia data, including:

[0020] Identify the text length delimiter and the voice-over delimiter in the multimedia data.

[0021] Optionally, continue to play the next segmented multimedia data, including:

[0022] If the next segmented multimedia data is an audible video, play the next segmented multimedia data and display the corresponding synthesized text.

[0023] Optionally, the method further includes:

[0024] The virtual avatar is in a breathing state when playing the silent video.

[0025] In addition, the present application also provides a multimedia playback device, including:

[0026] A delimiter recognition unit configured to obtain multimedia data according to a received multimedia playback request and identify the delimiters in the multimedia data;

[0027] A segmented multimedia data determination unit configured to determine each segmented multimedia data in the multimedia data based on the delimiters;

[0028] A feature acquisition unit, configured to play each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-visual video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data finishes playing, acquire the features of the last frame of the currently played segmented multimedia data.

[0029] The key frame determination unit is further configured to acquire the silent video corresponding to the virtual image corresponding to the features of the last frame, and match the features of the last frame with the features of each frame in the silent video to obtain the matched key frame.

[0030] A multimedia playback unit, configured to start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data.

[0031] Optionally, the multimedia playback device further includes a loop playback unit, configured to:

[0032] If the corresponding voice-over audio has not finished playing when the silent video finishes playing, acquire the time point where the target state of the virtual image in the silent video is recorded;

[0033] Loop play the silent video starting from the time point until the corresponding voice-over audio finishes playing, and stop loop playing the silent video.

[0034] Optionally, the loop playback unit is further configured to:

[0035] Intercept the video of the silent video from the time point to the end of the silent video as the target silent video;

[0036] Loop play the target silent video.

[0037] Optionally, the key frame determination unit is further configured to:

[0038] Perform cosine similarity matching between the mouth state features of the last frame and the mouth state features of each frame in the silent video to obtain each cosine similarity;

[0039] Determine the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame.

[0040] Optionally, the delimiter recognition unit is further configured to:

[0041] Identify the text length delimiter and the voice-over delimiter in the multimedia data.

[0042] Optionally, the multimedia playback unit is further configured to:

[0043] If the next segmented multimedia data is an audio-visual video, play the next segmented multimedia data and display the corresponding synthesized text.

[0044] Optionally, the multimedia playback unit is further configured to:

[0045] The virtual avatar is in a breathing state when playing a silent video.

[0046] In addition, the present application also provides a multimedia playback electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the multimedia playback method as described above.

[0047] In addition, the present application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the multimedia playback method as described above is implemented.

[0048] To achieve the above object, according to another aspect of the embodiments of the present application, a computer program product is provided.

[0049] A computer program product of an embodiment of the present application includes a computer program, and when the program is executed by a processor, the multimedia playback method provided by the embodiments of the present application is implemented.

[0050] One embodiment of the above invention has the following advantages or beneficial effects: The present application obtains multimedia data according to a received multimedia playback request, and identifies the delimiter in the multimedia data; based on the delimiter, determines each segmented multimedia data in the multimedia data; plays each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-visual video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data finishes playing, obtain the feature of the last frame of the currently played segmented multimedia data; obtain the silent video corresponding to the virtual avatar corresponding to the feature of the last frame, and match the feature of the last frame with the features of each frame in the silent video to obtain the matched key frame; start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the corresponding voice-over audio finishing playing, continue to play the next segmented multimedia data. To achieve saving server resources and improving multimedia playback efficiency.

[0051] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings are used to better understand the present application and do not constitute an improper limitation to the present application. Among them:

[0053] Figure 1It is a schematic diagram of the main process of a multimedia playback method according to an embodiment of the present application;

[0054] Figure 2 It is a schematic diagram of the main process of a multimedia playback method according to an embodiment of the present application;

[0055] Figure 3 It is a general module flowchart for playing voiceovers in a multimedia playback method according to an embodiment of the present application;

[0056] Figure 4 It is a script production flowchart for a multimedia playback method according to an embodiment of the present application;

[0057] Figure 5 It is a script playback flowchart for a multimedia playback method according to an embodiment of the present application;

[0058] Figure 6 It is a schematic diagram of the main units of a multimedia playback device according to an embodiment of the present application;

[0059] Figure 7 It is an exemplary system architecture diagram to which the embodiments of the present application can be applied;

[0060] Figure 8 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present application. Detailed implementation manners

[0061] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant provisions of national laws and regulations. It should be noted that in the embodiments of the present application, some existing industry solutions such as certain software, components, models, etc. may be mentioned, and they should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution. In the technical solution of the present application, in terms of the collection, collection, update, analysis, processing, use, transmission, storage, etc. of the user's personal information, it complies with the provisions of relevant laws and regulations, is used for legal and reasonable purposes, does not violate public order and good customs, is not shared, leaked, or sold outside these legal uses, and is subject to the supervision and management of regulatory authorities. Necessary measures are taken for the user's personal information to prevent illegal access to such user personal information data, safeguard the security of the user's personal information, network security, and national security, ensure that the personnel with the right to access personal information data comply with the provisions of relevant laws and regulations, and ensure the security of the user's personal information. Once these user personal information data are no longer needed, the risk should be minimized by restricting or even prohibiting data collection and / or deleting the data.

[0062] When in use, including in certain related applications, user privacy is protected by de-identifying the data, for example, by removing specific identifiers, controlling the amount or specificity of the stored data, controlling how the data is stored, and / or other de-identification methods when in use.

[0063] Figure 1 is a schematic diagram of the main process of a multimedia playback method according to an embodiment of the present application, as Figure 1 shown, the multimedia playback method mainly includes the following steps S101 - step S105.

[0064] Step S101, obtain multimedia data according to the received multimedia playback request, and identify the delimiter in the multimedia data.

[0065] In this embodiment, the execution subject of the multimedia playback method (for example, it can be a server) can receive a multimedia playback request through a wired connection or a wireless connection. After receiving the multimedia playback request, the execution subject can obtain multimedia data, which can include data such as script text, audio, and video. The present application embodiment does not make specific limitations on the multimedia data obtained according to the multimedia playback request.

[0066] Specifically, identifying the delimiters in the multimedia data includes: identifying the text length delimiter and the off - screen voice delimiter in the multimedia data.

[0067] In some embodiments, after receiving the multimedia request, the execution subject can identify the text length delimiter and the off - screen voice delimiter in the multimedia data. By way of example, the text length delimiter can be a number, such as 2 (only for example, not making specific limitations), representing that the text is delimited every 2 characters. The off - screen voice delimiter can be a letter or a preset symbol.

[0068] Step S102, based on the delimiters, determine each segmented multimedia data in the multimedia data.

[0069] According to the identified delimiters of each type, segment the multimedia data to obtain each segmented multimedia data.

[0070] Step S103, play each segmented multimedia data according to the time sequence. If the currently played segmented multimedia data is an audio - visual video and the next segmented multimedia data is an off - screen voice audio, when the currently played segmented multimedia data is played to completion, obtain the feature of the last frame of the currently played segmented multimedia data.

[0071] Each segmented multimedia data in the multimedia data is arranged in time sequence. One segment of multimedia data is played and then the next segment is played. If the currently played segmented multimedia data is an audio - visual video and the next segmented multimedia data is an off - screen voice audio, when the currently played segmented multimedia data is played to completion, obtain the feature of the last frame of the currently played segmented multimedia data. By way of example, the feature of the last frame can be the closed - mouth image feature of the digital human image (i.e., the virtual image in the present application embodiment) of the script video, or it can be the image feature of any state of the digital human image (i.e., the virtual image in the present application embodiment) of the script video from the open - mouth state to the closed - mouth state. By obtaining the feature of the last frame of the currently played segmented multimedia data, the continuity and rationality of subsequent video playback can be ensured.

[0072] Step S104, obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the features of each frame in the silent video to obtain the key frame that is matched.

[0073] In some embodiments, by obtaining the features of the last frame of the currently played segmented multimedia data (such as the closed - mouth image features) and matching them with the features of each frame in the silent video, the key frame (i.e., the frame where the closed - mouth image features in the silent video are located) can be accurately matched to ensure the continuity and rationality of subsequent video playback.

[0074] In some embodiments, if the features of the last frame of the currently played segmented multimedia data obtained are open - mouth image features, then the open - mouth image features can be matched with the features of each frame in the silent video to obtain the matched key frame, that is, the frame where the open - mouth image features in the silent video are located.

[0075] Step S105, starting from the key frame, play the silent video while playing the corresponding voice - over audio. In response to the completion of the playback of the corresponding voice - over audio, continue to play the next segmented multimedia data.

[0076] Specifically, the method further includes: if the corresponding voice - over audio has not finished playing when the silent video finishes playing, obtain the time point where the target state of the virtual avatar in the silent video is recorded; start playing the silent video in a loop from the time point until the corresponding voice - over audio finishes playing, and then stop playing the silent video in a loop.

[0077] In some embodiments, if after starting to play the silent video from the key frame, the voice - over audio also finishes playing at the same time, then continue to play the next segmented multimedia data. For example, the next segmented multimedia data can be a video with sound or a voice - over audio.

[0078] In some embodiments, the target state can be that the digital human image (i.e., the virtual avatar in the embodiments of the present application) in the silent video is in a closed - mouth state. If after starting to play the silent video from the key frame, the voice - over audio has not finished playing, then start playing the silent video in a loop from the time point when the digital human image (i.e., the virtual avatar in the embodiments of the present application) in the silent video is in a closed - mouth state until the voice - over audio finishes playing, and then stop playing the silent video in a loop. Specifically, the execution entity can detect the playback progress of the voice - over audio in real - time. When the playback progress reaches a preset percentage (such as 99%), start entering the preparation stage for stopping the loop, and when the playback progress is 100%, stop playing the silent video in a loop.

[0079] Specifically, starting to play the corresponding silent video in a loop from the time point includes: intercepting the video from the time point to the end of the silent video as the target silent video; playing the target silent video in a loop.

[0080] Take the silent video between the time point when the digital human image (i.e., the virtual image in the embodiments of the present application) in the silent video is in the closed - mouth state and the end time point of the silent video as the target silent video, and loop - play the target silent video to ensure the continuity and rationality of multimedia playback and improve the multimedia playback efficiency.

[0081] In this embodiment, multimedia data is obtained according to the received multimedia playback request, and the delimiters in the multimedia data are identified; based on the delimiters, each segmented multimedia data in the multimedia data is determined; each segmented multimedia data is played according to the time sequence. If the currently played segmented multimedia data is an audible video and the next segmented multimedia data is a voice - over audio, when the currently played segmented multimedia data is played to the end, the features of the last frame of the currently played segmented multimedia data are obtained; the silent video corresponding to the virtual image corresponding to the features of the last frame is obtained, and the features of the last frame are matched with the features of each frame in the silent video to obtain the key frames that are matched; starting from the key frames, the silent video is played while the corresponding voice - over audio is played. In response to the corresponding voice - over audio being played to the end, the next segmented multimedia data is continued to be played. To achieve saving server resources and improving multimedia playback efficiency.

[0082] Figure 2 It is a schematic diagram of the main process of a multimedia playback method according to an embodiment of the present application. As Figure 2 shown, the multimedia playback method mainly includes the following steps S201 - step S207.

[0083] Step S201, obtain multimedia data according to the received multimedia playback request, and identify the delimiters in the multimedia data.

[0084] The execution subject can receive multimedia playback requests in real - time or at a preset time point.

[0085] In some embodiments, the execution subject can input the multimedia data into a delimiter recognition model to output each delimiter in the multimedia data. Thus, the rate and accuracy of identifying delimiters in multimedia data are improved.

[0086] Step S202, based on the delimiters, determine each segmented multimedia data in the multimedia data.

[0087] The execution subject can first classify the delimiters and segment the multimedia data according to different categories of delimiters respectively to quickly and accurately obtain each segmented multimedia data.

[0088] Step S203: Play each segmented multimedia data according to the time sequence. If the currently played segmented multimedia data is an audio-video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data finishes playing, obtain the feature of the last frame of the currently played segmented multimedia data.

[0089] When the execution entity detects the playback completion flag, it can immediately obtain the feature of the last frame of the currently played segmented multimedia data.

[0090] Step S204: Obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame.

[0091] There can be multiple virtual images. The execution entity can determine the virtual image corresponding to the feature of the last frame obtained, and obtain the corresponding pre-produced silent video according to the determined virtual image, thereby improving the accuracy of obtaining the silent video.

[0092] Step S205: Perform cosine similarity matching between the mouth state feature of the last frame and the mouth state feature of each frame in the silent video to obtain each cosine similarity.

[0093] The feature of the last frame of the obtained currently played audio-video can be the image feature of any mouth state of the digital human image (i.e., the virtual image in the embodiments of the present application) from the open-mouth state to the closed-mouth state. After converting the mouth state feature of the last frame and the mouth state feature of each frame in the silent video into vectors respectively, perform cosine similarity matching based on the converted vectors to obtain each cosine similarity.

[0094] Step S206: Determine the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame.

[0095] By determining the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame, the rate and accuracy of determining the key frame can be improved.

[0096] Step S207: Start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data.

[0097] Specifically, continuing to play the next segmented multimedia data includes: if the next segmented multimedia data is an audio-video, play the next segmented multimedia data and display the corresponding synthesized text.

[0098] If the next segmented multimedia data is an audio-video, the execution entity can directly play the next segmented multimedia data without performing feature comparison, improving the multimedia playback efficiency.

[0099] Specifically, the method also includes: when playing a silent video, the virtual image is in a breathing state.

[0100] The virtual image in the breathing state may be that the virtual image does not open its mouth to speak but can breathe, and the abdomen of the virtual image may rise and fall with the breathing, thereby ensuring that the virtual image is dynamic and not stiff, thereby improving the user experience.

[0101] Figure 3 FIG. 1 is a general module flow chart of playing voice-over according to a multimedia playing method of an embodiment of the present application. Figure 3 As shown, first read the digital human image (i.e., the virtual image in the embodiment of the present application), for example, the digital human image (i.e., the virtual image in the embodiment of the present application), and make a silent digital human breathing state video according to the read digital human image (i.e., the virtual image in the embodiment of the present application). Specifically, the process of playing the voice-over of the multimedia playback method includes: reading the script, making a video, and playing the script.

[0102] Figure 4 FIG. 1 is a flowchart of a multimedia playback method according to an embodiment of the present application. Figure 4 As shown, when making a script, the production starts, the script for playback (i.e., the text script) is read, the text script is split according to the length of the text and the voice-over delimiter, the text script and the voice-over are obtained, and the video is produced according to the text script. Finally, the video is produced including the video from the last word of the text script to the closing of the mouth.

[0103] Figure 5 FIG. 1 is a flowchart of a multimedia playback method according to an embodiment of the present application. Figure 5 As shown, when the script is playing, the playback starts, and then the produced text video is read to determine whether there is a voice-over. If not, the next video (i.e., a video with sound) is played. If so, the last frame of the text video (text video can also be called a script video. In the embodiment of the present application, both the text video and the script video are videos with sound) is extracted, and the silent video corresponding to the corresponding image is read. Through this last frame, the key frame close to the silent video of the corresponding image is searched, and the silent video is played from the key frame, and the voice-over audio is played at the same time.

[0104] The embodiment of the present application provides a method for splitting and producing scripts and playing scripts during a virtual live broadcast. The played voice-over audio does not need to be produced separately for the voice-over audio and video while ensuring that the digital human image (i.e., the virtual image in the embodiment of the present application) is in a breathing state, thereby reducing resource consumption and improving the continuity of the live broadcast.

[0105] For example, this is achieved through the following three processing flows:

[0106] Pre - produce a silent video: For each digital human image (i.e., the virtual image in the embodiments of the present application), produce a silent video from opening the mouth to closing the mouth and then remaining in the closed - mouth state. And record the time point of the closed - mouth state from the start position of the video. For example, if the digital human image (i.e., the virtual image in the embodiments of the present application) (or referred to as the digital human image (i.e., the virtual image in the embodiments of the present application)) is open - mouthed after the script video finishes playing, then wait until the open - mouth state changes to the closed - mouth state, and then record the time point from the closed - mouth state to the start position of the silent video. When the silent video finishes playing but the voice - over has not finished playing, it can be looped from this time point. When looping, it is looped from this recorded time point, rather than looping from the open - mouth state to the closed - mouth state, so as to ensure the rationality and accuracy of multimedia playback.

[0107] Produce a script video (i.e., produce an audio - visual video): Split the text script into multiple segments according to the text length and the voice - over (as a separator). Then only synthesize the text script into an audio - visual video. After the text script is synthesized into an audio - visual video, if there is a voice - over behind the script, record the characteristics of the last frame of this audio - visual video.

[0108] In the embodiments of the present application, the preset dimensions may include audio - visual videos, silent videos, and audio; a virtual image is an interactive virtual image with a "human" appearance, behavior, and even thoughts created through technologies such as computer graphics, speech synthesis technology, and deep learning, and is widely used in various fields, such as virtual characters in concerts and variety shows interacting with users, virtual customer service on enterprise websites providing online consultations, virtual tutors in teaching software providing tutoring, and virtual doctors in medical software conducting disease consultations, etc. Voice - over: The digital human stops speaking and an external sound is played.

[0109] Play the script: First, read the produced audio - visual video corresponding to the script, the audio of the voice - over, and the pre - produced silent video. After playing a segment of the script video (in the embodiments of the present application, the script video is the audio - visual video), if the next segment is a voice - over, read the characteristics of the last frame of the script video (i.e., the audio - visual video) (referring to the characteristics of the last frame of the digital human image (i.e., the virtual image in the embodiments of the present application), such as the closed - mouth image characteristics of the digital human image (i.e., the virtual image in the embodiments of the present application) in the script video). Find the frame with the closest characteristics from the start to the closed - mouth time point of the pre - produced silent video, and then play the silent video (without sound) from here, while playing the voice - over audio. When the silent video finishes playing and the voice - over has not finished playing, loop the video from the closed - mouth position of the pre - produced silent video. After the voice - over finishes playing, continue to play the video script synthesized from the text.

[0110] Figure 6It is a schematic diagram of the main units of a multimedia playback device according to an embodiment of the present application. As Figure 6 shown, the multimedia playback device 600 includes a delimiter recognition unit 601, a segmented multimedia data determination unit 602, a feature acquisition unit 603, a key frame determination unit 604, and a multimedia playback unit 605.

[0111] The delimiter recognition unit 601 is configured to obtain multimedia data according to a received multimedia playback request and recognize delimiters in the multimedia data.

[0112] The segmented multimedia data determination unit 602 is configured to determine each segmented multimedia data in the multimedia data based on the delimiter.

[0113] The feature acquisition unit 603 is configured to play each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-visual video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data finishes playing, obtain the feature of the last frame of the currently played segmented multimedia data.

[0114] The key frame determination unit 604 is configured to obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the feature of each frame in the silent video to obtain the matched key frame.

[0115] The multimedia playback unit 605 is configured to start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data.

[0116] In some embodiments, the multimedia playback device further includes Figure 6 a loop playback unit (not shown in the figure), which is configured to: if the corresponding voice-over audio has not finished playing when the silent video finishes playing, obtain the time point where the target state of the virtual image in the silent video is recorded; start looping and playing the silent video from the time point until the corresponding voice-over audio finishes playing, and then stop looping and playing the silent video.

[0117] In some embodiments, the loop playback unit is further configured to: intercept the video from the time point to the end of the silent video as the target silent video; loop and play the target silent video.

[0118] In some embodiments, the key frame determination unit 604 is further configured to: perform cosine similarity matching between the mouth state feature of the last frame and the mouth state feature of each frame in the silent video to obtain each cosine similarity; determine the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame.

[0119] In some embodiments, the separator recognition unit 601 is further configured to: recognize the text length separator and the voiceover separator in the multimedia data.

[0120] In some embodiments, the multimedia playback unit 605 is further configured to: if the next segmented multimedia data is an audio-visual video, play the next segmented multimedia data and display the corresponding synthesized text.

[0121] In some embodiments, the multimedia playback unit 605 is further configured to: the virtual avatar is in a breathing state when playing a silent video.

[0122] It should be noted that the multimedia playback method and the multimedia playback device of the present application have a corresponding relationship in terms of specific implementation content, so the repeated content will not be described again.

[0123] Figure 7 An exemplary system architecture 700 to which the multimedia playback method or the multimedia playback device of the embodiments of the present application can be applied is shown.

[0124] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0125] Users can use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 701, 702, 703, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0126] The terminal devices 701, 702, 703 may be various electronic devices having a multimedia playback processing screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0127] The server 705 can be a server that provides various services. For example, it can be a background management server (only for illustration) that supports multimedia playback requests submitted by users using terminal devices 701, 702, and 703. The background management server can obtain multimedia data according to the received multimedia playback request, identify the delimiter in the multimedia data; based on the delimiter, determine each segmented multimedia data in the multimedia data; play each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data finishes playing, obtain the feature of the last frame of the currently played segmented multimedia data; obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the features of each frame in the silent video to obtain the key frame that is matched; start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data. This is to achieve saving server resources and improving multimedia playback efficiency.

[0128] It should be noted that the multimedia playback method provided by the embodiments of the present application is generally executed by the server 705. Correspondingly, the multimedia playback device is generally set in the server 705.

[0129] It should be understood that Figure 7 the numbers of the terminal devices, networks, and servers in

[0130] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 8 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 8 The terminal device shown is only an example and should not bring any limitations to the functions and usage scope of the embodiments of the present application.

[0131] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0132] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 810 as required so that a computer program read therefrom is installed into the storage section 808 as required.

[0133] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, the above-described functions defined in the system of the present application are executed.

[0134] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0136] The units involved in the embodiments described in this application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: A processor includes a delimiter recognition unit, a segmented multimedia data determination unit, a feature acquisition unit, a key frame determination unit, and a multimedia playback unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases.

[0137] As another aspect, this application also provides a computer-readable medium. This computer-readable medium can be included in the device described in the above embodiments; it can also exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device obtains multimedia data according to a received multimedia playback request, recognizes delimiters in the multimedia data; based on the delimiters, determines each segmented multimedia data in the multimedia data; plays each segmented multimedia data according to the time sequence. If the currently played segmented multimedia data is an audio-visual video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data is played to completion, obtains the feature of the last frame of the currently played segmented multimedia data; obtains the silent video corresponding to the virtual image corresponding to the feature of the last frame, and matches the feature of the last frame with the features of each frame in the silent video to obtain the matched key frame; starts playing the silent video from the key frame, and at the same time plays the corresponding voice-over audio. In response to the corresponding voice-over audio being played to completion, continues to play the next segmented multimedia data.

[0138] The computer program product of this application includes a computer program, and the computer program implements the multimedia playback method in the embodiments of this application when executed by a processor.

[0139] According to the technical solution of the embodiments of this application, it is possible to save server resources and improve the multimedia playback efficiency.

[0140] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A multimedia playing method, characterized in that, Including: Obtain multimedia data according to the received multimedia playback request, and identify the delimiter in the multimedia data; Based on the delimiter, determine each segmented multimedia data in the multimedia data; Play each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data is finished playing, obtain the feature of the last frame of the currently played segmented multimedia data; Obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the feature of each frame in the silent video to obtain the matched key frame; Start playing the silent video from the key frame, and at the same time play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data.

2. The method according to claim 1, wherein The method further includes: If the corresponding voice-over audio has not finished playing when the silent video finishes playing, obtain the time point where the target state of the virtual image in the silent video is recorded; Start playing the silent video in a loop from the time point until the corresponding voice-over audio finishes playing, and stop playing the silent video in a loop.

3. The method according to claim 2, wherein The starting to play the corresponding silent video in a loop from the time point includes: Intercept the video of the silent video from the time point to the end of the silent video as the target silent video; Play the target silent video in a loop.

4. The method according to claim 1, wherein The matching the feature of the last frame with the feature of each frame in the silent video to obtain the matched key frame includes: Perform cosine similarity matching on the mouth state feature of the last frame and the mouth state feature of each frame in the silent video to obtain each cosine similarity; Determine the frame in the silent video corresponding to the maximum cosine similarity as the matched key frame.

5. The method according to claim 1, characterized in that, The identifying the delimiter in the multimedia data includes: Identify the text length delimiter and the voice-over delimiter in the multimedia data.

6. The method according to claim 1, characterized in that, The continuing to play the next segmented multimedia data includes: If the next segmented multimedia data is an audio-video, play the next segmented multimedia data and display the corresponding synthesized text.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: The virtual image is in a breathing state when playing the silent video.

8. A multimedia playback device, characterized in that, Including: A delimiter identification unit configured to obtain multimedia data according to the received multimedia playback request and identify the delimiter in the multimedia data; A segmented multimedia data determination unit configured to determine each segmented multimedia data in the multimedia data based on the delimiter; A feature acquisition unit configured to play each segmented multimedia data in chronological order. If the currently played segmented multimedia data is an audio-video and the next segmented multimedia data is a voice-over audio, when the currently played segmented multimedia data is finished playing, obtain the feature of the last frame of the currently played segmented multimedia data; The key frame determination unit is further configured to obtain the silent video corresponding to the virtual image corresponding to the feature of the last frame, and match the feature of the last frame with the feature of each frame in the silent video to obtain the matched key frame; The multimedia playback unit is configured to start playing the silent video from the key frame, and simultaneously play the corresponding voice-over audio. In response to the completion of the playback of the corresponding voice-over audio, continue to play the next segmented multimedia data.

9. A multimedia playback electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-7 is implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.