Digital human broadcasting method and apparatus, electronic device and storage medium

By determining a smooth end frame and generating replacement data during the digital human's broadcast, the problem of abrupt interruptions in the digital human's broadcast was solved, achieving a smooth transition between sound and image.

WO2026051990A1PCT designated stage Publication Date: 2026-03-12JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

In existing technologies, digital human broadcasting suffers from abrupt interruptions, noticeable audio breaks, and obvious frame skipping in the digital human's visuals when paused or exited.

Method used

When an interruption request is received during the digital human's broadcast, the system determines the end frame in the data to be broadcast that meets the smooth end condition, and generates broadcast replacement data based on the end frame, including audio and video streams of the silent state or the end of an action, in order to achieve smooth interruption.

Benefits of technology

It achieves smooth interruption of digital human broadcasting, avoiding issues of audio breaks and video frame skipping, and providing a more natural broadcasting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025118887_12032026_PF_FP_ABST
    Figure CN2025118887_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer technology, and provides a digital human broadcasting method and apparatus, an electronic device, and a storage medium. The method comprises: in response to receiving an interruption request during a digital human broadcast process, determining current data to be broadcast, the data to be broadcast comprising at least one frame of audio and video stream to be broadcast starting from a current progress frame; traversing the data to be broadcast according to a time sequence, to determine an ending frame in the data to be broadcast satisfying a smooth ending condition, the smooth ending condition comprising the digital human being at a natural pause when broadcasting the ending frame; on the basis of the data to be broadcast in the portion from the current progress frame to the ending frame, determining broadcast replacement data; and performing digital human broadcasting of the broadcast replacement data. Embodiments of the present disclosure can achieve the effect of smooth interruption, avoiding problems such as noticeable sound breaks and frame skipping in the human avatar.
Need to check novelty before this filing date? Find Prior Art

Description

Digital human reporting method and device, electronic device, and storage medium

[0001] Cross-reference to Related Applications

[0002] This application claims priority to Chinese Patent Application No. 202411259281.2, filed September 9, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present disclosure relates to the field of computer technology, and in particular, to a digital human reporting method and device, electronic device, and storage medium. BACKGROUND

[0004] Digital human technology is a comprehensive technology that integrates machine learning, natural language processing, speech synthesis, and other technologies, aiming to create virtual characters that are lifelike, highly realistic, and have human characteristics and interactive capabilities. With the continuous development of computer technology, digital humans are gradually being applied to various aspects of daily life.

[0005] In related technologies, a digital human can be used as a virtual anchor or news reporter to report voice information. However, when a user performs a reporting pause or exit reporting operation, the digital human reporting method provided by related technologies will directly interrupt the digital human reporting and actions, resulting in a harsh interruption effect, with obvious sound breaks and obvious digital human frame skipping.

[0006] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0007] The present disclosure provides a digital human reporting method and device, electronic device, and storage medium, which at least partially overcome the problem of harsh interruption effect, obvious sound breaks, and obvious digital human frame skipping caused by directly interrupting digital human reporting and actions in related technologies.

[0008] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.

[0009] According to one aspect of the present disclosure, a digital human broadcasting is provided, comprising: in response to receiving a break request during a digital human broadcasting process, determining current to-be-broadcast data, the to-be-broadcast data comprising at least one frame of audio and video stream to be broadcast starting from a current progress frame; traversing the to-be-broadcast data in chronological order, determining an end frame in the to-be-broadcast data that satisfies a smooth end condition, the smooth end condition comprising being in a natural pause when the digital human broadcasting reaches the end frame; determining broadcasting replacement data according to the to-be-broadcast data from the current progress frame to the end frame; and performing digital human broadcasting on the broadcasting replacement data.

[0010] In some example embodiments, before determining the current to-be-broadcast data in response to receiving a break request during a digital human broadcasting process, the method further comprises: obtaining audio information and phoneme information; generating digital human lip shape parameters according to the phoneme information; rendering the digital human lip shape parameters to obtain a plurality of frames of digital human lip images; and sending an audio and video stream to the broadcasting end for digital human broadcasting, the audio and video stream comprising the plurality of frames of digital human lip images and audio data corresponding to the plurality of frames of digital human lip images.

[0011] In some example embodiments, the to-be-broadcast data further comprises action frame information corresponding to the at least one frame of audio and video stream to be broadcast, and after obtaining the audio information and the phoneme information, the method further comprises: obtaining action information; and arranging digital human broadcasting actions according to the action information to obtain a plurality of frames of action frame information, wherein any arranged action is composed of one action frame information or a plurality of consecutive action frame information, and the action comprises at least one of expression action and body action.

[0012] The sending of the audio and video stream to the broadcasting end for digital human broadcasting comprises: sending the audio and video stream and the action frame information to the broadcasting end for digital human broadcasting.

[0013] In some example embodiments, when the to-be-broadcast data further comprises action frame information corresponding to the at least one frame of audio and video stream to be broadcast, the smooth end condition further comprises not performing an action or performing an end of a target action when the digital human broadcasting reaches the end frame.

[0014] In some example embodiments, the determining of the broadcasting replacement data according to the to-be-broadcast data from the current progress frame to the end frame comprises: if no action is performed when the digital human broadcasting reaches the end frame, taking the to-be-broadcast data from the current progress frame to the end frame as the broadcasting replacement data.

[0015] In some example embodiments, the determining the playback replacement data according to the to-be-playback data from the current progress frame to the end frame comprises: if an end of a target action is not performed when the digital human plays to the end frame, obtaining the to-be-playback data from the current progress frame to the end frame; and adding action frame information of the target action end not performed and an audio and video stream in a mute state at the end of the to-be-playback data from the current progress frame to the end frame to obtain the playback replacement data, wherein a time length of the audio and video stream in the mute state is the same as a time length of the action frame information not performed.

[0016] In some example embodiments, the determining the current to-be-playback data in response to receiving a break request in the digital human playback process comprises: in response to receiving a break request in the digital human playback process, determining whether the digital human playback process supports the break; and if the break is supported, determining the current to-be-playback data.

[0017] According to another aspect of the present disclosure, a digital human playback device is also provided, comprising:

[0018] A to-be-playback data determining module is configured to determine current to-be-playback data in response to receiving a break request in a digital human playback process, the to-be-playback data comprising at least one frame of audio and video stream to be played starting from a current progress frame.

[0019] An end frame determining module is configured to traverse the to-be-playback data in chronological order and determine an end frame in the to-be-playback data that meets a smooth end condition, the smooth end condition comprising being in a natural pause when the digital human plays to the end frame.

[0020] A playback replacement data determining module is configured to determine playback replacement data according to to-be-playback data from the current progress frame to the end frame.

[0021] A playback replacement data sending module is configured to perform digital human playback on the playback replacement data.

[0022] In some example embodiments, the digital human playback device provided by the embodiments of the present disclosure further comprises: an information obtaining module configured to obtain audio information and phoneme information; a parameter generating module configured to generate digital human lip shape parameters according to the phoneme information; an image rendering module configured to render the digital human lip shape parameters to obtain a plurality of frames of digital human lip images; and an audio and video stream sending module configured to send an audio and video stream to the playback end for digital human playback, the audio and video stream comprising the plurality of frames of digital human lip images and audio data corresponding to the plurality of frames of digital human lip images.

[0023] In some example embodiments, the information obtaining module is further configured to obtain action information; the digital human broadcasting device provided by the embodiments of the present disclosure further comprises: an action frame information determining module configured to arrange the digital human broadcasting action according to the action information to obtain a plurality of action frame information, wherein any arranged action is composed of one action frame information or a plurality of continuous action frame information, and the action comprises at least one of expression action and body action; and an audio and video stream sending module configured to send the audio and video stream and the action frame information to the broadcasting end for digital human broadcasting.

[0024] In some example embodiments, when the data to be broadcast further comprises action frame information corresponding to the at least one frame of audio and video stream to be broadcast, the smooth end condition further comprises that no action is performed or the end of the target action is performed when the digital human broadcasts to the end frame.

[0025] In some example embodiments, the broadcasting replacement data determining module is configured to, if no action is performed when the digital human broadcasts to the end frame, determine the data to be broadcast from the current progress frame to the end frame as the broadcasting replacement data.

[0026] In some example embodiments, the broadcasting replacement data determining module is configured to, if the end of the target action is performed when the digital human broadcasts to the end frame, obtain the data to be broadcast from the current progress frame to the end frame; and add action frame information of an action that has not yet been performed and audio and video stream in a mute state to the end of the data to be broadcast from the current progress frame to the end frame to obtain the broadcasting replacement data, wherein the time length of the audio and video stream in the mute state is the same as the time length of the action frame information that has not yet been performed.

[0027] In some example embodiments, the data to be broadcast determining module is configured to, in response to receiving a break request during the digital human broadcasting process, determine whether the digital human broadcasting process supports the break; and if the break is supported, determine the current data to be broadcast.

[0028] According to another aspect of the present disclosure, an electronic device is also provided, which comprises: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the digital human broadcasting method of any of the above by executing the executable instructions.

[0029] According to another aspect of the present disclosure, a computer readable storage medium is also provided, which stores a computer program, and the computer program is executed by a processor to implement the digital human broadcasting method of any of the above.

[0030] According to another aspect of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the digital human broadcasting method provided in any of the various optional manners in the embodiments of the present disclosure.

[0031] The technical solution provided in the embodiments of the present disclosure can find an end frame meeting the smooth ending condition in the to-be-broadcast data when the digital human broadcasting needs to be interrupted, and determine the broadcasting replacement data according to the end frame, so as to realize smooth interruption and avoid problems such as obvious breakpoints of sound and frame skipping of human pictures.

[0032] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0033] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure. It is apparent that the accompanying drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those of ordinary skill in the art without creative labor on the basis of these drawings.

[0034] FIG. 1 shows a schematic diagram of a system architecture in an embodiment of the present disclosure;

[0035] FIG. 2 shows a flowchart of a digital human broadcasting method in an embodiment of the present disclosure;

[0036] FIG. 3 shows a flowchart of another digital human broadcasting method in an embodiment of the present disclosure;

[0037] FIG. 4 shows a flowchart of another digital human broadcasting method in an embodiment of the present disclosure;

[0038] FIG. 5 shows a schematic diagram of a digital human broadcasting device in an embodiment of the present disclosure;

[0039] FIG. 6 shows a structural block diagram of an electronic device in an embodiment of the present disclosure;

[0040] FIG. 7 shows a schematic diagram of a computer readable storage medium in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] Example implementations are now described with reference to the drawings. Example implementations can, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations.

[0042] In addition, the accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate examples of the present disclosure and, as such, a change in the size or proportion of some parts on the drawings can be exaggerated to clearly show the structure and features of the examples. Like reference numerals refer to like elements throughout the specification. It should be noted that the drawings are not necessarily drawn to scale. Identical reference numerals have been used, where possible, to designate corresponding elements that are common between the figures. Some of the blocks in the drawings are functional blocks that represent functions implemented by a processor, software, or hardware, or a combination of a processor, software, or hardware. These functional blocks can be implemented in software or hardware or a combination thereof.

[0043] The specific implementations of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0044] FIG. 1 shows an exemplary application system architecture diagram to which the digital human broadcasting method in the embodiments of the present disclosure can be applied. As shown in FIG. 1, the system architecture 100 can include a server 101, a network 102, and a broadcasting terminal 103.

[0045] In response to receiving a break request during digital human broadcasting, the server 101 can determine current to-be-broadcast data, which includes at least one frame of audio and video stream to be broadcasted starting from the current progress frame. Then, the server 101 can traverse the to-be-broadcast data in chronological order and determine an end frame in the to-be-broadcast data that satisfies a smooth end condition, which includes being in a natural pause when the digital human broadcasts to the end frame.

[0046] Then, the server 101 can determine broadcasting replacement data according to the to-be-broadcast data from the current progress frame to the end frame. Finally, the server 101 can send the broadcasting replacement data to the broadcasting terminal 103 to enable the broadcasting terminal 103 to replace the to-be-broadcast data with the broadcasting replacement data for digital human broadcasting.

[0047] The network 102 is a medium for providing a communication link between the server 101 and the broadcasting terminal 103, which can be a wired network or a wireless network.

[0048] Optionally, the wireless or wired networks described above use standard communications technologies and / or protocols. The networks typically carry Internet traffic, but can also include, without limitation, any combination of local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), proprietary networks, propriety peer-to-peer communications systems, or virtual private networks. In some embodiments, technologies and / or formats including, without limitation, Hyper Text Markup Language (HTML), Extensible Markup Language (XML), and the like are used to represent data exchanged over the networks. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPNs), Internet Protocol Security (IPSec), and the like can be used to encrypt all or some links. In other embodiments, custom and / or proprietary data communications technologies can be used in place of, or in addition to, the above-described data communications technologies.

[0049] The server 101 can be a server that provides various services, such as a background management server that provides support for operations performed by a user using the terminal device 101. The background management server can analyze and process received request data, and feed back the processing result to the terminal device.

[0050] Optionally, the server 101 can be a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.

[0051] The terminal device 103 can be a terminal device, which can be various electronic devices, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, and the like.

[0052] Optionally, the application clients installed in different broadcasting terminals 103 are the same, or are clients of the same type of application based on different operating systems. Depending on the terminal platform, the specific form of the application client can also be different, for example, the application client can be a mobile phone client, a PC client, etc.

[0053] Those skilled in the art can know that the number of servers 101, networks 102 and broadcasting terminals 103 in FIG. 1 is only illustrative, and any number of servers 101, networks 102 and broadcasting terminals 103 can be provided according to actual needs. The present disclosure does not limit this.

[0054] Under the above system architecture, the present disclosure provides a digital person broadcasting method, which can be executed by any electronic device with computing processing capability.

[0055] In some embodiments, the digital person broadcasting method provided in the present disclosure can be executed by the server of the above system architecture; in other embodiments, the digital person broadcasting method provided in the present disclosure can be realized by the server and the broadcasting terminal in the above system architecture through interaction.

[0056] FIG. 2 shows a flowchart of a digital person broadcasting method according to an embodiment of the present disclosure. As shown in FIG. 2, the digital person broadcasting method provided in the present disclosure includes the following steps S202 to S208.

[0057] S202, in response to receiving a break request in the digital person broadcasting process, determining the current to-be-broadcast data, the to-be-broadcast data including at least one frame of to-be-broadcast audio and video stream starting from the current progress frame.

[0058] In an exemplary embodiment, the to-be-broadcast audio and video stream can include video information corresponding to the digital person broadcasting picture and audio information and phoneme information corresponding to the broadcasting content, etc. The content of the audio and video stream can be determined according to the application scenario. The present disclosure does not limit the application scenario, which can be, for example, news broadcasting, entertainment live broadcast, etc.

[0059] Taking news broadcasting as an example, when a user is listening to news broadcasting, the user can pause or close the broadcasting at any time, at which time a break request for breaking the digital person broadcasting can be generated.

[0060] In some embodiments, in response to receiving the interrupt request in the digital human broadcasting process before determining the current to-be-broadcast data, the digital human broadcasting method provided by the embodiments of the present disclosure can further include: obtaining audio information and phoneme information; generating digital human lip shape parameters according to the phoneme information; rendering the digital human lip shape parameters to obtain a plurality of frames of digital human lip images; and sending an audio and video stream including the plurality of frames of digital human lip images and audio data corresponding to the plurality of frames of digital human lip images to the broadcasting end for digital human broadcasting.

[0061] It should be noted that the phoneme information can be distinguished according to the physical characteristics of the voice, for example, the phoneme information can be used to explain the pronunciation of the word and the composition of the syllable, etc. The audio information refers to a sound signal that can be recorded, stored, processed and transmitted by electronic means, which can be represented by a waveform, including parameters such as amplitude, frequency, duration, etc.

[0062] In an example embodiment, the best lip shape corresponding to each phoneme information can be determined through a predefined phoneme-lip shape mapping rule or model, so as to determine the corresponding digital human lip shape parameters based on the phoneme information. The digital human lip shape parameters may, for example, include a parameter matrix required by a digital human lip rendering model.

[0063] Then, the plurality of frames of digital human lip images can be obtained based on the digital human lip shape parameters. The number of frames of the digital human lip images may, for example, be consistent with and correspond to the audio data, so as to realize that the dynamic of the lips during digital human broadcasting can be matched with the audio data played at the time.

[0064] In some embodiments, after obtaining the audio and video stream including the plurality of frames of digital human lip images and the audio data corresponding to the plurality of frames of digital human lip images, the audio and video stream can be sent to the broadcasting end for digital human broadcasting. After receiving the interrupt request, the current playback progress of the broadcasting end, i.e., the current progress frame, can be determined. Finally, at least one frame of the audio and video stream obtained above, starting from the current progress frame to the end of the audio and video stream, can be taken as the to-be-broadcast data.

[0065] In an example embodiment, in addition to the lips that need to change with the audio data played at the time during digital human broadcasting, there can also be corresponding expression actions and body actions, etc.

[0066] In some embodiments, the to-be-broadcast data further includes action frame information corresponding to the at least one frame of to-be-broadcast audio and video stream.

[0067] In an example embodiment, after obtaining the audio information and the phoneme information, the digital human broadcasting method provided by the embodiments of the present disclosure can further include: obtaining action information; and arranging the digital human broadcasting action according to the action information to obtain a plurality of action frame information, wherein any arranged action is composed of one action frame information or a plurality of continuous action frame information, and the action includes at least one of expression action and body action.

[0068] In this case, the sending of the audio and video stream to the broadcasting end for digital human broadcasting includes: sending the audio and video stream and the action frame information to the broadcasting end for digital human broadcasting.

[0069] By way of example, the action frame information can be used for action identification of the audio and video stream, so that the digital human can make corresponding expression action, body action, etc. during broadcasting according to the played audio data. In some possible implementations, the action information can be preset, and the action arrangement is performed based on the action information to obtain the action frame information. By way of example, the expression action can include smiling, closing the mouth, etc. The body action can include waving the hand, etc.

[0070] Any action can be completed by at least one frame of digital human expression image or body image, so any action can be composed of one action frame information or a plurality of continuous action frame information. Any action frame information is used for describing the corresponding digital human expression image or body image of the frame. After generating the action frame information, each action frame information can be inserted into the corresponding position in the audio and video stream. By way of example, the insertion of each action frame information into the corresponding position in the audio and video stream can be achieved by action frame index.

[0071] In some embodiments, in response to receiving a break request during the digital human broadcasting process, determining the current to-be-broadcast data includes: in response to receiving a break request during the digital human broadcasting process, determining whether the digital human broadcasting process supports breaking; and if the digital human broadcasting process supports breaking, determining the current to-be-broadcast data.

[0072] In an example embodiment, whether the digital human broadcasting process supports breaking can be set based on experience or application scenarios. If the digital human broadcasting process supports breaking, smooth breaking can be achieved according to the method provided by the embodiments of the present disclosure. If the digital human broadcasting process does not support breaking, the playing end can send an instruction message to continue playing. In addition, the user can also be prompted that breaking is not currently supported through a pop-up prompt or the like.

[0073] S204, traversing the to-be-broadcast data in chronological order to determine an end frame in the to-be-broadcast data that meets a smooth end condition, the smooth end condition including that the digital human is in a natural pause when broadcasting to the end frame.

[0074] In an example embodiment, the audio information, phoneme information, action frame information, etc. can be buffered. When the interrupt request is obtained, the playing progress frame of the announcer can be determined and the information in the buffer can be obtained, so that the to-be-announced data including at least one frame of audio and video stream to be announced starting from the current progress frame can be obtained.

[0075] In an example, the current playing progress of the announcer, i.e., the playing progress frame, can be determined by the action frame index, and the information in the buffer can be obtained by the action frame index. Then, the action frame index can be pushed back according to the subsequent calculation time and the data transmission delay. The subsequent calculation time can represent the calculation time required to determine the end frame and obtain the announcement replacement data. The data transmission delay can represent the time of sending the announcement replacement data to the announcement end.

[0076] In an example, it can be determined whether each frame is a mute frame. If the current frame being traversed is a mute frame, the digital person is in a natural pause when announcing to the frame, and the frame can be taken as an end frame satisfying the smooth ending condition.

[0077] In some embodiments, when the to-be-announced data further includes the action frame information corresponding to the at least one frame of audio and video stream to be announced, the smooth ending condition further includes that no action is performed or the end of the target action is performed when the digital person announces to the end frame.

[0078] In an example, it can be determined whether each frame is inserted with action frame information. If the current frame being traversed is not inserted with action frame information, no action is performed when the digital person announces to the frame. In addition, if the current frame being traversed is inserted with action frame information of the target action, it can be determined whether the end of the target action is performed at the frame.

[0079] The disclosure embodiments do not limit the judgment method of whether the end of the target action is performed. The judgment method can be set based on experience or application scenarios. In an example, when the target action is currently performed for 90% or 95%, it is considered that the end of the target action is currently performed. Alternatively, if the number of frames in which the target action is not currently performed is not greater than 12 frames or 24 frames, it can be considered that the end of the target action is currently performed.

[0080] In a possible implementation, it can be first determined whether the current frame being traversed is a mute frame. If not, the traversal is continued. If it is a mute frame, it is determined whether no action is performed when the digital person announces to the frame. If no action is performed, the frame is determined as an end frame. If an action is performed, it is determined whether the frame is the end of the target action. If it is the end of the target action, the frame is determined as an end frame. If it is not the end of the target action, the traversal is continued.

[0081] It should be noted that the target action is not limited in the embodiments of the present disclosure. For example, the target action can be a mouth-closed expression action or a hand-spread body action.

[0082] In S206, the replacement data for broadcasting is determined according to the to-be-broadcast data from the current progress frame to the end frame.

[0083] In an example embodiment, since the to-be-broadcast data is directly sent to the broadcasting end for broadcasting, the breaking effect can be harsh, there can be an obvious breakpoint in the sound, and the digital human frame can be obviously skipped. Therefore, the replacement data for broadcasting that can achieve a smooth breaking effect can be determined, and the to-be-broadcast data can be discarded by the streaming end, and the replacement data for broadcasting can be sent to the broadcasting end.

[0084] In some embodiments, the replacement data for broadcasting is determined according to the to-be-broadcast data from the current progress frame to the end frame, including: if no action is performed when the digital human broadcasts to the end frame, the to-be-broadcast data from the current progress frame to the end frame is taken as the replacement data for broadcasting.

[0085] In an example embodiment, if a frame in which the digital human is in a broadcasting pause state and no action is performed can be found when the to-be-broadcast data is traversed, the effect of smooth breaking can be achieved by normally playing to the frame. Therefore, the frame can be directly taken as the end frame, and the to-be-broadcast data from the current progress frame to the end frame can be taken as the replacement data for broadcasting.

[0086] In some embodiments, the replacement data for broadcasting is determined according to the to-be-broadcast data from the current progress frame to the end frame, including: if the end of a target action is performed when the digital human broadcasts to the end frame, the to-be-broadcast data from the current progress frame to the end frame is obtained; at the end of the to-be-broadcast data from the current progress frame to the end frame, an action frame information and a mute state audio and video stream that have not been performed at the end of the target action are added to obtain the replacement data for broadcasting, wherein the time length of the mute state audio and video stream is the same as the time length of the action frame information that has not been performed.

[0087] In a possible implementation, on the premise that the digital human is in a natural pause when broadcasting to the frame, if the frame is at the end of a target action, the frame can be taken as the end frame, the to-be-broadcast data from the current progress frame to the end frame is obtained, and a mute state audio and video stream is supplemented at the end of the to-be-broadcast data to obtain the replacement data for broadcasting. For example, the mute state audio and video stream can include action frame information corresponding to a mouth-closed action and mute audio information.

[0088] The time length of the audio and video stream in the mute state is not limited in the embodiments of the present disclosure. For example, the time length of the audio and video stream in the mute state can be consistent with the time length of the motion frame information to be supplemented.

[0089] In a possible implementation, after receiving the interrupt request and querying the current playing progress, the original audio and video stream to be sent and the queue to be rendered can be discarded, and the audio information to be supplemented, the digital human lip shape parameter and the motion information are received again, so as to re-render and cover the new digital human lip image, and push the new broadcast replacement data.

[0090] S208, the digital human broadcasts the broadcast replacement data.

[0091] In an example embodiment, the architecture of the digital human broadcasting can include a scheduling system, a lip and motion arrangement service module and a picture rendering and pushing service module. The lip and motion arrangement service can be used to generate audio data to be broadcast according to text information and motion information. And the phoneme information can be obtained according to the text information, and the digital human lip shape parameter is generated according to the phoneme information. In addition, the lip and motion arrangement service module can also arrange the motion of each frame of the digital human according to the motion information to obtain the motion frame information.

[0092] The picture rendering and pushing service module can be used to render the digital human lip shape parameter to obtain multiple frames of digital human lip image. And the rendered digital human lip image can be fused with the complete digital human image. Then, the picture rendering and pushing service module can also cut the audio data and push the audio and video stream to the playing end frame by frame.

[0093] The scheduling system can determine the current data to be broadcast and determine the end frame after receiving the interrupt request. Then, the scheduling system can determine the broadcast replacement data according to the end frame. The scheduling system can send the broadcast replacement data to the broadcast end for digital human broadcasting.

[0094] For example, the digital human broadcasting of the broadcast replacement data can include: sending the broadcast replacement data to the broadcast end, so that the broadcast end replaces the data to be broadcast with the broadcast replacement data to perform digital human broadcasting.

[0095] If the end of the target motion is performed when the digital human broadcasting reaches the end frame, the scheduling system can receive the audio information to be supplemented, the digital human lip shape parameter and the motion information, and re-render and cover the new digital human lip image through the picture rendering and pushing service module, and push the new broadcast replacement data to the playing end.

[0096] The method provided by the embodiments of the present disclosure can find an end frame meeting the smooth end condition in the to-be-broadcast data when the digital human broadcast needs to be interrupted, and determine broadcast replacement data according to the end frame, so as to realize smooth interruption and avoid problems such as obvious breakpoints of sound and frame skipping of human pictures.

[0097] Exemplarily, a digital human broadcast method flowchart can be as shown in FIG. 3. The FIG. 3 can include steps S302 to S316. Among them, steps S308 to S314 are interruption processes.

[0098] S302, receiving a broadcast message.

[0099] S304, obtaining TTS (Text-to-Speech) audio information.

[0100] S306, obtaining multiple frames of digital human lip images according to phoneme information, and arranging digital human action frame information according to action information.

[0101] S308, determining whether the digital human broadcast process supports interruption. If yes, step S310 is performed. If no, step S314 is performed.

[0102] S310, buffering the audio information, the phoneme information and the action frame information.

[0103] S312, starting an interruption information monitoring thread to obtain broadcast replacement data.

[0104] Exemplarily, the life of the interruption information monitoring thread can be set according to a message duration. The message duration can refer to a broadcast duration corresponding to the to-be-broadcast data.

[0105] S314, pushing the broadcast message frame by frame.

[0106] Exemplarily, the broadcast message can be to-be-broadcast data or broadcast replacement data.

[0107] S316, ending.

[0108] It should be noted that the implementation of steps S302 to S316 can refer to the corresponding description of S202 to S208 described above, which will not be repeated here.

[0109] Exemplarily, a digital human broadcast method flowchart can be as shown in FIG. 4. The FIG. 4 can include steps S402 to S426.

[0110] S402, receiving an interruption request.

[0111] S404, obtaining a current streaming progress a.

[0112] It should be noted that the current push stream progress a is the current progress frame described above.

[0113] S406, according to the subsequent calculation time and information transmission delay, push the action frame index.

[0114] S408, obtaining the cache information according to the action frame index.

[0115] S410, traversing the cache frame.

[0116] S412, judging whether there is frame information that has not been traversed.

[0117] S414, taking out a frame.

[0118] S416, judging whether it is an action frame. If yes, S418 is executed. If no, S420 is executed.

[0119] S418, judging whether it is the end of the action frame. If yes, S420 is executed. If no, S412 is executed.

[0120] S420, judging whether it is a mute frame. If yes, S422 is executed. If no, S412 is executed.

[0121] S422, setting the end frame as b using the current traversal point index.

[0122] S424, determining the broadcast replacement data according to a and b, and pushing the broadcast replacement data to the broadcast end frame by frame.

[0123] S426, ending.

[0124] It should be noted that the implementation of steps S402 to S426 can refer to the corresponding description of S202 to S208 described above, which will not be repeated here.

[0125] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure comply with the relevant provisions of national laws and regulations. The various types of data such as personal identity data, operation data, and behavior data of individuals, customers, and crowds obtained in the embodiments of the present disclosure have been authorized.

[0126] Based on the same concept, the present disclosure also provides a digital human broadcast device, as described in the following embodiments. Since the principles of solving problems of the device embodiments are similar to those of the above-mentioned method embodiments, the implementation of the device embodiments can refer to the implementation of the above-mentioned method embodiments, and the repeated parts will not be repeated.

[0127] Figure 5 shows a digital human broadcast device in an embodiment of the present disclosure, as shown in Figure 5, the device comprises:

[0128] The to-be-broadcast data determination module 501 is configured to, in response to receiving a break request in a digital human broadcasting process, determine current to-be-broadcast data, the to-be-broadcast data including audio and video streams of at least one frame to be broadcast starting from a current progress frame;

[0129] The end frame determination module 502 is configured to traverse the to-be-broadcast data in chronological order, and determine an end frame in the to-be-broadcast data that meets a smooth end condition, the smooth end condition including being in a natural pause when the digital human broadcasts to the end frame;

[0130] The broadcast replacement data determination module 503 is configured to determine broadcast replacement data according to the to-be-broadcast data from the current progress frame to the end frame.

[0131] The broadcast replacement data sending module 504 is configured to broadcast the broadcast replacement data.

[0132] In some example embodiments, the digital human broadcasting apparatus provided by the embodiments of the present disclosure further includes:

[0133] The information acquisition module is configured to acquire audio information and phoneme information.

[0134] The parameter generation module is configured to generate digital human lip shape parameters according to the phoneme information.

[0135] The image rendering module is configured to render the digital human lip shape parameters to obtain a plurality of digital human lip image frames.

[0136] The audio and video stream sending module is configured to send audio and video streams to the broadcasting end for digital human broadcasting, the audio and video streams including the plurality of digital human lip image frames and audio data corresponding to the plurality of digital human lip image frames.

[0137] In some example embodiments, the information acquisition module is further configured to acquire action information, and the digital human broadcasting apparatus provided by the embodiments of the present disclosure further includes:

[0138] The action frame information determination module is configured to arrange digital human broadcasting actions according to the action information to obtain a plurality of action frame information, wherein any arranged action is composed of one action frame information or a plurality of continuous action frame information, and the action includes at least one of expression actions and limb actions.

[0139] The audio and video stream sending module is configured to send audio and video streams and the action frame information to the broadcasting end for digital human broadcasting.

[0140] In some example embodiments, when the to-be-broadcast data further includes action frame information corresponding to the audio and video streams of the at least one frame to be broadcast, the smooth end condition further includes not performing an action or performing an end of a target action when the digital human broadcasts to the end frame.

[0141] In some example embodiments, the broadcast replacement data determination module 503 is configured to, if no action is performed when the digital human broadcasts to the end frame, determine the to-be-broadcast data of the current progress frame to the end frame as the broadcast replacement data.

[0142] In some example embodiments, the broadcast replacement data determination module 503 is configured to, if the end of the target action is performed when the digital human broadcasts to the end frame, obtain the to-be-broadcast data of the current progress frame to the end frame; and add the action frame information and the mute state audio and video stream of the to-be-broadcast data of the current progress frame to the end frame to obtain the broadcast replacement data, wherein the time length of the mute state audio and video stream is the same as the time length of the action frame information.

[0143] In some example embodiments, the to-be-broadcast data determination module 501 is configured to, in response to receiving a break request during the digital human broadcast, determine whether the digital human broadcast supports the break; and if the digital human broadcast supports the break, determine the current to-be-broadcast data.

[0144] The apparatus provided by the embodiments of the present disclosure can find an end frame that meets the smooth end condition in the to-be-broadcast data when the digital human broadcast needs to be interrupted, and determine the broadcast replacement data according to the end frame, so as to realize smooth interruption and avoid problems such as obvious breakpoints of sound and frame skipping of human pictures.

[0145] It should be noted that the to-be-broadcast data determination module 501, the end frame determination module 502, the broadcast replacement data determination module 503, and the broadcast replacement data sending module 504 correspond to S202-S208 in the method embodiment, and the examples and application scenarios realized by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above method embodiment. It should be noted that the above modules as part of the apparatus can be executed in a computer system such as a group of computer executable instructions.

[0146] Those skilled in the art can understand that each aspect of the present disclosure can be implemented as a system, a method or a program product. Therefore, each aspect of the present disclosure can be embodied as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining software and hardware aspects, which can be collectively referred to as "circuitry", "module" or "system" herein.

[0147] The electronic device provided by the embodiments of the present disclosure includes a processor and a memory. The memory can be used to store executable instructions of the processor. The processor is configured to execute the digital human broadcasting method provided by the embodiments of the present disclosure by executing the executable instructions.

[0148] The electronic device 600 according to the embodiment of the present disclosure will be described below with reference to FIG. 6. FIG. 6 shows the electronic device 600 as an example only, and should not be taken as limiting the functions and usage range of the embodiments of the present disclosure.

[0149] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 can include, but are not limited to, at least one processing unit 610, at least one storage unit 620, and a bus 630 connecting different system components, including the storage unit 620 and the processing unit 610.

[0150] The storage unit stores program codes that can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present disclosure described in the “Exemplary Method” section of the present specification. For example, the processing unit 610 can execute the following steps of the above method embodiments:

[0151] In response to receiving an interruption request in the digital human broadcasting process, the current to-be-broadcast data is determined, the to-be-broadcast data including at least one frame of audio and video stream to be broadcasted starting from the current progress frame; the to-be-broadcast data is traversed in chronological order, and an end frame satisfying a smooth end condition in the to-be-broadcast data is determined, the smooth end condition including being in a natural pause when the digital human broadcasts to the end frame; based on the to-be-broadcast data from the current progress frame to the end frame, broadcast replacement data is determined; and the digital human broadcasts the broadcast replacement data.

[0152] The storage unit 620 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 6201 and / or a cache memory unit 6202, and can further include a read-only memory (ROM) 6203.

[0153] The storage unit 620 can further include program / utility 6204 having a set of at least one program modules 6205, including but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which can include implementation of a network environment, or some combination thereof.

[0154] Bus 630 can be one or more of several types of bus structure including a memory bus or memory controller, a peripheral bus, a graphics bus, a processor or local bus using any of a variety of bus architectures.

[0155] Electronic device 600 can also communicate with one or more external devices 640 such as a keyboard or pointing device, a Bluetooth device, etc.; other devices that enable a user to interact with electronic device 600; and / or any devices (e.g., a router, a modem, a printer, etc.) that enable electronic device 600 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 650. Still yet, electronic device 600 can communicate with one or more networks, such as one or more local area networks (LANs), one or more wide area networks (WANs), and / or the Internet, through network adapter 660. As an example, network adapter 660 can include a modem, a network card (wireless or wired), or other well-known interface devices. As shown, network adapter 660 communicates with the other components of electronic device 600 via bus 630. It should be appreciated that although not shown, other hardware and / or software components could be used in conjunction with electronic device 600. These include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0156] Those skilled in the art will readily understand that the example embodiments described herein can be implemented by software and / or by hardware coupled with software, as described above. Thus, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, etc.) or a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to perform the methods according to the embodiments of the present disclosure.

[0157] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer program product, which includes a computer program that, when executed by a processor, implements the digital person broadcasting method described above.

[0158] In the example embodiments of the present disclosure, a computer readable storage medium is also provided, which stores a computer program. The computer program, when executed by a processor, can implement the digital person broadcasting method provided by the embodiments of the present disclosure. The computer readable storage medium can be a readable signal medium or a readable storage medium.

[0159] FIG. 7 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. As shown in FIG. 7, the computer-readable storage medium 700 stores a program product capable of implementing the method of the present disclosure. In some possible implementation manners, each aspect of the present disclosure can also be implemented in the form of a program product including program codes for causing a terminal device to perform the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of the specification when the program product runs on the terminal device.

[0160] More specific examples of the computer-readable storage medium in the present disclosure can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination of the foregoing.

[0161] In the present disclosure, the computer-readable storage medium can include a data signal carried in a baseband or as part of a carrier wave propagating through the transmission medium, in which readable program codes are borne. Such a propagating data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any appropriate combination of the foregoing. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate or transmit programs for use by or in connection with an instruction execution system, apparatus or device.

[0162] Optionally, the program codes contained in the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any appropriate combination of the foregoing.

[0163] In specific implementation, the program codes for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and a conventional procedural programming language such as “C” language or similar programming languages. The program codes can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected through the Internet by using an Internet service provider).

[0164] It should be noted that, although several modules or units of the devices for action execution are mentioned in the above detailed description, the division into such modules or units is not mandatory. In fact, according to an embodiment of the present disclosure, features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, features and functionalities of one module or unit described above can be further divided into a plurality of modules or units.

[0165] Furthermore, although the various steps of the methods in the present disclosure are described in a particular order in the drawings, this is not required or implied as to the order of the steps or that all of the steps shown must be performed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, one step can be broken into multiple steps, etc.

[0166] From the above description of the embodiments, those skilled in the art will readily perceive that the example embodiments described herein can be implemented by software and / or by software in combination with the necessary hardware. Thus, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, U disk, mobile hard disk, etc.) or network, and includes a number of instructions to make a computing device (which can be a personal computer, server, mobile terminal, or network device, etc.) execute the method according to the embodiments of the present disclosure.

[0167] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including such modifications as would be apparent to those skilled in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope of the present disclosure is indicated by the appended claims.

Claims

1. A digital human broadcasting method, comprising: in response to receiving a break request during a digital human broadcasting process, determining current to-be-broadcast data, the to-be-broadcast data including at least one frame of audio and video stream to be broadcast from a current progress frame; traversing the to-be-broadcast data in chronological order to determine an end frame in the to-be-broadcast data that meets a smooth end condition, the smooth end condition including being in a natural pause when the digital human broadcasts to the end frame; determining broadcast replacement data according to the to-be-broadcast data from the current progress frame to the end frame; and broadcasting the digital human according to the broadcast replacement data. Before the determining the current to-be-broadcast data in response to receiving the break request during the digital human broadcasting process, the method further comprises:

2. The digital human anchoring method of claim 1, wherein, obtaining audio information and phoneme information; generating digital human lip shape parameters according to the phoneme information; rendering the digital human lip shape parameters to obtain a plurality of digital human lip images; and sending an audio and video stream to a broadcasting end for digital human broadcasting, the audio and video stream including the plurality of digital human lip images and audio data corresponding to the plurality of digital human lip images. The to-be-broadcast data further includes action frame information corresponding to the at least one frame of audio and video stream to be broadcast, and after the obtaining the audio information and the phoneme information, the method further comprises:

3. The digital human anchoring method of claim 2, wherein, obtaining action information; arranging digital human broadcasting actions according to the action information to obtain a plurality of action frame information, wherein any arranged action is composed of one action frame information or a plurality of continuous action frame information, and the action includes at least one of an expression action and a body action; The sending the audio and video stream to the broadcasting end for digital human broadcasting comprises: sending the audio and video stream and the action frame information to the broadcasting end for digital human broadcasting. When the to-be-broadcast data further includes the action frame information corresponding to the at least one frame of audio and video stream to be broadcast, the smooth end condition further includes not performing an action or performing an end of a target action when the digital human broadcasts to the end frame.

4. The digital human broadcasting method of claim 1 or 3, wherein, The determining the broadcast replacement data according to the to-be-broadcast data from the current progress frame to the end frame comprises:

5. The digital human anchoring method of claim 4, wherein, if no action is performed when the digital human broadcasts to the end frame, taking the to-be-broadcast data from the current progress frame to the end frame as the broadcast replacement data. The determining the broadcast replacement data according to the to-be-broadcast data from the current progress frame to the end frame comprises:

6. The digital human anchoring method of claim 4, wherein, if an end of a target action is performed when the digital human broadcasts to the end frame, obtaining the to-be-broadcast data from the current progress frame to the end frame; and adding action frame information of an action that has not yet been performed and audio and video stream of a mute state to an end of the to-be-broadcast data from the current progress frame to the end frame to obtain the broadcast replacement data, wherein a time length of the audio and video stream of the mute state is the same as a time length of the action frame information that has not yet been performed. The determining the current to-be-broadcast data in response to receiving the break request during the digital human broadcasting process comprises:

7. The digital human anchoring method of claim 1, wherein, in response to receiving the break request during the digital human broadcasting process, determining whether the digital human broadcasting process supports the break; and ​ If the interruption is supported, determine the current data to be broadcast.

8. A digital human broadcasting device, comprising: a data to be broadcast determination module configured to, in response to receiving an interruption request during digital human broadcasting, determine current data to be broadcast, the data to be broadcast comprising audio and video streams of at least one frame to be broadcast starting from a current progress frame; an end frame determination module configured to, in chronological order, traverse the data to be broadcast, and determine an end frame in the data to be broadcast that satisfies a smooth end condition, the smooth end condition comprising being in a natural pause when the digital human broadcasts to the end frame; a broadcast replacement data determination module configured to determine broadcast replacement data according to the data to be broadcast from the current progress frame to the end frame; a broadcast replacement data sending module configured to broadcast the broadcast replacement data.

9. An electronic device, comprising: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the digital human broadcasting method according to any one of claims 1-7 by executing the executable instructions.

10. A computer readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the digital human broadcasting method according to any one of claims 1-7.

11. A computer program product, the computer program product comprising computer instructions stored in a computer readable storage medium, a processor of an electronic device reading the computer instructions from the computer readable storage medium, the processor executing the computer instructions to cause the electronic device to execute the digital human broadcasting method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Voice interruption method and device

    CN111540349A

  • Voice interaction method and device and computer readable storage medium

    CN112637431A

  • Virtual digital human rendering method, rendering engine and system

    CN115550711A

  • Online digital human action generation method and system based on strategy gradient optimization

    CN117221653A

  • Digital people broadcasting method and device, electronic equipment and storage medium

    CN119128184A