Multi-role live streaming method, apparatus and electronic device

By acquiring audio and scripts, generating text through speech recognition, creating character timelines, and controlling visual animations, the problem of insufficient interactivity and poor synchronization in existing live streaming technologies has been solved. This enables multi-character interaction and an immersive experience, reducing costs and improving efficiency.

CN122120476APending Publication Date: 2026-05-29HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
Filing Date
2026-02-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing live streaming technologies, the playback of pre-recorded audio content lacks interactivity, live role-playing requires high costs and complex scheduling, and existing virtual character solutions have difficulty accurately identifying different characters and time points in the audio content. Audio playback and visual effects are difficult to synchronize precisely, affecting the user experience.

Method used

By acquiring target audio and script, performing speech recognition processing to generate text, generating character timelines, creating virtual accounts to simulate character behavior, and controlling visual animation display during live streaming, intelligent character recognition and timeline generation of audio content are achieved, and audio playback and visual animation are precisely synchronized.

Benefits of technology

It achieves intelligent character recognition and timeline generation for audio content, accurately marks the dialogue time range of different characters, provides an immersive user experience, supports multi-character interaction effects, improves automated processing efficiency, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120476A_ABST
    Figure CN122120476A_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-role live broadcast method and device and electronic equipment, obtaining target audio and corresponding target script, performing speech recognition processing on the target audio to generate target text corresponding to the target audio, the target text including multiple subtexts and time stamps corresponding to the multiple subtexts respectively; based on the target text and the target script, generating a role timeline, the role timeline including multiple role information and dialogue time ranges corresponding to the multiple role information respectively; creating multiple virtual accounts for simulating dialogue behaviors of virtual roles corresponding to the role information included in the role timeline according to the role information corresponding to the role timeline; playing the target audio during the live broadcast process, and controlling the virtual roles corresponding to the multiple virtual accounts to perform visual special effect display according to the role timeline during the playing of the target audio. This way converts the target audio live broadcast into a multi-role interactive form of live broadcast, improving user experience and live broadcast interactivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of live streaming technology, and in particular to a multi-role live streaming method, apparatus, and electronic device. Background Technology

[0002] In today's internet landscape, live streaming has become a crucial medium for information dissemination, fulfilling information needs in areas such as e-commerce, entertainment, and education. Among various live streaming scenarios, there exists one where pre-produced audio and video content is provided to viewers in a live format, such as replays of past live streams or pre-recorded concerts or films. However, directly playing pre-recorded audio content only allows users to passively listen, resulting in a rather monotonous experience. Summary of the Invention

[0003] In view of this, the purpose of this disclosure is to provide a multi-role live streaming method, device, and electronic device to enhance the interactivity of role-playing live streaming.

[0004] In a first aspect, embodiments of this disclosure provide a multi-role live streaming method, the method comprising: acquiring target audio and a corresponding target script; performing speech recognition processing on the target audio to generate target text corresponding to the target audio; wherein the target script includes multiple character information and dialogue text corresponding to each of the multiple character information; the target text includes multiple sub-texts and timestamps corresponding to each of the multiple sub-texts; generating a character timeline based on the target text and the target script; wherein the character timeline includes multiple character information and dialogue time ranges corresponding to each of the multiple character information; wherein the dialogue time ranges are determined based on the timestamps in the target text; creating multiple virtual accounts according to the character information corresponding to the character timeline; wherein the virtual accounts are used to simulate the dialogue behavior of virtual characters corresponding to the character information contained in the character timeline; playing the target audio during the live stream, and controlling the virtual characters corresponding to the multiple virtual accounts to perform visual animation display according to the character timeline during the playback of the target audio.

[0005] Secondly, this disclosure also provides a multi-role live streaming device, which includes: an audio processing module for acquiring target audio and a corresponding target script, performing speech recognition processing on the target audio, and generating target text corresponding to the target audio; wherein the target script includes multiple character information and dialogue text corresponding to the multiple character information; the target text includes multiple sub-texts and timestamps corresponding to the multiple sub-texts; a timeline generation module for generating a character timeline based on the target text and the target script; wherein the character timeline includes multiple character information and dialogue time ranges corresponding to the multiple character information; wherein the dialogue time ranges are determined based on the timestamps in the target text; an account creation module for creating multiple virtual accounts according to the character information corresponding to the character timeline; wherein the virtual accounts are used to simulate the dialogue behavior of virtual characters corresponding to the character information contained in the character timeline; and an audio playback module for playing the target audio during the live stream, and controlling the virtual characters corresponding to the multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio.

[0006] Thirdly, this disclosure provides an electronic device including a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the above-described multi-role live streaming method.

[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned multi-role live streaming method.

[0008] The embodiments disclosed herein bring the following beneficial effects: This disclosure provides a multi-role live streaming method, apparatus, and electronic device. First, it acquires target audio and its corresponding target script. Then, it performs speech recognition processing on the target audio to generate target text corresponding to the target audio. The target script includes multiple character information and corresponding dialogue texts for each character. The target text includes multiple sub-texts and corresponding timestamps for each sub-text. Based on the target text and target script, a character timeline is generated. The character timeline includes multiple character information and corresponding dialogue time ranges, where the dialogue time ranges are determined based on the timestamps in the target text. Next, multiple virtual accounts are created based on the character information corresponding to the character timeline. These virtual accounts simulate the dialogue behavior of virtual characters corresponding to the character information included in the character timeline. During the live stream, the target audio is played, and during playback, the virtual characters corresponding to the multiple virtual accounts are controlled to display visual animations based on the character timeline. This method achieves intelligent character recognition and timeline generation for audio content, accurately marking the dialogue time ranges of different characters. It also simulates realistic multi-role interaction effects through virtual accounts and achieves precise synchronization between audio playback and visual animations, providing an immersive user experience.

[0009] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objects and other advantages of this disclosure are realized and obtained through the structures particularly pointed out in the description, claims and drawings.

[0010] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a multi-role live streaming method provided in this embodiment of the disclosure; Figure 2 This is a schematic diagram illustrating the binding of a role and a virtual account, provided as an embodiment of this disclosure. Figure 3 A schematic diagram illustrating a preset animation effect provided in an embodiment of this disclosure; Figure 4 An exception handling method provided in this embodiment of the disclosure; Figure 5 This is a schematic diagram of the structure of a multi-role live streaming device provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0014] Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0015] With the explosive growth of AIGC (Artificial Intelligence Generated Content), more and more virtual humans are replacing human broadcasters in live streams, and viewers are no longer satisfied with simple static replays but demand interactive live content. The technical solutions provided in related technologies mainly include: Option 1, the traditional audio live streaming solution, involves directly playing pre-recorded audio content, where users can only passively listen and lack interactivity.

[0016] Option 2, live role-playing, requires multiple live streamers to perform role-playing simultaneously, resulting in high labor costs and scheduling difficulties.

[0017] Option 3, a simple robot playback solution, uses a single robot account to play audio content and cannot demonstrate the effect of multi-role interaction.

[0018] Option 4, the static role allocation scheme, pre-assigns fixed roles to different accounts, but cannot dynamically adjust role allocation based on content.

[0019] The interactive live streaming methods described above all have certain applicability, but they also have some issues that affect efficiency. For example, traditional audio playback lacks visual feedback and character differentiation, resulting in a monotonous user experience; live role-playing requires multi-person collaboration, which is complex and costly; existing solutions struggle to accurately identify different characters and corresponding time points in audio content; audio playback and visual effects are difficult to synchronize precisely, affecting the user experience; and existing solutions are unable to process large amounts of audio content in batches, resulting in low automation.

[0020] To address the aforementioned issues, this disclosure provides a multi-role live streaming method, apparatus, and electronic device, which is applicable to any scenario where live streaming is conducted using virtual characters.

[0021] To facilitate understanding of the embodiments of the present invention, a multi-role live streaming method is first described in detail, such as... Figure 1 As shown, the method includes the following specific steps: Step S102: Obtain the target audio and the corresponding target script, perform speech recognition processing on the target audio, and generate the target text corresponding to the target audio; wherein, the target script includes multiple character information and dialogue texts corresponding to the multiple character information; the target text includes multiple sub-texts and timestamps corresponding to the multiple sub-texts.

[0022] The target script is a script produced by the planners that includes character information and dialogue text for relevant characters. This character information may include, but is not limited to, character names, character descriptions, character inner thoughts, and character emotional descriptions. In one specific embodiment, the target audio can be audio content recorded by a real person or a robot based on the target script. Alternatively, it can be understood as audio generated by a real person or a robot recording the dialogue text corresponding to the character information in the target script.

[0023] In another specific embodiment, the target audio can also be any audio content, and the target script is generated after organizing and adjusting the audio content of the target audio.

[0024] The system first performs speech recognition processing on the input target audio, generating target text corresponding to the target audio. This target text includes multiple subtexts and a timestamp for each subtext. These subtexts correspond to the dialogue text in the target script; while the subtexts may not be entirely identical to the dialogue text, the semantics are the same. This speech recognition processing is essentially a method of converting audio into text.

[0025] For example, the target script's script format is: [Narrator] Hello everyone, welcome to our program; [Character A] Today we're going to discuss xxx; [Character B] This is a very interesting topic.

[0026] Based on the target script, target audio can be generated. Different audio frames in the target audio have different timestamps. By performing speech recognition on the target audio, sentence content containing timestamps can be obtained (equivalent to the sub-text mentioned above). For example, the timestamps for the sentence content "Hello everyone, welcome to our program" are 0 for the start time and 3000 for the end time; the timestamps for the sentence content "Today we are going to discuss the topic of xxx" are 3000 for the start time and 6000 for the end time; and the timestamps for the sentence content "This is a very interesting topic" are 6000 for the start time and 9000 for the end time.

[0027] Step S104: Generate a character timeline based on the target text and the target script; wherein, the character timeline includes: multiple character information and the dialogue time range corresponding to each of the multiple character information; wherein, the dialogue time range is determined based on the timestamp in the target text.

[0028] The aforementioned character timeline is a data structure that records the duration (equivalent to the dialogue time range) and content of different characters speaking at specific points in the audio content. Specifically, it identifies the correspondence between multiple sub-texts included in the target text and the dialogue text in the target script. Then, based on the dialogue text corresponding to the sub-text, it determines the character information corresponding to that sub-text, and further determines the timestamp corresponding to that sub-text as the timestamp of the corresponding character information. Finally, it determines the dialogue time range corresponding to that character information based on this timestamp. The timestamp includes a start time and an end time, thus defining the time range between the start and end times as the dialogue time range.

[0029] For example, the timeline for the above characters could be: the dialogue time range for the narrator character is 0-3000; the dialogue time range for character A is 3000-6000; and the dialogue time range for character B is 6000-9000.

[0030] Step S106: Create multiple virtual accounts based on the character information corresponding to the character timeline; wherein, the virtual accounts are used to simulate the dialogue behavior of the virtual characters corresponding to the character information contained in the character timeline.

[0031] In practice, the system automatically creates and manages multiple virtual accounts to bind them to the role information corresponding to the role timeline. This allows the multiple virtual accounts to simulate the behavior of real users. Alternatively, it can be understood that the multiple virtual accounts simulate the dialogue behavior of the virtual characters corresponding to the role information contained in the role timeline, so that the live stream viewers visually perceive it as dialogue behavior performed by virtual characters.

[0032] Step S108: Play the target audio during the live stream, and control the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio.

[0033] During the live stream, target audio is played, and visual effects are simultaneously displayed on the live stream interface. This ensures that the virtual character indicated by the visual effects matches the virtual character indicated by the character information corresponding to the currently playing target audio. The character displaying the visual effects on the live stream interface is determined based on the character's timeline. The visual effects could be displayed on the virtual character's avatar or on the corresponding virtual account's identifier, etc.

[0034] This invention provides a multi-role live streaming method that enables intelligent role recognition and timeline generation of audio content, accurately marks the dialogue time range of different roles, simulates real multi-role interaction effects through virtual accounts, and achieves precise synchronization between audio playback and visual effects, providing an immersive user experience.

[0035] The following examples illustrate how character timelines are generated.

[0036] Specifically, the process of generating a character timeline based on the target text and the target script can include: using a large language model, performing content matching between the dialogue text corresponding to each character information in the target script and the subtext in the target text to obtain the subtext corresponding to each character information; for each character information, determining the timestamp corresponding to the character information from the target text based on the subtext corresponding to the character information, determining the dialogue time range corresponding to the character information based on the timestamp corresponding to the character information; and generating a character timeline based on the dialogue time range corresponding to each character information in the target script.

[0037] In its implementation, the Large Language Model (MLM) is a deep learning-based natural language processing model capable of understanding and generating human-language text. Based on the MLM, natural language processing can be performed on the dialogue text and subtexts within the target script to obtain content-matched dialogue text and subtexts. From these matching dialogue texts, the corresponding character information can be obtained, leading to the subtexts for each character. Then, based on the timestamps of each subtext in the target text, the timestamps of the corresponding character information are determined. Finally, the start and end times of the timestamps are used to determine the dialogue time range corresponding to each character.

[0038] The above method relies on the text processing capabilities of a large language model to associate content and obtain the dialogue time range of the characters based on the timestamps in the target text, thereby generating an accurate character timeline. In other words, this method can achieve intelligent character recognition and timeline generation of audio content, accurately marking the speaking time points and durations of different characters.

[0039] The following examples are used to describe a way of displaying visual animation effects.

[0040] Specifically, after creating multiple virtual accounts based on the character information corresponding to the character timeline, the character information contained in the character timeline is bound to the multiple virtual accounts respectively, so that the virtual accounts can simulate the dialogue behavior of the virtual characters corresponding to the character information; wherein, each virtual account is used to bind at least one character information.

[0041] In practice, the system automatically creates multiple virtual accounts based on the role information corresponding to the role's timeline. Each virtual account has complete user attributes, meaning each virtual account can simulate user live-streaming behavior. Then, based on the role information contained in the role's timeline, the backend administrator binds the virtual accounts to the corresponding role information. Each virtual account can be bound to one or more role information. Specifically, the role information corresponding to multiple virtual characters with similar voices can be bound to the same virtual account. For example, the role information of multiple female characters can be bound to a first virtual account, and the role information of multiple male characters can be bound to a second virtual account; where the first virtual account and the second virtual account are different virtual accounts.

[0042] like Figure 2 The diagram illustrates a method for binding a character to a virtual account according to an embodiment of the present invention. After generating the character timeline, the system automatically extracts all character names from the character information contained in the timeline and automatically creates two virtual accounts, that is... Figure 2 Robot account A and robot account B are used in the system. The backend administrator binds the robot accounts to their corresponding role names; each robot account can be associated with multiple role names, such as... Figure 2 In the example, robot account A is associated with virtual characters named A, B, C, and D; robot account B is associated with virtual characters named 1, 2, 3, and 4.

[0043] Furthermore, after creating multiple virtual accounts based on the character information corresponding to the character's timeline, the live streaming interface displays the account identifiers corresponding to each virtual account. These account identifiers can be the virtual account's name or corresponding avatar, etc.

[0044] Based on the above description, the specific process of playing target audio during a live stream and controlling the visual animation display of virtual characters corresponding to multiple virtual accounts according to the character timeline during the playback of the target audio may include: playing target audio and synchronizing the playback progress information of the target audio during the live stream; determining the target character information corresponding to the target audio being played at the current time based on the character timeline and playback progress information, and controlling the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live stream interface.

[0045] In practice, the aforementioned playback progress information is used to indicate the progress of the target audio being played, that is, to indicate the audio content or corresponding time of the target audio currently being played.

[0046] In one specific embodiment, a target audio can be played through a live streaming server and the playback progress information of the target audio can be synchronized. By using the playback progress information and the character timeline, the virtual characters corresponding to multiple virtual accounts can be controlled to display visual animation effects.

[0047] In another optional embodiment, the target audio can be played through a first server, and the playback progress information of the target audio can be synchronized with a second server. During the live broadcast, the target audio needs to be played through the first server, and when the target audio starts playing, the playback progress information of the target audio is synchronized with the second server so that the second server can load the corresponding character timeline and virtual account. When the target audio finishes playing, the first server is automatically driven to load the next audio.

[0048] The first server and the second server mentioned above are two different servers, performing different functions during the live stream. Both the first server and the second server can be cloud servers. When the second server receives the playback progress information sent by the first server, it immediately determines the character information corresponding to the current time based on the character timeline, and identifies the virtual character corresponding to this character information as the target character information corresponding to the target audio being played at the current time in the live stream. A preset animation effect is then displayed on the account identifier of the virtual account corresponding to the target character information in the live stream interface, so that the live stream viewers visually perceive that the target audio being played at the current time is emitted by the target character information.

[0049] In an optional embodiment, the process of determining the target character information corresponding to the target audio being played at the current time based on the character timeline and playback progress information, and controlling the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, may include: generating animation triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information; if the current time matches the dialogue time range of the target character information, triggering the animation triggering instructions corresponding to the target character information to trigger the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

[0050] In practical implementation, the dialogue time range and playback progress information of each character information indicated on the character timeline can be matched to determine the dialogue start time of each character information when the target audio is played. At the dialogue start time, an animation trigger command is generated and triggered. When the animation trigger command is triggered, the preset animation effect is displayed on the account identifier of the virtual account in the live broadcast interface.

[0051] Furthermore, the process of generating motion effect triggering instructions corresponding to the dialogue time range of each character's information based on the character timeline and playback progress information, and triggering the motion effect triggering instructions to display preset motion effects on the account identifier of the virtual account corresponding to the target character's information in the live broadcast interface if the current time matches the dialogue time range of the target character's information, includes: generating delayed messages corresponding to each character's information based on the dialogue time range and playback progress information indicated by the character timeline; wherein, the delayed message includes the motion effect triggering instructions and the triggering time of the motion effect triggering instructions, which is the dialogue start time of the character's information; sending the delayed message to a preset message queue; responding to the current time being the dialogue start time of the target character's information, consuming the delayed message corresponding to the target character's information in the preset message queue to control the display of preset motion effects on the account identifier of the virtual account corresponding to the target character's information in the live broadcast interface; wherein, the dialogue start time of the target character's information is the minimum time value of the dialogue time range corresponding to the target character's information.

[0052] When playing the target audio, the system first matches the character timeline with the virtual account. Then, based on the dialogue time range and audio playback progress corresponding to each character's information as indicated by the character timeline, it generates a delay message for each character's information. This delay message includes the animation trigger command corresponding to the character's information and the trigger time of the animation trigger command, which is the start time of the character's dialogue. The delay message is consumed or executed when the trigger time indicated by the delay message is reached.

[0053] When consuming delayed messages corresponding to the target role information in the preset message queue, a preset animation effect is displayed on the account identifier of the virtual account corresponding to the target role information in the live broadcast interface; the preset animation effect can be a water ripple effect or an identifier shaking effect, etc.

[0054] like Figure 3 The image shown is a schematic diagram illustrating a preset animation effect provided in an embodiment of the present invention. Figure 3 The profile pictures of streamer 1, streamer 2, and streamer 3 represent the account identifiers of three different virtual accounts. Each virtual account has a different account identifier. Figure 3 The middle part is the water ripple effect displayed on the streamer's profile picture.

[0055] Specifically, preset animations are used in the live stream to indicate which character is speaking. These preset animations subtly suggest to users that the virtual character displaying the preset animations is speaking, creating the impression that many people are role-playing and recreating a radio drama in the live stream.

[0056] In practical applications, the system uses the cloud playback capabilities of RTC to control and establish a synchronization mechanism between playback progress and animation display. When a cloud playback task (here, the playback task is the task of playing the target audio) starts, the task start status is recorded, and a delay message is sent synchronously to a preset message queue. When consuming the delay message in the message queue, based on the role information bound to the virtual account, a preset animation is displayed on the account identifier of the virtual account corresponding to the role information.

[0057] The above methods can be applied to various scenarios, such as intelligent audiobook playback, which transforms audiobook content into a multi-character interactive format to enhance user experience; AI-enhanced radio drama playback, which uses AI technology to reproduce the multi-character dialogue effects of radio dramas; voice-enhanced social interaction, which provides richer interactive formats in voice-based social platforms; and interactive educational content, which transforms educational audio content into role-playing formats to increase the fun of learning.

[0058] In addition, the above method uses AI technology to automatically identify different characters in the audio with an accuracy rate of over 95%; moreover, this method supports batch processing of audio content, which is more than 10 times more efficient than manual processing; at the same time, this method does not require multiple people to collaborate, and a single person can manage multiple virtual character live broadcast rooms, and this method supports the expansion of any number of characters to adapt to different types of audio content.

[0059] The following examples illustrate exception handling and fault tolerance mechanisms.

[0060] During the live stream, it is determined in real time whether the target audio being played is consistent with the target audio displayed with preset animation effects in the live stream interface; if they are inconsistent, the currently playing audio is adjusted to be consistent with the target audio displayed with preset animation effects in the live stream interface.

[0061] In practical implementation, during live streaming, an audio playback task is executed to play the target audio. While playing the target audio, a delay message needs to be generated based on the character's timeline corresponding to the target audio and the synchronized playback progress information. This delay message is then used to control the account identifier of the virtual account displaying the preset animation effects on the live streaming interface. If the played target audio is inconsistent with the target audio of the preset animation effects displayed on the live streaming interface, a mismatch will occur between the virtual account associated with the preset animation effects displayed when playing the target audio and the virtual account bound to the character information of the target audio being played at the current time. This can lead to display errors or playback chaos. Therefore, it is necessary to monitor in real time whether the played target audio is consistent with the actual target audio of the preset animation effects to ensure the synchronization and consistency of the audio and animation display.

[0062] Furthermore, during the live stream, the target audio is played, and while playing the target audio, the virtual characters corresponding to multiple virtual accounts are controlled to display visual animations according to the character timeline. Then, the task status of the audio playback task corresponding to the target audio is set to the task start state. Here, the target audio is the audio that is currently being played. The task status of the audio playback task is checked periodically. If the task status is the task end state, the next audio playback task is automatically queried and started, so that the task status of the next audio playback task is set to the task start state.

[0063] In practical implementation, when playing the target audio during the live stream, the task status of the audio playback task corresponding to the target audio is set to the task start state. This disclosure sets up a timed task to periodically check the task status of the audio playback task. If the task status is in the task start state, no processing is required; if the task status is in the task end state, it means that no audio is being played in the live stream. In this case, the system cannot automatically switch to the next audio, and it is necessary to automatically query the next audio playback task and start the next audio playback task so that the audio corresponding to the next audio playback task can be played in the live stream.

[0064] In an optional embodiment, the target audio is the current task in a predefined audio playback task sequence. The aforementioned audio playback task can be understood as the current task, and the next audio playback task is the next audio playback task in the audio playback task sequence following the current task.

[0065] like Figure 4The illustration shows an exception handling method according to an embodiment of the present invention. A scheduled task queries the task status of the audio playback task corresponding to the currently playing target audio. If the task status is not in progress (equivalent to the task being started as described above), it indicates that the audio playback task callback has been lost. The audio playback task will enter an abnormal chain and cannot automatically switch to the next audio playback task. The scheduled inspection task will detect this exception during its scheduled scan. If the current task is not in progress, the scheduled task will proactively re-initiate the unstarted audio playback task to avoid task interruption caused by callback loss.

[0066] In optional embodiments, this disclosure also supports speech-to-text transcription failure handling, that is, supports resubmitting transcription tasks and provides a manual intervention interface; it also supports large language model timeout handling, that is, adopts an asynchronous processing mechanism to avoid synchronous call timeout; it can also perform motion effect synchronization deviation handling, that is, prevents repeated triggering of motion effects through delayed message idempotency verification.

[0067] Corresponding to the above method embodiments, this disclosure also provides a multi-role live streaming device, such as... Figure 5 As shown, the device includes: The audio processing module 50 is used to acquire the target audio and the corresponding target script, perform speech recognition processing on the target audio, and generate the target text corresponding to the target audio; wherein, the target script includes multiple character information and dialogue texts corresponding to the multiple character information; the target text includes multiple sub-texts and timestamps corresponding to the multiple sub-texts.

[0068] The timeline generation module 51 is used to generate a character timeline based on the target text and the target script; wherein, the character timeline includes: multiple character information and the dialogue time range corresponding to each of the multiple character information; wherein, the dialogue time range is determined based on the timestamp in the target text.

[0069] The account creation module 52 is used to create multiple virtual accounts based on the character information corresponding to the character timeline; among them, the virtual accounts are used to simulate the dialogue behavior of the virtual characters corresponding to the character information contained in the character timeline.

[0070] The audio playback module 53 is used to play the target audio during the live broadcast, and to control the virtual characters corresponding to multiple virtual accounts to display visual animation effects according to the character timeline during the playback of the target audio.

[0071] The aforementioned multi-role live streaming device achieves intelligent role recognition and timeline generation of audio content, accurately marks the dialogue time range of different roles, simulates real multi-role interaction effects through virtual accounts, and achieves precise synchronization between audio playback and visual effects, providing an immersive user experience.

[0072] Furthermore, the aforementioned timeline generation module 51 is used to: perform content matching between the dialogue text corresponding to each character information in the target script and the subtext in the target text based on the large language model, to obtain the subtext corresponding to each character information; for each character information, determine the timestamp corresponding to the character information from the target text based on the subtext corresponding to the character information, and determine the dialogue time range corresponding to the character information based on the timestamp corresponding to the character information; and generate a character timeline based on the dialogue time range corresponding to each character information in the target script.

[0073] Furthermore, the above-mentioned device also includes a role binding module, used to: after creating multiple virtual accounts based on the role information corresponding to the role timeline, bind the role information contained in the role timeline to the multiple virtual accounts respectively, so that the virtual accounts simulate the dialogue behavior of the virtual characters corresponding to the role information; wherein, each virtual account is used to bind at least one role information.

[0074] Furthermore, the aforementioned device also includes an identifier display module, used to: after creating multiple virtual accounts based on the role information corresponding to the role timeline, display the account identifiers corresponding to the multiple virtual accounts in the live streaming interface.

[0075] Furthermore, the aforementioned audio playback module 53 is used to: play target audio and synchronize the playback progress information of the target audio during the live broadcast; determine the target character information corresponding to the target audio played at the current time based on the character timeline and playback progress information, and control the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

[0076] Furthermore, the aforementioned audio playback module 53 is used to: generate motion effect triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information; if the current time matches the dialogue time range of the target character information, trigger the motion effect triggering instructions corresponding to the target character information to display preset motion effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

[0077] Furthermore, the aforementioned audio playback module 53 is also used to: generate delayed messages corresponding to each character information based on the dialogue time range and playback progress information corresponding to each character information indicated by the character timeline; wherein, the delayed message includes an animation trigger command and the trigger time of the animation trigger command, the trigger time being the dialogue start time of the character information; send the delayed message to a preset message queue; in response to the current time being the dialogue start time of the target character information, consume the delayed message corresponding to the target character information in the preset message queue, so as to control the display of a preset animation on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface; wherein, the dialogue start time of the target character information is the minimum time value of the dialogue time range corresponding to the target character information.

[0078] Furthermore, the aforementioned device also includes a real-time monitoring module, used to: determine in real time whether the currently playing target audio is consistent with the target audio of the preset animation displayed in the live broadcast interface; if they are inconsistent, adjust the currently playing audio and the target audio of the preset animation displayed in the live broadcast interface to be consistent with each other.

[0079] Furthermore, the aforementioned device also includes an exception handling module, used for: playing target audio during the live stream, and, during the playback of the target audio, controlling the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline, setting the task status of the audio playback task corresponding to the target audio to the task start state; wherein, the target audio is the currently playing audio; periodically checking the task status of the audio playback task, and if the task status is the task end state, automatically querying the next audio playback task and starting the next audio playback task so that the task status of the next audio playback task is the task start state.

[0080] The multi-role live streaming device provided in this disclosure has the same implementation principle and technical effects as the aforementioned method embodiments. For the sake of brevity, any parts not mentioned in the device embodiments can be referred to the corresponding content in the aforementioned method embodiments.

[0081] This embodiment also provides an electronic device, such as... Figure 6 As shown, the electronic device includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the aforementioned multi-role live streaming method. This electronic device can be a server or a terminal device.

[0082] Specifically, the aforementioned multi-role live streaming method includes: acquiring target audio and its corresponding target script; performing speech recognition processing on the target audio to generate target text corresponding to the target audio; wherein, the target script includes multiple character information and dialogue text corresponding to each of the multiple character information, and the target audio is generated based on the dialogue text; the target text includes multiple sub-texts and timestamps corresponding to each of the multiple sub-texts; generating a character timeline based on the target text and the target script; wherein, the character timeline includes multiple character information and dialogue time ranges corresponding to each of the multiple character information; wherein, the dialogue time ranges are determined based on the timestamps in the target text; creating multiple virtual accounts according to the character information corresponding to the character timeline; wherein, the virtual accounts are used to simulate the dialogue behavior of the virtual characters corresponding to the character information contained in the character timeline; playing the target audio during the live stream, and controlling the virtual characters corresponding to the multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio.

[0083] The aforementioned multi-role live streaming method achieves intelligent role recognition and timeline generation of audio content, accurately marks the dialogue time range of different roles, simulates real multi-role interaction effects through virtual accounts, and achieves precise synchronization between audio playback and visual effects, providing an immersive user experience.

[0084] In an optional embodiment, the step of generating a character timeline based on the target text and the target script includes: performing content matching on the dialogue text corresponding to each character information in the target script and the subtext in the target text based on a large language model to obtain the subtext corresponding to each character information; for each character information, determining the timestamp corresponding to the character information from the target text based on the subtext corresponding to the character information, determining the dialogue time range corresponding to the character information based on the timestamp corresponding to the character information; and generating a character timeline based on the dialogue time range corresponding to each character information in the target script.

[0085] In an optional embodiment, after creating multiple virtual accounts based on the role information corresponding to the role timeline, the method further includes: binding the role information contained in the role timeline to the multiple virtual accounts respectively, so that the virtual accounts simulate the dialogue behavior of the virtual characters corresponding to the role information; wherein, each virtual account is used to bind at least one role information.

[0086] In an optional embodiment, after creating multiple virtual accounts based on the role information corresponding to the role timeline, the above method further includes: displaying the account identifiers corresponding to the multiple virtual accounts in the live streaming interface.

[0087] In an optional embodiment, the steps of playing target audio during the live stream and controlling the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio include: playing target audio and synchronizing the playback progress information of the target audio during the live stream; determining the target character information corresponding to the target audio played at the current time based on the character timeline and playback progress information, and controlling the display of preset animations on the account identifier of the virtual account corresponding to the target character information in the live stream interface.

[0088] In an optional embodiment, the steps of determining the target character information corresponding to the target audio being played at the current time based on the character timeline and playback progress information, and controlling the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, include: generating animation triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information; if the current time matches the dialogue time range of the target character information, triggering the animation triggering instructions corresponding to the target character information to trigger the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

[0089] In an optional embodiment, the step of generating motion effect triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information includes: generating delayed messages corresponding to each character information based on the dialogue time range and playback progress information indicated by the character timeline; wherein, the delayed message includes a motion effect triggering instruction and a triggering time of the motion effect triggering instruction, the triggering time being the dialogue start time of the character information; and sending the delayed message to a preset message queue; based on this, the step of triggering a motion effect triggering instruction if the current time matches the dialogue time range of the target character information, so as to trigger the display of a preset motion effect on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, includes: in response to the current time being the dialogue start time of the target character information, consuming the delayed message corresponding to the target character information in the preset message queue, so as to control the display of the preset motion effect on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface; wherein, the dialogue start time of the target character information is the minimum time value of the dialogue time range corresponding to the target character information.

[0090] In an optional embodiment, the method further includes: determining in real time whether the currently playing target audio is consistent with the target audio of the preset animation displayed in the live broadcast interface; if they are inconsistent, adjusting the currently playing audio and the target audio of the preset animation displayed in the live broadcast interface to be consistent audio.

[0091] In an optional embodiment, after playing the target audio during the live stream and controlling the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio, the above method further includes: setting the task status of the audio playback task corresponding to the target audio to the task start state; wherein, the target audio is the audio currently being played; periodically checking the task status of the audio playback task, and if the task status is the task end state, automatically querying the next audio playback task and starting the next audio playback task so that the task status of the next audio playback task is the task start state.

[0092] Furthermore, Figure 6 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 101, the communication interface 103 and the memory 100 connected via the bus 102.

[0093] The memory 100 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0094] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0095] This disclosure also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are invoked and executed by a processor, they cause the processor to implement the aforementioned multi-role live streaming method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0096] Specifically, the aforementioned multi-role live streaming method includes: acquiring target audio corresponding to a target script; performing speech recognition processing on the target audio to generate target text corresponding to the target audio; wherein, the target script includes multiple character information and dialogue text corresponding to each of the multiple character information, and the target audio is generated based on the dialogue text; the target text includes multiple sub-texts and timestamps corresponding to each of the multiple sub-texts; generating a character timeline based on the target text and the target script; wherein, the character timeline includes multiple character information and dialogue time ranges corresponding to each of the multiple character information; wherein, the dialogue time ranges are determined based on the timestamps in the target text; creating multiple virtual accounts according to the character information corresponding to the character timeline; wherein, the virtual accounts are used to simulate the dialogue behavior of virtual characters corresponding to the character information contained in the character timeline; playing the target audio during the live stream, and controlling the virtual characters corresponding to the multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio.

[0097] The aforementioned multi-role live streaming method achieves intelligent role recognition and timeline generation of audio content, accurately marks the dialogue time range of different roles, simulates real multi-role interaction effects through virtual accounts, and achieves precise synchronization between audio playback and visual effects, providing an immersive user experience.

[0098] In an optional embodiment, the step of generating a character timeline based on the target text and the target script includes: performing content matching on the dialogue text corresponding to each character information in the target script and the subtext in the target text based on a large language model to obtain the subtext corresponding to each character information; for each character information, determining the timestamp corresponding to the character information from the target text based on the subtext corresponding to the character information, determining the dialogue time range corresponding to the character information based on the timestamp corresponding to the character information; and generating a character timeline based on the dialogue time range corresponding to each character information in the target script.

[0099] In an optional embodiment, after creating multiple virtual accounts based on the role information corresponding to the role timeline, the method further includes: binding the role information contained in the role timeline to the multiple virtual accounts respectively, so that the virtual accounts simulate the dialogue behavior of the virtual characters corresponding to the role information; wherein, each virtual account is used to bind at least one role information.

[0100] In an optional embodiment, after creating multiple virtual accounts based on the role information corresponding to the role timeline, the above method further includes: displaying the account identifiers corresponding to the multiple virtual accounts in the live streaming interface.

[0101] In an optional embodiment, the steps of playing target audio during the live stream and controlling the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio include: playing target audio and synchronizing the playback progress information of the target audio during the live stream; determining the target character information corresponding to the target audio played at the current time based on the character timeline and playback progress information, and controlling the display of preset animations on the account identifier of the virtual account corresponding to the target character information in the live stream interface.

[0102] In an optional embodiment, the steps of determining the target character information corresponding to the target audio being played at the current time based on the character timeline and playback progress information, and controlling the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, include: generating animation triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information; if the current time matches the dialogue time range of the target character information, triggering the animation triggering instructions corresponding to the target character information to trigger the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

[0103] In an optional embodiment, the step of generating motion effect triggering instructions corresponding to the dialogue time range of each character information based on the character timeline and playback progress information includes: generating delayed messages corresponding to each character information based on the dialogue time range and playback progress information indicated by the character timeline; wherein, the delayed message includes a motion effect triggering instruction and a triggering time of the motion effect triggering instruction, the triggering time being the dialogue start time of the character information; and sending the delayed message to a preset message queue; based on this, the step of triggering a motion effect triggering instruction if the current time matches the dialogue time range of the target character information, so as to trigger the display of a preset motion effect on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, includes: in response to the current time being the dialogue start time of the target character information, consuming the delayed message corresponding to the target character information in the preset message queue, so as to control the display of the preset motion effect on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface; wherein, the dialogue start time of the target character information is the minimum time value of the dialogue time range corresponding to the target character information.

[0104] In an optional embodiment, the method further includes: determining in real time whether the currently playing target audio is consistent with the target audio of the preset animation displayed in the live broadcast interface; if they are inconsistent, adjusting the currently playing audio and the target audio of the preset animation displayed in the live broadcast interface to be consistent audio.

[0105] In an optional embodiment, after playing the target audio during the live stream and controlling the virtual characters corresponding to multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio, the above method further includes: setting the task status of the audio playback task corresponding to the target audio to the task start state; wherein, the target audio is the audio currently being played; periodically checking the task status of the audio playback task, and if the task status is the task end state, automatically querying the next audio playback task and starting the next audio playback task so that the task status of the next audio playback task is the task start state.

[0106] The aforementioned function, when implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0108] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A multi-role live streaming method, characterized in that, The method includes: The target audio and its corresponding script are acquired, and speech recognition processing is performed on the target audio to generate target text corresponding to the target audio. The target script includes multiple character information and dialogue texts corresponding to the multiple character information. The target text includes multiple sub-texts and timestamps corresponding to the multiple sub-texts. Based on the target text and the target script, a character timeline is generated; wherein, the character timeline includes: multiple character information and dialogue time ranges corresponding to the multiple character information; wherein, the dialogue time ranges are determined based on timestamps in the target text; Based on the character information corresponding to the character timeline, multiple virtual accounts are created; wherein, the virtual accounts are used to simulate the dialogue behavior of the virtual characters corresponding to the character information contained in the character timeline; During the live stream, the target audio is played, and while playing the target audio, the visual animations of the virtual characters corresponding to the multiple virtual accounts are controlled to be displayed according to the character timeline.

2. The method according to claim 1, characterized in that, The step of generating a character timeline based on the target text and the target script includes: Based on the large language model, the dialogue text corresponding to each character information in the target script and the subtext in the target text are matched to obtain the subtext corresponding to each character information. For each piece of character information, the timestamp corresponding to the character information is determined from the target text based on the subtext corresponding to the character information, and the dialogue time range corresponding to the character information is determined based on the timestamp corresponding to the character information. A character timeline is generated based on the dialogue time range corresponding to each character information in the target script.

3. The method according to claim 1, characterized in that, After the step of creating multiple virtual accounts based on the character information corresponding to the character timeline, the method further includes: The character information contained in the character timeline is bound to the plurality of virtual accounts respectively, so that the virtual accounts simulate the dialogue behavior of the virtual characters corresponding to the character information; wherein, each virtual account is used to bind at least one of the character information.

4. The method according to any one of claims 1-3, characterized in that, After the step of creating multiple virtual accounts based on the character information corresponding to the character timeline, the method further includes: The live streaming interface displays the account identifiers corresponding to the multiple virtual accounts.

5. The method according to claim 4, characterized in that, The steps of playing the target audio during the live stream, and controlling the virtual characters corresponding to the multiple virtual accounts to display visual animations according to the character timeline during the playback of the target audio, include: During the live stream, the target audio is played and the playback progress information of the target audio is synchronized. Based on the character timeline and the playback progress information, the target character information corresponding to the target audio being played at the current time is determined, and a preset animation effect is displayed on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

6. The method according to claim 5, characterized in that, The step of determining the target character information corresponding to the target audio being played at the current time based on the character timeline and the playback progress information, and controlling the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live streaming interface, includes: Based on the character timeline and the playback progress information, generate motion effect triggering instructions corresponding to the dialogue time range of each character information; If the current time matches the dialogue time range of the target character information, trigger the animation trigger command corresponding to the target character information to display the preset animation on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface.

7. The method according to claim 6, characterized in that, The step of generating motion effect triggering instructions corresponding to the dialogue time range of each character based on the character timeline and the playback progress information includes: Based on the dialogue time range corresponding to each character information indicated by the character timeline and the playback progress information, a delay message corresponding to each character information is generated; wherein, the delay message includes a motion effect trigger command and the trigger time of the motion effect trigger command, and the trigger time is the dialogue start time of the character information; Send the delayed message to a preset message queue; The step of triggering the animation trigger command if the current time matches the dialogue time range of the target character information, so as to trigger the display of a preset animation on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface, includes: In response to the current time being the dialogue start time of the target character information, the delayed message corresponding to the target character information is consumed in the preset message queue to control the display of preset animation effects on the account identifier of the virtual account corresponding to the target character information in the live broadcast interface; wherein, the dialogue start time of the target character information is the minimum time value of the dialogue time range corresponding to the target character information.

8. A multi-role live streaming device, characterized in that, The device includes: An audio processing module is used to acquire target audio and its corresponding target script, perform speech recognition processing on the target audio, and generate target text corresponding to the target audio; wherein, the target script includes multiple character information and dialogue text corresponding to the multiple character information respectively; the target text includes multiple sub-texts and timestamps corresponding to the multiple sub-texts respectively; A timeline generation module is used to generate a character timeline based on the target text and the target script; wherein, the character timeline includes: multiple character information and dialogue time ranges corresponding to the multiple character information; wherein, the dialogue time ranges are determined based on timestamps in the target text; The account creation module is used to create multiple virtual accounts based on the character information corresponding to the character timeline; wherein, the virtual accounts are used to simulate the dialogue behavior of the virtual characters corresponding to the character information contained in the character timeline; The audio playback module is used to play the target audio during the live broadcast, and to control the virtual characters corresponding to the multiple virtual accounts to display visual animation effects according to the character timeline during the playback of the target audio.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the multi-role live streaming method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the multi-role live streaming method according to any one of claims 1-7.