Audiovisual broadcast system and method thereof

The audiovisual broadcast system addresses the limitation of conventional speech synthesis by using a speech synthesis model to generate customizable voice effects for specific characters, enhancing user expression and interaction in live broadcasts.

JP2026048139AActive Publication Date: 2026-03-17シャイン レイ カンパニー リミテッド
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Conventional speech synthesis technology lacks the ability to generate natural and customizable voice effects for specific characters, limiting user expression and interaction in live broadcasts.

Method used

An audiovisual broadcast system utilizing a speech synthesis model with multiple feature parameters, enabling real-time dialogue and interaction by associating preset or custom voices with input messages, and integrating these with video broadcasts.

Benefits of technology

Enables natural and smooth voice effects, supporting multiple character voices and reducing manual tagging, thus enhancing user expression and interaction in live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048139000001_ABST
    Figure 2026048139000001_ABST
Patent Text Reader

Abstract

This invention provides an audiovisual broadcast system and a method thereof. [Solution] The audiovisual broadcast system includes a speech synthesis model that includes a plurality of feature parameters, and a system backend that enables connection and login between at least one user terminal and at least one host terminal, and receives a host number, user number, input message and preset voice number or customer voice transmitted from the user terminal, wherein the speech synthesis model associates a preset voice number with at least some of the feature parameters, or selects at least one feature parameter based on the customer voice, generates an effect sound based on the corresponding or selected feature parameter and input message, and ensures that the audiovisual broadcast provided by the host terminal includes the host terminal's video, user number and effect sound.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio-visual broadcast system and method, and more particularly to an audio-visual broadcast system and method for generating sound effects based on an input message.

Background Art

[0002] Regarding text-to-speech (TTS) technology, many people may think of Google Assistant or Siri, which are well-known voice broadcast methods in the Internet audio-visual industry. For example, in live broadcasts, in addition to promoting gifts related to photos / animations for various viewers to donate gifts to the streamer, live broadcast platforms have gradually included gifts related to sound effects.

[0003] However, conventional speech synthesis technology is usually based on the synthesis of a single sound source, and the resulting voice effects may appear unnatural or not conform to the voice characteristics of a specific character. These sound effect gifts often only have fixed content such as pre-recorded voice files, and it is often difficult for users to express what they really want to express. When using sounds they like or are familiar with, they can be broadcast on the live broadcast platform for others to hear, and greater empathy can be obtained. Conventional live broadcast platforms focus on the styles of text, stamps, and gifts. Although various types of gifts, sound effects, or stamps with voices actively developed are all displayed as a pair of one voice and one image, in many cases, these sound effects do not necessarily represent what the viewer wants to say.

[0004] Therefore, conventional speech synthesis technology still lacks a wider range of expression methods, the speech synthesis model only converts text into speech, and it lacks specialized voice effects designed for specific roles and situations.

[0005] Therefore, the industry in this technological field needs new audiovisual broadcast systems and methods that combine speech synthesis technology and real-time transmission technology to enable real-time dialogue and interaction between users or between multiple users and a host, support speech synthesis technology for the voice features of multiple characters, and provide more natural and smooth voice effects. Providing a selectable custom voice messaging communication platform system that uses less text, eliminates the need for manual tagging to reduce labor costs, and offers a variety of custom voice messaging communication platform systems is an urgent challenge for the industry. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Yihan Wu et al., published April 2022, paper title: "AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios", URL: https: / / www.isca-speech.org / archive / interspeech_2022 / index.html [Overview of the project] [Problems that the invention aims to solve]

[0007] The objective of the present invention is to solve the above problems by providing an audiovisual broadcast system that generates sound effects from input messages and is used for real-time dialogue and interaction between users or between a user and a host. [Means for solving the problem]

[0008] To achieve the above objective, the audiovisual broadcast system of the present invention includes a speech synthesis model that includes a plurality of feature parameters, and a system backend that enables connection and login between at least one user terminal and at least one host terminal, and receives a host number, user number, preset voice number and input message transmitted by the at least one user terminal, wherein the preset voice number corresponds to one of a plurality of characters provided from a character menu, and the speech synthesis model associates the preset voice number with at least a portion of the feature parameters, generates an effect sound based on the corresponding feature parameters and the input message, and causes the audiovisual broadcast provided by the host terminal to include the host terminal's video, the user number and the effect sound.

[0009] According to one embodiment provided by the audiovisual broadcast system of the present invention, the speech synthesis model is located in the system backend. The system backend transmits the user number and the sound effect to a host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, and the sound effect. The system backend transmits the input message to a host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, the input message, and the sound effect.

[0010] According to another embodiment provided by the audiovisual broadcast system of the present invention, the speech synthesis model is located on the host terminal. The system backend transmits the user number and the preset voice number to the host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, and the sound effect. The system backend transmits the input message to the host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, the input message, and the sound effect. The audiovisual broadcast provided by the host terminal further includes a character image corresponding to the preset voice number. The input message is a text message. The speech synthesis model generates these feature parameters based on training of a supervised neural network and is used to generate the corresponding sound effects, respectively.

[0011] To achieve the above objectives, the present invention also provides an audiovisual broadcast system comprising a speech synthesis model including a plurality of feature parameters, and a system backend that enables connection and login between at least one user terminal and a host terminal, and receives a host number, user number, customer voice, and input message transmitted by at least the user terminal. The speech synthesis model selects at least one feature parameter based on the customer voice, generates an effect sound based on the selected at least one feature parameter and the input message, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, and the effect sound.

[0012] According to an embodiment provided by the audiovisual broadcast system of the present invention, the speech synthesis model is located in the system backend or the host terminal, and the input message is a text message.

[0013] To solve the above problems, the present invention also provides an audiovisual broadcast system including a speech synthesis model that includes a plurality of feature parameters, and a system backend that enables connection and login of at least two user terminals, receives a user number and an input message from the speaker's user terminal, and receives either a customer voice or a preset voice number. When the system backend receives the preset voice number from the speaker's user terminal, the speech synthesis model associates the preset voice number with one of the plurality of feature parameters, generates a first effect voice based on the corresponding feature parameter and the input message, and the user terminal corresponding to the user number receives the input message and the first effect voice. When the system backend receives the customer voice from the speaker's user terminal, the speech synthesis model selects several feature parameters based on the customer voice, generates a second effect voice based on the selected several feature parameters and the input message, and causes the user terminal corresponding to the user number to receive the input message and the second effect voice.

[0014] The present invention further provides an audiovisual broadcasting method for an audiovisual broadcasting system such that an audiovisual broadcast provided by at least one host terminal includes video of the host terminal. The method includes the steps of: constructing a speech synthesis model including a plurality of feature parameters; receiving a host number, a user number and an input message from at least one user terminal, and receiving one of customer voice and a preset voice number; when the preset voice number is received, the speech synthesis model associates the preset voice number with at least a portion of the feature parameters, generates a first effect sound based on the corresponding feature parameters and the input message, so that the audiovisual broadcast provided by the host terminal corresponding to the host number includes video of the host terminal corresponding to the host number and the first effect sound; and when the customer voice is received, the speech synthesis model selects several feature parameters based on the customer voice, generates a second effect sound based on the selected several feature parameters and the input message, so that the audiovisual broadcast provided by the host terminal corresponding to the host number includes video of the host terminal corresponding to the host number and the effect sound.

[0015] The present invention further provides an audiovisual broadcasting method for the audiovisual broadcasting system, wherein the speech synthesis model is located in the system backend or in a host terminal corresponding to the host number, and the input message is a text message.

[0016] The present invention further provides another audiovisual broadcast method for the audiovisual broadcast system, comprising the steps of: constructing a speech synthesis model including a plurality of feature parameters; receiving a user number and an input message from at least one user terminal, and receiving one of a customer voice and a preset voice number; when the preset voice number is received, the speech synthesis model associates the preset voice number with one of a plurality of feature parameters, generates a first effect voice based on the corresponding feature parameter and the input message, causing the user terminal corresponding to the user number to receive the input message and the first effect voice; and when the system backend receives the customer voice from the speaker's user terminal, the speech synthesis model selects several feature parameters based on the customer voice, generates a second effect voice based on the selected several feature parameters and the input message, causing the user terminal corresponding to the user number to receive the input message and the second effect voice. [Effects of the Invention]

[0017] In short, the audiovisual broadcast system and method implemented in the present invention combine speech synthesis technology and real-time transmission technology to enable real-time dialogue and interaction between users or between multiple users and a host, supports speech synthesis technology for the voice characteristics of multiple characters, provides more natural and smooth voice effects, and offers a selectable custom voice message exchange platform system. [Brief explanation of the drawing]

[0018] [Figure 1A] This is an architecture diagram of the audiovisual broadcast system of the present invention, showing the system backend architecture of the audiovisual broadcast system. [Figure 2A]A more detailed architecture diagram of the audio-visual broadcast system of the present invention, showing the system backend of the audio-visual broadcast system combining a host terminal and a user terminal. [Figure 2B] An architecture diagram of the audio-visual broadcast system according to another embodiment of the present invention, showing that the speech synthesis model is in the system backend. [Figure 2C] An architecture diagram of the audio-visual broadcast system according to a further embodiment of the present invention, showing that two user terminals execute an audio-visual broadcast via the system backend. [Figure 3A] A block diagram of the speech synthesis model of the audio-visual broadcast system of the present invention. [Figure 3B] A block diagram of the speech synthesis model of the audio-visual broadcast system of the present invention, which obtains feature parameters using preset voice numbers. [Figure 3C] A block diagram of the speech synthesis model of the audio-visual broadcast system of the present invention, which obtains feature parameters using custom voices. [Figure 4] A flowchart showing that the audio-visual broadcast system of the present invention executes an audio-visual broadcast using a preset character menu. [Figure 5] A flowchart showing that the audio-visual broadcast system of the present invention executes an audio-visual broadcast using custom voices. [Figure 6] A schematic diagram showing an embodiment of the user interface of the audio-visual broadcast system of the present invention. [Figure 7] A schematic diagram showing another embodiment of the user interface of the audio-visual broadcast system of the present invention.

Mode for Carrying Out the Invention

[0019] Hereinafter, the implementation of the present invention will be described through specific embodiments.

[0020] First, referring to FIG. 1, it is an architecture diagram of the audio-visual broadcast system of the present invention. The system backend 2 of the audio-visual broadcast system of the present invention enables the connection and login of at least one user terminal 1 and at least one host terminal 3. The system backend 2 includes at least: a reception queue module 21, a cloud storage server 22, and a model update server 23. The system backend 2 receives the information transmitted from the user terminal 1 via the reception queue module 21, and the information includes, but is not limited to, the host number of the host terminal 3 to be transmitted, the user number of the user terminal 1, the input text message, and one of the preset voice number and the custom voice. The cloud storage server 22 of the system backend 2 records one of the preset voice number and the custom voice, generates an effect voice based on one of the preset voice number and the custom voice and the input text message via the voice synthesis model 4, and makes the video of the host terminal and the effect voice be included in the audio-visual broadcast provided by the host terminal corresponding to the host number.

[0021] Referring to Figure 2A in conjunction with Figure 1, Figure 2A is a detailed architecture diagram of the audiovisual broadcast system of the present invention. In one embodiment of the present invention, the system backend 2 of the audiovisual broadcast system accepts connections from one or more user terminals 1. The user terminal 1 comprises a character menu 11, an input module 12, a user interface 13, and an input / output module 14. The user can select the character menu 11 by operating the user interface 13 of the user terminal 1, which is a selectable menu containing preset character numbers of several default characters, and the voices of the default characters are provided as reference voices to train the speech synthesis model 4 to generate effect voices. The user inputs a text message to transmit voice through the input module 12, and further operates the user interface 13 of the user terminal 1 to send the preset character number and text message to the receive queue module 21 of the system backend 2 via the input / output module 14.

[0022] The system backend 2 enables connection and login to at least one user terminal 1 and at least one host terminal 3, and receives information transmitted from the user terminal 1 via the receive queue module 21. This information includes, but is not limited to, the host number of the host terminal 3 to be transmitted, the user number of the user terminal 1, a preset voice number, and the entered text message. The cloud storage server 22 of the system backend 2 records the preset voice number and transmits the preset voice number and the entered text message to the designated host terminal 3 via the host number to be transmitted. The host terminal 3 comprises an input / output module 32, a user interface 33, and a mixer 34.

[0023] In this embodiment of the present invention, a speech synthesis model 4 is located in each host terminal 3, and the system backend 2 includes a model update server 23. The input / output module 32 of the host terminal 3 receives a preset voice number and an input text message transmitted by the system backend 2 and transmits them to the speech synthesis model 4. The speech synthesis model 4 generates a plurality of feature parameters by training on reference voices of different characters based on a neural network, and by selecting different feature parameters, the speech synthesis model 4 can generate a corresponding effect voice based on an input text message. Figures 3A and 3B below further illustrate the speech synthesis model 4. In a specific embodiment of the present invention, after the system backend 2 has completed training the speech synthesis model 4, the model update server 23 updates the speech synthesis model 4 in each host terminal 3, and the plurality of feature parameters generated by the latest training can be downloaded to each speech synthesis model 4, and the speech synthesis model 4 is configured to generate a corresponding effect voice using the plurality of feature parameters generated by the latest training.

[0024] For example, after the speech synthesis model 4 is pre-trained using Donald Duck or Mickey Mouse as the reference voice for a character, the speech synthesis model 4 can select Donald Duck as the character based on a preset voice number and synthesize an effect voice of Donald Duck speaking a text message based on the input text message. The host accesses the mixer 34 via the user interface 33, combines this effect voice with the video of the host terminal, and performs an audiovisual broadcast to the user terminal 1 via the input / output module 32, the video of the host terminal displays the host's live screen. See also Figure 6. This audiovisual broadcast includes an effect voice broadcast 511 of a text message spoken by the selected character on the host's live screen 51, and further displays an image R2 indicating the selected character and the user number, so that both the host and the user can confirm from which user number the effect voice broadcast 511 of the text message spoken by the selected character was sent.

[0025] In a further specific embodiment of the present invention, continuing to refer to Figure 2A, the system backend 2 of the audiovisual broadcast system accepts connections from one or more user terminals 1, and the user can operate the user interface 13 of the user terminal 1 to select a character menu 11, one of which is to provide the user with the option to upload a customer voice file themselves. The customer voice file is not limited to mp3, mp4, etc., but can be in various audiovisual file formats. For example, the user wants to generate an effect sound using Jay Chou's voice and send it to the host, but Jay Chou's voice is not one of the default characters in the character menu 11 and is not available in the preset character numbers, so the user can upload Jay Chou's voice file themselves. The user enters a text message to send the voice through the input module 12, and further operates the user interface 13 of the user terminal 1 to send the customer voice file and text message to the receive queue module 21 of the system backend 2 via the input / output module 14. Subsequently, the speech synthesis model 4 of the present invention can be used to generate an effect sound and simulate Jay Chou's voice speaking the text message.

[0026] In a further embodiment of the present invention, the system backend 2 enables connection and login of at least one user terminal 1 and at least one host terminal 3, and receives information transmitted from the user terminal 1 via the receiving queue module 21. The information includes, but is not limited to, the host number of the host terminal 3 to be transmitted, the user number of the user terminal 1, a customer voice file, and an entered text message. The cloud storage server 22 of the system backend 2 records the customer voice file and transmits the customer voice file and the entered text message to the designated host terminal 3 via the host number to be transmitted.

[0027] The input / output module 32 of the host terminal 3 receives the customer voice file and input text message transmitted by the system backend 2 and transmits them to the speech synthesis model 4. The speech synthesis model 4 generates multiple feature parameters by training on reference voices of different characters based on a neural network, and by selecting different feature parameters, the speech synthesis model 4 can generate a corresponding customer effect voice based on the input text message. Figure 3C below further illustrates how the speech synthesis model 4 processes the customer voice. The host accesses the mixer 34 via the user interface 33, combines this customer effect voice with the host terminal's video, and performs an audiovisual broadcast to the user terminal 1 via the input / output module 32, with the host terminal's video displaying the host's live screen. See also Figure 6. This audiovisual broadcast includes an effect voice broadcast 511 of a text message spoken by a selected character on the host's live screen 51, and further displays the user number so that both the host and the user can see which user number sent the effect voice broadcast 511 speaking the text message.

[0028] Referring to Figure 2B, Figure 2B shows another embodiment of the audiovisual broadcast system of the present invention. The user terminal 1 is the same as shown in Figure 2A, and the speech synthesis model 4 is located in the system backend 2. The speech synthesis model 4 connects to the cloud storage server 22 to access preset voice numbers and input text messages to perform speech synthesis, generate sound effects, and send them to the cloud storage server 22. The following Figures 3A and 3B further show the speech synthesis model 4. The cloud storage server 22 sends the sound effects and input text messages to the designated host terminal 3 via the host number to be transmitted. The host accesses the mixer 34 via the user interface 33, combines the sound effects with the video on the host terminal, and performs an audiovisual broadcast to the user terminal 1 via the input / output module 32, with the video on the host terminal displaying the host's live screen.

[0029] In a further embodiment of the present invention, continuing with reference to Figure 2B, the user terminal 1 is the same as that shown in Figure 2A, and the speech synthesis model 4 is located in the system backend 2. The speech synthesis model 4 connects to the cloud storage server 22 to access custom voice files, synthesizes them with the input text message to generate sound effects, and sends them to the cloud storage server 22. Figure 3C further illustrates how the speech synthesis model 4 processes the customer voice. The cloud storage server 22 sends the sound effects to the designated host terminal 3 via the host number to which it intends to transmit. The host accesses the mixer 34 via the user interface 33, combines the sound effects with the video on the host terminal, and performs an audiovisual broadcast to the user terminal 1 via the input / output module 32, where the video on the host terminal displays the host's live screen.

[0030] Referring to Figure 2C, Figure 2C is a schematic diagram of one embodiment of the audiovisual broadcast system of the present invention, in which effect sounds are received between user terminals via a cloud server 24. The system backend 2 of the audiovisual broadcast system accepts connections from one or more user terminals 1, and the user can operate the user interface 13 of the user terminal 1 to select a character menu 11, which is a selectable menu containing preset character numbers of several default characters. The user enters a text message to send sound via the input module 12, and further operates the user interface 13 of the user terminal 1 to send the preset character number and text message to the cloud server 24 of the system backend 2 via the input / output module 14.

[0031] The cloud server 24 of the system backend 2 receives a preset character number and an input text message uploaded by the user terminal 1. The speech synthesis model 4 extracts feature parameters corresponding to the preset character number, performs speech synthesis of the input text message, and generates an effect sound. Figures 3A and 3B below further illustrate the speech synthesis model 4. In a specific embodiment of the present invention, after the system backend 2 completes training the speech synthesis model 4, the model update server 23 updates the speech synthesis model 4 and downloads a plurality of feature parameters generated by the latest training to the speech synthesis model 4. The speech synthesis model 4 is then configured to generate a corresponding effect sound using the plurality of feature parameters generated by the latest training.

[0032] In a further embodiment of the present invention, continuing with reference to Figure 2C, the system backend 2 of the audiovisual broadcast system accepts connections from one or more user terminals 1, and the user can operate the user interface 13 of the user terminal 1 to select a character menu 11, one of which is to provide the user with the option to upload a customer voice file themselves. The customer voice file is not limited to mp3, mp4, etc., but can be in various audiovisual file formats. The user inputs a text message to send the voice through the input module 12, and further operates the user interface 13 of the user terminal 1 to send the customer voice file and text message to the cloud server 24 of the system backend 2 via the input / output module 14.

[0033] The cloud server 24 of the system backend 2 receives the customer voice file uploaded by the user terminal 1 and the entered text message. The speech synthesis model 4 connects to the cloud storage server 22 to access the custom voice file and synthesizes it with the entered text message to generate the sound effect. Figure 3C further illustrates how the speech synthesis model 4 processes the customer voice. Other users can obtain the sound effect via the cloud server 24 of the system backend 2 and play and use it on the user terminal 1.

[0034] The architecture of the speech synthesis model 4 of the present invention can be seen by referring to the block diagram shown in Figure 3A. In a specific embodiment of the present invention, the constituent blocks of the speech synthesis model 4 include phoneme embeddings, a speaker-state phoneme encoder 41, a discrimination adapter 42, and a speaker-state Mel decoder 43, wherein the speaker state of the phoneme encoder 41 and Mel decoder 43 is represented by a plurality of feature parameters extracted from different characters by a speech feature extractor 44. The present invention trains the speech synthesis model 4 and its speech feature extractor 44 using reference voices of different characters, and the speech synthesis model 4 can be trained based on a supervised or unsupervised neural network. After training, the speech feature extractor 44 of the speech synthesis model 4 can extract a plurality of corresponding feature parameters by specifying a character based on a preset character number, or it can extract a plurality of feature parameters selected based on a customer voice of a non-default character.

[0035] Each component block of the speech synthesis model 4 of the present invention can be implemented in accordance with the contents disclosed in Non-Patent Document 1. Non-Patent Document 1 discloses a speech synthesis technology in which a speech synthesis model is trained using a reference speech, and the speech synthesis model is trained based on a neural network.

[0036] Continuing to refer to Figure 3A, in a specific embodiment of the present invention, the present invention trains a speech synthesis model 4 using reference voices of different characters, and the speech synthesis model 4 can be trained based on a supervised or unsupervised neural network. The speech feature extractor 44 of the trained speech synthesis model 4 can extract multiple feature parameters based on the reference voices of different characters.

[0037] When the system of the present invention uses a preset voice number and an input text message, as shown in Figure 3B, the phonemes of the preset voice number and the input text message are input to the speech synthesis model 4 of the present invention. The phoneme encoder 41 and Mel decoder 43 use the character feature parameters corresponding to the specified voice number to represent the speaker state, via a pre-trained speech feature extractor 44. The pre-trained speech synthesis model 4 embeds phonemes based on the phonemes of the text message, encodes the phonemes using the phoneme encoder 41, and synthesizes the phonemes and feature parameter encodings of the input text message. Next, the discriminative adapter 42 adjusts the audio track adaptability, and finally the Mel decoder 43 decodes according to the speaker state to generate an effect sound, which is the voice of the character corresponding to the specified voice number speaking the text message.

[0038] Next, referring to Figure 3C, when a user uses a customer voice file and inputs the text message into the speech synthesis model 4 of the present invention, the speech synthesis model 4 performs preprocessing and interpolation feature extraction on the customer voice via the speech feature extractor 44. The preprocessed and interpolated customer voice uses one or more feature parameters generated upon completion of training to represent the speaker state of the phoneme encoder 41 and Mel decoder 43. The pre-trained speech synthesis model 4 embeds phonemes based on the phonemes of the text message, encodes the phonemes with the phoneme encoder 41, and synthesizes the phonemes of the input text message with the encoding of the feature parameters. Next, the discriminative adapter 42 adjusts the audio track adaptability, and finally the Mel decoder 43 decodes according to the speaker state to generate an effect voice, which becomes the voice of the customer voice speaker speaking the text message.

[0039] Next, referring to Figure 4 in conjunction with Figure 2A, Figure 4 is a flowchart of a specific embodiment in which the present invention performs an audiovisual broadcast using a preset character number. The user receives a preset voice number from the character menu 11 and an input message, such as a text message, from the input module 12 via the user interface 13 on the user terminal 1 (S11). The input / output module 14 outputs the host number, user number, the preset voice number, and the input message to the system backend (S12). These two steps are completed at the user terminal 1.

[0040] Next, the receiving queue module 21 of the system backend 2 receives the host number, user number, the preset voice number, and the input message (S13). The cloud storage server 22 receives the preset voice number and the input message and transmits them to the speech synthesis model 4 of the host terminal 3 (S14). These two steps are completed in the system backend 2.

[0041] Next, the speech synthesis model 4 of the host terminal 3 uses the character feature parameters corresponding to the preset voice number to represent the speaker state of the phoneme encoder 41 and Mel decoder 43, and performs speech synthesis with the input message to generate an effect sound (S15). The host terminal 3 mixes the effect sound using the mixer 34 (S18). The mixer outputs the effect sound, such as an audiovisual effect like a selected gift, along with an audiovisual screen, and the mixed effect sound is broadcast audiovisually via the input / output module 32 (S19). These three steps are completed at the host terminal 3.

[0042] When the speech synthesis model 4 receives a preset voice number, uses the character feature parameters corresponding to the preset voice number to represent the speaker's state, and performs speech synthesis with the input message to generate an effect voice, if the parameters for voice generation are optimized, the speech synthesis model 4 synchronizes and updates the optimized feature parameters with the model update server 23 of the system backend 2 (S16), and the model update server 23 updates the feature parameters to facilitate the next time the corresponding feature parameters are generated (S17).

[0043] Next, referring to Figure 5 in conjunction with Figure 2A, Figure 5 is a flowchart of another specific embodiment of performing an audiovisual broadcast using a customer voice file of the present invention. The user receives the customer voice file of the character menu 11 and the input message of the input module 12 via the user interface 13 on the user terminal 1 (S21). The input / output module 14 outputs the host number, user number, customer voice file, and input message to the system backend (S22). These two steps are completed at the user terminal 1.

[0044] Next, the receiving queue module 21 of the system backend 2 receives the host number, user number, customer voice file, and input message (S23). The cloud storage server 22 receives the preset voice number and the input message and transmits them to the speech synthesis model 4 of the host terminal 3 (S24). These two steps are completed in the system backend 2.

[0045] Next, the speech synthesis model 4 of the host terminal 3 performs preprocessing and interpolation feature extraction on the customer voice via the speech feature extractor 44. The customer voice obtained through preprocessing and interpolation uses one or more feature parameters generated upon completion of training to represent the speaker state in the phoneme encoder 41 and Mel decoder 43. The pre-trained speech synthesis model 4 embeds phonemes based on the phonemes of the text message, encodes the phonemes using the phoneme encoder 41, and synthesizes the phonemes of the input text message with the encoded feature parameters. Next, the discriminative adapter 42 adjusts the audio track adaptability, and finally, the Mel decoder 43 decodes according to the speaker state to generate effect sounds (S25), and the mixer 34 mixes the effect sounds (S26). The mixer outputs the effect sounds, such as an audiovisual effect like a selected gift, along with an audiovisual screen, and the mixed effect sounds are broadcast audiovisually via the input / output module 32 (S27). These three steps are completed in the host terminal 3.

[0046] In a specific embodiment of the present invention, the feature parameters selected through the customer voice file are not stored in the speech synthesis model 4 as default voice, and the voice feature extraction step must be performed again each time a customized effect voice is executed.

[0047] Next, referring to Figure 6, which is a schematic diagram of a specific embodiment of the user interface of the present invention. The screen is divided into three areas: an auxiliary list 52 on the left, a host live screen 51 in the center, and a chat room screen 53 on the right. The host live screen 51 in the center displays the interactive screen of the host live and plays the sound effect broadcast 511 of the present invention. The sound effect broadcast 511 includes sound effects and their corresponding text, an avatar of the character of the selected menu, the user number of the user who sent the sound effect, and a corresponding gift icon, etc.

[0048] The chat room screen 53 on the right displays the text chat content between multiple users and the host. In a specific embodiment, the text of the transmitted sound effect broadcast 511 is also displayed on the chat room screen 53, combined with a corresponding special color, appearance, etc. At the bottom of the chat room screen 53 are a character menu 531 and a text input module 532. Figure 6 shows only the display of the character menu 531 in a specific embodiment, but it can also be designed as a menu-style or pop-up window. For example, the character menu 531 may include multiple avatar images R1, R2, etc., to facilitate user identification and selection.

[0049] The auxiliary list 52 on the left stores lists of related interactive features, the top left provides a menu list 54 of various functions, the daily task list 521 contains tasks that need to be completed daily, such as sending five conversations to increase interaction between the user and the host, and events that can be designed to be held at specific festivals and listed in the event list 522. The bottom left is a list of other hosts 523, and by clicking on the screen, you can switch to live streaming rooms of different hosts for interaction.

[0050] Referring to Figure 7, Figure 7 is a schematic diagram of another embodiment of the user interface of the present invention, particularly for sending messages between mobile phone users. In this embodiment, as shown in Figure 7, the screen is a chat screen 63, with the user on the right side of the chat screen 63 and the other user on the left side of the chat screen 63. In a specific embodiment, the other user sends the message "I love you," and the user's avatar 64 is displayed on the right side of the screen. The user selects R2's character through the character menu 631 and sends the text "I love you" via the text input module 632. The audiovisual broadcast system of the present invention outputs the preset voice number and the entered text message as sound effects, and a sound effect broadcast 611 is made to the other user. [Explanation of symbols]

[0051] 1 User terminal 11. Character Menu 12 Input Modules 13 User Interface 14 Input / Output Modules 2 System Backend 21 Receiving Queue Module 22 Cloud Storage Servers 23 Model Update Servers 24 Cloud Servers 3 Host terminal 32 Input / Output Modules 33 User Interface 34 Mixer 4. Speech synthesis models 41 Phoneme Encoder 42 Discrimination Adapter 43 Mel Decoder 5. User Interface 51 Host Live Screen 511 Sound Effects Broadcast 52 Auxiliary List 521 Daily Task List 522 Event List 523 Other host lists 53 Chat room screen 531 Character Menu 532 Text Input Module 54 Menu List 6. User Interface 611 Sound Effects Broadcast 63 Chat screen 631 Character Menu 632 Text Input Module 64 user avatars

Claims

1. A speech synthesis model that includes multiple feature parameters, A system backend that enables connection and login between at least one user terminal and at least one host terminal, and receives the host number, user number, preset voice number corresponding to one of a plurality of characters provided from a character menu, and input message transmitted by the at least one user terminal. An audiovisual broadcast system including, An audiovisual broadcast system comprising: a speech synthesis model that associates the preset speech numbers with at least some of the feature parameters, generates sound effects based on the corresponding feature parameters and the input message, and causes the audiovisual broadcast provided by the host terminal to include the host terminal's video, the user number, and the sound effects.

2. The audiovisual broadcast system according to claim 1, wherein the speech synthesis model is located in the system backend.

3. The audiovisual broadcast system according to claim 2, wherein the system backend transmits the user number and the sound effect to a host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the video of the host terminal, the user number, and the sound effect.

4. The audiovisual broadcast system according to claim 3, wherein the system backend transmits the input message based on the host number to the host terminal corresponding to the host number, so that the audiovisual broadcast provided by the host terminal includes the video of the host terminal, the user number, the input message, and the sound effect.

5. The audiovisual broadcast system according to claim 1, wherein the speech synthesis model is located on the host terminal.

6. The audiovisual broadcast system according to claim 5, wherein the system backend transmits the user number and the preset audio number to the host terminal corresponding to the host number based on the host number, so that the audiovisual broadcast provided by the host terminal includes the video of the host terminal, the user number, and the sound effect audio.

7. The audiovisual broadcast system according to claim 6, wherein the system backend transmits the input message based on the host number to the host terminal corresponding to the host number, so that the audiovisual broadcast provided by the host terminal includes the video of the host terminal, the user number, the input message, and the sound effect.

8. The audiovisual broadcast system according to claim 1, wherein the audiovisual broadcast provided by the host terminal further includes a character image corresponding to the preset voice number.

9. The audiovisual broadcast system according to claim 1, wherein the input message is a text message.

10. The audiovisual broadcast system according to claim 1, wherein the speech synthesis model is used to generate the feature parameters based on training of a neural network and to generate corresponding effect sounds.

11. A speech synthesis model that includes multiple feature parameters, A system backend that enables connection and login between at least one user terminal and a host terminal, and receives the host number, user number, customer voice, and entered message transmitted by at least one user terminal, an audiovisual broadcast system including An audiovisual broadcast system comprising: a speech synthesis model that selects at least one feature parameter based on the customer speech; generates an effect sound based on the selected at least one feature parameter and the input message, so that the audiovisual broadcast provided by the host terminal includes the host terminal's video, the user number, and the effect sound.

12. The audiovisual broadcast system according to claim 11, wherein the speech synthesis model is located in the system backend or the host terminal.

13. The audiovisual broadcast system according to claim 11, wherein the input message is a text message.

14. A speech synthesis model that includes multiple feature parameters, A system backend that enables connection and login of at least two user terminals, receives the user number and entered message from the speaker's user terminal, and receives either a customer voice or a preset voice number. an audiovisual broadcast system including When the system backend receives the preset voice number from the speaker's user terminal, the speech synthesis model associates the preset voice number with at least some of the plurality of feature parameters, generates a first effect voice based on the corresponding feature parameters and the input message, and the user terminal corresponding to the user number receives the input message and the first effect voice. An audiovisual broadcast system in which, when the system backend receives the customer voice from the speaker's user terminal, the speech synthesis model selects several feature parameters based on the customer voice, generates a second effect sound based on the selected feature parameters and the input message, and causes the user terminal corresponding to the user number to receive the input message and the second effect sound.

15. The audiovisual broadcast system according to claim 14, wherein the speech synthesis model is located in the system backend.

16. The audiovisual broadcast system according to claim 14, wherein the input message is a text message.

17. An audiovisual broadcasting method such that the video of at least one host terminal is included in the audiovisual broadcast provided by that host terminal, The steps include building a speech synthesis model that includes multiple feature parameters, The steps include receiving a host number, user number, and entered message from at least one user terminal, and receiving one of the customer voice and preset voice numbers, When the preset voice number is received, the speech synthesis model associates the preset voice number with at least a portion of the feature parameters, generates a first effect sound based on the corresponding feature parameters and the input message, and causes the audiovisual broadcast provided by the host terminal corresponding to the host number to include the video of the host terminal corresponding to the host number and the first effect sound; When the customer voice is received, the speech synthesis model selects several feature parameters based on the customer voice, generates a second sound effect based on the selected feature parameters and the input message, and causes the audiovisual broadcast provided by the host terminal corresponding to the host number to include the video of the host terminal corresponding to the host number and the sound effect. Audiovisual broadcasting methods, including those mentioned above.

18. The audiovisual broadcast method according to claim 17, wherein the speech synthesis model is located in the system backend or in a host terminal corresponding to the host number.

19. The audiovisual broadcast method according to claim 17, wherein the input message is a text message.

20. The steps include building a speech synthesis model that includes multiple feature parameters, The steps include receiving a user number and an entered message from at least one user terminal, and receiving one of the customer voice and a preset voice number, When the preset voice number is received, the speech synthesis model associates the preset voice number with at least some of the plurality of feature parameters, generates a first effect voice based on the corresponding feature parameters and the input message, and causes the user terminal corresponding to the user number to receive the input message and the first effect voice; When the system backend receives the customer voice from the speaker's user terminal, the speech synthesis model selects several feature parameters based on the customer voice, generates a second effect voice based on the selected feature parameters and the input message, and causes the user terminal corresponding to the user number to receive the input message and the second effect voice. Audiovisual broadcasting methods, including those mentioned above.

Citation Information

Patent Citations

  • Voice guiding device, voice guidance system and program

    JP2006330440A

  • Speech learning / synthesis system and speech learning / synthesis method

    JP2010237307A

  • Program and game system

    JP2020168109A

  • Information processing device, information processing method, program, and information processing system

    WO2022118748A1