Audio-visual broadcasting system and method thereof

KR102999075B1Active Publication Date: 2026-08-03시너레이 컴퍼니 리미티드
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020240114048
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-08-03
Estimated Expiration
2044-08-26

Smart Images

  • Figure 112024092790377-PAT00002_ABST
    Figure 112024092790377-PAT00002_ABST
Patent Text Reader

Abstract

The video and audio broadcasting system of the present invention comprises a voice synthesis model including a plurality of feature parameters and a system backend that enables at least one user side to connect and log in with at least one streaming host side, and receives a streaming host number, a user number, an input message, and a preset voice number or a customized voice transmitted from the user side. The voice synthesis model selects at least one feature parameter corresponding to at least a portion of the feature parameters of the preset voice number or according to the customized voice, and generates an effect voice together with the input message according to the corresponding or selected feature parameters, thereby including the streaming host side image, the user number, and the effect voice in the video and audio broadcast provided by the streaming host side.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a video-audio broadcasting system and a method thereof, in particular to a video-audio broadcasting system and a method thereof that generates a single effect sound according to an input message. Background Technology

[0002] Text-to-speech (TTS) technology is a voice broadcasting method well-known in the internet audio-video industry, to the extent that most people associate it with Miss Google or Siri. In order to incentivize various customers to present gifts to streaming hosts during live broadcasts, streaming platforms are increasingly introducing voice effect-related gifts in addition to promoting gifts related to pictures and animations.

[0003] However, existing voice synthesis technology is generally based on a single sound source, and the generated voice effects may sound unnatural or not match the vocal characteristics of a specific character. These voice effect gifts are often fixed in content; for instance, when recording a high-quality audio file, it is often difficult to express what the user truly wants to convey. If the voice is one that the user likes or is familiar with, it can be broadcast on live streaming platforms to be heard by others, potentially having a greater impact on empathy. While existing streaming platforms focus on text, stickers, and gift styles, actively developing various gift types, voice effects, or audio stickers, these are often implemented using a single voice and single screen correspondence method, and such voice effects cannot express everything the audience wants to say.

[0004] Existing speech synthesis technology still has the disadvantage of lacking more diverse expressions; it merely converts text into speech along with the speech synthesis model and lacks unique sound effects designed for various characters and situations.

[0005] Therefore, the industry in this technology field requires new video-audio broadcasting systems and methods. Such systems combine speech synthesis technology and real-time transmission technology to enable immediate interaction between users or between multiple users and a streaming host, and support speech synthesis technology with multi-character voice characteristics to provide more natural and seamless voice effects. Reducing operating costs by using fewer phrases and eliminating the need for manual tagging, while simultaneously providing a selectable and customized voice message communication platform system, is an urgent problem that the industry must solve. The problem to be solved

[0006] To solve the above problems, the objective of the present invention is to provide a video-audio broadcasting system that applies effect voice generated in an input message to immediate interaction between users or between a user and a streaming host.

[0007] To achieve the above objective, the video and audio broadcasting system of the present invention includes a voice synthesis model and a system back end. The voice synthesis model includes a plurality of feature parameters, and the system back end enables at least one user side to connect and log in with at least one streaming host side, and receives a streaming host number, a user number, a preset voice number, and an input message transmitted by at least one user side. This preset voice number corresponds to one of a plurality of characters provided in a character menu, and the voice synthesis model corresponds the preset voice number to at least a portion of the feature parameters, and generates effect voice along with the input message according to the corresponding feature parameters, so that the streaming host image, the user number, and the effect voice are included in the video and audio provided by the streaming host side. means of solving the problem

[0008] In one embodiment of the video and audio broadcasting system of the present invention, a voice synthesis model is placed in a system back end, and the system back end transmits a user number and effect voice according to a streaming host number to a streaming host side corresponding to the streaming host number, so that the video and audio broadcast provided by the streaming host side includes the streaming host side image, user number, and effect voice. The system back end transmits an input message according to a streaming host number to a streaming host side corresponding to the streaming host number, so that the video and audio broadcast provided by the streaming host side includes the streaming host side image, user number, input message, and effect voice.

[0009] In one embodiment of the video-audio broadcasting system of the present invention, the speech synthesis model is configured on the streaming host side. The system backend transmits a user number and a preset voice number according to the streaming host number to the streaming host side corresponding to the streaming host number, so that the video-audio broadcast provided by the streaming host side includes the image, user number, and effect voice of the streaming host side. The system backend transmits an input message according to the streaming host number to the streaming host side corresponding to the streaming host number, so that the video-audio broadcast provided by the streaming host side includes the image, user number, input message, and effect voice of the streaming host side. The video-audio broadcast provided by the streaming host side also includes a character image corresponding to the preset voice number. This input message is a text message. The speech synthesis model is used to generate these feature parameters based on supervised neural network training and to generate effect voices corresponding to each.

[0010] To achieve the above objective, the present invention provides a video-audio broadcasting system comprising a voice synthesis model and a system backend. The voice synthesis model includes a plurality of feature parameters, and the system backend allows connection login between at least one user side and at least one streaming host side, and receives a streaming host number, a user number, a customized voice, and an input message transmitted from at least one user side. The voice synthesis model selects at least one feature parameter according to the customized voice, generates effect voice according to the selected at least one feature parameter and the input message, and includes the streaming host side image, user number, and effect voice in the video-audio broadcast provided by the streaming host side.

[0011] In one embodiment of the video-audio broadcasting system of the present invention, the voice synthesis model is placed on the back end or streaming host side of the system, and the input message is a text message.

[0012] To achieve the above objective, the present invention provides a video-audio broadcasting system comprising a speech synthesis model and a system back end. The speech synthesis model includes a plurality of feature parameters, and the system back end allows connection logins of at least two user sides, receives a user number and an input message from the speaking user side, and receives either a customized voice or a preset voice number. When the system back end receives the preset voice number from the speaking user side, the speech synthesis model assigns the preset voice number to one of the plurality of feature parameters, generates a first effect voice according to the corresponding feature parameter and the input message, and causes the user side corresponding to the user number to receive the input message and the first effect voice. When the system back end receives the customized voice from the speaking user side, the speech synthesis model selects and uses multiple feature parameters according to the customized voice, generates a second effect voice according to the selected multiple feature parameters and the input message, and causes the user side corresponding to the user number to receive the input message and the second effect voice.

[0013] The present invention further provides a video-audio broadcasting method of a video-audio broadcasting system that includes a video of one streaming host side in a video-audio broadcast provided by at least one streaming host side. The method comprises the following: establishing a voice synthesis model including a plurality of feature parameters; receiving a streaming host number, a user number, and an input message from at least one user side, and receiving one of a customized voice and a preset voice number; when the preset voice number is received, the voice synthesis model corresponds the preset voice number to at least a portion of the feature parameters and generates a first effect voice according to the corresponding feature parameters and the input message, thereby including an image of the streaming host side corresponding to the streaming host number and the first effect voice in a video-audio broadcast provided by the streaming host side corresponding to the streaming host number; and when the customized voice is received, the voice synthesis model selects and uses multiple feature parameters according to the customized voice, and generates a second effect voice according to the selected multiple feature parameters and the input message, thereby including an image of the streaming host side corresponding to the streaming host number and the effect voice in a video-audio broadcast provided by the streaming host side corresponding to the streaming host number.

[0014] The present invention further provides a video-audio broadcasting method of a video-audio broadcasting system, wherein the voice synthesis model is placed on the streaming host side corresponding to the system back end or streaming host number, and the input message is a text message.

[0015] The present invention further provides a video-audio broadcasting method of a video-audio broadcasting system. The method comprises the following: establishing a voice synthesis model including a plurality of feature parameters; receiving a user number and an input message from at least one user side and receiving either a customized voice or a preset voice number; when the preset voice number is received, the voice synthesis model maps the preset voice number to one of the plurality of feature parameters and generates a first effect voice according to the corresponding feature parameter and the input message, thereby causing the user side corresponding to the user number to receive the input message and the first effect voice; and when the system back end receives a customized voice from a speaking user side, the voice synthesis model selects and uses multiple feature parameters according to the customized voice, and generates a second effect voice according to the selected feature parameter and the input message, thereby causing the user side corresponding to the user number to receive the input message and the second effect voice.

[0016] To summarize the above description, the video-audio broadcasting system and method according to the present invention combine voice synthesis technology and real-time transmission technology to allow immediate interaction between users or between multiple users and a streaming host, and support voice synthesis of multiple character voice characteristics to provide more natural and smooth voice effects, and simultaneously provide a message communication platform system for selectable customized voices. Effects of the invention

[0017] Included in the contents of the present invention. Brief explanation of the drawing

[0018] Figure 1 is a structural diagram of the video-audio broadcasting system of the present invention, showing the system back end of the video-audio broadcasting system. FIG. 2a is a more detailed structural diagram of the video-audio broadcasting system of the present invention, showing a structure in which the system back end, streaming host side, and user side of the video-audio broadcasting system are combined. FIG. 2b is a structural diagram of another embodiment of the video-audio broadcasting system of the present invention, wherein the voice synthesis model is located at the system back end. FIG. 2c is a structural diagram of another embodiment of the video-audio broadcasting system of the present invention, showing two user sides performing video-audio broadcasting through the system back end. FIG. 3a is a block diagram of the voice synthesis model of the video-audio broadcasting system of the present invention. FIG. 3b is a block diagram of a voice synthesis model of the video-audio broadcasting system of the present invention acquiring characteristic parameters with a pre-set voice number. FIG. 3c is a block diagram of a voice synthesis model of the video-voice broadcasting system of the present invention acquiring feature parameters with customized voice. Figure 4 is a flowchart of the video and audio broadcasting system of the present invention performing video and audio broadcasting using a pre-set character menu. Figure 5 is a flowchart of the video-audio broadcasting system of the present invention performing video-audio broadcasting using customized voice. FIG. 6 is a figure showing an embodiment of the user interface of the video-audio broadcasting system of the present invention. FIG. 7 is a figure showing another embodiment of the user interface of the video-audio broadcasting system of the present invention. Specific details for implementing the invention

[0019] To explain the technical details of the present invention in detail, diagrams are incorporated into the detailed description of the embodiments below.

[0020] First, refer to FIG. 1, which shows the system back end of the video audio broadcasting system as a structural diagram of the video audio broadcasting system of the present invention. The system back end (2) of the video audio broadcasting system of the present invention enables at least one user side (1) to connect and log in with at least one streaming host side (3), and the system back end (2) includes at least a reception waiting module (21), a cloud storage server (22), and a model update server (23). The system back end (2) receives data transmitted by the user side (1) through the reception waiting module (21), and this data includes, but is not limited to, the streaming host number of the streaming host side (3) to be transmitted, the user number of the user side (1), an input text message, and one of a preset voice number and a customized voice. The cloud storage server (22) of the system back end (2) records one of a preset voice number and a customized voice, and generates an effect voice according to the preset voice number, one of the customized voice, and an input text message through a voice synthesis model (4), so that the streaming host side image and the effect voice are included in the video and audio broadcast provided by the streaming host side corresponding to the streaming host number.

[0021] Refer to FIG. 2a together with FIG. 1. FIG. 2a is a more detailed structural diagram of the video-audio broadcasting system of the present invention. In one embodiment of the present invention, the system back end (2) of the video-audio broadcasting system receives a connection from one or more user sides (1). The user side (1) includes a character menu (11), an input module (12), a user interface (13), and an input / output module (14). The user can select the character menu (11) by operating the user interface (13) of the user side (1). The character menu (11) is a selectable menu and includes a plurality of pre-set character numbers of pre-set characters, and the voice of the pre-set character is used as a reference voice provided to train the voice synthesis model (4) to generate effect voice. The user inputs a text message that transmits voice through the input module (12) and then operates the user interface (13) of the user side (1) to transmit the pre-set character number and the text message to the receiving waiting module (21) of the system back end (2) through the input / output module (14).

[0022] The system back end (2) allows connection login between at least one user side (1) and at least one streaming host side (3), and receives data transmitted by the user side (1) through a receiving waiting module (21). This data includes, but is not limited to, the streaming host number of the streaming host side (3) to be transmitted, the user number of the user side (1), a preset voice number, and an input text message. The cloud storage server (22) of the system back end (2) records the preset voice number and transmits the preset voice number and the input text message to the designated streaming host side (3) via the streaming host number to be transmitted. Here, the streaming host side (3) includes an input / output module (32), a user interface (33), and a voice mixing device (34).

[0023] In this embodiment of the present invention, a speech synthesis model (4) is placed on each streaming host side (3), and the system back end (2) includes a model update server (23). The input / output module (32) of the streaming host side (3) receives a situation setting voice number and an input text message transmitted by the system back end (2), and then transmits them to the speech synthesis model (4). The speech synthesis model (4) generates multiple feature parameters by training reference voices of various characters according to an artificial neural network. By selecting and using various feature parameters, the speech synthesis model (4) can generate effect voices corresponding to the input text message. Next, the speech synthesis model (4) will be further described with reference to FIGS. 3a and 3B. In a specific embodiment of the present invention, when the system back end (2) completes training of the speech synthesis model (4), the model update server (23) updates the speech synthesis model (4) of each streaming host side (3) and downloads a plurality of feature parameters generated through the most recent training to each speech synthesis model (4), enabling the speech synthesis model (4) to generate a corresponding effect voice using the plurality of feature parameters generated through the most recent training.

[0024] For example, the voice synthesis model (4) is trained in advance using a reference voice featuring Donald Duck or Mickey Mouse as a character. Then, the voice synthesis model (4) selects the corresponding Donald Duck as a character according to the situation setting voice number and synthesizes the voice according to the input text message to synthesize an effect voice in which Donald Duck speaks the text message. The streaming host operates the user interface (33) to enter the voice mixing device (34), combines the effect voice with the streaming host image, and then performs video-audio broadcasting to the user side (1) through the input / output module (32). The streaming host side image displays the live screen of the streaming host. Refer to Fig. 6 together. The video-audio broadcast includes an effect voice broadcast (511) in which the selected character speaks the text message on the streaming host live screen (51). The image (R2) of the selected character and the user number are additionally displayed so that both the streaming host and the user can see which user number the user belongs to, and the effect voice broadcast (511) in which the selected character speaks the text message is transmitted.

[0025] Referring further to FIG. 2a, in a specific embodiment of the present invention, the system back end (2) of the video audio broadcasting system receives a connection from one or more user sides (1). The user can select a character menu (11) by operating the user interface (13) of the user side (1). One option of the character menu (11) is to allow the user to upload a custom voice file. The custom voice file is not limited to MP3 or MP4 and is available in various video audio file formats. For example, if the user intends to generate an effect voice using the voice of BTS's Jungkook and send it to a streaming host, but Jungkook's voice is not a pre-set character in the character menu (11) and therefore does not exist in the pre-set character number, the user can upload the Jungkook voice file themselves. The user can send a text message of voice through the input module (12) and then operate the user interface (13) of the user side (1) to send the custom voice file and the text message to the receiving waiting module (21) of the system back end (2) through the input / output module (14). After that, the voice synthesis model (4) of the present invention can be used to generate an effect voice to simulate the sound of Jeong-guk reading the text message.

[0026] In a further embodiment of the present invention, the system back end (2) allows connection login between at least one user side (1) and at least one streaming host side (3) and receives data transmitted from the user side (1) through a receiving waiting module (21). This data includes, but is not limited to, the streaming host number of the streaming host side (3) to be transmitted, the user number of the user side (1), a custom voice file, and an input text message. The cloud storage server (22) of the system back end (2) records the custom voice file and transmits the custom voice file and the input text message to a designated streaming host side (3) through the streaming host side (3) to be transmitted.

[0027] The input / output module (32) of the streaming host side (3) receives a customized voice file and an input text message transmitted by the system back end (2) and then transmits them to the voice synthesis model (4). The voice synthesis model (4) trains reference voices of various characters according to an artificial neural network to generate multiple feature parameters. By selecting and using various feature parameters, the voice synthesis model (4) can generate effect voice corresponding to the input text message. Next, with reference to FIG. 3c, the method by which the voice synthesis model (4) processes the customized voice is explained. The streaming host enters the voice mixing device (34) through the user interface (33), combines the customized effect voice with the streaming host image, and then performs video-audio broadcasting to the user side (1) through the input / output module (32), and the streaming host side image represents the live screen of the streaming host. Refer to FIG. 6 together. The above video and audio broadcast includes an effect audio broadcast (511) in which a character selected from the streaming host live screen (51) speaks a text message, additionally displays a user number so that both the streaming host and the user can see which user number the user is, and transmits an effect audio broadcast (511) in which a text message is spoken.

[0028] Refer to FIG. 2b. FIG. 2b shows another embodiment of the video-audio broadcasting system of the present invention. The user side (1) corresponds to that described in FIG. 2a, and the voice synthesis model (4) is placed in the system back end (2). The voice synthesis model (4) connects to a cloud storage server (22) to access a preset voice number and input text message to perform voice synthesis, generate effect voice, and transmit it back to the cloud storage server (22). FIG. 3a and 3B below further describe the voice synthesis model (4). The cloud storage server (22) transmits the effect voice and input text message to a designated streaming host side (3) via the streaming host number to be transmitted. The streaming host uses a user interface (33) to enter a voice mixing device (34), combines the effect voice and the streaming host side image, and performs video-audio broadcasting to the user side (1) through an input / output module (32), wherein the streaming host side image represents the live screen of the streaming host.

[0029] Refer to FIG. 2b. In a more specific embodiment of the present invention, the user side (1) corresponds to FIG. 2a, and the voice synthesis model (4) is located at the system back end (2). The voice synthesis model (4) connects to a cloud storage server (22) to access a customized voice file, synthesizes the input text message with the voice to generate an effect voice, and then transmits it to the cloud storage server (22). FIG. 3c below further explains how this voice synthesis model (4) processes the customized voice. The cloud storage server (22) transmits the effect voice to a designated streaming host side (3) via the streaming host number to be transmitted. The streaming host enters the voice mixing device (34) through the user interface (33), combines the effect voice with the streaming host image, and then performs video-audio broadcasting to the user side (1) through the input / output module (32), and the streaming host side image represents the live screen of the streaming host.

[0030] Refer to FIG. 2c. FIG. 2c is a diagram showing an embodiment in which the user side of the video-audio broadcasting system of the present invention receives effect audio through a cloud (24). The system back end (2) of the video-audio broadcasting system receives a connection from one or more user sides (1). The user can select a character menu (11) by operating the user interface (13) of the user side (1). The character menu (11) is a selectable menu and includes a plurality of preset character numbers of preset characters.

[0031] The user inputs a text message that transmits voice through the input module (12) and then operates the user interface (33) on the user side (1) to transmit the preset character number and text message to the cloud server (24) of the system back end (2) through the input / output module (14).

[0032] The cloud server (24) of the system back end (2) receives a pre-set character number uploaded by the user side (1) and an input text message, and then the speech synthesis model (4) separates feature parameters corresponding to the pre-set character number and implements speech synthesis for the input text message to generate an effect voice. Figures 3a and 3B below further explain the speech synthesis model (4). In a specific embodiment of the present invention, after the system back end (2) completes training of the speech synthesis model (4), the model update server (23) can update the speech synthesis model (4) and download a plurality of feature parameters generated through the most recent training to the speech synthesis model (4), so that the speech synthesis model (4) can use the plurality of feature parameters generated through the most recent training to generate a corresponding effect voice.

[0033] Referring further to FIG. 3c. In a more specific embodiment of the present invention, the system back end (2) of the video audio broadcasting system receives a connection from one or more user sides (1). The user can select a character menu (11) by operating the user interface (13) of the user side (1). The character menu (11) provides an option for the user to upload a custom voice file. The custom voice file is not limited to MP3 or MP4 and is in various video audio file formats. The user inputs a text message of the transmitted voice through the input module (12) and then operates the user interface (33) of the user side (1) to transmit the custom voice file and the text message to the cloud server (24) of the system back end (2) through the input / output module (14).

[0034] The system back end (2) cloud server (24) receives a custom voice file uploaded by the user side (1) and an input text message, and the voice synthesis model (4) connects to the cloud storage server (22) to access the custom voice file and performs voice synthesis with the input text message to generate an effect voice. Figure 3c below further explains how the voice synthesis model (4) processes the custom effect voice. Another user can obtain the effect voice through the cloud server (24) of the system back end (2) and broadcast and use it on the user side (1).

[0035] Refer to the block diagram shown in FIG. 3a for the structure of the speech synthesis model (4) of the present invention. In a specific embodiment of the present invention, the constituent blocks of the speech synthesis model (4) include phoneme embeddings, a phoneme encoder (41) according to the speaker's situation, a differentiation adapter (42), and a Mel decoder (43) according to the speaker's situation. The speaker's situation in the phoneme encoder (41) and the Mel decoder (43) is expressed by a plurality of feature parameters separated according to various characters through speech feature extraction (44). The present invention trains the speech synthesis model (4) and its speech feature extraction (44) using reference voices of various characters. The speech synthesis model (4) may be trained based on a supervised or unsupervised artificial neural network. After training, the speech feature extraction (44) of the speech synthesis model (4) may separate a plurality of corresponding feature parameters by designating a character according to a preset character number, or separate a plurality of feature parameters selected and used according to the customized voice of a non-preset character.

[0036] Each constituent block of the speech synthesis model (4) of the present invention can be implemented according to the content presented in the paper published by Yihan Wu et al. in April 2022. The title of the paper is 'AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios', and the public website is https: / / www.isca-speech.org / archive / interspeech_2022 / index.html. The paper presents speech synthesis technology and allows the speech synthesis model to be trained using a reference speech. Such a speech synthesis model is trained based on an artificial neural network.

[0037] Referring further to FIG. 3a. In a specific embodiment of the present invention, the present invention trains a speech synthesis model (4) using reference voices of various characters. The speech synthesis model (4) may be trained based on a supervised or unsupervised artificial neural network. After training, the speech feature separation (44) of the speech synthesis model (4) can separate a plurality of feature parameters according to the reference voices of various characters.

[0038] Refer to FIG. 3b. When the system of the present invention uses a preset voice number and an input text message, the phonemes of the preset voice number and the input text message are input into the voice synthesis model (4) of the present invention, and the characteristic parameters of the character corresponding to the specified voice number are used to indicate the speaker's situation in the phoneme encoder (41) and Mel decoder (43) through the pre-trained voice feature separation (44). The pre-trained voice synthesis model (4) synthesizes the phonemes of the input text message and the characteristic parameter encoding by implementing phoneme encoding through the phoneme encoder (41) following phoneme embedding according to the phonemes of the text message. Then, voice track adaptability is adjusted through the differentiation adapter (42), and finally, an effect voice is generated by decoding according to the speaker's situation through the Mel decoder (43). This effect voice is the voice of the character corresponding to the specified voice number speaking the text message.

[0039] Refer to FIG. 3c. When a user inputs a customized voice file and an input text message into the voice synthesis model (4) of the present invention, the voice synthesis model (4) implements preprocessing and interpolation operation feature separation for the customized voice through voice feature separation (44), and selects and uses one or more feature parameters generated by training through the preprocessing and interpolation operation customized voice to indicate the speaker's situation in the phoneme encoder (41) and Mel decoder (43). The pre-trained voice synthesis model (4) synthesizes the phonemes of the input text message and feature parameter encoding by implementing phoneme encoding through the phoneme encoder (41) after undergoing phoneme embedding according to the phonemes of the text message. Then, it adjusts the voice track adaptability through the differentiation adapter (42) and finally decodes according to the speaker's situation through the Mel decoder (43) to generate an effect voice. This effect voice is the voice of the customized voice speaker speaking the text message.

[0040] Next, refer to FIG. 4. FIG. 2a may also be referenced at this time. FIG. 4 is a flowchart of a specific embodiment of performing video and audio broadcasting using a preset character menu of the present invention. The user receives one preset voice number of the character menu (11) and one input message, for example, a text message, from the input module (12) through the user interface (13) on the user side (1) (S11), and the input / output module (14) outputs the streaming host number, user number, the preset voice number, and the input message to the system back end (S12). These two steps are completed on the user side (1).

[0041] Then, the receiving module (21) of the system back end (2) receives the streaming host number, user number, the preset voice number, and the input message (S13), and the cloud storage server (22) transmits the received preset voice number and the input message to the voice synthesis model (4) of one streaming host side (3) (S14). These two steps are completed in the system back end (2).

[0042] Then, the voice synthesis model (4) of the streaming host side (3) displays the situation of the speaker of the phoneme encoder (41) and Mel decoder (43) as characteristic parameters of the character corresponding to the preset voice number, and generates effect voice by implementing voice synthesis with the input message (S15). The streaming host side (3) mixes the effect voice through the voice mixing device (34) (S18), and the voice mixing device outputs the effect voice along with a video screen, for example, a selected gift. The effect voice after voice mixing is broadcast video audio through the input / output module (32) (S19). These three steps are completed on the streaming host side (3).

[0043] When the voice synthesis model (4) receives a preset voice number and the preset voice number represents the situation of a speaker as a characteristic parameter of a corresponding character and implements voice synthesis with the input message to generate an effect voice, if the voice generation parameter is optimized, the voice synthesis model (4) updates the optimized characteristic parameter in synchronization with the model update server (23) of the system back end (2) (S16), and the model update server (23) updates the characteristic parameter (S17) to make it advantageous to implement the corresponding characteristic parameter next time (S17).

[0044] Next, refer to FIG. 5. FIG. 2a may also be referenced at this time. FIG. 5 is a flowchart of another specific embodiment in which the video-audio broadcasting system of the present invention performs video-audio broadcasting using customized voice. The user receives a customized voice file of the character menu (11) and one input message of the input module (12) through the user interface (13) on the user side (1) (S21), and the input / output module (14) outputs the streaming host number, user number, customized voice file, and input message to the system back end (S22). These two steps are completed on the user side (1).

[0045] Then, the receiving module (21) of the system back end (2) receives the streaming host number, user number, custom voice number, and input message (S23), and the cloud storage server (22) transmits the received custom voice file and input message to the voice synthesis model (4) of one streaming host side (3) (S24). These two steps are completed in the system back end (2).

[0046] Then, the speech synthesis model (4) on the streaming host side (3) implements preprocessing and interpolation operation feature separation on the customized speech through speech feature separation (44), and selects and uses one or more feature parameters generated by training through the preprocessing and interpolation operation customized speech to indicate the speaker's situation in the phoneme encoder (41) and Mel decoder (43). The pre-trained speech synthesis model (4) synthesizes the phonemes of the input text message and feature parameter encoding by implementing phoneme encoding through the phoneme encoder (41) after phoneme embedding according to the phonemes of the text message. Then, it adjusts the voice track adaptability through the differentiation adapter (42) and finally generates effect speech by decoding according to the speaker's situation through the Mel decoder (43) (S25), mixes the effect speech through the voice mixing device (34) (S26), and the voice mixing device outputs the effect speech along with a video screen, for example, a selected gift. The effect voice after voice mixing is broadcast via video-audio through the input / output module (32) (S27). These three steps are completed on the streaming host side (3).

[0047] In a specific embodiment of the present invention, the feature parameter selected through the customized voice file is not stored as a voice preset in the voice synthesis model (4), and the voice feature separation step must be performed again every time a customized effect voice is implemented.

[0048] Next, look at FIG. 6. FIG. 6 is a diagram showing a specific embodiment of the user interface of the video and audio broadcasting system of the present invention. The screen is divided into three areas: an auxiliary list (52) on the left, a streaming host live screen (51) in the middle, and a chat room screen (53) on the right. The streaming host live screen (51) in the middle plays an interactive screen of the streaming host live broadcast and an effect voice broadcast (511) of the effect voice of the present invention, and the effect voice broadcast (511) includes the effect voice and the character corresponding to the effect voice, an image of a character among the selected menus and the user number of the effect voice, and a corresponding gift picture.

[0049] The chat room screen (53) on the right displays the content of multiple users conversing with the streaming host via text. In a specific embodiment, text of the transmitted effect voice broadcast (411) is also displayed on the chat room screen (53), and a special color, appearance, etc. corresponding to it is combined. Below the chat room screen (53), a character menu (531) and a text input module (532) are included, and FIG. 6 shows only the representation of the character menu (531) in a specific embodiment. The character menu may be designed as a menu, a pop-up window, etc. For example, the character menu (531) may include multiple profile images (R1, R2). This facilitates user identification and selection.

[0050] The auxiliary list (52) on the left stores a list of related interactive functions, and the menu list (54) providing various functions at the top left and the daily task list (521) allow you to design tasks to be accomplished every day, such as sending five conversations to increase the conversationality between the user and the streaming host, and hosting an event at a specific festival and posting it on the event list (522). The bottom left is a list of other streaming hosts (523), and you can click the screen to switch to the live broadcast room of a different streaming host and converse.

[0051] Refer to FIG. 7. FIG. 7 is a diagram showing another embodiment of the user interface of the video-audio broadcasting system of the present invention, for example, message transmission between mobile phone users. The screen shown in this embodiment is a chat screen (63). On the screen, the user is located on the right side of the chat screen (63), and the target user is located on the left side of the chat screen (63). As can be seen in FIG. 7, in a specific embodiment, the target user sends the text message "I love you," and the user image (64) is displayed on the right side of the screen. The user selects the character of R2 through the character menu (631) and simultaneously sends the text "I love you" through the text input module (632). Then, the video-audio broadcasting system of the present invention performs effect voice output with a preset voice number and input text message to broadcast effect voice (611) to the target user. Explanation of the symbols

[0052] 1 User side 11 Character Menu 12 Input Modules 13 User Interface 14 Input / Output Module 2 System Backend 21 Receive waiting module 22 Cloud Storage Servers 23 Model Update Server 24 cloud storage servers 3. Streaming host side 32 I / O Module 33 User Interface 34 Voice Mixer 4 Speech Synthesis Models 41 Phoneme Encoder 42 Differentiated Adapter 43 Mel decoder 5 User Interface 51 Streaming Host Live Screen 511 effect voice broadcast 52 Auxiliary List 521 Daily Task List 522 Activity List 523 Other Streaming Host List 53 Chat room screen 531 Character Menu 532 Character Input Module 54 Menu List 6 User Interface 611 effect voice broadcast 63 Conversation Screen 631 Character Menu 632 Character Input Module 64 User Images

Claims

Claim 1 delete Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 A video audio broadcasting system comprising a voice synthesis model including a plurality of feature parameters, and a system back end that allows connection login between at least one user side and at least one streaming host side, and receives a streaming host number, a user number, a customized voice, and an input message transmitted from at least one user side, wherein the voice synthesis model selects at least one feature parameter associated with the result of voice feature separation from a plurality of feature parameters based on the result of voice feature separation for the customized voice, and generates effect voice according to the selected at least one feature parameter and the input message, thereby including a streaming host side image, a user number, and effect voice in the video audio broadcast provided by the streaming host side. Claim 12 In paragraph 11, the voice synthesis model is a video-audio broadcasting system placed on the system back end or streaming host side. Claim 13 In Clause 11, the above input message is a text message, a video-audio broadcasting system. Claim 14 A video audio broadcasting system comprising a voice synthesis model including multiple characteristic parameters, and a system back end that allows connection login of at least two user sides, receives a user number and an input message from a speaking user side, and receives one of a customized voice and a preset voice number, wherein when the system back end receives a preset voice number from the speaking user side, the voice synthesis model corresponds the preset voice number to at least a portion of the characteristic parameters, generates a first effect voice according to the corresponding characteristic parameters and the input message, and causes the user side corresponding to the user number to receive the input message and the first effect voice, wherein when the system back end receives a customized voice from the speaking user side, the voice synthesis model selects one or more characteristic parameters associated with the result of voice feature separation from a plurality of characteristic parameters based on the result of voice feature separation for the customized voice, and generates a second effect voice according to the one or more characteristic parameters and the input message, and causes the user side corresponding to the user number to receive the input message and the second effect voice. Claim 15 In Clause 14, the voice synthesis model is a video-audio broadcasting system placed at the back end of the system. Claim 16 In Clause 14, the above input message is a text message video-audio broadcasting system Claim 17 A video and audio broadcasting method comprising: a video and audio broadcast provided by at least one streaming host side having a streaming host side image; establishing a voice synthesis model including a plurality of feature parameters; receiving a streaming host number, a user number, and an input message from at least one user side, and receiving one of a customized voice and a preset voice number; when the preset voice number is received, the voice synthesis model corresponds the preset voice number to at least a portion of the feature parameters and generates a first effect voice according to the corresponding feature parameters and the input message, thereby including the image of the streaming host side corresponding to the streaming host number and the first effect voice in the video and audio broadcast provided by the streaming host side corresponding to the streaming host number; and when the customized voice is received, the voice synthesis model selects one or more feature parameters associated with the result of voice feature separation from a plurality of feature parameters based on the result of voice feature separation for the customized voice, and generates a second effect voice according to the one or more feature parameters and the input message, thereby including the image of the streaming host side corresponding to the streaming host number and the effect voice in the video and audio broadcast provided by the streaming host side corresponding to the streaming host number. method. Claim 18 In claim 17, the above-mentioned voice synthesis model is a video-audio broadcasting method in which the system back end or the streaming host side corresponding to the streaming host number is placed. Claim 19 In Clause 17, the above input message is a video-audio broadcasting method in which the text message is a text message. Claim 20 A video audio broadcasting method comprising establishing a voice synthesis model including a plurality of feature parameters, receiving a user number and an input message from at least one user side, and receiving one of a customized voice and a preset voice number; when the preset voice number is received, the voice synthesis model corresponds the preset voice number to one of the plurality of feature parameters and generates a first effect voice according to the corresponding feature parameter and the input message, thereby enabling the user side corresponding to the user number to receive the input message and the first effect voice; and when the customized voice is received from a speaking user side, the voice synthesis model selects one or more feature parameters associated with the result of voice feature separation from a plurality of feature parameters based on the result of voice feature separation for the customized voice, and generates a second effect voice according to the one or more feature parameters and the input message, thereby enabling the user side corresponding to the user number to receive the input message and the second effect voice.