Information processing system and information processing program

The system enhances the appeal of translated content by generating unique voice models for broadcasters, preserving their voice characteristics and speaking style, addressing the loss of appeal in cross-language content translation.

JP7894186B1Active Publication Date: 2026-07-23ANOTHERBALL PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ANOTHERBALL PTE LTD
Filing Date
2025-12-26
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing content translation technologies fail to preserve the voice characteristics and speaking style of broadcasters when translating content across different languages, leading to a loss of appeal in the translated content.

Method used

An information processing system that generates unique voice models for each broadcaster using machine learning, converting their voice into text, and then into audio in the viewer's language while reproducing the broadcaster's voice characteristics and speaking style.

Benefits of technology

Enhances the appeal of translated content by allowing viewers to experience the original voice and speaking style of the broadcaster, even when the content is streamed globally across different languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894186000001_ABST
    Figure 0007894186000001_ABST
Patent Text Reader

Abstract

To further enhance the appeal of translated content. [Solution] The information processing system S comprises a voice model DB372, a source language text conversion unit 115, a translated language text conversion unit 116, and a voice data generation unit 212. The voice model DB372 stores a unique voice model for each broadcaster, capable of generating voice data corresponding to the broadcaster based on text data. The source language text conversion unit 115 converts the broadcaster's voice into text data in a first language. The translated language text conversion unit 116 converts the text data in the first language converted by the source language text conversion unit 115 into text data in a second language. The voice data generation unit 212 generates voice data in a second language corresponding to the broadcaster using the text data in the second language converted by the translated language text conversion unit 116 and a unique voice model corresponding to the broadcaster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system and an information processing program.

Background Art

[0002] In recent years, using electronic devices such as personal computers and smartphones, users have become content providers and provided content to other users who are viewers. For example, content is provided by the distributor himself / herself as the provider or by delivering conversations or the like in real time using an avatar. An example of a related technology is disclosed in, for example, Citation 1. In the technology disclosed in this Citation 1, overseas content is translated into the native language of the viewer and then provided to the viewer.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By translating with a general technology such as the technology disclosed in Patent Document 1 described above, it becomes possible to enjoy content without overly worrying about the language provided. However, from the viewpoint of enjoying content, it is desired to further improve the interestingness.

[0005] The present invention has been made in view of such a situation, and an object thereof is to further improve the interestingness of the content to be translated.

Means for Solving the Problems

[0006] To achieve the above objective, an information processing system according to one aspect of the present invention is: A storage means that stores a unique voice model for each broadcaster, capable of generating voice data corresponding to the broadcaster based on text data, A first conversion means for converting the broadcaster's voice into text data in a first language, A second conversion means converts the text data in the first language converted by the first conversion means into text data in the second language, A generation means that generates audio data in the second language corresponding to the distributor using the second language text data converted by the second conversion means and a proprietary voice model corresponding to the distributor, It is characterized by being equipped with [the following features]. [Effects of the Invention]

[0007] According to the present invention, the appeal of translated content can be further enhanced. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing the overall configuration of information processing system S. [Figure 2] This is a conceptual diagram illustrating the concepts behind the "voice model generation process" and the "delivery translation process." [Figure 3] This block diagram shows the hardware configuration of the distribution terminal 10, the viewing terminal 20, and the server device 30. [Figure 4] This is a block diagram showing the functional configuration of the distribution terminal 10. [Figure 5] This is a block diagram showing the functional configuration of the viewing terminal 20. [Figure 6] This block shows the functional configuration of the server device 30. [Figure 7] This is a flowchart showing the workflow of the voice model generation process. [Figure 8] This is a flowchart showing the workflow of the delivery translation process. [Figure 9] This figure shows an example of a screen being delivered by the mechanism of the present invention. [Modes for carrying out the invention]

[0009] Embodiments of the present invention will be described below with reference to the drawings.

[0010] [Overall System Configuration] First, the overall configuration of this embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing the overall configuration of the information processing system S according to this embodiment. As shown in Figure 1, the information processing system S consists of a distribution terminal 10, n (where n is any integer greater than or equal to 1) viewing terminals 20, a server device 30, and a network N.

[0011] In this case, the broadcasting terminal 10 and the viewing terminal 20 are each implemented by, for example, a smartphone, a tablet, or a personal computer. The server device 30 is implemented by, for example, one or more server devices, or a cloud server that virtualizes multiple server devices. Furthermore, the network N is implemented by, for example, a LAN (Local Area Network), the Internet, a mobile phone network, or a network that combines these.

[0012] In the diagram, n viewing terminals 20 are shown as viewing terminal 20a, viewing terminal 20b, and viewing terminal 20n. However, in the following explanation, when these n viewing terminals 20 are described without distinction, some of the reference numerals will be omitted and they will simply be referred to as "viewing terminal 20".

[0013] In this configuration, the information processing system S realizes a client-server system by having the client-side terminal 10 and n viewing-side terminals 20 communicate with the server-side device 30 via a network N. In this embodiment, for the sake of example, it is assumed that the server device 30 provides an application function to realize the distribution of content called so-called live distribution. In this case, a user of the distribution-side terminal 10 (hereinafter referred to as the “distributor”) creates a live video, and via the server device 30, distributes this live video in real time to each user of n viewing-side terminals 20 (hereinafter referred to as the “viewers”).

[0014] Note that it is assumed that this live video is not a video of the viewer himself / herself, but is performed by an avatar, which is a virtual character reflecting the movement of the viewer. However, this embodiment is not limited to the distribution of a live video using such an avatar. For example, it may be a distribution using a live video of the distributor himself / herself, or a distribution in which the distributor and the viewer simultaneously view content such as a video shot in the past. That is, this embodiment can be applied to the entire system for providing content between users.

[0015] [Generation of Voice Model by TTS Server] As shown in the figure, a TTS (Text-to-Speech) server, which is a server that generates a TTS model, is also connected to the information processing system S. The TTS model is a model that generates voice data like a natural human voice from text data. Here, in this embodiment, by performing machine learning using the TTS server, it is possible to output a voice that comprehensively reproduces elements (for example, pitch (height of voice), tone (quality of voice), rhythm (rhythm of voice), intonation (intonation of voice), etc.) that are characteristics of the voice and speaking style of a specific person to be processed.

[0016] Specifically, first, after performing Fourier transform on the voice data of the target person, and then applying a mel filter bank, a mel spectrogram, which is a feature quantity of the voice of the target person, is extracted. Then, this Mel spectrogram and its corresponding text data are used as a pair to train a deep learning model such as a CNN (Convolutional Neural Network) or Transformer. In other words, supervised learning is performed with the text data as input and the Mel spectrogram as the training label, resulting in the generation of a TTS model that has learned the speech of the target person. When arbitrary text is input to this trained TTS model, a Mel spectrogram of the target person matching that text is generated. Furthermore, by inputting this Mel spectrogram into a vocoder, natural-sounding audio data can be generated.

[0017] Furthermore, in this embodiment, instead of using a typical TTS model, a general-purpose model capable of outputting multilingual audio is prepared in advance, and then fine-tuned (additional learning) is performed using data specific to the broadcaster. As a concrete example of the method, we first prepare a basic model that, when text data in a certain language (for example, English) is input, can output audio data in the same language (in this case, English) that corresponds to that text data. Similarly, we prepare a basic model like this for each language, such as multiple languages ​​(for example, Chinese and Spanish).

[0018] Next, each of the pre-prepared basic models is fine-tuned using the broadcaster's training audio data. This allows for the creation of a model that can output audio data similar to the broadcaster's voice for each language. By performing this process for each of the multiple broadcasters, a unique TTS model can be generated for each of the multiple broadcasters. In addition to preparing a basic model for each language, it is also possible to prepare different basic models for each gender, or to prepare basic models for different age groups such as adults and children, and then select the appropriate basic model for the broadcaster and fine-tune it.

[0019] Alternatively, another approach is to perform supervised learning on the target person's voice, as described above, on a TTS model that has been pre-trained in multiple languages ​​(at least two languages: the language used by the broadcaster and the language used by the viewers).

[0020] These methods generate a TTS model that, when text data in a language different from the source language (in this embodiment, the language in which viewers hear the broadcaster's voice during the broadcast) is input, is reproduced in the voice of the target person and is spoken in that other language. In the following, a TTS model capable of generating audio data that is reproduced in the voice of the target person and also spoken in another language will be referred to as a "voice model."

[0021] It should be noted that the above-described method for generating a speech model is merely one example, and the method for generating a speech model in this embodiment is not limited to the method described above.

[0022] [Outline of this embodiment] In this embodiment, a live video created by a broadcaster, who is the provider of the live video, is distributed to viewers, who are the recipients of the live video. By the way, when such a streaming service is deployed globally across national borders, the language used by the streamer and the language used by the viewers may differ, which can be a limitation in enjoying the content (in this case, live video). Therefore, in such cases, it is desirable to implement a language translation function.

[0023] However, simply transcribing a streamer's voice into text and then outputting a translated audio file fails to convey the streamer's voice characteristics and speaking style to the audience. This can lead to a loss of appeal in the streamed content. In consideration of these issues, this embodiment applies a voice model, as described above, that is capable of generating audio data that is reproduced in the voice of the target person and spoken in another language, to the translation of the broadcaster's voice. This not only translates the language used by the broadcaster into the language used by the viewers, but also reproduces the broadcaster's voice characteristics and speaking style in the translated audio. This means that when a streaming service is expanded globally, content can be delivered to viewers who speak different languages ​​without compromising the broadcaster's appeal.

[0024] Let's now explain the outline of this embodiment in more detail. Figure 2 is a conceptual diagram illustrating the concepts of the "voice model generation process" and the "distribution translation process" in this embodiment. Here, the voice model generation process is a series of processes that generate individual voice models for each broadcaster by performing machine learning using the broadcaster's voice as training data. The broadcast translation process is a series of processes that reproduce the characteristics of the broadcaster's voice and speaking style when outputting audio translated into the language used by the viewers.

[0025] First, please refer to Figure 2 to explain the speech model generation process. In the speech model generation process, the first step is to obtain training data for the target broadcaster (corresponding to (1) in the diagram). This training data consists of the target broadcaster's voice data and its corresponding text data. Next, by having the TTS server perform machine learning as described above, it generates individual voice models for each broadcaster that correspond to the target broadcaster (corresponding to (2) in the diagram).

[0026] Then, the speech model generated by the TTS server is obtained (corresponding to (3) in the diagram). Furthermore, the acquired voice model is linked to and stored in memory of the target broadcaster (corresponding to (4) in the diagram). In this way, the voice model generation process makes it possible to prepare individual voice models for each broadcaster that correspond to the target broadcaster.

[0027] Next, referring to Figure 2, we will explain the delivery translation process. Hereafter, the language used by the broadcaster will be referred to as the "original language," while the language used by the viewers will be referred to as the "translated language." The original language corresponds to the "first language" in the present invention, and the translated language corresponds to the "second language" in the present invention. In the streaming translation process, the first step is to obtain the original language audio from the content streamed by the broadcaster in the original language (in this case, live video) (corresponding to (1) in the diagram). Next, the acquired source language audio is converted into source language text (corresponding to (2) in the diagram).

[0028] Furthermore, the converted source text is converted into a translated text, which is the text of the translated words (corresponding to (3) in the diagram). Then, the translated text after conversion is input into the speech model corresponding to the broadcaster currently streaming, which was generated in the speech model generation process described above (corresponding to (4) in the diagram).

[0029] This generates audio data that is reproduced in the streamer's voice and spoken using the translated words. This audio data is then output to the viewer via a vocoder (corresponding to (5) in the diagram). This allows the viewer to watch the stream while listening to audio that is reproduced in the streamer's voice and spoken using the translated words. Therefore, the viewer can enjoy the streamer's voice characteristics and speaking style, which are major attractions of the streamed content. In other words, this embodiment makes it possible to solve the problem of "further improving the appeal of translated content." The above is an overview of this embodiment.

[0030] [Hardware configuration] Next, the hardware configuration of this embodiment will be described with reference to Figure 3. Figure 3 is a block diagram showing the hardware configuration of the distribution terminal 10, the viewing terminal 20, and the server device 30 according to an embodiment of the present invention. In the figure, reference numerals corresponding to the hardware of the distribution terminal 10 are written without parentheses, while reference numerals corresponding to the hardware configuration of the viewing terminal 20 and the hardware of the server device 30 are written with parentheses.

[0031] First, let's describe the hardware configuration of the distribution terminal 10. As shown in Figure 3, the distribution terminal 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a bus 14, an input unit 15, an output unit 16, a storage unit 17, a communication unit 18, and a drive 19.

[0032] The CPU 11 executes various processes according to the program recorded in the ROM 12 or the program loaded from the storage unit 17 into the RAM 13.

[0033] RAM13 also stores data and other information necessary for the CPU11 to perform various processes.

[0034] The CPU 11, ROM 12, and RAM 13 are interconnected via a bus 14. The input unit 15, output unit 16, storage unit 17, communication unit 18, and drive 19 are also connected to this bus 14.

[0035] The input unit 15 consists of various buttons, a touch panel, or a microphone, and inputs various information according to the user's instructions. The input unit 15 may also be implemented using an input device such as a keyboard or mouse, independent of the main unit that houses the other parts of the distribution terminal 10. Furthermore, the input unit 15 includes a camera for capturing images of the user, and inputs the captured still images or video image data.

[0036] The output unit 16 outputs image data and music data to a display, speaker, etc. The image data and music data output by the output unit 16 are output from the display, speaker, etc., in a way that the user can recognize them as images and music.

[0037] The memory unit 17 consists of semiconductor memory such as an SSD (Solid State Drive), HDD (Hard Disk Drive), or DRAM (Dynamic Random Access Memory), and stores various types of data.

[0038] The communication unit 18 enables communication with other devices via a network. For example, the communication unit 18 communicates with the distribution terminal 10 via the network N.

[0039] Drive 19 is provided as needed. A removable media 100, consisting of a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is appropriately mounted on drive 19. Various data, such as programs for running the game and image data, are stored on the removable media 100. The programs and various data, such as image data, read from the removable media 100 by drive 19 are installed in the storage unit 17 as needed. Please note that the above hardware configuration is merely a basic configuration, and it is possible to omit some hardware, add additional hardware, or change the hardware implementation.

[0040] Next, the hardware configuration of the viewing terminal 20 will be described. As shown in Figure 3, the viewing terminal 20 includes a CPU 21, ROM 22, RAM 23, bus 24, input unit 25, output unit 26, storage unit 27, communication unit 28, and drive 29. Each of these units has the same function as the units of the same name, differing only in their coding, that are present in the distribution terminal 10 described above. Therefore, redundant explanations will be omitted. Next, the hardware configuration of the server device 30 will be described. As shown in Figure 3, the server device 30 includes a CPU 31, a ROM 32, a RAM 33, a bus 34, an input unit 35, an output unit 36, a storage unit 37, a communication unit 38, and a drive 39. Each of these units has the same function as the units of the same name, differing only in their symbol, that are provided in the aforementioned distribution terminal 10. Therefore, redundant explanations will be omitted.

[0041] Furthermore, if the broadcasting terminal 10 and the viewing terminal 20 are configured as portable devices, the hardware components of the broadcasting terminal 10, along with input / output hardware such as a display and speakers, may be integrated into a single device.

[0042] [Functional configuration] Figure 4 is a block diagram showing the functional configuration of the distribution terminal 10. Figure 5 is a block diagram showing the functional configuration of the viewing terminal 20. Furthermore, Figure 6 is a block diagram showing the functional configuration of the server device 30. Figures 4 to 6 illustrate the functional blocks for executing the voice model generation process and the distribution translation process described above, with reference to Figure 2.

[0043] First, let's explain the functional configuration of the distribution terminal 10. When content provision control processing is executed, as shown in Figure 4, the following functions operate in the CPU 11: the distribution-side application control unit 111, the live distribution control unit 112, the learning data collection unit 113, the voice model generation instruction unit 114, the source language text conversion unit 115, and the translated language text conversion unit 116. Furthermore, one area of ​​the memory unit 17 is configured to include a distribution-side application data storage unit 171, a learning data storage unit 172, and a voice model storage unit 173.

[0044] The broadcasting application control unit 111 controls the entire process by which the broadcaster uses the broadcasting application functions to stream live video. The broadcasting application control unit 111 controls, for example, the process by which the broadcaster sets up their avatar, the avatar's accessories (e.g., clothes and accessories), and the background for the broadcast prior to broadcasting. These settings may require data from an external source, but they may also be completed entirely within the functions of the broadcasting application.

[0045] In this case, for example, an avatar can be generated by combining parts that make up the avatar, such as eyes, nose, mouth, and facial contours, with hairstyles and body types, and then adjusting colors, etc. Furthermore, clothing and backgrounds can be set by selecting from multiple options. In addition, the distribution application control unit 111 manages various information, such as the streamer's username (in this case, the name of the streamer displayed when streaming) and the password for logging into the server device 30.

[0046] The live streaming control unit 112 controls the streaming performed by the streamer using the app's functions. The live streaming control unit 112 generates a live video based on images representing, for example, an avatar, the avatar's accessories, or a background set by the user, and the streamer's voice collected by the microphone included in the input unit 15 of the streaming terminal 10. It then transmits this live video to the server device 30 in real time.

[0047] In this case, the avatar (and its accessories) may be a still image, but it is preferable that it be a video. To this end, the live streaming control unit 112 tracks the movements of the streamer's entire face, eyes, mouth, upper body, and other parts of their body in an image captured by the camera included in the input unit 15 of the streaming terminal 10. Then, by changing the movements of the avatar (and its accessories) in real time to correspond to the tracked movements, the avatar reflects the actual movements of the streamer. As a result, the avatar changes its facial expressions and posture just like the streamer, making it possible to create a live video that is more interesting to viewers. Furthermore, the live video being streamed is not only displayed on the receiving device 20, but is also displayed in real time on the streaming device 10 by the streaming application control unit 111.

[0048] The training data collection unit 113 collects training data to generate a speech model. To this end, the broadcaster performs a test mode prior to broadcasting. In test mode, the broadcaster speaks any phrase, or at least one of the phrases specified in test mode. The training data collection unit 113 collects the speech data of the voice spoken by the broadcaster in this test mode.

[0049] Furthermore, the learning data collection unit 113 also collects text data corresponding to this audio. Here, if the broadcaster speaks an arbitrary phrase, the source language text conversion unit 115 (described later) converts the audio data into source language text, thereby obtaining text data corresponding to the audio. On the other hand, if the broadcaster speaks a phrase specified in test mode, the text data of this specified phrase can be obtained as text data corresponding to the audio. The learning data collection unit 113 collects audio data and corresponding text data as pairs in this manner, and stores these pairs as learning data in the learning data storage unit 172.

[0050] The voice model generation instruction unit 114 instructs the TTS server to generate a voice model using the accumulated training data when the training data collection unit 113 has accumulated a predetermined amount of training data necessary for training. The method of generating the voice model by the TTS server is as described above in the section [Generation of voice model by TTS server]. This generates individual voice models corresponding to the broadcasters using the broadcasting terminal 10. This speech model is trained so that the features of the output speech data are as identical or similar as possible to the features of the speech data actually generated by the broadcaster. In this case, the main feature of the audio data is the waveform of the Mel spectrogram, which functions as a blueprint for the voice, and more specifically, the strength of each frequency component contained in that waveform. Other features of audio data include the fundamental frequency (FO), which indicates the pitch of the voice; the energy (Energy / Loudness), which indicates the volume of the voice; and the aperiodicity, which indicates the "hoarseness" and "breath" components of the voice. Furthermore, the speech model in this embodiment is a speech model that has been trained so that these features are as identical or similar as possible to the features of the speech data actually generated by the broadcaster.

[0051] The voice model generation instruction unit 114 then acquires this voice model and stores it in the voice model storage unit 173. The voice model generation instruction unit 114 also acquires this voice model and transmits it to the server device 30. The server device 30 stores the received voice model and associates it with the broadcaster using the broadcasting terminal 10. Furthermore, the training data collection unit 113 and the speech model generation instruction unit 114 may perform additional training data collection and issue instructions to the TTS server after the speech model has been completed. This allows for additional training of the completed speech model, further improving the accuracy of the speech model.

[0052] The source language text conversion unit 115 acquires the source language audio when the broadcaster starts streaming content (in this case, live video) in the source language. The source language text conversion unit 115 then converts the acquired source language audio into source language text. In this case, the source language text conversion unit 115 achieves this conversion by using a speech recognition engine that supports the source language. The source language text conversion unit 115 then outputs the converted source language text to the translated text conversion unit 116.

[0053] Furthermore, in test mode, when the broadcaster speaks an arbitrary phrase, the source language text conversion unit 115 also performs the process of converting the audio data into source language text to obtain text data corresponding to the audio. In this case, the source language text conversion unit 115 outputs the obtained text data to the audio model generation instruction unit 114.

[0054] The translated text conversion unit 116 converts the converted source text input from the source text conversion unit 115 into a translated text (i.e., translates it). This translation may be performed locally without using a network, or it may be performed using an external translation system via an API (Application Programming Interface) or the like.

[0055] In this case, the administrator of the information processing system S (for example, the provider of the distribution application) can arbitrarily set which language the translated text should be converted into. For example, the translated text may be converted into all of the most widely spoken languages, such as English, Chinese, and Spanish. Alternatively, the translated text may be converted into the language specified by the distributor at the start of distribution. Alternatively, during the broadcast, the viewer's terminal 20 may accept the viewer's request for a translation into a specific language. The server device 30 may then obtain information on which language the specified language is and convert it into a translated text. The translated text conversion unit 116 then transmits the converted translated text to each of the viewing terminals 20 that wish to receive the translation into that language, via the server device 30.

[0056] The distribution-side application data storage unit 171 is a storage unit that stores programs for operating each functional block of the distribution-side terminal 10, as well as data such as avatars used by the distribution-side application control unit 111 and the live distribution control unit 112.

[0057] The learning data storage unit 172 is a storage unit that stores the learning data collected by the learning data collection unit 113. The voice model storage unit 173 is a storage unit that stores the voice model acquired by the voice model generation instruction unit 114.

[0058] Next, we will explain the functional configuration of the viewing terminal 20. When content provision control processing is executed, the CPU 21 functions as follows: the viewer application control unit 211 and the audio data generation unit 212. Furthermore, one area of ​​the memory unit 27 is configured to store viewer-side application data 271 and audio model data 272.

[0059] The viewer-side application control unit 211 controls the entire process by which viewers watch live videos using the viewer-side application functions. Prior to broadcasting, the viewer-side application control unit 211 sets an icon representing the viewer (for example, a face icon). The viewer-side application control unit 211 also displays a list of broadcasters, allowing the viewer to select the broadcaster they wish to watch.

[0060] Furthermore, the viewer-side application control unit 211 manages the amount of virtual value elements (hereinafter referred to as "virtual coins") owned by the viewer within the game. Virtual coins can be obtained by paying real-world currency. However, this is not the only way; for example, users may be able to obtain virtual coins for free by fulfilling certain conditions, such as logging into the app or watching the stream to the end. These virtual coins can be used, for example, to send items or messages accompanied by images, or to provide gifts to the streamer while watching the stream. In addition, the viewer-side application control unit 211 manages various information, such as the viewer's username (in this case, the name of the viewer displayed when viewing the stream) and the password for logging into the server device 30.

[0061] The audio data generation unit 212 outputs the translated audio of the broadcaster during broadcasting. To this end, the audio data generation unit 212 receives from viewers, at the start of broadcasting, etc., their intention to use the translation function and the language they wish to be translated into. Then, it obtains an individual audio model corresponding to the broadcaster from the server device 30. The audio data generation unit 212 also stores the obtained audio model in the audio model storage unit 272.

[0062] The audio data generation unit 212 then inputs the translated text in the specified language, obtained from the distribution terminal 10 via the server device 30, into this audio model. As a result, audio data is generated that is reproduced in the broadcaster's voice and spoken using the translated words. The audio data generation unit 212 then outputs this audio data to the listener via a vocoder. This allows viewers to watch the stream while listening to audio reproduced in the streamer's voice and spoken using translated terms. Consequently, viewers can enjoy the streamer's distinctive voice and speaking style, which is a major draw of the streamed content.

[0063] The viewer-side application data storage unit 271 is a storage unit that stores programs for operating each functional block of the viewer-side terminal 20, as well as data such as the amount of virtual coins managed by the viewer-side application control unit 211.

[0064] The voice model storage unit 272 is a storage unit that stores the voice model acquired by the voice data generation unit 212.

[0065] Next, we will explain the functional configuration of the server device 30. When content provision control processing is executed, the content distribution unit 311 and the voice model management unit 312 function in the CPU 31, as shown in Figure 6. Furthermore, one area of ​​the memory unit 37 is configured with an application information database 371 and a voice model database 372.

[0066] Content distribution unit 311 provides live video as content. Live video is provided by the broadcaster to the viewers. When a broadcaster starts streaming live video from their broadcasting terminal 10, the content distribution unit 311 sends the broadcaster's live video to the viewer terminal 20 as one of the available streaming slots. If there are multiple broadcasters streaming, the content distribution unit 311 sends each broadcaster's slot to multiple viewer terminals 20. Viewers can compare the multiple broadcasters' slots displayed on their viewer terminals 20 and select the broadcaster they wish to join. Accordingly, the content distribution unit 311 identifies the viewer terminals 20 of viewers who wish to join the broadcast. The content distribution unit 311 then streams the live video to the identified viewer terminals 20.

[0067] To enable the provision of content such as live video, the content distribution unit 311 communicates with the distribution terminal 10 and the viewing terminal 20 as needed and centrally manages the data of the distribution application and the viewing application. For example, if data changes occur in the distribution application or the viewing application in response to operations by the broadcaster or viewers during or before / after live video distribution, the content distribution unit 311 first updates the application information DB 371. Then, it updates the distribution application data storage unit 171 of the distribution terminal 10 and the viewing application data storage unit 271 of the viewing terminal 20 in synchronization with the application information DB 371. This prevents inconsistencies in the data of the distribution application and the viewing application. Furthermore, the content distribution unit 311 also manages the login status of broadcasters and viewers.

[0068] The voice model management unit 312 manages voice models. When the voice model management unit 312 receives a voice model from the sending terminal 10, it stores this voice model in the voice model DB 372, linking it to the sender using the sending terminal 10. For example, it stores the voice model linked to the sender's identification information (e.g., the sender's username or the sender's management number). Furthermore, the voice model management unit 312, in response to a request for a voice model from the listener terminal 20, reads the corresponding voice model of the broadcaster from the voice model DB 372 and transmits it to the listener terminal 20. In addition, the audio model management unit 312 relays requests from the viewing terminal 20 to the distribution terminal 10 for translated text in the language specified by the viewer. It also relays the transmission of the translated text in the specified language from the distribution terminal 10 to the viewing terminal 20 in response to these requests.

[0069] The application information DB371 is a database that stores data for the distribution application on the distribution terminal 10 and the viewing application on the viewing terminal 20, respectively. By providing an application information DB 371 on the server device 30 in this way, and having businesses that provide distribution-side applications and viewing-side applications manage this information, it is possible to prevent data inconsistencies and data tampering by users.

[0070] The voice model DB372 is a database that stores and links voice models with the broadcasters who correspond to those voice models.

[0071] [Operation] The functional blocks of the broadcasting terminal 10, the viewing terminal 20, and the server device 30 have been described above. Next, the operation of the voice model generation process and broadcast translation process performed by each of these functional blocks of the broadcasting terminal 10, the viewing terminal 20, and the server device 30 will be described.

[0072] [Voice model generation process] Figure 7 is a flowchart showing the workflow of the voice model generation process. The voice model generation process is executed prior to distribution in response to the start command operation by the broadcaster.

[0073] In step S11, the training data collection unit 113 collects training data for generating a speech model. In step S12, the speech model generation instruction unit 114 determines whether the training data collected by the training data collection unit 113 has accumulated to a predetermined amount necessary for training. If the required amount for learning has been accumulated, the result is determined to be Yes in step S12, and the process proceeds to step S13. On the other hand, if the required amount for learning has not been accumulated, the result is determined to be No in step S12, and the process repeats step S11.

[0074] In step S13, the speech model generation instruction unit 114 instructs the TTS server to generate a speech model using the accumulated training data. In step S14, the voice model generation instruction unit 114 acquires the voice model generated by the TTS server.

[0075] In step S15, the voice model generation instruction unit 114 transmits this voice model to the server device 30. In step S16, the voice model management unit 312 of the server device 30 stores the voice model received from the distribution terminal 10 in the voice model DB 372, associating it with the distributor using the originating distribution terminal 10. This completes the voice model generation process.

[0076] [Translation processing for delivery] Figure 8 is a flowchart illustrating the workflow of the distribution translation process. The distribution translation process is executed in response to a start command from the distributor. Note that the speech model generation process is performed prior to the execution of the distribution translation process, so the speech model has already been generated.

[0077] In step S21, the live streaming control unit 112 requests the server device 30 to distribute content (in this case, live video). In step S22, the content distribution unit 311 of the server device 30 starts distributing content. Consequently, viewers using the viewing terminal 20 also begin viewing the content.

[0078] Here, Figure 9 shows an example of a screen being delivered in this embodiment. As shown in the figure, the display area of ​​the screen is broadly divided into display area AR1, display area AR2, and display area AR3. Display area AR1 shows the streamer's avatar and a background that corresponds to the stream. In live video, the avatar is displayed in this way while the streamer's voice is output, and the avatar's facial expressions and posture change in real time according to the streamer's movements. The display area AR2 shows each viewer's icon along with comments from viewers. Viewers can input comments for other viewers using the touch panel, keyboard, or voice on the viewer's terminal 20. In the AR3 display area, items that display images are shown along with the viewer's icon. Viewers can send such image-displaying items by, for example, spending a predetermined amount of virtual coins. This allows viewers to gain benefits such as being addressed by the broadcaster or being recognized by the broadcaster. In this way, live videos are streamed via a display screen, and broadcasters and viewers can enjoy the live videos by communicating with each other.

[0079] Furthermore, information regarding the translation function in this embodiment may be displayed in one of the three display areas mentioned above (for example, the area of ​​display area AR2 where viewer comments are not displayed). For example, information such as "What language is the source language used by the broadcaster, and to which language is it translated?" may be displayed as text. Alternatively, a user interface (for example, icons for selection and change) may be displayed to allow the user to select the target language for translation or to change a previously selected language to another language.

[0080] In general, the same content is displayed on the live streaming screen for both the streamer and the viewer. However, there may be differences in the display depending on the role of the streamer and the viewer. For example, the streamer's screen might display an icon to stop the stream, while the viewer's screen might display an icon to leave the stream.

[0081] In step S23, the audio data generation unit 212 determines whether it has received a request from the viewer to use the translation function and to specify the language to be translated into, which will trigger the translation to be performed. If the request to use the translation function has been received, the determination in step S22 is Yes, and the process proceeds to step S23. On the other hand, if the request to use the translation function has not been received, the determination in step S22 is No, and the process proceeds to step S29.

[0082] In step S24, the audio data generation unit 212 determines whether or not it has already acquired an individual audio model corresponding to the broadcaster currently broadcasting. If an audio model has already been acquired, the determination in step S24 is Yes, and the process proceeds to step S26. On the other hand, if an audio model has not been acquired, the determination in step S24 is No, and the process proceeds to step S25. The audio model may be deleted from the viewer's device 20 when the broadcast ends and retrieved each time the broadcaster broadcasts again. Alternatively, if the viewer is a follower of the broadcaster, the audio model may not be deleted even after the broadcast ends, but instead stored on the viewer's device 20 and linked to the broadcaster.

[0083] In step S25, the audio data generation unit 212 obtains an individual audio model from the server device 30 that corresponds to the broadcaster who is broadcasting. In step S26, the source language text conversion unit 115 converts the source language audio into source language text.

[0084] In step S27, the translated text conversion unit 116 converts the source text into a translated text (i.e., translates it). In step S28, the translated text in the specified language, obtained from the broadcasting terminal 10 via the server device 30, is input to the voice model corresponding to the broadcaster currently broadcasting. As a result, audio data is generated that is reproduced in the broadcaster's voice and spoken using the translated words, and the audio data generation unit 212 outputs this audio data to the viewers via a vocoder. Furthermore, the processing from step S22 to step S28 is performed in accordance with each viewing terminal 20.

[0085] In step S29, the content distribution unit 311 of the server device 30 determines whether or not to terminate the distribution. Distribution terminates when the distributor performs a distribution termination operation or when a predetermined distribution time has elapsed. If the distribution is to be terminated, the result in step S29 is determined to be Yes, and the content distribution unit 311 terminates the distribution process that has been continuously repeated since step S22. This terminates the distribution translation process. On the other hand, if the distribution is not to be terminated, the result in step S29 is determined to be No, and the process is repeated from step S23.

[0086] The voice model generation process and delivery translation process described above produce the various effects described in this embodiment. In other words, this embodiment solves the problem that the present invention aims to solve, which is to "further enhance the appeal of translated content."

[0087] [Differentiation] Although embodiments of the present invention have been described above, these embodiments are merely illustrative and do not limit the technical scope of the present invention. The present invention can take various other forms without departing from the spirit of the invention, and various modifications such as omissions and substitutions can be made. For example, it is possible not only to apply any of the modifications described below to the embodiments of the present invention, but also to combine some or all of the modifications described below as appropriate and apply them to the embodiments of the present invention.

[0088] [Example 1] In the above-described embodiment, functional blocks were implemented in the distribution terminal 10, the viewing terminal 20, and the server device 30 as shown in Figures 4 to 6. The distribution terminal 10, the viewing terminal 20, and the server device 30 each performed the operations they were required to perform. The following is not limited to this; the operations that the distribution terminal 10, the viewing terminal 20, and the server device 30 are to perform may differ from those in the above-described embodiment.

[0089] For example, the training data collection unit 113 and the speech model generation instruction unit 114, which were implemented in the distribution terminal 10, may be implemented in the server device 30. Furthermore, the server device 30 may generate the speech model using a TTS server. This eliminates the need for the distribution terminal 10 to generate the speech model using the TTS server and then transmit this speech model to the server device 30.

[0090] Furthermore, for example, the source language text conversion unit 115 and the translated language text conversion unit 116, which were implemented in the distribution terminal 10, may be implemented in the server device 30 or the viewing terminal 20. This reduces the load on the distribution terminal 10 during distribution. The server device 30 or the viewing terminal 20 may then perform the conversion using these functional blocks. However, in this case, if the source text conversion unit 115 and the translated text conversion unit 116 are implemented in the server device 30, the load will be concentrated on the server device 30, making it necessary to improve the processing capacity of the server device 30. Consequently, from the perspective of increasing the operating costs for the service provider providing the distribution service, it is preferable to implement these functions on the terminal side, such as the distribution terminal 10 or the viewing terminal 20.

[0091] Alternatively, the processing may be distributed such that the source language text conversion unit 115 is implemented on the distribution terminal 10 and the translated language text conversion unit 116 is implemented on the viewing terminal 20. In this case, the distribution terminal 10 sends the converted source language text to the viewing terminal 20. The viewing terminal 20 then converts the received source language text into translated text in the language selected by the viewer using the viewing terminal 20.

[0092] Alternatively, for example, the audio data generation unit 212, which was implemented in the viewing terminal 20, may be implemented in the distribution terminal 10 or server device 30. The distribution terminal 10 or server device 30 may then generate audio data corresponding to the language used by the viewer using an audio model and transmit it to the viewing terminal 20. Furthermore, if different viewers use different languages, similar processing can be performed for each viewer. This reduces the processing load on the viewing terminal 20, making it possible to provide translated audio using the voice model even if the viewing terminal 20 has low specifications. However, in this case, if audio data is generated to match the type and version of the OS (Operating System) of each of the multiple viewing terminals 20, or if audio data is generated for multiple languages, the load on the distribution terminal 10 or server device 30 will increase. Therefore, from the perspective of the distribution terminal 10 not being able to process everything, or the need to improve the processing power of the server device 30 as described above, it is preferable for each of the multiple viewing terminals 20 to generate audio data that matches its own environment.

[0093] [Differentiation 2] In the embodiments described above, the source text conversion unit 115 was assumed to use a general-purpose speech recognition engine. Similarly, the translated text conversion unit 116 was assumed to use a general-purpose translation engine. However, the invention is not limited to these, and these general-purpose speech recognition engines and general-purpose translation engines may be extended. For example, in live streaming videos, streamers often use a variety of unusual terms, such as their own unique vocabulary (e.g., self-introductions viewers say when they start watching, or original catchphrases), special expressions (e.g., unique expressions of gratitude used when receiving gifts), nouns specific to the subject being streamed (e.g., specific nouns within games or anime), or temporary slang. However, there is a risk that general-purpose speech recognition engines may not be able to properly recognize such unusual terms, and general-purpose translation engines may not be able to properly translate them.

[0094] Therefore, these general-purpose engines are pre-loaded with dictionary data of uncommon terms (so to speak, unnatural language) as exemplified above. This allows for accurate detection of word boundaries based on this dictionary data, as well as the selection of appropriate translation terms, thereby enabling accurate speech recognition and translation. Furthermore, such dictionary data may be common to all distributors, or it may differ in some or all of its content from one distributor to another.

[0095] [Difference 3] In the above-described embodiment, the audio translated into the target words was provided to the viewer during the distribution translation process. However, the distribution translation process may also provide other elements to the viewer. For example, the avatar's movements (such as mouth movements or overall facial expressions) could be adjusted in accordance with the translated audio.

[0096] First, for viewers watching the broadcast in the original language, as described in the above embodiment, the camera captures images of the broadcaster during the broadcast, and the movements of various parts of the broadcaster's face, eyes, mouth, and upper body are tracked. Then, by changing the movements of the avatar (and its accessories) in real time to correspond to the tracked movements, the actual movements of the broadcaster are reflected in the avatar. This makes the avatar's movements realistic and synchronized with the audio.

[0097] In contrast, when viewing a stream with translated words, if the tracked movements are reflected, the timing and tone of the translated words' pronunciation will not match the avatar's movements, resulting in an unnatural appearance. Therefore, speech analysis is performed, and the avatar's movements are adjusted according to the analysis results. To do this, the audio data of the translated words is analyzed to identify the loudness (volume), frequency (pitch), and phonemes (corresponding to vowels such as A, I, U, E, and O). Then, the avatar's movements (for example, changes in mouth shape) are controlled to match these identified results. This allows us to provide the translated audio to the audience in a more natural way.

[0098] [Differentiation Example 4] In the above embodiment, the audio of the original language spoken by the broadcaster was translated into audio of the translated language and provided to the viewers. However, the system is not limited to this, and other elements related to the broadcast may also be translated. For example, the viewer terminal 20 could be equipped with a text translation function equivalent to that of the translated language text conversion unit 116. Then, when a viewer comments using the translated language (i.e., the original language used by the viewer) via the viewer terminal 20, this could be translated back into the original language (i.e., the original language used by the broadcaster) before being displayed to the broadcaster and other viewers. Alternatively, comments from other streamers, which are in the original language, could be translated and displayed to this streamer.

[0099] [Difference 5] In the embodiment described above, a test mode was executed prior to distribution, and the training data collection unit 113 collected training data in this test mode. However, the training data may be collected by other methods as well. For example, the system collects audio data of speech uttered by the broadcaster during past broadcasts. The source-language text conversion unit 115 then converts this audio data into source-language text, thereby obtaining text data corresponding to the audio. In this way, the source-language text conversion unit 115 can collect pairs of audio data and their corresponding text data as training data without having to run a test mode. Furthermore, the collection of training data using the test mode in the above-described embodiment may be used in combination with the collection of training data from the speech uttered by the broadcaster during broadcasting, as in this modified example. In this case, first, a speech model is generated using the training data from the test mode. Then, this speech model is fine-tuned (additional training) using the training data from the speech uttered by the broadcaster. This makes it possible to generate a speech model that reproduces the broadcaster's voice with higher accuracy. Furthermore, the collection of training data from the speech uttered by the broadcaster may be limited to cases where the broadcaster's past broadcasts are less than a predetermined number of times, or where the broadcast duration of past broadcasts is less than a predetermined time. This prevents an increase in computational processing due to excessive fine-tuning. It also prevents overfitting due to excessive fine-tuning.

[0100] [Modification 6] The voice model generated in the above-described embodiment may be further fine-tuned from other perspectives. For example, a test mode can be performed with the broadcaster speaking in the original language to generate a voice model once. Then, a test mode can be performed with the broadcaster speaking in the translated language to fine-tune this initially generated voice model. This makes it possible to generate a voice model that reproduces the broadcaster's voice with greater accuracy. In this case, the generated speech model can be used to actually output the translated words, and after the broadcaster listens to this output, a test mode can be performed using the broadcaster's pronunciation of the translated words. This makes it possible to collect training data in which the broadcaster intentionally corrects parts that they find unnatural when listening to the output of the speech model themselves.

[0101] [Difference 7] In the above embodiment, the broadcast translation process was performed in real time during the broadcast by the broadcaster, translating the audio in the original language spoken by the broadcaster into audio in the target language and providing it to the viewers. However, the broadcast translation process may be performed at other times. For example, past broadcasts may be stored in a server device 30, and the broadcast translation process may be executed when a viewer plays back and watches the stored broadcast after the broadcast has ended.

[0102] [Differentiation 8] The above embodiment assumed the use of a TTS model that had been trained to handle multiple languages ​​in a general way. However, it is not limited to this, and a TTS model that supports only one source language may also be used. For example, if a broadcaster can speak two or more languages ​​(e.g., Japanese and English), they can generate a TTS model in one language (e.g., English). Then, when broadcasting in the other language (e.g., Japanese), they can use the pre-generated TTS model as the speech model. In this way, the broadcaster can provide audio with translations that more accurately reflect their own English voice and speaking style.

[0103] [Modification 9] In the above-described embodiment, a client-server system was constructed using the distribution terminal 10, the viewing terminal 20, and the server device 30, and the functions of the information processing system S were realized by further utilizing a TTS server. However, the functions of the information processing system S may be realized in other ways. For example, the functions of the TTS server may be realized by the server device 30 or by the distribution terminal 10. Alternatively, for example, one of the server device 30 or the distribution terminal 10 may have some or all of its functional blocks, while the other has some or all of them.

[0104] [Example Configuration] As described above, the information processing system S in this embodiment comprises a voice model DB372, a source language text conversion unit 115, a translated language text conversion unit 116, and a voice data generation unit 212. The voice model DB372 stores a unique voice model for each broadcaster, capable of generating broadcaster-specific voice data based on text data. The source language text conversion unit 115 converts the broadcaster's voice into text data in the first language. The translated text conversion unit 116 converts the text data in the first language converted by the source text conversion unit 115 into text data in the second language. The audio data generation unit 212 generates audio data in the second language corresponding to the broadcaster, using the second language text data converted by the translated text conversion unit 116 and a proprietary audio model corresponding to the broadcaster.

[0105] The information processing system S further comprises a learning data collection unit 113 and a speech model generation instruction unit 114. The learning data collection unit 113 collects the broadcaster's voice data from broadcasts the broadcaster has conducted in the past. The voice model generation instruction unit 114 generates a unique voice model for each broadcaster by performing machine learning using the collected broadcaster's voice data.

[0106] The learning data collection unit 113 collects the broadcaster's voice data when the broadcaster speaks any phrase or at least one of the specified phrases at a time when the broadcaster is not broadcasting. The voice model generation instruction unit 114 generates a unique voice model for each broadcaster by performing machine learning using the collected broadcaster's voice data.

[0107] The information processing system S comprises a distribution terminal 10 used by the distributor, a viewing terminal 20 used by the viewer, and a server device 30 that controls the viewing of the distribution from the distribution terminal 10 on the viewing terminal 20. The source language text conversion unit 115, the translated language text conversion unit 116, and the audio data generation unit 212 are provided in either the distribution terminal 10 or the viewing terminal 20, and operate while the distributor is performing the distribution.

[0108] The source language text conversion unit 115, or the source language text conversion unit 115 and the translated language text conversion unit 116, are provided in the distribution terminal 10.

[0109] When the audio data generation unit 212 outputs audio data in a second language, it adjusts the image displayed in the broadcast by the broadcaster according to the audio data in the second language before displaying it.

[0110] There are multiple types of second languages. The DB372 voice model stores a unique multilingual voice model for each broadcaster, capable of supporting multiple types of second languages. The audio data generation unit 212 uses a multilingual speech model to generate audio data corresponding to a second language of a type suitable for the listener.

[0111] The information processing system S comprises a distribution terminal 10 used by the distributor, a viewing terminal 20 used by the viewer, and a server device 30 that controls the viewing of the distribution from the distribution terminal 10 on the viewing terminal 20. The audio data generation unit 212 uses a multilingual speech model to generate multiple types of audio data tailored to each of the multiple viewers. The server device 30 includes an audio model management unit 312 that distributes to each of the multiple viewing terminals 20 used by multiple viewers the type of audio data from among the multiple types of audio data generated by the audio data generation unit 212 that is suitable for the viewer using that viewing terminal 20.

[0112] The information processing system S comprises a distribution terminal 10 used by the distributor, a viewing terminal 20 used by the viewer, and a server device 30 that controls the viewing of the distribution from the distribution terminal 10 on the viewing terminal 20. The viewing terminal 20 includes a viewing application control unit 211 that obtains a multilingual voice model corresponding to a second language of a type suitable for the viewer using the viewing terminal 20 from the voice model DB 372 via communication, and a voice data generation unit 212.

[0113] The distribution-side application data storage unit 171 stores dictionary data for converting text data in the first language to text data in the second language, and further stores dictionary data unique to each distribution provider. The translated text conversion unit 116 uses dictionary data to convert the text data of the first language converted by the source language text conversion unit 115 into text data of the second language.

[0114] Embodiments and modifications thereof of the present invention have been described above. However, the present invention is not limited to the embodiments and modifications described above, and any modifications, improvements, etc. that can achieve the objectives of the present invention are included within the scope of the present invention and its equivalents.

[0115] Furthermore, in the embodiments described above, the distribution terminal 10, viewing terminal 20, and server device 30 to which the present invention is applied were described as smartphones, game consoles, and server devices as examples, but are not particularly limited to these. The present invention can be applied to electronic devices in general that have information processing functions. In addition, the functional configuration of the server device 30 and the distribution terminal 10 and viewing terminal 20 may be realized in a single device. Alternatively, the functions of the server device 30 may be distributed among multiple server devices and realized in multiple devices.

[0116] Furthermore, the series of processes described above can be executed by hardware or by software. In other words, the functional configuration in the above-described embodiment is merely illustrative and not particularly limiting. That is, an information processing system has the functionality to execute the above-described series of processes as a whole. It is sufficient if it is provided in S, and the type of functional block used to realize this function is not limited to the examples of the embodiments described above. Furthermore, a single functional block may consist of hardware alone, software alone, or a combination of both.

[0117] When a series of processes are executed by software, the programs that make up that software are installed on a computer or other device from a network or storage medium. A computer may be a computer built into dedicated hardware. Alternatively, a computer may be a computer capable of performing various functions by installing various programs, such as a general-purpose personal computer.

[0118] The functional configuration of these computers is realized by processors that perform arithmetic processing. For example, this includes not only systems composed of various processing units such as single processors, multiprocessors, and multicore processors, but also systems that combine these various processing units with processing circuits such as ASICs (Application Specific Integrated Circuits) and FPGAs (Field-Programmable Gate Arrays).

[0119] The storage medium for storing programs consists of removable media distributed separately from the main unit, or storage media pre-installed in the main unit. Removable media consists of, for example, magnetic disks, optical disks, magneto-optical disks, or flash memory. Optical disks consist of, for example, CD-ROM (Compact Disk-Read Only Memory), DVD (Digital Versatile Disk), Blu-ray Disc (registered trademark), etc. Magneto-optical disks consist of, for example, MD (Mini-Disk). Flash memory consists of, for example, USB (Universal Serial Bus) memory or SD cards. Furthermore, storage media pre-installed in the main unit consists of, for example, ROM, SSD, HDD, etc., on which programs are stored.

[0120] In this specification, the step of describing a program to be recorded on a recording medium includes not only processes that are performed chronologically in that order, but also processes that are not necessarily performed chronologically, but are executed in parallel or individually. Furthermore, in this specification, the term "system" refers to an overall system composed of multiple devices, means, etc. [Explanation of symbols]

[0121] 10 Broadcasting terminal, 20 Viewing terminal, 30 Server device, 11,21,31 CPU, 12,22,32 ROM, 13,23,33 RAM, 14,24,34 Bus, 15,25,35 Input unit, 16,26,36 Output unit, 17,27,37 Storage unit, 18,28,38 Communication unit, 19,29,39 Drive, 100 Removable media, 111 Broadcasting application control unit, 112 Live streaming control unit, 113 Training data collection unit, 114 Voice model generation instruction unit, 115 Source language text conversion unit, 116 Translated language text conversion unit, 171 Broadcasting application data storage unit, 172 Training data storage unit, 173 Voice model storage unit, 211 Viewing application control unit, 212 Voice data generation unit, 271 Viewer-side application data storage unit, 272 Voice model storage unit, 311 Content distribution unit, 312 Voice model management unit, 371 Application information DB (database), 372 Voice model DB (database), N Network, S Information processing system

Claims

1. A storage means that stores a unique voice model for each broadcaster, capable of generating voice data corresponding to the broadcaster based on text data, A first conversion means for converting the broadcaster's voice into text data in a first language, A second conversion means converts the text data in the first language converted by the first conversion means into text data in the second language, The audio data generation means generates audio data in the second language corresponding to the broadcaster using the second language text data converted by the second conversion means and a proprietary voice model corresponding to the broadcaster, and when outputting the generated audio data in the second language, it displays the image displayed in the broadcast by the broadcaster after adjusting it according to the audio data in the second language. An information processing system characterized by comprising the following features.

2. A collection means for collecting audio data of the broadcaster from broadcasts previously conducted by the broadcaster, The system further includes a generation instruction means for generating a unique voice model for each broadcaster by performing machine learning using the collected voice data of the broadcasters. The information processing system according to feature 1.

3. A collection means for collecting audio data of the aforementioned broadcaster when the broadcaster utters any phrase or at least one of a specified phrase at a time when the broadcaster is not broadcasting, The system further includes a generation instruction means for generating a unique voice model for each broadcaster by performing machine learning using the collected voice data of the broadcasters. The information processing system according to feature 1.

4. The information processing system comprises a distribution terminal used by the broadcaster, a viewing terminal used by the viewer, and a server device that controls the viewing of the broadcast from the distribution terminal on the viewing terminal, as components of the information processing system. Each of the first conversion means, the second conversion means, and the voice data generation means is: The broadcasting terminal and the viewing terminal are equipped with and operate while the broadcaster is performing the broadcast, The information processing system according to claim 1 or 2, characterized in that it is the same as described in claim 1 or 2.

5. The first conversion means, or the first conversion means and the second conversion means, are provided by the distribution terminal. The information processing system according to feature 4.

6. There are multiple types of the aforementioned second language. The storage means stores a unique multilingual speech model for each of the distributors that is compatible with each of the multiple types of second languages. The audio data generation means uses the multilingual audio model to generate audio data corresponding to the second language of a type suitable for the listener. The information processing system according to claim 1 or 2, characterized in that it is the same as described in claim 1 or 2.

7. The information processing system comprises a distribution terminal used by the broadcaster, a viewing terminal used by the viewer, and a server device that controls the viewing of the broadcast from the distribution terminal on the viewing terminal, as components of the information processing system. The audio data generation means generates multiple types of audio data tailored to each of the multiple viewers using the multilingual audio model. The server device includes a management means for distributing, to each of the multiple viewing terminals used by the multiple viewers, the type of audio data from among the multiple types of audio data generated by the audio data generation means that is suitable for the viewer using that viewing terminal. The information processing system according to feature 6.

8. The information processing system comprises a distribution terminal used by the broadcaster, a viewing terminal used by the viewer, and a server device that controls the viewing of the broadcast from the distribution terminal on the viewing terminal, as components of the information processing system. The aforementioned viewing terminal is A communication means for acquiring, by communication, a multilingual voice model corresponding to the second language of a type suitable for the viewer using the viewing terminal, from the storage means; The aforementioned audio data generation means, Equipped with, The information processing system according to feature 6.

9. The storage means further stores dictionary data for converting text data in the first language to text data in the second language, and includes dictionary data specific to each distributor. The second conversion means uses the dictionary data to convert the text data in the first language converted by the first conversion means into text data in the second language. The information processing system according to claim 1 or 2, characterized in that it is the same as described in claim 1 or 2.

10. A memory function that stores a unique voice model for each broadcaster, capable of generating voice data corresponding to the broadcaster based on text data, A first conversion function that converts the broadcaster's voice into text data in a first language, A second conversion function converts the text data in the first language converted by the first conversion function into text data in the second language, The second conversion function generates audio data in the second language corresponding to the broadcaster using the second language text data converted by the second conversion function and a proprietary voice model corresponding to the broadcaster, and when outputting the generated audio data in the second language, it displays the image displayed in the broadcast by the broadcaster after adjusting it according to the audio data in the second language. An information processing program characterized by its ability to implement the following on a computer.