Live voice generation system
The live audio generation system addresses the challenge of providing real-time and flexible live commentary by using AI to generate commentary based on event status data, prioritizing timely explanations and allowing for natural, engaging commentary.
Patent Information
- Application Number
- JP2023188412
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-16
AI Technical Summary
Existing live commentary systems struggle to provide real-time, flexible, and subjective commentary that captures the nuances of event situations, such as racing games, while maintaining a clear and timely explanation of the situation.
A live audio generation system that utilizes an input unit to acquire status data, a first processing unit to select fixed speech data, and a second processing unit to generate flexible speech data using AI processing, prioritizing the output of fixed speech data for timely explanations and allowing flexible speech during less critical moments.
The system achieves real-time and highly flexible live commentary by prioritizing fixed situation explanations and allowing for natural commentary with minimal pauses, enhancing the understanding and engagement of event situations.
Smart Images

Figure 2025076663000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a live commentary sound generation system that outputs live commentary sound regarding an event situation. [Background technology]
[0002] Sports videos, including motorsports such as F1, are accompanied by live commentary such as "Next is the final straight, can you overtake?", explaining the situation and conveying the commentator's subjective comments to the viewer. By listening to the commentary while watching the video, viewers can understand the situation more deeply and enjoy watching the race more. Live commentary is added to many sports videos and video game videos, and it plays an important role in entertaining viewers and is also expected to increase the value of the video itself. However, since commentary requires knowledge of the target sport or event and appropriate speaking skills, many sports videos and video game videos available online do not have commentary. As an example of live commentary, we focus on live commentary of a racing game video. In such live commentary, it is necessary to speak about important events that occur in the video at the appropriate timing within a short period of time. In conventional language generation research, the problem of planning, i.e., "what to speak about," and the problem of language surfaceization, i.e., "how to speak," have been mainly studied separately. In live commentary generation, in addition to these problems that have been tackled traditionally, it is necessary to deal with problems that have not been considered in the past, such as "when to speak," "how long to speak," and "how detailed to speak." In order to identify the timing of speech and generate live commentary speech, it is necessary to consider time-series data expressed in videos, etc. This is different from the problem setting of language generation that does not consider the time axis, such as image caption generation, or the problem setting of story generation, where the timing of speech is given in advance. In conventional language generation research, there are many settings for inputting visual information such as video and images. However, it is not easy to correctly recognize video for generating live commentary. For example, live commentary often describes the detailed positional relationship between vehicles, and it is necessary to recognize slight differences in the positional relationship of vehicles from a series of similar video frames such as aerial footage. On the other hand, image caption generation assumes the output of text that does not necessarily need to capture minor differences between objects, such as "people are dancing." In this context, it is effective to use time-series situation data as input in addition to video in order to capture the situation of the race more accurately. The situation data includes multiple numerical data related to the circuit and the racing car, such as the coordinate position, speed, and steering angle of the racing car. In actual F1 races, more than 300 types of numerical data are measured in real time from multiple sensors, and even in fields other than motorsports, the location information of soccer players is obtained by GPS, and attempts to utilize situation data in various sports are spreading. Incidentally, Patent Document 1 proposes a game commentary broadcasting device that automatically broadcasts a commentary corresponding to the game development. In the game commentary broadcasting device of Patent Document 1, predetermined live broadcast audio data corresponding to a plurality of game development patterns is stored in advance in an audio data storage means, the game development pattern of the game system is determined, and a command to read out the audio data corresponding to the game development pattern is output. Patent Document 2, like the game commentary broadcasting device of Patent Document 1, pre-stores commentary data that specifies the commentary content according to the situation of the game as it progresses, but it also determines at least two commentary axes that match the focus of the game content, and by determining the commentary content according to the commentary data corresponding to these commentary axes, it proposes a device that can realize multifaceted commentary, such as commenting on a game from multiple perspectives, just like a real announcer. Patent Document 3 proposes a system that has a voice information generation unit configured to function based on a trained machine learning model, and can output voice for commentary, etc. in accordance with the game situation and the flow of conversation, such as commentary. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 8-215433 [Patent Document 2] JP 2013-111178 A [Patent Document 3] Patent Publication No. 2021-194229 Summary of the Invention [Problem to be solved by the invention]
[0004] When live commentary data is stored in advance, as in Patent Documents 1 and 2, it has an advantage in terms of real-time performance, but it is difficult to have the commentator speak subjective comments along with an explanation of the situation, such as "Next is the final straight zone. Can you overtake?" On the other hand, when generating voice information based on a trained model of machine learning as in Patent Document 3, it is possible to have the commentator speak subjective comments along with an explanation of the situation, but this is not suitable for real-time performance.
[0005] The present invention aims to provide a live voice generation system that can realize real-time speech and highly flexible speech by processing speech that is somewhat fixed in response to a situation and speech in flexible language differently. [Means for solving the problem]
[0006] The live commentary voice generation system of the present invention described in claim 1 is a live commentary voice generation system that outputs live commentary voice about an event situation, and comprises an input unit 1 that acquires situation data about the event situation, a first processing unit 10 that selects first utterance data based on the situation data acquired by the input unit 1, a second processing unit 20 that generates second utterance data based on the situation data acquired by the input unit 1, a first utterance data storage unit 3 that stores the first utterance data selected by the first processing unit 10 and the situation data conditions for selecting the first utterance data, a voice generation unit 11, 21 that generates a first voice based on the first utterance data and a second voice based on the second utterance data, and a voice output unit 2 that outputs the first voice and the second voice generated by the voice generation units 11, 21, wherein the first processing unit 10 selects the first utterance data from the first utterance data storage unit 3, the second processing unit 20 generates the second utterance data by a first AI processing 4, and the voice output unit 2 outputs the first voice in priority to the second voice. The present invention as set forth in claim 2 is characterized in that, in the live audio generation system as set forth in claim 1, the audio output unit 2 does not output the second audio while the first audio is being output. The present invention described in claim 3 is characterized in that, in the live voice generation system described in claim 1, when the second speech data is generated while the first voice is being output, the second processing unit 20 generates the second speech data based on new situation data without generating the second voice based on the generated second speech data. The present invention described in claim 4 is characterized in that, in the live voice generation system described in claim 1, when the second voice is generated while the first voice is being output, the second processing unit 20 generates the second speech data based on new situation data without outputting the generated second voice from the voice output unit 2. The present invention as described in claim 5 is characterized in that, in the live audio generation system as described in claim 1, the audio output unit 2 stops outputting the second audio when outputting the first audio while outputting the second audio. The present invention described in claim 6 is characterized in that, in the live voice generation system described in claim 1, the second processing unit 20 generates the second speech data in response to the output of the first voice as a trigger. The present invention described in claim 7 is characterized in that, in the live commentary voice generation system described in claim 1, it comprises a third processing unit 30 that generates third utterance data based on the situation data acquired by the input unit 1, and a second utterance data memory unit 5 that associates and stores the second utterance data generated by the second processing unit 20 and the situation data at the time the second utterance data is generated, wherein the voice generation unit 31 generates a third voice based on the third utterance data, the voice output unit 2 outputs the third voice generated by the voice generation unit 31, the third processing unit 30 generates the third utterance data by a second AI processing 6 using data stored in the second utterance data memory unit 5 as teacher data, and the voice output unit 2 outputs the first voice in priority to the second voice and the third voice. The present invention described in claim 8 is characterized in that, in the live audio generation system described in claim 7, the audio output unit 2 does not output the third audio while the first audio is being output. The present invention as described in claim 9 is characterized in that, in the live voice generation system as described in claim 7, when the third speech data is generated while the first voice is being output, the third processing unit 30 generates the third speech data based on new situation data without generating the third voice based on the generated third speech data. The present invention described in claim 10 is characterized in that, in the live voice generation system described in claim 7, when the third voice is generated while the first voice is being output, the third processing unit 30 generates the third speech data based on new situation data without outputting the generated third voice from the voice output unit 2. The present invention as described in claim 11 is characterized in that, in the live audio generation system as described in claim 7, the audio output unit 2 stops outputting the third audio when outputting the first audio while outputting the third audio. The present invention described in claim 12 is characterized in that, in the live voice generation system described in claim 7, the third processing unit 30 generates the third speech data when the first voice is output as a trigger. The live commentary voice generation system of the present invention described in claim 13 is a live commentary voice generation system that outputs live commentary voice about an event situation, and is characterized in that, when the acquired situation data matches a condition for selecting first utterance data, a computer executes a first processing step of selecting the matching first utterance data, a first voice generation step of generating a first voice using the first utterance data selected in the first processing step, a first voice output instruction step of instructing output using the first voice generated in the first voice generation step, and a voice output step of outputting the first voice instructed in the first voice output instruction step, and, when an instruction is given by the first voice output instruction step, a second processing step of generating second utterance data by a first AI processing 4, a second voice generation step of generating a second voice using the second utterance data generated in the second processing step, and, if output using the first voice is not being performed, a second voice output instruction step of instructing output using the second voice generated in the second voice generation step, and, if output using the first voice is being performed, the second utterance data using new situation data is generated in the second processing step without outputting the second voice generated in the second voice generation step. The present invention as described in claim 14 is characterized in that, in the live audio generation system as described in claim 13, when output using the first audio is instructed in the first audio output instruction step while output using the second audio is being performed in the second audio output instruction step, output using the second audio is stopped and output using the first audio is performed. The live commentary voice generation system of the present invention described in claim 15 is a live commentary voice generation system that outputs live commentary voice about an event situation, and includes a first processing step of selecting the first utterance data that matches the acquired situation data when the acquired situation data matches a condition for selecting first utterance data, a first voice generation step of generating a first voice using the first utterance data selected in the first processing step, a first voice output instruction step of instructing output using the first voice generated in the first voice generation step, a voice output step of outputting the first voice instructed in the first voice output instruction step, and a second processing step of generating second utterance data by a first AI processing 4 when an instruction is given in the first voice output instruction step. The method is characterized in that the control unit executes a step of generating a second voice using the second speech data generated in the second processing step, a third processing step of generating third speech data by a second AI processing 6 when an instruction is given by the first speech output instruction step, a third speech generation step of generating a third voice using the third speech data generated in the third processing step, a third speech output instruction step of outputting the third voice generated in the third speech generation step if output using the first voice is not being performed, and a second speech output instruction step of outputting the second voice generated in the second speech generation step if output using the first voice and output using the third voice are not being performed. The present invention as described in claim 16 is characterized in that, in the live voice generation system as described in claim 15, if output using the first voice is being performed, the second utterance data is generated using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. The present invention as described in claim 17 is characterized in that, in the live voice generation system as described in claim 15, if output using the first voice is being performed, the third speech data is generated using new situation data in the third processing step without outputting the third voice generated in the third voice generation step. The present invention as described in claim 18 is characterized in that, in the live voice generation system as described in claim 15, if output using the third voice is being performed, the second utterance data is generated using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. The present invention as described in claim 19 is characterized in that, in the live audio generation system as described in claim 15, when output using the first audio is instructed in the first audio output instruction step while output using the second audio is being performed in the second audio output instruction step, output using the second audio is stopped and output using the first audio is performed. The present invention as described in claim 20 is characterized in that, in the live audio generation system as described in claim 15, when output using the first audio is instructed in the first audio output instructing step while output using the third audio is being performed in the third audio output instructing step, output using the third audio is stopped and output using the first audio is performed.
[0007] According to the present invention, by using the first speech data and the second speech data, it is possible to output utterances that require real-time response and highly flexible utterances as live commentary audio, allowing for a deeper understanding of the event situation. By outputting the first audio prior to the second audio, it is possible to reliably output a relatively fixed explanation of the situation in a timely manner, and highly flexible utterances can be output during periods when no explanation of the situation is necessary, thereby achieving a natural live commentary audio with fewer silent periods. [Brief description of the drawings]
[0008] [Figure 1] A block diagram showing a live voice generation system according to an embodiment of the present invention, with functional implementation means. [Diagram 2] A flow chart showing the process flow of the live commentary voice generation system. [Diagram 3] FIG. 1 is an explanatory diagram showing the processing performed by the live commentary voice generation system; [Figure 4] A block diagram showing a live commentary sound generation system according to another embodiment of the present invention, with functional implementation means. [Diagram 5] A flow chart showing the process flow of the live commentary voice generation system. [Figure 6] FIG. 2 is an explanatory diagram showing the relationship between the second process and the third process in the live commentary voice generation system. [Figure 7] FIG. 11 is an explanatory diagram illustrating a case where a first voice output occurs during a second process in the live voice generation system. [Figure 8] FIG. 11 is an explanatory diagram showing a case where a first voice output occurs during output of a third voice in the live voice generation system; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] A live voice generation system according to a first embodiment of the present invention comprises an input unit that acquires situation data regarding an event situation, a first processing unit that selects first utterance data based on the situation data acquired by the input unit, a second processing unit that generates second utterance data based on the situation data acquired by the input unit, a first utterance data storage unit that stores the first utterance data selected by the first processing unit and situation data conditions for selecting the first utterance data, a voice generation unit that generates a first voice based on the first utterance data and a second voice based on the second utterance data, and a voice output unit that outputs the first voice and the second voice generated by the voice generation unit, wherein the first processing unit selects the first utterance data from the first utterance data storage unit, the second processing unit generates the second utterance data by first AI processing, and the voice output unit outputs the first voice in priority to the second voice. According to this embodiment, by using the first speech data stored in the first speech data storage unit and the second speech data generated by the first AI processing, it is possible to output speech that requires real-time processing and speech with high flexibility as live commentary, and to understand the event situation more deeply. Also, according to this embodiment, by outputting the first speech in preference to the second speech, it is possible to reliably output a certain degree of a fixed situation explanation in a timely manner, and to output a flexible speech during a period when a situation explanation is not necessary, thereby realizing a natural live commentary with few silent periods.
[0010] The second embodiment of the present invention is a commentary voice generation system according to the first embodiment, in which the voice output unit does not output the second voice while the first voice is being output. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner.
[0011] In the third embodiment of the present invention, in the live commentary generation system according to the first embodiment, when the second utterance data is generated while the first voice is being output, the second processor generates the second utterance data based on new situation data without generating the second voice based on the generated second utterance data. According to this embodiment, it is possible to reliably output a situation description that is somewhat fixed in a timely manner, and to output a flexible utterance during a period when a situation description is not required.
[0012] In the fourth embodiment of the present invention, in the live commentary voice generation system according to the first embodiment, when the second voice is generated while the first voice is being output, the generated second voice is not output from the voice output unit, and the second processing unit generates second speech data based on new situation data. According to this embodiment, it is possible to reliably output a situation explanation that is somewhat fixed in a timely manner, and to output a highly flexible utterance during a period when a situation explanation is not necessary.
[0013] The fifth embodiment of the present invention is a commentary voice generation system according to the first embodiment, in which the voice output unit stops outputting the second voice when outputting the first voice while outputting the second voice. According to this embodiment, it is possible to reliably output a situation explanation that is determined to a certain extent in a timely manner.
[0014] In the sixth embodiment of the present invention, in the live voice generation system according to the first embodiment, the second processing unit generates the second speech data in response to the output of the first voice as a trigger. According to this embodiment, flexible speech can be output during a period when a situation explanation is not required.
[0015] A seventh embodiment of the present invention is a live commentary voice generation system according to the first embodiment, which includes a third processing unit that generates third utterance data according to situation data acquired by the input unit, and a second utterance data storage unit that associates and stores the second utterance data generated by the second processing unit and the situation data when the second utterance data is generated, and the voice generation unit generates a third voice according to the third utterance data, the voice output unit outputs the third voice generated by the voice generation unit, the third processing unit generates the third utterance data by a second AI process using the data stored in the second utterance data storage unit as teacher data, and the voice output unit outputs the first voice in preference to the second voice and the third voice. According to this embodiment, the second utterance data that is repeatedly generated is generated as the third utterance data by the second AI process, so that the third utterance data can be generated in a shorter time than the generation time of the second utterance data, and it is easy to output a highly flexible utterance. Furthermore, according to this embodiment, by outputting the first voice in priority over the second and third voices, a relatively fixed situation explanation can be output in a timely manner with certainty, and flexible speech can be output during periods when a situation explanation is not necessary, thereby realizing a natural commentary voice with fewer silent periods.
[0016] The eighth embodiment of the present invention is a commentary voice generation system according to the seventh embodiment, in which the voice output unit does not output the third voice while the first voice is being output. According to this embodiment, a situation explanation that is determined to a certain extent can be output in a timely manner with certainty.
[0017] In the ninth embodiment of the present invention, in the live commentary generation system according to the seventh embodiment, when the third utterance data is generated while the first voice is being output, the third processing unit generates the third utterance data based on new situation data without generating the third voice based on the generated third utterance data. According to this embodiment, it is possible to reliably output a situation description that is somewhat fixed in a timely manner, and to output a flexible utterance during a period when a situation description is not required.
[0018] In the tenth embodiment of the present invention, in the live commentary voice generation system according to the seventh embodiment, when the third voice is generated while the first voice is being output, the generated third voice is not output from the voice output unit, and the third processing unit generates third speech data based on new situation data. According to this embodiment, it is possible to reliably output a certain degree of situation explanation in a timely manner, and to output a highly flexible utterance during a period when a situation explanation is not necessary.
[0019] The eleventh embodiment of the present invention is a commentary voice generation system according to the seventh embodiment, in which the voice output unit stops outputting the third voice when outputting the first voice while outputting the third voice. According to this embodiment, it is possible to reliably output a certain degree of a situation explanation in a timely manner.
[0020] A twelfth embodiment of the present invention is a commentary voice generation system according to the seventh embodiment, in which the third processing unit generates third speech data by using the output of the first voice as a trigger. According to this embodiment, flexible speech can be output during a period when a situation explanation is not required.
[0021] A live commentary voice generation system according to a thirteenth embodiment of the present invention executes a first processing step in which, when acquired situation data matches a condition for selecting first utterance data, a computer selects the matching first utterance data, a first voice generation step in which a first voice is generated using the first utterance data selected in the first processing step, a first voice output instruction step in which an instruction is given to output using the first voice generated in the first voice generation step, an audio output step in which the first voice instructed in the first voice output instruction step is output, a second processing step in which, when an instruction is given in the first voice output instruction step, a second voice generation step in which a second voice is generated using the second utterance data generated in the second processing step, and, if output using the first voice is not being performed, a second voice output instruction step in which an instruction is given to output using the second voice generated in the second voice generation step, and, if output using the first voice is being performed, generates second utterance data using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. According to this embodiment, by using the first speech data stored in advance and the second speech data generated by the first AI processing, it is possible to output a speech that requires real-time processing and a flexible speech as a live commentary voice, and to understand the event situation more deeply. Also, according to this embodiment, by outputting the first speech in preference to the second speech, it is possible to reliably output a certain degree of a fixed situation explanation in a timely manner, and to output a flexible speech during a period when a situation explanation is not necessary, thereby realizing a natural live commentary voice with few silent periods.
[0022] A fourteenth embodiment of the present invention is a commentary voice generation system according to the thirteenth embodiment, in which, when output by the first voice is instructed in the first voice output instructing step while output by the second voice is being performed in the second voice output instructing step, output by the second voice is stopped and output by the first voice is performed. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner.
[0023] A live commentary voice generation system according to a fifteenth embodiment of the present invention includes a first processing step in which, when acquired situation data matches a condition for selecting first utterance data, a first voice generation step in which a first voice is generated using the first utterance data selected in the first processing step, a first voice output instruction step instructing output using the first voice generated in the first voice generation step, a voice output step instructing output using the first voice instructed in the first voice output instruction step, a second processing step in which, when an instruction is given in the first voice output instruction step, second utterance data is generated by a first AI processing, and The second voice generation step generates a second voice by the second utterance data generated in the first voice output instruction step, a third processing step generates a third utterance data by the second AI processing when an instruction is given by the first voice output instruction step, a third voice generation step generates a third voice by the third utterance data generated in the third processing step, and if the output by the first voice is not performed, a third voice output instruction step performs output by the third voice generated in the third voice generation step, and if the output by the first voice is not performed, a second voice output instruction step performs output by the second voice generated in the second voice generation step. According to this embodiment, by using the first utterance data stored in advance and the second utterance data generated by the first AI processing, it is possible to output utterances that require real-time performance and highly flexible utterances as live voices, and it is possible to understand the event situation more deeply. According to the present embodiment, by outputting the first voice in preference to the second voice, a relatively fixed situation description can be output in a timely manner and a flexible utterance can be output during a period when a situation description is not necessary, thereby realizing a natural live commentary voice with few silent periods. According to the present embodiment, the second utterance data that is repeatedly generated is generated as the third utterance data by the second AI processing, so that the third utterance data can be generated in a shorter time than the second utterance data, and flexible utterance can be easily output.Furthermore, according to this embodiment, by outputting the first voice in priority over the second and third voices, a relatively fixed situation explanation can be output in a timely manner with certainty, and flexible speech can be output during periods when a situation explanation is not necessary, thereby realizing a natural commentary voice with fewer silent periods.
[0024] A 16th embodiment of the present invention is a commentary voice generation system according to the 15th embodiment, in which if the output of the first voice is being performed, the second processing step generates second utterance data based on new situation data without outputting the second voice generated in the second voice generation step. According to this embodiment, a situation explanation that is somewhat fixed can be output in a timely manner and reliably, and highly flexible utterances can be output during periods when situation explanation is not necessary.
[0025] The seventeenth embodiment of the present invention is a commentary voice generation system according to the fifteenth embodiment, in which, if the first voice is output, the third processing step generates third utterance data based on new situation data without outputting the third voice generated in the third voice generation step. According to this embodiment, a situation explanation that is somewhat fixed can be output in a timely manner and reliably, and flexible utterances can be output during periods when situation explanation is not necessary.
[0026] The 18th embodiment of the present invention is a commentary voice generation system according to the 15th embodiment, in which if the output of the third voice is being performed, the second processing step generates second utterance data based on new situation data without outputting the second voice generated in the second voice generation step. According to this embodiment, it is possible to reliably output a situation explanation that is somewhat fixed in a timely manner, and to output a highly flexible utterance during a period when a situation explanation is not necessary.
[0027] The 19th embodiment of the present invention is a commentary voice generation system according to the 15th embodiment, in which, when output by the first voice is instructed in the first voice output instructing step while output by the second voice is being performed in the second voice output instructing step, output by the second voice is stopped and output by the first voice is performed. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner.
[0028] The twentieth embodiment of the present invention is a commentary voice generation system according to the fifteenth embodiment, in which, when output by the first voice is instructed in the first voice output instructing step while output by the third voice is being performed in the third voice output instructing step, output by the third voice is stopped and output by the first voice is performed. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner. EXAMPLES
[0029] A commentary voice generating system according to an embodiment of the present invention will be described below with reference to FIGS. FIG. 1 is a block diagram showing a commentary voice generating system according to an embodiment of the present invention, represented by functional implementation means. The live commentary sound generation system in this embodiment outputs live commentary sound regarding an event situation, and includes a first processing unit 10 and a second processing unit 20. The input unit 1 acquires situation data about the event situation, and the audio output unit 2 outputs live audio. For example, in a racing game, situation data is a number of numerical data about the circuit and the racing car, such as the coordinate position, speed, and steering angle of the racing car. More specific types and values of situation data are "1" for "current lap number [0..]", "false" for "racing or not", "256" for "lap time of current lap (ms)", "156164" for "lap time of previous lap (ms)", "30 ... For example, the following data are obtained: "0.002780" for "difference from best lap time": "0.0" for "speed (km / h)": "177.693130" for "steering wheel angle (rad)": "-59.793526" for "altitude, latitude, longitude": "5.372770,64.056038,130.219971" for "lateral position on the circuit (L=-1,R=1)": "-0.515301" for "lateral position on the circuit (L=-1,R=1)": "0.854022" for "distance from theoretical fastest course (m)". Such situation data is obtained in chronological order. In actual F1 races, situational data can include data measured in real time by sensors, and in fields other than motorsports, GPS data obtained to capture the location information of soccer players.
[0030] The first processing unit 10 selects first utterance data from the first utterance data storage unit 3 according to the situation data acquired by the input unit 1. For example, in a racing game, the first utterance data is objective event occurrence data corresponding to a situation, such as "contact" between racing cars or a racing car "going off the course." The first utterance data storage unit 3 stores the first utterance data selected by the first processing unit 10 and the condition of the situation data for selecting the first utterance data. For example, based on the coordinate positions of the two racing cars, if the distance is less than a predetermined distance, the utterance data "contact" is selected, and if the distance to the immediately preceding car is decreasing and will be less than the predetermined distance, the utterance data "caught up with the car in front" is selected. The second processing unit 20 generates second speech data by performing the first AI processing 4 using the situation data acquired by the input unit 1. Here, AI is a computer system or software that can learn, recognize, and understand, and the AI processing is processing using such a computer system or software. By using a language generation model based on a neural network trained using large-scale data, the processing speed becomes slower, but flexible and natural language can be generated. The second processing unit 20 performs a process of transmitting to the AI, together with predetermined situation data, speech condition data such as the number of characters and the speech viewpoint (such as the player's viewpoint or an aerial viewpoint) for the second speech data obtained as a result of AI processing, and a process of receiving the second speech data as a processing result by the AI.
[0031] The first voice generating unit 11 generates a first voice based on the first utterance data selected by the first processing unit . The first audio output instruction unit 12 instructs the audio output unit 2 to output the first audio generated by the first audio generation unit 11, and sends data to the second processing unit 20, the second audio output determination unit 22, and the second audio output instruction unit 23 to determine that an instruction to output the first audio has been given. The second voice generating unit 21 generates a second voice based on the second utterance data generated by the second processing unit 20. The second voice output determination unit 22 determines whether or not to output the second voice generated by the second voice generation unit 21 from the voice output unit 2. When data is received from the first voice output instruction unit 12 and the first voice is being output, it determines not to output the second voice. When the first voice is not being output, it determines to output the second voice. The second voice output instruction unit 23 issues an instruction to output the second voice generated by the second voice generation unit 21 from the voice output unit 2. When the second processing unit 20 receives data from the first voice output instruction unit 12, it starts generating second utterance data based on new situation data from the input unit 1. The second utterance data by the first AI processing 4 that is already in progress in the second processing unit 20 does not generate a second voice in the second voice generation unit 21. When the second voice output determination unit 22 receives data from the first voice output instruction unit 12 and the second voice is generated by the second voice generation unit 21 while the first voice is being output, the generated second voice is not output from the voice output unit 2. When the second audio output instruction unit 23 receives data from the first audio output instruction unit 12 while outputting the second audio, it issues an instruction to stop outputting the second audio.
[0032] In this way, by using the first utterance data stored in the first utterance data memory unit 3 and the second utterance data generated by the first AI processing unit 4, utterances that require real-time performance and highly flexible utterances can be output as live audio, allowing for a deeper understanding of the event situation. Furthermore, the audio output unit 2 outputs the first audio in priority over the second audio, thereby enabling a more or less fixed situation explanation to be output reliably in a timely manner as the first audio, and a more flexible utterance to be output as the second audio during periods when no situation explanation is necessary, thereby achieving a natural commentary with fewer silent periods. Furthermore, since the voice output unit 2 does not output the second voice while the first voice is being output, it is possible to reliably output a relatively well-defined explanation of the situation in a timely manner. Furthermore, when second utterance data is generated while the first voice is being output, the second processing unit 20 generates second utterance data based on new situation data without generating a second voice based on the generated second utterance data, thereby making it possible to reliably output a relatively well-defined situation explanation in a timely manner, and to output highly flexible utterances during periods when a situation explanation is not necessary. Furthermore, if a second voice is generated while the first voice is being output, the second processing unit 20 generates second utterance data based on new situation data without outputting the generated second voice from the voice output unit 2, thereby making it possible to reliably output a relatively well-defined situation explanation in a timely manner, and to output highly flexible utterances during periods when a situation explanation is not necessary. Furthermore, in the audio output unit 2, when outputting the first audio while outputting the second audio, output of the second audio is stopped, thereby making it possible to reliably output a relatively well-defined explanation of the situation in a timely manner. Furthermore, in the second processing unit 20, by generating the second utterance data using the output of the first voice as a trigger, it is possible to output highly flexible utterances during a period when a situation explanation is not necessary.
[0033] FIG. 2 is a flow chart showing the process flow in the live commentary voice generation system in this embodiment. The computer executes the following process in accordance with the flow chart. If the acquired situation data matches the condition for selecting the first utterance data by the first processing unit 10 (Yes in S51), the matching first utterance data is selected (first processing step). If the acquired situation data does not match the condition for selecting the first utterance data (No in S51), new situation data is acquired from the input unit 1. Note that acquisition of situation data from the input unit 1 is performed periodically at a predetermined timing. The first utterance data selected in the first processing step in S51 is used by the first voice generating unit 11 to generate the first voice (S52). The first voice generated in the first voice generating step in S52 is instructed to be output by the first voice output instructing unit 12 (S53). When an output instruction is given in the first sound output instruction step in S53, the first sound is output by the sound output unit 2 (S54).
[0034] When an output instruction is given in the first voice output instruction step in S53, the second processing unit 20 uses the situation data to generate second speech data by the first AI processing 4 (S55). The second speech data generated in the second processing step in S55 is used by the second speech generating unit 21 to generate a second speech (S56). When the second voice is generated in the second voice generation step in S56, it is determined whether or not the first voice is being output in the first voice output step in S54 (S57). If the first voice is being output (No in S57), new situation data is acquired and new second speech data is generated (S55) without outputting the second voice generated in the second voice generation step in S56. If output is not being performed using the first voice (Yes in S57), an instruction to output the second voice generated in the second voice generation step is issued (S58). When an output instruction is given in the second sound output instruction step in S58, the second sound is output by the sound output unit 2 (S59). If an output instruction is given in the first voice output instruction step in S53 while the second voice is being output in the second voice output step in S59 (Yes in S60), the second voice output is stopped by an instruction from the second voice output instruction unit 23 (S61). Therefore, the first voice is output in the first voice output step in S54. If an output instruction is not given in the first voice output instruction step in S53 while the second voice is being output in the second voice output step in S59 (No in S60), the output of the second voice by the voice output unit 2 continues (S62).
[0035] FIG. 3 is an explanatory diagram showing the process in the live commentary voice generation system in this embodiment. FIG. 3(a) shows the start of the second process. If the acquired situation data by the first processing unit 10 matches the conditions for selecting first utterance data (Yes in S51 in FIG. 2), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 2), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 2). In the second processing unit 20, second speech data is generated in response to the output of the first voice, and a second voice is generated from the generated second speech data (S56 in FIG. 2), and the generated second voice is output by the voice output unit 2 (S59 in FIG. 2). The first voice utterance time is preferably shorter than the second voice utterance time and shorter than the second utterance data generation time required to generate the second utterance data. Here, the second utterance data generation time is the data processing time by the second processing unit 20 or the time from the start of data processing by the second processing unit 20 to the generation of the second voice by the second voice generating unit 21. It is preferable that the start of data processing by the second processing unit 20 is triggered by an output instruction in the first voice output instruction step in S53, but if the second utterance data generation time is short, the trigger may be during the first voice output or at the end of the first voice output.
[0036] FIG. 3(b) shows a case where the first audio output occurs during the second process. As in FIG. 3(a), when the acquired situation data matches the conditions for selecting first utterance data by the first processing unit 10 (Yes in S51 in FIG. 2), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 2), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 2). The second processing unit 20 generates second utterance data in response to the output of the first voice, and if new first utterance data is selected by the first processing unit 10 during the time when the second utterance data is being generated, the first voice based on the new first utterance data is output by the voice output unit 2. In the second processing unit 20, in place of processing the second utterance data that was being generated, due to the occurrence of new first utterance data, new situation data is acquired and processing of the new second utterance data is started. Second utterance data is generated based on the new situation data, and a second voice based on the second utterance data is output by the voice output unit 2 (S59 in FIG. 2).
[0037] FIG. 3(c) shows a case where the first voice is output while the second voice is being output. As in FIG. 3(a), when the acquired situation data matches the conditions for selecting first utterance data by the first processing unit 10 (Yes in S51 in FIG. 2), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 2), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 2). In the second processing unit 20, second speech data is generated in response to the output of the first voice, and a second voice is generated from the generated second speech data (S56 in FIG. 2), and the generated second voice is output by the voice output unit 2 (S59 in FIG. 2). When new first utterance data is selected by the first processing unit 10 while the second voice is being output by the voice output unit 2, the output of the second voice is stopped and the first voice based on the new first utterance data is output by the voice output unit 2. In response to the occurrence of new first utterance data, the second processing unit 20 acquires new situation data and starts processing the new second utterance data. Second utterance data is generated based on the new situation data, and a second voice based on the second utterance data is output by the voice output unit 2 (S59 in FIG. 2).
[0038] Hereinafter, a commentary voice generating system according to another embodiment of the present invention will be described with reference to FIGS. 4 is a block diagram showing a commentary sound generating system according to another embodiment of the present invention, expressed by functional implementation means. The same functions as those in FIG. 1 are given the same reference numerals and the description thereof will be omitted. The live commentary voice generation system in this embodiment includes, in addition to the configuration of the live commentary voice generation system of the embodiment shown in Figures 1 to 3, a third processing unit 30 that generates third speech data based on situation data acquired by the input unit 1, and a second utterance data memory unit 5 that associates and stores the second utterance data generated by the second processing unit 20 and the situation data at the time the second utterance data is generated. Then, in the third processing unit 30, the third utterance data is generated by the second AI processing 6 using the data stored in the second utterance data storage unit 5 as teacher data. The first audio output instruction unit 12 sends to the third processing unit 30, the third audio output determination unit 32, and the third audio output instruction unit 33 data that enables it to be determined that an instruction to output the first audio has been issued. The third voice generating unit 31 generates a third voice based on the third utterance data generated by the third processing unit 30. The third voice output determination unit 32 determines whether or not to output the third voice generated by the third voice generation unit 31 from the voice output unit 2. When data is received from the first voice output instruction unit 12 and the first voice is being output, it determines not to output the third voice. When the first voice is not being output, it determines to output the third voice. The third voice output instruction unit 33 issues an instruction to output the third voice generated by the third voice generation unit 31 from the voice output unit 2 .
[0039] When the third processing unit 30 receives data from the first voice output instruction unit 12, it starts generating third utterance data based on new situation data from the input unit 1. The third utterance data by the second AI processing 6 already in progress in the third processing unit 30 does not generate a third voice in the third voice generation unit 31. When the third audio output determination unit 32 receives data from the first audio output instruction unit 12 and the third audio generation unit 31 generates a third audio while the first audio is being output, the generated third audio is not output from the audio output unit 2. When the third sound output instruction unit 33 receives data from the first sound output instruction unit 12 while outputting the third sound, it issues an instruction to stop outputting the third sound.
[0040] In this way, the second utterance data repeatedly generated is generated as the third utterance data by the second AI processing 6, so that the third utterance data can be generated in a shorter time than the time required to generate the second utterance data, and it becomes easier to output highly flexible utterances. Also, by outputting the first voice in preference to the second voice and the third voice, a relatively fixed situation explanation can be output in a timely manner and with certainty, and highly flexible utterances can be output during periods when situation explanation is not necessary, so that a natural live commentary with fewer silent periods can be realized. Since the voice output unit 2 does not output the third voice while the first voice is being output, it is possible to reliably output a relatively well-defined explanation of the situation in a timely manner. Furthermore, when the third utterance data is generated while the first voice is being output, the third processing unit 30 generates the third utterance data based on new situation data without generating a third voice based on the generated third utterance data, thereby making it possible to reliably output a relatively well-defined situation explanation in a timely manner, and to output highly flexible utterances during periods when a situation explanation is not necessary. Furthermore, if a third voice is generated while the first voice is being output, the third processing unit 30 generates third utterance data based on new situation data without outputting the generated third voice from the voice output unit 2, thereby making it possible to reliably output a relatively well-defined situation explanation in a timely manner, and to output highly flexible utterances during periods when a situation explanation is not necessary. Furthermore, in the audio output unit 2, when outputting the first audio while outputting the third audio, output of the third audio is stopped, thereby making it possible to reliably output a certain degree of a defined situation explanation in a timely manner. Furthermore, in the third processing unit 30, generating the third utterance data is triggered by the output of the first voice, so that highly flexible utterances can be output during a period when a situation explanation is not required.
[0041] 5 is a flow chart showing the process flow in the live commentary voice generation system in this embodiment. The same processes as those in FIG. 2 are assigned the same reference numerals and the description thereof will be omitted. The computer executes the following process in accordance with the flow chart. The live commentary voice generation system in this embodiment performs the following process in addition to the process of the live commentary voice generation system according to the embodiment shown in FIGS. When an output instruction is given in the first voice output instruction step in S53, the third processing unit 30 uses the situation data to generate third utterance data by the second AI processing 6 (S63). The third utterance data generated in the third processing step in S63 is used by the third voice generating unit 31 to generate a third voice (S64). When the third voice is generated in the third voice generation step in S64, it is determined whether or not the first voice is being output in the first voice output step in S54 (S65). If the first voice is being output (No in S65), new situation data is acquired and new third speech data is generated (S63) without outputting the third voice generated in the third voice generation step in S64. If output is not being performed using the first voice (Yes in S65), an instruction to output the third voice generated in the third voice generation step is issued (S66). When an output instruction is given in the third sound output instruction step in S66, the third sound is output by the sound output unit 2 (S67). If an output instruction is given in the first audio output instruction step in S53 while the third audio is being output in the third audio output step in S67 (Yes in S68), the third audio output is stopped by an instruction from the third audio output instruction unit 33 (S69). Therefore, the first audio is output in the first audio output step in S54. If an output instruction is not given in the first voice output instruction step in S53 while the third voice is being output in the third voice output step in S67 (No in S68), the output of the third voice by the voice output unit 2 continues (S70).
[0042] 6 to 8 are explanatory diagrams showing the processing in the live commentary voice generation system in this embodiment. FIG. 6 shows the relationship between the second process and the third process. FIG. 6(a) shows a case where the third utterance data is not generated as a result of the third processing, and FIG. 6(b) shows a case where the third utterance data is generated as a result of the third processing. The third utterance data generation time required to generate the third utterance data is longer than the first voice utterance time, and shorter than the second utterance data generation time required to generate the second utterance data. This is because, although both the second processing unit 20 and the third processing unit 30 perform AI processing, the third processing unit 30 performs AI processing based on the teacher data stored in the second utterance data storage unit 5. Note that, in the third processing unit 30, if the probability of the match rate between the generated third utterance data and the teacher data is equal to or lower than a predetermined value, the third utterance data is not generated. In this way, by not generating the third utterance data for data that has a low match rate with the teacher data as a result of the AI processing, it is possible to prevent the generation of erroneous live voice.
[0043] If the acquired situation data by the first processing unit 10 matches the conditions for selecting first utterance data (Yes in S51 in FIG. 5), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 5), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 5). The second processing unit 20 generates the second utterance data in response to the output of the first voice, and the third processing unit 30 generates the third utterance data in response to the output of the first voice.
[0044] As shown in Figure 6(a), if the third processing unit 30 does not generate the third utterance data, a second voice is generated from the second utterance data generated by the second processing unit 20 (S56 in Figure 5), and the generated second voice is output by the voice output unit 2 (S59 in Figure 5). As shown in Figure 6 (b), when third speech data is generated by the third processing unit 30, a third voice is generated from the third speech data generated by the third processing unit 30 (S64 in Figure 5), and the generated third voice is output by the voice output unit 2 (S67 in Figure 5). The third utterance data generation time is the data processing time by the third processing unit 30, or the time from the start of data processing by the third processing unit 30 to the generation of the third voice by the third voice generating unit 31. It is preferable that the start of data processing by the third processing unit 30 is triggered by an output instruction in the first voice output instruction step in S53, but if the third utterance data generation time is short, the trigger may be during the first voice output or at the end of the first voice output.
[0045] FIG. 7 shows a case where the first audio output occurs during the second process. FIG. 7(a) shows a case where the third utterance data is not generated as a result of the third processing, and FIG. 7(b) shows a case where the third utterance data is generated as a result of the third processing. As in FIG. 6, when the acquired situation data matches the conditions for selecting first utterance data by the first processing unit 10 (Yes in S51 in FIG. 5), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 5), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 5). The second processing unit 20 generates second utterance data in response to the output of the first voice, and if new first utterance data is selected by the first processing unit 10 during the time when the second utterance data is being generated, the first voice based on the new first utterance data is output by the voice output unit 2. In the second processing unit 20, in place of processing the second utterance data that was being generated, due to the occurrence of new first utterance data, new situation data is acquired and processing of the new second utterance data is started. Like the second processing unit 20, the third processing unit 30 generates third speech data in response to the output of the first voice, but if new first speech data is selected by the first processing unit 10 during the generation of the third speech data, the first voice based on the new first speech data is output by the voice output unit 2. In the third processing unit 30, in place of processing the third utterance data that was being generated, due to the occurrence of new first utterance data, new situation data is acquired and processing of the new third utterance data is started.
[0046] Then, as shown in FIG. 7(a), if the third processing unit 30 does not generate the third utterance data, a second voice is generated from the second utterance data generated by the second processing unit 20 (S56 in FIG. 5), and the generated second voice is output by the voice output unit 2 (S59 in FIG. 5). Also, as shown in FIG. 7(b), when third speech data is generated by the third processing unit 30, a third voice is generated from the third speech data generated by the third processing unit 30 (S64 in FIG. 5), and the generated third voice is output by the voice output unit 2 (S67 in FIG. 5).
[0047] FIG. 8 shows a case where the first voice is output while the third voice is being output. If the acquired situation data by the first processing unit 10 matches the conditions for selecting first utterance data (Yes in S51 in FIG. 5), the matching first utterance data is selected, a first voice is generated from the selected first utterance data by the first voice generation unit 11 (S52 in FIG. 5), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 5). In the third processing unit 30, third speech data is generated in response to the output of the first voice, and a third voice is generated from the generated third speech data (S64 in FIG. 5), and the generated third voice is output by the voice output unit 2 (S67 in FIG. 5). If new first utterance data is selected by the first processing unit 10 while the third voice is being output by the voice output unit 2, the output of the third voice is stopped and the first voice based on the new first utterance data is output by the voice output unit 2. In response to the occurrence of new first utterance data, the third processing unit 30 acquires new situation data and starts processing new third utterance data. Third utterance data is generated based on the new situation data, and a third voice based on the third utterance data is output by the voice output unit 2 (S67 in FIG. 5). [Industrial Applicability]
[0048] The present invention is particularly suited to racing game footage as an event situation, but can also be applied to other video game footage and sports footage that generates live commentary of swimming competitions and the like in real time. [Explanation of symbols]
[0049] 1 Input section 2 Audio output section 3. First speech data storage unit 4. First AI Processing 5. Second speech data storage unit 6 Second AI processing 10 First Processing Section 11 First voice generation unit 12 First audio output instruction unit 20 Second Processing Section 21 Second voice generation unit 22 Second audio output determination unit 23 Second audio output instruction unit 30 Third Processing Section 31 Third voice generation unit 32 Third audio output determination unit 33 Third audio output instruction unit
Claims
1. A live commentary sound generation system for outputting live commentary sound about an event situation, an input unit for acquiring situation data regarding the event situation; a first processing unit that selects first utterance data according to the situation data acquired by the input unit; a second processing unit that generates second utterance data based on the situation data acquired by the input unit; a first utterance data storage unit that stores the first utterance data selected by the first processing unit and a condition of the situation data for selecting the first utterance data; a voice generating unit that generates a first voice based on the first utterance data and a second voice based on the second utterance data; a voice output unit that outputs the first voice and the second voice generated by the voice generation unit; Equipped with The first processing unit selects the first utterance data from the first utterance data storage unit, The second processing unit generates the second speech data by a first AI process, The audio output unit outputs the first audio in priority to the second audio. A live voice generation system comprising:
2. The audio output unit does not output the second audio while the first audio is being output.
2. The live commentary sound generation system according to claim 1 .
3. When the second utterance data is generated while the first voice is being output, the second processing unit generates the second utterance data based on new situation data without generating the second voice based on the generated second utterance data.
2. The live commentary sound generation system according to claim 1 .
4. When the second voice is generated while the first voice is being output, the second processing unit generates the second speech data based on new situation data without outputting the generated second voice from the voice output unit.
2. The live commentary sound generation system according to claim 1 .
5. When the audio output unit is outputting the first audio while the second audio is being output, the output of the second audio is stopped.
2. The live commentary sound generation system according to claim 1 .
6. The second processing unit generates the second speech data in response to the output of the first voice.
2. The live commentary sound generation system according to claim 1 .
7. a third processing unit that generates third utterance data based on the situation data acquired by the input unit; a second utterance data storage unit that stores the second utterance data generated by the second processing unit and the situation data at the time when the second utterance data is generated in association with each other; Equipped with the speech generation unit generates a third speech based on the third speech data; The audio output unit outputs the third audio generated by the audio generation unit; The third processing unit generates the third utterance data by a second AI process using the data stored in the second utterance data storage unit as teacher data, The audio output unit outputs the first audio in priority to the second audio and the third audio.
2. The live commentary sound generation system according to claim 1 .
8. The audio output unit does not output the third audio while the first audio is being output.
8. The live commentary sound generation system according to claim 7.
9. When the third utterance data is generated while the first voice is being output, the third processing unit generates the third utterance data based on new situation data without generating the third voice based on the generated third utterance data.
8. The live commentary sound generation system according to claim 7.
10. When the third voice is generated while the first voice is being output, the third processing unit generates the third speech data based on new situation data without outputting the generated third voice from the voice output unit.
8. The live commentary sound generation system according to claim 7.
11. When the audio output unit is outputting the first audio while the third audio is being output, the output of the third audio is stopped.
8. The live commentary sound generation system according to claim 7.
12. The third processing unit generates the third speech data in response to the output of the first voice.
8. The live commentary sound generation system according to claim 7.
13. A live commentary sound generation system for outputting live commentary sound about an event situation, The computer a first processing step of selecting the first utterance data when the acquired situation data matches a condition for selecting the first utterance data; a first speech generating step of generating a first speech based on the first speech data selected in the first processing step; a first voice output instruction step of instructing output of the first voice generated in the first voice generating step; a voice output step of outputting the first voice instructed in the first voice output instructing step; a second processing step of generating second speech data by a first AI processing step when an instruction is given by the first voice output instruction step; a second speech generating step of generating a second speech based on the second speech data generated in the second processing step; a second voice output instruction step of instructing output by the second voice generated in the second voice generating step if output by the first voice is not being performed; Run If the output of the first voice is being performed, the second processing step generates the second speech data based on new situation data without performing the output of the second voice generated in the second voice generating step. A live voice generation system comprising:
14. When the output by the first voice is instructed in the first voice output instructing step while the output by the second voice is being performed in the second voice output instructing step, the output by the second voice is stopped and the output by the first voice is performed. The live audio generation system according to claim 13 .
15. A live commentary sound generation system for outputting live commentary sound about an event situation, The computer a first processing step of selecting the first utterance data when the acquired situation data matches a condition for selecting the first utterance data; a first speech generating step of generating a first speech based on the first speech data selected in the first processing step; a first voice output instruction step of instructing output of the first voice generated in the first voice generating step; a voice output step of outputting the first voice instructed in the first voice output instructing step; a second processing step of generating second speech data by a first AI processing step when an instruction is given by the first voice output instruction step; a second speech generating step of generating a second speech based on the second speech data generated in the second processing step; a third processing step of generating third speech data by a second AI processing step when an instruction is given by the first voice output instruction step; a third speech generating step of generating a third speech based on the third speech data generated in the third processing step; a third voice output instruction step of outputting the third voice generated in the third voice generating step if the output of the first voice is not being performed; a second voice output instruction step of outputting the second voice generated in the second voice generating step if the output using the first voice and the output using the third voice have not been performed; Run A live voice generation system comprising:
16. If the output of the first voice is being performed, the second processing step generates the second speech data based on new situation data without performing the output of the second voice generated in the second voice generating step.
16. The live audio generation system according to claim 15.
17. If the output of the first voice is being performed, the third processing step generates the third speech data based on new situation data without performing the output of the third voice generated in the third voice generating step.
16. The live audio generation system according to claim 15.
18. If the output of the third voice is being performed, the second processing step generates the second speech data based on new situation data without performing the output of the second voice generated in the second voice generating step.
16. The live audio generation system according to claim 15.
19. When the output by the first voice is instructed in the first voice output instructing step while the output by the second voice is being performed in the second voice output instructing step, the output by the second voice is stopped and the output by the first voice is performed.
16. The live audio generation system according to claim 15.
20. When the output by the first voice is instructed in the first voice output instructing step while the output by the third voice is being performed in the third voice output instructing step, the output by the third voice is stopped and the output by the first voice is performed.
16. The live audio generation system according to claim 15.
Citation Information
Patent Citations
On the scene broadcasting device for games
JP1996215433A
Game apparatus, method for controlling game, and game control program
JP2013111178A
Computer system and audio information generation method
JP2021194229A