Live audio generation system
The live audio generation system addresses the challenge of generating real-time and flexible live commentary by using status data to select and generate speech, prioritizing timely explanations and allowing for natural, engaging commentary.
Patent Information
- Application Number
- PCT/JP2024/029371
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-08
AI Technical Summary
Existing technologies struggle to generate live commentary that is both real-time and flexible, particularly in event situations like sports and video games, due to challenges in recognizing video content and integrating subjective comments effectively.
A live audio generation system that utilizes input units to acquire status data, with processing units selecting and generating speech data based on this information. The system prioritizes first speech data for timely situation explanations and generates second speech data flexibly using AI processes, allowing for natural commentary with minimal silence.
The system achieves real-time and highly flexible live commentary, providing deeper understanding of event situations while ensuring timely and reliable fixed situation explanations, resulting in a natural and engaging commentary experience.
Smart Images

Figure JP2024029371_08052025_PF_FP_ABST
Abstract
Description
Live voice generation system
[0001] The present invention relates to a live audio generation system that outputs live audio about an event situation.
[0002] Sports videos, including motorsports like Formula 1, often include commentary audio, such as "Next is the final straight. Can you overtake?", providing viewers with situational explanations and subjective comments from the commentator. By listening to the commentary while watching the video, viewers gain a deeper understanding of the situation and enjoy the event more. Commentary is often added to many sports and video game videos, playing an important role in entertaining viewers and potentially enhancing the value of the video itself. However, commentary requires knowledge of the sport or event in question and appropriate speaking skills, so most online sports and video game videos do not include commentary. As an example of commentary, we focus on commentary for racing game videos. Such commentary requires the commentator to speak about important events occurring in the video at the appropriate time within a short period of time. Previous language generation research has primarily focused on the problem of planning (what to say) and the problem of language surfaceization (how to say it). In addition to these traditionally addressed issues, commentary generation must address previously unaccounted for issues such as when to speak, how long to speak, and how detailed to provide the information. Identifying the timing of utterances and generating commentary utterances requires considering time-series data, such as video. This differs from language generation settings that do not consider a time axis, such as image caption generation, or problem settings such as story generation, where the timing of utterances is given in advance. In conventional language generation research, many settings use visual information, such as video and images, as input. However, accurately recognizing video for commentary generation is not easy. For example, commentary often includes detailed descriptions of vehicle positional relationships, and it is necessary to recognize subtle differences in vehicle positional relationships from a series of similar video frames, such as aerial footage. On the other hand, image caption generation assumes the output of text that does not necessarily capture subtle differences between objects, such as "people are dancing." Given this background, using time-series situation data in addition to video as input is effective for more accurately capturing the racing situation.Situation data includes multiple numerical data related to the circuit and the racing car, such as the coordinate position, speed, and steering angle of the racing car. In actual F1 races, more than 300 types of numerical data are measured in real time from multiple sensors. Even in fields other than motorsports, the location information of soccer players is obtained via GPS, and attempts to utilize situation data in various sports are spreading. Meanwhile, Patent Document 1 proposes a live game broadcasting device that automatically broadcasts a live commentary corresponding to the game's development. In the live game broadcasting device of Patent Document 1, predetermined live commentary audio data corresponding to multiple game development patterns is stored in an audio data storage means, the game system determines the game development pattern, and outputs a command to read the audio data corresponding to the game development pattern. Patent Document 2, like the live game broadcasting device of Patent Document 1, pre-stores commentary data that defines commentary content according to the game's progress. However, Patent Document 2 proposes a device that determines at least two commentary axes that match the focus of the game content and determines the commentary content according to the commentary data corresponding to these commentary axes, thereby realizing multifaceted commentary, such as reporting the game from multiple perspectives, like a real announcer. Patent Document 3 proposes a system that has an audio information generation unit configured to function based on a trained machine learning model, and that can output audio such as commentary in accordance with the game situation and the flow of conversation such as commentary.
[0003] Japanese Patent Application Laid-Open No. 8-215433 Japanese Patent Application Laid-Open No. 2013-111178 Japanese Patent Application Laid-Open No. 2021-194229
[0004] When live commentary data is stored in advance, as in Patent Document 1 and Patent Document 2, it is excellent in real-time performance, but it is difficult to have the commentator speak subjective comments along with a situation explanation, such as "Next is the final straight zone. Can you overtake?" On the other hand, when voice information is generated based on a trained model of machine learning, as in Patent Document 3, it is possible to have the commentator speak subjective comments along with a situation explanation, but it is inferior in real-time performance.
[0005] The present invention aims to provide a live voice generation system that can realize real-time speech and highly flexible speech by processing speech that is somewhat fixed in accordance with the situation and speech using flexible language differently.
[0006] The live audio generation system of the present invention described in claim 1 is a live audio generation system that outputs live audio about an event situation, and comprises an input unit 1 that acquires situation data about the event situation, a first processing unit 10 that selects first utterance data based on the situation data acquired by the input unit 1, a second processing unit 20 that generates second utterance data based on the situation data acquired by the input unit 1, a first utterance data storage unit 3 that stores the first utterance data selected by the first processing unit 10 and the situation data conditions for selecting the first utterance data, audio generation units 11, 21 that generate a first voice based on the first utterance data and a second voice based on the second utterance data, and an audio output unit 2 that outputs the first voice and the second voice generated by the audio generation units 11, 21, wherein the first processing unit 10 selects the first utterance data from the first utterance data storage unit 3, the second processing unit 20 generates the second utterance data by a first AI processing 4, and the audio output unit 2 outputs the first voice in priority to the second voice. The present invention of claim 2 is characterized in that, in the live commentary speech generation system of claim 1, the audio output unit 2 does not output the second audio while the first audio is being output. The present invention of claim 3 is characterized in that, in the live commentary speech generation system of claim 1, when the second utterance data is generated while the first audio is being output, the second processing unit 20 generates the second utterance data based on new situation data without generating the second audio based on the generated second utterance data. The present invention of claim 4 is characterized in that, in the live commentary speech generation system of claim 1, when the second audio is generated while the first audio is being output, the second processing unit 20 generates the second utterance data based on the new situation data without outputting the generated second audio from the audio output unit 2. The present invention of claim 5 is characterized in that, in the live commentary speech generation system of claim 1, when the audio output unit 2 outputs the first audio while outputting the second audio, the output of the second audio is stopped.The present invention of claim 6 is characterized in that, in the live commentary voice generation system of claim 1, the second processing unit 20 generates the second utterance data triggered by output of the first voice. The present invention of claim 7 is the live commentary voice generation system of claim 1, further comprising: a third processing unit 30 that generates third utterance data based on the situation data acquired by the input unit 1; and a second utterance data storage unit 5 that associates and stores the second utterance data generated by the second processing unit 20 with the situation data at the time the second utterance data was generated, wherein the voice generation unit 31 generates a third voice based on the third utterance data, the voice output unit 2 outputs the third voice generated by the voice generation unit 31, the third processing unit 30 generates the third utterance data by second AI processing 6 using data stored in the second utterance data storage unit 5 as training data, and the voice output unit 2 outputs the first voice in priority to the second voice and the third voice. The present invention of claim 8 is characterized in that, in the live commentary speech generation system of claim 7, the audio output unit 2 does not output the third audio while the first audio is being output. The present invention of claim 9 is characterized in that, in the live commentary speech generation system of claim 7, when the third utterance data is generated while the first audio is being output, the third processing unit 30 generates the third utterance data based on new situation data without generating the third audio based on the generated third utterance data. The present invention of claim 10 is characterized in that, in the live commentary speech generation system of claim 7, when the third audio is generated while the first audio is being output, the third processing unit 30 generates the third utterance data based on the new situation data without outputting the generated third audio from the audio output unit 2. The present invention of claim 11 is characterized in that, in the live commentary speech generation system of claim 7, when the audio output unit 2 outputs the first audio while outputting the third audio, the output of the third audio is stopped.The present invention as set forth in claim 12 is characterized in that in the live commentary voice generation system as set forth in claim 7, the third processing unit 30 generates the third speech data using the output of the first speech as a trigger. The live commentary voice generation system of the present invention as set forth in claim 13 is a live commentary voice generation system that outputs live commentary voice about an event situation, and includes a first processing step of selecting the matching first speech data when acquired situation data matches a condition for selecting first speech data, a first speech generation step of generating first speech using the first speech data selected in the first processing step, a first speech output instruction step of instructing output using the first speech generated in the first speech generation step, a speech output step of outputting the first speech instructed in the first speech output instruction step, and a first speech output instruction step of outputting the first speech instructed in the first speech output instruction step. When an instruction is given by the display step, the system executes a second processing step of generating second utterance data by a first AI processing 4, a second voice generation step of generating a second voice using the second utterance data generated in the second processing step, and a second voice output instruction step of issuing an output instruction using the second voice generated in the second voice generation step if output using the first voice is not being performed, and if output using the first voice is being performed, the system generates the second utterance data using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. The present invention as set forth in claim 14 is characterized in that, in the live voice generation system as set forth in claim 13, when output using the first voice is instructed in the first voice output instruction step while output using the second voice is being performed in the second voice output instruction step, output using the second voice is stopped and output using the first voice is performed.A live voice generation system of the present invention as set forth in claim 15 is a live voice generation system that outputs live voice about an event situation, and includes a first processing step in which, when acquired situation data matches a condition for selecting first utterance data, a computer selects the matching first utterance data; a first voice generation step in which a first voice is generated using the first utterance data selected in the first processing step; a first voice output instruction step in which an instruction is given to output the first voice generated in the first voice generation step; a voice output step in which the first voice instructed in the first voice output instruction step is output; and a second processing step in which a first AI process 4 generates second utterance data when an instruction is given in the first voice output instruction step. a second voice generation step of generating a second voice using the second utterance data generated in the second processing step, a third processing step of generating third utterance data by second AI processing 6 when an instruction is given in the first voice output instruction step, a third voice generation step of generating a third voice using the third utterance data generated in the third processing step, a third voice output instruction step of outputting the third voice generated in the third voice generation step if output using the first voice has not been performed, and a second voice output instruction step of outputting the first voice and, if output using the third voice has not been performed, outputting the second voice generated in the second voice generation step. The present invention as set forth in claim 16 is characterized in that, in the live voice generation system as set forth in claim 15, if output using the first voice is being performed, the second utterance data is generated using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. The present invention described in claim 17 is characterized in that, in the live audio generation system described in claim 15, if output is being performed using the first audio, the third speech data is generated using new situation data in the third processing step without outputting the third audio generated in the third audio generation step.The present invention of claim 18 is characterized in that, in the live commentary voice generation system of claim 15, if output using the third voice is being performed, output using the second voice generated in the second voice generation step is not performed, and the second utterance data is generated based on new situation data in the second processing step. The present invention of claim 19 is characterized in that, in the live commentary voice generation system of claim 15, if output using the first voice is instructed in the first voice output instructing step while output using the second voice is being performed in the second voice output instructing step, output using the second voice is stopped and output using the first voice. The present invention of claim 20 is characterized in that, in the live commentary voice generation system of claim 15, if output using the first voice is instructed in the first voice output instructing step while output using the third voice is being performed in the third voice output instructing step, output using the third voice is stopped and output using the first voice.
[0007] According to the present invention, by using the first speech data and the second speech data, it is possible to output speech that requires real-time processing and highly flexible speech as live commentary audio, allowing for a deeper understanding of the event situation.By outputting the first speech in priority over the second speech, it is possible to reliably output a fairly fixed explanation of the situation in a timely manner, and highly flexible speech can be output during periods when no explanation of the situation is necessary, thereby achieving natural live commentary audio with fewer periods of silence.
[0008] 1 is a block diagram of a live commentary sound generation system according to an embodiment of the present invention, expressed by means of functional realization; 2 is a flow diagram showing the processing flow in the live commentary sound generation system; 3 is an explanatory diagram showing the processing in the live commentary sound generation system; 4 is a block diagram of a live commentary sound generation system according to another embodiment of the present invention, expressed by means of functional realization; 5 is a flow diagram showing the processing flow in the live commentary sound generation system; 6 is an explanatory diagram showing the relationship between the second processing and the third processing in the live commentary sound generation system; 7 is an explanatory diagram showing the case where a first audio output occurs during the second processing in the live commentary sound generation system;
[0009] A live voice generation system according to a first embodiment of the present invention comprises an input unit that acquires situation data regarding an event situation, a first processing unit that selects first utterance data based on the situation data acquired by the input unit, a second processing unit that generates second utterance data based on the situation data acquired by the input unit, a first utterance data storage unit that stores the first utterance data selected by the first processing unit and the situation data conditions for selecting the first utterance data, a voice generation unit that generates a first voice based on the first utterance data and a second voice based on the second utterance data, and a voice output unit that outputs the first voice and the second voice generated by the voice generation unit, wherein the first processing unit selects the first utterance data from the first utterance data storage unit, the second processing unit generates the second utterance data by first AI processing, and the voice output unit outputs the first voice in priority to the second voice. According to this embodiment, by using the first utterance data stored in the first utterance data storage unit and the second utterance data generated by the first AI processing, it is possible to output utterances that require real-time processing and highly flexible utterances as live commentary audio, thereby enabling a deeper understanding of the event situation. Furthermore, according to this embodiment, by outputting the first audio prior to the second audio, it is possible to reliably output a fairly fixed situation explanation in a timely manner and output highly flexible utterances during periods when no situation explanation is necessary, thereby realizing a natural live commentary audio with few silent periods.
[0010] In the second embodiment of the present invention, in the live voice generation system according to the first embodiment, the voice output unit does not output the second voice while the first voice is being output. According to this embodiment, it is possible to reliably output a certain amount of defined situation explanation in a timely manner.
[0011] In a third embodiment of the present invention, in the live speech generation system according to the first embodiment, when second utterance data is generated while the first speech is being output, the second processing unit generates second utterance data based on new situation data without generating the second speech based on the generated second utterance data. According to this embodiment, it is possible to reliably output a certain degree of situation description in a timely manner, and it is also possible to output highly flexible utterances during periods when a situation description is not necessary.
[0012] In a fourth embodiment of the present invention, in the live voice generation system according to the first embodiment, when a second voice is generated while a first voice is being output, the generated second voice is not output from the voice output unit, and the second processing unit generates second utterance data based on new situation data. According to this embodiment, it is possible to reliably output a certain degree of situation description in a timely manner, and it is also possible to output highly flexible utterances during periods when a situation description is not necessary.
[0013] In the fifth embodiment of the present invention, in the live audio generation system according to the first embodiment, the audio output unit stops outputting the second audio when outputting the first audio while the second audio is being output. According to this embodiment, it is possible to reliably output a certain amount of defined situation explanation in a timely manner.
[0014] In a sixth embodiment of the present invention, in the live speech generation system according to the first embodiment, the second processing unit generates second speech data in response to the output of the first speech as a trigger. According to this embodiment, highly flexible speech can be output during a period when a situation explanation is not required.
[0015] A seventh embodiment of the present invention provides a live commentary speech generation system according to the first embodiment, further comprising: a third processing unit that generates third utterance data using situation data acquired by the input unit; and a second utterance data storage unit that stores the second utterance data generated by the second processing unit and situation data at the time the second utterance data was generated in association with each other, wherein the sound generation unit generates a third voice using the third utterance data, the sound output unit outputs the third voice generated by the sound generation unit, the third processing unit generates the third utterance data by a second AI process using data stored in the second utterance data storage unit as training data, and the sound output unit outputs the first voice in priority to the second voice and the third voice. According to this embodiment, by generating the repeatedly generated second utterance data as third utterance data by the second AI process, the third utterance data can be generated in a shorter time than the generation time of the second utterance data, making it easier to output highly flexible speech. Furthermore, according to this embodiment, by outputting the first voice in priority to the second and third voices, a more or less fixed explanation of the situation can be output in a timely manner with reliability, and highly flexible speech can be output during periods when no explanation of the situation is necessary, thereby realizing natural commentary audio with fewer periods of silence.
[0016] In the eighth embodiment of the present invention, in the live voice generation system according to the seventh embodiment, the voice output unit does not output the third voice while outputting the first voice. According to this embodiment, it is possible to reliably output a certain amount of defined situation explanation in a timely manner.
[0017] In a ninth embodiment of the present invention, in the live speech generation system according to the seventh embodiment, when third utterance data is generated while the first speech is being output, the third processing unit generates third utterance data based on new situation data without generating the third speech based on the generated third utterance data. According to this embodiment, it is possible to reliably output a certain degree of situation description in a timely manner, and it is also possible to output highly flexible utterances during periods when a situation description is not necessary.
[0018] In a tenth embodiment of the present invention, in the live speech generation system according to the seventh embodiment, when a third speech is generated while the first speech is being output, the generated third speech is not output from the speech output unit, and the third processing unit generates third utterance data based on new situation data. According to this embodiment, it is possible to reliably output a certain degree of situation description in a timely manner, and it is also possible to output highly flexible utterances during periods when a situation description is not necessary.
[0019] In the eleventh embodiment of the present invention, in the live voice generation system according to the seventh embodiment, the voice output unit stops outputting the third voice when outputting the first voice while the third voice is being output. According to this embodiment, it is possible to reliably output a certain amount of defined situation explanation in a timely manner.
[0020] A twelfth embodiment of the present invention is a commentary speech generation system according to the seventh embodiment, in which the third processing unit generates third speech data in response to the output of the first speech as a trigger. According to this embodiment, highly flexible speech can be output during periods when a situation explanation is not required.
[0021] A live voice generation system according to a thirteenth embodiment of the present invention executes the following steps: a first processing step in which, when acquired situation data matches the conditions for selecting first utterance data, a computer selects the matching first utterance data; a first voice generation step in which a first voice is generated using the first utterance data selected in the first processing step; a first voice output instruction step in which an instruction is given to output the first voice generated in the first voice generation step; a voice output step in which the first voice instructed in the first voice output instruction step is output; a second processing step in which, when an instruction is given in the first voice output instruction step, a second voice generation step in which a second voice is generated using the second utterance data generated in the second processing step; and a second voice output instruction step in which, if output using the first voice is not being performed, an instruction is given to output using the second voice generated in the second voice generation step; and, if output using the first voice is being performed, second utterance data is generated using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. According to this embodiment, by using the first speech data stored in advance and the second speech data generated by the first AI processing, it is possible to output speech that requires real-time processing and highly flexible speech as commentary audio, thereby enabling a deeper understanding of the event situation. Furthermore, according to this embodiment, by outputting the first speech prior to the second speech, it is possible to reliably output a fairly fixed situation explanation in a timely manner and output highly flexible speech during periods when no situation explanation is necessary, thereby realizing a natural commentary audio with few silent periods.
[0022] A fourteenth embodiment of the present invention is a commentary voice generation system according to the thirteenth embodiment, in which, when an instruction to output the first voice is given in the first voice output instruction step while outputting the second voice is being given in the second voice output instruction step, outputting the second voice is stopped and outputting the first voice. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner.
[0023] A live voice generation system according to a fifteenth embodiment of the present invention includes a first processing step in which, when acquired situation data matches a condition for selecting first utterance data, a computer selects matching first utterance data; a first voice generation step in which a first voice is generated using the first utterance data selected in the first processing step; a first voice output instruction step instructing output using the first voice generated in the first voice generation step; a voice output step instructing output of the first voice instructed in the first voice output instruction step; a second processing step in which, when an instruction is given in the first voice output instruction step, second utterance data is generated by first AI processing; The present embodiment executes the following steps: a second voice generation step of generating a second voice using second utterance data generated in the first voice output instruction step; a third processing step of generating third utterance data by second AI processing when an instruction is given in the first voice output instruction step; a third voice generation step of generating a third voice using the third utterance data generated in the third processing step; a third voice output instruction step of outputting the third voice generated in the third voice generation step if output by the first voice has not been performed; and a second voice output instruction step of outputting the first voice and, if output by the third voice has not been performed, outputting the second voice generated in the second voice generation step. According to this embodiment, by using the first utterance data stored in advance and the second utterance data generated by the first AI processing, it is possible to output utterances that require real-time processing and highly flexible utterances as live voice, thereby enabling a deeper understanding of the event situation. Furthermore, according to this embodiment, by outputting the first voice prior to the second voice, a fairly fixed situation description can be reliably output in a timely manner, and highly flexible utterances can be output during periods when no situation description is necessary, thereby realizing a natural live commentary voice with fewer silent periods. Furthermore, according to this embodiment, the repeatedly generated second utterance data can be generated as third utterance data by the second AI processing, thereby enabling the third utterance data to be generated in a shorter time than the generation time of the second utterance data, making it easier to output highly flexible utterances.Furthermore, according to this embodiment, by outputting the first voice in priority to the second and third voices, a more or less fixed explanation of the situation can be output in a timely manner with reliability, and highly flexible speech can be output during periods when no explanation of the situation is necessary, thereby realizing natural commentary audio with fewer periods of silence.
[0024] A sixteenth embodiment of the present invention is a commentary speech generation system according to the fifteenth embodiment, in which, if output using the first voice is being performed, second utterance data is generated using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. According to this embodiment, it is possible to reliably output a situation description that is somewhat fixed in a timely manner, and to output highly flexible utterances during periods when situation explanation is not necessary.
[0025] A seventeenth embodiment of the present invention is a commentary speech generation system according to the fifteenth embodiment, in which, if output using the first voice is being performed, third utterance data is generated using new situation data in the third processing step without outputting the third voice generated in the third voice generation step. According to this embodiment, it is possible to reliably output a situation description that is fairly well-defined in a timely manner, and to output highly flexible utterances during periods when a situation description is not necessary.
[0026] In the eighteenth embodiment of the present invention, in the live commentary speech generation system according to the fifteenth embodiment, if output using the third voice is being performed, second utterance data is generated using new situation data in the second processing step without output using the second voice generated in the second voice generation step. According to this embodiment, it is possible to reliably output a situation description that is somewhat fixed in a timely manner, and it is also possible to output highly flexible utterances during periods when a situation explanation is not necessary.
[0027] A 19th embodiment of the present invention is a commentary voice generation system according to the 15th embodiment, in which, if an instruction to output the first voice is given in the first voice output instruction step while outputting the second voice is being given in the second voice output instruction step, outputting the second voice is stopped and outputting the first voice. According to this embodiment, a situation explanation that is determined to a certain extent can be reliably output in a timely manner.
[0028] A twentieth embodiment of the present invention is a commentary voice generation system according to the fifteenth embodiment, in which, if an instruction to output the first voice is given in the first voice output instruction step while outputting the third voice in the third voice output instruction step, outputting the third voice is stopped and outputting the first voice. According to this embodiment, a situation explanation that is somewhat fixed can be reliably output in a timely manner.
[0029] A live commentary audio generation system according to an embodiment of the present invention will be described below with reference to Figures 1 to 3. Figure 1 is a block diagram showing a live commentary audio generation system according to an embodiment of the present invention, represented by functional implementation means. The live commentary audio generation system according to this embodiment outputs live commentary audio about an event situation, and includes a first processing unit 10 and a second processing unit 20. An input unit 1 acquires situation data about the event situation, and an audio output unit 2 outputs live commentary audio. For example, in a racing game, the situation data is a plurality of numerical data about the circuit and the racing car, such as the coordinate position, speed, and steering angle of the racing car. More specific types and values of the situation data include "1" for "current lap count [0..]", "false" for "whether or not a race is in progress", "256" for "lap time (ms) of the current lap", "156164" for "lap time (ms) of the previous lap", and "0" for "progress of the current lap [0,1]". The following situation data are acquired in chronological order: "0.002780" for "Difference from Best Lap Time"; "0.0" for "Speed (km / h)"; "177.693130" for "Speed (km / h)"; "-59.793526" for "Steering Wheel Angle (rad)"; "5.372770, 64.056038, 130.219971" for "Altitude, Latitude, Longitude"; "-0.515301" for "Lateral Position on the Circuit (L=-1, R=1)"; and "0.854022" for "Distance from the Theoretical Fastest Course (m)." Situation data like this is acquired in chronological order. In actual F1 races, situation data can be measured in real time by sensors, or, in fields outside of motorsports, GPS data acquired as the location information of soccer players.
[0030] The first processing unit 10 selects first utterance data from the first utterance data storage unit 3 based on the situation data acquired by the input unit 1. For example, in a racing game, the first utterance data is objective event occurrence data corresponding to a situation, such as "contact" between racing cars or a racing car "going off the track." The first utterance data storage unit 3 stores the first utterance data selected by the first processing unit 10 and the situation data conditions for selecting the first utterance data. For example, if the coordinate positions of the two racing cars are within a predetermined distance, utterance data indicating "contact" is selected. If the distance to the immediately preceding car is decreasing and will be within the predetermined distance, utterance data indicating "caught up with the car in front" is selected. The second processing unit 20 generates second utterance data by performing a first AI process 4 based on the situation data acquired by the input unit 1. Here, AI refers to a computer system or software capable of learning, recognition, and understanding, and AI processing refers to processing using such a computer system or software. By using a language generation model based on a neural network trained using large amounts of data, the processing speed is slower, but flexible and natural language can be generated. The second processing unit 20 performs a process of transmitting to the AI, together with predetermined situation data, utterance condition data such as the number of characters and the utterance viewpoint (player viewpoint or aerial viewpoint, etc.) for the second utterance data obtained as a result of the AI processing, and a process of receiving the second utterance data as a result of the AI processing.
[0031] The first voice generation unit 11 generates a first voice based on the first utterance data selected by the first processing unit 10. The first voice output instruction unit 12 instructs the voice output unit 2 to output the first voice generated by the first voice generation unit 11, and sends data that can determine that an instruction to output the first voice has been issued to the second processing unit 20, the second voice output determination unit 22, and the second voice output instruction unit 23. The second voice generation unit 21 generates a second voice based on the second utterance data generated by the second processing unit 20. The second voice output determination unit 22 determines whether or not the second voice generated by the second voice generation unit 21 should be output from the voice output unit 2. Upon receiving data from the first voice output instruction unit 12, if the first voice is being output, it determines not to output the second voice. If the first voice is not being output, it determines to output the second voice. The second voice output instruction unit 23 instructs the voice output unit 2 to output the second voice generated by the second voice generation unit 21. When the second processing unit 20 receives data from the first voice output instruction unit 12, it starts generating second utterance data based on new situation data from the input unit 1. The second utterance data by the first AI processing 4 that is already in progress in the second processing unit 20 does not cause the second voice generation unit 21 to generate the second voice. When the second voice output determination unit 22 receives data from the first voice output instruction unit 12 and the second voice generation unit 21 generates the second voice while the first voice is being output, the generated second voice is not output from the voice output unit 2. When the second voice output instruction unit 23 receives data from the first voice output instruction unit 12 while the second voice is being output, it issues an instruction to stop outputting the second voice.
[0032] In this way, by using the first utterance data stored in the first utterance data storage unit 3 and the second utterance data generated by the first AI processing unit 4, utterances that require real-time processing and highly flexible utterances can be output as live commentary audio, allowing for a deeper understanding of the event situation. Furthermore, by outputting the first voice in priority over the second voice, the voice output unit 2 can reliably output a relatively well-defined situation description as the first voice in a timely manner and output highly flexible utterances as the second voice during periods when a situation explanation is not necessary, thereby achieving a natural live commentary audio with fewer silent periods. Furthermore, since the voice output unit 2 does not output the second voice while outputting the first voice, it is possible to reliably output a relatively well-defined situation explanation in a timely manner. Furthermore, if second utterance data is generated while the first voice is being output, the second processing unit 20 does not generate the second voice based on the generated second utterance data, but instead generates second utterance data based on new situation data, thereby reliably outputting a relatively well-defined situation explanation in a timely manner and outputting highly flexible utterances during periods when a situation explanation is not necessary. Furthermore, when the second voice is generated while the first voice is being output, the second processing unit 20 generates second utterance data based on new situation data without outputting the generated second voice from the voice output unit 2, thereby enabling a more or less definite situation description to be output in a timely manner and with high flexibility, during periods when a situation description is not necessary. Furthermore, when the voice output unit 2 outputs the first voice while the second voice is being output, the output of the second voice is stopped, enabling a more or less definite situation description to be output in a timely manner and with high flexibility. Furthermore, the second processing unit 20 generates the second utterance data using the output of the first voice as a trigger, enabling a more or less definite situation description to be output in a timely manner.
[0033] FIG. 2 is a flow chart showing the processing flow of the live voice generation system in this embodiment. The computer executes the following processing according to the flow chart. If the acquired situation data matches the condition for selecting first utterance data (Yes in S51), the first processing unit 10 selects the matching first utterance data (first processing step). If the acquired situation data does not match the condition for selecting first utterance data (No in S51), new situation data is acquired from the input unit 1. Note that acquisition of situation data from the input unit 1 is performed periodically at predetermined timings. The first utterance data selected in the first processing step of S51 is used to generate a first voice by the first voice generation unit 11 (S52). An output instruction is issued by the first voice output instruction unit 12 for the first voice generated in the first voice generation step of S52 (S53). When an output instruction is issued in the first voice output instruction step of S53, the first voice is output by the voice output unit 2 (S54).
[0034] When an output instruction is issued in the first voice output instruction step of S53, the second processing unit 20 generates second utterance data using the situation data by the first AI processing 4 (S55). The second utterance data generated in the second processing step of S55 is used by the second voice generation unit 21 to generate second voice (S56). When the second voice is generated in the second voice generation step of S56, it is determined whether the first voice is being output in the first voice output step of S54 (S57). If the first voice is being output (No in S57), new situation data is acquired and new second utterance data is generated (S55) without outputting the second voice generated in the second voice generation step of S56. If output of the first voice is not being performed (Yes in S57), an output instruction is issued using the second voice generated in the second voice generation step (S58). When an output instruction is issued in the second voice output instruction step of S58, the voice output unit 2 outputs the second voice (S59). If an output instruction is given in the first audio output instruction step of S53 while the second audio is being output in the second audio output step of S59 (Yes in S60), the second audio output is stopped by an instruction from the second audio output instruction unit 23 (S61). Therefore, the first audio is output in the first audio output step of S54. If an output instruction is not given in the first audio output instruction step of S53 while the second audio is being output in the second audio output step of S59 (No in S60), the output of the second audio by the audio output unit 2 continues (S62).
[0035] FIG. 3 is an explanatory diagram showing the processing of the live speech generation system in this embodiment. FIG. 3(a) shows the start of the second processing. If the situation data acquired by the first processing unit 10 matches the condition for selecting first speech data (Yes in S51 in FIG. 2), the matching first speech data is selected, and the selected first speech data is used to generate a first speech by the first speech generation unit 11 (S52 in FIG. 2), and the generated first speech is output by the speech output unit 2 (S54 in FIG. 2). In the second processing unit 20, the output of the first speech triggers the generation of second speech data, and the generated second speech is generated from the generated second speech data (S56 in FIG. 2), and the generated second speech is output by the speech output unit 2 (S59 in FIG. 2). Note that the first speech utterance time is preferably shorter than the second speech utterance time and shorter than the second speech data generation time required to generate the second speech data. Here, the second utterance data generation time is the data processing time by the second processing unit 20, or the time from the start of data processing by the second processing unit 20 to the generation of the second voice by the second voice generation unit 21. The start of data processing by the second processing unit 20 is preferably triggered by the output instruction in the first voice output instruction step in S53, but if the second utterance data generation time is short, the trigger may be during the output of the first voice or the end of the output of the first voice.
[0036] 3B illustrates a case where a first voice output occurs during the second processing. Similar to FIG. 3A, when the situation data acquired by the first processing unit 10 matches the condition for selecting first utterance data (Yes in S51 in FIG. 2), the matching first utterance data is selected, and the first voice generation unit 11 generates a first voice from the selected first utterance data (S52 in FIG. 2), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 2). The second processing unit 20 generates second utterance data in response to the output of the first voice. However, if new first utterance data is selected by the first processing unit 10 during the second utterance data generation, the first voice based on the new first utterance data is output by the voice output unit 2. In response to the occurrence of the new first utterance data, the second processing unit 20 acquires new situation data and starts processing the new second utterance data, instead of processing the second utterance data that was being generated. Second utterance data is generated based on the new situation data, and the second voice based on the second utterance data is output by the voice output unit 2 (S59 in FIG. 2).
[0037] 3C illustrates a case where a first speech is output while a second speech is being output. Similar to FIG. 3A, when the acquired situation data by the first processing unit 10 matches the condition for selecting first speech data (Yes in S51 in FIG. 2), the matching first speech data is selected, and the first speech generation unit 11 generates a first speech from the selected first speech data (S52 in FIG. 2), and the generated first speech is output by the speech output unit 2 (S54 in FIG. 2). The second processing unit 20 generates second speech data in response to the output of the first speech, generates a second speech based on the generated second speech data (S56 in FIG. 2), and outputs the generated second speech from the speech output unit 2 (S59 in FIG. 2). When new first speech data is selected by the first processing unit 10 while the speech output unit 2 is outputting the second speech, the second speech output is stopped, and the first speech based on the new first speech data is output by the speech output unit 2. In response to the occurrence of new first utterance data, the second processing unit 20 acquires new situation data and starts processing new second utterance data. The new situation data generates second utterance data, and the second speech based on the second utterance data is output by the speech output unit 2 (S59 in FIG. 2).
[0038] A commentary voice generation system according to another embodiment of the present invention will be described below with reference to FIGS. 4 to 8. FIG. 4 is a block diagram showing a commentary voice generation system according to another embodiment of the present invention, represented by functional implementation means. Functions identical to those in FIG. 1 are assigned the same reference numerals and will not be described again. In addition to the configuration of the commentary voice generation system according to the embodiment shown in FIGS. 1 to 3, the commentary voice generation system according to this embodiment includes a third processing unit 30 that generates third utterance data based on situation data acquired by the input unit 1, and a second utterance data storage unit 5 that stores the second utterance data generated by the second processing unit 20 in association with situation data at the time the second utterance data was generated. The third processing unit 30 then generates the third utterance data by a second AI process 6 using the data stored in the second utterance data storage unit 5 as training data. The first voice output instruction unit 12 sends data that can determine whether an instruction to output the first voice has been issued to the third processing unit 30, the third voice output determination unit 32, and the third voice output instruction unit 33. The third voice generation unit 31 generates a third voice based on the third speech data generated by the third processing unit 30. The third voice output determination unit 32 determines whether or not the third voice generated by the third voice generation unit 31 should be output from the voice output unit 2. Upon receiving data from the first voice output instruction unit 12, if the first voice is being output, it determines not to output the third voice. If the first voice is not being output, it determines to output the third voice. The third voice output instruction unit 33 gives an instruction to output the third voice generated by the third voice generation unit 31 from the voice output unit 2.
[0039] When the third processing unit 30 receives data from the first voice output instruction unit 12, it starts generating third utterance data based on new situation data from the input unit 1. The third utterance data by the second AI processing 6 that is already in progress in the third processing unit 30 does not result in the generation of a third voice by the third voice generation unit 31. When the third voice output determination unit 32 receives data from the first voice output instruction unit 12 and the third voice is generated by the third voice generation unit 31 while the first voice is being output, the generated third voice is not output from the voice output unit 2. When the third voice output instruction unit 33 receives data from the first voice output instruction unit 12 while the third voice is being output, it issues an instruction to stop outputting the third voice.
[0040] In this way, by generating the repeatedly generated second utterance data as third utterance data using the second AI processing 6, the third utterance data can be generated in a shorter time than the generation time of the second utterance data, making it easier to output highly flexible utterances. Furthermore, by outputting the first voice prior to the second and third voices, a relatively well-defined situation description can be reliably output in a timely manner, and highly flexible utterances can be output during periods when a situation description is not necessary, thereby achieving natural commentary with fewer silent periods. Furthermore, since the voice output unit 2 does not output the third voice while outputting the first voice, a relatively well-defined situation description can be reliably output in a timely manner. Furthermore, if third utterance data is generated while the first voice is being output, the third processing unit 30 generates third utterance data based on new situation data without generating the third voice based on the generated third utterance data, thereby reliably outputting a relatively well-defined situation description in a timely manner, and highly flexible utterances can be output during periods when a situation description is not necessary. Furthermore, when the third voice is generated while the first voice is being output, the third processing unit 30 generates third utterance data based on new situation data without outputting the generated third voice from the voice output unit 2, thereby enabling a more or less definite situation description to be output in a timely manner and highly flexible utterances to be output during periods when a situation description is not necessary. Furthermore, when the voice output unit 2 outputs the first voice while the third voice is being output, it stops outputting the third voice, enabling a more or less definite situation description to be output in a timely manner and flexible utterances to be output during periods when a situation description is not necessary. Furthermore, the third processing unit 30 generates the third utterance data using the output of the first voice as a trigger, enabling a more or less flexible utterance to be output during periods when a situation description is not necessary.
[0041] FIG. 5 is a flow chart showing the processing flow of the commentary voice generation system in this embodiment. Processes identical to those in FIG. 2 are assigned the same reference numerals and will not be described again. The computer executes the following processes according to the flow chart. The commentary voice generation system in this embodiment executes the following processes in addition to the processes of the commentary voice generation system in the embodiments shown in FIGS. 1 to 3 . When an output instruction is issued in the first voice output instruction step of S53, the third processing unit 30 generates third speech data by the second AI processing 6 using situation data (S63). The third speech data generated in the third processing step of S63 is used by the third speech generation unit 31 to generate third speech (S64). When the third speech is generated in the third speech generation step of S64, it is determined whether the first speech is being output in the first speech output step of S54 (S65). If the first speech is being output (No in S65), new situation data is acquired and new third speech data is generated (S63) without outputting the third speech generated in the third speech generation step of S64. If the first voice is not being output (Yes in S65), an output instruction is given for the third voice generated in the third voice generation step (S66). When an output instruction is given in the third voice output instruction step in S66, the third voice is output by the voice output unit 2 (S67). If an output instruction is given in the first voice output instruction step in S53 while the third voice is being output in the third voice output step in S67 (Yes in S68), the third voice output is stopped by an instruction from the third voice output instruction unit 33 (S69). Therefore, the first voice is output in the first voice output step in S54. If an output instruction is not given in the first voice output instruction step in S53 while the third voice is being output in the third voice output step in S67 (No in S68), the voice output unit 2 continues to output the third voice (S70).
[0042] 6 to 8 are explanatory diagrams showing the processing of the commentary voice generation system in this embodiment. FIG. 6 shows the relationship between the second process and the third process. FIG. 6(a) shows a case where third utterance data is not generated as a result of the third process, and FIG. 6(b) shows a case where third utterance data is generated as a result of the third process. The third utterance data generation time required to generate the third utterance data is longer than the first voice utterance time and shorter than the second utterance data generation time required to generate the second utterance data. This is because, while both the second processing unit 20 and the third processing unit 30 perform AI processing, the third processing unit 30 performs AI processing based on the training data stored in the second utterance data storage unit 5. Note that the third processing unit 30 does not use the generated third utterance data as third utterance data if the probability of match rate with the training data is below a predetermined value. In this way, by not generating third utterance data for data with a low match rate with the training data as a result of AI processing, the generation of incorrect commentary voice can be prevented.
[0043] When the acquired situation data matches the condition for selecting first utterance data by the first processing unit 10 (Yes in S51 in FIG. 5), the matching first utterance data is selected, the selected first utterance data is used to generate a first voice by the first voice generation unit 11 (S52 in FIG. 5), and the generated first voice is output by the voice output unit 2 (S54 in FIG. 5). The second processing unit 20 generates second utterance data in response to the output of the first voice, and the third processing unit 30 generates third utterance data in response to the output of the first voice.
[0044] As shown in FIG. 6A, when the third processing unit 30 does not generate third utterance data, the second speech is generated from the second utterance data generated by the second processing unit 20 (S56 in FIG. 5), and the generated second speech is output by the speech output unit 2 (S59 in FIG. 5). As shown in FIG. 6B, when the third processing unit 30 generates third utterance data, the third speech is generated from the third utterance data generated by the third processing unit 30 (S64 in FIG. 5), and the generated third speech is output by the speech output unit 2 (S67 in FIG. 5). The third speech data generation time is the data processing time by the third processing unit 30 or the time from the start of data processing by the third processing unit 30 to the generation of the third speech by the third speech generation unit 31. The start of data processing by the third processing unit 30 is preferably triggered by the output instruction in the first speech output instruction step in S53. However, if the third speech data generation time is short, the start of data processing by the third processing unit 30 may be triggered during the output of the first speech or when the output of the first speech ends.
[0045] FIG. 7 illustrates a case where first speech output occurs during the second process. FIG. 7(a) illustrates a case where third speech data is not generated as a result of the third process, and FIG. 7(b) illustrates a case where third speech data is generated as a result of the third process. As in FIG. 6, when the situation data acquired by the first processing unit 10 matches the condition for selecting first speech data (Yes in S51 in FIG. 5), the matching first speech data is selected, and the first speech generation unit 11 generates first speech from the selected first speech data (S52 in FIG. 5), and the generated first speech is output by the speech output unit 2 (S54 in FIG. 5). The second processing unit 20 generates second speech data in response to the output of the first speech. However, if new first speech data is selected by the first processing unit 10 during the second speech data generation time, the first speech based on the new first speech data is output by the speech output unit 2. In the second processing unit 20, in response to the occurrence of new first utterance data, instead of processing the second utterance data that was being generated, new situation data is acquired and processing of the new second utterance data is started. Similar to the second processing unit 20, the third processing unit 30 generates third utterance data triggered by the output of the first voice, but if new first utterance data is selected by the first processing unit 10 during the third utterance data generation time, the first voice based on the new first utterance data is output by the voice output unit 2. In the third processing unit 30, in response to the occurrence of new first utterance data, instead of processing the third utterance data that was being generated, new situation data is acquired and processing of the new third utterance data is started.
[0046] 7(a), when the third processing unit 30 does not generate the third utterance data, the second speech is generated from the second utterance data generated by the second processing unit 20 (S56 in FIG. 5), and the generated second speech is output by the speech output unit 2 (S59 in FIG. 5). Also, when the third processing unit 30 generates the third utterance data, as shown in FIG. 7(b), when the third processing unit 30 generates the third utterance data, the third speech is generated from the third processing unit 30 (S64 in FIG. 5), and the generated third speech is output by the speech output unit 2 (S67 in FIG. 5).
[0047] FIG. 8 illustrates a case where the first speech is output while the third speech is being output. When the acquired situation data by the first processing unit 10 matches the condition for selecting the first speech data (Yes in S51 in FIG. 5 ), the matching first speech data is selected, and the selected first speech data is used by the first speech generation unit 11 to generate the first speech (S52 in FIG. 5 ). The generated first speech is output by the speech output unit 2 (S54 in FIG. 5 ). The third processing unit 30 generates third speech data in response to the output of the first speech, and the third speech is generated based on the generated third speech data (S64 in FIG. 5 ). The generated third speech is output by the speech output unit 2 (S67 in FIG. 5 ). When new first speech data is selected by the first processing unit 10 while the speech output unit 2 is outputting the third speech, the third speech is stopped, and the first speech based on the new first speech data is output by the speech output unit 2. In response to the occurrence of new first utterance data, the third processing unit 30 acquires new situation data and starts processing new third utterance data. The new situation data generates third utterance data, and the third speech based on the third utterance data is output by the speech output unit 2 (S67 in FIG. 5 ).
[0048] The present invention is particularly suited to racing game footage as an event situation, but can also be applied to other video game footage and sports footage that generates live commentary of swimming competitions and the like in real time.
[0049] REFERENCE SIGNS LIST 1 Input unit 2 Audio output unit 3 First speech data storage unit 4 First AI processing 5 Second speech data storage unit 6 Second AI processing 10 First processing unit 11 First speech generation unit 12 First speech output instruction unit 20 Second processing unit 21 Second speech generation unit 22 Second speech output determination unit 23 Second speech output instruction unit 30 Third processing unit 31 Third speech generation unit 32 Third speech output determination unit 33 Third speech output instruction unit
Claims
1. A live-action voice generation system that outputs live audio about an event situation, comprising: an input unit that acquires situation data about the event situation; a first processing unit that selects first utterance data based on the situation data acquired by the input unit; a second processing unit that generates second utterance data based on the situation data acquired by the input unit; a first utterance data storage unit that stores the first utterance data selected by the first processing unit and conditions of the situation data for selecting the first utterance data; a voice generation unit that generates a first voice based on the first utterance data and a second voice based on the second utterance data; and a voice output unit that outputs the first voice and the second voice generated by the voice generation unit, wherein the first processing unit selects the first utterance data from the first utterance data storage unit; the second processing unit generates the second utterance data by first AI processing; and the voice output unit outputs the first voice in priority to the second voice.
2. The live commentary sound generation system according to claim 1, characterized in that the sound output unit does not output the second sound while the first sound is being output.
3. The live voice generation system described in claim 1, characterized in that when the second speech data is generated while the first voice is being output, the second processing unit generates the second speech data based on new situation data without generating the second voice based on the generated second speech data.
4. The live voice generation system described in claim 1, characterized in that when the second voice is generated while the first voice is being output, the second processing unit generates the second speech data based on new situation data without outputting the generated second voice from the voice output unit.
5. The live audio generation system according to claim 1, characterized in that the audio output unit stops outputting the second audio when outputting the first audio while outputting the second audio.
6. The live voice generation system according to claim 1, characterized in that the second processing unit generates the second speech data in response to the output of the first voice as a trigger.
7. A live commentary voice generation system as described in claim 1, comprising: a third processing unit that generates third utterance data based on the situation data acquired by the input unit; and a second utterance data storage unit that associates and stores the second utterance data generated by the second processing unit and the situation data at the time the second utterance data was generated, wherein the voice generation unit generates a third voice based on the third utterance data, the voice output unit outputs the third voice generated by the voice generation unit, the third processing unit generates the third utterance data by a second AI processing using data stored in the second utterance data storage unit as teacher data, and the voice output unit outputs the first voice in priority to the second voice and the third voice.
8. The live commentary sound generation system according to claim 7, characterized in that the sound output unit does not output the third sound while the first sound is being output.
9. The live voice generation system described in claim 7, characterized in that when the third speech data is generated while the first voice is being output, the third processing unit generates the third speech data based on new situation data without generating the third voice based on the generated third speech data.
10. The live voice generation system described in claim 7, characterized in that when the third voice is generated while the first voice is being output, the third processing unit generates the third speech data based on new situation data without outputting the generated third voice from the voice output unit.
11. The live audio generation system as described in claim 7, characterized in that the audio output unit stops outputting the third audio when outputting the first audio while outputting the third audio.
12. The live voice generation system according to claim 7, characterized in that the third processing unit generates the third speech data in response to the output of the first voice as a trigger.
13. A live voice generation system for outputting live voice regarding an event situation, wherein a computer executes the following steps: a first processing step of selecting first utterance data when acquired situation data matches a condition for selecting first utterance data; a first voice generation step of generating a first voice using the first utterance data selected in the first processing step; a first voice output instruction step of instructing output using the first voice generated in the first voice generation step; a voice output step of outputting the first voice instructed in the first voice output instruction step; a second processing step of generating second utterance data by a first AI processing when an instruction is given in the first voice output instruction step; a second voice generation step of generating a second voice using the second utterance data generated in the second processing step; and a second voice output instruction step of instructing output using the second voice generated in the second voice generation step if output using the first voice is not being performed; and if output using the first voice is being performed, generating the second utterance data using new situation data in the second processing step without outputting the second voice generated in the second voice generation step. A live voice generation system comprising:
14. The live commentary voice generation system as described in claim 13, characterized in that when output of the first voice is instructed in the first voice output instructing step while output of the second voice is being performed in the second voice output instructing step, output of the second voice is stopped and output of the first voice is performed.
15. A live audio generation system for outputting live audio about an event situation, comprising: a computer that performs a first processing step of selecting first utterance data when acquired situation data matches a condition for selecting first utterance data; a first audio generation step of generating a first audio using the first utterance data selected in the first processing step; a first audio output instruction step of instructing output using the first audio generated in the first audio generation step; an audio output step of outputting the first audio instructed in the first audio output instruction step; a second processing step of generating second utterance data by first AI processing when an instruction is given in the first audio output instruction step; a second audio generation step of generating a second audio using the second utterance data generated in the second processing step; a third processing step of generating third utterance data by second AI processing when an instruction is given in the first audio output instruction step; and a third audio generation step of generating a third audio using the third utterance data generated in the third processing step. a third voice output instruction step of outputting the third voice generated in the third voice generation step if output using the first voice is not being performed, and a second voice output instruction step of outputting the second voice generated in the second voice generation step if output using the first voice and output using the third voice are not being performed.
16. The live voice generation system described in claim 15, characterized in that if output using the first voice is being performed, then in the second processing step, the second speech data is generated using new situation data without outputting the second voice generated in the second voice generation step.
17. The live commentary voice generation system described in claim 15, characterized in that if output using the first voice is being performed, then in the third processing step, the third speech data is generated using new situation data without outputting the third voice generated in the third voice generation step.
18. The live voice generation system described in claim 15, characterized in that if output using the third voice is being performed, then in the second processing step, the second speech data is generated using new situation data without outputting the second voice generated in the second voice generation step.
19. The live commentary voice generation system as described in claim 15, characterized in that when output of the first voice is instructed in the first voice output instructing step while output of the second voice is being performed in the second voice output instructing step, output of the second voice is stopped and output of the first voice is performed.
20. The live commentary voice generation system as described in claim 15, characterized in that when output using the first voice is instructed in the first voice output instructing step while output using the third voice is being performed in the third voice output instructing step, output using the third voice is stopped and output using the first voice is performed.
Citation Information
Patent Citations
Voice output program, voice output method and video game device
JP2003024626A
Computer system and audio information generation method
JP2021194229A