Audio commentary production device and program

The commentary audio production device addresses limitations in existing systems by leveraging multiple data sources and real-time inputs to generate versatile and synchronized commentary audio, enhancing scalability and reliability.

JP7840208B2Active Publication Date: 2026-04-03NIPPON HOSO KYOKAI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing commentary audio systems are limited to specific events and data sources, lack versatility, and fail to provide real-time commentary due to reliance on a single information source and potential delays in data delivery.

Method used

A commentary audio production device that utilizes data from multiple information sources, including real-time and operator-input sources, to generate scalable and versatile commentary audio by analyzing and labeling text elements, managing them in an information management table, and ensuring timely delivery through an order discard control unit.

Benefits of technology

Enables highly scalable and versatile real-time commentary audio generation, utilizing diverse data sources to ensure reliable and synchronized commentary delivery across various sports events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007840208000001
    Figure 0007840208000001
  • Figure 0007840208000002
    Figure 0007840208000002
  • Figure 0007840208000003
    Figure 0007840208000003
Patent Text Reader

Abstract

To provide an explanation voice with high expandability and versatility at a real time while using data of a plurality of information sources.SOLUTION: An analysis part 11 of an explanation voice production device 1 extracts a text element by analyzing data input from a plurality of information sources 2, and a storage part 12 stores the data to an information management table 13 by applying a label to the text element. A reading part 16 reads out the updated text element from the information management table 13 in each text element of the label defined to a template 14. A format conversion part 17 converts the label and the text element into a Json file in each speech of the read text element. A text generation part 18 generates a text for an explanation voice from the Json file. An order disposal control part 19 determines an order of the speech on the basis of a priority of a presentation timing contained in the level of the Json file, and disposes the Json file of the speech when an output processing related to the speech in ahead is completed.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an explanatory voice production device and a program for generating text for explanatory voice.

Background Art

[0002] Conventionally, there has been known an explanatory voice service that broadcasts a sports relay broadcast program and provides the explanatory voice of the broadcast program to viewers (see, for example, Patent Document 1).

[0003] FIG. 15 is a diagram for explaining the outline of a system that provides an explanatory voice service. This system includes a broadcast transmission device 101, a broadcast reception device 102, an explanatory voice production and distribution device 103, an app server 104, and a mobile terminal 105.

[0004] [[ID=2G]]The broadcast transmission device 101, the explanatory voice production and distribution device 103, and the app server 104 are installed, for example, at a broadcasting station, and the broadcast reception device 102 is installed, for example, in the home of the viewer 100. The mobile terminal 105 is used by the viewer 100 who watches the broadcast program at home.

[0005] With this explanatory voice service of the system, the viewer 100 can receive the provision of the explanatory voice together with the broadcast program of the voice and video that explains the game situation by the live broadcast of the announcer and the explanation of the commentator. [[ID=Z5]]

[0006] The broadcast transmission device 101 transmits broadcast program content to the broadcast reception device 102 via a terrestrial digital broadcast wave. The broadcast reception device 102 is, for example, a television receiver, receives the broadcast program content transmitted from the broadcast transmission device 101 via the terrestrial digital broadcast wave, and reproduces the received broadcast program content.

[0007] The commentary audio production and distribution device 103 produces commentary audio for the broadcast program content transmitted by the broadcast transmission device 101 and sends the commentary audio to the mobile terminal 105. The application server 104 stores applications that run on the mobile terminal 105 and sends applications to the mobile terminal 105 in response to requests from the mobile terminal 105. "Application" is an abbreviation for application, and in this case, it is a program that receives and plays the commentary audio.

[0008] The mobile device 105 is, for example, a smartphone or a PDA (Personal Digital Assistant), and it plays commentary audio for the broadcast program content in synchronization with the broadcast receiving device 102. When playing the commentary audio, the mobile device 105 changes the playback speed and other settings according to the viewer's (100) input.

[0009] For example, if the broadcast program is a baseball game, viewers 100 can receive not only video and audio of the baseball game, but also commentary that explains the game situation in detail, allowing them to understand the details of the game. Baseball commentary may include information such as the pitcher's actions, pitch type, speed, and location, as well as information about the batter, the batter's actions, and the score, depending on the game situation.

[0010] An example of a commentary audio production and distribution device 103 that realizes such a commentary audio service is a system that receives data in accordance with the ODF (Olympic Data Feed) specification, produces commentary audio using that data, and distributes it (see, for example, Non-Patent Document 1).

[0011] The commentary audio production and distribution device 103 described in Non-Patent Literature 1 sequentially receives data such as scores and fouls from a single source of Olympic data. The commentary audio production and distribution device 103 then generates commentary text according to the match situation by applying variables to a pre-set template, converts the text into speech using a speech synthesizer, and transmits the commentary audio file to a mobile terminal 105. [Prior art documents] [Patent Documents]

[0012] [Patent Document 1] Japanese Patent Publication No. 2017-203827 [Non-patent literature]

[0013] [Non-Patent Document 1] Masashi Kumano, "Audio Guide Generation Technology for Sports Programs," NHK Science & Technology Research Laboratories R&D, No. 154, pp. 12-20, 2017. [Overview of the Initiative] [Problems that the invention aims to solve]

[0014] As mentioned above, the technology described in Non-Patent Document 1 can only be used for specific Olympic Games, and generates text in a limited format according to the ODF specifications of the Games, and generates audio files for commentary.

[0015] Therefore, the technology described in Non-Patent Document 1 could not be directly used in other competitions, resulting in problems with its scalability and versatility.

[0016] Furthermore, in the technology described in Non-Patent Document 1, since there is only one information source for generating the explanatory audio, even if there is information that the user wants to convey to the viewer 100 as explanatory audio, that information source is not necessarily available. For this reason, there has been a demand for a technology that can utilize multiple information sources.

[0017] Furthermore, the technology described in Non-Patent Document 1 does not guarantee that data will be delivered from the information source at the necessary time. Therefore, for data requiring real-time delivery, there was a problem in that the commentary audio service could not function unless the data was delivered from the information source at the time it needed to be conveyed to the viewer 100.

[0018] Thus, the technology described in Non-Patent Document 1, which generates text and explanatory audio using data distributed from a single source, is insufficient as an explanatory audio service and cannot adequately meet the requirements of the audience 100.

[0019] Therefore, the present invention has been made to solve the aforementioned problems, and its objective is to provide a commentary audio production device and program that can utilize data from multiple information sources and provide highly scalable and versatile commentary audio in real time. [Means for solving the problem]

[0020] To solve the aforementioned problems, the commentary audio production device of claim 1 is a commentary audio production device that generates text for commentary audio of a live-streamed sports program for each utterance, comprising: a template that defines utterance definition data for each utterance, including one or more labels corresponding to one or more text elements when the text is composed of one or more text elements; an information management table in which the text elements are stored; an analysis unit that inputs data according to the match situation of the sports program from each of a plurality of information sources and analyzes the data according to a pre-set data format of the information source to extract the text elements from the data; a storage unit that assigns labels to the text elements extracted by the analysis unit, including the priority of the timing for presenting the utterance of the text element, and stores the labeled text elements in the information management table; an update monitoring unit that monitors whether the text elements stored in the information management table have been updated and outputs the labels assigned to the text elements when it is determined that they have been updated; and the label output by the update monitoring unit. The device is characterized by comprising: a reading unit that reads from the information management table one or more corresponding text elements to which one or more labels included in the utterance definition data are assigned for the utterance including a bell, and outputs one or more labels for the utterance and one or more corresponding text elements thereto; a format conversion unit that converts the one or more labels for the utterance and one or more corresponding text elements output by the reading unit into a file including a predetermined playback time; a text generation unit that extracts the one or more text elements from the file converted by the format conversion unit, generates the text, and outputs it; and an order discard control unit that inputs the file for each utterance converted by the format conversion unit, determines the order of the utterances based on the priority included in the labels in the file for each utterance, resets the playback time included in the file for each utterance according to the order, outputs the playback time included in the file of the first utterance whose order has been determined, and discards the file of the first utterance.

[0021] Further, the commentary voice production device according to claim 2 is the commentary voice production device according to claim 1, wherein the plurality of information sources include an information source that transmits real-time data according to the game situation of the sports program, and further, an information source that transmits data of the sports program according to an input operation of an operator, an information source that transmits data obtained by analyzing an image of the game situation of the sports program, and at least one of an information source that transmits data obtained by recognizing the voice of the game situation of the sports program is included.

[0022] Further, the commentary voice production device according to claim 3 is the commentary voice production device according to claim 1, wherein in the template, for each utterance, in addition to the one or more labels, one of the one or more labels is defined as a trigger label, and the update monitoring unit monitors whether or not the text element to which the trigger label is assigned is updated in the information management table, and outputs the trigger label when it is determined that an update has occurred.

[0023] Further, the commentary voice production device according to claim 4 is the commentary voice production device according to claim 1, wherein the label is composed of respective numerical values indicating the type of the information source, the event type of the sports program, the priority, the group to which the text element belongs, and the item within the group, and when reading out the corresponding one or more text elements to which the one or more labels included in the utterance definition data are assigned, the label is a read target label, and in addition to the read target label, a label having the same group and item within the group to which the text element constituting the read target label belongs and a different information source type is defined as a same type label, and when a plurality of text elements to which the same type label is assigned are stored in the information management table, the reading unit reads out the text element that was stored first among the plurality of text elements to which the same type label is assigned from the information management table.

[0024] Further, the explanatory voice production device according to claim 5 is the explanatory voice production device according to claim 4, wherein when the reading unit does not store a text element with the reading target label in the information management table and stores a text element with the same type label other than the reading target label, the reading unit reads out a text element with the same type label other than the reading target label from the information management table.

[0025] Further, the explanatory voice production device according to claim 6 is the explanatory voice production device according to claim 1, wherein the label includes the competition event of the sports program in addition to the priority, and for the text element read out from the information management table by the reading unit, the reading unit corrects the text element according to the competition event included in the label given to the text element, and outputs one or more labels and one or more text elements (if there is a corrected text element, the text element) of the utterance.

[0026] Further, the explanatory voice production device according to claim 7 is the explanatory voice production device according to claim 1, wherein the one or more labels included in the utterance definition data include labels corresponding to a predetermined particle or word, the reading unit reads out one or more text elements including the predetermined particle or word from the information management table, the text generation unit extracts one or more text elements including the predetermined particle or word from the file, and generates and outputs a text including the predetermined particle or word.

[0027] Furthermore, the explanatory voice production device of claim 8 is an explanatory voice production device of claim 1, wherein the priority included in the label is any of the information indicating immediate, semi-immediate, periodic, and other, with immediate having the highest priority, semi-immediate the next highest, and other the lowest, and the order discard control unit determines the order of the utterances of the label including the periodic priority as the first utterance of the label including the immediate or semi-immediate priority, and the second utterance of the label including the periodic priority, which is arranged at a predetermined time interval, and determines the order of the utterances such that the higher the priority of the first utterance, the closer it is to the beginning, and determines the order of the utterances such that the second utterance is arranged after the first utterance at the predetermined time interval.

[0028] Furthermore, the commentary voice production device of claim 9 is the commentary voice production device of claim 1, wherein the label consists of the type of information source, the sport of the sports program, the priority, the group to which the text element belongs, and a numerical value indicating the item within the group, and the label to which the one or more labels included in the utterance definition data are assigned to the corresponding one or more text elements is defined as the target label to be read, and in addition to the target label to be read, labels to which the group to which the text element constituting the target label belongs and the item within the group are the same, but to which the type of information source is different are defined as the same type label, and when the order discard control unit inputs a new file containing the same type label in conjunction with the update monitoring unit's determination of an update, and determines the order of utterances for each utterance file, if there are multiple files containing the same type label, the order discard control unit discards all files containing the same type label except for the new file.

[0029] Furthermore, the commentary voice production device of claim 10 is characterized in that, in the commentary voice production device of claim 1, the sequence discard control unit discards files from the files for each utterance after a predetermined time has elapsed.

[0030] Furthermore, the program of claim 11 comprises a computer constituting a commentary audio production device that generates commentary audio text for each utterance of a live-streamed sports program, a template defining utterance definition data including one or more labels corresponding to one or more text elements when the text is composed of one or more text elements, an information management table in which the text elements are stored, an analysis unit that inputs data according to the match situation of the sports program from each of a plurality of information sources and analyzes the data according to a pre-set data format of the information source to extract the text elements from the data, an analysis unit that assigns labels to the text elements extracted by the analysis unit, including the priority of the timing for presenting the utterance of the text element, and stores the labeled text elements in the information management table, an update monitoring unit that monitors whether the text elements stored in the information management table have been updated and outputs the labels assigned to the text elements when it is determined that they have been updated, and the labels output by the update monitoring unit The system is characterized by functioning as an order discard control unit that, with respect to the utterances in the utterance definition data, reads from the information management table one or more text elements to which one or more labels included in the utterance definition data are assigned, outputs one or more labels for the utterance and one or more text elements corresponding thereto; converts the one or more labels for the utterance and one or more text elements corresponding thereto output by the read unit into a file containing a predetermined playback time; extracts the one or more text elements from the file converted by the format conversion unit, generates and outputs the text; and inputs the files for each utterance converted by the format conversion unit, determines the order of the utterances based on the priority included in the labels in the files for each utterance, resets the playback time included in the files for each utterance according to the order, outputs the playback time included in the file of the first utterance whose order has been determined, and discards the file of the first utterance. [Effects of the Invention]

[0031] As described above, according to the present invention, it is possible to utilize data from multiple information sources and provide highly scalable and versatile explanatory audio in real time. [Brief explanation of the drawing]

[0032] [Figure 1] This is a schematic diagram illustrating an example of the overall configuration of a commentary audio production and distribution system including a commentary audio production device according to an embodiment of the present invention. [Figure 2] This is a block diagram showing an example configuration of a commentary audio production device according to an embodiment of the present invention. [Figure 3] This flowchart shows examples of processing in the analysis and storage units. [Figure 4] This is a diagram explaining the labels. [Figure 5] The diagram shows an example of a template. [Figure 6] This diagram illustrates an example of speech definition data defined in a template, labels and text elements stored in an information management table, a JSON file, and text for explanatory audio. [Figure 7] This flowchart shows an example of processing by the update monitoring unit. [Figure 8] This diagram illustrates an example of the processing performed by the update monitoring unit. [Figure 9] This flowchart shows an example of the processing in the reading section. [Figure 10] This diagram illustrates an example of reading text elements (step S905). [Figure 11] This flowchart shows an example of the format conversion process. [Figure 12] This flowchart shows an example of the text generation process. [Figure 13] This flowchart shows an example of the processing performed by the sequence discarding control unit. [Figure 14] This diagram illustrates the changes in speech data within the sequence in the sequence discard control unit. [Figure 15] This diagram illustrates the system's overview for providing explanatory audio services. [Modes for carrying out the invention]

[0033] The embodiments for carrying out the present invention will be described in detail below with reference to the drawings. [Audio commentary production and distribution system] First, we will describe a commentary audio production and distribution system that realizes commentary audio services. Figure 1 is a schematic diagram illustrating an example of the overall configuration of a commentary audio production and distribution system including a commentary audio production device according to an embodiment of the present invention.

[0034] This commentary audio production and distribution system 10 comprises a commentary audio production device 1, multiple information sources 2, a speech synthesis device 3, a distribution device 4, and a mobile terminal 5. The commentary audio production and distribution system 10 corresponds to the commentary audio production and distribution device 103 and the mobile terminal 105 of the system that provides commentary audio services as shown in Figure 15.

[0035] The commentary audio production device 1 is a device that generates commentary text for each utterance when producing commentary audio for live-streamed sports programs. The commentary audio production device 1 receives real-time data corresponding to the game situation of the live-streamed sports program from multiple information sources 2. The commentary audio production device 1 then extracts text elements by analyzing the data according to the unique data format of the information sources 2, assigns labels to the text elements, and stores them in the information management table 13, which will be described later.

[0036] Here, a text element is one or more elements that make up the text for the explanatory audio you want to generate (the text of what you want to say). A label is information used to identify the content of a text element. More details will be provided later.

[0037] The commentary audio production device 1 generates commentary audio text for a single utterance by reading updated text elements from the information management table 13 according to the utterance definition data defined in the template 14 described later, generating a JSON file including the playback time, assigning an utterance ID, and generating the commentary audio text. The commentary audio production device 1 also determines the order of utterances and resets the playback time. The playback time is the time when the mobile terminal 5 plays the audio file of the commentary audio text.

[0038] The commentary audio production device 1 outputs the utterance ID and the text for the commentary audio to the speech synthesis device 3 for each utterance, and also outputs the utterance ID and the playback time to the distribution device 4.

[0039] Information source 2 consists of multiple information sources for each sport, for example. As shown in Figure 1, the multiple information sources 2 for baseball include, for example, information source 2-1 which distributes Olympic-related data in accordance with the ODF specification, information source 2-2 which distributes professional baseball-related data in accordance with the BIS specification, information source 2-3 which distributes professional baseball-related data in accordance with the BIP specification, and information source 2-4 which distributes high school baseball-related data in accordance with the SIGN specification.

[0040] Furthermore, baseball information sources 2 include information sources 2-5, which transmit baseball-related data according to predetermined specifications based on input operations by operators watching broadcast programs; information sources 2-6, which generate baseball-related data by analyzing images of baseball game situations and transmit baseball-related data according to predetermined specifications; and information sources 2-7, which generate baseball-related data by recognizing audio of baseball game situations and transmit baseball-related data according to predetermined specifications. In addition, there are multiple information sources 2 that distribute tennis data.

[0041] Thus, in addition to information sources 2-1, ..., 2-4 that transmit real-time data according to the match situation in sports programs, by using information source 2-5 obtained from operator input, information source 2-6 obtained from image analysis, and information source 2-7 obtained from speech recognition, or at least one of information sources 2-5, 2-6, or 2-7, it is possible to supplement the data that is not available from information sources 2-1, ..., 2-4 and broaden the range of commentary audio that can be presented.

[0042] Furthermore, by using data from multiple information sources 2, a highly versatile commentary audio production and distribution system 10 can be realized without being dependent on a specific program, tournament, or competition. In addition, since the necessary data for commentary audio can be obtained from multiple information sources 2, reliable commentary audio can be presented, and real-time capabilities can also be achieved.

[0043] The speech synthesis device 3 receives the utterance ID and the text for the commentary voice from the commentary voice production device 1, and generates an audio file by generating synthesized sound from the commentary voice text using existing technology. The speech synthesis device 3 then outputs the utterance ID and the audio file to the distribution device 4.

[0044] The distribution device 4 receives the utterance ID and playback time from the commentary audio production device 1, as well as the utterance ID and audio file from the speech synthesis device 3. The distribution device 4 then adjusts the playback time so that the utterances of the commentary audio do not overlap with the utterances of the main broadcast audio. Specifically, the distribution device 4 predicts the end time of the utterances of the main broadcast audio, and if the time period between the start and end times of the main audio overlaps with the time period during which the commentary audio file is played, it changes the playback time to a time after the end time of the main audio.

[0045] The distribution device 4 distributes the audio file and playback time of the same speech ID to the mobile terminal 5.

[0046] The mobile terminal 5 receives the audio file and playback time distributed from the distribution device 4, and plays the audio file at the playback time.

[0047] For real-time audio broadcasts, it's possible that multiple voices may overlap. Furthermore, elderly individuals or those with hearing impairments may have difficulty hearing female voices or may miss audio during fast-paced competitions. In such cases, the mobile device 5 selects the speaker and playback speed of the audio according to the viewer's (100) input. This ensures that the audio is easily understandable for each individual's needs.

[0048] [Commentary Audio Production Device 1] Next, the commentary audio production device 1 shown in Figure 1 will be described in detail. Figure 2 is a block diagram showing an example of the configuration of the commentary audio production device 1 according to an embodiment of the present invention.

[0049] This commentary audio production device 1 comprises an analysis unit 11, a storage unit 12, an information management table 13, a template 14, an update monitoring unit 15, a reading unit 16, a format conversion unit 17, a text generation unit 18, and an order discard control unit 19.

[0050] <Analysis section 11> Figure 3 is a flowchart showing an example of processing by the analysis unit 11 and the storage unit 12. The analysis unit 11 inputs data from multiple information sources 2 according to the match status of the live-streamed sports program (step S301). The input data is defined in various formats such as fixed length, CSV, XML, and JSON.

[0051] The analysis unit 11 identifies the type of information source 2, which is the input source of the data, and generates identification information. It also analyzes the input data from information source 2 according to a pre-set data format of information source 2, thereby extracting text elements from the data (step S302). In addition, when extracting text elements, the analysis unit 11 generates analysis results indicating the type, content, and other information of the text elements, and outputs the text elements, analysis results, and identification information to the storage unit 12.

[0052] For example, the analysis unit 11 takes in professional baseball-related data (such as "Pitcher Suzuki" and "He took a stance") from information source 2-2 in accordance with the BIS specifications, generates identification information indicating that the type of information source 2 is "BIS", and analyzes the data to match the data format of information source 2-2, thereby extracting the text elements "Pitcher Suzuki" and "He took a stance" from the data. The analysis unit 11 also generates results indicating that the sport is "baseball", "Pitcher Suzuki" is the name of the pitcher, and "He took a stance" is the action of the pitcher.

[0053] <Storage section 12> The storage unit 12 receives text elements, analysis results, and identification information from the analysis unit 11, and assigns labels to the text elements according to the analysis results and identification information (step S303). Then, the storage unit 12 stores the labeled text elements along with a timestamp in the information management table 13 (placed in the position corresponding to the label) (step S304). The timestamp is information about the time when the text elements are stored in the information management table 13.

[0054] As mentioned above, labels are information used to identify the content of text elements. As shown in Figure 4 below, labels consist of a total of five numerical values ​​from the first to the fifth column. Specifically, the first column indicates the type of information source 2 from which the text element was obtained, the second column indicates the sport of the text element, and the third column indicates the priority (priority) of the timing of presentation when the text element is presented as an utterance. The fourth and fifth columns indicate the group and item when the text element is classified according to its content.

[0055] For example, when the storage unit 12 receives the text element "Pitcher Suzuki", the analysis result (results indicating that the sport is "baseball" and that "Pitcher Suzuki" is the name of a pitcher), and identification information (information indicating that the type of information source 2 is "BIS") from the analysis unit 11, it assigns the label "2-1-3-9-1" to the text element "Pitcher Suzuki" as a label corresponding to the analysis result and identification information.

[0056] As shown in Figure 4 below, the first column "2" indicates that the type of information source 2 is "BIS", the second column "1" indicates that the sport is "baseball", and the third column "3" indicates that the presentation timing is "regular". In addition, the fourth and fifth columns "9-1" indicate that the group is "pitcher information" and the item is "name" (indicating the pitcher's name).

[0057] Furthermore, when the storage unit 12 receives the text element "kamagaeta", the analysis result (a result indicating that information source 2 is "BIS", the sport is "baseball", and "kamagaeta" is the pitcher's motion, etc.) and identification information (information indicating that the type of information source 2 is "BIS") from the analysis unit 11, it assigns "2-1-1-11-1" to the text element "kamagaeta" as a label corresponding to the analysis result and identification information.

[0058] As shown in Figure 4 below, the first column "2" indicates that the type of information source 2 is "BIS", the second column "1" indicates that the sport is "baseball", and the third column "1" indicates that the presentation timing is "immediate". In addition, the fourth and fifth columns "11-1" indicate that the group is "pitcher's actions" and the item is "prepared".

[0059] As a result, text elements are assigned common labels, allowing text elements extracted from data in the data formats of Information Source 2, which are inherently different, to be centrally managed in Information Management Table 13. Similarly, data from different sports can also be centrally managed.

[0060] Here, in the storage unit 12, the first column of the label is assigned a numerical value corresponding to the identification information, the second, fourth, and fifth columns of the label are assigned numerical values ​​corresponding to the analysis results, and at the timing of the presentation of the third column, the numerical value in the third column of the label defined in the template 14 described later is assigned. Specifically, for the numerical values ​​in the first to fifth columns of the label to be assigned, the storage unit 12 first determines the numerical values ​​in the first, second, fourth, and fifth columns according to the analysis results and identification information. Then, for the third column, the storage unit 12 identifies a label from the labels defined in the template 14 described later that has the same numerical values ​​in the first, second, fourth, and fifth columns as the determined numerical values ​​in the first, second, fourth, and fifth columns, extracts the numerical value in the third column of the identified label, and determines that numerical value as the numerical value in the third column of the label to be assigned.

[0061] Furthermore, the template 14, described later, may be configured to define the presentation timing of the third column corresponding to the fourth and fifth columns of the label. In this case, the template 14 contains numerical values ​​for the presentation timing of the third column for each of the fourth and fifth columns of the label. The storage unit 12 determines the numerical values ​​for the first, second, fourth, and fifth columns of the label according to the analysis results and identification information, then reads the numerical values ​​for the presentation timing of the third column corresponding to the fourth and fifth columns of the label from the template 14, and determines the read numerical values ​​as the numerical values ​​for the third column of the label to be assigned.

[0062] <Information Management Table 13> The information management table 13 stores text elements with labels along with timestamps. In other words, the information management table 13 consists of labels, text elements, and timestamps. A text element is the smallest unit of text used to create the explanatory audio.

[0063] Figure 4 illustrates the labels. As shown in Figure 4(1), a label is assigned to a single text element and consists of a total of five numbers, from the first to the fifth column.

[0064] The first column of the label indicates the type of information source 2, which is the source of the text element, as shown in Figure 4(2). The type of information source 2 is information used to distinguish its origin. The number "1" indicates "ODF", ​​the number "2" indicates "BIS", the number "3" indicates "BIP", the number "4" indicates "image analysis tool", the number "5" indicates "input tool", and so on.

[0065] The second column of the label, as shown in Figure 4(3), indicates the sport represented by the text element. The sport is information used when unique utterances or conditions differ for each sport. The number "1" represents "baseball," "2" represents "tennis," "3" represents "table tennis," "4" represents "badminton," "5" represents "basketball," and so on.

[0066] The third column of the label, as shown in Figure 4(4), indicates the priority of presentation timing when presenting text elements as utterances. Presentation timing is information used to control when the explanatory audio is presented. The number "1" indicates "immediate," the number "2" indicates "semi-immediate," the number "3" indicates "periodic," and the number "4" indicates "other."

[0067] The commentary audio must be presented one at a time, and especially under conditions where it is presented in conjunction with the broadcast, the presentation timing is predetermined according to the text element, taking into account whether or not it is acceptable for the broadcast to overlap with the commentary audio.

[0068] "Immediate" means that synchronization with the video is crucial, and there is no consideration for overlap with the broadcast audio. As soon as the audio files for the commentary are distributed from the distribution device 4, the app on the mobile device 5 plays them immediately. For this reason, it has the highest priority. For example, if the commentary audio says "The pitcher is ready" or "He threw," these are meaningless unless they are played in sync with the video.

[0069] "Semi-immediate" means that the app on the mobile device 5 plays the audio file of the commentary within a predetermined time, taking into account any overlap with the broadcast audio. For example, if the commentary audio says "Suzuki vs. Yamada 10-6" when a move is made in a table tennis match, the app on the mobile device 5 will either play the audio within a time frame of two seconds when there is no overlap with the broadcast audio, or play it immediately if two seconds have passed, in order to ensure that the commentary does not overlap with the broadcast audio.

[0070] The "Regular" setting is used when the commentary audio, such as the match title, opponents, and current score, is not immediate and should be spoken periodically. The app on the mobile device 5 plays the commentary audio file at predetermined time intervals or under predetermined conditions.

[0071] In this case, for example, the distribution device 4 shown in Figure 1 receives the utterance ID and playback time from the commentary audio production device 1, along with the label corresponding to the utterance. If it determines that the timing of the label presentation is "periodic," it changes the playback time so that the utterance of the commentary audio does not overlap with the utterance of the main audio broadcast.

[0072] The fourth column of the label, as shown in Figure 4(5), shows the groups formed when text elements are classified according to their content. The group is information indicating the category of the text element. The number "1" indicates "game information", the number "2" indicates "type of game", ..., the number "9" indicates "pitcher information", the number "10" indicates "batter information", the number "11" indicates "pitcher's actions", ....

[0073] This allows text elements to be managed as a group, enabling collectively controlling the information in the fifth column, which is defined below.

[0074] The fifth column of the label, as shown in Figure 4(5), indicates the items within the group. These items represent information that further subdivides the categories of the text elements, and are the most specific representations of the information.

[0075] For example, if the fourth column of the label is group "Match Information" with the numerical value "1", then the numerical value "1" indicates "Tournament Name", the numerical value "2" indicates "Match Name (e.g., X vs Y)", the numerical value "3" indicates "Venue (e.g., Z Stadium)", and so on. Also, for example, if the fourth column of the label is group "Pitcher Information" with the numerical value "9", then the numerical value "1" indicates "Name (e.g., Suzuki)", the numerical value "2" indicates "Season Record (e.g., 5 wins and 2 losses for this season)", and the numerical value "3" indicates "Today's Record (e.g., Today's ERA 0.50)". Also, for example, if the fourth column of the label is group "Pitcher Actions" with the numerical value "11", then the numerical value "1" indicates "Prepared", the numerical value "2" indicates "Throwed", and the numerical value "3" indicates "Pickoff Attempt".

[0076] In the fourth column of the label, "Group," the numbers "1" through "5" and "18" are common information for all sports and are not used if this type of text element cannot be obtained from Information Source 2. In the fourth column of the label, "Group," the numbers "6," "7," and "15" through "17" are common information for racket sports and are used, for example, when "Sport" is "Table Tennis," "Badminton," and "Tennis." When "Sport" is "Table Tennis," "Badminton," and "Tennis," there are many common "Items," so this common information is used.

[0077] The numbers "8" through "14" in the "Group" column of the label represent information when the "Sport" is "Baseball," but since there are common "items" when the "Sport" is "Softball," it may be considered information common to both "Baseball" and "Softball."

[0078] In this way, by using multiple information sources 2, text elements can be stored in many "items" of the information management table 13, allowing for the generation of many types of explanatory audio text and broadening the range of explanatory audio that can be expressed.

[0079] Furthermore, if the fourth column group of the label is "Particle" with the numerical value "40", then the numerical value "1" of the item represents "wa", the numerical value "2" represents "no", the numerical value "3" represents "e", and the numerical value "4" represents "ga". Also, if the fourth column group of the label is "Word (Position)" with the numerical value "41", then the numerical value "1" of the item represents "direction", the numerical value "2" represents "back", and the numerical value "3" represents "front".

[0080] If the fourth column group of the label is the numerical value "40" and is a "particle," and if the fourth column group of the label is the numerical value "41" and is a "word (position)," these text elements are not obtained from information source 2, but are pre-stored as fixed strings in the information management table 13.

[0081] By using such "particles" or "words (positions)" as text elements, that is, by using fixed text elements that are not obtained from information source 2 and are pre-stored in information management table 13, it is possible to generate flexible explanatory audio text. The app on the mobile device 5 can then play audio files of explanatory voice that closely resemble human speech, and the listener 100 can easily recognize the explanatory voice.

[0082] <Template 14> Figure 5 shows an example of template 14. This template 14 defines speech definition data for each utterance text generated by the narration audio production device 1, i.e., for each utterance, consisting of an utterance number, utterance content, a combination of labels, and a trigger label. To increase the number of narration audio types, one simply needs to add the utterance definition data for the new utterance, i.e., the utterance number, utterance content, a combination of labels, and a trigger label, to this template 14.

[0083] The speech definition data defined in template 14 is set by key input from the user operating the commentary voice production device 1. Note that the configuration of template 14 shown in Figure 5 is just one example, and other configurations are also possible.

[0084] The utterance number is a number used to identify the utterance definition data for each utterance. The utterance content is the content to be uttered and consists of one or more text elements, the "items" in the 5th column of the label. The label combination consists of one or more labels corresponding to the utterance content. The trigger label is a label corresponding to a text element whose update is monitored by the update monitoring unit 15, which will be described later.

[0085] In the example in Figure 5, utterance number 1 defines the following information: the utterance content is "name of pitcher information" and "pitcher's action (prepared)", the label combination is "4-1-3-9-1" and "5-1-1-11-1", and the trigger label is "5-1-1-11-1".

[0086] This indicates that when the text element "11-1" in the 4th and 5th columns of the trigger label "5-1-1-11-1" stored in the information management table 13 is updated, it generates explanatory audio text consisting of the text elements "Pitcher Information Name" and "Pitcher's Action (Prepared)" from the 4th and 5th columns of the labels "4-1-3-9-1" and "11-1" of "5-1-1-11-1" stored in the information management table 13.

[0087] Furthermore, as utterance number 2, the following information is defined: the utterance content is "Tournament name," "Match name," "Country name 1," and "Country name 2," and the label combinations and trigger labels are "1-1-3-1-1," "1-1-3-1-2," "1-1-3-3-1," and "1-1-3-3-2."

[0088] This indicates that when all (or at least one) text elements in the 4th and 5th columns "1-1" of the trigger label "1-1-3-1-1", "1-2" of the 4th and 5th columns of "1-1-3-1-2", "3-1" of the 4th and 5th columns of "1-1-3-3-1", and "3-2" of the 4th and 5th columns of "1-1-3-3-2" stored in the information management table 13 are updated, the system generates commentary audio text consisting of these text elements "Tournament Name", "Match Name", "Country Name 1", and "Country Name 2" stored in the information management table 13.

[0089] Furthermore, under utterance number 3, the following information is defined: the utterance content is "Pitch type (curveball)", the label combination is "5-1-1-12-1", and the trigger label is "5-1-1-12". In addition, under utterance number 4, the following information is defined: the utterance content is "Pitch type (fastball)", the label combination is "5-1-1-12-2", and the trigger label is "5-1-1-12".

[0090] This indicates that when the text element "Curveball" of the "Item" label "5-1-1-12-1" or the text element "Straight" of the "Item" label "5-1-1-12-2" belonging to the "Pitch Type" "Group" of the trigger label "5-1-1-12" stored in the information management table 13 is updated, explanatory audio text consisting of the text element "Curveball" of the label "5-1-1-12-1" or the text element "Straight" of the label "5-1-1-12-2" stored in the information management table 13 is generated.

[0091] Furthermore, the reason the trigger label consists of the label "5-1-1-12" in columns 1-4, rather than the labels "5-1-1-12-1" and "5-1-1-12-2" in columns 1-5, is that from the information management table 13, only one of the updated text elements, "curveball" from label "5-1-1-12-1" or "straight" from label "5-1-1-12-2", will be read, and both text elements will never be read simultaneously.

[0092] Furthermore, as utterance number 5, the following information is defined: the utterance content is "defensive position," "word (direction)," and "batting result (hit)," the label combinations are "4-1-3-20-1," "5-1-4-41-1," and "5-1-1-21-1," and the trigger label is "5-1-1-21-1."

[0093] This indicates that when the text element "Hit" of the trigger label "5-1-1-21-1" stored in the information management table 13 is updated (when "Hit" is stored), explanatory audio text is generated consisting of the text elements "Defensive position", "Word (direction)", and "Hit result (hit)" of the labels "4-1-3-20-1", "5-1-4-41-1", and "5-1-1-21-1" stored in the information management table 13.

[0094] This example of utterance number 5 adds a fixed string (in this example, "word (direction)") between the text elements "defensive position" and "batting result (hit)" obtained from information source 2, which are not obtained from information source 2. This generates, for example, the commentary text "Hit to left field," which is easy for viewer 100 to understand and eliminates any sense of incongruity when viewer 100 listens to the commentary.

[0095] Figure 6 illustrates an example of speech definition data defined in template 14, labels and text elements stored in information management table 13, a JSON file, and text for explanatory audio.

[0096] As shown in Figure 6(1), the utterance definition data for utterance number 1 in template 14 defines the following information: utterance content "name of pitcher information" and "pitcher action (prepared)", label combination "4-1-3-9-1" and "5-1-1-11-1", and trigger label "5-1-1-11-1". This is the same as the utterance definition data for utterance number 1 shown in Figure 5.

[0097] Here, in the information management table 13, the text element "Pitcher's actions (positioned)" for the label "5-1-1-11-1" is newly stored as "positioned," and this data is updated.

[0098] At this time, as shown in Figure 6(2), when the information management table 13 is updated, the text element labeled "4-1-3-9-1" contains "Pitcher Suzuki" as the "Name of pitcher information". Also, the information management table 13 contains the text element labeled "5-1-1-11-1" contains "Pitcher's action (positioned)" as the "Pitcher's action (positioned)".

[0099] In this case, according to the utterance definition data for utterance number 1 defined in template 14, the information management table 13 determines whether the text element with the trigger label "5-1-1-11-1" has been updated, and the text element "Pitcher Suzuki" with the label "4-1-3-9-1" and the text element "Prepared" with the label "5-1-1-11-1" are read out.

[0100] Then, the JSON files and commentary audio text shown in Figures 6(3) and (4) described later are generated, and the audio file "Pitcher Suzuki takes his position" is played on the mobile device 5 as the commentary audio.

[0101] As mentioned above, the speech definition data for speech number 1 of template 14 shown in Figure 5 defines the speech content as "name of pitcher information" and "pitcher's action (prepared)", the label combination as "4-1-3-9-1" and "5-1-1-11-1", and the trigger label as "5-1-1-11-1". This speech definition data generates explanatory audio text consisting of the text elements "name of pitcher information" and "pitcher's action (prepared)".

[0102] Alternatively, the speech definition data for speech number 1 may be modified to include additional information such as generating explanatory audio text consisting of the text elements "Pitcher's Name" and "Pitcher's Action (Prepared)" once out of five updates, and generating explanatory audio text consisting only of the text element "Pitcher's Action (Prepared)" four times.

[0103] <Update monitoring section 15> Figure 7 is a flowchart showing an example of processing by the update monitoring unit 15. The update monitoring unit 15 receives a trigger label from the reading unit 16 and monitors whether the text element to which the trigger label is assigned has been updated in the information management table 13 (step S701). The trigger label is read from the template 14 by the reading unit 16 and output from the reading unit 16 to the update monitoring unit 15.

[0104] If the update monitoring unit 15 determines in step S701 that the text element of the trigger label has not been updated (step S701: no update), it continues the processing in step S701.

[0105] Meanwhile, if the update monitoring unit 15 determines in step S701 that the text element of the trigger label has been updated (step S701: update detected), it outputs "update detected" and the trigger label to the reading unit 16 (step S702). The processing shown in Figure 7 is performed for each utterance (each utterance definition data) defined in the template 14.

[0106] Figure 8 illustrates an example of processing by the update monitoring unit 15. Figure 8(1) shows an example of processing for utterance number 1 of template 14 shown in the examples in Figures 5 and 6. The update monitoring unit 15 receives the trigger label "5-1-1-11-1" from the reading unit 16. Then, in the information management table 13, the text element "pitcher's action (prepared)" to which the trigger label "5-1-1-11-1" is assigned is updated, and "prepared" is stored (see α in Figure 8).

[0107] In this case, the update monitoring unit 15 determines whether there has been an update to "Pitcher's Action (Prepared)" by checking for updates to the timestamp of the text element with the trigger label "5-1-1-11-1" in the information management table 13, and outputs the update status and the trigger label "5-1-1-11-1" to the reading unit 16. In this case, the update monitoring unit 15 may also determine whether there has been an update to "Pitcher's Action (Prepared)" by checking for updates to the text element with the trigger label "5-1-1-11-1" itself (an update due to a change from a state where nothing is stored in the area of ​​"Pitcher's Action (Prepared)" to a state where "Prepared" is stored).

[0108] This results in the text elements "Pitcher Suzuki" and "Prepared" corresponding to the utterance content "Name of pitcher information" and "Pitcher's action (prepared)" defined in utterance number 1 of template 14.

[0109] Figure 8(2) shows an example of processing for utterance number 4 of template 14 shown in the example in Figure 5. The update monitoring unit 15 receives the trigger label "5-1-1-12" from the reading unit 16. Then, in the information management table 13, the text elements "Pitch type (curveball)" and "Pitch type (straight)", which are assigned the labels "5-1-1-12-1" and "5-1-1-12-2" corresponding to the trigger label "5-1-1-12", are updated, and the area for "Pitch type (curveball)" is empty, and "straight" is stored in the area for "Pitch type (straight)" (see β in Figure 8). Alternatively, the timestamp of the text element "Pitch Type (Straight)" to which the label "5-1-1-12-2" is assigned is updated, and the new text "Straight" is stored in the area of ​​"Pitch Type (Straight)".

[0110] In this case, the update monitoring unit 15 determines whether an update has occurred by checking whether the text elements "Pitch type (curveball)" and "Pitch type (straight)" attached to the labels "5-1-1-12-1" and "5-1-1-12-2" corresponding to the trigger label "5-1-1-12" in the information management table 13 have been updated, or whether the timestamp of the text element "Pitch type (straight)" attached to the label "5-1-1-12-2" has been updated. It then outputs "Update occurred" and the trigger label "5-1-1-12" to the reading unit 16.

[0111] This results in the text element "straight" corresponding to the utterance content "pitch type (straight)" defined in utterance number 4 of template 14. In this case, by determining updates using the timestamp, it is possible to determine that there have been consecutive updates even if the text element "straight" has been updated consecutively.

[0112] <Readout section 16> Figure 9 is a flowchart showing an example of processing by the reading unit 16. The reading unit 16 reads utterance definition data (utterance number, utterance content, label combination, and trigger label) for each utterance from the template 14 (step S901). Then, the reading unit 16 outputs the trigger label for each utterance to the update monitoring unit 15 (step S902).

[0113] The reading unit 16 determines whether or not an update has been input from the update monitoring unit 15 (step S903). If the reading unit 16 determines in step S903 that no update has been input (step S903:N), it continues the process in step S903.

[0114] On the other hand, if the reading unit 16 determines in step S903 that an update has been input from the update monitoring unit 15 (step S903: Y), it identifies the utterance corresponding to the trigger label input along with the update. Then, for each combination of labels in the utterance (each label), the reading unit 16 reads the label and the text element to which the label is assigned from the information management table 13 (step S904).

[0115] Here, if the reading unit 16 finds that the information management table 13 contains multiple text elements of the same type as the label to be read (identical labels (labels with the same numerical values ​​in columns 4 and 5)), it reads the label and text element that was stored first (step S905). Identical labels are those that, in addition to the label to be read, have the same numerical values ​​in columns 4 and 5 as those in columns 4 and 5 (group and item) of the label to be read, and have a different numerical value in column 1 than that in column 1 (type of information source 2) of the label to be read. A specific example will be explained in Figure 10 below.

[0116] Furthermore, if the information management table 13 does not contain text elements of the labels defined in the template 14 (labels to be read), and only contains text elements of similar labels other than the label to be read, the reading unit 16 reads the other similar labels and their text elements.

[0117] This allows the commentary audio production device 1 to retrieve data (data related to the labels defined in template 14) from the primary information source 2, but to retrieve data from another information source 2 (data related to labels of the same type (same type labels) with a different type of information source 2 in the first column of the labels defined in template 14), read the same type label and text element from the information management table 13, and generate commentary audio text that reflects this. In other words, the mobile terminal 5 can play an audio file of commentary audio that reflects data retrieved from information source 2 other than the primary information source 2.

[0118] Figure 10 illustrates an example of reading text elements (step S905). It assumes that the label "5-1-1-11-1" is defined in the label combination of template 14, and that the information management table 13 stores two text elements as text elements where the 4th and 5th columns of the label are "11-1".

[0119] Suppose that one is the text element "assumed" TA of the label "5-1-1-11-1" and its timestamp is t1, and the other is the text element "assumed" TB of the label "4-1-1-11-1" and its timestamp is t2. Assume that t1 < t2 and t2 is closer to the current time. The timestamp indicates the time when the label and the text element were stored in the information management table 13.

[0120] These labels have the same "11-1" in the 4th and 5th columns. The label "5-1-1-11-1" is the read target label, and the labels "5-1-1-11-1" and "4-1-1-11-1" are of the same type. Both labels are the same in that the utterance content corresponding to the "11-1" in the 4th and 5th columns of the label is "assumed", and they are different in that the types of information sources 2 indicated in the 1st column of the "5-1-1" and "4-1-1" in the 1st to 3rd columns of the label are "input tool" and "image analysis tool", respectively.

[0121] In this case, the reading unit 16 compares the timestamps t1 and t2 of the labels "5-1-1-11-1" and "4-1-1-11-1" of the same type with respect to the label (read target label) "5-1-1-11-1" defined in the label combination of the template 14, and identifies the earliest (oldest, most past) timestamp t1. Then, the reading unit 16 reads out the label "5-1-1-11-1" with the earliest stored timestamp t1 and its text element "assumed" TA.

[0122] As a result, the mobile terminal 5 can play the audio file of the explanatory voice corresponding to the earliest stored text element, so that real-time performance according to the video can be realized.

[0123] Returning to FIG. 9, for the label and the text element read from the information management table 13, the reading unit 16 modifies the text element according to a preset rule according to the "competition event" indicated by the numerical value in the 2nd column of the label (step S906).

[0124] For example, the utterance definition data for template 14 defines the utterance content "Score 1, Word (pair), Score 2" and the corresponding label combination, and the reading unit 16 reads "15" as the text element "Score 1" for the label "4-2-1-5-1" (the number "2" in the second column indicates that the "sport" is "tennis") from the information management table 13, and also reads "15" as the text element "Score 2" for the label "4-2-1-5-2".

[0125] The reading unit 16 determines that the "sport" in the utterance is "tennis" according to pre-set rules, and modifies the text elements "15", "pair", and "15" (in this case, the explanatory audio text is "15 vs 15") that make up the utterance to the text elements "fifteen" and "all" (in this case, the explanatory audio text is "fifteen all").

[0126] Furthermore, the reading unit 16 reads "15" as the text element "Score 1" for the label "4-4-1-5-1" (the number "4" in the second column indicates the "sport" of "badminton") from the information management table 13, and also reads "15" as the text element "Score 2" for the label "4-4-1-5-2".

[0127] In this case, the reading unit 16 determines that the "sport" is "badminton" according to a predetermined rule for the utterance, and modifies the text elements "15", "pair", and "15" (in this case, the text for the explanatory audio is "15 vs 15") that make up the utterance to the text elements "15", "pair", "15", and "tie" (in this case, the text for the explanatory audio is "15 vs 15 tie").

[0128] This allows for differentiation of phrasing when the scores are the same across multiple sports but the wording used differs for each sport, enabling the playback of audio files with commentary tailored to each sport.

[0129] Furthermore, for example, the reading unit 16 reads labels "5-2-1-16-1" to "5-2-1-16-4" and their text elements "decisive move (smash)"..., as well as labels "5-2-1-17-1" and "5-2-1-17-2" and their text elements "result (success)" and "result (failure)" from the information management table 13, determines that the "sport" is "tennis" according to pre-set rules, and automatically adds up the score based on these combinations.The reading unit 16 may also generate new labels and text elements "Suzuki", "vs", "Tanaka", "thirty", and "fifteen" (in this case, the text for the commentary audio would be "Suzuki vs Tanaka Thirty Fifteen") when the score is 30 to 15.

[0130] In this way, if, for example, a text element of a player who is serving (e.g., "Suzuki") is identified, the text element can be modified or a new text element can be generated in accordance with a pre-set rule in the reading unit 16, such as text element "Suzuki" "double fault", "Suzuki" "service ace", "Suzuki" "return ace", etc.

[0131] Returning to Figure 9, the reading unit 16 outputs one or more labels and one or more corresponding text elements to the format conversion unit 17 for each utterance (step S907).

[0132] <Format conversion unit 17> Figure 11 is a flowchart showing an example of the processing of the format conversion unit 17. The format conversion unit 17 receives one or more labels and one or more corresponding text elements from the reading unit 16 for each utterance (step S1101).

[0133] The format conversion unit 17 assigns an ID to each utterance to identify the JSON file described later (step S1102). The format conversion unit 17 then sets the playback time based on the time when the JSON file is generated in step S1104 described later, taking into account delays due to the speech synthesis processing time by the speech synthesizer 3 (step S1103). The playback time is the time when the mobile terminal 5 plays the audio file of the commentary voice, and may be reset by the subsequent sequence discard control unit 19 and the distribution device 4 shown in Figure 1.

[0134] The format conversion unit 17 generates a JSON file for each utterance according to a pre-set data format, containing an ID, playback time, and one or more labels and one or more corresponding text elements. The format conversion unit 17 then outputs the JSON file for each utterance to the text generation unit 18 and the sequence discard control unit 19 (step S1104).

[0135] Referring to Figures 6(1) to (3), the reading unit 16 reads the label "4-1-3-9-1" and its corresponding text element "Pitcher Suzuki", and the label "5-1-1-11-1" and its corresponding text element "Prepared", from the information management table 13. The format conversion unit 17 then inputs the label "4-1-3-9-1" and its corresponding text element "Pitcher Suzuki", and the label "5-1-1-11-1" and its corresponding text element "Prepared" for the utterance. The format conversion unit 17 then assigns an ID to the utterance and sets the playback time.

[0136] The format conversion unit 17 generates a JSON file containing the ID "000···2724", the playback time "2021-03-23···2233705Z", the first label "4-1-3-9-1" and the text element "Pitcher Suzuki", and the second label "5-1-1-11-1" and the text element "Prepared".

[0137] This allows for the unification of data in different data formats obtained from multiple sources 2, resulting in the generation of a JSON file that does not reflect the characteristics of source 2.

[0138] <Text generation unit 18> Figure 12 is a flowchart showing an example of processing by the text generation unit 18. The text generation unit 18 receives a JSON file for each utterance from the format conversion unit 17 (step S1201).

[0139] The text generation unit 18 sequentially extracts one or more text elements from the JSON file (step S1202) and generates explanatory audio text consisting of one or more text elements (step S1203).

[0140] In the example in Figure 6, the text elements "Pitcher Suzuki" and "He's in position" are extracted from the JSON file shown in Figure 6(3), and the text for the commentary audio, "Pitcher Suzuki, He's in position," is generated as shown in Figure 6(4).

[0141] Returning to Figure 12, the text generation unit 18 assigns a unique utterance ID to the JSON file (to the utterance) to identify the JSON file (utterance) (step S1204). Then, the text generation unit 18 determines the number of characters in the explanatory audio text generated in step S1203, and calculates the duration of the audio file (wav (Waveform Audio File Format) file) of the explanatory audio text using a predetermined calculation process based on the number of characters (step S1205). The process for calculating the duration of the audio file from the number of characters is known, so its explanation is omitted here.

[0142] The text generation unit 18 outputs the utterance ID and duration to the sequence discard control unit 19 (step S1206), and outputs the utterance ID and explanatory audio text to the speech synthesizer 3 (step S1207).

[0143] This allows the speech synthesis device 3 to manage the audio files for explanatory text using the utterance ID. Furthermore, the subsequent sequence discard control unit 19 can reset the playback time for each utterance using the time length.

[0144] In addition, although the text generation unit 18 generates explanatory audio text based on the JSON file input from the format conversion unit 17, it may also be configured to generate explanatory audio text by inputting one or more text elements for each utterance from the reading unit 16.

[0145] <Order Discard Control Unit 19> Figure 13 is a flowchart showing an example of the processing of the sequence discard control unit 19. The sequence discard control unit 19 determines whether or not there is input from the format conversion unit 17 (whether or not it is the input timing) (step S1301). If the sequence discard control unit 19 determines in step S1301 that there is input (step S1301:Y), it proceeds to step S1302. On the other hand, if the sequence discard control unit 19 determines in step S1301 that there is no input (step S1301:N), it proceeds to step S1308.

[0146] The sequence discard control unit 19 moves from step S1301(Y) and receives a JSON file from the format conversion unit 17, and also receives the utterance ID and duration corresponding to the JSON file from the text generation unit 18 (step S1302).

[0147] The sequence discard control unit 19 extracts the playback time and one or more labels from the input JSON file, generates speech data consisting of the utterance ID, playback time, one or more labels, and duration, and adds it to the end of the array (step S1303). As a result, speech data corresponding to the updated text elements in the information management table 13 is added to the array. Note that the speech data may consist of the utterance ID and the JSON file.

[0148] Here, the array consists of one or more utterance data for each utterance, that is, one or more utterance data corresponding to one or more audio files to be spoken as explanatory audio. If there is no explanatory audio to be spoken, no utterance data is present in the array. Each time a JSON file is input from the format conversion unit 17, utterance data corresponding to that JSON file is added to the array. Furthermore, utterance data in the array is discarded by processing such as step S1306, which will be described later.

[0149] The order discard control unit 19 determines and rearranges the order of multiple speech data in the array based on the "presentation timing" in the third column of the label, so that the priority order is "immediate" > "semi-immediate" > "periodic" > "other" (step S1304).

[0150] Here, the sequence discard control unit 19 treats the "presentation timing" of the speech data as "periodic" if the speech data contains multiple labels and at least one of the labels has a numerical value of "3" in its third column, which represents "periodic". Furthermore, the sequence discard control unit 19 treats the "presentation timing" of the speech data as "immediate" if none of the labels have a numerical value of "3" for "periodic" and at least one has a numerical value of "1" for "immediate". Also, the sequence discard control unit 19 treats the "presentation timing" of the speech data as "immediate" if none of the labels have a numerical value of "3" for "periodic" and at least one has a numerical value of "2" for "semi-immediate".

[0151] While video and images can convey multiple pieces of information simultaneously, even if audio presents multiple pieces of information at the same time, it is difficult for viewers 100 to understand and receive the content. In particular, in audio commentary services using the audio commentary production device 1, even if there are multiple pieces of information to convey simultaneously, they must be conveyed one by one in order.

[0152] Therefore, in the process of step S1304, the order of the utterance data is rearranged according to priority, and in the processes of step S1306 and others described later, the utterance data is discarded to form an array consisting of one or more utterance data.

[0153] The sequence discard control unit 19 resets the playback time for multiple speech data in the sequence (step S1305). Specifically, the sequence discard control unit 19 resets the playback time for the speech data based on its order in the sequence, as well as the playback time and duration of the speech data in the sequence.

[0154] For example, suppose that the speech data added in step S1303 is determined to be the second in the sequence in step S1304. In this case, the sequence discard control unit 19 adds the duration of the first speech data after the rearrangement to the playback time of the first speech data before the rearrangement, and resets the resulting time as the playback time of the second speech data determined after the rearrangement.

[0155] Furthermore, although the sequence discard control unit 19 uses the time length calculated by the text generation unit 18, it may also use the time length calculated by the speech synthesizer 3 to reset the playback time of the speech data. Since the speech synthesizer 3 calculates the time length when generating the audio file, it can calculate a time length with higher accuracy than the text generation unit 18. Therefore, the sequence discard control unit 19 can reset the playback time with higher accuracy by using the time length calculated by the speech synthesizer 3.

[0156] The sequence discard control unit 19 discards older speech data if there are multiple speech data with the same type of label in the 4th and 5th columns of the array (step S1306).

[0157] Specifically, the sequence discard control unit 19 identifies speech data in the array that contain labels with the same numerical value in the 4th and 5th columns (speech data with the same type of label) from among multiple speech data. Then, the sequence discard control unit 19 keeps the speech data containing the most recent playback time from among the identified multiple speech data, and discards one or more older speech data containing playback times (older than the most recent playback time). In this case, the sequence discard control unit 19 resets the playback times of the speech data in the array after the discard process using the same process as in step S1305.

[0158] As a result, in step S1309 described later, during the waiting period before outputting the utterance ID and playback time of the first utterance data in the array, if text elements of the same label are updated in the information management table 13 and this utterance data is added to the array, the old utterance data containing the playback time is discarded, and the audio file of the commentary corresponding to the utterance data containing the latest playback time, i.e., the added utterance data, is played on the mobile terminal 5. Therefore, when playing the audio file of commentary synchronized with the video, the viewer 100 can be presented with commentary that reflects the latest match situation, thereby achieving real-time synchronization with the video.

[0159] The sequence discard control unit 19 discards speech data in the sequence that has been added to the sequence for a certain period of time (a predetermined time) (step S1307), and proceeds to step S1308. In this case as well, the sequence discard control unit 19 resets the playback time for the speech data in the sequence after the discard process using the same process as in step S1305.

[0160] The sequence discard control unit 19 moves from step S1301(N) or step S1307 to determine whether or not there is an output from the sequence discard control unit 19 (whether or not it is the output timing) (step S1308).

[0161] If the sequence discard control unit 19 determines in step S1308 that there is an output (step S1308:Y), it proceeds to step S1309. On the other hand, if the sequence discard control unit 19 determines in step S1308 that there is no output (step S1308:N), it terminates the process and restarts the process from step S1301.

[0162] The sequence discard control unit 19, moving from step S1308(Y), extracts the utterance ID and playback time from the first utterance data in the sequence and outputs the utterance ID and playback time to the distribution device 4 (step S1309). Then, the sequence discard control unit 19 discards the first utterance data in the sequence (step S1310).

[0163] Figure 14 illustrates the transition of speech data within an array in the sequence discard control unit 19. Assume that the array (A) is configured. The beginning of this array is the speech data "pitcher name (name of pitcher information) + ready (pitcher's action (ready))". The third column of the label, "Presentation Timing," is "Immediate" ([1]). This array also contains the utterance data "Game Name + Inning + Score (Score 1, Score 2)." This includes the label, and the third column of the label, “Presentation Timing,” is “Regular” ([3]). Utterance data The text for the commentary audio is "X vs Y, bottom of the 6th inning, 7-0".

[0164] As shown in (B), in step S1309, the sequence discard control unit 19 discards the speech data at the beginning of the sequence. The speech ID and playback time are output, and in step S1310, the speech data Discard it. This removes the speech data from array (A). The elements are discarded, and the array (C) is formed.

[0165] Then, as shown in (D), the sequence discard control unit 19, in step S1302, receives the speech data "batter name (name of batter information) + hit (batter's action)" JSON file, speech data for "ball speed" <c>The JSON file and the speech data of "defensive position + direction (word (direction)) + hit (batting result (hit))" <d>Enter the JSON file. Speech data The third column of the labels included, “Presentation Timing”, is “Immediate” ([1]), and the utterance data <c>The third column of the labels included, "Presentation Timing," is "Semi-immediate" ([2]). Also, the speech data <d>The third column of the labels included in is “Presentation Timing” which is “Immediate” ([1]).

[0166] As shown in (E), in step S1304, the sequence discard control unit 19 determines and rearranges the order of multiple utterance data in the array based on the "presentation timing" in the third column of the label, so that the priority is "immediate" > "semi-immediate" > "periodic" > "other". In step S1305, it resets the playback time of the utterance data to form the array in (F).

[0167] The beginning of the array (F) is the speech data. The third column of the label, "Presentation Timing," is "Immediate" ([1]). The second is the utterance data. <d>The third column of the label, "Presentation Timing," is "Immediate" ([1]). The third is the utterance data. <c>Therefore, the "Presentation Timing" in the third column of the label is "Semi-immediate" ([2]). For the fourth and subsequent utterance data, the "Presentation Timing" in the third column of the label is assumed to be "Semi-immediate" ([2]), "Regular" ([3]), or "Other" ("4").

[0168] Then, as shown in (G), in step S1309, the sequence discard control unit 19 discards the speech data at the beginning of the sequence. The speech ID and playback time are output, and in step S1310, the speech data Discard it. This removes the utterance data from the array (F). The sequence discard control unit 19, in step S1309, discards the speech data at the beginning of the sequence. <d>The speech ID and playback time are output, and in step S1310, the speech data <d>Discard it. This removes the utterance data from the array. <d>It will be discarded.

[0169] Then, as shown in (H), the sequence discard control unit 19, in step S1302, receives the speech data "throw from left field towards third base". <e>The JSON file and the speech data in the format "game name + inning + score" <f>Enter the JSON file. Speech data <e>The third column of the labels included, “Presentation Timing”, is “Immediate” ([1]), and the utterance data <f>The third column of the labels included in is “Presentation Timing” which is “Regular” ([3]).

[0170] Here, speech data <f> and speech data< / f> < / f> < / e> < / f> < / e> < / d> < / d> < / d> < / c> < / d> < / d> < / c> < / d> < / c> The labels are homogeneous labels. (See the aforementioned speech data) The text for the commentary audio is "X vs Y, bottom of the 6th inning, 7-0", whereas the updated speech data <f>The text for the commentary audio is "X vs Y, bottom of the 6th inning, 7-1".

[0171] As shown in (I), in step S1304, the sequence discard control unit 19 determines and rearranges the order of the multiple speech data in the array based on the "presentation timing" in the third column of the label, so that the priority is "immediate" > "semi-immediate" > "periodic" > "other", and in step S1305, it resets the playback time of the speech data.

[0172] In this case, the sequence discard control unit 19 will discard the updated speech data whose "presentation timing" is "periodic". <f>Regarding this, the "presentation timing" is set to "immediate" to determine the order of the speech data and rearrange them. Then, the order discard control unit 19 discards the speech data for which the "presentation timing" is set to "immediate". <f>For the period from the next utterance to the last utterance data, the utterance data <f>While keeping the "presentation timing" set to "periodically," the speech data will be set to a predetermined time interval. <f>Insert.

[0173] This will result in speech data <f>The text for the commentary audio, "X vs. Y, bottom of the 6th inning, 7-1," is generated immediately and will be generated periodically thereafter.

[0174] Then, as shown in (J), the sequence discard control unit 19, in step S1306, discards speech data with the same label in the array. <f>< / f> < / f> < / f> < / f> < / f> < / f> < / f> There is speech data is speech data <f>The speech data has been updated, so the newly added speech data <f> Other speech data< / f> < / f> Discard the old audio text for the commentary "X vs Y 6th inning bottom 7-0" from the array. The data has been discarded, and the speech data corresponding to the commentary audio text for the latest playback time, "X vs Y, bottom of the 6th inning, 7-1", has been discarded. <f>This will remain. Also, in step S1307, the sequence discard control unit 19 discards speech data that has been in the sequence for a certain amount of time. <c>Discard it. This creates the array (K).

[0175] The beginning of the array (K) is the utterance data. <e>The third column of the label, "Presentation Timing," is "Immediate" ([1]). The second is the utterance data. <f>Therefore, the "Presentation Timing" in the third column of the label is "Immediate" ([1]). This utterance data <f>Regarding this, the "presentation timing" when a JSON file is input is "periodically" ([3]), so the utterance data should be set to a predetermined time interval.<f’> It is inserted into the array as such.

[0176] This will result in speech data <f>The text for the commentary audio, "X vs. Y, bottom of the 6th inning, 7-1," will be generated immediately and periodically. Viewers 100 will be able to listen to this commentary audio when it is updated, and will be able to listen to it periodically thereafter unless it is updated again.

[0177] Here, for speech data containing multiple labels, if it includes a label whose "presentation timing" is "periodic" and also includes a label whose "presentation timing" is "immediate" or "semi-immediate," the "presentation timing" of that speech data will be treated as "periodic."

[0178] In this case, as described in (I) above, when the sequence discard control unit 19 constructs a new sequence, it sets the "presentation timing" of the first updated speech data (speech data from the newly input JSON file) whose "presentation timing" is "periodically" to "immediate" or "semi-immediate" (if other speech data in the sequence contains the label "immediate" and does not contain the label "semi-immediate", it sets it to "immediate", and if other speech data contains the label "semi-immediate" and does not contain the label "immediate", it sets it to "semi-immediate") to determine the order and rearrange the speech data in the sequence. Then, the sequence discard control unit 19 rearranges the speech data in the sequence by determining the order so that the "presentation timing" of subsequent speech data remains "periodically" and matches the pre-set time intervals.

[0179] As described above, according to the explanatory audio production device 1 of the embodiment of the present invention, the analysis unit 11 inputs data from a plurality of information sources 2 and extracts text elements by analyzing the input data according to a pre-set data format of the information sources 2.

[0180] The storage unit 12 assigns labels to text elements and stores the labeled text elements along with timestamps in the information management table 13.

[0181] The update monitoring unit 15 monitors whether the text elements of the trigger labels included in the speech definition data defined in the template 14 have been updated in the information management table 13.

[0182] The reading unit 16 reads one or more text elements of a combination of labels (one or more labels) included in the utterance definition data from the information management table 13 for the utterance of the trigger label of the updated text element.

[0183] The format conversion unit 17 converts one or more labels and their corresponding one or more text elements into a JSON file that includes the playback time.

[0184] The text generation unit 18 extracts one or more text elements from the JSON file, generates text for explanatory audio, assigns an utterance ID to the JSON file, and calculates the length of the audio file based on the number of characters in the explanatory audio text.

[0185] When the sequence discard control unit 19 receives a JSON file as input from the format conversion unit 17, it extracts the playback time and one or more labels from the JSON file, generates speech data consisting of an utterance ID, playback time, one or more labels, and duration, and determines the order of the speech data based on the priority of presentation timings included in the labels to construct an array. The sequence discard control unit 19 also resets the playback time of the utterance data based on the order of the constructed array, the playback time of the speech data, and the duration.

[0186] The sequence discard control unit 19 outputs the utterance ID and playback time for the first utterance data in the sequence and discards that utterance data. In addition, if there are multiple utterance data with the same label in the sequence, the sequence discard control unit 19 discards the utterance data with the oldest playback time and also discards utterance data that has been played for a certain period of time.

[0187] As a result, according to the utterance definition data for each utterance defined in template 14, the text for the explanatory audio to be spoken is generated, the order of the utterances is determined based on the priority of "presentation timing" in the third column of the label, and the playback times of the utterances are reset.

[0188] The audio files of the commentary text generated in this way are played at the scheduled time, allowing real-time commentary to be provided to 100 viewers in sync with the live broadcast of the sports program.

[0189] Furthermore, by acquiring data from multiple information sources 2 and storing the extracted text elements in the information management table 13 along with labels that unify the data formats of the multiple information sources 2, the information management table 13 can be centrally managed. In other words, a highly versatile commentary audio production and distribution system 10 can be realized.

[0190] In this way, by centrally managing the information management table 13 using multiple information sources 2, real-time functionality is ensured while generating explanatory audio text for each utterance without overlapping text elements.

[0191] If you want to increase the number of utterances in this explanatory audio text, you can simply add the utterance definition data for those utterances to template 14, thus realizing a highly scalable explanatory audio production and distribution system 10.

[0192] Therefore, by utilizing data from multiple information sources 2, it is possible to provide highly scalable and versatile explanatory audio in real time.

[0193] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments and can be modified in various ways without departing from the technical concept.

[0194] For example, the format conversion unit 17 is configured to generate a JSON file, but it may also generate a file other than a JSON file. The present invention does not limit the files input to the text generation unit 18 and the order discard control unit 19 to JSON files, but may be files conforming to other data formats.

[0195] Furthermore, in the explanatory audio production and distribution system 10 shown in Figure 1, the speech synthesis device 3 generates an audio file through speech synthesis processing, and the distribution device 4 distributes the audio file and playback time to the mobile terminal 5.

[0196] In contrast, the distribution device 4 may distribute to the mobile terminal 5 the text for the commentary audio generated by the commentary audio production device 1 and the playback time, instead of the audio file and playback time. In this case, the commentary audio production and distribution system 10 does not need to have a speech synthesis device 3, and the mobile terminal 5 receives the text for the commentary audio, performs speech synthesis processing to generate an audio file, and plays the audio file at the playback time.

[0197] Furthermore, as shown in Figure 5, template 14 defines the content of the utterance, combinations of labels, etc. Alternatively, template 14 may define information such as captions. In this case, a highly versatile item such as captions is prepared as the fifth column of the labels, and the storage unit 12 stores the text elements of the information to which such labels are assigned in the information management table 13. Then, the text generation unit 18 generates text for commentary audio that includes text elements such as captions. This makes it possible to generate text for commentary audio that includes text elements such as captions even when it is unclear what kind of information should be presented as commentary audio, thereby improving the rate at which commentary is provided in a program.

[0198] For example, text elements of data entered from information source 2-5 of the input tool are stored in information management table 13 along with labels such as captions. As a result, the character information entered by the operator of information source 2-5 of the input tool is reflected in the text for the explanatory audio, and the audio file of the explanatory audio reflecting that character information is played.

[0199] Furthermore, a standard computer can be used as the hardware configuration for the commentary audio production device 1 according to the embodiment of the present invention. The commentary audio production device 1 is composed of a computer equipped with a CPU, a volatile storage medium such as RAM, a non-volatile storage medium such as ROM, and an interface.

[0200] The functions of the commentary audio production device 1, including the analysis unit 11, storage unit 12, information management table 13, template 14, update monitoring unit 15, reading unit 16, format conversion unit 17, text generation unit 18, and order discard control unit 19, are each realized by having the CPU execute a program that describes these functions.

[0201] These programs are stored in the aforementioned storage medium and are read and executed by the CPU. These programs can also be stored and distributed on storage media such as magnetic disks (floppy disks, hard disks, etc.), optical disks (CD-ROMs, DVDs, etc.), and semiconductor memory, and can be transmitted and received via a network. [Explanation of symbols]

[0202] 1. Audio commentary production device 2 Sources of information 3. Speech synthesis device 4 Distribution device 5,105 Mobile devices 10. Commentary Audio Production and Distribution System 11 Analysis Department 12 Storage Unit 13 Information Management Table 14 Templates 15 Update monitoring department 16 Reading section 17 Format conversion section 18 Text generation unit 19. Sequence Discard Control Unit 100 viewers 101 Broadcasting transmission equipment 102 Broadcast receiving equipment 103 Commentary Audio Production and Distribution Device 104 Application Server< / f> < / f> < / f> < / e> < / c> < / f>

Claims

1. In a commentary audio production device that generates text for commentary audio for live-streamed sports programs, For each utterance, a template is defined which includes utterance definition data that includes one or more labels corresponding to the one or more text elements when the text is composed of one or more text elements, An information management table in which the aforementioned text elements are stored, An analysis unit inputs data corresponding to the match status of the sports program from each of multiple information sources, and analyzes the data according to a pre-set data format of the information source to extract the text elements from the data. A storage unit assigns labels to the text elements extracted by the analysis unit, including the priority of the timing for presenting the utterance of the text element, and stores the labeled text elements in the information management table. An update monitoring unit monitors whether the text element stored in the information management table has been updated, and outputs the label assigned to the text element if it is determined that it has been updated. A reading unit reads from the information management table one or more corresponding text elements to which one or more labels are assigned in the utterance definition data output by the update monitoring unit, for the utterance in the utterance definition data that includes the label, and outputs one or more labels of the utterance and one or more corresponding text elements. A format conversion unit that formats one or more labels of the utterance output by the reading unit and one or more corresponding text elements into a file containing a predetermined playback time, A text generation unit extracts one or more text elements from the file format-converted by the format conversion unit, generates the text, and outputs it. The format conversion unit inputs the files for each utterance, determines the order of the utterances based on the priority included in the label of the file for each utterance, and resets the playback time included in the file for each utterance according to the order. A sequence discard control unit outputs the playback time included in the file of the first utterance whose order has been determined, and discards the file of the first utterance. A commentary audio production device characterized by having the following features.

2. In the explanatory audio production device according to claim 1, The commentary audio production device is characterized in that the plurality of information sources include an information source that transmits real-time data corresponding to the match status of the sports program, and further includes at least one of the following: an information source that transmits data of the sports program according to the operator's input operation, an information source that transmits data obtained by analyzing images of the match status of the sports program, and an information source that transmits data obtained by recognizing the audio of the match status of the sports program.

3. In the explanatory audio production device according to claim 1, The template includes, for each utterance, one of the one or more labels, in addition to the one or more labels, defined as a trigger label. The aforementioned update monitoring unit, An explanatory audio production device characterized by monitoring whether the text element to which the trigger label has been assigned has been updated in the information management table, and outputting the trigger label if it is determined that it has been updated.

4. In the explanatory audio production device according to claim 1, The label shall consist of the type of information source, the sport of the sports program, the priority, the group to which the text element belongs, and a numerical value indicating the item within the group. When reading the corresponding one or more text elements to which the one or more labels are assigned within the speech definition data, the label is defined as the target label to be read, and in addition to the target label to be read, labels that belong to the same group as the text elements constituting the target label, and whose items within the group are the same but whose information sources are of different types are defined as same-type labels. The aforementioned reading unit is, If the information management table contains multiple text elements to which the same type of label has been assigned, An explanatory audio production device characterized by reading the text element that was stored first among a plurality of text elements to which the same label has been assigned from the information management table.

5. In the explanatory audio production device according to claim 4, The aforementioned reading unit is, If the information management table does not contain any text elements to which the target label is assigned, but does contain text elements to which other similar labels are assigned, An explanatory audio production device characterized by reading text elements with the same type of label other than the target label from the information management table.

6. In the explanatory audio production device according to claim 1, The label includes, in addition to the priority, the sport of the sports program. The aforementioned reading unit is, A commentary audio production device characterized by modifying a text element read from the information management table according to the sport included in the label assigned to the text element, and outputting one or more labels and one or more text elements (or the modified text element if any) of the utterance.

7. In the explanatory audio production device according to claim 1, The one or more labels included in the speech definition data include labels corresponding to predetermined particles or words. The aforementioned reading unit is, From the information management table, read one or more text elements containing the predetermined particle or word, The text generation unit, An explanatory audio production device characterized by extracting one or more text elements containing the predetermined particle or word from the aforementioned file, and generating and outputting text containing the predetermined particle or word.

8. In the explanatory audio production device according to claim 1, The priority included in the label is one of the following pieces of information: immediate, semi-immediate, periodic, and other, with immediate having the highest priority, semi-immediate the next highest, and other the lowest. The sequence discarding control unit, The utterance of the label including the periodic priority is defined as the first utterance of the label including the immediate or semi-immediate priority, and the second utterance of the label including the periodic priority is arranged at predetermined time intervals. With respect to the first utterance, the order of the utterances is determined such that those with higher priority are placed closer to the beginning, A commentary audio production device characterized by determining the order of the second utterances so that they are placed after the first utterance at a predetermined time interval.

9. In the explanatory audio production device according to claim 1, The label shall consist of the type of information source, the sport of the sports program, the priority, the group to which the text element belongs, and a numerical value indicating the item within the group. When reading the corresponding one or more text elements to which the one or more labels are assigned within the speech definition data, the label is defined as the target label to be read, and in addition to the target label to be read, labels that belong to the same group as the text elements constituting the target label, and whose items within the group are the same but whose information sources are of different types are defined as same-type labels. The sequence discarding control unit, When the update monitoring unit determines that an update has occurred, and a new file containing the same label is input, and the order of utterances for each utterance file is determined, if there are multiple files containing the same label, A commentary audio production device characterized by discarding files other than the new file from among a plurality of files containing the same type of label.

10. In the explanatory audio production device according to claim 1, The sequence discarding control unit, A commentary audio production device characterized by discarding files from the aforementioned utterance files after a predetermined time has elapsed.

11. A computer that constitutes a commentary audio production device that generates text for each utterance of commentary audio for live-streamed sports programs, For each utterance, a template is defined which includes utterance definition data that includes one or more labels corresponding to the one or more text elements when the text is composed of one or more text elements. The information management table in which the aforementioned text elements are stored, An analysis unit inputs data corresponding to the match status of the sports program from each of multiple information sources, and analyzes the data according to a pre-set data format of the information source to extract the text elements from the data. A storage unit assigns labels to the text elements extracted by the analysis unit, including the priority of the timing for presenting the utterance of the text element, and stores the labeled text elements in the information management table. An update monitoring unit monitors whether the text element stored in the information management table has been updated, and outputs the label assigned to the text element if it is determined that it has been updated. A reading unit reads from the information management table one or more corresponding text elements to which one or more labels are assigned in the utterance definition data output by the update monitoring unit, for the utterance in the utterance definition data that includes the label, and outputs one or more labels of the utterance and one or more corresponding text elements. A format conversion unit converts one or more labels and corresponding one or more text elements of the utterance output by the reading unit into a file containing a predetermined playback time. A text generation unit extracts one or more text elements from the file formatted by the format conversion unit, generates the text, and outputs it; The format conversion unit inputs the files for each utterance, determines the order of the utterances based on the priority included in the label of the file for each utterance, and resets the playback time included in the file for each utterance according to the order. A program to function as an order discard control unit that outputs the playback time contained in the file of the first utterance whose order has been determined, and discards the file of the first utterance.

Citation Information

Patent Citations

  • Explanation voice reproduction device and program thereof

    JP2017203827A

  • Audio guidance generation device, audio guidance generation method, and broadcasting system

    WO2018216729A1