Real-time commentary voice generation system

The real-time commentary audio generation system addresses the accessibility issue for visually impaired individuals by using machine learning to generate play-by-play audio based on game video and data, facilitating their enjoyment of sports.

JP7768852B2Active Publication Date: 2025-11-12DENTSU INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022112830
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2025-11-12
Estimated Expiration
2042-07-14

AI Technical Summary

Technical Problem

Conventional sports commentary systems, particularly those using mobile terminals, are inaccessible to visually impaired individuals due to the reliance on animated commentary images, making it difficult for them to enjoy watching sports.

Method used

A real-time commentary audio generation system utilizing machine learning to analyze the relationship between video and commentary information, enabling automatic generation of play-by-play audio based on video and factual game data, and selecting appropriate camera angles for live broadcast.

Benefits of technology

Enables real-time generation of live audio commentary for visually impaired individuals, allowing them to enjoy sports games without visual barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007768852000001
    Figure 0007768852000001
  • Figure 0007768852000002
    Figure 0007768852000002
  • Figure 0007768852000003
    Figure 0007768852000003
Patent Text Reader

Abstract

To provide a live audio real time generation system capable of automatically generating live audio of a target sport game in real time.SOLUTION: Upon input of video of a prescribed scene of a target sport game for which live audio is to be generated real time, a live audio real time generation system 1 infers, on the basis of a relationship analyzed by a first machine learning unit 14 and by using the inputted video of the prescribed scene of the target sport game as input, live information on the prescribed scene of the target sport game, and outputs the inferred live information. Upon acquisition of fact information relating to the target sport game, the system infers, on the basis of a relationship analyzed by a second machine learning unit 15 and by using the acquired fact information and outputted live information as input, live audio in the prescribed scene of the target sport game, and outputs the inferred the live audio.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a real-time commentary audio generation system that generates commentary audio for a predetermined scene of a sports game in real time. [Background technology]

[0002] Conventionally, a sports commentary system using a mobile terminal has been proposed (see, for example, Patent Document 1). In the conventional system, animated commentary images are displayed on the mobile device, and sports commentary is performed in real time. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-260554 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional systems display animated baseball commentary screens (for example, a screen that displays animations of the pitcher throwing the ball and the batter swinging the bat), but this makes it difficult for visually impaired people (those who have difficulty seeing the screen) to enjoy watching sports. Up until now, no systems have been proposed that enable visually impaired people to enjoy watching sports (for example, a system that generates live audio of the scene in real time).

[0005] The present invention has been made in view of the above-mentioned problems, and aims to provide a real-time commentary audio generation system that can generate commentary audio for a specific scene of a sports game in real time. [Means for solving the problem]

[0006] The real-time commentary audio generation system of the present invention includes a first machine learning unit that analyzes, using machine learning, the relationship between video of a predetermined scene in a past sports game and commentary information for the predetermined scene in the past sports game; a second machine learning unit that analyzes, using machine learning, the relationship between fact information about the past sports game, the commentary information for the predetermined scene in the past sports game, and commentary audio for the predetermined scene in the past sports game; a video input unit that receives video of the predetermined scene in a target sports game for which commentary audio is to be generated in real time; a first estimation unit that, based on the relationship analyzed by the first machine learning unit, estimates and outputs commentary information for the predetermined scene in the target sports game using the video of the predetermined scene in the target sports game input from the video input unit; a fact information acquisition unit that acquires fact information about the target sports game; and a second estimation unit that, based on the relationship analyzed by the second machine learning unit, estimates and outputs commentary audio for the predetermined scene in the target sports game using the fact information about the target sports game acquired by the fact information acquisition unit and the commentary information for the predetermined scene in the target sports game output from the first estimation unit.

[0007] According to this configuration, first, when a video of a predetermined scene in a target sports game (e.g., baseball) for which a play-by-play voice is to be generated in real time is input, play-by-play information for the predetermined scene in the target sports game (e.g., "The pitcher has thrown. The batter has struck out and missed.") is estimated using the relationships analyzed by the first machine learning unit. Next, when fact information about the target sports game (e.g., player names "Pitcher A, Batter B," ball count "One ball, one strike," out count "No outs," score "0-0," etc.) is acquired, play-by-play voice for the predetermined scene in the target sports game (e.g., "Pitcher A has thrown. Batter B has struck out and missed. One ball, one strike.") is estimated using the relationships analyzed by the second machine learning unit. In this way, a play-by-play voice for the target sports game can be automatically generated in real time.

[0008] Furthermore, the real-time live audio generation system of the present invention includes a third machine learning unit that analyzes, through machine learning, the relationship between a plurality of predetermined scene videos captured by a plurality of cameras installed at the venue of a past sports game and a predetermined scene video from the plurality of predetermined scene videos to be used for live broadcast of the past sports game; and a third estimation unit that, based on the relationship analyzed by the third machine learning unit, inputs the plurality of predetermined scene videos captured by a plurality of cameras installed at the venue of the target sports game and estimates and outputs a predetermined scene video from the plurality of predetermined scene videos to be used for live broadcast of the target sports game, and the predetermined scene video to be used for live broadcast of the target sports game output from the third estimation unit may be input to the video input unit.

[0009] With this configuration, when a plurality of videos of predetermined scenes captured by a plurality of cameras installed at the venue of the target sports game are input, the relationship analyzed by the third machine learning unit is used to estimate the video of the predetermined scene to be used for live broadcast of the target sports game. In this way, the video of the predetermined scene to be input to the video input unit (the video of the predetermined scene to be used for real-time generation of the live audio of the target sports game) can be appropriately selected from the plurality of videos of predetermined scenes captured by the plurality of cameras installed at the venue of the target sports game.

[0010] The method for generating commentary audio in real time of the present invention includes a first machine learning step of analyzing, using machine learning, the relationship between video of a predetermined scene of a past sports game and commentary information for the predetermined scene of the past sports game; a second machine learning step of analyzing, using machine learning, the relationship between fact information about the past sports game, the commentary information for the predetermined scene of the past sports game, and commentary audio for the predetermined scene of the past sports game; a video input step of inputting video of the predetermined scene of a target sports game for which commentary audio is to be generated in real time; a first estimation step of estimating and outputting commentary information for the predetermined scene of the target sports game using the video of the predetermined scene of the target sports game input from the video input step based on the relationship analyzed in the first machine learning step; a fact information acquisition step of acquiring fact information about the target sports game; and a second estimation step of estimating and outputting commentary audio for the predetermined scene of the target sports game using the fact information about the target sports game acquired in the fact information acquisition step and the commentary information for the predetermined scene of the target sports game output from the first estimation step based on the relationship analyzed in the second machine learning step.

[0011] Similar to the above system, this method also uses the relationships analyzed by the first machine learning unit to estimate commentary information for a given scene in a target sports game (e.g., baseball) for which real-time commentary audio is to be generated. The first machine learning unit then uses the relationships analyzed by the second machine learning unit to estimate commentary information for the given scene in the target sports game (e.g., "The pitcher threw the ball. The batter struck out and missed."). Next, factual information about the target sports game (e.g., player names "Pitcher A, Batter B," ball count "One ball, one strike," out count "No outs," score "0-0") is acquired. The second machine learning unit then uses the relationships analyzed by the second machine learning unit to estimate commentary information for the given scene in the target sports game (e.g., "Pitcher A threw the ball. Batter B struck out and missed. One ball, one strike."). In this way, commentary audio for a target sports game can be automatically generated in real time. [Effects of the Invention]

[0012] According to the present invention, live audio of a target sports game can be automatically generated in real time. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a block diagram showing a configuration of a real-time live audio generation system according to an embodiment of the present invention; [Figure 2] 1 is a diagram schematically illustrating an example of a plurality of cameras installed at a sports game venue in an embodiment of the present invention. [Figure 3] 3 is a flow diagram for explaining the operation of the real-time live audio generation system according to the embodiment of the present invention. FIG. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, a real-time commentary audio generation system according to an embodiment of the present invention will be described with reference to the drawings. In this embodiment, a real-time commentary audio generation system used in a system for allowing visually impaired people to enjoy watching sports will be exemplified.

[0015] The configuration of a real-time commentary audio generation system according to an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing the configuration of the real-time commentary audio generation system according to this embodiment. As shown in FIG. 1, the real-time commentary audio generation system 1 is connected to a sports game filming system 2 via a network N. The sports game filming system 2 includes a plurality of cameras 21 that capture video of a game such as baseball, and a video distribution unit 22 that distributes the captured video via the network N. As shown in FIG. 2, a plurality of cameras 21 (camera A, camera B, camera C, etc.) are set up at various positions in a sports game venue S so that video of various scenes of the sports game can be acquired.

[0016] 1, the live commentary real-time production system 1 includes a video acquisition unit 11 that acquires video (camera video) of the game being played delivered from a sports game filming system 2, a video input unit 12 that receives video of a predetermined scene from the sports game (target sports game) for which live commentary is to be generated in real time, and a fact information acquisition unit 13 that acquires fact information about the target sports game. For example, if the target sports game for which live commentary is to be generated in real time is "baseball," the fact information may include "player names, ball count, out count, score," etc. The fact information can be acquired, for example, from an information database (not shown) of the sports game organizer.

[0017] The real-time live commentary audio production system 1 also includes a video storage unit 3 that stores video data of past matches. Note that the video storage unit 3 may also store video of the match distributed from the sports game filming system 2.

[0018] Furthermore, the real-time live audio generation system 1 has three machine learning units (first machine learning unit 14, second machine learning unit 15, and third machine learning unit 16) and three estimation units (first estimation unit 17, second estimation unit 18, and third estimation unit 19).

[0019] The first machine learning unit 14 analyzes the relationship between video of a predetermined scene from a past sports game and commentary information about that predetermined scene from the past sports game through machine learning. Any method, such as deep learning using a neural network, can be used for this machine learning. For example, a neural network is configured to input video of a predetermined scene from a past sports game into an input layer and output commentary information about that predetermined scene from the past sports game from an output layer. Then, weighting coefficients between neurons in the neural network are optimized through supervised learning using analysis data that links data input to the input layer and data output from the output layer.

[0020] Based on the relationship analyzed by the first machine learning unit 14, the first estimation unit 17 receives as input a video of a predetermined scene of the target sports game from the video input unit 12, and estimates and outputs commentary information for the predetermined scene of the target sports game. For example, in the case of the above-mentioned neural network, estimation is performed by inputting the video of the predetermined scene of the target sports game from the video input unit 12 into the input layer and outputting commentary information for the predetermined scene of the target sports game from the output layer. For example, if the target sports game for which commentary audio is to be generated in real time is "baseball," the commentary information may be, "The pitcher has thrown. The batter has struck out and missed."

[0021] The second machine learning unit 15 analyzes the relationship between fact information about past sports games, commentary information about predetermined scenes in past sports games, and commentary audio for the predetermined scenes in the past sports games through machine learning. Any method, such as deep learning using a neural network, can be used for this machine learning. For example, a neural network is configured such that fact information about past sports games and commentary information about predetermined scenes in past sports games are input to an input layer, and commentary audio for the predetermined scenes in the past sports games is output from an output layer. Then, weighting coefficients between neurons in the neural network are optimized through supervised learning using analysis data that links data input to the input layer with data output from the output layer.

[0022] Based on the relationship analyzed by the second machine learning unit 15, the second estimation unit 18 receives as input the fact information about the target sports game acquired by the fact information acquisition unit 13 and the commentary information about a predetermined scene of the target sports game output from the first estimation unit 17, and estimates and outputs a commentary audio for the predetermined scene of the target sports game. For example, in the case of the above-mentioned neural network, the fact information about the target sports game acquired by the fact information acquisition unit 13 and the commentary information about the predetermined scene of the target sports game output from the first estimation unit 17 are input to the input layer, and commentary audio for the predetermined scene of the target sports game is output from the output layer, thereby performing estimation. For example, if the target sports game for which commentary audio is to be generated in real time is "baseball," the commentary audio may be something like, "Pitcher A (player name) has pitched. Batter B (player name) has struck out and missed. One ball, one strike."

[0023] The third machine learning unit 16 uses machine learning to analyze the relationship between videos of multiple predetermined scenes captured by multiple cameras 21 installed at the venues of past sports games and videos of the predetermined scenes used in live broadcasts of those past sports games. Any method, such as deep learning using a neural network, can be used for this machine learning. For example, a neural network is configured to input videos of multiple predetermined scenes captured by multiple cameras 21 installed at the venues of past sports games into an input layer, and output videos of the predetermined scenes used in live broadcasts of those past sports games from an output layer. Then, weighting coefficients between neurons in the neural network are optimized through supervised learning using analysis data that links data input to the input layer with data output from the output layer.

[0024] Based on the relationship analyzed by the third machine learning unit 16, the third estimation unit 19 receives as input videos of multiple predetermined scenes captured by multiple cameras 21 installed at the venue of the target sports game, estimates and outputs videos of predetermined scenes to be used for live broadcast of the target sports game from the multiple videos of predetermined scenes. For example, in the case of the above-mentioned neural network, the videos of multiple predetermined scenes captured by multiple cameras 21 installed at the venue of the target sports game are input to the input layer, and the videos of predetermined scenes to be used for live broadcast of the target sports game are output from the output layer, thereby performing estimation. The videos of predetermined scenes to be used for live broadcast of the target sports game output from the third estimation unit 19 are input to the video input unit 12.

[0025] The operation of the real-time commentary audio production system 1 configured as above will be described with reference to the flow chart of FIG.

[0026] In the commentary audio real-time generation system 1 of this embodiment, first, as a preliminary step, the first machine learning unit 14 analyzes, by machine learning, the relationship between video of a predetermined scene in a past sports game and commentary information for that predetermined scene in the past sports game (first machine learning step). The second machine learning unit 15 also analyzes, by machine learning, the relationship between fact information about the past sports game, commentary information for a predetermined scene in the past sports game, and commentary audio for that predetermined scene in the past sports game (second machine learning step). The third machine learning unit 16 also analyzes, by machine learning, the relationship between video of a plurality of predetermined scenes captured by a plurality of cameras 21 installed at the venue of the past sports game and video of the predetermined scene used in the commentary broadcast of the past sports game from among the plurality of predetermined scene videos (third machine learning step).

[0027] 3, when generating live commentary audio of a target sports game in real time, a plurality of camera images of the game distributed from the sports game filming system 2 are acquired (S1), and based on the relationship analyzed by the third machine learning unit 16, a plurality of camera images (images of a plurality of predetermined scenes) installed at the venue of the target sports game are input, and a predetermined scene image (broadcast image) to be used for live broadcast of the target sports game from among the plurality of predetermined scene images is estimated and output (S2). The estimated predetermined scene image (broadcast image) is then input to the image input unit 12. Note that if a live broadcast of the target sports game is actually being performed, step S2 of estimating the broadcast image is unnecessary, and the actual broadcast image is input to the image input unit 12.

[0028] Next, based on the relationship analyzed by first machine learning unit 14, commentary information for the predetermined scene of the target sports game is estimated and output using as input the video of the predetermined scene of the target sports game input from video input unit 12 (S3). Thereafter, fact information regarding the target sports game is acquired by fact information acquisition unit 13 (S4), and based on the relationship analyzed by second machine learning unit 15, commentary audio for the predetermined scene of the target sports game is estimated and output using as input the fact information regarding the target sports game acquired by fact information acquisition unit 13 and commentary information for the predetermined scene of the target sports game output from first estimation unit 17 (S5).

[0029] According to the real-time commentary generation system 1 of this embodiment, first, when a video of a predetermined scene of a target sports game (e.g., baseball) for which commentary is to be generated in real time is input, commentary information (e.g., "The pitcher has thrown. The batter has struck out and missed") for the predetermined scene of the target sports game is estimated using the relationships analyzed by the first machine learning unit 14. Next, when fact information about the target sports game (e.g., player names "Pitcher A, Batter B," ball count "One ball, one strike," out count "No outs," score "0-0," etc.) is acquired, commentary information (e.g., "Pitcher A has thrown. Batter B has struck out and missed. One ball, one strike," etc.) for the predetermined scene of the target sports game is estimated using the relationships analyzed by the second machine learning unit 15. In this way, commentary for the target sports game can be automatically generated in real time.

[0030] Furthermore, in this embodiment, when a plurality of videos of predetermined scenes captured by a plurality of cameras 21 installed at the venue of the target sports game are input, the relationship analyzed by the third machine learning unit 16 is used to estimate the video of the predetermined scene to be used for live broadcast of the target sports game from among the plurality of videos of the predetermined scenes. In this way, the video of the predetermined scene to be input to the video input unit 12 (the video of the predetermined scene to be used for real-time generation of live audio for the target sports game) can be appropriately selected from the plurality of videos of the predetermined scene captured by the plurality of cameras 21 installed at the venue of the target sports game.

[0031] Although the embodiments of the present invention have been described above by way of example, the scope of the present invention is not limited to these, and can be modified and changed according to the purpose within the scope of the claims.

[0032] In the above explanation, we have explained an example in which the sports game for which commentary audio is to be generated in real time is "baseball," where the fact information is "player name, ball count, out count, score," the commentary information is "The pitcher has thrown. The batter has struck out and missed," and the commentary audio is "Pitcher A (player name) has thrown. Batter B (player name) has struck out and missed. One ball, one strike," but this can also be implemented for other sports games.

[0033] For example, if the sports game for which commentary audio is generated in real time is "horse racing," the fact information may be "horse name, horse number, jockey name, ranking," etc., the commentary information may be "All horses start at the same time," etc., and the commentary audio may be "All horses start at the same time. C (horse name) is in the lead," etc.

[0034] Also, if the sports game for which commentary audio is generated in real time is "motorsports," the fact information may be "team name, driver name, ranking," etc., the commentary information may be "the leader was overtaken by the second-place driver on the back straight," etc., and the commentary audio may be "the leader, D (driver name), was overtaken by the second-place driver, E (driver name), on the back straight." [Industrial Applicability]

[0035] As described above, the real-time commentary audio generation system of the present invention has the effect of automatically generating commentary audio for a target sports game in real time, and is useful as a system for visually impaired people to enjoy watching sports. [Explanation of symbols]

[0036] 1. Real-time commentary voice generation system 2. Sports Game Filming System 3. Video storage unit 11 Video acquisition unit 12 Video input section 13 Fact Information Acquisition Unit 14 Machine Learning Department 1 15 Second Machine Learning Department 16 Third Machine Learning Department 17 1st estimation part 18 Second estimation part 19 Third estimation part 21 Camera 22 Video Distribution Department N Network

Claims

1. a first machine learning unit that analyzes, by machine learning, a relationship between a video of a predetermined scene of a past sports game and live commentary information of the predetermined scene of the past sports game; a second machine learning unit that analyzes, by machine learning, a relationship between fact information regarding the past sports games, commentary information for a predetermined scene of the past sports games, and commentary audio for the predetermined scene of the past sports games; a video input unit to which video of a predetermined scene of a target sports game for which live commentary audio is to be generated in real time is input; a first estimation unit that estimates and outputs commentary information for a predetermined scene of the target sports game based on the relationship analyzed by the first machine learning unit, using as input a video of the predetermined scene of the target sports game input from the video input unit; and a fact information acquisition unit that acquires fact information regarding the target sports game; a second estimation unit that estimates and outputs a commentary audio for a predetermined scene of the target sports game based on the relationship analyzed by the second machine learning unit, using as input the fact information related to the target sports game acquired by the fact information acquisition unit and commentary information for a predetermined scene of the target sports game output from the first estimation unit; A real-time commentary audio generation system comprising:

2. a third machine learning unit that analyzes, by machine learning, a relationship between a plurality of predetermined scene images captured by a plurality of cameras installed at the venue of a past sports game and a predetermined scene image used for live broadcast of the past sports game among the plurality of predetermined scene images; a third estimation unit that receives as input videos of a plurality of predetermined scenes captured by a plurality of cameras installed at the venue of the target sports game based on the relationship analyzed by the third machine learning unit, estimates and outputs a video of a predetermined scene to be used for live broadcast of the target sports game from among the videos of the plurality of predetermined scenes; Equipped with The live audio real-time generation system according to claim 1 , wherein the video of a predetermined scene used in live broadcast of the target sports game output from the third estimation unit is input to the video input unit.

3. a first machine learning step of analyzing, by machine learning, a relationship between a video of a predetermined scene of a past sports game and live commentary information of the predetermined scene of the past sports game; a second machine learning step of analyzing, by machine learning, a relationship between fact information regarding the past sports games, commentary information for a predetermined scene of the past sports games, and commentary audio for the predetermined scene of the past sports games; a video input step of inputting a video of a predetermined scene of a target sports game for which live commentary audio is to be generated in real time; a first estimation step of estimating and outputting commentary information for a predetermined scene of the target sports game using the video of the predetermined scene of the target sports game input from the video input step based on the relationship analyzed in the first machine learning step; a fact information acquisition step of acquiring fact information about the target sports game; a second estimation step of estimating and outputting a commentary audio for a predetermined scene of the target sports game based on the relationship analyzed in the second machine learning step, using as input the fact information regarding the target sports game acquired in the fact information acquisition step and commentary information for a predetermined scene of the target sports game output from the first estimation step; A method for generating commentary audio in real time, including:

Citation Information

Patent Citations

  • Sport play-by-play broadcasting system utilizing portable equipment

    JP2006260554A

  • Computer system and audio information generation method

    JP2021194229A

  • Information processing program, device, and method

    JP2022067478A

  • Server, method and computer program for providing sports broadcasting script

    KR1020210026650A

  • Audio guidance generation device, audio guidance generation method, and broadcasting system

    WO2018216729A1