Method and apparatus for generating audio commentary of sports game, and instruction-recorded recording medium
The system addresses the limitations of existing sports commentary by generating real-time, personalized commentary using multiple data sources and learning models, improving viewer engagement and enjoyment.
Patent Information
- Application Number
- PCT/KR2025/007118
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-26
- Publication Date
- 2025-11-27
AI Technical Summary
Existing sports commentary systems lack the ability to provide real-time, personalized, and comprehensive analysis of sports events, limiting viewer engagement and enjoyment, especially for those without extensive sports knowledge.
A system that analyzes sports games in real-time using multiple cameras, microphones, and sensors to generate first data, combines this with statistical data and user preferences to determine broadcast screens and generate voice commentary using learning models and text-to-speech technology.
Enriches the viewer experience by providing personalized and comprehensive commentary, enhancing engagement and enjoyment of sports events through real-time analysis and tailored commentary.
Smart Images

Figure KR2025007118_27112025_PF_FP_ABST
Abstract
Description
Method for generating voice commentary for a sports game, device and recording medium recording commands
[0001] The present disclosure relates to a technique for generating audio commentary for a sports game.
[0002] With the advancement of communication technology using electronic devices, people around the world are increasingly watching sports events through mass media. When viewers access mass media, including TV and online platforms like streaming services, audio commentary is provided to enrich their experience. Audio commentary provides information that allows even those without extensive sports knowledge to enjoy the game, and by providing appropriate commentary based on the flow of the game, it plays a crucial role in enhancing viewers' enjoyment of the sport.
[0003] Audio commentary can be provided by commentators broadcasting sports events. For example, commentators can use information gathered directly from the field to provide commentary based on their own perspectives, and this commentary can be provided to viewers in the form of audio data.
[0004] At least one embodiment of the present disclosure provides a technique for providing audio commentary of a sports event.
[0005] At least one embodiment of the present disclosure can analyze a sports game in real time by acquiring data related to the game situation.
[0006] At least one embodiment of the present disclosure can analyze a sports game in real time to determine which broadcast screen to broadcast.
[0007] According to one aspect of the present disclosure, a method performed in a device including one or more processors and one or more memories storing instructions to be executed by the one or more processors may include: the one or more processors obtaining, from one or more cameras, image data of a sports event; generating, based on the image data, first data related to a game situation of the sports event; obtaining second data including statistical data regarding an object of the sports event; generating, based on the first data and the second data, first analysis data indicating an analysis result for the sports event; determining, based on the first analysis data, at least a portion of the image data as a broadcast screen to be transmitted to a user's terminal; and generating, based on the first analysis data, a voice commentary corresponding to the broadcast screen.
[0008] In one embodiment, the one or more cameras include a first camera and a second camera that capture an object of the sports event from different angles, and the step of generating the first data may include: determining a first position of the object in a capture screen of the first camera based on the image data; determining a second position of the object in a capture screen of the second camera based on the image data; determining a three-dimensional position coordinate of the object based on the first position and the second position; and generating the first data based on a change in the three-dimensional position coordinate of the object over time.
[0009] In one embodiment, the step of generating the first data may further include the step of receiving voice data generated inside the playing space where the sports game is being played from one or more microphones; the step of receiving sensor data including information on temperature and humidity of the playing space from one or more sensors; and the step of generating the first data based on the voice data and the sensor data.
[0010] In one embodiment, the first data may include at least one of a player's position and movement in the sports game, a ball's position and movement, a player's facial expression, or a crowd reaction.
[0011] In one embodiment, the second data may include at least one of information regarding the tactics of each team, the characteristics of each player, the referee's tendencies, the opponent's record, interviews of each team before the sports game, or analysis regarding the sports game.
[0012] In one embodiment, the step of generating the first analysis data may include a step of analyzing at least one of the tactics of each team in the sports game, the movements of each player in the sports game, an event occurring in the sports game, or statistics related to the sports game, based on the first data and the second data.
[0013] In one embodiment, the step of determining the relay screen may further include: obtaining first image data from the one or more cameras at a first point in time; obtaining second image data from the one or more cameras at a second point in time that is after the first point in time; generating the first analysis data based at least on the first image data and the second image data; and determining the first image data as the relay screen based on the first analysis data.
[0014] In one embodiment, the step of generating the voice commentary may include the step of inputting the first analysis data for the sports game into a learning model learned based on second analysis data of one or more sports games and a commentary script corresponding to the second analysis data, thereby generating a commentary script for the sports game.
[0015] In one embodiment, the step of generating the voice commentary may further include the step of generating the voice commentary based on the commentary script using a text-to-speech technology.
[0016] In one embodiment, the step of generating the voice commentary may include the step of generating a voice corresponding to the analysis result for the sports game among at least one pre-stored voice as the voice commentary.
[0017] In one embodiment, the step of generating the voice commentary may include: collecting viewing history information including information about one or more sports games played on the user's terminal and voice commentary for each of the one or more sports games; generating preference data by analyzing characteristics of voice commentary preferred by the user based on the viewing history information; and generating the voice commentary based on the preference data.
[0018] In one embodiment, the preference data may include at least one of whether the user prefers commentary on a specific player, whether the user prefers commentary on a specific team, whether the user prefers commentary from a neutral standpoint, or whether the user prefers voice commentary with an average volume greater than a predetermined value.
[0019] According to another aspect of the present disclosure, a device includes: a communication circuit; one or more processors; and one or more memories storing instructions executed by the one or more processors, wherein the one or more processors are configured to obtain image data of a sports event from one or more cameras, generate first data related to a game situation of the sports event based on the image data, obtain second data including statistical data regarding an object of the sports event, generate first analysis data indicating an analysis result for the sports event based on the first data and the second data, determine at least a portion of the image data as a broadcast screen to be transmitted to a user's terminal based on the first analysis data, and generate a voice commentary corresponding to the broadcast screen based on the first analysis data.
[0020] In one embodiment, the one or more cameras include a first camera and a second camera that capture an object of the sports event from different angles, and the one or more processors can determine a first position of the object in a capture screen of the first camera based on the image data, determine a second position of the object in a capture screen of the second camera based on the image data, determine a three-dimensional position coordinate of the object based on the first position and the second position, and generate the first data based on a change in the three-dimensional position coordinate of the object over time.
[0021] In one embodiment, the one or more processors may obtain first image data from the one or more cameras at a first point in time, obtain second image data from the one or more cameras at a second point in time that is after the first point in time, generate the first analysis data based at least on the first image data and the second image data, and determine the first image data as a relay screen based on the first analysis data.
[0022] In one embodiment, the one or more processors may input the first analysis data for the sports game into a learning model learned based on second analysis data of one or more sports games and a commentary script corresponding to the second analysis data, thereby generating a commentary script for the sports game.
[0023] In one embodiment, the one or more processors may generate a voice commentary based on the commentary script using speech synthesis technology.
[0024] In one embodiment, the one or more processors may generate, as the voice commentary, a voice corresponding to the analysis result for the sports game among at least one pre-stored voice.
[0025] In one embodiment, the one or more processors may collect viewing history information including information about one or more sports games played on the user's terminal and audio commentary for each of the one or more sports games, generate preference data by analyzing characteristics of audio commentary preferred by the user based on the viewing history information, and generate the audio commentary based on the preference data.
[0026] According to another aspect of the present disclosure, a non-transitory computer-readable recording medium having recorded thereon instructions to be executed by one or more processors, wherein the instructions, when executed, cause the one or more processors to obtain image data of a sports event from one or more cameras, generate first data related to a game situation of the sports event based on the image data, obtain second data including statistical data regarding an object of the sports event, generate first analysis data indicating an analysis result for the sports event based on the first data and the second data, determine at least a portion of the image data as a broadcast screen to be transmitted to a user's terminal based on the first analysis data, and generate a voice commentary corresponding to the broadcast screen based on the first analysis data.
[0027] According to various embodiments of the present disclosure, audio commentary of a sports game can be provided.
[0028] According to various embodiments of the present disclosure, data related to game situations can be acquired to analyze a sports game in real time.
[0029] According to various embodiments of the present disclosure, a sports game can be analyzed in real time to determine a broadcast screen to be transmitted.
[0030] The effects according to the technical idea of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person skilled in the art from the description of the specification.
[0031] Figure 1 is a drawing showing the operation process of a device according to one embodiment of the present disclosure.
[0032] FIG. 2 is a block diagram of a device according to one embodiment of the present disclosure.
[0033] FIG. 3 is a diagram illustrating a process for generating position data and expression data of an object according to one embodiment of the present disclosure.
[0034] FIG. 4 is a diagram illustrating a process for generating movement data of an object according to one embodiment of the present disclosure.
[0035] FIG. 5 is a diagram illustrating a process for generating movement data of an object according to one embodiment of the present disclosure.
[0036] FIG. 6 is a diagram illustrating a process for generating position data of an object according to one embodiment of the present disclosure.
[0037] FIG. 7 is a diagram illustrating a process for generating movement data of an object according to one embodiment of the present disclosure.
[0038] FIG. 8 is a diagram illustrating an example of first analysis data generated by a device according to one embodiment of the present disclosure based on first data and second data.
[0039] FIG. 9 is a diagram illustrating a process of obtaining a commentary script using a learning model learned to generate a script of voice commentary in a sports game according to one embodiment of the present disclosure.
[0040] FIG. 10 is a flowchart of a method for generating voice commentary for a sports game according to one embodiment of the present disclosure.
[0041] FIG. 11 is a flowchart of a method for generating voice commentary based on viewing history information according to one embodiment of the present disclosure.
[0042] The various embodiments described in this document are exemplified for the purpose of clearly explaining the technical concept of the present disclosure and are not intended to limit it to a specific embodiment. The technical concept of the present disclosure includes various modifications, equivalents, alternatives, and embodiments selectively combining all or part of each embodiment described in this document. Furthermore, the scope of the technical concept of the present disclosure is not limited to the various embodiments presented below or the specific descriptions thereof.
[0043] Terms used in this document, including technical or scientific terms, unless otherwise defined, may have the meaning commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0044] The expressions "includes," "may include," "comprises," "may have," "have," and "may have" used in this document imply the presence of a target feature (e.g., a function, operation, or component), but do not exclude the presence of other additional features. In other words, such expressions should be understood as open-ended terms that imply the possibility of including other embodiments.
[0045] The singular forms used in this document may include the plural form unless the context clearly indicates otherwise, and this also applies to the singular forms set forth in the claims.
[0046] The expressions "first," "second," or "first", "second", etc. used in this document, unless the context indicates otherwise, are used to refer to multiple similar objects and to distinguish one object from another, and do not limit the order or importance among the objects.
[0047] As used herein, the expressions "A, B, and C", "A, B, or C", "A, B, and / or C", or "at least one of A, B, and C", "at least one of A, B, or C", "at least one of A, B, and / or C", etc., may refer to each of the listed items or to all possible combinations of the listed items. For example, "at least one of A or B" may refer to (1) at least one A, (2) at least one B, (3) at least one A and at least one B.
[0048] The expression "based on" as used in this document is used to describe one or more factors that influence the decision, act of judgment, or action described in the phrase or sentence containing the expression, and this expression does not exclude additional factors that influence the decision, act of judgment, or action.
[0049] In the present disclosure, determining C information based on A information and B information means taking A information and B information into consideration in determining C information, and does not exclude that information other than A information and B information is additionally considered.
[0050] As used herein, the expression that a component (e.g., a first component) is “connected” or “connected” to another component (e.g., a second component) may mean that the component is directly connected or connected to the other component, as well as connected or connected via a new other component (e.g., a third component).
[0051] The expression "configured to" as used in this document can mean "set to do", "having the ability to do", "modified to do", "made to do", "capable of doing", etc., depending on the context. The expression is not limited to the meaning of "specifically designed in hardware", and for example, a processor configured to perform a specific operation can mean a generic-purpose processor that can perform the specific operation by executing software.
[0052] In the present disclosure, a "learning model" may be designed to implement the structure of a human brain on a computer, and may include a plurality of network nodes that simulate neurons of a human neural network and have weights. The plurality of network nodes simulate the synaptic activity of neurons that exchange signals through synapses and may have connections among themselves. In the learning model, the plurality of network nodes may be located at layers of different depths and exchange data according to convolutional connections. For example, the learning model may be an artificial neural network model, a time series analysis model, etc. Here, the time series analysis model is a model that analyzes time series data, and may be a recurrent neural network (RNN) model, a long short-term memory (LSTM) model, or an Autoregressive Integrated Moving Average (ARIMA) model. Meanwhile, the present disclosure is not limited thereto, and various types of models for analyzing data may be used.
[0053] In the present disclosure, a “training process” may mean a process in which a learning model extracts and analyzes features (patterns) of input data and output data pairs of learning data, repeats the process of deriving correlations between input and output data, and optimizes parameters of the learning model based on the correlations between input and output data.
[0054] In the present disclosure, an “inference process” may mean a process in which a learning model applies a previously learned pattern to new input data to generate output data as a result of prediction or classification of the input data.
[0055] Hereinafter, various embodiments of the present disclosure will be described with reference to the attached drawings. In the attached drawings and the description of the drawings, identical or substantially equivalent components may be assigned the same reference numerals. Furthermore, in the description of various embodiments below, duplicate descriptions of identical or corresponding components may be omitted, but this does not mean that the corresponding components are not included in the embodiments.
[0056] FIG. 1 is a diagram illustrating the operation process of a device (100) according to one embodiment of the present disclosure. In the present disclosure, the device (100) may be a device (e.g., a server) that generates audio commentary for a sports game. In the present disclosure, the sports games may include various ball games and non-ball games, but are not limited to any one of them.
[0057] In the present disclosure, a camera group (130) may include one or more cameras (131, 132, 133, 134, 135, 136). The camera group (130) may capture a game space (110) of a sporting event from one or more angles through one or more cameras (131, 132, 133, 134, 135, 136), thereby generating image data (138) of the sporting event.
[0058] For example, the game space (110) may refer to a predetermined three-dimensional space related to a sports game. Specifically, the game space (110) may include an area where a sports game takes place (e.g., a court), an area where objects (120) of the sports game may be located, an area where spectators, analysts, or referees of the sports game may be located, an area where stadium facilities (e.g., an electronic scoreboard) may be located, etc. Meanwhile, the object (120) may include at least one of objects or a ball located in the game space (110), such as players and referees of the sports game, for example.
[0059] For example, at least some of the cameras (131, 132, 133, 134, 135, 136) of the camera group (130) may be fixed cameras, and the others may be tracking cameras. The fixed camera may refer to a camera that photographs at least a portion of the playing space (110) at a predetermined angle, magnification, etc. The tracking camera may be a camera that tracks and photographs at least one object (120) within the playing space (110). The tracking camera may track and photograph the object (120) while moving according to a preset, and the preset may be one or more settings for camera parameters such as a shooting angle, magnification, etc. Meanwhile, FIG. 1 illustrates six cameras (131, 132, 133, 134, 135, 136) that photograph the playing space (110), but the present disclosure is not limited thereto. The number and arrangement of cameras within a camera group (130) may change depending on the type of sport, the shape of the stadium (e.g., size), the number of spectators, etc.
[0060] For example, each camera of the camera group (130) can capture a game space (110) of a sports game at a predetermined frame rate (e.g., 120 fps) to generate image data (138) of the corresponding sports game. Here, the image data (138) can include an image set obtained by dividing the captured video of the sports game into units of shooting frames (e.g., 120 fps).
[0061] The microphone (140) can generate voice data (142) for voices generated within the playing space (110). The microphone (140) may refer to a microphone group including one or more microphones (140) placed within the playing space (110) to receive voice data (142) generated within the playing space (110). The microphone (140) can generate voice data (142) for various voices generated during a sports game and transmit the data to the device (100).
[0062] The sensor (150) may be a device that generates or acquires sensor data (152) corresponding to a current state. The sensor (150) may generate sensor data (152) corresponding to an environmental state within the playing space (110). For example, the sensor (150) may include a temperature sensor, a pressure sensor, a humidity sensor, or an illuminance sensor. The sensor (150) may generate sensor data (152) including information on at least one of temperature, pressure, humidity, or illuminance within the current playing space (110) and transmit the data to the device (100).
[0063] In the present disclosure, the external device (160) may provide second data (162) containing various data regarding a sports game to the device (100). In one embodiment, the second data (162) may refer to information collected from one or more previously held sports games, rather than information collected in real time from a currently ongoing sports game. The second data (162) may be statistical data regarding the sports game. The external device (160) may statistically process the data regarding one or more received sports games to generate the second data (162). Alternatively, the external device (160) may receive the second data (162) and provide it to the device (100).
[0064] In the present disclosure, the user terminal (170) may be a terminal of a user (e.g., analyst, viewer) using a sports game broadcasting service.
[0065] In one embodiment, the device (100) may receive image data (138) for a target sporting event from a camera group (130), audio data (142) from a microphone (140), sensor data (152) from a sensor (150), and second data (162) from an external device (160). The device (100) may generate first analysis data (172) indicating an analysis result on the progress of the sport based on at least one of the first data (154) including the received image data (138), audio data (142), and sensor data (152), and the second data (162). Here, the target sporting event may mean a sport for which an audio commentary is to be generated among one or more sporting events. The device (100) can generate relay screen data (174) and voice commentary data (176) to be transmitted to the user terminal (170) based on the first analysis data (172).
[0066] The first data (154) may include at least one of position data (125) indicating the position of a player and / or a ball in the sports game, movement data (180) indicating movement, facial expression data (127) regarding the facial expression of a player, or audience reactions. The device (100) may generate frame-by-frame position data of an object (120) in the target sports game based on image data (138). Here, the frame-by-frame position data may include three-dimensional coordinates where the object (120) is located.
[0067] For example, the device (100) may generate movement data regarding the movement of an object (120) in a target sports game based on frame-by-frame position data. Here, the movement data of the object (120) may include metrics indicating the movement of the object (120), such as the movement direction, movement speed, movement time, and movement trajectory of the object (120) in the target sports game.
[0068] The second data (162) may include at least one of the following: main tactics, characteristics of each player, referee tendencies, records of opponents, interviews before a sports game, or analysis data from experts.
[0069] The first analysis data (172) may indicate the current progress of a sports game, analyzed based on the first data (154) and the second data (162). For example, the first analysis data (172) may include an analysis result for at least one of the tactics (or tactical changes) of each team in the sports game, the movement of each player in the sports game (e.g., overlapping, volley shot, dribbling breakthrough, etc.), an event occurring in the sports game, or statistics related to the sports game (e.g., ball possession, pass success rate, number of effective shots, etc.). The events occurring in the sports game may include events related to the rules of the target sports game (e.g., free kick, penalty kick, corner kick, offside, tackle, transfer of possession, foul, goal, goal conceded).
[0070] In one embodiment, the device (100) may generate relay screen data (174) for a relay screen to be transmitted to the user terminal (170) based on the first analysis data (172). According to one embodiment, the device (100) may determine, as the relay screen data (174), the image data (138) that best represents the current progress of the target sporting event among the image data (138) received from the camera group (130). According to one embodiment, the device (100) may generate relay screen data (174) including a replay scene as an event occurs in the target sporting event. For example, when the device (100) receives first image data at a first time point and receives second image data at a second time point that is after the first time point, the device (100) may generate relay screen data (174) including the first image data based on the first analysis data (172) indicating that an event occurred at the first time point.
[0071] In one embodiment, the device (100) may generate a commentary script corresponding to the broadcast screen data (174) of the target sports game based on the first analysis data (172). For example, the device (100) may input the first analysis data (172) for the target sports game into a learning model trained using second analysis data for one or more sports games as learning data, thereby generating a commentary script corresponding to the broadcast screen data (174).
[0072] In one embodiment, the device (100) may generate a commentary script further based on user preference data. The device (100) may collect viewing history information including information about one or more sports games played on the user's terminal (170) and the audio commentary for the sports games. Based on the viewing history information, the device (100) may generate preference data by analyzing the characteristics of the audio commentary preferred by the user. In one embodiment, the preference data may include information about the user's preferred tone of voice (or intonation) of the commentary. For example, the preference data may include information about whether the user prefers a calm tone of voice or an excited tone of voice. For example, the device (100) may generate preference data based on the average volume of audio commentaries played on the user's terminal (170) based on the user's viewing history. In one embodiment, the preference data may include information about whether the user has a preference for a specific team or a specific player. For example, the device (100) may determine that the user prefers a particular team or a particular player based on the user's determination that the user watches sports games of a particular team or sports games featuring a particular player more frequently than a predetermined frequency. The device (100) may then generate a commentary script based on this preference data.
[0073] In one embodiment, the device (100) may generate voice commentary data (176) corresponding to the broadcast screen data (174) based on the generated commentary script and transmit the same to the user terminal (170). For example, the device (100) may select voice recording data corresponding to the commentary script from among one or more pre-recorded voice recording data to generate the voice commentary data (176). According to another embodiment, the device (100) may generate voice commentary data (176) corresponding to the commentary script using a voice synthesis technology (Text To Speech, TTS). The device (100) may transmit the generated voice commentary data (176) to the user terminal (170), thereby providing an environment in which the user can watch a sports game with voice commentary.
[0074] FIG. 2 is a block diagram of a device (100) according to one embodiment of the present disclosure. In one embodiment, the device (100) may include one or more processors (210) and / or one or more memories (220) as elements. In one embodiment, at least one of the elements of the device (100) may be omitted, or another element may be added to the device (100). In one embodiment, additionally or alternatively, some of the elements may be implemented in an integrated manner or implemented as a single or multiple entities. One or more processors may be referred to as processors (210). The expression “processor (210)” may mean a set of one or more processors, unless the context clearly indicates otherwise. One or more memories may be referred to as memories (220). The expression “memory (220)” may mean a set of one or more memories, unless the context clearly indicates otherwise. At least some of the components inside / outside the device (100) are connected to each other through a bus, GPIO (general purpose input / output), SPI (serial peripheral interface), MIPI (mobile industry processor interface), etc., and can exchange information (data, signals, etc.).
[0075] In one embodiment, the processor (210) may execute instructions (e.g., code, software, program, etc.) to control at least one component of the device (100) connected to the processor (210). In addition, the processor (210) may perform various operations such as calculations, processing, data generation, and processing related to the present disclosure. In addition, the processor (210) may load data, etc. from the memory (220) or store data, etc. in the memory (220). For example, the processor (210) may receive image data (138) from a camera group (130) that captures a target sporting event, receive voice data (142) from a microphone (140), receive sensor data (152) from a sensor (150), and receive second data (162) from an external device (160). The processor (210) can generate first analysis data (172) indicating the current progress of the target sports game based on the received data, and can generate at least one of the broadcast screen data (174) and the voice commentary data (176) based on the first analysis data (172) and transmit the generated data to the user terminal (170).
[0076] In one embodiment, the memory (220) may store various data. The data stored in the memory (220) may include instructions (e.g., code, software, programs, etc.) as data acquired, processed, or used by at least one component of the device (100). The memory (220) may include volatile and / or non-volatile memory. The instructions or programs are software stored in the memory (220) and may include an operating system for controlling the resources of the device (100), applications, and / or middleware that provides various functions to applications so that the applications can utilize the resources of the device (100). For example, the memory (220) may store instructions that, when executed by the processor (210), cause the processor (210) to perform operations. The memory (220) can store image data (138), location data (125), movement data (180), voice data (142), sensor data (152), second data (162), first analysis data (172), broadcast screen data (174), and voice commentary data (176) of the target sporting event. In addition, the memory (220) can store image data, movement data, voice data, sensor data, second data, first analysis data, broadcast screen data, and voice commentary data of one or more other sporting events. In one embodiment, the processor (210) can control the communication circuit to obtain information from another server or device. The information obtained in this way can also be stored in the memory (220).
[0077] In one embodiment, the device (100) may further include a communication circuit (230). The communication circuit (230) may be omitted from the device (100) depending on the embodiment. The communication circuit (230) may perform wireless or wired communication between the device (100) and another server, or between the device (100) and another device. For example, the communication circuit (230) may perform wireless communication according to a method such as eMBB (enhanced Mobile Broadband), URLLC (Ultra Reliable Low-Latency Communications), MMTC (Massive Machine Type Communications), LTE (Long-Term Evolution), LTE-A (LTE Advance), NR (New Radio), UMTS (Universal Mobile Telecommunications System), GSM (Global System for Mobile communications), CDMA (Code Division Multiple Access), WCDMA (Wideband CDMA), WiBro (Wireless Broadband), WiFi (Wireless Fidelity), Bluetooth (Bluetooth), NFC (Near Field Communication), GPS (Global Positioning System), or GNSS (Global Navigation Satellite System). For example, the communication circuit (230) may perform wired communication according to a method such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), RS-232 (Recommended Standard-232), or POTS (Plain Old Telephone Service).For example, the communication circuit (230) may be used to communicate with a camera group (130), a microphone (140), a sensor (150), an external device (160), or a user terminal (170). The communication circuit (230) may be implemented as a circuit or chip configured to transmit and receive data.
[0078] In one embodiment, the device (100) may further include an input / output interface. The input / output interface may be omitted from the device (100) depending on the embodiment. For example, the input / output interface may include an input device and / or an output device. The input device may receive various information from an e-commerce service provider and transmit the input information to at least one component of the device (100). The output device may receive various information output by at least one component of the device (100) and provide (display) the information to the service provider in an audiovisual form. For example, the input device may include a mouse, a keyboard, a touch pad, etc. For example, the output device may include a display, a projector, a hologram, etc.
[0079] In one embodiment, the device (100) may be a device of various forms. For example, the device (100) may be a computer device, a back-end server, a front-end server, a portable communication device, a portable multimedia device, a wearable device, a device according to a combination of the aforementioned devices, or a chip, board, circuit, etc. within the aforementioned devices. However, the device (100) of the present disclosure is not limited to the aforementioned devices.
[0080] Hereinafter, a method for generating position data (125) and movement data (180) of an object (120) in FIGS. 3 to 7 will be described in detail. The position data (125) and movement data (180) may be generated based on image data (138) and may be included in the first data (154).
[0081] FIG. 3 is a diagram illustrating a process for generating position data (125) and expression data (127) of an object (120) according to one embodiment of the present disclosure. In one embodiment, the processor (210) may receive image data (138) including an image set in which an image of a target sporting event is divided into frames from a camera group (130), and may generate frame-by-frame position data (125) of an object (120) within a game space (110) of the target sporting event based on the image set. For convenience of explanation, it is assumed that the object (120) of FIG. 3 is a player of the target sporting event. In addition, it will be readily understood that position data (125) of a plurality of players in the present disclosure may be generated as described in FIG. 3.
[0082] In one embodiment, the processor (210) may determine frame-by-frame three-dimensional coordinates of an object (120) from a plurality of images constituting an image set, and generate frame-by-frame position data (125) including the determined three-dimensional coordinates and the corresponding time. Here, the three-dimensional coordinates may be composed of X (x-axis), Y (y-axis), and Z (z-axis) values in pixel units.
[0083] For example, the processor (210) may select one or more two-dimensional images corresponding to a specific frame (time) among a plurality of two-dimensional images in an image set, and determine two-dimensional coordinates of an object (120) in each of the selected one or more two-dimensional images.
[0084] For example, the two-dimensional coordinates of the object (120) may be one or more two-dimensional coordinates forming an outline of the object (120) within a two-dimensional image, or may be a center point of one or more such two-dimensional coordinates.
[0085] For example, the two-dimensional coordinates of the object (120) may be one of one or more two-dimensional coordinates indicating a specific part of the object (120) or a center point obtained by averaging one or more two-dimensional coordinates. As a specific example, if the target sport is a ball game, the specific part of the object (120) may be a part where the ball and the object (120) come into direct or indirect contact. That is, if the target sport is table tennis, the specific part of the object (120) may be a hand holding a racket. Alternatively, if the target sport is a non-ball game, the specific part of the object (120) may be a part that the object (120) mainly uses during the target sport. For example, if the target sport is boxing, the specific part of the object (120) may be a hand that the object (120) mainly uses during the target sport.
[0086] For example, the processor (210) can determine the three-dimensional coordinates of the object (120) corresponding to a specific frame (time) based on the two-dimensional coordinates of the object (120) in each of one or more images.
[0087] Meanwhile, the processor (210) can determine the three-dimensional coordinates of the object (120) from the image set using various computer vision algorithms. However, the present disclosure is not limited thereto.
[0088] In one embodiment, the processor (210) may analyze a facial expression of a player included in one or more two-dimensional images corresponding to a specific frame (time) among a plurality of two-dimensional images within an image set. The processor (210) may extract one or more features of the player's face included in one or more two-dimensional images using a facial expression recognition technology (e.g., MPEG-4 face model, Active Shape Model, Bayesian Shape Model), and may analyze the player's facial expression and emotion based on the extracted features to generate facial expression data (127).
[0089] FIG. 4 is a diagram illustrating a process for generating movement data (180) of an object (120) according to one embodiment of the present disclosure. In one embodiment, the processor (210) may generate movement data (180) regarding the movement of an object (120) based on frame-by-frame position data (125) of the object (120) within the game space (110) of the target sports game.
[0090] In one embodiment, the movement data (180) of the object (120) may include indicators indicating the movement of the object (120), such as the direction of movement, the speed of movement, the time of movement (e.g., t seconds), and the movement trajectory of the object (120). Here, each indicator included in the movement data (180) may be numerical data or categorical data.
[0091] For example, movement data (180) of an object (120) may include at least one of a movement direction, a movement speed, a movement time, or a movement trajectory regarding the movement of the object (120) from a first position (X, Y, Z) to a second position (X', Y', Z').
[0092] FIG. 5 is a diagram illustrating a process of generating movement data (180) of an object (120) according to one embodiment of the present disclosure. In one embodiment, the processor (210) may generate movement data (180) regarding the movement of a specific portion of the object (120) based on frame-by-frame position data (125) of a specific portion of the object (120) within the game space (110) of the target sports game.
[0093] In one embodiment, movement data (180) of an object (120) may include indicators indicating movement of a specific part of the object (120), such as a movement direction, movement speed, movement time (e.g., t seconds), and movement trajectory of a specific part of the object (120). Here, each indicator included in the movement data (180) may be numerical data or categorical data.
[0094] For example, movement data (180) of an object (120) may include at least one of a movement direction, a movement speed, a movement time, or a movement trajectory regarding a specific part of the object (120) moving from a first position (X, Y, Z) to a second position (X', Y', Z').
[0095] FIG. 6 is a diagram illustrating a process for generating location data (125) of an object (120) according to one embodiment of the present disclosure. In one embodiment, the processor (210) may receive image data (138) including an image set in which an image of a target sporting event is divided into frames from a camera group (130), and generate frame-by-frame location data (125) of an object (120) within a game space (110) of the target sporting event based on the image set. For convenience of explanation, the object (120) of FIG. 6 is assumed to be a ball of the target sporting event.
[0096] In one embodiment, the processor (210) may determine frame-by-frame three-dimensional coordinates of an object (120) from a plurality of images constituting an image set, and generate frame-by-frame position data (125) including the determined three-dimensional coordinates and the corresponding time. Here, the three-dimensional coordinates may be composed of X (x-axis), Y (y-axis), and Z (z-axis) values in pixel units.
[0097] For example, the processor (210) may select one or more two-dimensional images corresponding to a specific frame (time) among a plurality of two-dimensional images in an image set, and determine two-dimensional coordinates of an object (120) in each of the selected one or more two-dimensional images.
[0098] For example, the two-dimensional coordinates of the object (120) may be any one of one or more two-dimensional coordinates that constitute the outline of the object (120) within the two-dimensional image, or may be the midpoint of one or more such two-dimensional coordinates.
[0099] For example, the processor (210) can determine the three-dimensional coordinates of the object (120) corresponding to a specific frame (time) based on the two-dimensional coordinates of the object (120) in each of one or more images.
[0100] Meanwhile, the processor (210) can determine the three-dimensional coordinates of the object (120) from the image set using various computer vision algorithms. However, the present disclosure is not limited to using computer vision algorithms for such determination.
[0101] FIG. 7 is a diagram illustrating a process of generating movement data (180) of an object (120) according to one embodiment of the present disclosure. In one embodiment, the processor (210) may generate movement data (180) regarding the movement of an object (120) based on frame-by-frame position data (125) of the object (120) within the game space (110) of the target sports game.
[0102] In one embodiment, the movement data (180) of the object (120) may include indicators indicating the movement of the object (120), such as the direction of movement, the speed of movement, the time of movement (e.g., t seconds), the trajectory of movement, the direction of spin, the number of spins, etc. Here, each indicator included in the movement data (180) may be numerical data or categorical data.
[0103] For example, movement data (180) of an object (120) may include at least one of a movement direction, a movement speed, a movement time, a movement trajectory, a spin direction, or a spin count regarding the movement of the object (120) from a first position (X, Y, Z) to a second position (X', Y', Z').
[0104] Meanwhile, the processor (210) may obtain movement data (180) of an object (120) of a target sport depending on the type of the target sport. For example, if the target sport is a ball game, the processor (210) may generate frame-by-frame position data (125) of a player (120) and frame-by-frame position data (125) of a ball (120) of the target sport, respectively, and may generate movement data (180) of the player and movement data (180) of the ball based on the frame-by-frame position data (125) of the player (120) and frame-by-frame position data (125) of the ball (120), respectively. Alternatively, if the target sport is a non-ball game, the processor (210) may generate frame-by-frame position data (125) of a player of the target sport, and generate movement data (180) of the player based on the frame-by-frame position data (125) of the player.
[0105] FIG. 8 is a diagram for explaining an example of first analysis data generated by a device according to an embodiment of the present disclosure based on first data and second data. As described above, the processor (210) can generate position data (125) and movement data (180) of a player and / or a ball. The processor (210) can generate first analysis data (172) based on the generated position data (125) and movement data (180), the acquired facial expression data (127), voice data (142), and sensor data (152), and the first data (154), and the second data (162). For convenience of explanation, it is explained that the target sport is a team game, a ball game, and that objects (120a), (120b), and (120c) are players belonging to the first team, and that objects (120d), (120e), and (120f) are players belonging to the second team.
[0106] According to one embodiment, the processor (210) may analyze various progresses of the target sports game based on the first data (154) and the second data (162) to generate the first analysis data (172). The processor (210) may analyze the following various aspects of the target sports game using the first data (154) and the second data (162), which include location data (125), movement data (180), facial expression data (127), image data (138), voice data (142), and sensor data (152).
[0107] For example, the processor (210) may analyze the tactics and strategies of the first and second teams based at least on the location data (125), movement data (180), and second data (162). For example, the processor (210) may analyze whether the first and second teams have chosen an offensive or defensive tactic. For example, the processor (210) may analyze the compatibility between the tactics of the first and second teams.
[0108] For example, the processor (210) can analyze that an event (e.g., a goal, offside, foul, corner kick, free kick) has occurred in a target sporting event (e.g., soccer) based at least on location data (125), voice data (142), movement data (180), and reactions of spectators (122).
[0109] For example, the processor (210) may analyze the skills (e.g., dribbling, passing, shooting, tackling) of players (120a to 120f) of both teams based at least on location data (125), facial expression data (127), movement data (180), and reactions of spectators (122). For example, the processor (210) may analyze the accuracy, difficulty, success, and whether a chance was created of the skill.
[0110] For example, the processor (210) may generate various statistical data regarding the progress of a target sports game based at least on the location data (125) and movement data (180). For example, various statistical data regarding possession, number of shots, number of effective shots, attack direction, pass success rate, number of fouls, number of turnovers, etc. may be generated.
[0111] For example, the processor (210) may analyze the overall flow of the target sports game based on the first data (154). For example, the processor may analyze which team, the first or second team, has the upper hand and whether the momentum of the target sports game is shifting.
[0112] For example, the processor (210) can analyze the atmosphere, momentum, and reactions of the audience (122) of the game. For example, the processor (210) can analyze the reactions of the audience based on the facial expressions and voice data (142) of the audience (122).
[0113] Examples of first analysis data (172) that the processor (210) can generate based on the first data (154) and the second data (162) are not limited to those mentioned above. The processor (210) can generate first analysis data (172) by analyzing various aspects of the progress of the target sports game.
[0114] FIG. 9 is a diagram illustrating a process of obtaining a commentary script using a learning model (900) trained to generate a script of voice commentary in a sports game according to one embodiment of the present disclosure.
[0115] In one embodiment, the learning model (900) may be a model that is trained to generate voice commentary corresponding to a sports game using learning data (910) for other sports games during the learning process, and generates output data (930), which is a commentary script corresponding to a broadcast screen of the target sports game, from input data (920) for the target sports game during the inference process.
[0116] For example, the processor (210) may input at least one of the first data (154) or the second data (162) from a target sports game as input data (910) to the learning model (900), and obtain a commentary script for the target sports game as output data (930) of the learning model (900).
[0117] In one embodiment, the learning model (900) may be a model trained using commentary scripts from one or more other sporting events as training data (910). The one or more other sporting events may be sporting events distinct from the target sporting event.
[0118] FIG. 10 is a flowchart of a method for generating voice commentary for a sports game according to one embodiment of the present disclosure.
[0119] In step S1010, the device (100) can obtain image data (138) of a sports game from a camera group (130) including one or more cameras. The device (100) can obtain voice data (142) from a microphone (140) and sensor data (152) from a sensor (150).
[0120] In step S1020, the device (100) may generate first data (154) related to the game situation of a sports game based on the image data (138). The first data (154) may include at least one of player and ball position data (125), player facial expression data (127), movement data (180), voice data (142), or sensor data (152).
[0121] In step S1030, the device (100) may obtain second data (162) including statistical data regarding an object (120) of a sports game. The device (100) may obtain the second data (162) from an external device (160). The second data (162) may be statistical data that statistically organizes data regarding one or more sports games.
[0122] In step S1040, the device (100) may generate first analysis data (172) indicating analysis results for a sports game based on the first data (154) and the second data (162). The first analysis data (172) may include various analyses of the progress of the target sports game.
[0123] In step S1050, the device (100) may generate relay screen data (174) to be transmitted to the user's terminal, including at least a portion of the image data (138), based on the first analysis data (172). In step S1060, the device (100) may generate voice commentary data (176) corresponding to the relay screen based on the first analysis data (172). For example, the device (100) may generate a commentary script corresponding to the relay screen based on the first analysis data (172), and generate the voice commentary data (176) based on the commentary script. According to one embodiment, the device (100) may generate the voice commentary data (176) further based on user preference data.
[0124] FIG. 11 is a flowchart of a method for generating voice commentary based on viewing history information according to one embodiment of the present disclosure.
[0125] In step S1110, the device may collect viewing history information from the user's terminal. The viewing history information may include information about one or more sports games played on the user's terminal.
[0126] In step S1120, the device may generate preference data by analyzing the characteristics of the user's preferred voice commentary based on viewing history information. The preference data may include at least one of information regarding the user's preferred speaking style (or tone) of the commentary or the user's preference for a specific team or player. In step S1130, the device may generate a voice commentary based on the preference data.
[0127] Although the steps of the method or algorithm according to the present disclosure are described in a sequential order in the flowcharts illustrated in FIGS. 10 and 11, the steps may be performed in any order that can be arbitrarily combined according to the present disclosure, in addition to being performed sequentially. The description according to this flowchart does not exclude changes or modifications to the method or algorithm, and does not imply that any step is essential or desirable. In one embodiment, at least some of the steps may be performed in parallel, iteratively, or heuristically. In one embodiment, at least some of the steps may be omitted, or other steps may be added.
[0128] Various embodiments of the present disclosure may be implemented as software on a machine-readable storage medium. The software may be software for implementing various embodiments of the present disclosure. The software may be inferred from various embodiments of the present disclosure by programmers skilled in the art to which the present disclosure pertains. For example, the software may be a program including machine-readable instructions (e.g., code or code segments). The device may be a device capable of operating according to instructions called from a storage medium, such as a computer. In one embodiment, the device may be a device (100) according to embodiments of the present disclosure. In one embodiment, the processor of the device may execute the called instructions, causing components of the device to perform functions corresponding to the instructions. In one embodiment, the processor may be a processor (210) according to embodiments of the present disclosure. The storage medium may refer to any type of recording medium that stores data and can be read by the device. The storage medium may include, for example, ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage, etc. In one embodiment, the storage medium may be memory (220). In one embodiment, the storage medium may be implemented in a distributed form, such as in a network-connected computer system. The software may be distributed and stored and executed in a computer system, etc. The storage medium may be a non-transitory storage medium. A non-transitory storage medium means a tangible medium regardless of whether data is stored semi-permanently or temporarily, and does not include a signal that is propagated transitorily.
[0129] While the technical concept of the present disclosure has been described through various embodiments, it should be understood that the technical concept of the present disclosure encompasses various substitutions, modifications, and variations that can be made within the scope of those skilled in the art to which the present disclosure pertains. Furthermore, it should be understood that such substitutions, modifications, and variations are encompassed within the scope of the appended claims.
Claims
A method performed in a device comprising one or more processors and one or more memories storing instructions to be executed by the one or more processors, One or more of the above processors, A step of acquiring image data capturing a sports event from one or more cameras; A step of generating first data related to the game situation of the sports game based on the image data; A step of acquiring second data including statistical data regarding an object of the above sports game; A step of generating first analysis data indicating analysis results for the sports game based on the first data and the second data; A step of determining a relay screen to be transmitted to the user's terminal at least part of the image data based on the first analysis data; and A method comprising the step of generating a voice commentary corresponding to the relay screen based on the first analysis data. In the first paragraph, The one or more cameras include a first camera and a second camera that photograph objects of the sports event from different angles, The step of generating the above first data is: A step of determining a first position of the object in the shooting screen of the first camera based on the image data; A step of determining a second position of the object in the shooting screen of the second camera based on the image data; A step of determining the three-dimensional position coordinates of the object based on the first position and the second position; A method comprising a step of generating the first data based on a change in the three-dimensional position coordinates of the object over time. In the first paragraph, The step of generating the above first data is: A step of receiving voice data generated inside a game space where a sports game is being played from one or more microphones; A step of receiving sensor data including information on temperature and humidity of the game space from one or more sensors; and A method further comprising a step of generating the first data based on the voice data and the sensor data. In the first paragraph, A method wherein the first data includes at least one of a player's position and movement, a ball's position and movement, a player's facial expression, or an audience reaction in the sports game. In the first paragraph, A method wherein the second data includes at least one of information regarding the tactics of each team, the characteristics of each player, the referee's tendencies, the opponent's record, interviews of each team before the sports game, or analysis regarding the sports game. In the first paragraph, The step of generating the above first analysis data is: A method comprising a step of analyzing at least one of the tactics of each team in the sports game, the movements of each player in the sports game, an event occurring in the sports game, or statistics related to the sports game, based on the first data and the second data. In the first paragraph, The step of determining the above relay screen is: A step of acquiring first image data at a first point in time from the one or more cameras; A step of acquiring second image data from the one or more cameras at a second time point after the first time point; A step of generating the first analysis data based at least on the first image data and the second image data; A method further comprising a step of determining the first image data as a relay screen based on the first analysis data. In the first paragraph, The steps for generating the above voice commentary are: A method comprising the step of inputting first analysis data for one or more sports matches into a learning model learned based on second analysis data for the sports matches and a commentary script corresponding to the second analysis data, thereby generating a commentary script for the sports match. In paragraph 8, The steps for generating the above voice commentary are: A method further comprising the step of generating a voice commentary based on the commentary script using a text-to-speech (Text To Speech) voice synthesis technology. In the first paragraph, The steps for generating the above voice commentary are: A method comprising the step of generating a voice corresponding to the analysis result for the sports game from among at least one pre-stored voice as the voice commentary. In the first paragraph, The steps for generating the above voice commentary are: A step of collecting viewing history information including information about one or more sports games played on the user's terminal and voice commentary for each of the one or more sports games; A step of generating preference data by analyzing the characteristics of the voice commentary preferred by the user based on the viewing history information; and A method comprising the step of generating the voice commentary based on the preference data. In Article 11, A method wherein the preference data includes at least one of whether the user prefers commentary on a specific player, whether the user prefers commentary on a specific team, whether the user prefers commentary from a neutral standpoint, or whether the user prefers voice commentary with an average volume greater than a predetermined value. communication circuit; one or more processors; and comprising one or more memories storing instructions executed by the one or more processors; One or more of the above processors, Acquire image data of a sports event from one or more cameras, Based on the above image data, first data related to the game situation of the sports game is generated, Obtaining second data including statistical data regarding objects of the above sports game, Based on the first data and the second data, first analysis data indicating analysis results for the sports game is generated, Based on the first analysis data, at least a portion of the image data is determined as a relay screen to be transmitted to the user's terminal, and A device that generates a voice commentary corresponding to the relay screen based on the first analysis data. In Article 13, The one or more cameras include a first camera and a second camera that photograph objects of the sports event from different angles, One or more of the above processors, Based on the image data, the first position of the object is determined in the shooting screen of the first camera, Based on the image data, the second position of the object is determined in the shooting screen of the second camera, Based on the first position and the second position, the three-dimensional position coordinates of the object are determined, and A device that generates the first data based on changes in the three-dimensional position coordinates of the object over time. In Article 13, One or more of the above processors, Acquire first image data from one or more cameras at a first point in time, Acquire second image data from the one or more cameras at a second time point after the first time point, Generating the first analysis data at least based on the first image data and the second image data, and A device that determines the first image data as a relay screen based on the first analysis data. In Article 13, One or more of the above processors, A device for generating a commentary script for a sports game by inputting the first analysis data for the sports game into a learning model learned based on second analysis data of one or more sports games and a commentary script corresponding to the second analysis data. In Article 16, One or more of the above processors, A device that generates voice commentary based on the commentary script using voice synthesis technology. In Article 13, One or more of the above processors, A device that generates a voice corresponding to the analysis result for the sports game among at least one pre-stored voice as the voice commentary. In Article 13, One or more of the above processors, Collecting viewing history information including information about one or more sports games played on the user's terminal and audio commentary for each of the one or more sports games, Based on the above viewing history information, preference data is generated by analyzing the characteristics of the voice commentary preferred by the user, and A device for generating the voice commentary based on the above preference data. In a non-transitory computer-readable recording medium having recorded thereon instructions to be executed by one or more processors, The above instructions, when executing the above instructions, cause the one or more processors to: Acquire image data of a sports event from one or more cameras, Based on the above image data, first data related to the game situation of the sports game is generated, Obtaining second data including statistical data regarding objects of the above sports game, Based on the first data and the second data, first analysis data indicating analysis results for the sports game is generated, Based on the first analysis data, at least a portion of the image data is determined as a relay screen to be transmitted to the user's terminal, and A non-transitory computer-readable recording medium that generates a voice commentary corresponding to the relay screen based on the first analysis data.
Citation Information
Patent Citations
Method and system for servicing relay broadcast for athletics
KR1020120004631A
Well screen welding apparauts
KR1020220122256A
Apparatus for midair extinguishing using firefighting helicopter
KR1020230153728A
Manufacturing method of bread using sugar-resistant yeast
KR1020250109296A
A system and method to analyze and improve sports performance using monitoring devices
US20190347956A1