Program
By using a trained model to automatically generate live commentary for game videos, the program addresses the high user workload in conventional techniques, resulting in a more efficient and streamlined process for broadcasting live commentary.
Patent Information
- Application Number
- JP2023187262
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-15
- Estimated Expiration
- 2043-10-31
AI Technical Summary
Conventional techniques require significant user workload for editing live commentary into game videos, which is time-consuming and labor-intensive.
A program utilizing a trained model to generate and output live commentary data based on image data from game actions, reducing the need for manual editing.
This approach significantly reduces the user workload in broadcasting live commentary, enabling faster and more efficient content delivery.
Smart Images

Figure 2025075825000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a program and a system. [Background technology]
[0002] 2. Description of the Related Art Conventionally, there has been known a technique for adding commentary to the play of a game or the like and distributing the resulting video or the like.
[0003] For example, a distribution server first distributes a live commentary of a game. Next, the system manages information related to the game. Furthermore, the system detects the viewing status of the viewers watching the live commentary of the game. Next, viewer benefits are given to the viewers according to the viewing status and game play information. In this way, a technique is known that motivates viewers to watch the live commentary of a game and gets viewers interested in playing the game (for example, Patent Document 1, etc.).
[0004] Furthermore, in online games in which many people participate, game situations occur simultaneously due to multiple players. Audio game commentary corresponding to the game situations also occurs simultaneously. This results in continuous playback of multiple game commentaries, which impedes the player's progress in the game. Therefore, when the number of commentary playbacks exceeds an upper limit, audio data to be played within a predetermined time is controlled. In this way, a technology for providing a suitable game commentary to a player is known (for example, Patent Document 2, etc.). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2023-048505 [Patent Document 2] JP 2009-125108 A Summary of the Invention [Problem to be solved by the invention]
[0006] In conventional technology, a human commentator provides commentary on the play of a game or the like. For example, commentary is added by editing a video of the play, adding commentary in the form of voice or text (sometimes using characters, etc.). Thus, in conventional technology, editing is required to add commentary in order to distribute the commentary, which increases the workload of the user.
[0007] The present invention aims to reduce the workload of a user when delivering a live report of game play or the like. [Means for solving the problem]
[0008] In order to solve the above-mentioned problem, the present invention provides a program that causes a computer to function as: unknown data input means for inputting unknown data including unknown image data in which the correct answer for a user's action is unknown; generation means for generating output data relating to a commentary on the action based on the unknown data using a trained model trained using learning data including image data showing the action and commentary data relating to a commentator's commentary on the action; and output means for performing output based on the output data. Effect of the Invention
[0009] According to the present invention, the workload of a user can be reduced when delivering a live report of game play or the like. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a system configuration according to an embodiment of the present invention. [Diagram 2] FIG. 13 is a diagram illustrating an example of pre-processing. [Diagram 3] FIG. 11 illustrates an example of an execution process. [Figure 4] FIG. 1 is a diagram showing an example of the overall process of AI learning and execution. [Diagram 5] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 6]FIG. 1 is a network diagram showing an example of an AI configuration. [Figure 7] FIG. 13 is a diagram illustrating an example of a learning process for a game. [Figure 8] FIG. 11 is a diagram illustrating an example of an execution process for a game. [Figure 9] FIG. 13 is a diagram illustrating an example of the overall process. [Figure 10] FIG. 2 is a diagram illustrating an example of a functional configuration. [Figure 11] FIG. 13 is a diagram showing an example of a screen for performing a selection operation. [Figure 12] FIG. 13 is a diagram showing a configuration example in which an auxiliary device is used. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, an embodiment will be described with reference to the drawings.
[0012] [System configuration example] Fig. 1 is a diagram showing an example of a system configuration according to the present embodiment. For example, as shown in Fig. 1, a system 1 mainly includes user terminals 20A, 20B, and 20C (hereinafter, these may be collectively referred to as "user terminal 20") and a server 11.
[0013] Hereinafter, the person who manages the server 11 will be referred to as the "administrator 5." In addition, the people who operate the user terminals 20A, 20B, and 20C will be referred to as the "user 4A," the "user 4B," and the "user 4C" (hereinafter, these may be collectively referred to as the "user 4").
[0014] The administrator 5 is a person who operates the information processing service provided by the system 1. On the other hand, the user 4 is a person who uses the information processing service provided by the system 1. The administrator 5 and the user 4 differ in which information processing device they operate: the server 11, which is an example of a management device, or the user terminal 20. Hereinafter, the user 4 will be a player of the game, and the administrator 5 will manage the game and the server 11.
[0015] Although the example shown in FIG. 1 has three user terminals 20 and one server 11, the number of servers 11, the number of user terminals 20, the number of administrators 5, and the number of users 4 are not important.
[0016] The server 11 and the user terminal 20 are connected to each other so as to be able to communicate with each other via a communication network 2. For example, the communication network 2 is the Internet, a mobile communication system (for example, a public line based on 4G (4th Generation, fourth generation mobile communication standard) or 5G (5th Generation, fifth generation mobile communication standard)), a wireless network such as Wi-Fi (registered trademark) (Wireless Fidelity), or a combination of these.
[0017] The user terminal 20 downloads a program for playing a game (hereinafter referred to as a "game program") from the server 11, or accesses the server 11 to provide a game service. Note that communication with the server 11 is not required to play a game. In other words, the user terminal 20 may download a program or install it from a medium to create an environment for playing a game.
[0018] [Examples of AI (Artificial Intelligence) learning and execution] Hereinafter, AI learns through "pre-processing." AI in the learning stage, i.e., "pre-processing," will be referred to as the "learning model A1." Then, once learning has progressed to a certain extent, the learning model A1 becomes the "trained model A2." Hereinafter, the execution stage in which output processing is executed using the trained model A2 will be referred to as the "execution process."
[0019] The "pre-processing" is performed before the "execution processing". However, the "pre-processing", that is, the trained model A2 may continue to train before the "execution processing".
[0020] [Pre-processing example] 2 is a diagram showing an example of pre-processing. For example, the pre-processing is performed by the server 11.
[0021] The learning model A1 performs learning by inputting learning data D1. That is, the learning model A1 performs so-called "supervised" learning.
[0022] In the pre-processing, image data D10 is input. For example, the image data D10 is generated by capturing an output screen when a user 4 plays a game on a user terminal 20. Below, an example will be described in which playing a game is an "action."
[0023] However, the image data D10 is not limited to being generated by capturing the output screen, and may be generated by another information processing device, or may be generated by capturing an image of the output screen during play on the user terminal 20 with a photographing device such as a camera. In the learning data D1, the correct answer data D11 is associated with the image data D10. Hereinafter, image data for which the "correct answer" is known will simply be referred to as "image data D10."
[0024] The learning model A1 learns a correspondence relationship in which, based on the input of learning data D1, image data D10 is input and answer data D11 is output.
[0025] The learning data D1, the image data D10, and the answer data D11 will be described in detail later.
[0026] Furthermore, it is preferable that the learning model A1 learns using big data D4. For example, the big data D4 is data on the Internet. However, the big data D4 may be data input by an administrator 5 or the like.
[0027] [Example of execution process] 3 is a diagram showing an example of the execution process. For example, the execution process is performed by the user terminal 20, or by the user terminal 20 and the server 11 working together.
[0028] The trained model A2 is in a state where the trained model A1 has been trained through pre-processing. That is, when the pre-processing shown in FIG. 2 is executed, the trained model A2 is generated.
[0029] When unknown data D2 is input, the trained model A2 generates output data D3 for the unknown data D2.
[0030] The unknown data D2 is unknown image data (hereinafter referred to as "unknown image data D20"), that is, data for which the "correct answer" for the image data is unknown at the time of input.
[0031] The output data D3 is data that is composed of commentaries D30. Note that the trained model A2 may generate the output data D3 so that the output data D3 is composed of a plurality of commentaries D30.
[0032] The output data D3 and the commentary D30 will be described in detail later.
[0033] Fig. 4 is a diagram showing an example of the overall process of AI learning and execution. The relationship between the pre-processing shown in Fig. 2 and the execution process shown in Fig. 3 is as shown in Fig. 4.
[0034] Note that the pre-processing and the execution processing do not have to be executed in the consecutive order as illustrated in the figure. Therefore, it is not essential that the period in which preparation is performed by the pre-processing and the period in which the execution processing is performed thereafter are consecutive. Therefore, the execution processing may be executed after a time has elapsed from the pre-processing, once the trained model A2 has been created. Also, once the trained model A2 has been generated, the execution processing may be executed by repurposing the trained model A2.
[0035] The learning data D1 and unknown data D2 are different between the learning process and the execution process. Also, the AI is a learning model A1 in the learning stage, but when learning progresses to a certain extent, it becomes a trained model A2. In this way, the trained model A2 trained using big data D4 as training data is a so-called "generative AI."
[0036] Image data D10 included in the learning data D1 and unknown image data D20 included in the unknown data D2 are the same data type. That is, the image data D10 and the unknown image data D20 are moving images generated by capturing an output screen of a game played by the user 4. Note that a dedicated device may be used for capturing.
[0037] In the learning data D1, the "correct answer" is known, whereas in the unknown data D2, the "correct answer" is unknown. Specifically, the learning data D1 includes the correct answer data D11, whereas the unknown data D2 does not include the correct answer data D11. Therefore, in the learning data D1, the relationship between the image data D10 and the correct answer data D11 is known.
[0038] On the other hand, the unknown data D2 does not include the correct answer data D11, and the “correct answer” for the unknown data D2 is unknown. The trained model A2 generates output data D3 for the unknown data D2 based on the correlation between the training data D1 trained in pre-processing and the correct answer data D11.
[0039] The execution process may be partly a process using a table, etc. In this way, in a configuration using a table, that is, a so-called rule-based configuration, the pre-processing is a process of preparing to input a table (also called a look-up table (LUT) or a mathematical expression, etc.).
[0040] [Example of hardware configuration of information processing device] 5 is a hardware configuration diagram of an information processing device. The information processing device is a server 11, a user terminal 20, etc. Hereinafter, the information processing device is assumed to have the same hardware configuration as the server 11. For example, the information processing device is a workstation or a general-purpose computer such as a personal computer. However, each information processing device may have a different hardware configuration.
[0041] The server 11 mainly includes a processor 111, a memory 112, a storage 113, an input / output interface 114, and a communication interface 115. In addition, each component of the server 11 is connected to a communication bus .
[0042] The processor 111 executes a series of instructions contained in a server program 11P stored in the memory 112 or the storage 113, thereby achieving processing and control.
[0043] The processor 111 is, for example, an arithmetic device and a control device such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an MPU (Micro Processing Unit), an FPGA (Field-Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a combination thereof.
[0044] The memory 112 is a main storage device that stores the server program 11P and data, etc. For example, the server program 11P is loaded from the storage 113. The data includes data input to the server 11 and data generated by the processor 111. For example, the memory 112 is a RAM (Random Access Memory) or other volatile memory.
[0045] The storage 113 is an auxiliary storage device that stores the server program 11P and data, etc. The storage 113 is, for example, a ROM (Read-Only Memory), a hard disk drive, a flash memory, or other non-volatile storage device. The storage 113 may also be a removable storage device such as a memory card. As another example, the storage 113 may be an external storage device. With this configuration, for example, in a situation where multiple user terminals 20 are used, such as an amusement facility, it becomes possible to collectively update the server program 11P or data.
[0046] The input / output interface 114 is an interface that connects external devices, such as a monitor, an input device (for example, a keyboard or a pointing device), an external storage device, a speaker, a camera, a microphone, and a sensor, to the server 11.
[0047] The processor 111 also communicates with external devices via an input / output interface 114. The input / output interface 114 is, for example, a Universal Serial Bus (USB), a Digital Visual Interface (DVI), a High-Definition Multimedia Interface (HDMI (registered trademark)), wireless, or other terminal.
[0048] The communication interface 115 communicates with other devices (e.g., the user terminal 20, etc.) connected to the communication network 2. For example, the communication interface 115 is a wired communication interface such as a Local Area Network (LAN), or a wireless communication interface such as Wi-Fi (registered trademark) (Wireless Fidelity), Bluetooth (registered trademark), or Near Field Communication (NFC).
[0049] However, the information processing device is not limited to the above hardware configuration. For example, the user terminal 20 may further include a sensor such as a camera. Various data acquired by the user terminal 20 using the sensor may be transmitted to the server 11.
[0050] [Examples of learning model and trained model configuration] 6 is a network diagram showing an example of the configuration of an AI. The learning model A1 and the trained model A2 are, for example, AIs having a configuration shown in the following network.
[0051] Hereinafter, the learning model A1 and the trained model A2 will be described as being implemented on the server 11, i.e., on the cloud. However, a part or all of the learning model A1 and the trained model A2 may be implemented on a user terminal 20, etc. For example, the AI may have a network structure having an input layer L1, a hidden layer L2, and an output layer L3.
[0052] Specifically, the AI has a network structure having a CNN (Convolution Neural Network) as shown in the figure.
[0053] The input layer L1 is a layer that receives the input image DIN.
[0054] The hidden layer L2 is a layer that performs processing such as convolution, pooling, normalization, or a combination of these on the input image DIN input from the input layer L1.
[0055] The output layer L3 is a layer that outputs the result of processing in the hidden layer L2 as an output image DOUT. For example, the output layer L3 is composed of a fully connected layer, etc.
[0056] Convolution is a process of generating a feature map by performing a filter process on an image or a feature map generated by performing a predetermined process on an image based on, for example, a filter, a mask, or a kernel (hereinafter simply referred to as a "filter").
[0057] Specifically, a filter is data used to perform a calculation in which a filter coefficient (sometimes called a "weight" or "parameter") is multiplied by a pixel value of an image or a feature map. Note that the filter coefficient is a value determined by learning, settings, or the like.
[0058] The convolution process is a process of multiplying the pixel values of each pixel constituting an image or feature map by a filter coefficient, and generating a feature map whose components are the calculation results.
[0059] Thus, the convolution process can extract features of the image or feature map, such as edge components or statistical results of the neighborhood of a pixel of interest.
[0060] Furthermore, when convolution processing is performed, similar features can be extracted even if the subject shown in the target image or feature map is shifted up and down, left and right, diagonally, rotated, or a combination of these.
[0061] Pooling is a process that performs processing such as calculating the average, extracting the minimum value, or extracting the maximum value for a target range, extracting features, and generating a feature map. That is, pooling can be max pooling, avg pooling, or the like.
[0062] Note that the convolution and pooling may include pre-processing such as zero padding.
[0063] By using the above-described convolution, pooling, or a combination of these, it is possible to obtain the so-called data amount reduction effect, composition property, translation invariance, and the like.
[0064] Normalization is, for example, a process of aligning variance and average value. Note that normalization includes cases where normalization is performed locally. Normalization means that data becomes a value within a predetermined range. This makes it easier to handle the data in subsequent processes.
[0065] Fully connected is the process of reducing data such as feature maps to output.
[0066] For example, the output is in the form of a binary output, such as "YES" or "NO". In such an output format, full connection is a process of connecting nodes based on the features extracted in the hidden layer L2 so that one of two conclusions is reached.
[0067] On the other hand, when there are three or more types of output, the full connection is a process that performs a so-called softmax function, etc. In this way, classification (including the case where an output showing a probability is performed) can be performed by a maximum likelihood estimation method, etc., using the full connection. Note that the AI network configuration is not limited to that described above. In other words, the AI may be realized by other machine learning methods.
[0068] For example, the AI may be configured to perform preprocessing such as dimensionality reduction (for example, a process that converts a relationship of three or more dimensions into a relationship that can be obtained by simple calculations of three dimensions or less) using "unsupervised" machine learning. It is desirable to process the relationship between input and output using simple calculations such as linear expressions. This type of calculation can reduce calculation costs.
[0069] In addition, the AI may undergo processing to reduce overfitting, such as dropout. In addition, preprocessing such as dimensionality reduction and normalization may be performed.
[0070] The AI may have a network structure such as CNN. In addition, for example, the network structure may have a configuration such as LLM (Large Language Model), RNN (Recurrent Neural Network), or LSTM (Long Short-Term Memory). In other words, the AI may have a network structure other than deep learning.
[0071] The AI may also have hyperparameters. In other words, the AI may be configured such that some of the settings are performed by a user, etc. Furthermore, the AI may specify features to be learned, or the user may set some or all of the features to be learned.
[0072] Furthermore, the learning model A1 and the trained model A2 may use other machine learning. For example, the learning model A1 and the trained model A2 may be preprocessed by normalization using an unsupervised model. Furthermore, the learning may be reinforcement learning (a learning method in which an AI is made to make a choice and an evaluation (reward) is given for the choice, resulting in a larger evaluation), etc.
[0073] In the learning, data expansion or the like may be performed. That is, in order to increase the amount of learning data used in the learning of the learning model A1, a single piece of experimental data or the like may be expanded and pre-processed to generate multiple pieces of learning data. In this way, if the amount of learning data can be increased, the learning of the learning model A1 can be further advanced.
[0074] Furthermore, the learning model A1 and the trained model A2 may be configured to perform transfer learning, fine tuning, or the like. That is, since the user terminal 20 often has a different execution environment for each device, the settings may be different for each device according to the execution environment. For example, the basic configuration of the AI is trained on another information processing device. After that, each information processing device may be further trained or configured to be further optimized for each execution environment.
[0075] In addition, when the learning data D1 and the unknown data D2 include data other than images, such as text, sensor detection results, or audio, the network configuration may be other than the above.
[0076] [Examples of image data, examples of correct answer data, examples of live commentary, and examples of training data] 7 is a diagram showing an example of a learning process for a game. For example, while a user 4 is playing a game on the user terminal 20, a screen on which the game is being played on the user terminal 20 (hereinafter simply referred to as "output screen 6") is captured, and image data D10 is generated. That is, the image data D10 is a video or the like showing the output screen 6 in multiple frames.
[0077] For example, a commentator 8 watches the play of the image data D10 and inputs a voice (hereinafter referred to as "commentary voice 9") commenting on the play in the image data D10. Note that the commentary voice 9 is not limited to voice, and may be in a format such as text. The commentary voice 9 is then input via a microphone or the like and becomes voice data, and this voice data becomes the correct answer data D11. However, the commentary voice 9 may also be converted into text format by voice recognition processing and become the correct answer data D11.
[0078] It is desirable for the correct answer data D11 to be in the form of the commentary audio 9, i.e., to include audio data. Even with the same wording, audio data may be able to express pitch (frequency), volume, voice speed, intonation, or atmosphere better than text format. In other words, when the commentator 8 is excited, the volume will often increase or the voice speed will increase, i.e., the commentator will speak quickly. Therefore, with audio data, the situation can be better inferred from the audio, making it easier to determine the excitement of the play.
[0079] The commentator 8 may be the same person as the user 4 who played the game, or may be a different person. Also, there may be multiple commentators 8.
[0080] Other big data D4 and the like may also be learned. For example, the Internet may provide information such as reputation, game strategy information, or highlights. If such information is in the big data D4, more useful information is learned, so that the learning model A1 can learn more about the content of the game, etc.
[0081] The learning data D1 and the unknown data D2 may include auxiliary data, such as game specifications, rules, log data, or statistical data.
[0082] The log data is generated in accordance with the progress of the game when the game is played on the user terminal 20. Note that the log data may be generated based on an operation by the user 4 (for example, an operation to give an instruction to generate log data, etc.), or may be automatically generated in the background by the user terminal 20 in accordance with the progress of the game.
[0083] The log data is data indicating peripheral information such as, for example, the situation, the controller operation by the user 4, self-reported information of the user 4, information on the progress of the play, parameters not displayed to the user 4 (so-called "hidden parameters" may be included), and parameters of the character operated in the play. When such log data is input, even the peripheral information is learned and taken into account in the commentary D30, and an output that takes into account more information can be made.
[0084] The statistical data is, for example, the result of calculating the average, median, or variance for the entire game. For example, in a racing game, the times and the like of various players are the subject of statistical processing. With such statistical data, it is possible to compare the play of user 4 with others and determine whether the time and the like is an excellent result.
[0085] [Examples of unknown data, unknown image data, and output data] 8 is a diagram showing an example of execution processing for a game. The execution processing is executed when unknown data D2 is input to a trained model A2.
[0086] As with the image data D10 in the learning process, the unknown data D20 is generated by capturing the output screen 6 while the user 4 is playing a game on the user terminal 20. The unknown data D2 may include auxiliary data.
[0087] The output data D3 is generated to show a commentary D30 for the unknown image data D20. Specifically, the commentary D30 is generated in a format in which commentary audio is input to a video represented by the unknown image data D20 (hereinafter, the video with commentary audio input is referred to as "commentary video D31").
[0088] The commentary video D31 is generated so that audio explaining the content of the play etc. is played in accordance with the image showing the play indicated by the unknown image data D20.
[0089] The explanatory video D31 may have the same images as the unknown image data D20 or may have different images. For example, the explanatory video D31 may have a new image for explanation (hereinafter, referred to as an "explanation image").
[0090] The explanatory image is, for example, a mark added to a portion of the screen that is of interest, or the number of frames per second is increased compared to the playback speed of other scenes (so-called "slow motion"). Note that the explanatory image may be included in the explanatory video D31, or may be separate data.
[0091] Furthermore, the commentary D30 may be generated using a voice other than that input in the learning data D1. For example, the commentary D30 may be dubbed with a voice other than the commentary voice 9 (i.e., the voice of the commentator 8). Furthermore, the commentary D30 is not limited to adding a voice, and may include effects such as adding a character or adding subtitles.
[0092] The output data D3 may include a summary D32. For example, the summary D32 is data summarizing scenes of interest in the play indicated by the unknown data D2. Specifically, the summary D32 is composed of a heading, the time of the scene of interest, an excerpt image, a summary, and the like.
[0093] A heading is a short sentence or word that summarizes the entire play.
[0094] The times of interest are text that indicate periods of time during the play that are considered to be important.
[0095] An excerpt image is a still image or a video of a few seconds that captures an action considered to be important during gameplay.
[0096] The summary is a short text that describes the action depicted in the image excerpt.
[0097] When the summary D32 is generated separately from the commentary D30, the user 4 can quickly understand the contents by looking at the summary D32. Also, since the summary D32 indicates so-called "highlights," the user 4 can quickly determine which points to pay attention to. However, since the summary D32 may contain so-called "spoilers," it is preferable that the viewer can select whether or not to use the summary D32.
[0098] The point in time when the summary D32 is generated is determined, for example, around the point in time when the number of viewers increases, or when there is a large change in the scores or the like in the game content.
[0099] [Overall processing example] 9 is a diagram showing an example of the overall process. In the following example, the overall process is a sequence of pre-processing and execution processing. Specifically, the pre-processing is steps S01 and S02. Furthermore, the execution processing is steps S03 and S04. However, the overall process may include other processes.
[0100] In step S01, the server 11 inputs learning data D1 to the learning model A1. For example, in step S01, when a user 4 plays a game on a user terminal 20 and captures an output screen 6 on which the user 4 is playing, image data D10 is generated.
[0101] Next, the commentator 8 provides commentary on the image data D10, that is, the play by the user 4. For example, when the commentator 8 provides commentary using only audio, the audio data becomes the answer data D11.
[0102] After the image data D10 and the correct answer data D11 are generated in this manner, the image data D10 and the correct answer data D11 are associated with each other to generate the learning data D1. After generating the learning data D1 in this manner, the server 11 inputs the learning data D1 including the image data D10 and the correct answer data D11 to the learning model A1.
[0103] In step S02, the server 11 trains the learning model A1. Specifically, as shown in Fig. 7, the server 11 trains the learning model A1 using the learning data D1 in which the image data D10 and the correct answer data D11 are associated with each other, which is generated in step S01, to generate a trained model A2.
[0104] When a certain amount of learning is performed through steps S01 and S02 as described above and a trained model A2 is generated, the following steps S03 and S04 are executed.
[0105] In step S03, the server 11 inputs unknown data D2. For example, as shown in Fig. 8, when a user 4 plays a game on a user terminal 20 and captures an output screen 6 during the game, unknown image data D20 is generated. After that, the unknown data D2 is input to the trained model A2.
[0106] In step S04, when the unknown data D2 is input to the server 11, for example, as shown in FIG. 8, the server 11 uses the trained model A2 to generate output data D3 including a commentary D30 for the play indicated by the unknown data D2.
[0107] Next, in step S04, the server 11 performs output based on the output data D3. For example, the server 11 transmits the output data D3 to the user terminal 20 or the like so that a video with commentary D30 as shown in FIG. 8 is displayed on the user terminal 20 or the like.
[0108] However, the output based on the output data D3 may be uploaded to a video site, or displayed on an output device such as a public view, other than the user terminal 20. Also, the output may be almost real-time (there may be a certain delay due to processing), or may start when the output data D3 is saved and the user 4 performs an operation to play it later.
[0109] In the overall process described above, the server 11 executes all of the learning process and the execution process, but a part or all of the process may be executed by the user terminal 20. In other words, the learning model A1 and the trained model A2 may be implemented in an information processing device other than the server 11.
[0110] [Example of functional configuration] 10 is a diagram showing an example of a functional configuration. For example, system 1 is an example of an AI evaluation system including a learning device 31 and an execution device 32. Note that learning device 31 and execution device 32 may be the same information processing device or may be different information processing devices.
[0111] The learning device 31 includes an image data generating means 1F11, a live data input means 1F12, a learning data input means 1F13, and a learning means 1F14.
[0112] The execution device 32 comprises unknown data input means 1F21, output data generation means 1F22, and output means 1F23.
[0113] The image data generating means 1F11 performs an image data generating procedure for generating image data D10 showing an action by the user 4. For example, the image data generating means 1F11 is realized by the communication interface 115 or the like.
[0114] The live data input means 1F12 performs a live data input procedure for inputting live data such as the live audio 9. For example, the live data input means 1F12 is realized by the input / output interface 114 or the like.
[0115] The learning data input means 1F13 performs a learning data input procedure for inputting the learning data D1 including the image data D10 and the real-time data. For example, the learning data input means 1F13 is realized by the input / output interface 114 or the like.
[0116] The learning means 1F14 performs a learning procedure to train the learning model A1 using the learning data D1. Then, when the learning model A1 has trained to a certain extent, a trained model A2 is generated. For example, the learning means 1F14 is realized by the processor 111 or the like.
[0117] The unknown data input means 1F21 performs an unknown data input procedure for inputting the unknown data D2. For example, the unknown data input means 1F21 is realized by the communication interface 115 or the like.
[0118] When the unknown data D2 is input, the output data generation means 1F22 performs an output data generation procedure to generate output data D3 indicating a commentary D30 for the action indicated by the unknown data D2, based on the unknown data D2, using the trained model A2. For example, the output data generation means 1F22 is realized by the processor 111 or the like.
[0119] The output means 1F23 performs an output procedure for outputting the output data D3 to the user terminal 20, etc. For example, the output means 1F23 is realized by the communication interface 115, etc.
[0120] For example, on a video site, a video of a game being played is uploaded with commentary D30 added. In this way, adding commentary D30 requires a lot of editing work, such as adjusting the timing of the audio, which places a heavy workload on the user 4.
[0121] On the other hand, in the above configuration, when a video or the like is input, a commentary D30 is output by the AI. In this way, by using the AI, the workload of the user 4 can be reduced when delivering a commentary of sports, martial arts, games, or the like.
[0122] In addition, even if the user 4 is the one who performs the action and also serves as the commentator 8, the commentary video, etc. can be output by AI without editing by the user 4, so that the commentary video, etc. can be distributed at a speed close to real time in relation to the action.
[0123] Furthermore, if the commentator 8 is a so-called professional commentator or announcer, the AI can learn the actual commentary by the professional commentator and reproduce the commentary D30. In this way, for example, it is possible to enjoy the game as if a professional commentator were playing it.
[0124] The subject of live streaming and the subject to be filmed are not limited to games. For example, the action may be sports, etc. Furthermore, the game is not limited to a video game, but may be a table game such as actual shogi, actual chess, or actual billiards, or a board game, etc.
[0125] For example, when the action is sports, the user 4 becomes an actual athlete, and the play in the actual competition is filmed by a photographing device such as a camera to produce a video or the like. The user terminal 20 then acquires the video of the actual competition and performs an AI learning process. Meanwhile, the execution process sets the unknown data D2 as unknown image data D20 obtained by filming the game play. In this way, a commentary D30 as would be performed in an actual competition can be reproduced and output for the game play.
[0126] Besides, the action is not limited to sports, martial arts, games, etc. For example, at an exhibition or a retail store, the premises may be photographed by a fixed camera or the like, and the live commentary D30 may be output as an in-premises announcement or on a display installed in the premises.
[0127] Specifically, when a fixed camera captures the start of an event such as a seminar or time sale within the premises, a guide to the event (for example, the location where the event is being held) may be broadcast as a live commentary D30 and distributed as audio within the premises. In addition, if there is a commemorative event at an exhibition or the like, such as "You are the 100th visitor," this may be output as a live commentary D30.
[0128] [Variations for pre-processing] It is desirable that the unknown data D2 be preprocessed and input to the trained model A2. For example, the preprocessing is realized by an AI or an editing process based on an operation by the user 4. Note that the preprocessing may be performed by the user terminal 20 or may be executed by an information processing device other than the one on which the game is played, such as the server 11.
[0129] Hereinafter, the video before preprocessing, which is the unknown image data D20 in the unknown data D2, will be referred to as the "first video data." On the other hand, the video after the first video data is preprocessed and input to the trained model A2 and accompanied by commentary D30 will be referred to as the "second video data."
[0130] The pre-processing is a process of deleting or extracting a part of the first moving image data, that is, a process of extracting a scene that is a part of interest from all of the data of the first moving image data, and generating extracted data.
[0131] Specifically, the first video data often includes so-called "useless" images. For example, if the first video data includes an image with no subject at all at the beginning, it is desirable to delete the image because it is a scene that does not interest the viewer.
[0132] Alternatively, the pre-processing may be a process of extracting scenes of interest from the first video data, or a process of making the first video data into a so-called "highlight" format.
[0133] For example, if the first video data is 15 minutes in total, it is desirable to generate the second video data so that it is shorter, such as about 3 minutes. The length of the second video data may be specified by the user 4, or the length of the video may not be specified in advance and may be made as short as possible.
[0134] By shortening the video through preprocessing in this way, the number of objects that the trained model A2 has to process can be reduced, thereby reducing the processing load.
[0135] The pre-processing may also include, for example, processing for improving image quality, etc. For example, the pre-processing may include processing for brightening the entire image of the first moving image data, or for emphasizing edges, etc.
[0136] [Variations that change the nature of the commentary] It is desirable to generate a plurality of types of trained models A2 so that the characteristics of the commentary D30 are different. For example, the commentary D30 has different characteristics even for the same play if the commentator 8 is different. The characteristics are the personality (characteristics) of the commentator 8.
[0137] Depending on the commentator 8, there may be various characteristics, such as "enthusiastic" commentary, "calm" commentary, etc. Therefore, in order to impart such characteristics to the trained model A2, it is preferable that the live commentary voice 9 to be trained by the learning model A1, that is, the correct answer data D11, is different from the commentator 8.
[0138] It is preferable that the user 4 who is the viewer can select the trained model A2 in the execution process. Hereinafter, the operation of the viewer selecting the trained model A2 is referred to as a "selection operation." Also, the trained model A2 selected by the selection operation is referred to as a "selected model."
[0139] FIG. 11 is a diagram showing an example of a screen for performing a selection operation. For example, the selection operation is performed on a selection screen 61 or the like for selecting multiple trained models A2. It is preferable that the selection screen 61 displays the properties of each trained model A2. Therefore, the viewer selects the selection model by looking at the properties, etc. In this way, when the selection model can be selected, different live commentaries D30 are performed even for the same play, so that the viewer can enjoy various live commentaries D30.
[0140] For example, the selection screen 61 is output at the start of execution processing, etc. Then, the user 4 can select a selection model by operating a GUI (Graphical User Interface) such as a push button.
[0141] [Other variations] The big data D4 is desirably selected according to the type of object to be live-streamed and the object to be photographed, and becomes the learning data D1. Specifically, when the object is a game live-stream, it is desirably selected as the learning data D1 big data D4 to be used as the learning data D1 to be game live videos, or to weight the game live videos more heavily for learning, etc.
[0142] The words that are frequently used are different between game conversations (conversations are posts on the Internet, etc.) and non-game conversations (hereinafter referred to as "general conversations"; in other words, if the data in big data D4 is not specifically selected, the data will consist of general conversations).
[0143] For example, in the case of soccer games, the word "soccer" is often used. Meanwhile, in general conversation, words such as "hill" or "hacker", which are similar to the sound of "soccer", are also often used. Therefore, if general conversation is used as the learning data D1, words such as "soccer" may be easily recognized by speech recognition as words with similar sounds such as "hill" or "hacker". Therefore, by selecting a specific field (in this case, the field of soccer games) or by increasing the weight of words such as industry jargon in the specific field, misrecognition can be reduced.
[0144] In addition, even when a game is the subject of live commentary, conversations may be used as references without being limited to conversations about the game. For example, when a soccer game is the subject of live commentary, not only conversations about the soccer game but also conversations about actual soccer competitions (for example, professional soccer matches broadcast on television) may be used as references.
[0145] In addition, after downloading the trained model A2, the user 4 may perform further learning processing to allow the trained model A2 to undergo additional learning so that it follows the training method preferred by the user 4.
[0146] [Other embodiments] In the above example, the information processing device performs both pre-processing for the learning model and execution processing using the learned model. However, the pre-processing and execution processing do not have to be performed by the same information processing device. In addition, the pre-processing and execution processing do not have to be consistently performed by one information processing device. In other words, each process and data storage, etc. may be performed by an information system, etc., composed of multiple information processing devices.
[0147] The learning process may be additionally performed after the execution process or before the execution process.
[0148] The above-mentioned processing may be performed by an information processing device other than the server 11 and the user terminal 20 in an auxiliary manner.
[0149] Fig. 12 is a diagram showing a configuration example using an auxiliary device. Compared to the example shown in Fig. 1, the configuration shown in Fig. 12 differs in that an auxiliary device 60 is added. Note that the auxiliary device 60 may be a configuration that is used temporarily.
[0150] The auxiliary device 60 is an information processing device installed near the user terminal 20 (in this example, it is installed near the user terminal 20A, but it may be installed near other devices), etc. The auxiliary device 60 executes a part or all of a specific process on behalf of the user terminal 20 or the server 11.
[0151] For example, the auxiliary device 60 is provided with a device specialized for graphic processing and performs graphic processing at high speed. In this way, so-called edge computing or the like may be performed by installing the auxiliary device 60 or the like. In this way, the above-mentioned processing may be executed by utilizing the hardware resources of various information processing devices. Therefore, the above-mentioned processing may be executed by an information processing device different from the above-mentioned one.
[0152] The above-described processes and data used in the processes executed in this embodiment may be executed and stored by an information processing system. For example, the information processing system may execute or store the data in a plurality of information processing devices in order to realize redundant, distributed, parallel, or a combination of these processes or storage. Therefore, the present invention may be realized in devices other than the hardware configurations shown above and in systems other than the devices shown above.
[0153] In addition, the program according to the present invention is not limited to a single program, but may be a collection of multiple programs. In addition, the program according to the present invention is not limited to being executed by a single device, but may be executed by sharing among multiple information processing devices. Furthermore, the role sharing among the information processing devices is not limited to the above example. In other words, a part or all of the above-mentioned processing may be executed by an information processing device different from the above-mentioned information processing device.
[0154] Furthermore, a part or all of the means realized by the program can be realized by hardware such as an integrated circuit. Furthermore, the program may be provided by being recorded on a non-transitory recording medium readable by a computer. The recording medium refers to, for example, a hard disk, an SD card (registered trademark), an optical disk such as a DVD, or a server on the Internet. Therefore, the program may be distributed via a telecommunication line such as the Internet.
[0155] Furthermore, the information processing devices constituting the information processing system may be located overseas, i.e., an information processing device that executes some of the processes executed by the information processing system may be located overseas.
[0156] The present invention is not limited to the above-mentioned embodiments. Therefore, the present invention can be modified or modified without departing from the technical gist of the invention. Therefore, all technical matters included in the technical ideas described in the claims are the subject of the present invention. The above-mentioned embodiments are preferred examples. Those skilled in the art can realize various modifications from the disclosed contents, and such modifications are included in the technical scope described in the claims. [Explanation of symbols]
[0157] 1: System 1F11: Image data generating means 1F12: Live data input means 1F13: Learning data input means 1F14: Learning methods 1F21: Means for inputting unknown data 1F22: Output data generation means 1F23: Output method 2: Communication network 4: User 5:Administrator 6: Output screen 8: Commentator 9: Live audio 11: Server 20: User terminal 31: Learning device 32: Execution device 60: Auxiliary equipment 61: Selection screen D1: Training data D10: Image data D11: Correct data D2: Unknown data D20: Unknown image data D3: Output data D30: Live D31:Explanatory video D32: Summary D4: Big Data A1: Learning model A2: Trained model
Claims
1. Computer, an unknown data input means for inputting unknown data including unknown image data for which a correct answer to a user's action is unknown; A generation means for generating output data related to a commentary on the action based on the unknown data, using a trained model trained using training data including image data showing the action and commentary data related to a commentator's commentary on the action; a program that causes the program to function as an output means for performing output based on the output data;
2. The action is: It's about playing the game. The image data is 1 is a video showing an output screen of the game. The program according to claim 1.
3. The image data and the live data are For the actual competition, The unknown image data is Involves playing the same games as the competition The program according to claim 1.
4. The trained model is A plurality of types of the commentary having different characteristics are generated, The generating means includes: When a viewer who watches the live broadcast inputs a selection operation to select the trained model, the output data is generated using a selection model, which is the trained model selected by the selection operation. The program according to claim 1.
5. The unknown data is First video data is included; The output data is Second video data is included; The first video data is pre-processed to extract a portion of interest from all data, and then the extracted data is input to the trained model; The second video data is a video shorter than the first video data. The program according to claim 1.
6. an image data generating means for generating image data indicative of a user's action; a commentary data input means for inputting commentary data relating to a commentator's commentary on the action; learning data input means for inputting learning data including the image data and the real-time data; A learning means for learning a learning model using the learning data to generate a learned model; an unknown data input means for inputting unknown data including unknown image data indicating the action and a correct answer to the action being unknown; An output data generation means for generating output data relating to a commentary on the action based on the unknown data using the trained model; an output means for performing an output based on the output data; A system comprising:
Citation Information
Patent Citations
Game commentary information generation method, device and equipment and readable storage medium
CN116943247A
Computer system and audio information generation method
JP2021194229A
Server system, program, and live reporting distribution method of game live reporting play
JP2023146391A
Live audio real time generation system
JP2024011105A
Commentary video generation method and apparatus, server, and storage medium
US20230018621A1