Method, system, and program for inferring audience evaluation of performance data

A learning model-based system infers audience evaluations from performance data, addressing the inability of existing technologies to predict audience reception, enabling users to improve their performance effectively.

JP7718536B2Active Publication Date: 2025-08-05YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024075706
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-04
Filing Date
2024-05-08
Publication Date
2025-08-05
Estimated Expiration
2041-02-02

AI Technical Summary

Technical Problem

Existing performance evaluation technologies do not allow users to infer how their performance will be evaluated by an audience, hindering appropriate improvement.

Method used

A method and system that utilizes a learning model to infer audience evaluations by analyzing the relationship between performance data and audience feedback, incorporating sound, video, and operation data, and presenting inferred evaluations to users.

Benefits of technology

Enables users to predict how their performance will be received by an audience, facilitating targeted improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718536000001
    Figure 0007718536000001
  • Figure 0007718536000002
    Figure 0007718536000002
  • Figure 0007718536000003
    Figure 0007718536000003
Patent Text Reader

Abstract

To provide a method, system, and program by which evaluations on performance data are appropriately inferred.SOLUTION: In acquiring a learning model that has learned a relation between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience who received the performance, acquiring second performance data, processing the second performance data using the learning model to infer an evaluation for the second performance data, and outputting second evaluation data indicating an inference result, the first performance data includes sound data indicating sound played or operation data indicating performer's performance operation in the performance and video data indicating a video of the performer in the performance.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method, system, and program for inferring an audience's evaluation of performance data. [Background technology]

[0002] Conventionally, performance evaluation devices have been used to evaluate performance operations performed by users. For example, Patent Document 1 discloses a technology for evaluating performance operations by selectively targeting a portion of the entirety of a performed piece of music. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 3678135 Summary of the Invention [Problem to be solved by the invention]

[0004] Patent Document 1 discloses a technology for evaluating the accuracy of a user's performance, but not a technology for inferring how well a performance will be evaluated by an audience (whether it will be well-received by the audience). In order for a user to appropriately improve their performance, they need to be able to infer in advance how their performance will be evaluated.

[0005] An object of the present invention is to provide a method, system, and program for appropriately inferring an evaluation of performance data. [Means for solving the problem]

[0006] In order to achieve the above object, one aspect of the present invention provides a method implemented by a computer, which includes: acquiring a learning model that has learned the relationship between first performance data representing a performance by a performer and first evaluation data representing evaluations by an audience that has received the performance; acquiring second performance data; processing the second performance data using the learning model to infer evaluations of the second performance data; and outputting second evaluation data representing the inference results; the first performance data is divided into a series of performance pieces, and the first evaluation data includes a plurality of evaluation pieces corresponding to any of the series of performance pieces; The first performance data includes sound data indicating the played sound or operation data indicating the performance operation of the player during the performance, and video data showing an image of the player during the performance. [Effects of the Invention]

[0007] According to the present invention, an evaluation of performance data can be appropriately inferred. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is an overall configuration diagram showing an information processing system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing a hardware configuration of the information processing device. [Figure 3] FIG. 2 is a block diagram showing the hardware configuration of a learning server. [Figure 4] 1 is a block diagram showing a functional configuration of an information processing system according to an embodiment of the present invention. [Figure 5] FIG. 10 is a sequence diagram showing machine learning processing in the information processing system according to the embodiment of the present invention. [Figure 6] FIG. 10 is a sequence diagram showing an inference presentation process in the information processing system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Each embodiment described below is merely an example of a configuration that can realize the present invention. Each of the following embodiments can be modified or changed as appropriate depending on the configuration of the device to which the present invention is applied and various conditions. Furthermore, not all combinations of elements included in each of the following embodiments are necessarily essential for realizing the present invention, and some elements can be omitted as appropriate. Therefore, the scope of the present invention is not limited to the configurations described in each of the following embodiments. Furthermore, configurations that combine multiple configurations described in the embodiments can also be adopted as long as they are not mutually contradictory.

[0010] Fig. 1 is an overall configuration diagram showing an information processing system S according to an embodiment of the present invention. As shown in Fig. 1, the information processing system S of this embodiment includes an information processing device 100 and a learning server 200. The information processing device 100 and the learning server 200 can communicate with each other via a network NW. A distribution server DS, which will be described later, may be connected to the network NW.

[0011] The information processing device 100 is an information terminal used by a user, such as a personal device such as a tablet terminal, a smartphone, a personal computer (PC), etc. The information processing device 100 may be connected wirelessly or by wire to an electronic musical instrument EM, which will be described later.

[0012] The learning server 200 is a cloud server connected to the network NW, and can train a learning model M (described later) and supply the trained learning model M to other devices such as the information processing device 100. The server 300 is not limited to a cloud server, but may be a server on a local network. Furthermore, the functions of the server 300 in this embodiment may be realized by cooperative operation between the cloud server and a server on the local network.

[0013] In the information processing system S of this embodiment, the performance data A to be inferred is input to a learning model M that has machine-learned the relationship between performance data A indicating a performance by a performer and evaluation data B indicating an evaluation of the performance, and an evaluation of the input performance data A is inferred.

[0014] Fig. 2 is a block diagram showing the hardware configuration of the information processing device 100. As shown in Fig. 2, the information processing device 100 includes a CPU (Central Processing Unit) 101, a RAM (Random Access Memory) 102, a storage 103, an input / output unit 104, a sound collection unit 105, an imaging unit 106, a transmission / reception unit 107, and a bus 108.

[0015] The CPU 101 is a processing circuit that executes various calculations in the information processing device 100. The RAM 102 is a volatile storage medium that stores setting values used by the CPU 101 and functions as a working memory on which various programs are loaded. The storage 103 is a non-volatile storage medium that stores various programs and data used by the CPU 101.

[0016] The input / output unit 104 is an element (user interface) that accepts user operations on the information processing device 100 and displays various information, and is configured by, for example, a touch panel.

[0017] The sound collection unit 105 is an element, such as a microphone, that converts collected sound into an electrical signal and supplies it to the CPU 101. The sound collection unit 105 may be built into the information processing device 100, or may be connected to the information processing device 100 via an interface (not shown).

[0018] The imaging unit 106 is, for example, a digital camera, and is an element that converts captured images into electrical signals and supplies the signals to the CPU 101. The imaging unit 106 may be built into the information processing device 100, or may be connected to the information processing device 100 via an interface (not shown).

[0019] The transceiver 107 is an element that transmits and receives data to and from other devices such as the learning server 200. The transceiver 107 can connect to and transmit and receive data from an electronic musical instrument EM that a user uses to play music. The transceiver 107 may include multiple modules (for example, a Bluetooth (registered trademark) module and a Wi-Fi (registered trademark) module used for short-range wireless communication).

[0020] The bus 108 is a signal transmission path that interconnects the hardware elements of the information processing device 100 described above.

[0021] 3 is a block diagram showing the hardware configuration of the learning server 200. As shown in FIG. 3, the learning server 200 includes a CPU 201, a RAM 202, a storage 203, an input unit 204, an output unit 205, a transmission / reception unit 206, and a bus 207.

[0022] The CPU 201 is a processing circuit that executes various calculations in the learning server 200. The RAM 202 is a volatile storage medium that stores setting values used by the CPU 201 and functions as a working memory on which various programs are deployed. The storage 203 is a non-volatile storage medium that stores various programs and data used by the CPU 201.

[0023] The input unit 204 is an element that accepts operations on the learning server 200, and accepts input signals from a keyboard and mouse connected to the learning server 200, for example.

[0024] The output unit 205 is an element that displays various information, and outputs a video signal to a liquid crystal display connected to the learning server 200, for example.

[0025] The transmitting / receiving unit 206 is an element that transmits and receives data to and from other devices such as the information processing device 100, and is, for example, a network interface card (NIC).

[0026] The bus 207 is a signal transmission path that interconnects the hardware elements of the learning server 200 described above.

[0027] The CPUs 101, 201 of the above-described devices 100, 200 read out programs stored in the storages 103, 203 into the RAMs 102, 202 and execute the programs, thereby realizing the following functional blocks (controllers 150, 250, etc.) and various processes according to this embodiment. Each CPU is not limited to a normal CPU, but may be a DSP or an inference processor, or any combination of two or more of these. Furthermore, the various processes according to this embodiment may be realized by one or more processors, such as a CPU, DSP, inference processor, or GPU, executing a program.

[0028] FIG. 4 is a block diagram showing the functional configuration of an information processing system S according to an embodiment of the present invention.

[0029] The learning server 200 has a control unit 250 and a memory unit 260. The control unit 250 is a functional block that comprehensively controls the operation of the learning server 200. The memory unit 260 is composed of RAM 202 and storage 203, and stores various data used by the control unit 250 (particularly performance data A and evaluation data B). The control unit 250 has sub-functional blocks: a server authentication unit 251, a data acquisition unit 252, a data preprocessing unit 253, a learning processing unit 254, and a model distribution unit 255.

[0030] The server authentication unit 251 is a functional block that authenticates a user in cooperation with the information processing device 100 (authentication unit 151). The server authentication unit 251 determines whether or not the authentication data supplied from the information processing device 100 matches the authentication data stored in the storage unit 260, and transmits the authentication result (permission or denial) to the information processing device 100.

[0031] The data acquisition unit 252 is a functional block that receives distribution data from an external distribution server DS via the network NW and acquires performance data A and evaluation data B. The distribution server DS is a server that distributes video data containing video and sound, such as live video, as distribution data. The distribution data includes video data (e.g., video data) showing the performer's performance, sound data (e.g., audio data), and operation data (e.g., MIDI data). The distribution data also includes subjective data on the performance. The subjective data is an evaluation value assigned by a viewer to the performer's performance and is chronologically associated with the video. For example, the evaluation value of the evaluation data may be assigned a time in the corresponding video or a serial number (frame number) of the video. The video and the subjective data may be integrated. It is preferable that the distribution data include operation data, such as MIDI data, that indicates the performance operations performed by the performer during the performance. The operation data may include pedal operations on an electronic piano or operation of an effector on an electric guitar.

[0032] The data acquisition unit 252 acquires performance data A by dividing the video data and sound data contained in the received distribution data into a plurality of performance pieces in time series, and stores the data in the storage unit 260. The data acquisition unit 252 may divide the video data and sound data into performance pieces for each phrase indicated by a break in the performance, or may divide the video data and sound data into performance pieces based on a performance motif, or may divide the video data and sound data into performance pieces based on a chord pattern.

[0033] The performance data A may include operation data divided in time series instead of or in addition to the sound data divided in time series. That is, the performance data A includes either or both of sound data representing sounds produced by the performance and operation data generated based on the performance of the electronic musical instrument EM.

[0034] Furthermore, the data acquisition unit 252 acquires evaluation data B including evaluation pieces indicating evaluations for each divided performance piece based on the subjective data and evaluation times included in the received distribution data, and stores the evaluation data B in the storage unit 260. The evaluation data B is data indicating the time-series progression of evaluations for the performance data A, which is configured in a time series. The evaluation data B may include the time of the performance piece corresponding to the evaluation piece, or may be assigned a serial number corresponding to the performance piece and the evaluation piece, or the evaluation piece may be embedded in the corresponding performance piece. The data acquisition unit 252 stores the acquired performance data A and evaluation data B in the storage unit 260.

[0035] The data pre-processing unit 253 is a functional block that performs data pre-processing such as scaling on the performance data A and evaluation data B stored in the memory unit 260 so that they are in a format suitable for training the learning model M (machine learning).

[0036] The learning processing unit 254 is a functional block that trains a learning model M using the performance data A after data preprocessing as input data and the evaluation data B after data preprocessing as training data. Any machine learning model can be adopted for the learning model M of this embodiment. Preferably, a recurrent neural network (RNN) adapted to time-series data and its derivatives (long short-term memory (LSTM), gated recurrent unit (GRU), etc.) are adopted for the learning model M. The learning model M may also be configured according to an attention-based algorithm.

[0037] The model distribution unit 255 is a functional block that supplies the learning model M trained by the learning processing unit 254 to the information processing device 100.

[0038] The information processing device 100 has a control unit 150 and a memory unit 160. The control unit 150 is a functional block that comprehensively controls the operation of the information processing device 100. The memory unit 160 is configured with a RAM 102 and a storage 103, and stores various data used by the control unit 150. The control unit 150 has sub-functional blocks, an authentication unit 151, a performance acquisition unit 152, a video acquisition unit 153, a data preprocessing unit 154, an inference processing unit 155, and an evaluation presentation unit 156.

[0039] The authentication unit 151 is a functional block that authenticates users in cooperation with the learning server 200 (server authentication unit 251). The authentication unit 151 transmits authentication data, such as a user identifier and password, entered by the user using the input / output unit 104 to the learning server 200, and permits or denies access to the user based on the authentication result received from the learning server 200. The authentication unit 151 can supply the user identifier of the authenticated (access-permitted) user to other functional blocks.

[0040] The performance acquisition unit 152 is a functional block that acquires either or both of sound data and operation data representing a user's performance. Both the sound data and operation data are data (sound characteristic data) indicating the characteristics (e.g., onset time and pitch) of multiple sounds included in the music piece being performed, and are a type of high-dimensional time-series data that expresses the user's performance. The performance acquisition unit 152 may acquire sound data based on electrical signals generated by the sound collection unit 105 by collecting sounds produced by the user's performance. The performance acquisition unit 152 may also acquire operation data generated based on the user's performance of the electronic musical instrument EM from the electronic musical instrument EM via the transmission / reception unit 107. The electronic musical instrument EM may be, for example, an electronic keyboard instrument such as an electronic piano, an electronic string instrument such as an electric guitar, or an electronic wind instrument such as a wind synthesizer. The performance acquisition unit 152 supplies the acquired sound characteristic data to the data preprocessing unit 154. The performance acquisition unit 152 may also assign a user identifier supplied from the authentication unit 151 to the sound characteristic data and transmit the data to the learning server 200.

[0041] The video acquisition unit 153 is a functional block that acquires video data showing the user's performance. The video data is motion data that indicates the characteristics of the user's (performer's) movements during performance, and is a type of high-dimensional time-series data that expresses the user's performance. The video acquisition unit 153 may acquire the motion data based on electrical signals generated by the imaging unit 106 capturing an image of the user performing. The motion data is, for example, data that captures the user's skeleton in time series. The video acquisition unit 153 supplies the acquired video data to the data preprocessing unit 154. Note that the video acquisition unit 153 can also assign a user identifier supplied from the authentication unit 151 to the video data and transmit it to the learning server 200.

[0042] The data preprocessing unit 154 is a functional block that performs data preprocessing such as scaling on the performance data A, which includes the sound characteristic data supplied from the performance acquisition unit 152 and the video data supplied from the video acquisition unit 153, so that the performance data A is in a format suitable for inference by the learning model M.

[0043] The inference processing unit 155 is a functional block that inputs the preprocessed performance data A as input data to the learning model M trained by the learning processing unit 254, and infers evaluation data B indicating an evaluation of the performance data A. As mentioned above, the evaluation data B includes evaluation pieces indicating evaluations for each of the multiple performance pieces included in the performance data A.

[0044] The evaluation presentation unit 156 is a functional block that presents the evaluation data B inferred by the inference processing unit 155 to the user. For example, the evaluation presentation unit 156 causes the input / output unit 104 to display the evaluations for each of the multiple performance pieces included in the performance data A in chronological order. Note that instead of or in addition to visually presenting the evaluation data B, the evaluation presentation unit 156 may also audibly or tactilely present the evaluation data B to the user. The evaluation presentation unit 156 may also display the evaluations on a display unit of another device, for example, the electronic musical instrument EM.

[0045] 5 is a sequence diagram showing machine learning processing in an information processing system S according to an embodiment of the present invention. The machine learning processing of this embodiment is executed in a learning server 200. Note that the machine learning processing of this embodiment may be executed periodically, or may be executed in response to a request from the information processing device 100 based on a user instruction.

[0046] In step S510, the data acquisition unit 252 acquires performance data A and evaluation data B based on the distribution data received from the distribution server DS, and stores them in the storage unit 260. Note that the distribution data may be acquired in advance by the data acquisition unit 252 and stored in the storage unit 260, or may be acquired by the data acquisition unit 252 in this step.

[0047] In step S520, the data pre-processing unit 253 reads out the data set including the performance data A and the evaluation data B stored in the storage unit 260, and performs data pre-processing.

[0048] In step S530, the learning processing unit 254 trains the learning model M based on the data set preprocessed in step S520, using the performance data A as input data and the evaluation data B as teacher data, and stores the trained learning model M in the storage unit 260. For example, if the learning model M is a neural network system, the learning processing unit 254 may perform machine learning of the learning model M using backpropagation or the like.

[0049] In step S540, the model distribution unit 255 supplies the learning model M trained in step S530 to the information processing device 100 via the network NW. The control unit 150 of the information processing device 100 stores the received learning model M in the storage unit 160.

[0050] 6 is a sequence diagram showing an inference presentation process in the information processing system S according to an embodiment of the present invention. In this embodiment, the information processing device 100 infers an evaluation for each performance piece and visually presents the inferred evaluation to the user.

[0051] In step S610, the performance acquisition unit 152 acquires either or both of sound data and operation data (sound characteristic data) from the electronic musical instrument EM or the like, as described above, and supplies them to the data pre-processing unit 154.

[0052] In step S620, the moving image acquisition unit 153 acquires the video data as described above and supplies it to the data preprocessing unit 154.

[0053] In step S630, the data preprocessing unit 154 performs data preprocessing on the performance data A, which includes the sound characteristic data supplied from the performance acquisition unit 152 in step S610 and the video data supplied from the video acquisition unit 153 in step S620, and supplies the preprocessed performance data A to the inference processing unit 155.

[0054] In step S640, the inference processing unit 155 inputs the performance data A supplied from the data preprocessing unit 154 as input data to the trained learning model M stored in the memory unit 160. The learning model M processes the input performance data A and infers the audience's evaluation of each performance piece included in the performance data A. The inferred value indicating the evaluation may be a discrete value or a continuous value. The inferred evaluation (evaluation data B) for each performance piece is supplied from the inference processing unit 155 to the evaluation presentation unit 156.

[0055] In step S650, the evaluation presenting unit 156 presents to the user the evaluation data B inferred by the inference processing unit 155 in step S640. Various modes can be envisioned for presenting the evaluation data B to the user.

[0056] For example, consider an application that simulates and displays the reactions of a virtual audience (e.g., avatars in a virtual reality (VR) space) to a user's performance. In the above application, the evaluation presentation unit 156 causes the input / output unit 104 to display the reactions of the virtual audience based on the evaluation data B in synchronization with the playback of performance data A. The evaluation presentation unit 156 displays reactions that indicate excitement, such as standing up or cheering, when the inferred evaluation is higher than a threshold, and displays reactions that indicate decline, such as sitting down, remaining silent, or booing, when the inferred evaluation is lower than the threshold.

[0057] Also, for example, consider an application that objectively displays a user's performance by quantifying and graphing it. In the above application, the evaluation presenting unit 156 causes the input / output unit 104 to display a waveform representing performance data A and a graph showing the transition of evaluation data B corresponding to the performance data A.

[0058] The inference display process of steps S610 to S650 described above may be executed in real time in parallel with the input of the performance data A to the information processing device 100, or may be executed after the performance data A has been stored in the information processing device 100.

[0059] As described above, in the information processing system S of this embodiment, the trained learning model M appropriately infers the evaluation corresponding to each of the multiple performance pieces included in the performance data A. The information processing device 100 presents the inferred evaluation for each performance piece to the user. As a result, the user can predict how the audience will evaluate their own performance.

[0060] <Modification> The above embodiments can be modified in various ways. Specific modified embodiments are exemplified below. Two or more embodiments arbitrarily selected from the above embodiments and the following examples may be combined as long as they are not mutually inconsistent.

[0061] In the above embodiment, the performance data A is divided into multiple performance pieces in time series and used for the learning process and the inference process. However, the performance data A may not be divided and may correspond to a single piece of music.

[0062] In the above-described embodiment, various methods may be used to divide the performance data A. For example, the multiple performance pieces may be multiple performance sections obtained by dividing the music piece at predetermined time intervals, or multiple phrases identified based on the performance data A.

[0063] The evaluation data B in the above-described embodiment is subjective data indicating the evaluation value given by the viewer to the performance of the performer shown in the distribution data, but other information may be used as the evaluation data B.

[0064] For example, posting data regarding the number of posts posted by viewers related to the performer's performance may be used as evaluation data B. The posting data is, for example, text information associated with a video fragment included in a video, and is included in the distribution data, and the number of posts is tallied for each performance fragment.

[0065] Alternatively, for example, reaction data indicating the behavior of the audience during the performance may be used as evaluation data B. Reaction data is information indicating characteristics of the audience's movements during the performance. The data acquisition unit 252 can acquire reaction data by analyzing footage (audience footage) of the music performance video included in the distribution data during a period in which the audience is displayed. The reaction data may be, for example, data acquired over time of the skeletons of each audience member, data indicating the magnitude of the movement of the entire audience, data indicating the facial expressions of each audience member, or data indicating the body temperature of the audience acquired by an infrared camera or the like.

[0066] In the above-described embodiment, the evaluation presenting unit 156 visually presents the evaluation data B to the user. Instead of or in addition to presenting the evaluation data B, the control unit 150 may present candidate visual effects for the video shown in the performance data A so as to improve the inferred evaluation. Visual effects for the video are, for example, information indicating the timing of switching camera angles when the video is shot with multiple cameras, or the start and end timing of a fade-out.

[0067] In the above-described embodiment, the information processing device 100 infers a rating using the learning model M supplied from the learning server 200. However, each process related to rating inference may be executed by any device constituting the information processing system S. For example, the learning server 200 may preprocess the performance data A supplied from the information processing device 100 and input the preprocessed performance data A as input data into the learning model M stored in the storage unit 260, thereby inferring a rating for the performance data A. According to the configuration of this modified example, the learning server 200 can execute inference processing using the learning model M with the performance data A as input data. As a result, the processing load on the information processing device 100 is reduced.

[0068] Furthermore, the electronic musical instrument 100 of the above-described embodiment may have the functions of the control device 200 , or the control device 200 may have the functions of the electronic musical instrument 100 .

[0069] The same effect may be achieved by reading a storage medium storing each control program represented by software for achieving the present invention into each device. In this case, the program code itself read from the storage medium realizes the novel functions of the present invention, and the non-transitory computer-readable recording medium storing the program code constitutes the present invention. The program code may also be supplied via a transmission medium, in which case the program code itself constitutes the present invention. In these cases, storage media may include, in addition to ROM, floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tape, and non-volatile memory cards. The term "non-transitory computer-readable recording medium" also includes volatile memory (e.g., DRAM (Dynamic Random Access Memory)) within a computer system that acts as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and that retains the program for a certain period of time.

[0070] While the present invention has been described in detail above based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate. [Explanation of symbols]

[0071] 100 Information processing device 150 control section 160 Storage section 200 Learning Server 250 control section 260 Storage section A Performance Data B. Evaluation Data DS distribution server EM Electronic Instrument M Learning Model S Information Processing System

Claims

1. obtaining a learning model that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is divided into a series of performance pieces; the first evaluation data includes a plurality of evaluation pieces associated with any of the series of performance pieces; the first performance data includes sound data indicating a played sound or operation data indicating a performance operation of a player during a performance, and video data indicating a video of the player during a performance; A computer-implemented method.

2. The method according to claim 1 , wherein the second performance data includes sound data indicating the sounds played or operation data indicating the performance operations of the player during the performance, and video data showing an image of the player during the performance.

3. The method according to claim 1 or 2, wherein the video data is movement data that indicates characteristics of the movement of the performer during the performance.

4. 4. The method according to claim 1, wherein the first evaluation data includes at least one of subjective data indicating an evaluation given by the audience to the performance, reaction data indicating the audience's reaction to the performance, and posting data regarding the amount of posts to the performance.

5. The method of claim 4, wherein the reaction data is at least one of data obtained over time of the skeletons of each spectator, data indicating the magnitude of the movement of the entire spectator, data indicating the facial expressions of each spectator, and data indicating the body temperature of the spectators obtained by an infrared camera.

6. A method according to any one of claims 1 to 3, wherein the type and timing of application of video effects to the video data included in the second performance data are presented in a user interface as candidates for the video effects so as to improve the evaluation indicated by the second evaluation data.

7. 7. The method according to claim 1, wherein the first performance data and the first evaluation data are received as distribution data from an external source.

8. 8. The method according to claim 1, further comprising displaying a simulated reaction of a virtual audience on a user interface based on the second evaluation data.

9. A learning model is obtained that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is a single piece of music that is not divided, the first evaluation data includes a plurality of evaluation pieces corresponding to any of a series of performance pieces obtained by dividing the first performance data; the first performance data includes sound data indicating a played sound or operation data indicating a performance operation of a player during a performance, and video data indicating a video of the player during a performance; A computer-implemented method.

10. a memory for storing a program; one or more processors that execute the program; The one or more processors execute the program stored in the memory, obtaining a learning model that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is divided into a series of performance pieces; the first evaluation data includes a plurality of evaluation pieces associated with any of the series of performance pieces; The first performance data includes sound data indicating the sounds played or operation data indicating the performance operations of the player during the performance, and video data showing an image of the player during the performance.

11. On the computer, obtaining a learning model that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is divided into a series of performance pieces; the first evaluation data includes a plurality of evaluation pieces associated with any of the series of performance pieces; the first performance data includes sound data indicating a played sound or operation data indicating a performance operation of a player during a performance, and video data indicating a video of the player during a performance; A program for executing a process.

12. A memory for storing a program; one or more processors that execute the program; The one or more processors execute the program stored in the memory, obtaining a learning model that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is a single piece of music that is not divided, the first evaluation data includes a plurality of evaluation pieces corresponding to any of a series of performance pieces obtained by dividing the first performance data; The first performance data includes sound data indicating the sounds played or operation data indicating the performance operations of the player during the performance, and video data showing an image of the player during the performance.

13. A computer comprising: obtaining a learning model that learns the relationship between first performance data indicating a performance by a performer and first evaluation data indicating an evaluation by an audience that has received the performance; Acquire second performance data; using the learning model to process the second performance data and infer an evaluation of the second performance data; outputting second evaluation data indicating the inference result; the first performance data is a single piece of music that is not divided, the first evaluation data includes a plurality of evaluation pieces corresponding to any of a series of performance pieces obtained by dividing the first performance data; the first performance data includes sound data indicating a played sound or operation data indicating a performance operation of a player during a performance, and video data indicating a video of the player during a performance; A program for executing a process.

Citation Information

Patent Citations

  • Big data based audio evaluation method and system, equipment and storage medium

    CN110675879A

  • Performance evaluation device and performance evaluation system

    JP3678135B2

  • Method and System for Assessing a Musical Performance

    US20070256543A1