Music processing device and music processing application
The music processing device addresses the challenge of supporting image outputs based on music development over time by estimating music characteristics and providing these estimates to an effect processing device, enhancing analysis speed and accuracy.
Patent Information
- Application Number
- JP2024187457
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The existing music image output devices and programs cannot support the output of images based on the development of music after a predetermined time.
A music processing device that acquires music data and processes it to estimate the characteristics of music that will occur at a prediction target time point after a first predetermined time point, and provides these estimates to an effect processing device to select recognizable effects.
Enables the support of effects based on the development of music over time, improving analysis speed and accuracy while reducing data storage needs.
Smart Images

Figure 0007698833000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a music processing apparatus and a music processing application for processing music data of music being played.
Background Art
[0002] An example of a music processing apparatus (music image output apparatus) and a music processing application (program) for processing music data of music being played is described in Patent Document 1. The music image output apparatus described in Patent Document 1 includes a music storage unit for storing music, an output instruction receiving unit for receiving an output instruction of music, a music output unit for outputting music in response to the output instruction, an attribute value acquisition unit for acquiring one or more attribute values based on an analysis result of the music, an image acquisition unit for acquiring an image using the one or more attribute values, and an image output unit for outputting the image.
[0003] Further, in Patent Document 1, a computer-accessible recording medium is described as "a program for causing the computer to function as a music storage unit for storing music, an output instruction receiving unit for receiving an output instruction of music, a music output unit for outputting the music in response to the output instruction, an attribute value acquisition unit for acquiring one or more attribute values based on an analysis result of the music, an image acquisition unit for acquiring an image using the one or more attribute values, and an image output unit for outputting the image."
[0004] Furthermore, Patent Document 1 describes that "the attribute value acquisition unit acquires one or more attribute values based on the analysis result of music. The analysis of music may be, for example, the analysis of sound or the analysis of lyrics. It is preferable that the analysis of music is the analysis of both sound and lyrics. The analysis of sound may be, for example, to acquire the feature amount of sound. The feature amount may be, for example, the change in amplitude in the sound waveform, the change in frequency components constituting the sound waveform, etc. The feature amount may be represented by, for example, an acoustic feature amount vector. The acoustic feature amount vector is a vector having two or more feature amounts such as the change in amplitude and the change in frequency components as components. However, the expression form of the feature amount is not limited."
[0005] Furthermore, Patent Document 1 describes that "the analysis of lyrics may be, for example, to acquire content words from the lyrics by natural language processing using machine learning such as deep learning, SVM, decision trees, or morphological analysis. ··· Information for specifying the surface scene is, for example, terms such as "coast", "fireworks", "Christmas", "graduation ceremony", etc. Also, the inner scene is a scene related to the user's inner self and may be called a subjective scene."
[0006] Furthermore, Patent Document 1 describes that "information for specifying the inner scene is, for example, terms such as "date", "lover", "nervous", "relax". The impression is the impression held by the user. Information for specifying the impression is, for example, terms such as "happy", "sad", "lonely", "fun". ··· And the attribute value acquisition unit, for example, uses the learning information in the storage unit to apply the vector of the sentence constituting the lyrics or the sentence obtained by morphological analysis of the sentence to the learning information for each term, determines whether it corresponds to each term, and acquires one or more terms determined to correspond as attribute values. " Furthermore, Patent Document 1 describes that "with such a configuration, an image corresponding to the music can be output during the output of the music."
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] The inventor of the present application recognized the problem that the "music image output device and program" described in Patent Document 1 cannot support the output of an image according to the development of music after a predetermined time.
[0009] An object of the present embodiment is to provide a music processing device and a music processing application capable of supporting an effect according to the development of music that occurs after a predetermined time.
Means for Solving the Problems
[0010] This embodiment is a music processing device having a processor that acquires music data and executes processing. The processor includes a data acquisition process for acquiring music data of music from a first predetermined time point, and by processing the acquired music data, estimates the characteristics of the music that will occur at a prediction target time point after the first predetermined time point. Reason and A provision process of providing the estimation result of the characteristics of the music obtained in the estimation process step to an effect processing device that selects an effect recognizable by a human according to the characteristics of the music is executed. The inference process executed by the processor includes a first process of processing the acquired music data, and a second process that is performed after the first process and infers the characteristics of the music generated at the music venue at the time of the inference target by processing the acquired music data with a learned model. The first process executed by the processor includes a process of generating a plurality of analysis data by performing a clipping process on a plurality of music data with different positions in the time axis direction within a total predetermined time range from the first predetermined time point to the inference target time point, and for each of the plurality of generated analysis data, comparing the output data obtained by processing with a model before or during learning that infers the characteristics of the music generated at the music venue at the inference target time point, and the correct answer data including the music data from the first predetermined time point to the future section after the inference target time point, and using the error of the music characteristics as the comparison result to perform machine learning on the model before or during learning, and generating a plurality of the learned models corresponding to each of the plurality of analysis data. A music processing device having such a configuration is provided.
Effects of the Invention
[0011] According to the present embodiment, it is possible to support an effect according to the development of music at the first time point.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0013] (Overview) The system has a music processing device and a performance processing device. The music processing device processes music data of music played at a concert venue. The music processing device generates information necessary for performing an effect corresponding to the music and automatically provides it to the performance processing device in real time. In addition, some specific examples of the system and the music processing device, and some specific examples of music processing applications are disclosed with reference to the drawings.
[0014] (Specific Example 1) (1) Overall Configuration As shown in FIG. 1, the system 10 has a music processing device 11 and a performance processing device 12. The specific example 1 shown in FIG. 1 is an example in which the music processing device 11 and the performance processing device 12 are arranged in the concert venue A1. There are humans (audience) in the concert venue A1 who recognize the live performance of music by hearing. The music processing device 11 is communicably connected to the network 13. In addition, the music processing device 11 is communicably connected to the performance processing device 12 via the network 15. The microphone 14 is a device provided in the concert venue A1 for music. The microphone 14 acquires the sound of the music actually played in the concert venue A1 and converts it into an electrical signal for output.
[0015] The performance venue A1 can be any of a concert hall, a music studio, a multipurpose hall, a convention hall, a gymnasium, an outdoor theater, etc. The sound source 16 of the music performed in the performance venue A1 includes humans (singers) and musical instruments. The musical instruments can be either acoustic instruments or electronic instruments. An acoustic instrument is an instrument equipped with a mechanical vibrating part. An electronic instrument is an instrument that does not have a mechanical vibrating part and uses an oscillating sound generated by an electronic circuit.
[0016] Furthermore, a computer 62 may be provided in the performance venue A1, and the music may be reproduced by the speaker 62A of the computer 62. The music data of the performance reproduced by the speaker 62A of the computer 62 is returned to the computer 62 by a loopback function, and the returned music data is sent to the music processing device 11 via the network 13 without passing through the microphone 14. The genre of the music handled by the music processing device 11 can be any of pop, rock, dance, Latin, classical, march, vocal music, Japanese traditional music, etc.
[0017] The networks 13 and 15 are each constituted by one or more communication systems among a wireless communication system or a wired communication system. The wireless communication system includes radio wave communication, optical communication, infrared communication, radio communication, satellite communication, etc. The wired communication system includes a communication circuit, a communication cable, and an antenna. The network includes one or more networks among the Internet, an intranet, a wide area network, and an intranet. The networks 13 and 15 include short-range wireless communication. The short-range wireless communication includes, for example, wireless LAN (Wi-Fi (registered trademark)) and Bluetooth (registered trademark).
[0018] The music processing device 11 can acquire music through the network 13 and execute various processes. The music processing device 11 executes preprocessing for performing an effect with the effect equipment 50 in accordance with the music played or reproduced at the concert venue A1. The effect processing device 12 is configured to automatically select and manage the effects to be executed by the effect equipment 50 based on the estimation results acquired from the music processing device 11. The effect equipment 50 is configured to execute effects in real time in accordance with the music played or reproduced at the concert venue A1. The effect content executed by the effect equipment 50 is recognizable by humans (the audience) through vision, touch, smell, etc.
[0019] (2) Configuration of the music processing device The music processing device 11 is a computer including a main body (casing), a processor 17, a main memory 18, an auxiliary memory 19, an operation device 20, a display device 21, a communication device 22, etc. The processor 17 is provided inside the main body and is constituted by a central processing unit (CPU (Central Processing Unit)) in which an arithmetic unit (arithmetic circuit) and a control unit (control circuit) are integrated. The processor 17 is communicably connected to the main memory 18, the auxiliary memory 19, the operation device 20, the display device 21, the communication device 22, etc. via a bus 23.
[0020] The processor 17 controls including other devices and circuits provided inside the main body and devices and circuits provided outside the main body. Further, in addition to the central processing unit, the processor 17 has arithmetic processing circuits such as a digital signal processor, an application specific integrated circuit (ASIC), and a GPU.
[0021] The GPU is an abbreviation for (Graphics Processing Unit), and the GPU is a graphic controller that performs arithmetic processing necessary for performing 3D graphic image processing, etc. Further, the GPU has a configuration for executing machine learning for constructing an acoustic model, specifically, deep learning, in the stage of processing and analyzing music data.
[0022] The processor 17 executes various processes by operating non-temporary applications stored in the auxiliary memory 19. The processes executed by the processor 17 include the processes themselves performed within the processor 17, judgments, analyses, inferences, control and instructions for other elements, acquisition of information and signals from other elements, storage processing of information in the auxiliary memory 19, and the like.
[0023] The main memory 18 is a volatile storage device, and the main memory 18 functions as a work area and a buffer area when the processor 17 executes processes. The auxiliary memory 19 is a non-volatile storage device, that is, a non-temporary storage medium. A non-temporary application is stored in the auxiliary memory 19. The application includes programs, configuration files, files for storing data, various libraries, and the like.
[0024] Also, various information used by the processor 17 to execute various processes, various information as a result of the processor 17 executing various processes, and the like are stored in the auxiliary memory 19. The various information stored in the auxiliary memory 19 includes, here, the information itself, data, graphs, maps, charts, and the like.
[0025] The auxiliary memory 19 has a larger capacity than the main memory 18, and the auxiliary memory 19 operates in accordance with input commands and output commands from the processor 17. The non-temporary storage medium, the auxiliary memory 19, is composed of, for example, a magnetic disk, an optical disk, a flash memory, or the like. As the magnetic disk, a hard disk drive is exemplified. As the optical disk, a compact disk, a digital video disk, a Blu-ray disk, or the like is exemplified.
[0026] The flash memory is a type of semiconductor memory, and examples of the flash memory include an SD memory card, a USB flash drive, a solid state drive, and the like. One or more elements included in the auxiliary memory 19 can be defined as a storage medium 19A that can be attached to and detached from the main body.
[0027] The operation device 20 is operated by a "music processing administrator" who uses the music processing device 11. When switching the operation and stop of the music processing device 11, when causing the processor 17 to execute various processes, when causing the display device 21 to display information, when acquiring a signal including music data via the network 13, when sending information to the effect processing device 12, etc., the operation device 20 is operated.
[0028] The operation device 20 includes at least any one of elements such as a keyboard, a touch pad, a mouse, a liquid crystal display, an organic electro-luminescence display, etc. These elements are appropriately selected depending on whether the music processing device 11 is a portable computer or a stationary computer, and the attachment structure to the main body is appropriately selected. That is, the operation device 20 has a structure directly attached to the main body or a structure connected to the main body via a cable.
[0029] The display device 21 is directly attached to the main body or connected to the main body via a cable. The display device 21 is a display visually observed by the "music processing administrator", and the display includes structures such as a liquid crystal display, an organic electro-luminescence display, etc. For these displays, the connection structure to the main body is appropriately selected depending on whether the music processing device 11 is a portable computer or a stationary computer. Note that the display may be defined as a monitor. Various information is displayed on the screen of the display device 21. Also, the display device 21 can display operation buttons, operation tabs, etc. on the screen. That is, the display device 21 can also serve as the operation device 20.
[0030] The communication device 22 includes devices, equipment, and specifications for connecting the music processing device 11 to the networks 13 and 15 by at least one of a wireless communication system or a wired communication system. The communication device 22 includes a communication circuit, a cable, an antenna, a communication port, a communication connector, a communication hub, etc.
[0031] The configuration of the processor 17 will be specifically described. The processor 17 functions as a music data processing unit 24, an expansion prediction unit 27, and an artificial intelligence unit 28 by operating non-temporary applications stored in the auxiliary memory 19.
[0032] The music data processing unit 24 is configured to classify and analyze music data of music acquired via the network 13, perform clipping processing on the music data, generate analysis data, analyze the analysis data, and so on. Specifically described, the music data processing unit 24 processes the acquired music data based on the acquired music data and the information stored in the auxiliary memory 19. The processing of the music data performed by the music data processing unit 24 includes classifying and analyzing the characteristics of the music, and specifically includes processing for estimating one or more items among the state of the music, determination of parts, type of musical instrument, human emotion expressed by the music, melody, etc. Note that the clipping processing, generation of analysis data, and analysis of analysis data performed by the music data processing unit 24 will be described later.
[0033] In addition, the music data processing unit 24 generates inference data by collaborating with the artificial intelligence unit 28 based on the analysis result of the music data and the information stored in the auxiliary memory 19. The inference data is data used when the expansion prediction unit 27 predicts the expansion of the music after a predetermined time, that is, predicts the characteristics of the music.
[0034] The expansion prediction unit 27 predicts the future of the music, that is, the expansion after a predetermined time, by collaborating with the artificial intelligence unit 28 based on the inference data and the information stored in the auxiliary memory 19. The details of the processing performed by the expansion prediction unit 27 will be described later.
[0035] The artificial intelligence unit 28 analyzes the music data in cooperation with the music data processing unit 24 based on various information acquired by the processor 17 and various information stored in the auxiliary memory 19 And perform machine learning on a model before or during learning to generate a learned model, and use the learned model It has a configuration for predicting the expansion of the music after a predetermined time.
[0036] The artificial intelligence unit 28 is equipped with a language model such as a Transformer and a neural network, and can perform processing as generative artificial intelligence. The Transformer includes GPT (Generative Pre-trained Transformer), BERT (Bidirectional Encoder Representations from Transformers), etc. The neural network is composed of, for example, a convolutional neural network (CNN). The convolutional neural network mainly has a convolutional layer, a pooling layer, a fully connected layer, etc.
[0037] The language model is an example of a learning model by a machine learning algorithm. Specific algorithms of machine learning include the nearest neighbor method, the naive Bayes method, decision trees, support vector machines, deep learning using neural networks, etc. The artificial intelligence unit 28 can appropriately apply the above algorithms.
[0038] The machine learning executed by the artificial intelligence unit 28 includes three types: supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning. Supervised learning makes a machine learn using learning data including teacher data. The teacher data is a method of making a machine (processor 17) learn using learning music data (input data) and music data including correct answer data (output data), and forms a learned model based on a music data set (sample data). Supervised learning includes classification and regression learning models. Classification is to classify (identify) the features of the acquired music data, and regression is to infer (predict) the future features of the acquired music data, which can also be said to be a task of time series analysis. Unsupervised learning is a method of making a machine learn using music data without correct answer information, and the machine forms a learned model based on the regularity and similarity of the music data set. Also, unsupervised learning includes classification (identification) of the features of the music included in the music data.
[0039] Semi-supervised learning is a type of machine learning that combines supervised learning and unsupervised learning. Semi-supervised learning uses both labeled music data and unlabeled music data to train artificial intelligence models for classification and regression tasks. Labels are tags or marks attached to music data, and labels contain music features. Reinforcement learning is a method of allowing a machine to learn through trial and error so as to maximize a set "score (the inference result of music features)".
[0040] The artificial intelligence unit 28 analyzes and processes various information and data stored in the auxiliary memory 19 and newly acquired information and data to execute machine learning. In machine learning, for example, deep learning using a neural network is performed to generate a trained model. The generated trained model is stored in the auxiliary memory 19. Note that the music processing device 11 may be composed of a plurality of single computers or may be composed of a plurality of computers.
[0041] If the music processing device 11 is composed of a plurality of computers, the music data processing unit 24 can be separately provided in different computers, and the computer that clips the music data and the computer that generates analysis data can be configured separately. Also, at least one of the functional units of the music data processing unit 24, the expansion inference unit 27, and the artificial intelligence unit 28 included in the processor 17 may be configured by an application programming interface operating on the website of the network 13. For example, when music data is transmitted (POST) from the music processing device 11 to a specified URL (: Uniform Resource Locator) in the network 13, a list of parameters in JSON (JavaScript Object Notation) format, performance processing content, etc. may be returned to the music processing device 11 as a response from the specified URL.
[0042] (3) Configuration of the performance processing device The performance processing device 12 is a computer operated by a "performance processor" that manages the performances executed by the performance equipment 50. The performance processing device 12 includes, for example, a portable computer and a stationary computer. The performance processing device 12 is connected to the performance equipment 50 via a network. The performance processing device 12 includes a main body (casing), a processor 30, a memory 31, an operation device 32, a display device 33, a communication device 34, etc. The processor 30 is provided inside the main body and is constituted by a central processing unit (CPU (Central Processing Unit)) in which an arithmetic unit (arithmetic circuit) and a control unit (control circuit) are integrated.
[0043] The processor 30 is communicably connected to the memory 31, the operation device 32, the display device 33, the communication device 34, etc. via a bus 35. The processor 30 has a configuration to control including other devices and circuits provided inside the main body and devices and circuits provided outside the main body.
[0044] The processor 30 executes processing based on the acquired information. The processing executed by the processor 30 includes processing to read information from the memory 31, processing of the information acquired from the music processing device 11, processing of the operation content of the operation device 32, processing of the information to be displayed on the display device 33, output of a control signal to the performance equipment 50, etc. A non-temporary application is stored in the memory 31. The processor 30 reads the non-temporary application and executes processing. Also, the information processed by the processor 30 is stored in the memory 31.
[0045] Furthermore, the memory 31 includes information for the effect processing device 12 to select effect processing based on the "inferred result of music development" obtained from the music processing device 11. This information includes labels indicating various parameters that are the characteristics of the music and information associating the effect processing with them. These pieces of information are classified and stored in the memory 31 for each type of music characteristic. Here, if the processor 30 includes an artificial intelligence unit, the processor 30 can automatically perform the classification process by the function of the artificial intelligence unit. Also, when an effect producer operates the operation device 32 to specify, for example, various parameters that are the characteristics of the music in advance for the effect content, the artificial intelligence unit of the processor 30 can also use them as data for the classification process. For example, in the case of a multimodal AI, when specifying parameters, if the descriptions of various parameters and the effect content to be used are input, a JSON file representing the degree of match for each parameter is output. The processor 30 can use the JSON file as data for the classification process.
[0046] The communication device 34 includes devices, equipment, and specifications for connecting the effect processing device 12 to the network 15 and the effect equipment 50 by at least one of a wireless communication system or a wired communication system. The communication device 34 includes a communication circuit, a cable, an antenna, a communication port, a communication connector, a communication hub, and the like.
[0047] The operation device 32 is operated when inputting information and commands, etc. to the effect processing device 12, when causing the effect processing device 12 to execute various processes, when sending various information from the effect processing device 12 to the effect equipment 50, when the effect processing device 12 obtains various information via the network 15, and so on. The operation device 32 includes at least one element among, for example, a keyboard, a touch pad, a mouse, a liquid crystal display, an organic electroluminescence display, and the like.
[0048] The display device 33 is either directly attached to the main body or connected to the main body via a cable. The display device 33 is a display visible to the "production processor", and the display includes structures such as a liquid crystal display and an organic electroluminescent display. Note that the display may be defined as a monitor. Operation buttons, operation tabs, etc. can be displayed on the screen of the display device 33. That is, the display device 33 can also serve as the operation device 32.
[0049] (4) Composition of the production equipment The production equipment 50 is arranged in the concert hall A1 or in a place visible to the people (audience) in the concert hall A1. The production equipment 50 is equipment that automatically performs an effect suitable for the live music performed in the concert hall A1 by operating according to the control signal sent from the production processing device 12. An effect suitable for the music means an effect that is visible to the people listening to the music performed in the concert hall A1 and is suitable for the characteristics of the music, an effect similar to the characteristics of the music, an effect that enlivens the music, etc.
[0050] The production equipment 50 is composed of, for example, a display, a laser light emitting device, a smoke generating device, a fireworks launching device, a water jet device, a vibration device, a lighting device, a fragrance generating device, etc. The display can display images according to the characteristics of the music, such as still images, moving images, illustrations, etc. The laser light emitting device can emit laser light of colors, directions, numbers, etc. according to the characteristics of the music. The smoke generating device can generate smoke of colors, amounts, directions, etc. according to the music. The fireworks launching device can launch fireworks of types, directions, colors, etc. according to the music. The water jet device can jet water of colors, directions, numbers, etc. according to the characteristics of the music.
[0051] The vibration device has an electric motor that applies vibrations according to the characteristics of music. The vibration device is provided in concert hall A1 on the chairs where the audience sits, on the devices distributed to the audience in concert hall A1, and on portable terminals owned by the audience, such as smartphones or tablet terminals, etc. Then, the vibration device can be operated according to the characteristics of music by an application pre-installed on the portable terminal or an application operating on a website to which the portable terminal is connected. The lighting device is provided in concert hall A1 and can adjust the color of the lighting, etc., according to the characteristics of music. The fragrance generator can spray aromatic mists according to the characteristics of music. The production facility 50 executes a production based on the control signal output from the production processing device 12.
[0052] (5) Example of overall processing executed by the system Figure 2 is a flowchart showing an example of the processing executed by system 10, that is, a music processing method. The music processing device 11 performs the process of acquiring music data in step S10. And perform machine learning on a model before or during learning to generate a learned model Also, the music processing device 11 performs the process of predicting the development of music after a predetermined time in step S20. Furthermore, the music processing device 11 provides, that is, outputs, the prediction result of the development of music to the production processing device 12 in step S30.
[0053] Based on the acquired prediction result of the development of music, the production processing device 12 selects a production corresponding to the development of music in step S40. The music processing device 11 can repeatedly execute the process of step S10 a plurality of times (many times) at a predetermined time interval for learning processing. Also, the music processing device 11 can repeatedly execute the process of step S20 a plurality of times while music is being performed live or played back in concert hall A1. Therefore, when the music processing device 11 executes the process of step S20 this time, it treats the result of the process performed in step S20 in the past as the result of the process in step S10.
[0054] (6) An example of learning control executed by the music processing device The process of step S10 performed by the music processing device 11 includes a process of analyzing music to be learned to obtain music data, a process of analyzing the music data to generate analysis data, a process of analyzing the analysis data, and a process of generating inference data based on the analysis result of the analysis data. The processes of analyzing the music data to be learned, generating the analysis data, and generating the inference data are performed to obtain the information and data necessary for the "process of inferring the development of music" performed in step S20. The meaning of the inference data will be described later.
[0055] Also, the music processing device 11 can perform the process of step S10 in either the first environment or the second environment. The first environment is an environment in which music is performed live at the concert hall A1 shown in FIG. 1, and the music data is sent to the music processing device 11 via the microphone 14 and the network 13. The second environment is an environment in which the music data of the music reproduced by the computer 62 is sent to the music processing device 11 via the network 13. For example, it includes an environment in which the music recorded on the recording medium of the computer 62 shown in FIG. 1 is reproduced by the speaker 62A. Also, it includes an environment in which music is provided on the website of the network 13, and the computer 62 acquires the music from the website and reproduces it with the speaker 62A. In the second environment, the music data of the performance reproduced by the speaker 62A of the computer 62 is returned to the computer 62 by the loopback function. The computer 62 sends the returned music data to the music processing device 11 via the network 13 without passing through the microphone 14.
[0056] The music processing device 11 acquires music (audio) via the network 13 and records, that is, stores, the acquired music in the auxiliary memory 19 in step S11. Also, the music data processing unit 24 of the music processing device 11 collaborates with the auxiliary memory 19 in step S12 to analyze and parameterize (numerify) the features included in the music data of the music. FIG. 3(A) shows a screen 61 illustrating an example of the features of the music data.
[0057] This screen 61 can be displayed on the display device 21. The characteristics of the music data have a plurality of major items 63, and each major item 63 is subdivided into minor items 64. The level of the parameters included in each minor item 64 is shown in the graph 65 on the horizontal axis. The major items 63 include the state of the music, part discrimination, melody, musical instrument, and emotion. In addition, the screen 61 also includes a pie chart 66 showing the parameter ratio of the designated genre of the music and a line graph 67 showing the time-series transition of a plurality of parameters included in the designated genre.
[0058] The state of the music includes minor items 64 such as the estimation of whether it is silent or not and the estimation of whether it is a piece of music or not. Part discrimination includes the discrimination of minor items 64 such as the prelude, interlude, postlude, part A, part B, part C, part D, part E, etc. The melody is the atmosphere conveyed from the tune of the music to human sensibility, and the melody includes minor items 64 such as high tempo, low tempo, break, conversation, narration, noise, etc. Musical instruments include minor items 64 such as the presence or absence of a human voice and the type of musical instrument. The types of musical instruments are, for example, string instruments, woodwind instruments, brass instruments, percussion instruments, synthesizers, keyboard instruments, etc. Emotions include minor items 64 such as emotionless, joy, anger, sadness, fun, relief, horror, feeling good, feeling bad, being excited (excited or exhilarated), being calm, etc. During the analysis of the music data, the level (numerical value) of the parameters included in the minor item 64 changes in the horizontal axis direction of the graph 65, that is, increases or decreases. Also, when the screen 61 is displayed on the display device 21, the parameters of each minor item 64 change in color according to the change in the level.
[0059] The technology in which the music processing device 11 analyzes music data to estimate the melody is well-known as described in, for example, Japanese Patent No. 7176133, Japanese Unexamined Patent Application Publication No. 2006-23524, Japanese Unexamined Patent Application Publication No. 2008-250113, etc., and thus a detailed description is omitted. In addition, the music processing device 11 can also estimate the melody of the music from the genre of the music. The technology for analyzing music data to estimate the genre of the music is well-known as described above.
[0060] The estimation of musical instruments performed by the music processing device 11 is the estimation of the presence or absence of a human voice and the type of musical instrument. The types of musical instruments include, for example, sub-items such as string instruments, woodwind instruments, brass instruments, percussion instruments, synthesizers, keyboard instruments, etc. Note that since the technology for estimating the types of musical instruments used by analyzing music data is well-known as described in, for example, Japanese Patent Application Laid-Open No. 2003-15684, Japanese Patent Application Laid-Open No. 2005-49859, Japanese Patent Application Laid-Open No. 2006-508390, etc., detailed explanations are omitted.
[0061] Joy, which is a human emotion estimated by the music processing device 11, can be estimated from the fact that the melody of the music is at a high tempo, the music is in a relatively high pitch range, that is, above a predetermined frequency, etc. Anger can be estimated from the fact that the music is intense, the music is powerful, the music has a heavy bass, the genre of the music is rock, etc. Sadness can be estimated from the fact that the music is at a low tempo, the type of musical instrument, etc.
[0062] Happiness can be estimated from the fact that the melody of the music is at a low tempo, the music is in a relatively high pitch range, the genre of the music is a ballad, etc. Excitement can be estimated from the fact that the volume of the music is relatively large, the music is relatively high in pitch, etc. Calmness can be estimated from the fact that the change in the frequency of the music has regularity, is not excessively high or low, the volume of the music is relatively small, etc. Also, the music processing device 11 can also estimate emotions from the genre of the music.
[0063] Note that since the technology for obtaining information regarding emotions from music data is well-known as described in, for example, Japanese Patent Application Laid-Open No. 9-230857, Japanese Patent Application Laid-Open No. 2002-366173, Japanese Patent Application Laid-Open No. 2004-61666, Japanese Patent Application Laid-Open No. 2006-23524, etc., detailed explanations are omitted. Also, the music processing device 11 can also estimate the emotions contained in the music from the genre of the music. Since the technology for estimating the genre of music by analyzing music data is well-known as described in, for example, Japanese Patent Application Laid-Open No. 2002-215195, Japanese Patent Application Laid-Open No. 2015-79110, Japanese Patent Application Laid-Open No. 2017-54121, etc., detailed explanations are omitted.
[0064] Also, the music processing device 11 further clips the music data analyzed as described above in step S12 and generates analysis data. An example of the music processing device 11 generating analysis data is shown in FIG. 3(B). The music data 60 shown in FIG. 3(B) is created corresponding to any one of a plurality of parameters included in the above-mentioned sub-items. In the music data 60, time is shown on the horizontal axis, and the level of the parameter is shown on the vertical axis. As the vertical axis moves relatively away from the starting point P10, it means that the numerical value of the parameter is relatively high. The display device 21 can display the music data 60.
[0065] In the clipping process, the music data is divided into a plurality of types of analysis layers in different time regions, for example, as shown in FIG. 3(B), into predetermined times (predetermined intervals) TM1, TM2, TM3. The predetermined time TM1 is an interval from the start time point T0 (0 seconds) to the time point T12 15 seconds later. The predetermined time TM1 includes the overall characteristics of the music data. The start time point T0 is the time point when the acquisition of the music data starts. The time point T1 is the time point when, for example, 12 seconds have elapsed from the start time point T0. The predetermined time TM2 is an interval from the time point T1 to the time point T2, and the predetermined time TM2 is, for example, 3 seconds. The predetermined time TM2 includes the most recent development of the music data. The time point T4 is the time point that is the target for estimating the characteristics of the music. After the time point TY is the future interval, and the time point T2 and the time point TY are set simultaneously. The time point T4 is the time point when, for example, 3 seconds have elapsed from the time point T2. The music data included in the predetermined time TM1 is the analysis data. The music data included in the predetermined time TM5 is the correct answer data used for the learning process.
[0066] The predetermined time TM3 is the interval from time point T2 to time point T4 and is a future interval (future time). Also, the predetermined time TM5 from the start time point T0 to time point T4 is the limit time (maximum time) for acquiring music data, and the predetermined time TM5 is, for example, 18 seconds. Therefore, after time point T4, music data is not acquired. When performing clipping processing on music data, the range B1 corresponding to the predetermined time TM5 can be set at different positions on the time axis with respect to the reference time point TX. In the example shown in FIG. 3(B), the start time point T0 is set simultaneously with the time point TX. In step S12, analysis data is generated for one or more parameters among all the parameters included in the music data.
[0067] The music processing device 11 executes the process of step S13 in cooperation with the artificial intelligence unit 28 and the auxiliary memory 19. The process that the music processing device 11 performs in step S13 is to perform a learning process by processing and analyzing various music data obtained in step S12 and generating inference data. The learning process is Of the model before or during learning machine learning, and the inference data is a learned model generated by the artificial intelligence unit 28. The inference data generated by the music processing device 11 in step S13 is used when the music processing device 11 executes the process of step S20. Since the music processing device 11 can repeatedly execute step S10, the inference data generated in step S13 is updated (replaced) with the latest learned model. When executing the process of step S13, the music processing device 11 uses a multi-layer configuration in a neural network, for example, a convolutional layer with a 5-layer configuration.
[0068] The learning process that the music processing device 11 performs in step S13 includes a process of creating a large amount of data that sets both correct data including future time and analysis data not including future time. Also, the learning process that the music processing device 11 performs in step S13 is The output data obtained by processing the analysis data with a model before or during learning, andIt includes a process of comparing with the correct data and generating a learned model by utilizing the comparison result, that is, the error in the characteristics of the music. The analysis data used in the process of step S13 is the data obtained in step S12.
[0069] In step S13, the music processing apparatus 11 generates a learned model for at least one or more of all the parameters included in the music data. Prepare the analysis data. Then, input the analysis data into a model before or during learning, and perform machine learning on the model before or during learning by comparing the output data of the model before or during learning with the correct answer data. Generate a learned model. An example of performing machine learning on a model before or during learning to generate a plurality of learned models is Refer to FIGS. 4(A), 4(B), 5(A), 5(B), 6(A), and 6(B) for explanation. In each generation example of the learned model, the analysis layer used at the predetermined times TM2 and TM3 shares at least a part of the convolutional layer but does not share only the final layer. In FIGS. 3(B), 4(A), 4(B), 5(A), 5(B), 6(A), and 6(B), the same reference numerals are assigned to the common technical matters. To explain In the generation example of the learned model shown in FIG. 4(A), the start time point T0 is set after the reference time point TX. And between the time point T1 and the time point T2, the time point T3 is set. Also, the time point T3 and the time point TY are set simultaneously. The time point T3 is the time point when the characteristics of the music data are inferred. The time point T4 is set after a predetermined time TM4 from the time point T3. The predetermined time TM4 exceeds each of the predetermined times TM2 and TM3 and is less than the predetermined time TM1.
[0070] At the time point T3, the characteristics of the music data at the time point T4 are inferred, but there is no music data at the time point T4. Therefore, in the example of FIG. 4(A), in fact, The model before or during learning The characteristics of the music data cannot be inferred. The model before or during learning is In the generation example of the learned model shown in FIG. 4(B), the reference time point TX and the start time point T0 are set simultaneously, and the time points T2, T3, and TY are set simultaneously. Also, the predetermined times TM3 and TM4 are of the same length. In the example shown in FIG. 4(B), The model before or during learning is At the time point T3, the characteristics of the music data at the time point T4 are inferred.
[0071] At the time point T3, the characteristics of the music data at the time point T4 are inferred. The model before or during learning is At the time point T3, the characteristics of the music data at the time point T4 are inferred. And output the output dataIt is done. The predetermined time TM3 is entirely in the future interval.
[0072] In the generation example of the learned model shown in FIG. 5(A), the starting point T0 is set before the reference point TX. The point in time T3 is set between the point in time T2 and the point in time T4. That is, the point in time T3 is located within the range of the predetermined time TM3. Also, the points in time T3 and TY are set simultaneously. In the example shown in FIG. 5(A), The model before or during learning is From the point in time T3, the characteristics of the music data at the point in time T4 after the predetermined time TM4 are inferred And output the output data It is done. Note that the predetermined time TM4 is less than the predetermined time TM3. The predetermined time TM3 includes a partial future interval.
[0073] In the generation example of the learned model shown in FIG. 5(B), the starting point T0 is set before the reference point TX. Also, the points in time T3, T4, and TY are set simultaneously. The predetermined time TM3 does not include a future interval. In the example shown in FIG. 5(B), The model before or during learning is At the point in time T3, the characteristics of the music data at the point in time T4 are judged in real time And output the output data It is done.
[0074] In the generation example of the learned model shown in FIG. 6(A), the starting point T0 is set before the reference point TX. The starting point T0 shown in FIG. 6(A) is before the starting point T0 shown in FIG. 5(B). Also, the points in time T3 and T4 are set simultaneously, and the points in time T3 and T4 are set before the point in time TY. For this reason, the predetermined time TM3 does not include a future interval. In the example shown in FIG. 6(A), the point in time T4 is set before the point in time TY. In that sense, The model before or during learning is It can also be seen that the characteristics of the music in the past, that is, before the point in time TY, are being inferred.
[0075] The example of FIG. 6(B) shows the case where the time axis of the music data that can be acquired is less than the predetermined time TM5. In this case, silent music data is added (inserted) before the starting point T0, and the reference point TX is set as the virtual starting point T00. By this process, the predetermined time TM5 can be secured between the starting point T00 and the point in time T4. Then, the points in time T2, T3, and TY are set simultaneously,The model before or during learning is Estimate the characteristics of the music from time point T3 to time point T4 And compare the output data of the model before or during learning with the correct answer data, Generate a learned model. In this way, using a plurality of analysis data with different positions of range B1 in the time axis direction Perform machine learning on the model before or during learning, and for each of the plurality of analysis data, a plurality of learned A model can be generated.
[0076] In addition, when the music data processing unit 24 executes the process of step S12, for example, any one or more of a spectrogram, a mel spectrogram, MFCC (Mel Frequency Cepstral Coefficient), etc. can be used.
[0077] (7) An example of the inference process executed by the music processing device A specific example in which the music processing device 11 infers the development of music in step S20 of FIG. 2 will be described. The music processing device 11 acquires "the music to be inferred" in step S21. Further, in step S22, the music processing device 11 analyzes and clips the music data of the music to be estimated, and generates analysis data (current data - estimation target data) for the music to be estimated. The process of step S22 is substantially the same as the process of step S12. Note that the analysis data generated in step S22 is different from the analysis data (past data) generated in step S12 in that there is no music data at a predetermined time TM3.
[0078] In step S23 following step S22, the music processing device 11 infers the development of the music at the inference target time point, that is, the characteristics of the music. The music processing device 11 performs the process of step S23 based on the analysis result of the music data obtained in step S22 and the inference data generated in step S13 described above. In step S23, the music processing device 11 first compares the characteristics of the music that is the target of the development estimation, that is, the analysis result of the analysis data and the inference data. Next, the music processing device 11 Based on the characteristics of the music in [context], and the speculation data among a plurality of speculation data that has the characteristics of music most similar to the characteristics of the music in the analysis data of the music to be estimated Infer the characteristics of the music at the inference target time point. The music processing device 11 outputs the inference result of the characteristics of the music obtained by executing the process of step S23 to the rendering processing device 12 in step S30.
[0079] (8) Processing examples performed by the performance processing device In step S40 of FIG. 2, the performance processing device 12 acquires the estimation result of the characteristics of the music, and based on the estimation result of the characteristics of the music, selects a performance corresponding to the characteristics of the music, and sends a control signal to the performance equipment 50 so as to execute the selected performance. The performance selected by the performance processing device 12 is recognizable by humans visually. The selection of the performance performed by the performance processing device 12 in step S40 includes displaying an image on the display of the performance equipment 50 and emitting a laser beam with the laser beam emitting device of the performance equipment 50.
[0080] Among the images displayed on the display of the performance equipment 50, the images representing joy include, for example, images including confetti, weddings, mountaintops, sunrises, etc. The images representing anger include, for example, images of volcanic eruptions, fists, flames, etc. The images representing sadness include images such as falling leaves, depopulated areas, people looking up at the sky, etc. Fun includes images such as people's smiles, goal scenes in ball games, etc. Excitement includes images such as the ascent of an airplane, the flow of clouds, etc. Calmness includes images such as a calm sea, a desert, etc. Furthermore, the images representing high tempo are exemplified by images of short-distance running in track and field competitions. The images representing low tempo are exemplified by images of marathon competitions.
[0081] In addition, the selection of the performance performed by the performance processing device 12 in step S40 includes generating a smoke screen with the smoke screen generating device of the performance equipment 50, launching fireworks with the fireworks launching device of the performance equipment 50, and injecting water with the water injection device of the performance equipment 50.
[0082] (Specific Example 2) A specific example 2 of the system 10 is shown in FIG. 4(A). The music processing device 11 may be arranged in the concert hall A1 or may be provided in an environment different from the concert hall A1. The music processing device 11 is connected to the server 42 via the network 41. The network 41 is configured in the same way as the network 13.
[0083] Server 42 is a computer having a processor 42A, a non-transitory memory 42B, a communication device, etc. Server 42 is configured to provide a website to network 41. In the memory 42B of server 42, a non-transitory application for operating the music processing device 11 to execute steps S10, S20, and S30 in FIG. 2 is stored.
[0084] In the system 10 shown in FIG. 4(A), the music processing device 11 can operate a web browser to connect to the website provided by the server 42. And the music processing device 11 can operate an application on the web browser. Also, the music processing device 11 can download and install the application from the server 42 to operate the application within the music processing device 11. The music processing device 11 can execute steps S10, S20, and S30 shown in FIG. 2 by operating the application.
[0085] Thus, the application for operating the music processing device 11 to execute the processing in FIG. 2 may be any of a native application installed and operating on the music processing device 11, a web application operating on a web browser, and a hybrid application. The hybrid application has the properties of both a native application and a web application. That is, as described above, at least one functional unit among the music data processing unit 24, the unfolding speculation unit 27, and the artificial intelligence unit 28, which functions by the processor 17 operating the application, may be configured by an application programming interface provided on the website of the network 13 and operating.
[0086] (Specific Example 3) A specific example 3 of the system 10 is shown in FIG. 4(B). The music processing device 11 may be either a computer arranged in the concert hall A1 of FIG. 1 or a server arranged in an environment different from the concert hall A1. The production processing device 12 has a plurality of user terminals 12A. The plurality of user terminals 12A are respectively provided in an environment different from the concert hall A1, for example, in the homes of the users.
[0087] The plurality of user terminals 12A are configured to be individually connected to the music processing device 11 via the network 15. The plurality of user terminals 12A are computers each having a processor, a memory, an operating device, a voice output device 43, a display device 44, a communication device, etc. individually. The voice output device 43 converts an electrical signal into voice and outputs it, such as a speaker, headphones, earphones, etc. The display device 44 has the same configuration as the display device 33.
[0088] In addition, the plurality of user terminals 12A can display images on the display devices 44 they each have. The music processing device 11 has a configuration of transmitting the audio data acquired from the concert hall A1 to the user terminals 12A via the network 15 respectively. When the user terminal 12A acquires the audio data acquired from the music processing device 11, it can output music from the voice output device 43. The user can listen to the music being performed live in the concert hall A1 from the voice output device 43 connected to the user terminal 12A.
[0089] When the music processing device 11 executes steps S10, S20, and S30 of FIG. 2, each user terminal 12A executes step S40. That is, it is selected to display an image corresponding to the development of the music on the display device 44.
[0090] (Effect of this embodiment) The music processing device 11 can assist with an effect according to the development of music after a predetermined time during a performance at the performance venue A1. Further, the music processing device 11 infers the development of music after the time point T3 by means of a convolutional neural network. For this reason, compared with the case of analyzing music data using a sequential estimation model, it is not necessary to infer all the characteristics of the music after the time point T3, and the development of the music after the time point T3 can be inferred with a smaller amount of data than in the sequential estimation model. Therefore, an increase in the amount of data stored in the auxiliary memory 19 can be suppressed, and the analysis process can be speeded up and the estimation accuracy can be improved.
[0091] The music processing device 11 can infer the characteristics of the music at the time point T4 at the time point T3, as shown in FIG. 4(B) or FIG. 6(B). Here, the predetermined time TM3 corresponds to the time lag from the time point T3 when the effect processing device 12 acquires the inference result to the time point when an effect according to the inference result of the development of the music is selected. Therefore, the development of the music after the time point T3 and the timing of the effect performed by the effect equipment 50 are more likely to be matched.
[0092] (Other explanations) In the present embodiment, the time interval from the start time point T0 to the time point T1 is not limited to 12 seconds. Further, the time interval from the start time point T0 to the time point T2 is not limited to 15 seconds. Furthermore, the predetermined time TM5 is not limited to 18 seconds. Also, an example of the technical meaning of the matters described in the present embodiment is as follows. The music processing device 11 is an example of a music processing device and a computer. The processor 17 is an example of a processor. The start time points T0 and T00 are examples of a first predetermined time point. The time point T4 is an example of a time point to be inferred. Step S10 is an example of a first process (learning stage), and step S20 is an example of a second process (inference stage). The predetermined time TM5 is an example of a total predetermined time. The range B1 is an example of a range. The predetermined time TM3 is an example of a predetermined time. The predetermined time point T3 is an example of a second predetermined time point. The performance venue A1 is an example of a concert venue. The state of the music, the part, the type of musical instrument, the human emotion expressed by the music, the melody, etc. are examples of the characteristics of the music.
[0093] The application described in this embodiment can also be regarded as an "application product". Also, the generated music includes live performances and music output from speakers. Furthermore, the operating device 20 and the display device 21 shown in FIG. 1 may not be provided. The predetermined time TM1 is not limited to 15 seconds. Also, the predetermined times TM2 and TM4 are not limited to 3 seconds each. The predetermined time TM5 can be set to, for example, 3 seconds.
[0094] This embodiment also describes the following subject matter. For example, a system having a music processing device that acquires music data and executes processing, and an effect processing device that selects an effect recognizable by a human according to the characteristics of the music, wherein the music processing device includes a data acquisition process of acquiring music data of music from a predetermined time point, a speculation process step of speculating the characteristics of the music generated at the music venue at a first time point after the predetermined time point by processing the acquired music data, and a provision process of providing the speculation result of the characteristics of the music obtained in the speculation process step to an effect processing device that selects an effect recognizable by a human according to the characteristics of the music.
[0095] Also, a music processing method executed by a music processing device that acquires music data, wherein the music processing device includes a data acquisition process of acquiring music data of music from a predetermined time point, a speculation process step of speculating the characteristics of the music generated at the music venue at a first time point after the predetermined time point by processing the acquired music data, and a provision process of providing the speculation result of the characteristics of the music obtained in the speculation process step to an effect processing device that selects an effect recognizable by a human according to the characteristics of the music.
[0096] Furthermore, there is provided a non-transitory recording medium having recorded thereon a music processing application for causing a computer that acquires music data to execute processing, the music processing application causing the computer to execute: data acquisition processing for acquiring music data of music from a predetermined time point; estimation processing for estimating characteristics of music generated at a concert venue at a first time point after the predetermined time point by processing the acquired music data; and provision processing for providing an estimation result of the characteristics of the music obtained in the estimation processing to an effect processing device that selects an effect recognizable by a human according to the characteristics of the music.
Industrial Applicability
[0097] This embodiment can be used as a music processing device that processes music data and a music processing application that causes a computer to execute processing.
Explanation of Signs
[0098] 10…System, 11…Music processing device, 12…Effect processing device, 17…Processor, A1…Concert venue, B1…Range, T0, T00…Start time point, T3, T4…Time point, TM3, TM5…Predetermined time
Claims
1. A music processing device having a processor that acquires music data and executes processing, The processor, a data acquisition process for acquiring music data of music from a first predetermined time point; an estimation process for estimating characteristics of music occurring at the music venue at an estimation target time point after the first predetermined time point by processing the acquired music data; a process of providing the estimation result of the music characteristics obtained in the estimation process step to a performance processing device that selects a performance that can be recognized by a human being according to the music characteristics; Run The inference process performed by the processor includes: A first process for processing the acquired music data; A second process that is performed after the first process and processes the acquired music data with a trained model to estimate characteristics of music occurring at the music venue at the time of estimation; Including, The first process executed by the processor includes: a process of generating a plurality of pieces of analysis data by clipping a plurality of pieces of music data that are different in position along a time axis within a total predetermined time range from the first predetermined time point to the estimation target time point; A process of comparing output data obtained by processing each of the generated multiple analysis data with a pre-learning or training model that predicts the characteristics of music generated at a music venue at the target time with correct answer data including music data from the first predetermined time to a future section after the target time, and machine learning the pre-learning or training model using an error in the music characteristics that is the comparison result, to generate multiple trained models corresponding to each of the multiple analysis data; A music processing device comprising:
2. 2. The music processing device according to claim 1, A music processing device, wherein the characteristics of the music estimated by the processor include one or more of the following: whether or not it is silent, part identification, type of instrument, melody, and human emotion expressed by the music.
3. 2. The music processing device according to claim 1, A music processing device, wherein the first process executed by the processor includes a process of generating one of the trained models by processing analysis data to which silent music data has been added for a predetermined time from the first predetermined time point with the pre-trained or trained model and machine learning the pre-trained or trained model.
4. 2. The music processing device according to claim 1, the inference process executed by the processor includes a process of inferring the characteristics of the music at a second predetermined time point that is after the first predetermined time point and before the inference target time point; The specified time from the second specified point in time to the target point in time corresponds to a time lag from the point in time when the performance processing device obtains a prediction result of the characteristics of the music to the point in time when the performance processing device selects a performance that is recognizable by humans in accordance with the characteristics of the music.
5. A non-transient music processing application for causing a computer that acquires music data to execute a process, comprising: The computer includes: a data acquisition process for acquiring music data of music from a first predetermined time point; an estimation process for estimating characteristics of music occurring at the music venue at an estimation target time point after the first predetermined time point by processing the acquired music data; a process of providing the result of the estimation of the characteristics of the music obtained by the estimation process to a performance processing device that selects a performance that can be recognized by a human being according to the characteristics of the music; Run the command, The inference process executed by the computer includes: A first process for processing the acquired music data; A second process that is performed after the first process and processes the acquired music data with a trained model to estimate characteristics of music occurring at the music venue at the time of estimation; Including, The first process to be executed by the computer includes: a process of generating a plurality of pieces of analysis data by clipping a plurality of pieces of music data that are different in position along a time axis within a total predetermined time range from the first predetermined time point to the estimation target time point; A process of comparing output data obtained by processing each of the generated multiple analysis data with a pre-learning or training model that predicts the characteristics of music generated at a music venue at the target time with correct answer data including music data from the first predetermined time to a future section after the target time, and machine learning the pre-learning or training model using an error in the music characteristics that is the comparison result, to generate multiple trained models corresponding to each of the multiple analysis data; A music processing application comprising:
Citation Information
Patent Citations
Information processing apparatus and method, and program
JP2010134790A
Information processing method
WO2019156091A1
Device, ensemble system, audio reproduction method, and program
WO2022269796A1
Music image output device, music image output method and program
JP2018136363A