Music processing equipment and music processing applications

The music processing device estimates future music characteristics to enable synchronized visual and sensory effects during performances.

JP2026076677AActive Publication Date: 2026-05-12株式会社RAW +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
株式会社RAW
Filing Date
2024-10-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing music processing systems fail to support the output of images corresponding to the development of music after a predetermined time.

Method used

A music processing device that acquires and processes music data, estimating music characteristics at a future time point and providing these estimates to an effect processing device to select human-recognizable effects.

Benefits of technology

Enables support for performances that correspond to the development of music at a future time point.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026076677000001_ABST
    Figure 2026076677000001_ABST
Patent Text Reader

Abstract

The present invention provides a music processing device capable of supporting the staging of music in a concert venue in accordance with its development from a predetermined point in time onward. [Solution] A music processing device 11 having a processor 17 that acquires and processes music data, wherein the processor 17 is configured to perform a data acquisition process that acquires music data of music from a first predetermined time point, an estimation process step that estimates the characteristics of music that will occur at a music venue at an estimated target time point after the first predetermined time point by processing the acquired music data, and a provision process that provides the estimation results of the music characteristics obtained in the estimation process step to an effect processing device 12 that selects an effect that can be recognized by a human according to the characteristics of the music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a music processing apparatus and a music processing application that process music data of music to be played.

Background Art

[0002] An example of a music processing apparatus (music image output apparatus) and a music processing application (program) that process music data of music to be played is described in Patent Document 1. The music image output apparatus described in Patent Document 1 includes a music storage unit that stores music, an output instruction reception unit that receives an output instruction of music, a music output unit that outputs music in response to the output instruction, an attribute value acquisition unit that acquires one or more attribute values based on an analysis result of music, an image acquisition unit that acquires an image using the one or more attribute values, and an image output unit that outputs the image.

[0003] Further, in Patent Document 1, a computer-readable recording medium is described as "a program for causing the computer to function as a music storage unit that stores music, an output instruction reception unit that receives an output instruction of music, a music output unit that outputs the music in response to the output instruction, an attribute value acquisition unit that acquires one or more attribute values based on an analysis result of the music, an image acquisition unit that acquires an image using the one or more attribute values, and an image output unit that outputs the image."

[0004] Furthermore, Patent Document 1 states that "the attribute value acquisition unit acquires one or more attribute values ​​based on the music analysis results. The music analysis may be, for example, a sound analysis or a lyrics analysis. It is preferable that the music analysis includes both sound and lyrics. The sound analysis may be, for example, an acquisition of sound features. Features are, for example, changes in amplitude in the sound waveform or changes in frequency components that make up the sound waveform. Features may be represented, for example, by an acoustic feature vector. An acoustic feature vector is a vector whose components are two or more features such as changes in amplitude and changes in frequency components. However, the representation format of the features is not limited."

[0005] Furthermore, Patent Document 1 states that "Lyric analysis may involve obtaining independent words from lyrics using, for example, machine learning methods such as deep learning, SVM, or decision trees, or natural language processing methods such as morphological analysis. ... Information that identifies the surface screen may include terms such as 'beach,' 'fireworks,' 'Christmas,' and 'graduation ceremony.' Also, the internal scene is a scene related to the user's inner world, and may be called a subjective scene."

[0006] Furthermore, Patent Document 1 states, "Information that identifies an internal scene is, for example, terms such as 'date,' 'lover,' 'nervous,' and 'relaxed.' An impression is the impression the user has. Information that identifies an impression is, for example, terms such as 'happy,' 'sad,' 'lonely,' and 'joyful.' ...The attribute value acquisition unit, for example, uses the learning information of the storage unit to apply the vectors of the sentences that make up the lyrics or the sentences obtained by morphological analysis of those sentences to the learning information for each term, determines whether or not each term corresponds, and acquires one or more terms that are determined to correspond as attribute values." Furthermore, Patent Document 1 states, "With this configuration, an image corresponding to the music can be output while the music is being output. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2018-136363 [Overview of the project] [Problems that the invention aims to solve]

[0008] The inventors of this application recognized a problem in the "music image output device and program" described in Patent Document 1: it is not possible to support the output of images corresponding to the development of music being played after a predetermined time.

[0009] The objective of this embodiment is to provide a music processing device and a music processing application capable of supporting performances that correspond to the development of music that occurs after a predetermined time. [Means for solving the problem]

[0010] This embodiment provides a music processing device having a processor that acquires and processes music data, wherein the processor is configured to perform a data acquisition process that acquires music data from a predetermined time point in time; an estimation process step that estimates the characteristics of music that will occur at a music venue at a first time point after the predetermined time point by processing the acquired music data; and a provision process that provides the estimation results of the music characteristics obtained in the estimation process step to an effect processing device that selects human-recognizable effects according to the characteristics of the music. [Effects of the Invention]

[0011] According to this embodiment, it is possible to support the performance in accordance with the development of the music at the first point in time. [Brief explanation of the drawing]

[0012] [Figure 1] This is a schematic diagram showing a specific example of a system including a music processing device. [Figure 2] This is a flowchart showing examples of processes executed across the entire system. [Figure 3]Figure 3(A) is a schematic diagram showing an example of a screen displayed on a display device, and Figure 3(B) is a diagram showing an example of music data clipping processing. [Figure 4] Figures 4(A) and 4(B) show examples of generating a trained model. [Figure 5] Figures 5(A) and 5(B) show examples of generating a trained model. [Figure 6] Figures 6(A) and 6(B) show examples of generating a trained model. [Figure 7] Figure 7(A) is a schematic diagram showing specific example 2 of a system including a music processing device, and Figure 7(B) is a schematic diagram showing specific example 3 of a system including a music processing device. [Modes for carrying out the invention]

[0013] (overview) The system comprises a music processing device and a performance processing device. The music processing device processes music data of music performed at a concert venue. The music processing device generates information necessary for performing performances corresponding to the music and provides it to the performance processing device in real time and automatically. In addition, several specific examples of the system and music processing device, as well as several specific examples of music processing applications, are disclosed in this embodiment with reference to the drawings.

[0014] (Specific example 1) (1) Overall structure As shown in Figure 1, System 10 includes a music processing device 11 and a performance processing device 12. Specific example 1 shown in Figure 1 is an example where the music processing device 11 and the performance processing device 12 are located in a performance venue A1. In performance venue A1, there are people (audience) who perceive live music performances through their hearing. The music processing device 11 is connected to a network 13 for communication. The music processing device 11 is also connected to the performance processing device 12 via a network 15 for communication. The microphone 14 is a device installed in the music performance venue A1. The microphone 14 acquires the sound of the music actually played in performance venue A1 and converts it into an electrical signal for output.

[0015] The concert venue A1 can be any of a concert hall, a music studio, a multipurpose hall, a public hall, a gymnasium, an outdoor theater, etc. The sound source 16 of the music played in the concert venue A1 includes humans (singers) and musical instruments. The musical instruments can be either acoustic instruments or electronic instruments. An acoustic instrument is an instrument equipped with a mechanical vibrating part. An electronic instrument is an instrument that does not have a mechanical vibrating part and uses an oscillating sound generated by an electronic circuit.

[0016] Furthermore, a computer 62 may be provided in the concert venue A1, and the music may be reproduced by the speaker 62A of the computer 62. The music data of the performance reproduced by the speaker 62A of the computer 62 is returned to the computer 62 by a loopback function, and the returned music data is sent to the music processing device 11 via the network 13 without passing through the microphone 14. The genre of the music handled by the music processing device 11 can be any of pop, rock, dance, Latin, classical, march, vocal music, Japanese traditional music, etc.

[0017] The networks 13 and 15 are each constituted by one or more communication systems, either a wireless communication system or a wired communication system. The wireless communication system includes radio wave communication, optical communication, infrared communication, radio communication, satellite communication, etc. The wired communication system includes a communication circuit, a communication cable, and an antenna. The network includes one or more networks among the Internet, an intranet, a wide area network, and an intranet. The networks 13 and 15 include short-range wireless communication. The short-range wireless communication includes, for example, wireless LAN (Wi-Fi (registered trademark)) and Bluetooth (registered trademark).

[0018] The music processing device 11 can acquire music through the network 13 and execute various processes. The music processing device 11 executes preprocessing for performing an effect with the effect equipment 50 in accordance with the music performed or reproduced at the concert hall A1. The effect processing device 12 is configured to automatically select and manage the effects to be executed by the effect equipment 50 based on the estimation results acquired from the music processing device 11. The effect equipment 50 is configured to execute an effect in real time in accordance with the music performed or reproduced at the concert hall A1. The effect content executed by the effect equipment 50 is recognizable by a human (audience) through vision, touch, smell, etc.

[0019] (2) Configuration of Music Processing Device The music processing device 11 is a computer including a main body (casing), a processor 17, a main memory 18, an auxiliary memory 19, an operation device 20, a display device 21, a communication device 22, etc. The processor 17 is provided inside the main body and is constituted by a central processing unit (CPU (Central Processing Unit)) in which an arithmetic unit (arithmetic circuit) and a control unit (control circuit) are integrated. The processor 17 is communicably connected to the main memory 18, the auxiliary memory 19, the operation device 20, the display device 21, the communication device 22, etc. via a bus 23.

[0020] The processor 17 controls including other devices and circuits provided inside the main body and devices and circuits provided outside the main body. Further, in addition to the central processing unit, the processor 17 has arithmetic processing circuits such as a digital signal processor, an application specific integrated circuit (ASIC), and a GPU.

[0021] The GPU is an abbreviation of (Graphics Processing Unit), and the GPU is a graphic controller that performs arithmetic processing necessary when performing image processing of 3D graphics, etc. Further, the GPU has a configuration for executing machine learning for constructing an acoustic model, specifically, deep learning, in the stage of processing and analyzing music data.

[0022] The processor 17 performs various processes by running non-temporary applications stored in the auxiliary memory 19. The processes performed by the processor 17 include the processes themselves, judgment, analysis, inference, control and instruction of other elements, acquisition of information and signals from other elements, and storage of information in the auxiliary memory 19.

[0023] The main memory 18 is a volatile storage device and functions as a work area and buffer area when the processor 17 performs processing. The auxiliary memory 19 is a non-volatile storage device, that is, a non-temporary storage medium. Non-temporary applications are stored in the auxiliary memory 19. These applications include programs, configuration files, data storage files, various libraries, etc.

[0024] Furthermore, the auxiliary memory 19 stores various information used by the processor 17 to perform various processes, various information as a result of the processor 17 performing various processes, and so on. The various information stored in the auxiliary memory 19 includes the information itself, data, graphs, maps, charts, and so on.

[0025] The auxiliary memory 19 has a larger capacity than the main memory 18, and operates according to input and output commands from the processor 17. The auxiliary memory 19, which is a non-temporary storage medium, is composed of, for example, a magnetic disk, an optical disk, flash memory, etc. A hard disk drive is an example of a magnetic disk. Compact discs, digital video discs, Blu-ray discs, etc. are examples of optical disks.

[0026] Flash memory is a type of semiconductor memory, and examples of flash memory include SD memory cards, USB flash drives, and solid-state drives. One or more elements included in the auxiliary memory 19 can be defined as a storage medium 19A that can be attached to and removed from the main unit.

[0027] The operating device 20 is operated by a “music processing administrator” who uses the music processing device 11. The operating device 20 is operated when switching the operation and stopping of the music processing device 11, when executing various processes on the processor 17, when displaying information on the display device 21, when acquiring signals including music data via the network 13, when sending information to the performance processing device 12, etc.

[0028] The operating device 20 includes at least one element from, for example, a keyboard, touchpad, mouse, liquid crystal display, organic electroluminescent display, etc. These elements are appropriately selected depending on whether the music processing device 11 is a portable computer or a fixed computer, and the mounting structure to the main unit is also appropriately selected. In other words, the operating device 20 has a structure that is directly attached to the main unit, or a structure that is connected to the main unit via a cable.

[0029] The display device 21 is either directly attached to the main unit or connected to the main unit via a cable. The display device 21 is a display that is visually observed by the "music processing manager," and the display includes structures such as liquid crystal displays and organic electroluminescent displays. The connection structure to the main unit is appropriately selected depending on whether the music processing device 11 is a portable computer or a fixed computer. Note that the display may also be defined as a monitor. Various information is displayed on the screen of the display device 21. The display device 21 can also display operation buttons, operation tabs, etc. on its screen. In other words, the display device 21 can also perform the functions of the operation device 20.

[0030] The communication device 22 includes devices, equipment, and standards for connecting the music processing device 11 to networks 13 and 15, respectively, by at least one of either a wireless communication system or a wired communication system. The communication device 22 includes communication circuits, cables, antennas, communication ports, communication connectors, communication hubs, etc.

[0031] The configuration of the processor 17 will be described in detail. The processor 17 is configured to function as a music data processing unit 24, an expansion prediction unit 27, and an artificial intelligence unit 28 by running non-temporary applications stored in the auxiliary memory 19.

[0032] The music data processing unit 24 is configured to perform functions such as classifying and analyzing music data acquired via the network 13, clipping music data, generating data for analysis, and analyzing the data for analysis. Specifically, the music data processing unit 24 processes music data based on the acquired music data and information stored in the auxiliary memory 19. The music data processing performed by the music data processing unit 24 includes classifying and analyzing the characteristics of the music, and specifically includes estimating one or more items from the following: the state of the music, the determination of parts, the type of instrument, the human emotions expressed by the music, and the musical style. The clipping process, generation of data for analysis, and analysis of the data for analysis performed by the music data processing unit 24 will be described later.

[0033] Furthermore, the music data processing unit 24 generates prediction data in cooperation with the artificial intelligence unit 28 based on the analysis results of the music data and the information stored in the auxiliary memory 19. The prediction data is used by the development prediction unit 27 when predicting the development of the music being played after a predetermined time, that is, the characteristics of the music.

[0034] The development prediction unit 27, in cooperation with the artificial intelligence unit 28, predicts the future of the music, that is, its development after a predetermined time, based on prediction data and information stored in the auxiliary memory 19. Details of the processing performed by the development prediction unit 27 will be described later.

[0035] The artificial intelligence unit 28, based on various information acquired by the processor 17 and various information stored in the auxiliary memory 19, collaborates with the music data processing unit 24 to analyze music data and perform machine learning, generates a trained model, and predicts the development of the music after a predetermined time.

[0036] The artificial intelligence unit 28 is equipped with transformers and language models such as neural networks, and is capable of performing generative artificial intelligence processing. The transformers include GPT (Generative Pre-trained Transformer), BERT (Bidirectional Encoder Representations from Transformers), etc. The neural network is composed of, for example, a convolutional neural network (CNN). The convolutional neural network mainly has convolutional layers, pooling layers, fully connected layers, etc.

[0037] The language model is an example of a learning model using a machine learning algorithm. Specific machine learning algorithms include nearest neighbors, naive Bayes, decision trees, support vector machines, and deep learning using neural networks. The artificial intelligence unit 28 can apply the above algorithms as appropriate.

[0038] The machine learning performed by the Artificial Intelligence Unit 28 includes three types: supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning. Supervised learning involves training the machine using training data that includes training data. Training data is a method of training the machine (processor 17) using training music data (input data) and music data that includes correct answer data (output data), forming a trained model based on the music dataset (sample data). Supervised learning includes classification and regression training models. Classification is the classification (identification) of the features of acquired music data, and regression is the estimation (prediction) of future features of acquired music data, which can also be called a time series analysis task. Unsupervised learning is a method of training using music data that does not have correct answer information attached, and the machine forms a trained model based on the regularity and similarity of the music dataset. Unsupervised learning also includes the classification (identification) of the musical features contained in the music data.

[0039] Semi-supervised learning is a type of machine learning that combines supervised and unsupervised learning. Semi-supervised learning uses both labeled and unlabeled music data to train artificial intelligence models for classification and regression tasks. Labels are tags or markers assigned to music data and contain musical features. Reinforcement learning is a method of training a machine through trial and error to maximize a set "score" (an inferred result of musical features).

[0040] The artificial intelligence unit 28 analyzes and processes various information and data stored in the auxiliary memory 19, as well as newly acquired information and data, to perform machine learning. In machine learning, for example, deep learning using a neural network is performed to generate a trained model. The generated trained model is stored in the auxiliary memory 19. The music processing unit 11 may be composed of multiple single computers or multiple computers.

[0041] If the music processing device 11 is composed of multiple computers, the music data processing unit 24 can be separated and provided on different computers, and the computer that performs clipping processing of music data and the computer that generates data for analysis can be configured separately. Alternatively, at least one of the functions of the processor 17, including the music data processing unit 24, the expansion and prediction unit 27, and the artificial intelligence unit 28, may be configured as an application programming interface that operates on a website on the network 13. For example, when music data is sent (POST) from the music processing device 11 to a specified URL (:Uniform Resource Locator) on the network 13, the music processing device 11 may be configured to receive a response from the specified URL containing a list of parameters in JSON (JavaScript Object Notation) format and details of the performance processing.

[0042] (3) Configuration of the performance processing device The performance processing unit 12 is a computer operated by a "performance processor" who manages the performances executed by the performance equipment 50. The performance processing unit 12 includes, for example, a portable computer and a fixed computer. The performance processing unit 12 is connected to the performance equipment 50 via a network. The performance processing unit 12 comprises a main unit (casing), a processor 30, memory 31, an operating device 32, a display device 33, a communication device 34, etc. The processor 30 is located inside the main unit and consists of a central processing unit (CPU) in which an arithmetic unit (arithmetic circuit) and a control device (control circuit) are integrated.

[0043] The processor 30 is connected to the memory 31, operating device 32, display device 33, communication device 34, etc. via the bus 35. The processor 30 has a configuration that comprehensively controls other devices and circuits provided inside the main unit, as well as devices and circuits provided outside the main unit.

[0044] The processor 30 executes processing based on the acquired information. The processing performed by the processor 30 includes reading information from the memory 31, processing information acquired from the music processing device 11, processing the operation content of the control device 32, processing information to be displayed on the display device 33, outputting control signals to the performance equipment 50, etc. Non-temporary applications are stored in the memory 31. The processor 30 reads non-temporary applications and executes processing. The memory 31 also stores information to be processed by the processor 30.

[0045] Furthermore, memory 31 contains information for the performance processing device 12 to select a performance process based on the "predicted music development" obtained from the music processing device 11. This information includes labels indicating various parameters that are characteristics of the music, and information associating these with performance processes. This information is classified according to the type of musical characteristic and stored in memory 31. If the processor 30 is equipped with an artificial intelligence unit, the processor 30 can automatically perform the classification process using the functions of the artificial intelligence unit. Also, if the performance processor operates the control device 32 to specify various parameters that are characteristics of the music for the performance content in advance, the artificial intelligence unit of the processor 30 can use this as data for classification processing. For example, if it is a multimodal AI, when specifying parameters, if the description of each parameter and the performance content to be used are input, a JSON file representing the degree of match for each parameter is output. The processor 30 can use the JSON file as data for classification processing.

[0046] The communication device 34 includes devices, equipment, and standards for connecting the performance processing device 12 to the network 15 and the performance equipment 50 by at least one of either a wireless communication system or a wired communication system. The communication device 34 includes communication circuits, cables, antennas, communication ports, communication connectors, communication hubs, etc.

[0047] The operating device 32 is operated when inputting information and commands to the performance processing device 12, when the performance processing device 12 executes various processes, when the performance processing device 12 sends various information to the performance equipment 50, when the performance processing device 12 acquires various information via the network 15, etc. The operating device 32 includes at least one element from among, for example, a keyboard, touchpad, mouse, liquid crystal display, organic electroluminescent display, etc.

[0048] The display device 33 is either directly attached to the main unit or connected to the main unit via a cable. The display device 33 is a display that is viewed by the "performance processor," and the display includes structures such as liquid crystal displays and organic electroluminescent displays. The display may also be defined as a monitor. The screen of the display device 33 can display operation buttons, operation tabs, etc. In other words, the display device 33 can also perform the functions of the operation device 32.

[0049] (4) Configuration of the stage equipment The performance equipment 50 is either placed in the performance venue A1 or in a location visible to people (audience) in the performance venue A1. The performance equipment 50 operates using control signals sent from the performance processing device 12, and automatically performs effects appropriate to the music being played live in the performance venue A1. Effects appropriate to the music mean effects that are visible to people listening to the music being played in the performance venue A1, and that are appropriate to the characteristics of the music, effects that are similar to the characteristics of the music, effects that enhance the music, etc.

[0050] The performance equipment 50 consists of, for example, a display, a laser beam emitter, a smoke screen generator, a fireworks launcher, a water sprayer, a vibration device, a lighting device, a fragrance generator, etc. The display can show images according to the characteristics of the music, such as still images, videos, illustrations, etc. The laser beam emitter can emit laser beams of color, direction, number, etc. according to the characteristics of the music. The smoke screen generator can generate a smoke screen of color, amount, direction, etc. according to the music. The fireworks launcher can launch fireworks of type, direction, color, etc. according to the music. The water sprayer can spray water of color, direction, number, etc. according to the characteristics of the music.

[0051] The vibration device has an electric motor that applies vibrations according to the characteristics of the music. The vibration device is installed in the chairs where the audience sits in the performance venue A1, in devices distributed to the audience in the performance venue A1, and in portable devices owned by the audience, such as smartphones or tablet devices. The vibration device can be operated according to the characteristics of the music by an application pre-installed on the portable device, or by an application that runs on a website to which the portable device is connected. The lighting device is installed in the performance venue A1 and can adjust the color of the lighting according to the characteristics of the music. The fragrance generator can spray a fragrant mist according to the characteristics of the music. The performance equipment 50 executes the performance based on control signals output from the performance processing device 12.

[0052] (5) Example of the overall process performed by the system Figure 2 is a flowchart illustrating an example of processing performed by system 10, i.e., a music processing method. In step S10, the music processing device 11 acquires and learns music data. In step S20, the music processing device 11 predicts the development of the music after a predetermined time. Furthermore, in step S30, the music processing device 11 provides, i.e., outputs, the result of the music development prediction to the performance processing device 12.

[0053] Based on the acquired prediction of the musical development, the performance processing device 12 selects a performance corresponding to the musical development in step S40. The music processing device 11 can repeat the process in step S10 multiple times at predetermined time intervals in order to perform learning. In addition, the music processing device 11 can repeat the process in step S20 multiple times while music is being played or reproduced in the performance venue A1. For this reason, when the music processing device 11 performs the process in step S20 this time, it treats the results of the process performed in step S20 in the past as the result of the process in step S10.

[0054] (6) An example of learning control performed in a music processing device The processing in step S10 performed by the music processing device 11 includes: a process of analyzing the music to be learned and acquiring music data; a process of analyzing the music data and generating data for analysis; a process of analyzing the data for analysis; and a process of generating inference data based on the analysis results of the data for analysis. The processes of analyzing the music to be learned, generating data for analysis, and generating inference data are performed to obtain the information and data necessary for the "process of inferring the development of the music" performed in step S20. The meaning of the inference data will be explained later.

[0055] Furthermore, the music processing device 11 can perform the processing in step S10 in either the first or second environment. The first environment is one in which music is performed live at the performance venue A1 shown in Figure 1, and the music data is sent to the music processing device 11 via the microphone 14 and the network 13. The second environment is one in which music data of music played by the computer 62 is sent to the music processing device 11 via the network 13. For example, this includes an environment in which music recorded on the recording medium of the computer 62 shown in Figure 1 is played back by the speaker 62A. It also includes an environment in which music is provided on a website on the network 13, and the computer 62 obtains music from the website and plays it back by the speaker 62A. In the second environment, the music data of the performance played back by the speaker 62A of the computer 62 is returned to the computer 62 by the loopback function. The computer 62 sends the returned music data to the music processing device 11 via the network 13 without going through the microphone 14.

[0056] The music processing device 11 acquires music (sound) via the network 13 and records the acquired music to the auxiliary memory 19 in step S11, i.e., stores it. In addition, the music data processing unit 24 of the music processing device 11 works in cooperation with the auxiliary memory 19 in step S12 to analyze and parameterize (quantify) the features contained in the music data of the music. Figure 3(A) shows a screen 61 that illustrates an example of the features of music data.

[0057] This screen 61 can be displayed on the display device 21. The characteristics of the music data include multiple major categories 63, each of which is subdivided into minor categories 64. The levels of the parameters included in each minor category 64 are shown on the horizontal axis graph 65. The major categories 63 include the state of the music, part identification, musical style, instruments, and emotions. The screen 61 also includes a pie chart 66 showing the parameter proportions of a specified genre of music, and a line graph 67 showing the time-series changes of multiple parameters included in the specified genre.

[0058] The state of the music includes 64 sub-items such as estimation of whether it is silent or not, estimation of whether it is a musical piece or not. Part identification includes identification of 64 sub-items such as intro, interlude, outro, A part, B part, C part, D part, E part, etc. The mood of the music is the atmosphere conveyed to human sensibilities from the tone of the music, and includes 64 sub-items such as high tempo, low tempo, break, dialogue, narration, noise, etc. Instruments include 64 sub-items such as the presence or absence of human voices and the type of instrument. The types of instruments include, for example, string instruments, woodwind instruments, brass instruments, percussion instruments, synthesizers, keyboard instruments, etc. Emotions include 64 sub-items such as emotionless, joy, anger, sadness, happiness, relief, fear, feeling good, feeling bad, excited (exhilarated or thrilled), calm, etc. During the analysis of the music data, the levels (numerical values) of the parameters included in the sub-items 64 change along the horizontal axis in graph 65, that is, increase or decrease. Furthermore, when screen 61 is displayed on the display device 21, the color of each sub-item 64 changes according to the change in level.

[0059] The technique by which the music processing device 11 analyzes music data to estimate the musical style is publicly known, as described in Japanese Patent Publication No. 7176133, Japanese Patent Publication No. 2006-23524, Japanese Patent Publication No. 2008-250113, etc., so a detailed explanation will be omitted. The music processing device 11 can also estimate the musical style from the genre of music. The technique of analyzing music data to estimate the genre of music is publicly known, as mentioned above.

[0060] The instrument estimation performed by the music processing device 11 includes the presence or absence of human voices and the estimation of the type of instrument. The type of instrument includes subcategories such as string instruments, woodwind instruments, brass instruments, percussion instruments, synthesizers, and keyboard instruments. Since the technique for estimating the type of instrument used by analyzing music data is publicly known, as described in Japanese Patent Publication No. 2003-15684, Japanese Patent Publication No. 2005-49859, Japanese Patent Publication No. 2006-508390, etc., a detailed explanation will be omitted.

[0061] The music processing device 11 estimates human emotions such as joy, which can be estimated from factors such as the music being fast-paced, the music being in a relatively high-frequency range, i.e., above a certain frequency. Anger can be estimated from factors such as the music being intense, powerful, having deep bass tones, and being in the rock genre. Sadness can be estimated from factors such as the music being slow-paced, the type of instruments used, etc.

[0062] Enjoyment can be inferred from factors such as the slow tempo of the music, the relatively high-pitched range of the music, and the genre of the music being a ballad. Excitement can be inferred from factors such as the relatively loud volume of the music and the relatively high pitch of the music. Calmness can be inferred from factors such as the regularity of the frequency changes in the music, the absence of excessively high or low pitches, and the relatively low volume of the music. Furthermore, the music processing device 11 can also infer emotions from the genre of music.

[0063] Furthermore, since the techniques for obtaining emotional information from music data are publicly known, as described in Japanese Patent Publication No. 9-230857, Japanese Patent Publication No. 2002-366173, Japanese Patent Publication No. 2004-61666, Japanese Patent Publication No. 2006-23524, etc., a detailed explanation will be omitted. In addition, the music processing device 11 can also estimate the emotions contained in the music from the genre of the music. Since the techniques for analyzing music data to estimate the genre of music are publicly known, as described in Japanese Patent Publication No. 2002-215195, Japanese Patent Publication No. 2015-79110, Japanese Patent Publication No. 2017-54121, etc., a detailed explanation will be omitted.

[0064] Furthermore, in step S12, the music processing device 11 further clips the music data analyzed as described above and generates analysis data. An example of how the music processing device 11 generates analysis data is shown in Figure 3(B). The music data 60 shown in Figure 3(B) is created to correspond to one of the multiple parameters included in the sub-items mentioned above. In the music data 60, time is shown on the horizontal axis and the level of the parameter is shown on the vertical axis. As the relative distance from the starting point P10 on the vertical axis increases, it means that the numerical value of the parameter is relatively high. The display device 21 can display the music data 60.

[0065] In the clipping process, the music data is divided into multiple types of analysis layers with different time domains, for example, predetermined time intervals TM1, TM2, and TM3, as shown in Figure 3(B). Predetermined time interval TM1 is the interval from the start time T0 (0 seconds) to time T12, 15 seconds later. Predetermined time interval TM1 includes the overall characteristics of the music data. The start time T0 is the time when the acquisition of music data begins. Time T1 is, for example, 12 seconds after the start time T0. Predetermined time interval TM2 is the interval from time T1 to time T2, and predetermined time interval TM2 is, for example, 3 seconds. Predetermined time interval TM2 includes the most recent development of the music data. Time T4 is the time when the characteristics of the music are to be estimated. From time TY onward is a future interval, and time T2 and time TY are set simultaneously. Time T4 is, for example, 3 seconds after time T2. The music data included in predetermined time interval TM1 is the data for analysis. The music data contained within the predetermined time TM5 is the ground truth data used in the learning process.

[0066] The predetermined time TM3 is the interval from time T2 to time T4, and is a future interval (future time). The predetermined time TM5 from the start time T0 to time T4 is the limit time (maximum time) during which music data can be acquired, and the predetermined time TM5 is, for example, 18 seconds. Therefore, music data is not acquired after time T4. When clipping music data, the range B1 corresponding to the predetermined time TM5 can be set at a different position on the time axis with respect to the reference time TX. In the example shown in Figure 3(B), the start time T0 is set simultaneously with time TX. In step S12, analysis data is generated for one or more parameters among all the parameters included in the music data.

[0067] The music processing unit 11 works in cooperation with the artificial intelligence unit 28 and the auxiliary memory 19 to execute the process in step S13. The process that the music processing unit 11 performs in step S13 is to process and analyze the various music data obtained in step S12 to perform learning processing and generate inference data. The learning processing is machine learning, and the inference data is a trained model generated by the artificial intelligence unit 28. The inference data generated by the music processing unit 11 in step S13 is used when the music processing unit 11 executes the process in step S20. Since the music processing unit 11 can repeatedly execute step S10, the inference data generated in step S13 is updated (replaced) with the latest trained model. When executing the process in step S13, the music processing unit 11 uses a multilayer configuration in a neural network, for example, a five-layer convolutional network.

[0068] The learning process performed by the music processing device 11 in step S13 includes creating a large amount of data consisting of both ground truth data including future time and analysis data not including future time. The learning process performed by the music processing device 11 in step S13 also includes comparing the ground truth data and the analysis data and generating a trained model using the comparison result, in other words, the error in the musical features. The analysis data used in the process of step S13 is the data obtained in step S12.

[0069] In step S13, the music processing device 11 generates a trained model for at least one parameter out of all the parameters included in the music data. Examples of trained model generation will be explained with reference to Figures 4(A), 4(B), 5(A), 5(B), 6(A), and 6(B). In each example of trained model generation, the analysis layers used for predetermined times TM2 and TM3 share at least a portion of the convolutional layers, but the final layer is not shared. Note that in Figures 3(B), 4(A), 4(B), 5(A), 5(B), 6(A), and 6(B), common technical elements are denoted by the same reference numerals.

[0070] In the example of generating a trained model shown in Figure 4(A), the starting time T0 is set after the reference time TX. Time T3 is set between time T1 and time T2. Time T3 and time TY are set simultaneously. Time T3 is the time when the features of the music data are estimated. Time T4 is set a predetermined time TM4 after time T3. The predetermined time TM4 is greater than either predetermined time TM2 or predetermined time TM3, and less than predetermined time TM1. At time T3, the features of the music data at time T4 are estimated, but there is no music data at time T4. Therefore, in the example in Figure 4(A), it is effectively impossible to estimate the features of the music data.

[0071] In the example of generating a trained model shown in Figure 4(B), the reference time TX and the start time T0 are set simultaneously, and times T2, T3, and TY are also set simultaneously. Furthermore, predetermined times TM3 and TM4 are of the same length. In the example shown in Figure 4(B), at time T3, the features of the music data at time T4 are inferred. All predetermined time intervals TM3 are future intervals.

[0072] In the example of generating a trained model shown in Figure 5(A), the starting time T0 is set before the reference time TX. Time T3 is set between time T2 and time T4. That is, time T3 is within the range of a predetermined time TM3. Also, time T3 and TY are set simultaneously. In the example shown in Figure 5(A), the characteristics of the music data at time T4 after a predetermined time TM4 are inferred from time T3. Note that the predetermined time TM4 is less than the predetermined time TM3. The predetermined time TM3 includes a portion of the future interval.

[0073] In the example of generating a trained model shown in Figure 5(B), the starting time T0 is set before the reference time TX. Also, times T3, T4, and TY are set simultaneously. The predetermined time TM3 does not include any future intervals. In the example shown in Figure 5(B), at time T3, the characteristics of the music data at time T4 are judged in real time.

[0074] In the example of generating a trained model shown in Figure 6(A), the starting time T0 is set before the reference time TX. The starting time T0 shown in Figure 6(A) is before the starting time T0 shown in Figure 5(B). Also, times T3 and T4 are set simultaneously, and times T3 and T4 are set before time TY. Therefore, the predetermined time TM3 does not include a future interval. In the example shown in Figure 6(A), time T4 is set before time TY. In that sense, it can be seen as inferring the characteristics of the music before time TY, that is, in the past.

[0075] The example in Figure 6(B) shows a case where the time axis of the acquired music data is less than a predetermined time TM5. In this case, silent music data is added (inserted) before the start time T0 to set the reference time TX as a virtual start time T00. This process ensures that the predetermined time TM5 is secured between the start time T00 and time T4. Then, time points T2, T3, and TY are set simultaneously, and the characteristics of the music from time T3 to time T4 are estimated to generate a trained model. In this way, a trained model can be generated using multiple analysis data with different positions in the range B1 along the time axis.

[0076] When the music data processing unit 24 performs the processing in step S12, it may use one or more of the following methods: a spectrogram, a Mel spectrogram, an MFCC (Mel-frequency cepstrum coefficient), etc.

[0077] (7) An example of inference processing performed by a music processing device A specific example of how the music processing device 11 infers the development of music in step S20 of Figure 2 will be explained. In step S21, the music processing device 11 acquires the "music to be inferred". In step S22, the music processing device 11 analyzes and clips the music data of the music to be estimated and generates analysis data (current data and estimated data) of the music to be estimated. The processing in step S22 is substantially the same as the processing in step S12. The difference between the analysis data generated in step S22 and the analysis data generated in step S12 (past data) is that there is no music data at a predetermined time TM3.

[0078] In step S23, following step S22, the music processing device 11 estimates the development of the music at the time of estimation, that is, the characteristics of the music. The music processing device 11 performs the processing in step S23 based on the analysis results of the music data obtained in step S22 and the estimation data generated in step S13 described above. In step S23, the music processing device 11 first compares the characteristics of the music that are the target of estimation, that is, the analysis results of the analysis data, with the estimation data. Next, the music processing device 11 estimates the characteristics of the music at the time of estimation based on the estimation data that most closely approximates the analysis data of the music that is the target of estimation. In step S30, the music processing device 11 outputs the estimation result of the music characteristics obtained by performing the processing in step S23 to the performance processing device 12.

[0079] (8) Examples of processing performed by the performance processing device In step S40 of Figure 2, the performance processing device 12 acquires the result of predicting the characteristics of the music, and based on the result of predicting the characteristics of the music, selects a performance appropriate to the characteristics of the music and sends a control signal to the performance equipment 50 to execute the selected performance. The performance selected by the performance processing device 12 is one that can be visually recognized by humans. The performance selection performed by the performance processing device 12 in step S40 includes displaying an image on the display of the performance equipment 50 and emitting laser light with the laser light emitter of the performance equipment 50.

[0080] Among the images displayed on the 50 display units, examples of images representing joy include confetti, weddings, mountain peaks, and sunrises. Examples of images representing anger include volcanic eruptions, fists, and flames. Examples of images representing sadness include fallen leaves, sparsely populated areas, and people looking up at the sky. Images representing happiness include people smiling and goals scored in ball games. Images representing excitement include airplanes ascending and clouds moving. Images representing calmness include calm seas and deserts. Furthermore, examples of images representing a fast tempo include sprint races in track and field. Examples of images representing a slow tempo include marathon races.

[0081] Furthermore, the selection of effects performed by the effects processing device 12 in step S40 includes generating a smoke screen with the smoke screen generating device of the effects equipment 50, launching fireworks with the fireworks launching device of the effects equipment 50, and spraying water with the water spraying device of the effects equipment 50.

[0082] (Specific example 2) A specific example of system 10, part 2, is shown in Figure 4(A). The music processing device 11 may be located in the performance venue A1, or it may be located in an environment different from the performance venue A1. The music processing device 11 is connected to the server 42 via the network 41. The network 41 is configured similarly to network 13.

[0083] Server 42 is a computer having a processor 42A, non-temporary memory 42B, communication equipment, etc. Server 42 is configured to provide a website to network 41. The memory 42B of server 42 stores non-temporary applications that are operated by the music processing unit 11 to perform steps S10, S20, and S30 in Figure 2.

[0084] In the system 10 shown in Figure 4(A), the music processing unit 11 can run a web browser and connect to a website provided by the server 42. The music processing unit 11 can then run applications on the web browser. The music processing unit 11 can also download and install applications from the server 42 and run those applications within the music processing unit 11. By running applications, the music processing unit 11 can execute steps S10, S20, and S30 shown in Figure 2.

[0085] Thus, the application on which the music processing device 11 operates to perform the processing shown in Figure 2 may be a native application installed and running on the music processing device 11, a web application running on a web browser, or a hybrid application. A hybrid application possesses the properties of both a native application and a web application. In other words, as mentioned above, at least one of the music data processing unit 24, the expansion prediction unit 27, and the artificial intelligence unit 28, on which the processor 17 operates the application, may be configured as an application programming interface that is provided to and operates on a website of the network 13.

[0086] (Specific example 3) A specific example of system 10, part 3, is shown in Figure 4(B). The music processing device 11 may be a computer located in the performance venue A1 shown in Figure 1, or a server located in an environment different from the performance venue A1. The performance processing device 12 has a plurality of user terminals 12A. The plurality of user terminals 12A are each located in an environment different from the performance venue A1, for example, in the user's home.

[0087] Multiple user terminals 12A are configured to connect to the music processing device 11 independently via the network 15. Each of the multiple user terminals 12A is a computer, each having its own processor, memory, operating device, audio output device 43, display device 44, communication device, etc. The audio output device 43 converts electrical signals into sound and outputs them, similar to a speaker, headphones, or earphones. The display device 44 has the same configuration as the display device 33.

[0088] Furthermore, multiple user terminals 12A can display video on their respective display devices 44. The music processing device 11 is configured to transmit audio data acquired from the performance venue A1 to each user terminal 12A via the network 15. When a user terminal 12A acquires audio data from the music processing device 11, it can output music from the audio output device 43. Users can listen to the music being performed live in the performance venue A1 through the audio output device 43 connected to their user terminal 12A.

[0089] When the music processing device 11 executes steps S10, S20, and S30 in Figure 2, each user terminal 12A executes step S40, which means selecting to display video on the display device 44 that corresponds to the progression of the music.

[0090] (Effects of this embodiment) The music processing device 11 can support the staging of the music performed in the performance venue A1 according to its development after a predetermined time. Furthermore, the music processing device 11 uses a convolutional neural network to predict the development of the music from time T3 onward. Therefore, compared to analyzing music data with a sequential estimation model, it does not need to predict all the features of the music from time T3 onward, and the development of the music from time T3 onward can be predicted with a smaller amount of data compared to a sequential estimation model. Consequently, the increase in the amount of data stored in the auxiliary memory 19 can be suppressed, and the analysis process can be sped up and the estimation accuracy can be improved.

[0091] The music processing device 11 can estimate the characteristics of the music at time T4 at time T3, as shown in Figure 4(B) or Figure 6(B). Here, the predetermined time TM3 corresponds to the time lag from "time T3 when the performance processing device 12 acquires the estimation result to the time when it selects a performance according to the estimation result of the music's development." Therefore, the development of the music from time T3 onward and the timing of the performances performed by the performance equipment 50 become easier to match.

[0092] (Other explanations) In this embodiment, the time interval from the start time T0 to time T1 is not limited to 12 seconds. Also, the time interval from the start time T0 to time T2 is not limited to 15 seconds. Furthermore, the predetermined time TM5 is not limited to 18 seconds. An example of the technical meaning of the matters described in this embodiment is as follows: Music processing device 11 is an example of a music processing device and a computer. Processor 17 is an example of a processor. Start times T0, T00 are examples of first predetermined time points. Time point T4 is an example of a time point to be estimated. Step S10 is an example of a first process (learning stage), and step S20 is an example of a second process (estimate stage). The predetermined time TM5 is an example of a total predetermined time. Range B1 is an example of a range. Predetermined time TM3 is an example of a predetermined time. Predetermined time point T3 is an example of a second predetermined time point. Performance venue A1 is an example of a music venue. The state of the music, parts, types of instruments, human emotions expressed by the music, melody, etc. are examples of characteristics of the music.

[0093] The application described in this embodiment can also be understood as an "application product." Furthermore, the music generated includes live performances and music output from speakers. Additionally, the operating device 20 and display device 21 shown in Figure 1 are not required. The predetermined time TM1 is not limited to 15 seconds. Also, the predetermined times TM2 and TM4 are not limited to 3 seconds each. The predetermined time TM5 can be set to, for example, 3 seconds.

[0094] This embodiment also describes the following subject matter. For example, a system comprising a music processing device that acquires and processes music data, and an effects processing device that selects effects that can be recognized by humans according to the characteristics of the music, wherein the music processing device is configured to perform a data acquisition process that acquires music data of music from a predetermined time, an estimation process step that estimates the characteristics of music that will occur at a music venue at a first time point after the predetermined time by processing the acquired music data, and a provision process that provides the estimation results of the music characteristics obtained in the estimation process step to the effects processing device that selects effects that can be recognized by humans according to the characteristics of the music.

[0095] The invention also describes a music processing method performed by a music processing device that acquires music data, wherein the music processing device performs a data acquisition process to acquire music data of music from a predetermined time point; an estimation process step to estimate the characteristics of music that will occur at a music venue at a first time point after the predetermined time point by processing the acquired music data; and a provision process to provide the estimation results of the musical characteristics obtained in the estimation process step to a performance processing device that selects performances that can be recognized by humans according to the characteristics of the music.

[0096] Furthermore, the recording medium contains a non-temporary music processing application that causes a computer to acquire music data to perform processing, wherein the music processing application includes a data acquisition process that acquires music data of music from a predetermined time point; an estimation process that processes the acquired music data to estimate the characteristics of music that will occur at a music venue at a first time point after the predetermined time point; and a provision process that provides the estimation results of the musical characteristics obtained in the estimation process to a performance processing device that selects performances that can be recognized by humans according to the characteristics of the music. [Industrial applicability]

[0097] This embodiment can be used as a music processing device for processing music data, and as a music processing application that causes a computer to perform processing. [Explanation of Symbols]

[0098] 10...System, 11...Music processing device, 12...Performance processing device, 17...Processor, A1...Performance venue, B1...Range, T0, T00...Start time, T3, T4...Time, TM3, TM5...Determined time

Claims

1. A music processing apparatus having a processor that acquires and processes music data, The aforementioned processor, A data acquisition process that acquires music data from a first predetermined point in time, An estimation processing step which involves processing acquired music data to estimate the characteristics of music that will occur at a music venue at an estimated target time after the first predetermined time, A providing process that provides the results of the music characteristics inferred in the aforementioned inference processing step to a performance processing device that selects a performance that can be recognized by a human according to the music characteristics, A music processing device configured to perform the following actions.

2. A music processing apparatus according to claim 1, A music processing device wherein the characteristics of the music that the processor infers include one or more of the following: whether or not it is silent, part identification, type of instrument, musical style, and the human emotions expressed by the music.

3. A music processing apparatus according to claim 1, The inference process performed by the aforementioned processor is: The first process involves processing the acquired music data, A second process is performed after the first process and processes the acquired music data, Includes, The first process performed by the processor includes a process of generating a trained model for estimating the characteristics of music at the time of estimation for each of a plurality of analysis data points that are located at different positions in the time axis within a total predetermined time range from the first predetermined time to the time of estimation.

4. A music processing apparatus according to claim 3, The music processing apparatus includes a process in which the second process performed by the processor includes a process of inferring the characteristics of music that will occur at the music venue at the time to be inferred by comparing the characteristics of music data from the first predetermined time to the second predetermined time between the first predetermined time and the time to be inferred with the trained model.

5. A music processing apparatus according to claim 3, The first process performed by the processor includes a process of generating a trained model by adding silent music data over a predetermined period of time starting from a first predetermined time.

6. A music processing apparatus according to claim 3, The inference process performed by the processor includes a process for inferring the characteristics of the music at a second predetermined time point, which is after the first predetermined time point and before the time point to be inferred. The predetermined time from the second predetermined time to the estimated target time is the time lag from the time the performance processing device acquires the estimated result of the musical characteristics to the time the performance processing device selects a performance that can be recognized by a human according to the musical characteristics, in the music processing device.

7. A non-temporary music processing application that causes a computer to perform processing on music data, To the aforementioned computer, A data acquisition process that acquires music data from a predetermined point in time, An estimation processing step that estimates the characteristics of music that will occur at a music venue at a first time point after a predetermined time point by processing the acquired music data, A providing process that provides the results of the music characteristics inferred in the aforementioned inference processing step to a performance processing device that selects a performance that can be recognized by a human according to the music characteristics, A music processing application configured to execute [this].