Acoustic analysis method, acoustic analysis system, and program
The acoustic analysis system addresses the issue of erroneous beat estimation and user burden by enabling user-driven beat adjustments, using a deep neural network and Hidden Semi-Markov Model to align beats with user intent.
Patent Information
- Application Number
- JP2025179286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-08
AI Technical Summary
Conventional techniques for estimating beats in music may erroneously identify backbeats as beats or fail to match the user's intended tempo, and adjusting individual beats throughout a piece of music is burdensome.
An acoustic analysis system that estimates multiple beats, allows user input to change beat positions, and updates the beat positions based on user instructions, utilizing a deep neural network and Hidden Semi-Markov Model to refine beat estimation.
The system accurately aligns beat positions with user intent, reducing user burden by allowing intuitive adjustment and improving beat estimation accuracy.
Smart Images

Figure 2026002982000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to techniques for analyzing acoustic signals. [Background technology]
[0002] Analysis techniques have been proposed for estimating beats of a piece of music by analyzing audio signals that represent the sounds of the music being played. For example, Patent Document 1 discloses a technique for estimating beats of a piece of music using a probabilistic model such as a hidden Markov model. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-114361 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional techniques for estimating beats in music may erroneously estimate, for example, a backbeat as a beat, or may erroneously estimate a beat corresponding to a tempo twice the original tempo of the music. Furthermore, the estimated beat may not match the user's intention, such as when a backbeat is estimated when the user is expecting a downbeat. Considering these circumstances, it is important to have a configuration that allows the user to change the positions on the time axis of multiple beats estimated from an audio signal. However, there is a problem in that the task of changing individual beats throughout a piece of music to desired times is excessively burdensome. In consideration of these circumstances, one aspect of the present disclosure aims to obtain a time series of beats that matches the user's intention, while reducing the burden on the user of instructing changes to the positions of each beat. [Means for solving the problem]
[0005] In order to solve the above problems, an acoustic analysis system according to one aspect of the present disclosure estimates multiple beats of a piece of music by analyzing an acoustic signal representing the sound of the piece of music being played, accepts instructions from a user to change the positions of some of the multiple beats, and updates the positions of the multiple beats in accordance with the instructions from the user.
[0006] An acoustic analysis system according to one aspect of the present disclosure includes an analysis processing unit that estimates multiple beats of a piece of music by analyzing an acoustic signal that represents the sound of the piece of music being played, an instruction receiving unit that receives instructions from a user to change the positions of some of the multiple beats, and a beat update unit that updates the positions of the multiple beats in accordance with the instructions from the user.
[0007] A program according to one aspect of the present disclosure causes a computer system to function as an analysis processing unit that estimates multiple beats of a piece of music by analyzing an audio signal representing the sound of the piece of music being played, an instruction receiving unit that receives instructions from a user to change the positions of some of the multiple beats, and a beat update unit that updates the positions of the multiple beats in response to instructions from the user. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram illustrating the configuration of an acoustic analysis system according to a first embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an example of the functional configuration of an acoustic analysis system. [Figure 3] FIG. 10 is an explanatory diagram of an operation in which a feature extraction unit generates feature data. [Figure 4] FIG. 1 is a block diagram illustrating the configuration of an estimation model. [Figure 5] FIG. 1 is an explanatory diagram of machine learning for establishing an estimation model. [Figure 6] 10 is a flowchart illustrating a specific procedure of a probability calculation process. [Figure 7] FIG. 2 is an explanatory diagram of a state transition model. [Figure 8] FIG. 10 is an explanatory diagram of a beat point estimation process. [Figure 9] 10 is a flowchart illustrating a specific procedure of beat point estimation processing. [Figure 10] FIG. 10 is a schematic diagram of an analysis screen. [Figure 11] FIG. 10 is an explanatory diagram of an estimation model update process. [Figure 12] 10 is a flowchart illustrating a specific procedure of an estimation model update process. [Figure 13] 4 is a flowchart illustrating a specific procedure of a process executed by a control device. [Figure 14] 10 is a flowchart illustrating a specific procedure of an initial analysis process. [Figure 15] 10 is a flowchart illustrating a specific procedure of a beat point updating process. [Figure 16] FIG. 10 is a block diagram illustrating an example of the functional configuration of an acoustic analysis system according to a second embodiment. [Figure 17] FIG. 10 is a schematic diagram of an analysis screen in the second embodiment. [Figure 18] FIG. 1 is an explanatory diagram of an estimated tempo curve, a maximum tempo curve, and an initial tempo curve. [Figure 19] 10 is a flowchart illustrating a specific procedure of beat point estimation processing in the second embodiment. [Figure 20] FIG. 11 is an explanatory diagram of a process for generating output data in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] A: First embodiment FIG. 1 is a block diagram illustrating the configuration of an acoustic analysis system 100 according to a first embodiment. The acoustic analysis system 100 is a computer system that estimates multiple beats of a piece of music by analyzing an audio signal A that represents the performance sounds of the piece of music. The acoustic analysis system 100 includes a control device 11, a storage device 12, a display device 13, an operation device 14, and a sound emission device 15. The acoustic analysis system 100 is realized, for example, by a portable information device such as a smartphone or tablet terminal, or a portable or stationary information device such as a personal computer. The acoustic analysis system 100 can be realized as a single device, or as multiple devices configured separately from each other.
[0010] The control device 11 is composed of one or more processors that control each element of the acoustic analysis system 100. For example, the control device 11 is composed of one or more types of processors, such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).
[0011] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured from a known storage medium such as a magnetic storage medium or a semiconductor storage medium, or a combination of multiple types of storage media. Note that the storage device 12 may be a portable storage medium that is detachable from the acoustic analysis system 100, or a storage medium (e.g., cloud storage) that the control device 11 can write to or read from via a communication network such as the Internet.
[0012] The storage device 12 stores the audio signal A. The audio signal A is a series of samples representing the waveform of the performance sound of a piece of music. Specifically, the audio signal A represents at least one of the instrument sounds and the vocal sounds of the piece of music. The audio signal A may have any data format. The audio signal A may be supplied to the acoustic analysis system 100 from a signal supply device separate from the acoustic analysis system 100. The signal supply device is, for example, a playback device that supplies the audio signal A recorded on a recording medium to the acoustic analysis system 100, or a communication device that supplies the audio signal A received from a distribution device (not shown) via a communication network to the acoustic analysis system 100.
[0013] The display device 13 displays images under the control of the control device 11. For example, various display panels such as a liquid crystal display panel or an organic EL (Electroluminescence) panel are used as the display device 13. Note that the display device 13, which is separate from the acoustic analysis system 100, may be connected to the acoustic analysis system 100 by wire or wirelessly. The operation device 14 is an input device that accepts instructions from a user. The operation device 14 is, for example, an operator operated by the user or a touch panel that detects contact by the user.
[0014] The sound emitting device 15 reproduces sound under the control of the control device 11. For example, a speaker or a headphone is used as the sound emitting device 15. Note that the sound emitting device 15, which is separate from the acoustic analysis system 100, may be connected to the acoustic analysis system 100 by wire or wirelessly.
[0015] 2 is a block diagram illustrating an example of the functional configuration of the acoustic analysis system 100. The control device 11 executes a program stored in the storage device 12 to realize multiple functions for processing the acoustic signal A (an analysis processing unit 20, a display control unit 24, a playback control unit 25, an instruction receiving unit 26, and an estimation model updating unit 27).
[0016] The analysis processing unit 20 estimates multiple beats in a piece of music by analyzing the audio signal A. Specifically, the analysis processing unit 20 generates beat data B from the audio signal A. The beat data B is data that represents each beat in a piece of music. Specifically, the beat data B is time-series data that specifies the time of each beat in a piece of music. For example, the beat data B specifies the time of each beat relative to the start point of the audio signal A. The analysis processing unit 20 of the first embodiment includes a feature extraction unit 21, a probability calculation unit 22, and an estimation processing unit 23.
[0017] [Feature Extraction Unit 21] FIG. 3 is an explanatory diagram of the operation of the feature extraction unit 21. The feature extraction unit 21 generates a feature f[m] (m=1 to M) of the audio signal A for each of M time points (hereinafter referred to as "analysis time points") t[m] on the time axis. Each analysis time point t[m] is a time point set on the time axis at a predetermined interval. The feature f[m] is an index representing the acoustic feature of the audio signal A. Specifically, a feature f[m] that tends to fluctuate significantly before and after a beat is used. For example, information on the intensity of the audio signal A, such as volume and amplitude, is exemplified as the feature f[m]. Information on the frequency characteristics (timbre) of the audio signal A, such as MFCC (Mel-Frequency Cepstrum Coefficients), MSLS (Mel-Scale Log Spectrum), or Constant-Q Transform (CQT), is also used as the feature f[m]. However, the type of feature f[m] is not limited to the above examples. Furthermore, the feature f[m] may be a combination of multiple types of information related to the audio signal A.
[0018] The feature extraction unit 21 generates feature data F[m] for each analysis time point t[m]. The feature data F[m] corresponding to an arbitrary analysis time point t[m] is a time series of multiple feature amounts f[m] within a period U (hereinafter referred to as a "unit period") that includes the analysis time point t[m]. FIG. 3 illustrates an example in which one unit period U includes five analysis time points t[m-2] to t[m+2], with the mth analysis time point t[m] at the center. Therefore, the feature data F[m] is a time series of five feature amounts f[m-2] to f[m+2] within the unit period U. Note that the unit period U may include only one analysis time point [m]. In other words, the feature data F[m] may be composed of only one feature amount f[m]. As can be understood from the above explanation, the feature extraction unit 21 generates feature data F[m] including the feature amount f[m] of the audio signal A for each analysis time point t[m].
[0019] [Probability Calculation Section 22] 2 generates output data O[m] from feature data F[m], which indicates the probability P[m] that each analysis time point t[m] corresponds to a beat point in the music. The generation of output data O[m] is repeated for each analysis time point t[m]. The larger the probability P[m], the higher the likelihood that analysis time point t[m] corresponds to a beat point. The probability calculation unit 22 uses an estimation model 50 to generate the output data O[m].
[0020] There is a correlation between the feature data F[m] at each analysis time point t[m] of the audio signal A and the likelihood that the analysis time point t[m] corresponds to a beat. The estimation model 50 is a statistical model that has learned this correlation. Specifically, the estimation model 50 is a trained model that has learned the relationship between the feature data F[m] and the output data O[m] through machine learning.
[0021] The estimation model 50 is configured, for example, by a deep neural network (DNN). The estimation model 50 is realized by a combination of a program that causes the control device 11 to execute a calculation to generate output data O[m] from feature data F[m] and a plurality of variables (specifically, weights and biases) that are applied to the calculation. The program and the plurality of variables that realize the estimation model 50 are stored in the storage device 12. The numerical values of each of the plurality of variables that define the estimation model 50 are set in advance by machine learning.
[0022] 4 is a block diagram illustrating a specific configuration of the estimation model 50. The estimation model 50 is configured as a convolutional neural network including an input layer 51, multiple intermediate layers 52 (52a, 52b), and an output layer 53. Multiple feature amounts f[m-2] to f[m+2] included in one piece of feature data F[m] are input to the input layer 51 in parallel.
[0023] The multiple intermediate layers 52 are hidden layers located between the input layer 51 and the output layer 53. The multiple intermediate layers 52 include multiple intermediate layers 52a and multiple intermediate layers 52b. The multiple intermediate layers 52a are located between the input layer 51 and the multiple intermediate layers 52b. Each intermediate layer 52a is configured, for example, by combining a convolutional layer and a pooling layer. Each intermediate layer 52b is a fully connected layer that uses, for example, ReLU as an activation function. The output layer 53 outputs output data O[m].
[0024] The estimation model 50 is divided into a first portion 50a and a second portion 50b. The first portion 50a is the input side portion of the estimation model 50. Specifically, the first portion 50a is the first half portion consisting of an input layer 51 and multiple intermediate layers 52a. The second portion 50b is the output side portion of the estimation model 50. Specifically, the second portion 50b is the second half portion consisting of multiple intermediate layers 52b and an output layer 53. The first portion 50a is a portion that generates intermediate data D[m] according to feature data F[m]. The intermediate data D[m] is data that represents the features of the feature data F[m]. Specifically, the intermediate data D[m] is data that represents the features that contribute to outputting statistically valid output data O[m] for the feature data F[m]. The second portion 50b is a portion that generates output data O[m] according to the intermediate data D[m].
[0025] 5 is an explanatory diagram of machine learning that establishes the estimation model 50. For example, the estimation model 50 is established through machine learning by a machine learning system 200 that is separate from the acoustic analysis system 100, and the estimation model 50 is provided to the acoustic analysis system 100. For example, the estimation model 50 is transmitted from the machine learning system 200 to the acoustic analysis system 100.
[0026] A plurality of pieces of training data Z are used for machine learning of the estimation model 50. Each of the plurality of pieces of training data Z is composed of a combination of training feature data Ft and training output data Ot. The feature data Ft represents a feature amount at a specific time point in the acoustic signal A prepared for training. Specifically, like the above-mentioned feature data F[m], the feature data Ft is composed of a time series of a plurality of feature amounts corresponding to different time points on the time axis. The training output data Ot corresponding to a specific time point is data (i.e., a correct value) representing the probability that the time point corresponds to a beat point in the music. A plurality of pieces of training data Z are prepared for a large number of known music pieces.
[0027] The machine learning system 200 calculates an error function that represents the error between output data O[m] output by an initial or provisional model (hereinafter referred to as the "provisional model") 59 when feature data Ft of each training data Z is input, and output data Ot of the training data Z. Then, the machine learning system 200 updates multiple variables of the provisional model 59 so as to reduce the error function. The provisional model 59 at the time when the above process is repeated for each of the multiple training data Z is determined as the estimation model 50.
[0028] Therefore, the estimation model 50 outputs statistically valid output data O[m] for unknown feature data F[m] based on the underlying relationship between the feature data Ft and the output data Ot in the multiple pieces of training data Z. In other words, the estimation model 50 is a trained model that has learned the relationship between the training feature data Ft corresponding to each time point on the time axis and the training output data Ot that represents the probability that that time point corresponds to a beat. The probability calculation unit 22 inputs the feature data F[m] for each analysis time point t[m] to the estimation model 50 established by the above procedure, thereby generating output data O[m] that represents the probability P[m] that the analysis time point t[m] corresponds to a beat.
[0029] 6 is a flowchart illustrating a specific procedure of the process (hereinafter referred to as "probability calculation process") Sa executed by the probability calculation unit 22. The control device 11 functions as the probability calculation unit 22 to execute the probability calculation process Sa.
[0030] When the probability calculation process Sa is started, the probability calculation unit 22 inputs the feature data F[m] corresponding to the analysis time point t[m] to the estimation model 50 (Sa1). The probability calculation unit 22 acquires the intermediate data D[m] output from the first part 50a of the estimation model 50 and stores the intermediate data D[m] in the storage device 12 (Sa2). The probability calculation unit 22 also acquires the output data O[m] output from the estimation model 50 (second part 50b) and stores the output data O[m] in the storage device 12 (Sa3).
[0031] The probability calculation unit 22 determines whether or not the above processing has been performed for M analysis time points t[1] to t[M] in the music piece (Sa4). If the determination result is negative (Sa4: NO), the probability calculation unit 22 generates intermediate data D[m] and output data O[m] for the unprocessed analysis time point t[m] (Sa1 to Sa3). If the processing has been performed for M analysis time points t[1] to t[M] (Sa4: YES), the probability calculation unit 22 ends the probability calculation processing Sa. As can be understood from the above explanation, as a result of the probability calculation processing Sa, M pieces of intermediate data D[1] to D[M] corresponding to different analysis time points t[m] and M pieces of output data O[1] to O[M] corresponding to different analysis time points t[m] are saved in the storage device 12.
[0032] [Estimation processing unit 23] 2 estimates multiple beat positions in a piece of music from M pieces of output data O[m] calculated for different analysis times t[m] by the probability calculation unit 22. Specifically, as described above, the estimation processing unit 23 generates beat position data B that indicates the time of each beat position in the piece of music. A state transition model 60 is used to generate the beat position data B by the probability calculation unit 22.
[0033] 7 is a block diagram illustrating the configuration of a state transition model 60. The state transition model 60 is a statistical model composed of a plurality (N) of states Q. Specifically, the state transition model 60 is composed of a Hidden Semi-Markov Model (HSMM), and a plurality of beats are estimated by the Viterbi algorithm, which is an example of dynamic programming.
[0034] FIG. 7 illustrates beats on a time axis. The time length of the interval δ between two successive beats on the time axis (hereinafter referred to as the "beat interval") is a variable value according to the tempo of the music. Specifically, the faster the tempo, the shorter the beat interval δ. Multiple time points (hereinafter referred to as "elapsed points") Y[j] are set within the beat interval δ. Each elapsed point Y[i] (i = 1 to 4) is set on the time axis based on the beat. Specifically, elapsed point Y[0] is the time point corresponding to a beat (the beginning of a beat), and elapsed points Y[1] to Y[4] are each time points that equally divide the beat interval δ. Elapsed point Y[3] is located after elapsed point Y[4], elapsed point Y[2] is located after elapsed point Y[3], and elapsed point Y[1] is located after elapsed point Y[2]. Elapsed point Y[0] corresponds to the end point (start point or end point) of the beat interval δ. The time length from each beat point (elapsed point Y[0]) to each elapsed point Y can also be expressed as the phase based on the beat point. For example, time progresses in the order of elapsed point Y[4] → elapsed point Y[3] → elapsed point Y[2] → elapsed point Y[1], and reaches elapsed point Y[0] (beat point) after elapsed point Y[1].
[0035] Each of the N states Q of the state transition model 60 corresponds to one of a plurality of tempos X[i] (i = 1, 2, 3, ...). Specifically, the N states Q correspond to a different combination of each of the plurality of tempos X[i] and each of the plurality of elapsed points Y[0] to Y[4]. That is, for each tempo X[i], there are five time series of states Q corresponding to different elapsed points Y[j]. In the following description, a state Q corresponding to a combination of a tempo X[i] and a elapsed point Y[j] may be referred to as a "state Q[i,j]." On the other hand, when the distinction between the tempo X[i] and the elapsed point Y[j] is not particularly important, it will be referred to simply as a "state Q." Note that the distinction between states Q according to the elapsed point Y[j] may be omitted. That is, a configuration in which each of the plurality of states Q corresponds to a different tempo X[i] is also conceivable. In a form in which the progress point Y[j] is not distinguished, for example, a Hidden Markov Model (HMM) is used as the state transition model 60.
[0036] In the first embodiment, it is assumed that the tempo X changes only at beat points on the time axis (i.e., at the elapsed point Y[0]). Under the above assumption, the state Q[i,j] corresponding to each elapsed point Y[j] other than the elapsed point Y[0] transitions only to the state Q[i,j-1] corresponding to the immediately succeeding elapsed point Y[j-1]. For example, the state Q[i,4] transitions to the state Q[i,3], the state Q[i,3] transitions to the state Q[i,2], and the state Q[i,2] transitions to the state Q[i,1]. On the other hand, the state Q[i,0] corresponding to the beat point transitions from multiple states Q[i,1] (Q[1,1], Q[2,1], Q[3,1], ...) corresponding to different tempos X[i].
[0037] Fig. 8 is an explanatory diagram of a process Sb (hereinafter referred to as "beat point estimation process") in which the estimation processing unit 23 estimates multiple beat points in a piece of music using the state transition model 60. Fig. 9 is a flowchart illustrating a specific procedure of the beat point estimation process Sb. The control device 11 functions as the estimation processing unit 23 to execute the beat point estimation process Sb.
[0038] When the beat point estimation process Sb starts, the estimation processing unit 23 calculates an observation likelihood Λ[m] for each of M analysis time points t[1] to t[M] (Sb1). The observation likelihood Λ[m] for each analysis time point t[m] is set to a numerical value corresponding to the probability P[m] represented by the output data O[m] for that analysis time point t[m]. For example, the observation likelihood Λ[m] is set to the probability P[m] represented by the output data O[m] or a numerical value calculated by a predetermined calculation for the probability P[m].
[0039] The estimation processing unit 23 calculates (Sb2) a path p[i,j] and a likelihood λ[i,j] for each analysis time point t[m] for each state Q[i,j] of the state transition model 60. The path p[i,j] is a path from another state Q to the state Q[i,j], and the likelihood λ[i,j] is an index of the probability that the state Q[i,j] will be observed.
[0040] As mentioned above, only one-way transitions occur among the multiple states Q[i,0] to Q[i,4] corresponding to any given tempo X[i]. Therefore, as can be seen from Figure 8, the only path p[1,1] that reaches state Q[1,1] corresponding to tempo X[1] and passing point Y[1] at analysis time t[m] is the path p from state Q[1,2] corresponding to the given tempo X[1] and the immediately preceding passing point Y[2]. Furthermore, the likelihood λ[1,1] of state Q[1,1] at analysis time t[m] is set to the likelihood corresponding to time t1, which is the time length d[1] corresponding to the given tempo X[1]. Specifically, the likelihood λ[1,1] of state Q[1,1] is calculated by interpolation (e.g., linear interpolation) between the observation likelihood Λ[mA] at the analysis time t[mA] immediately before time t1 and the observation likelihood Λ[mB] at the analysis time t[mB] immediately after time t1.
[0041] On the other hand, at the elapsed point Y[0], the tempo X[i] may change. Therefore, as can be seen from FIG. 8 , for example, state Q[1,0], which corresponds to the tempo X[1] and the elapsed point Y[0], is reached by a separate path p from each of the multiple states Q[i,1] corresponding to different tempos X[i]. For example, state Q[1,0] is reached by path p1 from state Q[1,1], which corresponds to the combination of the tempo X[1] and the immediately preceding elapsed point Y[1], as well as path p2 from state Q[2,1], which corresponds to the combination of the tempo X[2] and the immediately preceding elapsed point Y[1]. The likelihood λ1 for path p1 from state Q[1,1] to state Q[1,0] is calculated by interpolating (e.g., linearly interpolating) the observation likelihood Λ[mA] at the analysis time t[mA] immediately before time t1 and the observation likelihood Λ[mB] at the analysis time t[mB] immediately after time t1, as in the previous example. Furthermore, the likelihood λ2 for path p2 from state Q[2,1] to state Q[1,0] is set to the likelihood at time t2, which is the time length d[2] corresponding to the tempo X[2] of state Q[2,1], before the analysis time t[m]. Specifically, the likelihood λ2 is calculated by interpolating (for example, linearly interpolating) the observation likelihood Λ[mC] at the analysis time t[mC] immediately before time t2 and the observation likelihood Λ[mA] at the analysis time t[mA] immediately after time t2. The estimation processor 23 selects the maximum value of the likelihoods λ(λ1, λ2, ...) calculated for different tempos X[i] as the likelihood λ[1,0] of state Q[1,0] at analysis time t[m], and determines the path p corresponding to the likelihood λ[1,0] among the paths p(p1, p2, ...) reaching state Q[1,0] as the path p[1,0] to state Q[1,0]. Through the above procedure, the process of calculating the path p[i,j] and likelihood λ[i,j] for each of N states Q is performed at each analysis time t[m] along the forward direction of the time axis. That is, the path p[i,j] and likelihood λ[i,j] of each state Q are calculated for each of M analysis times t[1] to t[M].
[0042] The estimation processing unit 23 generates a time series of M states Q corresponding to different analysis time points t[m] (hereinafter referred to as a "state series") (Sb3). Specifically, the estimation processing unit 23 connects paths p[i,j] in order along the reverse direction of the time axis, starting from the state Q[i,j] corresponding to the maximum value of the N likelihoods λ[i,j] calculated for the last analysis time point t[M] of the song, and generates a state series from M states Q located on the series of connected paths (i.e., the maximum likelihood path). In other words, a series in which states Q with the largest likelihood λ[i,j] among the N states Q are arranged for each analysis time point t[m] is generated as the state series.
[0043] The estimation processing unit 23 estimates each analysis time point t[m] at which a state Q corresponding to the elapsed point Y[0] is observed among the M states Q constituting the state sequence as a beat, and generates beat data B specifying the time of each beat (Sb4). As can be understood from the above explanation, the analysis time point t[m] at which the probability P[m] represented by the output data O[m] is high and the tempo transitions naturally to the ear is estimated as a beat in the music piece.
[0044] As described above, in the first embodiment, feature data F[m] for each analysis time point t[m] is input to the estimation model 50 to generate output data O[m] for each analysis time point t[m], and multiple beats are estimated from the output data O[m]. Therefore, it is possible to generate output data O[m] that is statistically valid for unknown feature data F[m] based on the underlying relationship between the learning feature data Ft and the learning output data Ot. A specific example of the configuration of the analysis processing unit 20 has been described above.
[0045] The display control unit 24 in Fig. 2 causes the display device 13 to display an image. Specifically, the display control unit 24 causes the display device 13 to display an analysis screen 70 in Fig. 10. The analysis screen 70 is an image showing the results of the analysis of the acoustic signal A by the analysis processing unit 20.
[0046] The analysis screen 70 includes a first area 71 and a second area 72. The first area 71 displays a waveform 711 of the audio signal A. The second area 72 displays the results of an analysis of a portion of the audio signal A that is specified in the first area 71 (hereinafter referred to as the "specified period") 712. The second area 72 includes a waveform area 73, a probability area 74, and a beat area 75.
[0047] A common time axis is set for the waveform area 73, the probability area 74, and the beat point area 75. The waveform area 73 displays a waveform 731 of the audio signal A within a specified period 712 and onsets 732 in the audio signal A. The probability area 74 displays a time series 741 of the probability P[m] represented by the output data O[m] at each analysis time point t[m]. Note that the time series 741 of the probability P[m] represented by the output data O[m] may be displayed in the waveform area 73 superimposed on the waveform 731 of the audio signal A.
[0048] The beat point area 75 displays multiple beat points in the music piece estimated by analyzing the audio signal A. Specifically, a time series of multiple beat images 751 corresponding to different beat points in the music piece is displayed in the beat point area 75. Beat images 751 corresponding to one or more beat points (hereinafter referred to as "candidate correction points") that satisfy predetermined conditions among the multiple beat points in the music piece are highlighted in a display mode separate from the other beat images 751. Candidate correction points are beat points that the user is likely to instruct to change.
[0049] 2 controls the reproduction of sound by the sound emitting device 15. Specifically, the reproduction control unit 25 causes the sound emitting device 15 to reproduce the performance sound represented by the sound signal A. In parallel with the reproduction of the sound signal A, the reproduction control unit 25 reproduces a predetermined notification sound at a time point corresponding to each of the multiple beat points. Furthermore, the display control unit 24 highlights one beat image 751, among the multiple beat images 751 in the beat point area 75, that corresponds to the time point being reproduced by the sound emitting device 15, in a display mode separate from the other beat images 751 in the beat point area 75. That is, in parallel with the reproduction of the sound signal A, each of the multiple beat images 751 is highlighted in chronological order.
[0050] In the process of estimating multiple beats in a piece of music from audio signal A, for example, a backbeat of the piece of music may be erroneously estimated as a beat. Furthermore, the estimated beat may not match the user's intention, such as when a backbeat of the piece of music is estimated when the user is expecting a downbeat to be estimated. By operating the operation device 14, the user can instruct the change of the position on the time axis of any of the multiple beats in the piece of music. Specifically, the user moves one of the multiple beat images 751 in the beat area 75 along the time axis to instruct the change of the position of the beat corresponding to that beat image 751. For example, the user instructs the change of the position of a correction candidate point among the multiple beats.
[0051] 2 receives from the user an instruction to change the positions of some of the beats in a piece of music (hereinafter referred to as a "change instruction"). In the following explanation, it is assumed that the instruction receiving unit 26 receives an instruction to move one beat on the time axis from analysis time t[m1] to analysis time t[m2] (m1, m2 = 1 to M, m1 ≠ m2). Analysis time t[m1] is the beat initially estimated by the analysis processing unit 20 (i.e., the beat before the change due to the change instruction), and analysis time t[m2] is the beat after the change due to the change instruction from the user.
[0052] 2 updates the estimation model 50 in response to a change instruction from a user. Specifically, the estimation model update unit 27 updates the estimation model 50 so that the change in beat point according to the change instruction is reflected in the estimation of multiple beat points throughout the entire piece of music.
[0053] 11 is an explanatory diagram of the process Sc (hereinafter referred to as the "estimation model update process") in which the estimation model update unit 27 updates the estimation model 50. The estimation model update process Sc is a process (additional learning) in which the estimation model 50 trained by the machine learning system 200 is updated so as to reflect a change instruction from the user.
[0054] In the estimation model update process Sc, an adaptation block 55 is added between the first part 50a and the second part 50b of the estimation model 50. The adaptation block 55 is configured, for example, with an attention block whose activation function is initialized to an identity function. Therefore, the initial adaptation block 55 supplies the intermediate data D[m] output from the first part 50a to the second part 50b without changing it.
[0055] The estimation model update unit 27 sequentially inputs feature data F[m1] at analysis time point t[m1], where the pre-change beat point is located, and feature data F[m2] at analysis time point t[m2], where the post-change beat point is located, to the first portion 50a (input layer 51). The first portion 50a generates intermediate data D[m1] corresponding to the feature data F[m1] and intermediate data D[m2] corresponding to the feature data F[m2]. The intermediate data D[m1] and the intermediate data D[m2] are sequentially input to the adaptation block 55.
[0056] Furthermore, the estimation model update unit 27 sequentially supplies each of the M pieces of intermediate data D[1] to D[M] calculated in the immediately preceding probability calculation process Sa (Sa2) to the adaptation block 55. That is, intermediate data D[m] (D[m1], D[m2]) corresponding to some analysis time points t[m] related to the change instruction among the M analysis time points t[1] to t[M] in the music, and each of the M pieces of intermediate data D[1] to D[M] throughout the music are input to the adaptation block 55. The adaptation block 55 calculates the similarity between the intermediate data D[m] (D[m1], D[m2]) corresponding to the analysis time point t[m] related to the change instruction and the intermediate data D[m] supplied from the estimation model update unit 27.
[0057] As mentioned above, analysis time t[m2] was estimated not to be a beat in the immediately preceding probability calculation process Sa, but was designated as a beat in the change instruction. That is, the probability P[m2] represented by the output data O[m2] at analysis time t[m2] was set to a small value in the immediately preceding probability calculation process Sa, but should be set to a value close to 1 in response to the user's change instruction. Furthermore, not only for analysis time t[m2], but also for each analysis time t[m] among the M analysis times t[1] to t[M] in the song where intermediate data D[m] similar to intermediate data D[m2] at analysis time t[m2] is observed, the probability P[m] represented by the output data O[m] at that analysis time t[m] should be set to a value close to 1. Therefore, when the similarity between the intermediate data D[m] and the intermediate data D[m2] exceeds a predetermined threshold, the estimation model update unit 27 updates the multiple variables of the estimation model 50 so that the probability P[m] of the output data O[m] approaches a sufficiently large numerical value (e.g., 1). Specifically, the estimation model update unit 27 updates the coefficients defining each of the first portion 50a, the adaptive block 55, and the second portion 50b so as to reduce the error between the probability P[m] of the output data O[m] generated by the estimation model 50 from each piece of intermediate data D[m] whose similarity with the intermediate data D[m2] exceeds the threshold and the numerical value indicating the beat point (i.e., 1).
[0058] On the other hand, analysis time t[m1] is a time point that was estimated to correspond to a beat in the immediately preceding probability calculation process Sa, but was designated not to correspond to a beat in the change instruction. That is, the probability P[m1] represented by the output data O[m1] at analysis time t[m1] was set to a large value in the immediately preceding probability calculation process Sa, but should be set to a value close to 0 under the user's change instruction. Furthermore, not only for analysis time t[m1], but also for each analysis time t[m] among the M analysis times t[1] to t[M] in the song where intermediate data D[m] similar to intermediate data D[m1] at analysis time t[m1] is observed, the probability P[m] represented by the output data O[m] at that analysis time point t[m] should be set to a value close to 0. Therefore, when the similarity between the intermediate data D[m] and the intermediate data D[m1] exceeds a predetermined threshold, the estimation model update unit 27 updates the multiple variables of the estimation model 50 so that the probability P[m] of the output data O[m] approaches a sufficiently small numerical value (e.g., 0). Specifically, the estimation model update unit 27 updates the coefficients defining each of the first portion 50a, the adaptive block 55, and the second portion 50b so as to reduce the error between the probability P[m] of the output data O[m] generated by the estimation model 50 from each piece of intermediate data D[m] whose similarity with the intermediate data D[m1] exceeds the threshold and the numerical value indicating that the data does not correspond to a beat point (i.e., 0).
[0059] As can be understood from the above explanation, in the first embodiment, not only the intermediate data D[m1] and intermediate data D[m2] directly related to the change instruction, but also intermediate data D[m] similar to intermediate data D[m1] or intermediate data D[m2] among the M pieces of intermediate data D[1] to D[M] throughout the music piece are used to update the estimation model 50. Therefore, even though the beats at which the user instructs changes are only some of the beats within the music piece, after execution of the estimation model update process Sc, the estimation model 50 can generate M pieces of output data O[1] to O[M] that reflect the change instruction throughout the music piece.
[0060] 12 is a flowchart illustrating a specific procedure of the estimation model update process Sc. The control device 11 functions as the estimation model update unit 27 to execute the estimation model update process Sc.
[0061] When the estimation model update process Sc starts, the estimation model update unit 27 determines whether an adaptive block 55 has already been added to the estimation model 50 (Sc1). If an adaptive block 55 has not been added to the estimation model 50 (Sc1: NO), the estimation model update unit 27 adds a new initial adaptive block 55 between the first portion 50a and the second portion 50b of the estimation model 50 (Sc2). On the other hand, if an adaptive block 55 has already been added in a previous estimation model update process Sc (Sc1: YES), the addition of the adaptive block 55 (Sc2) is not executed.
[0062] When an adaptive block 55 is newly added, the estimation model 50 including the new adaptive block 55 is updated by the following process. When an adaptive block 55 has already been added, the estimation model 50 including the existing adaptive block 55 is updated by the following process. That is, in a state where the adaptive block 55 has been added to the estimation model 50, the estimation model update unit 27 updates multiple variables of the estimation model 50 by performing additional learning (Sc3 and Sc4) that applies the beat positions before and after the change in response to a change instruction from the user. Note that when the user instructs to change the positions of two or more beats, additional learning (Sc3 and Sc4) is performed for each beat related to the change instruction.
[0063] The estimation model update unit 27 updates the multiple variables of the estimation model 50 using the feature data F[m1] for the analysis time point t[m1], which corresponds to the beat point before the change due to the change instruction (Sc3). Specifically, the estimation model update unit 27 sequentially supplies each of the M pieces of intermediate data D[1] to D[M] to the adaptation block 55 in parallel with the supply of the feature data F[m1] to the estimation model 50, and updates the multiple variables of the estimation model 50 so that the probability P[m] of output data O[m] generated from each piece of intermediate data D[m] similar to the intermediate data D[m1] of the feature data F[m1] approaches 0. Therefore, the estimation model 50 is trained so as to generate output data O[m] with a probability P[m] close to 0 when feature data F[m] similar to the feature data F[m1] for the analysis time point t[m1] is input.
[0064] The estimation model update unit 27 also updates the multiple variables of the estimation model 50 using the feature data F[m2] for the analysis time point t[m2], which corresponds to the beat point after the change in response to the change instruction (Sc4). Specifically, while supplying the feature data F[m2] to the estimation model 50, the estimation model update unit 27 sequentially supplies each of the M pieces of intermediate data D[1] to D[M] to the adaptation block 55, and updates the multiple variables of the estimation model 50 so that the probability P[m] of output data O[m] generated from each piece of intermediate data D[m] similar to the intermediate data D[m2] of the feature data F[m2] approaches 1. Therefore, the estimation model 50 is trained to generate output data O[m] with a probability P[m] close to 1 when feature data F[m] similar to the feature data F[m2] for the analysis time point t[m2] is input.
[0065] In addition to updating the estimation model 50 in accordance with the change instruction by the estimation model update process Sc exemplified above, in the first embodiment, multiple updated beat positions are estimated by executing the beat position estimation process Sb under constraint conditions in accordance with the change instruction.
[0066] As described above, among the five elapsed points Y[0] to Y[4] within the beat interval δ, the elapsed point Y[0] corresponds to a beat, while the remaining four elapsed points Y[1] to Y[4] do not. The analysis time point t[m2] on the time axis corresponds to the beat after the change in response to the change instruction. Therefore, the estimation processing unit 23 forcibly sets the likelihood λ[i,j'] corresponding to the elapsed points Y[j'] (j' = 1 to 4) other than the elapsed point Y[0] to 0 among the N likelihoods λ[i,j] corresponding to the different states Q at the analysis time point t[m2]. Furthermore, the estimation processing unit 23 maintains the likelihood λ[i,0] corresponding to the elapsed point Y[0] at the value calculated by the above-described method among the N likelihoods λ[i,j] at the analysis time point t[m2]. Therefore, in the generation of the state sequence (Sb3), the maximum likelihood path that always passes through the state Q at the elapsed point Y[0] at the analysis time point t[m2] is estimated. That is, the analysis time t[m2] is estimated to correspond to a beat point. As can be understood from the above explanation, the beat point estimation process Sb is executed under the constraint that the state Q of the elapsed point Y[0] is observed at the analysis time t[m2] of the beat point after the change in response to a change instruction from the user.
[0067] On the other hand, the analysis time point t[m1] on the time axis does not correspond to the beat point after the change due to the change instruction. Therefore, the estimation processing unit 23 forcibly sets the likelihood λ[i,0] corresponding to the progress point Y[0] among the N likelihoods λ[i,j] corresponding to the different states Q at the analysis time point t[m1] to 0. In addition, the estimation processing unit 23 maintains the likelihood λ[i,j'] corresponding to the progress point Y[j'] other than the progress point Y[0] among the N likelihoods λ[i,j] at the analysis time point t[m1] at a significant value calculated by the above-mentioned method. Therefore, in generating the state sequence (Sb3), a maximum likelihood path that does not pass through the state Q of the progress point Y[0] at the analysis time point t[m1] is estimated. In other words, the analysis time point t[m1] is estimated not to correspond to a beat point. As can be understood from the above explanation, the beat point estimation process Sb is executed under the constraint that the state Q of the elapsed point Y[0] is not observed at the analysis time point t[m1] before the change due to the change instruction from the user.
[0068] As described above, the most likely path throughout the entire piece of music changes by setting the likelihood λ[i,0] of the passing point Y[0] at analysis time t[m1] to 0, and the likelihood λ[i,j'] of the passing point Y[j'] other than the passing point Y[0] at analysis time t[m2] to 0. In other words, even though the beats that the user instructs to change are only a portion of the beats in the piece of music, the change instruction is reflected in multiple beats throughout the piece of music.
[0069] Fig. 13 is a flowchart illustrating a specific procedure of the process executed by the control device 11. For example, the process of Fig. 13 is started in response to an instruction from a user via the operation device 14. When the process is started, the control device 11 executes a process (hereinafter referred to as "initial analysis process") of estimating multiple beats of a piece of music by analyzing the audio signal A (S1).
[0070] 14 is a flowchart illustrating a specific procedure of the initial analysis process. When the initial analysis process starts, the control device 11 (feature extraction unit 21) generates feature data F[m] for each of M analysis points t[1] to t[M] on the time axis (S11). As described above, the feature data F[m] is a time series of a plurality of feature amounts f[m] within a unit period U that includes the analysis point t[m].
[0071] The control device 11 (probability calculation unit 22) generates M pieces of output data O[m] corresponding to different analysis times t[m] by executing the probability calculation process Sa illustrated in Fig. 6 (S12). In addition, the control device 11 (estimation processing unit 23) estimates multiple beat points in the music piece by executing the beat point estimation process Sb illustrated in Fig. 9 (S13).
[0072] The control device 11 (display control unit 24) identifies one or more correction candidate points from among the multiple beat points estimated by the beat point estimation process Sb (S14). Specifically, a beat point whose beat interval δ with the immediately preceding or succeeding beat point deviates from the average value in the music piece, or a beat point whose duration of the beat interval δ is significantly different from the beat interval δ before and after the beat point, is identified as a correction candidate point. Alternatively, a beat point whose probability P[m] is below a predetermined value may be identified as a correction candidate point from among the multiple beat points. The control device 11 (display control unit 24) causes the display device 13 to display an analysis screen 70, as shown in FIG. 10 (S15).
[0073] After the initial analysis process exemplified above has been executed, the control device 11 (instruction receiving unit 26) waits until it receives an instruction from the user to change some of the beats in the music piece (S2: NO), as illustrated in Fig. 13. If an instruction to change is received (S2: YES), the control device 11 (estimation model update unit 27 and analysis processing unit 20) executes a beat point update process (S3) to update the positions of the multiple beats estimated in the initial analysis process in accordance with the change instruction from the user.
[0074] 15 is a flowchart illustrating a specific procedure of the beat point update process. The control device 11 (estimation model update unit 27) executes the estimation model update process Sc illustrated in FIG. 12 to update multiple variables of the estimation model 50 in response to a change instruction from the user (S31).
[0075] The control device 11 (probability calculation unit 22) generates M pieces of output data O[1] to O[M] by executing the probability calculation process Sa of FIG. 6 using the estimation model 50 updated by the estimation model update process Sc (S32). Furthermore, the control device 11 (analysis processing unit 20) generates beat point data B by executing the beat point estimation process Sb of FIG. 9 using the M pieces of output data O[1] to Q[M] (S33). In other words, multiple beat points in the music are estimated. The beat point estimation process Sb in the beat point update process is executed under the aforementioned constraint conditions in accordance with the change instruction.
[0076] As can be understood from the above explanation, updated beat positions are estimated by the estimation model update process Sc that updates the estimation model 50, the probability calculation process Sa that uses the updated estimation model 50, and the beat position estimation process Sb that uses the output data O[m] generated by the probability calculation process Sa. That is, the estimation model update unit 27, the probability calculation unit 22, and the analysis processing unit 20 implement an element (beat position update unit) that updates the positions of estimated beat positions.
[0077] The control device 11 (display control unit 24) identifies one or more correction candidate points from among the multiple beat points estimated by the beat point estimation process Sb, similar to step S14 described above (S34). The control device 11 (display control unit 24) causes the display device 13 to display the analysis screen 70 of Fig. 10, which includes beat images 751 representing each updated beat point (S35).
[0078] After the beat position update process illustrated above is executed, the control device 11 determines whether or not the user has instructed to end the process (S4), as illustrated in FIG. 13. If the user has not instructed to end the process (S4: NO), the control device 11 proceeds to wait for a change instruction from the user (S2). The control device 11 executes the beat position update process in response to another change instruction from the user (S3). In the estimation model update process Sc (S31) of the second or subsequent beat position update process, the result of the determination (Sc1) of the presence or absence of an adaptive block 55 is positive, so no new adaptive block 55 is added. That is, the estimation model 50 to which the adaptive block 55 was added in the first beat position update process is cumulatively updated with each subsequent execution of the estimation model update process Sc. On the other hand, if the user has instructed to end the process (S4: YES), the control device 11 ends the process of FIG. 13.
[0079] As described above, in the first embodiment, in response to a user's instruction to change some of the beats estimated by analyzing the audio signal A, the positions of the multiple beats in the piece of music, including the beats other than the certain beats, are updated. In other words, an instruction to change a certain part of the piece of music is reflected in the entire piece of music. Therefore, compared to a configuration in which the user must instruct a change in the position of each beat in the piece of music, it is possible to obtain a time series of beats that is in line with the user's intentions, while reducing the burden on the user of instructing a change in the position of each beat.
[0080] With the adaptation block 55 added between the first portion 50a and the second portion 50b in the estimation model 50, the estimation model 50 is updated by additional learning that applies the beat positions before and after the change in response to a change instruction from the user. Therefore, the estimation model 50 can be specialized to a state in which it can estimate beat positions that are in line with the user's intentions or preferences.
[0081] Furthermore, multiple beat points are estimated using a state transition model 60 configured with multiple states Q corresponding to any of multiple tempos X[i]. Therefore, multiple beat points can be estimated so that the tempo X[i] transitions naturally. In the first embodiment, in particular, the multiple states Q of the state transition model 60 correspond to different combinations of each of the multiple tempos X[i] and each of the multiple elapsed points Y[j] within the beat interval δ, and the beat point estimation process Sb is executed under the constraint that the state Q corresponding to the elapsed point Y[0] is observed at the analysis time point t[m] of the beat point after the change in response to a change instruction from the user. Therefore, multiple beat points can be estimated, including the time point after the change in response to a change instruction from the user as the beat point.
[0082] B: Second embodiment A second embodiment will be described. In each of the following exemplary embodiments, elements that have the same functions as those in the first embodiment will be denoted by the same reference numerals as those used in the description of the first embodiment, and detailed descriptions of each element will be omitted as appropriate.
[0083] 16 is a block diagram illustrating the functional configuration of an acoustic analysis system 100 according to the second embodiment. The control device 11 according to the second embodiment functions as a curve setting unit 28 in addition to the same elements as those in the first embodiment (an analysis processing unit 20, a display control unit 24, a playback control unit 25, an instruction receiving unit 26, and an estimation model updating unit 27).
[0084] The analysis processing unit 20 of the second embodiment estimates the tempo T[m] of a piece of music in addition to estimating multiple beats in the piece of music. That is, the analysis processing unit 20 analyzes the audio signal A to estimate a time series of M tempos T[1] to T[M] corresponding to different analysis points t[m] on the time axis.
[0085] 17 is a schematic diagram of an analysis screen 70 in the second embodiment. In addition to the same elements as in the first embodiment, the analysis screen 70 of the second embodiment includes an estimated tempo curve CT, a maximum tempo curve CH, and a minimum tempo curve CL. Specifically, the waveform area 73 of the analysis screen 70 displays a waveform 731 of audio signal A, the estimated tempo curve CT, the maximum tempo curve CH, and the minimum tempo curve CL on a common time axis. Note that in FIG. 17, the display of onset points 732 in audio signal A has been omitted for convenience.
[0086] FIG. 18 is a schematic diagram focusing on the estimated tempo curve CT, maximum tempo curve CH, and minimum tempo curve CL. The estimated tempo curve CT is a curve representing the time series of the tempo T[m] estimated by the analysis processing unit 20. The maximum tempo curve CH is a curve representing the time change of the maximum value H[m] of the tempo T[m] estimated by the analysis processing unit 20 (hereinafter referred to as the "maximum tempo"). In other words, the maximum tempo curve CH represents the time series of M maximum tempos H[1] to H[M] corresponding to different analysis times t[m] on the time axis. The minimum tempo curve CL is a curve representing the time change of the minimum value L[m] of the tempo T[m] estimated by the analysis processing unit 20 (hereinafter referred to as the "minimum tempo"). In other words, the minimum tempo curve CL represents the time series of M minimum tempos L[1] to L[M] corresponding to different analysis times t[m] on the time axis.
[0087] As can be understood from the above explanation, the analysis processing unit 20 estimates the tempo T[m] of a piece of music for each analysis time point t[m] within the range R[m] between the maximum tempo H[m] and the minimum tempo L[m] (hereinafter referred to as the "limited range"). Therefore, the estimated tempo curve CT is located between the maximum tempo curve CH and the minimum tempo curve CL. The position and width of the limited range R[m] change over time.
[0088] The curve setting unit 28 in FIG. 16 sets a maximum tempo curve CH and a minimum tempo curve CL. For example, the user can specify a maximum tempo curve CH of a desired shape and a minimum tempo curve CL of a desired shape by operating the operation device 14. The curve setting unit 28 sets the maximum tempo curve CH and the minimum tempo curve CL in response to a user's instruction on the analysis screen 70 (waveform area 73). For example, the curve setting unit 28 sets a continuous curve that passes through multiple points specified by the user in time series within the waveform area 73 as the maximum tempo curve CH or the minimum tempo curve. Furthermore, the user can specify, in the waveform area 73, changes to the previously set maximum tempo curve CH and minimum tempo curve CL by operating the operation device 14. The curve setting unit 28 changes the maximum tempo curve CH and the minimum tempo curve CL in response to a user's instruction on the analysis image (waveform area 73). As can be understood from the above description, according to the second embodiment, the user can easily change the maximum tempo curve CH and the minimum tempo curve CL while checking the analysis screen 70.
[0089] In the second embodiment, the waveform 731 of audio signal A and the maximum tempo curve CH and minimum tempo curve CL are displayed on a common time axis, making it easy for the user to visually grasp the relationship between changes over time in the maximum tempo H[m] or minimum tempo L[m] and the waveform 731 of audio signal A. Furthermore, because the estimated tempo curve CT is displayed together with the maximum tempo curve CH and minimum tempo curve CL, the user can visually grasp changes over time in the tempo T[m] of the music piece estimated between the maximum tempo curve CH and minimum tempo curve CL.
[0090] 19 is a flowchart illustrating the specific steps of the beat point estimation process Sb in the second embodiment. After setting the observation likelihood Λ[m] for each analysis time point t[m] in the same manner as in the first embodiment (Sb1), the estimation processor 23 calculates the path p[i,j] and the likelihood λ[i,j] for each state Q[i,j] of the state transition model 60 for each analysis time point t[m] (Sb2). The estimation processor 23 in the second embodiment sets, for each analysis time point t[m], the likelihood λ[i,j] corresponding to each tempo X[i] that exceeds the maximum tempo H[m] among the multiple tempos X[i] and the likelihood λ[i,j] corresponding to each tempo X[i] that is below the minimum tempo L[m] to 0. That is, among the N states Q of the state transition model 60, states Q corresponding to tempos X[i] outside the limit range R[m] are set to an invalid state. Furthermore, for each analysis time point t[m], the estimation processing unit 23 sets the likelihood λ[i,j] corresponding to each tempo X[i] inside the limit range R[m] to a significant value, as in the first embodiment. That is, among the N states Q of the state transition model 60, the state Q corresponding to the tempo X[i] inside the limit range R[m] is set to a valid state.
[0091] The estimation processing unit 23 generates a state sequence using the same method as in the first embodiment (Sb3). That is, a sequence in which the states Q with the largest likelihood λ[i,j] among the N states Q are arranged for each analysis time t[m] is generated as the state sequence. As described above, the likelihood λ[i,j] of the state Q[i,j] corresponding to the tempo X[i] outside the limit range R[m] at the analysis time t[m] is set to 0. Therefore, the state Q corresponding to the tempo X[i] outside the limit range R[m] is not selected as an element of the state sequence. As can be understood from the above explanation, the invalid state of each state Q means that the state Q is not selected.
[0092] The estimation processing unit 23 generates beat data B (Sb4) as in the first embodiment, and identifies the tempo T[m] at each analysis time point t[m] from the state series (Sb5). That is, the tempo X[i] of the state Q corresponding to the analysis time point t[m] in the state series is set as the tempo T[m]. As described above, a state Q corresponding to a tempo X[i] outside the limit range R[m] is not selected as an element of the state series, so the tempo T[m] is limited to a numerical value inside the limit range R[m].
[0093] As described above, in the second embodiment, the maximum tempo curve CH and the minimum tempo curve CL are set in response to instructions from the user. The tempo T[m] of the music piece is then estimated within a limited range R[m] between the maximum tempo H[m] represented by the maximum tempo curve CH and the minimum tempo L[m] represented by the minimum tempo curve CL. This reduces the possibility of estimating a tempo that deviates excessively from the tempo intended by the user (for example, a tempo that is twice or half the value intended by the user). In other words, the tempo T[m] of the music piece represented by the audio signal A can be estimated with high accuracy.
[0094] In the second embodiment, a state transition model 60 consisting of a plurality of states Q corresponding to any of a plurality of tempos X[i] is used to estimate a plurality of beats. Therefore, a tempo T[m] that naturally transitions over time is estimated. Moreover, a tempo T[m] restricted to within the restricted range R[m] can be estimated by a simple process of invalidating a state Q corresponding to a tempo X[i] outside the restricted range R[m].
[0095] C: Third embodiment In the first embodiment, the output data O[m] representing the probability P[m] calculated by the probability calculation unit 22 using the estimation model 50 is applied to the beat position estimation process Sb by the estimation processing unit 23. In the third embodiment, the probability P[m] calculated by the estimation model 50 (hereinafter referred to as "probability P1[m]") is adjusted in response to an operation from the user to the operation device 14, and the output data O[m] representing the adjusted probability P2[m] is applied to the beat position estimation process Sb.
[0096] 20 is an explanatory diagram of the process in which the probability calculation unit 22 of the third embodiment generates output data O[m]. While listening to the musical performance sound that the playback control unit 25 causes the sound emitting device 15 to play, the user operates the operation device 14 at each time point that the user recognizes as a beat. For example, the user performs a tap operation on the touch panel of the operation device 14 at each time point that the user recognizes as a beat, while the musical piece is being played. In FIG. 20, the time point τ at which the user performs an operation (hereinafter referred to as the "operation time point") is illustrated on the time axis.
[0097] The probability calculation unit 22 sets a unit distribution W for each operation time point τ. The unit distribution W is a distribution of weighted values w[m] on the time axis. For example, a probability distribution such as a normal distribution with a predetermined variance is used as the unit distribution W. In each unit distribution W, the weighted value w[m] is maximum at the operation time point τ, and decreases as the distance from the operation time point τ increases.
[0098] The probability calculation unit 22 calculates the adjusted probability P2[m] by multiplying the probability P1[m] generated by the estimation model 50 for the analysis time point t[m] by the weight w[m] at the analysis time point t[m]. Therefore, even if the probability P1[m] generated by the estimation model 50 is small for the analysis time point t[m], if the analysis time point t[m] is close to the operation time τ, the adjusted probability P2[m] is set to a large value. The probability calculation unit 22 supplies output data O[m] representing the adjusted probability P2[m] to the estimation processing unit 23. The procedure of the beat point estimation process Sb, in which the estimation processing unit 23 estimates multiple beat points using the output data O[m], is the same as in the first embodiment.
[0099] The third embodiment also achieves the same effects as the first embodiment. Furthermore, in the third embodiment, the weighted value w[m] of the unit distribution W set at the time τ of the user's operation is multiplied by the probability P1[m], which has the advantage of being able to estimate a beat point that fully reflects the user's intention or preference. The configuration of the second embodiment is also applicable to the third embodiment.
[0100] D: Modification Specific modified embodiments that can be added to each of the embodiments exemplified above are exemplified below. Two or more embodiments arbitrarily selected from the following examples may be appropriately combined within the scope of not being mutually contradictory.
[0101] (1) The configuration of the estimation model 50 is not limited to the example shown in FIG. 4. For example, the estimation model 50 may include a recurrent neural network. The estimation model 50 may also include additional elements such as a long short-term memory (LSTM). The estimation model 50 may also be configured by combining multiple types of deep neural networks.
[0102] (2) The specific procedure for estimating multiple beats in a piece of music by analyzing the audio signal A is not limited to the examples given in the above embodiments. For example, the analysis processing unit 20 may estimate, as a beat, the analysis time t[m] at which the probability P[m] represented by the output data O[m] is maximized. In other words, the use of the state transition model 60 is omitted. Alternatively, the analysis processing unit 20 may estimate, as a beat, the time at which a feature value f[m], such as the volume of the audio signal A, significantly increases. In other words, the use of the estimation model 50 is omitted.
[0103] (3) The configuration of the first embodiment, which updates the multiple beat points estimated by the initial analysis process, may be omitted in the second embodiment. In other words, the configuration of the first embodiment, which updates the multiple beat points throughout the entire piece of music in response to an instruction to change some of the multiple estimated beat points, and the configuration of the second embodiment, which estimates the tempo T[m] of the piece of music within a limited range R[m] in response to an instruction from the user, may be established independently of each other.
[0104] (4) The acoustic analysis system 100 may be realized by a server device that communicates with an information device such as a smartphone or a tablet terminal. For example, the acoustic analysis system 100 generates beat data B by analyzing an acoustic signal A received from the information device and transmits the beat data B to the information device. Similarly, the reception of a change instruction from a user (S2) and the beat update process (S3) are also performed by the acoustic analysis system 100 that communicates with the information device.
[0105] (5) As described above, the functions of the acoustic analysis system 100 illustrated above are realized by the cooperation of one or more processors constituting the control device 11 and a program stored in the storage device 12. The program according to the present disclosure may be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium may be, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but may also include any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transient, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage device 12 storing the program in the distribution device corresponds to the non-transitory recording medium described above.
[0106] E: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.
[0107] An acoustic analysis method according to one aspect (aspect 1) of the present disclosure estimates multiple beats of a piece of music by analyzing an audio signal representing the music being played, receives an instruction from a user to change the positions of some of the multiple beats, and updates the positions of the multiple beats in accordance with the user's instruction. In the above aspect, in accordance with an instruction to change the positions of some of the multiple beats estimated by analyzing the audio signal, the positions of the multiple beats, including those other than the some of the multiple beats, are updated. Therefore, compared to a configuration in which the user must change the positions of all of the multiple beats, it is possible to obtain a time series of beats that is in line with the user's intentions, while reducing the burden on the user of instructing changes to the positions of each beat.
[0108] In a specific example (Aspect 2) of Aspect 1, the estimation of beat points includes a feature extraction process that generates feature data including features of the audio signal for each of a plurality of analysis time points on a time axis, a probability calculation process that generates output data indicating the probability that the analysis time point corresponds to a beat by inputting the feature data generated for the analysis time point by the feature extraction process into an estimation model that has learned the relationship between the training feature data corresponding to the time point on the time axis and training output data indicating the probability that the time point corresponds to a beat, and a beat point estimation process that estimates the plurality of beat points from the output data generated by the probability calculation process. According to the above aspect, it is possible to generate output data that is statistically valid for unknown feature data based on the underlying relationship between the training feature data and the training output data.
[0109] In a specific example (aspect 3) of aspect 2, the updating of the positions of the multiple beats involves adding an adaptive block between the first part on the input side and the second part on the output side of the estimation model, and then performing additional learning to apply the beat positions before or after the change in accordance with the instruction from the user to update the estimation model, and then estimating the multiple updated beats by the probability calculation process using the updated estimation model and the beat estimation process using the output data generated by the probability calculation process. According to the above aspect, the estimation model is updated by additional learning to apply the beat positions before or after the change in accordance with the instruction from the user. Therefore, the estimation model can be specialized to a state in which beats can be estimated according to the user's intention or preference.
[0110] The adaptive block is a block that calculates the similarity between first intermediate data generated by the first portion from feature data corresponding to the beat positions before or after the change in response to a user instruction, and second intermediate data corresponding to the feature data at each of multiple analysis points in the music. The entire estimation model, including the adaptive block, is updated so that the output data at an analysis point corresponding to the second intermediate data similar to the first intermediate data at the beat positions before the change in response to a user instruction approaches a numerical value indicating that the output data at an analysis point corresponding to the second intermediate data similar to the first intermediate data at the beat positions after the change approaches a numerical value indicating that the output data at the analysis point corresponds to a beat.
[0111] In a specific example (Aspect 4) of Aspect 2 or Aspect 3, the beat point estimation process estimates the multiple beat points using a state transition model configured with multiple states corresponding to any of multiple tempos. According to the above aspect, multiple beat points are estimated using a state transition model configured with multiple states corresponding to any of multiple tempos. Therefore, multiple beat points are estimated so that the tempo transitions naturally over time.
[0112] In a specific example (aspect 5) of aspect 4, the multiple states of the state transition model correspond to different combinations of each of the multiple tempos and each of multiple elapsed points within a beat interval, and in the beat point estimation process, a time point at which a state of the multiple elapsed points corresponding to an end point of the beat interval is observed is estimated as a beat point, and in updating the positions of the multiple beat points, the multiple updated beat points are estimated by executing the beat point estimation process under a constraint that a state corresponding to the end point of the beat interval is observed at the time point of the beat point after the change based on the instruction from the user. According to the above aspect, it is possible to estimate multiple beat points including the beat point after the change based on the instruction from the user.
[0113] An acoustic analysis system according to one aspect (aspect 6) of the present disclosure includes an analysis processing unit that estimates multiple beats of a piece of music by analyzing an acoustic signal representing the performance sound of the piece of music, an instruction receiving unit that receives instructions from a user to change the positions of some of the multiple beats, and a beat update unit that updates the positions of the multiple beats in accordance with instructions from the user.
[0114] A program according to one aspect (aspect 7) of the present disclosure causes a computer system to function as an analysis processing unit that estimates multiple beats of a piece of music by analyzing an audio signal representing the sound of the piece of music being played, an instruction receiving unit that receives instructions from a user to change the positions of some of the multiple beats, and a beat update unit that updates the positions of the multiple beats in accordance with instructions from the user.
[0115] It should be noted that "tempo" in this specification is an arbitrary numerical value that indicates the speed of a performance, and is not limited to tempo in the narrow sense of the number of beats per minute (BPM). [Explanation of symbols]
[0116] 100...acoustic analysis system, 11...control device, 12...storage device, 13...display device, 14...operation device, 15...sound emission device, 20...analysis processing unit, 21...feature extraction unit, 22...probability calculation unit, 23...estimation processing unit, 24...display control unit, 25...playback control unit, 26...instruction reception unit, 27...estimation model update unit, 28...curve setting unit, 50...estimation model, 50a...first part, 50b...second part, 51...input layer, 52 (52a, 52b)...intermediate layer, 53...output layer, 55...adaptation block, 59...tentative model, 60...state transition model.
Claims
[Claim 1] estimating a plurality of beats of a piece of music by analyzing an audio signal representing a performance sound of the piece of music; receiving an instruction from a user to change the positions of some of the beats; Update the positions of the plurality of beats in response to an instruction from the user An acoustic analysis method implemented by a computer system.
Citation Information
Patent Citations
Acoustic signal analysis device and acoustic signal analysis program
JP2015114361A