Acoustic analysis system, acoustic analysis method and program

The acoustic analysis system enables users to adjust and update beat positions in music analysis, using a machine-learning trained model to align beats with their intentions, addressing errors in conventional estimation methods.

JP7800324B2Active Publication Date: 2026-01-16YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022106820
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-01
Publication Date
2026-01-16
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

Conventional techniques for estimating beats in music may erroneously identify backbeats as beats or estimate tempos twice the original tempo, and fail to align with user intentions, necessitating a system that allows users to adjust beat positions on the time axis.

Method used

An acoustic analysis system with a beat estimation unit, beat editing unit, and update processing unit that allows users to move and update beat positions based on their intentions, using an estimation model trained through machine learning to refine beat estimation.

Benefits of technology

The system accurately aligns beat positions with user intentions by updating the estimation model to reflect user adjustments, ensuring accurate and intuitive beat estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007800324000001
    Figure 0007800324000001
  • Figure 0007800324000002
    Figure 0007800324000002
  • Figure 0007800324000003
    Figure 0007800324000003
Patent Text Reader

Abstract

To estimate a beat point that is appropriately suitable to user's intention.SOLUTION: An information analysis system 100 includes a beat point estimation part 21 for estimating a plurality of beat points B by estimation processing to an acoustic signal A, a beat point edition part 24 for moving a target beat point selected by a user among the plurality of beat points B and one or more adjacent beat points located around the target beat point among the plurality of beat points on a temporal axis according to an instruction from the user, an update processing part 25 for updating estimation processing in accordance with movement of the target beat point and the one or more adjacent beat points. The beat point estimation part 21 re-estimates the plurality of beat points B by executing updated estimation processing to the acoustic signal A.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for analyzing acoustic signals. [Background technology]

[0002] Analysis techniques have been proposed for estimating beats of a piece of music by analyzing audio signals that represent the sounds of the music being played. For example, Patent Document 1 discloses a technique for estimating beats of a piece of music using a probabilistic model such as a hidden Markov model. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-114361 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques for estimating beats in music may erroneously estimate, for example, a backbeat as a beat, or may erroneously estimate a beat corresponding to a tempo twice the original tempo of the music. Furthermore, there is a possibility that the estimated beat may not match the user's intention, such as when a backbeat is estimated in a situation where the user is expecting a downbeat to be estimated. Considering the above circumstances, it is important to have a configuration that allows the user to change the positions on the time axis of multiple beats estimated from an audio signal. In consideration of the above circumstances, one aspect of the present disclosure aims to estimate beats that appropriately match the user's intention. [Means for solving the problem]

[0005] In order to solve the above problems, an acoustic analysis system according to one aspect of the present disclosure includes: a beat estimation unit that estimates multiple first beats by performing estimation processing on an acoustic signal; a beat editing unit that moves a target beat selected by a user from among the multiple first beats and one or more adjacent beats located around the target beat from among the multiple first beats on a time axis in accordance with instructions from the user; and an update processing unit that updates the estimation processing in accordance with the movement of the target beat and the one or more adjacent beats, and the beat estimation unit estimates multiple second beats by performing the updated estimation processing on the acoustic signal.

[0006] An acoustic analysis method according to one aspect of the present disclosure estimates multiple first beat points by performing an estimation process on an acoustic signal, moves a target beat point selected by a user from among the multiple first beat points and one or more adjacent beat points from among the multiple first beat points located around the target beat point on a time axis in accordance with an instruction from the user, updates the estimation process in accordance with the movements of the target beat point and the one or more adjacent beat points, and estimates multiple second beat points by executing the updated estimation process on the acoustic signal.

[0007] A program according to one aspect of the present disclosure causes a computer system to function as: a beat estimation unit that estimates multiple first beat points by performing estimation processing on an audio signal; a beat editing unit that moves a target beat point selected by a user from among the multiple first beat points and one or more adjacent beat points located around the target beat point from among the multiple first beat points on a time axis in accordance with instructions from the user; and an update processing unit that updates the estimation processing in accordance with the movement of the target beat point and the one or more adjacent beat points, wherein the beat estimation unit estimates multiple second beat points by performing the updated estimation processing on the audio signal. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating the configuration of an acoustic analysis system according to a first embodiment. [Figure 2] FIG. 1 is a block diagram illustrating an example of the functional configuration of an acoustic analysis system. [Figure 3] 10 is a flowchart of an estimation process. [Figure 4] FIG. 1 is an explanatory diagram of machine learning for establishing an estimation model. [Figure 5] FIG. 10 is a schematic diagram of a confirmation image. [Figure 6] FIG. 10 is an explanatory diagram of a beat point movement and update process. [Figure 7] 10 is a flowchart of an update process. [Figure 8] 10 is a flowchart of an acoustic analysis process. [Figure 9] FIG. 10 is a block diagram illustrating an example of the functional configuration of an acoustic analysis system according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] A: First embodiment 1 is a block diagram illustrating the configuration of an acoustic analysis system 100 according to a first embodiment. The acoustic analysis system 100 is a computer system that estimates multiple beats B of a piece of music by analyzing an audio signal A that represents the performance sounds of the piece of music.

[0010] The acoustic analysis system 100 includes a control device 11, a storage device 12, a display device 13, an operation device 14, and a sound emission device 15. The acoustic analysis system 100 is realized by an information device such as a smartphone, a tablet terminal, or a personal computer. The acoustic analysis system 100 may be realized as a single device, or may be realized as multiple devices configured separately from each other.

[0011] The control device 11 is one or more processors that control each element of the acoustic analysis system 100. Specifically, the control device 11 is configured by one or more types of processors, such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0012] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. For example, a known recording medium such as a semiconductor recording medium or a magnetic recording medium, or a combination of multiple types of recording media, is used as the storage device 12. Note that, for example, a portable recording medium that is detachable from the acoustic analysis system 100, or a recording medium that the control device 11 can access via a communication network (e.g., cloud storage) may also be used as the storage device 12.

[0013] The storage device 12 stores the audio signal A. The audio signal A is a series of samples representing the waveform of the performance sound of a piece of music. Specifically, the audio signal A represents at least one of the instrument sounds and the vocal sounds of the piece of music. The audio signal A may have any data format. The audio signal A may be supplied to the acoustic analysis system 100 from a signal supply device separate from the acoustic analysis system 100. The signal supply device is, for example, a playback device that supplies the audio signal A recorded on a recording medium to the acoustic analysis system 100, or a communication device that supplies the audio signal A received from a distribution device (not shown) via a communication network to the acoustic analysis system 100.

[0014] The display device 13 displays images under the control of the control device 11. For example, various display panels such as a liquid crystal display panel or an organic EL (Electroluminescence) panel are used as the display device 13. Note that the display device 13, which is separate from the acoustic analysis system 100, may be connected to the acoustic analysis system 100 by wire or wirelessly. The operation device 14 is an input device that accepts instructions from a user. The operation device 14 is, for example, an operator operated by the user or a touch panel that detects contact by the user.

[0015] The sound emitting device 15 reproduces sound under the control of the control device 11. For example, a speaker or a headphone is used as the sound emitting device 15. Note that the sound emitting device 15, which is separate from the acoustic analysis system 100, may be connected to the acoustic analysis system 100 by wire or wirelessly.

[0016] 2 is a block diagram illustrating an example of the functional configuration of the acoustic analysis system 100. The control device 11 executes a program stored in the storage device 12 to realize a plurality of functions for processing the acoustic signal A (a beat estimation unit 21, a display control unit 22, a playback control unit 23, a beat editing unit 24, and an update processing unit 25).

[0017] The beat estimation unit 21 estimates multiple beats B in a piece of music by analyzing an audio signal A. Specifically, the beat estimation unit 21 generates time-series data that specifies the time of each of the multiple beats B in the piece of music. The beat estimation unit 21 of the first embodiment includes a feature extraction unit 30, a first processing unit 31, and a second processing unit 32.

[0018] The feature extraction unit 30 calculates a feature value F(t) of the audio signal A for each of a plurality of time points t on the time axis (hereinafter referred to as "analysis time points"). The analysis time points t are set at predetermined intervals on the time axis. The intervals between the analysis time points t are sufficiently smaller than the intervals between beat points B expected in the music piece.

[0019] The feature F(t) is information representing the acoustic features of the audio signal A at analysis time t. For example, the feature F(t) at each analysis time t is a time series of audio information within a predetermined period including the analysis time t. The audio information is, for example, information about the intensity of the audio signal A, such as volume and amplitude. Information about the frequency characteristics (timbre) of the audio signal A is also used as audio information. Examples of information about frequency characteristics include Mel-Frequency Cepstrum Coefficients (MFCC), Mel-Scale Log Spectrum (MSLS), and Constant-Q Transform (CQT). Note that multiple pieces of audio information corresponding to one analysis time t may be used as the feature F(t). The types of audio information are not limited to the above examples. The audio information may be a combination of multiple types of audio information related to the audio signal A.

[0020] The first processing unit 31 and the second processing unit 32 estimate multiple beat positions B from each feature F(t) of the audio signal A. Fig. 3 is a flowchart of the process S2 for estimating multiple beat positions B (hereinafter referred to as "estimation process"). The estimation process S2 includes a first process S21 and a second process S22. The first processing unit 31 executes the first process S21, and the second processing unit 32 executes the second process S22.

[0021] The first process S21 is a process for generating a probability P(t) that each analysis time point t corresponds to a beat point B of the music. The larger the probability P(t) of each analysis time point t, the higher the likelihood that the analysis time point t corresponds to a beat point B. The first processing unit 31 generates a time series of the probability P(t) by repeating the first process S21 for each analysis time point t. An estimation model M is used in the first process S21.

[0022] There is a correlation between the feature F(t) at each analysis time point t of the audio signal A and the probability P(t) that the analysis time point t corresponds to the beat B. The estimation model M is a statistical model that has learned this correlation. In other words, the estimation model M is a trained model that has learned the relationship between the feature F(t) and the probability P(t) through machine learning. The estimation model M can also be expressed as a trained model that has acquired the relationship between the feature F(t) and the probability P(t) through training (machine learning). The first processing unit 31 processes the feature F(t) of the audio signal A at each analysis time point t using the estimation model M to generate the probability P(t). Specifically, the first processing unit 31 generates the probability P(t) by inputting input data including the feature F(t) into the estimation model M.

[0023] The estimation model M is configured, for example, by a deep neural network (DNN). For example, any type of deep neural network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN), is used as the estimation model M. The estimation model M may be configured by combining multiple types of deep neural networks. In addition, the estimation model M may be equipped with additional elements such as a long short-term memory (LSTM) or attention.

[0024] The estimation model M is realized by a combination of a program that causes the control device 11 to execute a calculation to generate a probability P(t) from the feature F(t) and a plurality of variables (specifically, weights and biases) that are applied to the calculation. The program and the plurality of variables that realize the estimation model M are stored in the storage device 12. The numerical values ​​of each of the plurality of variables that define the estimation model M are set in advance by machine learning.

[0025] The second process S22 in Fig. 3 is a process for estimating multiple beats B in a piece of music from the time series of probabilities P(t) generated by the first process S21. Various state transition models are used in the second process S22. The state transition model is configured, for example, by a Hidden Semi-Markov Model (HSMM), and multiple beats B are estimated using the Viterbi algorithm, which is an example of dynamic programming. For example, the time point at which the probability P(t) is maximized is estimated as the beat B.

[0026] 4 is an explanatory diagram of machine learning that establishes the estimation model M. For example, the estimation model M is established through machine learning by a machine learning system 200 that is separate from the acoustic analysis system 100. The estimation model M is provided from the machine learning system 200 to the acoustic analysis system 100. Note that the functions of the machine learning system 200 may be incorporated into the acoustic analysis system 100.

[0027] Multiple sets of training data Z are used for machine learning of the estimation model M. Each of the multiple sets of training data Z is composed of a combination of a machine learning feature Fm and a machine learning probability Pm. The feature Fm is a feature F(t) at a specific time point in an audio signal Am prepared for machine learning. The audio signal Am is a signal recorded from sound emitted in an acoustic space, or a signal synthesized using a known audio synthesis process. The machine learning probability Pm corresponding to a specific time point is the probability that the time point corresponds to beat B in the music (i.e., the correct value). Multiple sets of training data Z are prepared for a large number of songs for which beat B is known. The audio signal Am is an example of a "learning audio signal."

[0028] The machine learning system 200 calculates an error function that represents the error between the probability P(t) of an output from an initial or provisional model (hereinafter referred to as the "provisional model") M0 when a feature value Fm of each training data Z is input, and the probability Pm of the training data Z. Then, the machine learning system 200 updates multiple variables of the provisional model M0 so as to reduce the error function. The provisional model M0 at the time when the above process is repeated for each of the multiple training data Z is determined as the estimated model M.

[0029] Therefore, the estimation model M outputs a statistically valid probability P(t) for an unknown feature F(t) based on the underlying relationship between the feature Fm and the probability Pm in the multiple pieces of training data Z. In other words, the estimation model M is a trained model that has learned the relationship between the feature Fm of the machine learning audio signal Am and the probability Pm that the time point at which the feature is observed corresponds to the beat B. The first processing unit 31 processes the feature F(t) for each analysis time point t using the estimation model M established by the above procedure, thereby generating the probability P(t) that the analysis time point t corresponds to the beat B of the music.

[0030] As explained above, in the first embodiment, multiple beats B are estimated from an audio signal A using an estimation model M that has learned the relationship between the feature Fm of an audio signal Am for machine learning and the probability Pm that the analysis time point t at which the feature Fm is observed corresponds to a beat B. Therefore, multiple beats B can be estimated with high accuracy for an unknown audio signal A whose feature F(t) varies in a variety of ways.

[0031] The display control unit 22 in Fig. 2 displays an image on the display device 13. Specifically, the display control unit 22 displays the confirmation image G in Fig. 5 on the display device 13. The confirmation image G includes a waveform area Ga and a beat area Gb. A common time axis is set for the waveform area Ga and the beat area Gb.

[0032] The waveform area Ga displays a waveform within a specific range (hereinafter referred to as the "display range") of the audio signal A. The display control unit 22 changes the display range of the audio signal A in response to an instruction from the user via the operation device 14. The beat point area Gb displays multiple beat points B estimated from the audio signal A by the beat point estimation unit 21. Specifically, multiple beat points B within the display range of the audio signal A are displayed in the beat point area Gb. The beat point area Gb is an example of a "beat point image."

[0033] The user can instruct playback of the audio signal A by operating the operation device 14. The playback control unit 23 in FIG. 2 supplies the audio signal A to the sound emitting device 15, thereby playing back the audio represented by the audio signal A. As illustrated in FIG. 5, the display control unit 22 displays a playback position Gc on the confirmation image G in parallel with the playback of the audio signal A. The playback position Gc is the point in time in the audio signal A that is being played back by the sound emitting device 15. Therefore, the playback position Gc progresses in the direction of the time axis in parallel with the playback of the audio signal A. The user can confirm the position of the beat point B estimated by the immediately preceding estimation process S2 by visually checking the beat point area Gb while listening to the sound played back by the sound emitting device 15. If the current position of the beat point B does not match the user's intention, the user can instruct a correction of the estimated position of the beat point B by operating the operation device 14.

[0034] The beat editing unit 24 in Fig. 2 moves each beat B on the time axis in response to an instruction from the user. Moving the beat B is a process of changing the position of the beat B on the time axis. Fig. 6 is an explanatory diagram regarding the movement of the beat B. State 1 in Fig. 6 is a state in which multiple beats B have been estimated by the estimation process S2 described above. Fig. 6 also shows a time series of the probability P(t) calculated in the first process S21.

[0035] The user can select one of the multiple beats B displayed in the beat area Gb (hereinafter referred to as a "target beat Bn") by operating the operation device 14 while checking the beat area Gb. The user can also instruct the movement of the target beat Bn on the time axis by operating the operation device 14 while checking the beat area Gb. Specifically, the user can instruct the direction (forward / backward) and amount of movement δ of the target beat Bn. For example, the user can instruct the movement of the target beat Bn to a point that the user deems appropriate. As illustrated as state 2 in FIG. 6 , the beat editing unit 24 moves the target beat Bn on the time axis in the direction (forward / backward) instructed by the user by the amount of movement δ instructed by the user. Note that while FIG. 6 illustrates the case where the target beat Bn moves forward, the target beat Bn may also move backward.

[0036] As illustrated as state 3 in FIG. 6 , the beat editing unit 24 moves, on the time axis, the beat B (hereinafter referred to as "adjacent beat Bn-1") located immediately before the target beat Bn among the multiple beats B and the beat B (hereinafter referred to as "adjacent beat Bn+1") located immediately after the target beat Bn, in conjunction with the target beat Bn. Specifically, the beat editing unit 24 moves the adjacent beats Bn-1 and Bn+1 on the time axis in the direction (forward / backward) of movement instructed by the user for the target beat Bn by the movement amount δ instructed by the user for the target beat Bn. That is, the three beats B, namely the target beat Bn and the adjacent beats Bn±1 before and after it, move similarly on the time axis in accordance with the user's instruction. Therefore, the temporal relationship between the target beat Bn and each of the adjacent beats Bn±1 is maintained before and after the movement.

[0037] The beat point area Gb displayed on the display device 13 includes the target beat point Bn and adjacent beat points Bn±1. The display control unit 22 reflects the movement of each beat point B by the beat point editing unit 24 in the beat point area Gb displayed on the display device 13. Specifically, the display control unit 22 moves the target beat point Bn and each adjacent beat point Bn±1 in the beat point area Gb on the time axis in response to an instruction from the user.

[0038] 2 updates the estimation model M in accordance with the movement of the target beat Bn and each of the adjacent beats Bn±1. Specifically, the update processing unit 25 updates the estimation model M through machine learning in accordance with the movement of the target beat Bn and each of the adjacent beats Bn±1.

[0039] 7 is a flowchart of the process S8 (hereinafter referred to as the "update process") in which the control device 11 (update processing unit 25) updates the estimated model M. The update process S8 is started in response to the movement of the target beat Bn and each of the adjacent beats Bn±1.

[0040] When the update process S8 starts, the update processing unit 25 sets a numerical sequence C corresponding to the target beat Bn after the shift and each adjacent beat Bn±1 on the time axis (S81). As illustrated as State 4 in Fig. 6, the numerical sequence C is a time series of numerical values ​​Q(t) set for each analysis time point t on the time axis.

[0041] The numerical sequence C includes a numerical distribution D corresponding to the target beat Bn after the shift and each adjacent beat Bn±1. The numerical distribution D is a distribution of the numerical value Q(t) in a specific range on the time axis. The numerical distribution D is expressed by a probability distribution function defined on the time axis with time t as a variable. The numerical distribution D in the first embodiment is a line-symmetric triangular distribution over a predetermined distribution width. A numerical distribution D is set individually for each beat B. The position on the time axis of the numerical distribution D corresponding to each beat B is determined so that it has a maximum value at that beat B. For example, the numerical distribution D corresponding to the target beat Bn has a maximum value at the target beat Bn, and the numerical distribution D corresponding to each adjacent beat Bn±1 has a maximum value at the adjacent beat Bn±1. The numerical values ​​Q(t) of the numerical sequence C at each analysis time point t other than those of the numerical distribution D are set to zero.

[0042] As illustrated as state 5 in FIG. 6, the update processing unit 25 calculates the error e(t) for each analysis time point t within the application section T on the time axis (S82). The application section T is a series of sections including adjacent beat points Bn-1 and Bn+1. Specifically, the period on the time axis with adjacent beat points Bn-1 and Bn+1 as its end points is set as the application section T. The error e(t) for each analysis time point t is a numerical value corresponding to the difference between the probability P(t) at that analysis time point t and the numerical value Q(t) at that analysis time point t in the numerical sequence C. For example, the square of the difference between the probability P(t) and the numerical value Q(t) (={P(t) - Q(t)} 2 ) is calculated as the error e(t).

[0043] The update processing unit 25 calculates an error function E from multiple errors e(t) calculated for different analysis time points t within the application interval T (S83). The error function E is an objective function that represents the difference between the probability P(t) and the numerical value Q(t) within the application interval T. For example, the sum of multiple errors e(t) within the application interval T is calculated as the error function E.

[0044] The update processing unit 25 updates the estimation model M so as to minimize the error function E (S84). Any known technique may be used to update the estimation model M. For example, an adaptation process using self-attention may be used to update the estimation model M. The adaptation process for the estimation model M is described, for example, in Kazuhiko Yamamoto, "HUMAN-IN-THE-LOOP ADAPTATION FOR INTERACTIVE MUSICAL BEAT TRACKING," Proceedings of the 22nd ISMIR Conference, Online, November 7-12, 2021.

[0045] As can be understood from the above explanation, the update processing unit 25 updates the estimation model M so as to reduce the error e(t) between the numerical distribution D (numerical values ​​Q(t)) corresponding to the shifted target beat Bn and each adjacent beat Bn±1 and the time series of the probability P(t) estimated by the previous estimation process S2 (first process S21). Therefore, the shifts of the target beat Bn and each adjacent beat Bn±1 can be appropriately reflected in the estimation model M.

[0046] 8 is a flowchart of a process (hereinafter referred to as "acoustic analysis process") executed by the control device 11. For example, the acoustic analysis process is started in response to an instruction from the user via the operation device .

[0047] When the acoustic analysis process starts, the control device 11 (feature extraction unit 30) calculates a feature value F(t) of the acoustic signal A for each analysis time point t on the time axis (S1). The control device 11 (beat estimation unit 21) estimates multiple beats B from each feature value F(t) of the acoustic signal A by an estimation process S2 illustrated in FIG. 3. The first step S21 of the estimation process S2 uses an estimation model M that has learned the relationship between the feature value F(t) and the probability P(t) by machine learning. The control device 11 (display control unit 22) displays a confirmation image G on the display device 13 (S3). The multiple beats B estimated by the estimation process S2 are displayed in the beat area Gb of the confirmation image G.

[0048] The control device 11 determines whether or not a termination condition is met (S4). The termination condition is, for example, when the user issues an instruction to terminate the acoustic analysis process by operating the operation device 14. If the termination condition is met (S4: YES), the control device 11 terminates the acoustic analysis process. If the termination condition is not met (S4: NO), the control device 11 (beat editing unit 24) determines whether or not an instruction to move the target beat Bn has been received from the user (S5). If an instruction to move the target beat Bn has not been received (S5: NO), the control device 11 transitions the process to step S4. That is, the control device 11 waits for an instruction to terminate the acoustic analysis process or an instruction to move the target beat Bn.

[0049] If an instruction to move the target beat Bn is received (S5: YES), the control device 11 (beat editing unit 24) moves the target beat Bn and the adjacent beats Bn±1 on the time axis in accordance with the instruction from the user (S6). Also, the control device 11 (display control unit 22) moves the target beat Bn and the adjacent beats Bn±1 in the beat area Gb in accordance with the instruction from the user (S7).

[0050] The control device 11 (update processing unit 25) updates the estimation model M by the update process S8 illustrated in FIG. 7. Once the estimation model M is updated, the control device 11 shifts the process to the estimation process S2. That is, the control device 11 (beat estimation unit 21) estimates multiple beats B by executing the estimation process S2 using the updated estimation model M on the acoustic signal A. In the second and subsequent estimation processes S2, the feature F(t) calculated immediately after the start of the acoustic analysis process is applied.

[0051] As can be understood from the above explanation, each time the target beat Bn moves, the update process S8 of the estimated model M and the estimation process S2 using the updated estimated model M are repeated. Therefore, with each repetition of the estimation process S2, the position of each beat B estimated by the estimation process S2 approaches a position that reflects the instruction from the user. The beat B estimated by any one iteration of the estimation process S2 is an example of a "first beat," and the beat B estimated by the next estimation process S2 after the estimation model M is updated is an example of a "second beat."

[0052] As described above, in the first embodiment, the estimation model M is updated in accordance with the movement of the target beat Bn selected by the user among the multiple beats B estimated by the estimation process S2 and the adjacent beats Bn±1 around the target beat Bn, and the multiple beats B are re-estimated by the estimation process S2 applying the updated estimation model M. That is, when the estimation model M is updated, not only the movement of the target beat Bn but also the temporal relationship between the target beat Bn and each of the adjacent beats Bn±1 is reflected in the estimation model M. Therefore, compared to a configuration in which only the movement of the target beat Bn is reflected in the estimation model M (hereinafter referred to as the "comparative example"), it is possible to estimate a beat B that more appropriately matches the user's intention.

[0053] Specifically, in the comparative example, the estimation model M reflects a shortened interval between the target beat Bn and the immediately preceding adjacent beat Bn-1 and an expanded interval between the target beat Bn and the immediately succeeding adjacent beat Bn+1. Therefore, the estimation model M is given a tendency for the performance speed to decrease after the target beat Bn has passed (ritardando). However, when the target beat Bn is moved, the user is more likely to intend to modify the beats B throughout the entire piece of music than to intend to change the performance speed. In the first embodiment, the temporal relationship between the target beat Bn and each adjacent beat Bn±1 is reflected in the estimation model M, thereby eliminating the problem of the comparative example in which the performance speed decreases after the target beat Bn has passed. That is, as described above, compared to the comparative example, the estimation model M can estimate a beat B that more appropriately matches the user's intention. Ultimately, the user can be provided with a customer experience in which the beat B is estimated in a way that appropriately reflects the user's intention.

[0054] Furthermore, in the first embodiment, the beat area Gb is displayed on the display device 13, so the user can visually confirm how the target beat Bn and each of the adjacent beats Bn±1 move in response to instructions from the user. Therefore, the user can instruct the movement of the target beat Bn and each of the adjacent beats Bn±1 while predicting the beat B estimated by the updated estimation model M.

[0055] B: Second embodiment A second embodiment will be described. In each of the following exemplary embodiments, elements that have the same functions as those in the first embodiment will be denoted by the same reference numerals as those used in the description of the first embodiment, and detailed descriptions of each element will be omitted as appropriate.

[0056] 9 is a block diagram illustrating the functional configuration of an acoustic analysis system 100 according to the second embodiment. The control device 11 according to the second embodiment executes a program stored in the storage device 12, thereby functioning as a section setting unit 26 in addition to the same elements as those in the first embodiment (a beat estimation unit 21, a display control unit 22, a playback control unit 23, a beat editing unit 24, and an update processing unit 25).

[0057] The section setting unit 26 sets a section (hereinafter referred to as a "specific section") of the audio signal A on the time axis. Specifically, the section setting unit 26 sets the specific section in response to an instruction from the user. For example, the user can specify a specific section of the audio signal A displayed in the waveform area Ga by operating the operation device 14. The section setting unit 26 sets the section specified by the user as the specific section.

[0058] The control device 11 of the second embodiment executes the acoustic analysis process of Fig. 8 for a specific section of the audio signal A. For example, the estimation process S2 by the beat estimation unit 21 is executed in a limited manner for the specific section. That is, multiple beats B are estimated for the specific section of the music piece.

[0059] The specific procedure of the acoustic analysis process is the same as in the first embodiment. Therefore, the second embodiment also achieves the same effects as in the first embodiment. Furthermore, in the second embodiment, beat points B can be estimated in a limited manner for a partial section (specific section) of the acoustic signal A.

[0060] Although the above description illustrates an example in which a specific section is set in response to a user instruction, the method for setting the specific section is arbitrary and is not limited to the above example. For example, the section setting unit 26 may set the specific section according to a predetermined rule without requiring a user instruction. For example, the section setting unit 26 may set one of multiple structural sections of the music represented by the audio signal A as the specific section. A structural section is a section obtained by dividing the music on the time axis according to musical meaning. Examples of structural sections include an intro, an A-melody section, a B-melody section, a chorus section, and an outro section. The section setting unit 26 analyzes the audio signal A to divide the audio signal A into multiple structural sections and sets a specific structural section from among the multiple structural sections as the specific section. With the above configuration, it is possible to estimate the beat B in a limited manner for a specific structural section.

[0061] C: Modified Example Specific modified embodiments that can be added to each of the embodiments exemplified above are exemplified below. Two or more embodiments arbitrarily selected from the following examples may be appropriately combined within the scope of not being mutually contradictory.

[0062] (1) In the above-described embodiments, the adjacent beat Bn-1 immediately preceding the target beat Bn and the adjacent beat Bn+1 immediately following the target beat Bn are moved together with the target beat Bn. However, embodiments in which only one of the adjacent beat Bn-1 and the adjacent beat Bn+1 is moved together with the target beat Bn are also contemplated. For example, the beat editing unit 24 may move only the target beat Bn and the immediately preceding adjacent beat Bn-1 on the time axis in response to a user's instruction, and the update processing unit 25 may calculate the error e(t) within the application section T between the adjacent beat Bn-1 and the target beat Bn. Similarly, the beat editing unit 24 may move only the target beat Bn and the immediately following adjacent beat Bn+1 on the time axis in response to a user's instruction, and the update processing unit 25 may calculate the error e(t) within the application section T between the target beat Bn and the adjacent beat Bn+1. As can be understood from the above explanation, the beat editing unit 24 is expressed as an element that moves one or more adjacent beats Bn±1 located around the target beat Bn among the multiple beats B on the time axis.

[0063] In each of the above-mentioned forms, not only the movement of the target beat Bn, but also the temporal relationship between the target beat Bn and the immediately preceding adjacent beat Bn-1 and the temporal relationship between the target beat Bn and the immediately succeeding adjacent beat Bn+1 are reflected in the estimation model M. Therefore, compared to a form in which only the target beat Bn and one surrounding adjacent beat B are reflected in the estimation model M, the estimation model M can be updated so that a beat B that appropriately matches the user's intention can be estimated.

[0064] (2) In the above-described embodiments, a triangular distribution is used as an example of the numerical distribution D corresponding to the target beat Bn and each adjacent beat Bn±1 after the shift, but the type or shape of the numerical distribution D is not limited to the above examples. For example, a probability distribution such as a normal distribution, or a pulse-like distribution may also be used as the numerical distribution D.

[0065] (3) The type of feature F(t) calculated by the feature extraction unit 30 from the acoustic signal A is not limited to the examples in the above-described embodiments. For example, a time series of a predetermined number of samples constituting the acoustic signal A may be applied to the estimation process S2 as the feature F(t). While the above embodiments can be interpreted as the feature extraction unit 30 extracting a time series of samples from the acoustic signal A, they can also be interpreted as embodiments in which the feature extraction unit 30 is omitted, from the viewpoint that the acoustic signal A itself is partially applied to the estimation process S2.

[0066] (4) In each of the above-described embodiments, the target beat point Bn and each adjacent beat point Bn±1 are moved according to the direction of movement (forward / backward) and amount of movement δ instructed by the user. However, the method by which the user instructs the movement of the target beat point Bn and each adjacent beat point Bn±1 is not limited to the above examples.

[0067] For example, the beat editing unit 24 may move the target beat Bn and each adjacent beat Bn±1 according to the sign (±) and numerical value input by the user. When the user inputs a negative number, the beat editing unit 24 moves the target beat Bn and each adjacent beat Bn±1 forward on the time axis by a shift amount δ corresponding to the absolute value of the negative number. On the other hand, when the user inputs a positive number, the beat editing unit 24 moves the target beat Bn and each adjacent beat Bn±1 backward on the time axis by a shift amount δ corresponding to the positive number.

[0068] Furthermore, the beat editing unit 24 may move the target beat Bn and each adjacent beat Bn±1 on the time axis by a predetermined unit amount the number of times specified by the user. For example, each time a movement instruction is received from the user, the beat editing unit 24 moves the target beat Bn and each adjacent beat Bn±1 by the unit amount in the direction (forward / backward) specified by the user. Therefore, the target beat Bn and each adjacent beat Bn±1 move on the time axis by a movement amount δ that corresponds to the product of the predetermined unit amount and the number of movement instructions.

[0069] As can be understood from the above explanation, in this disclosure, "moving a beat point in response to an instruction from a user" means that the conditions for moving the beat point (for example, the direction and amount of movement) change in response to an instruction from the user, and the method of instruction and the items instructed by the user are arbitrary in this disclosure. Also, "moving a beat point" means changing the position of the beat point on the time axis.

[0070] (5) In the above-described embodiments, a deep neural network is exemplified as the estimation model M, but the configuration of the estimation model M is not limited to the above examples. For example, a statistical model such as a hidden Markov model (HMM) or a support vector machine (SVM) may also be used as the estimation model M. In the above-described embodiments, the estimation model M is updated by the update process S8. Since the estimation model M is applied to the estimation process S2, the update process S8 can also be expressed as a process of updating the estimation process S2.

[0071] (6) The acoustic analysis system 100 may be realized by a server device that communicates with an information device such as a smartphone or a tablet terminal. For example, the acoustic analysis system 100 estimates multiple beats B by analyzing an acoustic signal A received from the information device, and transmits data representing the multiple beats B to the information device.

[0072] (7) As described above, the functions of the acoustic analysis system 100 illustrated above are realized by the cooperation of one or more processors constituting the control device 11 and a program stored in the storage device 12. The program according to the present disclosure may be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium may be, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but may also include any known type of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transient, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage device in the distribution device that stores the program corresponds to the non-transitory recording medium described above.

[0073] D: Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0074] An acoustic analysis system according to one aspect (aspect 1) of the present disclosure includes: a beat estimation unit that estimates a plurality of first beats by performing estimation processing on an acoustic signal; a beat editing unit that moves a target beat selected by a user from among the plurality of first beats and one or more adjacent beats located around the target beat from among the plurality of first beats on a time axis in accordance with an instruction from the user; and an update processing unit that updates the estimation processing in accordance with the movement of the target beat and the one or more adjacent beats, and the beat estimation unit estimates a plurality of second beats by performing the updated estimation processing on the acoustic signal.

[0075] According to the above aspect, the estimation process is updated in accordance with the movement on the time axis of the target beat selected by the user and one or more adjacent beats located around the target beat, and multiple second beats are estimated by the updated estimation process. In updating the estimation process, not only the movement of the target beat but also the temporal relationship between the target beat and one or more adjacent beats is reflected in the estimation process. Therefore, compared to a configuration in which only the movement of the target beat is reflected in the estimation process, it is possible to estimate a second beat that more appropriately matches the user's intention. Note that the acoustic analysis system may also be referred to as an acoustic analysis device. It does not matter whether the "acoustic analysis system" or the "acoustic analysis device" is composed of a single device or multiple separate devices.

[0076] "Estimation processing" is processing for estimating multiple beats (first beat / second beat) from an audio signal. For example, an example of "estimation processing" is processing that uses an estimation model that has learned the relationship between the feature of a training audio signal and the probability that the time point at which the feature is observed corresponds to a beat. Specifically, by processing the feature at a specific time point in the audio signal to be processed using the estimation model, the probability that the time point corresponds to a beat is output.

[0077] "Updating the estimation process" refers to the process of updating the elements applied to the estimation process. For example, assuming an estimation process that uses an estimation model, machine learning that updates the variables that define the estimation model corresponds to "updating the estimation process."

[0078] In a specific example of aspect 1 (aspect 2), the one or more adjacent beat points include a first beat point among the plurality of first beat points that is located immediately before the target beat point, and a first beat point among the plurality of first beat points that is located immediately after the target beat point. In the above aspect, not only the movement of the target beat point but also the temporal relationship between the target beat point and the immediately preceding first beat point and the temporal relationship between the target beat point and the immediately succeeding first beat point are reflected in the estimation process. Therefore, compared to a configuration in which only the target beat point and one surrounding adjacent beat point are reflected in the estimation process, the estimation process can be updated so that a second beat point that appropriately matches the user's intention can be estimated.

[0079] An acoustic analysis system according to a specific example (aspect 3) of aspect 1 or aspect 2 further includes a display control unit 22 that displays a beat image representing the target beat and the one or more adjacent beats on a display device 13 and moves the target beat and the one or more adjacent beats included in the beat image in response to an instruction from the user. In the above aspect, the user can visually confirm how the target beat and the one or more adjacent beats move in response to an instruction from the user. Therefore, the user can instruct the movement of the target beat and the one or more adjacent beats while predicting the second beat estimated by the updated estimation process.

[0080] In a specific example (Aspect 4) of any of Aspects 1 to 3, the estimation process includes a first process of generating a probability that a time point at which the feature is observed corresponds to a beat by processing the feature at each time point of the audio signal using an estimation model that has learned the relationship between the feature of the training audio signal and the probability that the time point at which the feature is observed corresponds to a beat, and a second process of identifying the multiple first beat points from the time series of the probabilities generated by the first process. In the above aspect, multiple beat points are estimated from the audio signal using the estimation model that has learned the relationship between the feature of the training audio signal and the probability that the time point at which the feature corresponds to a beat. Therefore, multiple beat points (first beat points / second beat points) can be estimated with high accuracy for an unknown audio signal whose feature changes in a variety of ways.

[0081] In a specific example (Aspect 5) of Aspect 4, the update processing unit updates the estimation model so as to reduce an error between a numerical distribution set on the time axis corresponding to the target beat point and the one or more adjacent beat points after the movement and a time series of the probability estimated by the first process. In the above aspect, the estimation model is updated so as to reduce an error between a numerical distribution corresponding to the target beat point and adjacent beat points and a time series of the probability estimated by the first process, so that the movement of the target beat point and adjacent beat points can be appropriately reflected in the estimation model.

[0082] A "numerical distribution" is a distribution of numerical values ​​on a time axis. The type and shape of the numerical distribution are arbitrary. For example, triangular distribution, normal distribution, or pulse-like distribution are examples of "numerical distribution." With regard to a numerical distribution, "set corresponding to (target / adjacent) beat points" means that the positions of the beat points on the time axis and the positions of the numerical distribution on the time axis correspond to each other. In other words, the positions of the numerical distribution on the time axis change in conjunction with changes in the positions of the beat points on the time axis. For example, a relationship in which the maximum point of the numerical distribution coincides with a beat point is a typical example of the relationship "set corresponding to a beat point."

[0083] An acoustic analysis system according to a specific example (Aspect 6) of any of Aspects 1 to 5 further includes a section setting unit 26 that sets a specific section that is a portion of the acoustic signal on the time axis, and the beat estimation unit performs estimation processing on the specific section. According to the above aspect, it is possible to estimate the second beat in a limited manner for a portion of the acoustic signal.

[0084] A "specific section" is any part of an audio signal on the time axis. For example, a section specified by a user is an example of a "specific section." Furthermore, beat points may be estimated by designating one of multiple structural sections of a piece of music represented by an audio signal as a "specific section." A structural section is a section into which a piece of music is divided on the time axis according to musical meaning. Examples of structural sections include an intro, verse, bridge, chorus, and outro.

[0085] An acoustic analysis method according to one aspect of the present disclosure estimates multiple first beat points by performing an estimation process on an acoustic signal, moves a target beat point selected by a user from among the multiple first beat points and one or more adjacent beat points from among the multiple first beat points located around the target beat point on a time axis in accordance with an instruction from the user, updates the estimation process in accordance with the movements of the target beat point and the one or more adjacent beat points, and estimates multiple second beat points by performing the updated estimation process on the acoustic signal. Note that each aspect exemplified for the acoustic analysis system is similarly applicable to the acoustic analysis method according to the present disclosure.

[0086] A program according to one aspect of the present disclosure causes a computer system to function as: a beat estimation unit that estimates multiple first beat points by performing estimation processing on an audio signal; a beat editing unit that moves a target beat point selected by a user from among the multiple first beat points and one or more adjacent beat points located around the target beat point from among the multiple first beat points on a time axis in accordance with instructions from the user; and an update processing unit that updates the estimation processing in accordance with the movements of the target beat point and the one or more adjacent beat points, wherein the beat estimation unit estimates multiple second beat points by performing the updated estimation processing on the audio signal. Note that each of the aspects exemplified for the acoustic analysis system is similarly applicable to the program according to the present disclosure. [Explanation of symbols]

[0087] 100...acoustic analysis system, 11...control device, 12...storage device, 13...display device, 14...operation device, 15...sound emission device, 21...beat point estimation unit, 22...display control unit, 23...playback control unit, 24...beat point editing unit, 25...update processing unit, 26...section setting unit, 30...feature extraction unit, 31...first processing unit, 32...second processing unit.

Claims

1. a beat point estimation unit that estimates a plurality of first beat points by performing estimation processing on an audio signal; a beat editing unit that moves a target beat selected by a user from among the plurality of first beats and one or more adjacent beats located around the target beat from among the plurality of first beats on a time axis in response to an instruction from the user; an update processing unit that updates the estimation process in accordance with the movement of the target beat point and the one or more adjacent beat points; The beat estimation unit estimates a plurality of second beats by performing the updated estimation process on the audio signal. Acoustic analysis system.

2. The one or more adjacent beat points are a first beat point located immediately before the target beat point among the plurality of first beat points; a first beat point located immediately after the target beat point among the plurality of first beat points The acoustic analysis system of claim 1.

3. a display control unit that displays a beat image representing the target beat and the one or more adjacent beats on a display device, and moves the target beat and the one or more adjacent beats included in the beat image in response to an instruction from the user; 3. The acoustic analysis system according to claim 1, further comprising:

4. The estimation process includes: a first process for generating a probability that a time point at which the feature is observed corresponds to a beat by processing the feature at each time point of the training audio signal using an estimation model that has learned the relationship between the feature of the training audio signal and the probability that the time point at which the feature is observed corresponds to a beat; and a second process for identifying the plurality of first beat points from the time series of probabilities generated by the first process.

3. The acoustic analysis system according to claim 1.

5. The update processing unit: a numerical distribution set on a time axis corresponding to the target beat point and the one or more adjacent beat points after the movement; a time series of probabilities estimated by the first process; Update the estimation model so that the error in The acoustic analysis system of claim 4.

6. a section setting unit that sets a specific section that is a part of the acoustic signal on a time axis; Further comprising: The beat point estimation unit performs the estimation process for the specific section. The acoustic analysis system of claim 1.

7. estimating a plurality of first beat points by performing estimation processing on the acoustic signal; moving a target beat point selected by a user from among the plurality of first beat points and one or more adjacent beat points located around the target beat point from among the plurality of first beat points on a time axis in accordance with an instruction from the user; updating the estimation process in accordance with the movement of the target beat point and the one or more adjacent beat points; The updated estimation process is performed on the acoustic signal to estimate a plurality of second beat points. An acoustic analysis method implemented by a computer system.

8. The one or more adjacent beat points are a first beat point located immediately before the target beat point among the plurality of first beat points; a first beat point located immediately after the target beat point among the plurality of first beat points The acoustic analysis method according to claim 7.

9. moreover, displaying beat images representing the target beat and the one or more adjacent beats on a display device; The target beat point and the one or more adjacent beat points included in the beat point image are moved in accordance with an instruction from the user. The acoustic analysis method according to claim 7 or 8.

10. The estimation process includes: a first process for generating a probability that a time point at which the feature is observed corresponds to a beat by processing the feature at each time point of the training audio signal using an estimation model that has learned the relationship between the feature of the training audio signal and the probability that the time point at which the feature is observed corresponds to a beat; and a second process for identifying the plurality of first beat points from the time series of probabilities generated by the first process. The acoustic analysis method according to claim 7 or 8.

11. In updating the estimation process, a numerical distribution set on a time axis corresponding to the target beat point and the one or more adjacent beat points after the movement; a time series of probabilities estimated by the first process; Update the estimation model so that the error in The acoustic analysis method of claim 10.

12. moreover, a specific section is set as a section of the acoustic signal on a time axis; The estimation process is performed for the specific section. The acoustic analysis method according to claim 7.

13. a beat point estimation unit that estimates a plurality of first beat points by performing estimation processing on the audio signal; a beat editing unit that moves a target beat selected by a user from among the plurality of first beats and one or more adjacent beats located around the target beat from among the plurality of first beats on a time axis in response to an instruction from the user; and an update processing unit that updates the estimation process in accordance with the movement of the target beat point and the one or more adjacent beat points; A program that causes a computer system to function as The beat estimation unit estimates a plurality of second beats by performing the updated estimation process on the audio signal. program.

Citation Information

Patent Citations

  • Music information processing apparatus and music information processing method

    JP2013171070A

  • Acoustic signal analysis device and acoustic signal analysis program

    JP2015114361A