Signal processing system, signal processing method, and program

The signal processing system synchronizes audio or video playback with user performance by detecting and adjusting the playback speed and volume, addressing the challenge of synchronizing with user actions.

JP7679870B2Active Publication Date: 2025-05-20YAMAHA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023505085
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-09
Filing Date
2021-06-23
Publication Date
2025-05-20
Estimated Expiration
2041-06-23

AI Technical Summary

Technical Problem

Existing techniques fail to synchronize the playback of audio or video signals with a user's performance, such as playing a piece of music, effectively following the user's actions.

Method used

A signal processing system that includes an acquisition unit to detect the user's performance position and a control unit to perform time stretching of the audio or video signal accordingly, ensuring the playback follows the user's actions.

Benefits of technology

The system enables the playback to accurately follow the user's performance, maintaining auditory or visual naturalness by adjusting the playback speed and volume based on the user's actions, thereby enhancing the synchronization experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679870000010
    Figure 0007679870000010
  • Figure 0007679870000011
    Figure 0007679870000011
  • Figure 0007679870000012
    Figure 0007679870000012
Patent Text Reader

Abstract

This signal processing system causes a playback device to play back a time-series signal following the playback of music, and is equipped with: an acquisition unit that acquires a position designated by a user when playing back music; and a control unit that executes time expansion and contraction of the time-series signal according to the designated position.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a technique for processing time domain signals (hereinafter referred to as "time series signals"), such as audio signals or video signals. [Background technology]

[0002] Various techniques have been proposed for estimating the position on the time axis where a user is playing a piece of music (hereinafter referred to as the "playing position"). For example, Patent Document 1 discloses a technique for estimating the playing position by analyzing an audio signal representing the playing sound of a piece of music. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2015-79183 A Summary of the Invention [Problem to be solved by the invention]

[0004] For example, there is a demand for the playback of audio represented by an audio signal or video represented by a video signal to follow (synchronize) with a performance by a user. In consideration of the above circumstances, one aspect of the present disclosure has an object to make a time-series signal, such as an audio signal or a video signal, follow the action of a user. [Means for solving the problem]

[0005] In order to solve the above problems, a signal processing system according to one aspect of the present disclosure is a signal processing system that causes a playback device to play a time series signal in response to the playback of a piece of music, and includes an acquisition unit that acquires a position indicated by a user during the playback of the piece of music, and a control unit that performs time stretching of the time series signal in accordance with the indicated position.

[0006] A signal processing method according to one aspect of the present disclosure is a method for having a playback device play a time-series signal in response to playback of a piece of music, the method obtaining a position indicated by a user during playback of the piece of music, and performing time stretching of the time-series signal in accordance with the indicated position.

[0007] A program according to one aspect of the present disclosure is a program for causing a playback device to play a time series signal in response to the playback of a piece of music, and causes a computer to function as an acquisition unit that acquires a position indicated by a user during the playback of the music, and a control unit that performs time stretching of the time series signal in accordance with the indicated position. [Brief description of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating the configuration of a performance system according to a first embodiment. [Diagram 2] FIG. 2 is a block diagram illustrating a functional configuration of a signal processing system. [Diagram 3] 11 is an explanatory diagram of a process executed by an acquisition unit and an identification unit; FIG. [Figure 4] 10 is a flowchart illustrating a specific procedure of a control process. [Diagram 5] FIG. 11 is an explanatory diagram of a process for identifying a playback position. [Figure 6] 11 is a flowchart illustrating a specific procedure of a specification process. [Figure 7] 13 is a flowchart illustrating some specific steps of the probability setting process. [Figure 8] 13 is a flowchart illustrating some other specific steps of the probability setting process. [Figure 9] FIG. 13 is an explanatory diagram of an inter-utterance period. [Figure 10] 10 is a flowchart illustrating a specific procedure of a playback process. [Figure 11] FIG. 13 is an explanatory diagram of operation strength. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] A: First Embodiment FIG. 1 is a block diagram illustrating the configuration of a performance system 100 according to the first embodiment. The performance system 100 is a computer system for a user to perform a piece of music (hereinafter referred to as “target music”), and includes a keyboard instrument 10 and a signal processing system 20. The keyboard instrument 10 and the signal processing system 20 are connected to each other, for example, by wire or wirelessly.

[0010] The keyboard instrument 10 is an electronic instrument having a plurality of keys corresponding to different pitches. The user performs the target music by sequentially operating each key of the keyboard instrument 10. Specifically, the user performs, by the keyboard instrument 10, one or more specific performance parts among a plurality of performance parts constituting the target music. The keyboard instrument 10 emits a sound of the pitch played by the user (for example, an instrument sound). In addition, in parallel with the emission of the sound according to the performance by the user, the keyboard instrument 10 supplies performance data D representing the performance to the signal processing system 20. The performance data D is instruction data specifying the pitch corresponding to the key operated by the user and the key pressing intensity, and is generated for each operation of the keyboard instrument 10 by the user. That is, a time series of the performance data D is supplied from the keyboard instrument 10 to the signal processing system 20. The performance data D is, for example, event data conforming to the MIDI (Musical Instrument Digital Interface) standard.

[0011] The signal processing system 20 includes a control device 21, a storage device 22, and a sound emission device 23. The signal processing system 20 is realized by, for example, a portable information device such as a smartphone or a tablet terminal, or a portable or stationary information device such as a personal computer. Note that the signal processing system 20 is realized not only as a single device but also as a plurality of devices separately configured from each other. In addition, the signal processing system 20 may be mounted on the keyboard instrument 10.

[0012] The control device 21 is composed of one or more processors that control each element of the signal processing system 20. For example, the control device 21 is composed of one or more types of processors such as a CPU (Central Processing Unit), an SPU (Sound Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0013] The storage device 22 is one or more memories that store the programs executed by the control device 21 and various data used by the control device 21. The storage device 22 is configured with a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. Note that the storage device 22 may be a portable recording medium that is detachable from the signal processing system 20, or a recording medium (e.g., cloud storage) that the control device 21 can write or read via a communication network such as the Internet.

[0014] The storage device 22 stores an audio signal X representing the performance sound of the target song. The audio signal X is a time-series signal (i.e., a sample sequence) representing the waveform of the performance sound of the target song. Specifically, the audio signal X represents musical sounds produced by various musical instruments when the target song is played, or a singing voice produced by a singer when singing the target song. For example, the audio signal X represents the performance sounds of one or more performance parts other than the performance part played by the user on the keyboard instrument 10 among the multiple performance parts constituting the target song.

[0015] The sound emitting device 23 reproduces the sound instructed by the control device 21. The sound emitting device 23 is, for example, a speaker or a headphone. Note that the sound emitting device 23, which is separate from the signal processing system 20, may be connected to the signal processing system 20 by wire or wirelessly.

[0016] The control device 21 of the first embodiment reproduces the audio signal X on the sound output device 23 in accordance with the performance of the target music piece by the user. Specifically, the control device 21 estimates a position (performance position P[t]) of the target music piece corresponding to the performance by the user, and reproduces a portion Y of the audio signal X corresponding to a position on the time axis (playback position R[t]) corresponding to the estimated position on the time axis on the sound output device 23. That is, the audio signal X is expanded or contracted (time stretched) on the time axis in accordance with the performance of the target music piece by the user. For example, when the speed of the performance by the user is lower than a predetermined standard speed (hereinafter referred to as "standard speed") P0, the audio signal X is extended on the time axis. That is, the slower the movement speed of the performance position P[t], the slower the movement speed of the play position R[t] on the time axis, and as a result, the audio signal X is extended on the time axis. On the other hand, when the speed of the performance by the user is higher than the standard speed P0, the audio signal X is shortened on the time axis. That is, the faster the performance position P[t] moves, the faster the playback position R[t] moves on the time axis, resulting in a shortening of the audio signal X on the time axis. As described above, the playback of the audio signal X by the sound emitting device 23 follows the performance by the user, creating an atmosphere as if the signal processing system 20 and the user were playing together in a cooperative manner.

[0017] 2 is a block diagram illustrating an example of a functional configuration of the signal processing system 20. The control device 21 executes a program stored in the storage device 22 to realize a plurality of functions (an analysis unit 31, an acquisition unit 32, and a control unit 33) for reproducing an audio signal X in response to a performance of the keyboard instrument 10 by a user.

[0018] The analysis unit 31 generates indexes W[n] (Wa[n], Wb[n], Wc[n]) by analyzing the acoustic signal X. Indexes W[n] (n=1 to N) are generated for each of N periods (hereinafter referred to as "unit periods") U[1] to U[N] into which the acoustic signal X is divided on the time axis. Each unit period U[n] has a predetermined length. The symbol n means the number (frame number) of the unit period U[n]. Unit periods U[n-1] and U[n] that are successive on the time axis partially overlap each other. However, unit periods U[n-1] and U[n] may be consecutive to each other without overlapping.

[0019] Each index W[n] is a variable (feature amount) related to the acoustic characteristics of the audio signal X in the unit period U[n]. Before playing back the audio signal X, the analysis unit 31 generates the index W[n] (W[1] to W[N]) for each unit period U[n] and stores each index W[n] in the storage device 22. Specifically, the analysis unit 31 calculates the sound presence index Wa[n], the fluctuation index Wb[n], and the onset point index Wc[n] as the index W[n] for each unit period U[n].

[0020] The voice activity index Wa[n] is a variable that binary-value indicates whether the audio signal X is voiced or unvoiced in the unit period U[n]. That is, the voice activity index Wa[n] is set to a value of "1" when the unit period U[n] is voiced, and is set to a value of "0" when the unit period U[n] is unvoiced. The voice activity index Wa[n] is calculated using a well-known voice activity detection (VAD: Voice Activity Detection). Note that the probability that the audio signal X is voiced in the unit period U[n] (for example, a value between 0 and 1) may be used as the voice activity index Wa[n].

[0021] The fluctuation index Wb[n] is a variable that indicates the degree of fluctuation of the acoustic characteristics in the audio signal X. For example, the amount of fluctuation of the acoustic characteristics between the adjacent unit periods U[n-1] and U[n] is calculated as the fluctuation index Wb[n] of the unit period U[n]. Therefore, the more the acoustic characteristics of the audio signal X are likely to fluctuate, the larger the fluctuation index Wb[n] is set to. The acoustic characteristics are, for example, the intensity spectrum or frequency characteristics such as MFCC (Mel-Frequency Cepstrum Coefficients) of the audio signal X. Note that the amount of fluctuation of the acoustic characteristics such as the fundamental frequency of the audio signal X may be used as the fluctuation index Wb[n]. A known analysis technique such as discrete Fourier transform is used to calculate the fluctuation index Wb[n]. The fact that the acoustic characteristics are likely to fluctuate means that the acoustic characteristics of the audio signal X are likely to fluctuate unstably. Therefore, the fluctuation index Wb[n] can also be said as an index of the stability or instability of the acoustic characteristics in the audio signal X.

[0022] The onset point index Wc[n] is a variable that binary-value indicates whether or not the unit period U[n] of the audio signal X corresponds to an onset point. The onset point is the point at which the audio component included in the audio signal X starts to be produced (onset), or in other words, the point at which the audio component rises (attack). Any known analysis technique may be used to calculate the onset point index Wc[n]. For example, the point at which the volume of the audio signal X increases sharply is detected as the onset point. Note that the probability that the unit period U[n] of the audio signal X is an onset point (for example, a numerical value between 0 and 1) may be used as the onset point index Wc[n].

[0023] FIG. 3 is an explanatory diagram of an overview of the processing of the acquisition unit 32 and the control unit 33 in FIG. 2. The acquisition unit 32 acquires a performance position P[t] as time passes. Specifically, the acquisition unit 32 identifies a performance position P[t] in a target piece of music by analyzing the time series of the performance data D sequentially supplied from the keyboard instrument 10. The symbol t means any one of a plurality of time points set at equal intervals on the time axis. That is, the acquisition unit 32 identifies a performance position P[t] for each of a plurality of time points t on the time axis. The time point t is expressed by the number of each time point set on the time axis. The performance position P[t] means an elapsed time (for example, seconds) based on the start point of the audio signal X. The identification of the performance position P[t] by the acquisition unit 32 is repeated in parallel with the performance of the target piece of music by the user and the reproduction of the audio signal X. The speed at which the performance position P[t] moves on the time axis is a variable value according to the performance by the user.

[0024] The acquisition unit 32 of the first embodiment estimates (i.e., predicts) a playing position P[t+d] at a time point (t+d) that is a predetermined length d forward from the time point t on the time axis. The predetermined length d is a predetermined positive number corresponding to an integer number of times t. A publicly known analysis technique (score alignment technique) is arbitrarily adopted for the estimation of the playing position P[t] by the acquisition unit 32. For example, the analysis technique disclosed in JP 2016-099512 A is used for estimating the playing position P[t]. The acquisition unit 32 may also estimate the playing position P[t] by using a statistical estimation model such as a deep neural network (DNN) or a hidden Markov model (HMM).

[0025] 2 executes time expansion / contraction of the audio signal X in accordance with the performance position P[t]. The control unit 33 of the first embodiment includes a determination unit 331 and a reproduction unit 332.

[0026] The determination unit 331 in Fig. 2 determines a playback position R[t] according to the performance position P[t]. The determination unit 331 determines the playback position R[t] for each of a plurality of time points t on the time axis. The playback position R[t] is an elapsed time (e.g., seconds) based on the start point of the audio signal X. In other words, the playback position R[t] means that a point in the audio signal X where the time R[t] has elapsed from the start point should be reproduced at one time point t on the time axis. The determination unit 331 determines the playback position R[t] from the performance position P[t] so that the playback position R[t] is roughly close to the performance position P[t] and the auditory naturalness of the reproduced sound of the audio signal X is maintained.

[0027] Figure 3 shows a processing period Q and an analysis period q. The processing period Q is a period between time point t1 and time point t2 on the time axis. Time point t1 corresponds to the current time during the reproduction of the acoustic signal X. Time point t2 is located behind time point t1. Specifically, time point t2 is time point t which is behind time point t1 by a predetermined length d. That is, the processing period Q is a period of a predetermined length d. As described above, at time point t1, the performance position P[t] up to time point (t1 + d) has been estimated by the acquisition unit 32. That is, at time point t1, the performance position P[t] has been estimated for each time point t within the processing period Q starting from the time point t1. On the other hand, when time point t1 arrives, the reproduction position R[t] has not been specified for each time point t within the processing period Q. Note that time point t1 is an example of the "first time point", and time point t2 is an example of the "second time point".

[0028] The analysis period q is a period from time point t1 to time point t3. Time point t3 is located between time point t1 and time point t2. Specifically, time point t3 is time point t which is behind time point t1 by the number of time points t less than the predetermined length d. That is, the analysis period q is a partial period on the start point (t1) side of the processing period Q. In FIG. 3, a case where time point t3 is closer to time point t2 than time point t1 is illustrated, but the position of time point t3 within the processing period Q is arbitrary. For example, the time point t immediately after time point t1 may be set as time point t3. Time point t3 is an example of the "third time point".

[0029] The specifying unit 331 estimates the time series of the reproduction position R[t] for each time point t within the analysis period q among the processing period Q in which the performance position P[t] has been estimated, according to the time series of the performance position P[t] within the processing period Q. That is, for each analysis period q on the time axis, the time series of the reproduction position R[t] corresponding to each time point t within the analysis period q is specified. Note that in the form where time point t3 is the time point t immediately after time point t1, the reproduction position R[t] is specified for each time point t on the time axis.

[0030] Incidentally, the accuracy with which the acquisition unit 32 estimates the playing position P[t] decreases as the time t moves away from the current time t1 on the time axis. In consideration of the above circumstances, in the first embodiment, the time series of the playing position R[t] in the analysis period q from time t1 to time t3 is estimated according to the time series of the playing position P[t] in the processing period Q from time t1 to time t2. Therefore, the influence (noise) of the estimation error of the playing position P[t] in the period near the end point of the processing period Q is reduced. In other words, the playing position R[t] can be appropriately specified compared to a configuration in which the time series of the playing position P[t] in the processing period Q is used to specify the time series of the playing position R[t] throughout the processing period Q.

[0031] The reproducing unit 332 in Fig. 2 causes the sound emitting device 23 to reproduce a portion Y of the audio signal X corresponding to a reproduction position R[t]. Specifically, at each of a plurality of time points t on the time axis, the reproducing unit 332 causes the sound emitting device 23 to reproduce a portion Y of the audio signal X including the reproduction position R[t] at the time point t. The portion Y is composed of a time series of samples within a period corresponding to the reproduction position R[t] of the audio signal X. For convenience, a D / A converter that converts the portion Y of the audio signal X from digital to analog and an amplifier that amplifies the converted signal are omitted from the illustration. In the following description, it is assumed that the audio signal X is reproduced in units of a predetermined time length (hop length) Ht.

[0032] 4 is a flowchart illustrating a specific procedure of a process (hereinafter referred to as a "control process") S executed by the control device 21 to play back an audio signal X. For example, the control process S is started in response to an instruction from a user. When the control process S is started, the analysis unit 31 generates an index W[n] (Wa[n], Wb[n], Wc[n]) for each of N unit periods U[1] to U[N] by analyzing the audio signal X stored in the storage device 22 (Sa).

[0033] The determination unit 331 sets the transition probability τ[n1, n2] by analyzing the audio signal X (Sb). The transition probability τ[n1, n2] is the probability that when a unit period U[n1] of the audio signal X is reproduced at a time point (t-1) on the time axis, a unit period U[n2] of the audio signal X is reproduced at the immediately following time point t (n1, n2 = 1 to N). In other words, the transition probability τ[n1, n2] means the likelihood that the reproduction position R[t] transitions from the unit period U[n1] of the audio signal X to the unit period U[n2]. The determination unit 331 calculates the transition probability τ[n1, n2] for all combinations of selecting two unit periods U[n] (U[n1] and U[n2]) from the N unit periods U[1] to U[N] of the audio signal X. In addition, the unit period U[n2] is a unit period U[n] (n2>n1) located after the unit period U[n1], or a unit period U[n] (n2=n1) that coincides with the unit period U[n1]. The closer the unit period U[n1] and the unit period U[n2] related to the transition probability τ[n1,n2] are on the time axis, the greater the degree of extension of the audio signal X. In addition, the transition probability τ[n,n] (n1=n2) where the number n1 and the number n2 are common means the probability that the playback position R[t] stays at the unit period U[n]. As can be understood from the above explanation, the playback position R[t] moves backward on the time axis. However, the movement of the playback position R[t] in the retroactive direction (past) on the time axis may be allowed.

[0034] The calculation of the index W[n] (Sa) and the setting of the transition probability τ[n1, n2] (Sb) may be performed before the start of the control process S. The calculation of the index W[n] (Sa) and the setting of the transition probability τ[n1, n2] (Sb) may be performed in reverse order. The index W[n] and the transition probability τ[n1, n2] are stored in the storage device 22. After performing the preparatory processes (Sa, Sb) described above, the acquisition unit 32 estimates the performance position P[t+d] for each time point t on the time axis (Sc).

[0035] The identification unit 331 executes the identification process Sd. The identification process Sd is a process for identifying a time series of a playback position R[t] within an analysis period q according to each index W[n] of the audio signal X and a time series of a performance position P[t] within a processing period Q. The identification process Sd is executed for each analysis period q on the time axis. The reproduction unit 332 causes the sound emitting device 23 to reproduce (Se) a portion Y of the audio signal X that corresponds to each playback position R[t] identified by the identification process Sd.

[0036] The control device 21 judges whether a predetermined end condition is met (Sf). The end condition is, for example, when an end instruction is received from the user or when the entire reproduction of the audio signal X is completed. If the end condition is not met (Sf: NO), the control device 21 shifts the process to step SC. That is, the estimation of the performance position P[t+d] (Sc), the determination of the reproduction position R[t] within the analysis period q (Sd), and the reproduction of the portion Y of the audio signal X (Se) are repeated. On the other hand, if the end condition is met (Sf: YES), the control device 21 ends the control process S.

[0037] Each time the control device 21 shifts the process to step SC (Sf: NO), it sets an immediately following processing period Q starting from the end point of the current analysis period q (i.e., the period in which the time series of the playback position R[t] has been identified), and further sets an analysis period q within the processing period Q. That is, for each of the multiple processing periods Q on the time axis, the identification unit 331 identifies the time series of the playback position R[t] within the analysis period q of the processing period Q.

[0038] As described above, in the first embodiment, the portion Y of the audio signal X corresponding to the playback position R[t] according to the user's performance position P[t] is played by the sound emitting device 23. That is, the audio signal X is expanded or contracted on the time axis according to the performance of the target song by the user. Therefore, it is possible to make the playback of the audio signal X by the sound emitting device 23 follow the performance of the target song by the user.

[0039] The determination of the playback position R[t] will be described in detail below. In the following description, functions F(P[t]) and E(n) are used. The function F(P[t]) is a function for converting the performance position P[t] (seconds) into the number n of the unit period U[n] in the audio signal X, and is expressed, for example, by the following formula (1).

number

[0040] On the other hand, the function E(n) is a function for converting the number n of the unit period U[n] into an elapsed time (e.g., seconds) based on the start point of the acoustic signal X, and is expressed, for example, by the following equation (2).

number

[0041] FIG. 5 is an explanatory diagram of the aforementioned identification process Sd. FIG. 5 illustrates each time point t (..., t-2, t-1, t, t+1, t+2,...) on the time axis and each unit period U[n] (..., U[n-2], U[n-1], U[n], U[n+1], U[n+2],...) of the acoustic signal X. The identification process Sd of the first embodiment includes a process (hereinafter referred to as "path search") Sd2 for searching a most likely path (hereinafter referred to as "most likely path") C consisting of different combinations of each unit period U[n] and each time point t. The most likely path C is expressed by a time series of multiple position variables c[t] corresponding to different time points t on the time axis. The position variable c[t] specifies one of N unit periods U[1] to U[N] of the acoustic signal X (c[t]=1 to N). The route search Sd2 utilizes dynamic programming such as the Viterbi algorithm or beam search.

[0042] 6 is a flowchart illustrating a specific procedure of the identification process Sd. When the identification process Sd is started, the identification unit 331 calculates an observation likelihood L[t,n] for each time point t in the processing period Q (Sd1). The observation likelihood L[t,n] is the likelihood that the n-th unit period U[n] of N unit periods U[1] to U[N] of the acoustic signal X should be reproduced at the time point t. In other words, the observation likelihood L[t,n] means the probability that each unit period U[n] of the acoustic signal X corresponds to the reproduction position R[t] at the time point t.

[0043] The identification unit 331 estimates the maximum likelihood path C by the path search Sd2. The path search Sd2 applies the observation likelihood L[t, n] at each time point t in the processing period Q and the transition probability τ[n1, n2] of the audio signal X. As described above, in the first embodiment, the time series of the playback position R[t] can be appropriately identified by the path search Sd2 that applies the transition probability τ[n1, n2] for each combination of two unit periods U[n] (U[n1], U[n2]) of the audio signal X.

[0044] In the path search Sd2, the determination unit 331 searches for the most likely path C under a constraint condition that fixes the position variable c[t1] at the start point (time t1) of the processing period Q and the position variable c[t2] at the end point (time t2) of the processing period Q. Specifically, the position variable c[t1] at time t1 is fixed to a numerical value F(P[t1]) obtained by converting the performance position P[t1] estimated for the time t1 by the function F(P[t]) of Formula (1). In addition, the position variable c[t2] at time t2 is fixed to a numerical value F(P[t2]) obtained by converting the performance position P[t2] estimated for the time t2 by the function F(P[t]) of Formula (1).

[0045] As described above, the maximum likelihood path C is expressed by a time series of position variables c[t] corresponding to different time points t in the analysis period q. The determination unit 331 calculates the playback position R[t] for each time point t in the analysis period q by converting the number n of the unit period U[n] specified by each position variable c[t] by the function E(n) (Sd3). That is, the determination unit 331 of the first embodiment determines the time series of the playback position R[t] in the analysis period q under the constraint condition that the playback position R[t1] at the time point t1 of the analysis period q is fixed to the performance position P[t1] at the time point t1, and the playback position R[t2] at the time point t2 of the analysis period q is fixed to the performance position P[t2] at the time point t2, as illustrated in FIG. 3. With the above configuration, the possibility that the playback position R[t] deviates excessively from the performance position P[t] in the analysis period q is reduced.

[0046] As described above, in the first embodiment, the path search Sd2 for identifying the time series of the playback position R[t] is executed for each processing period Q on the time axis. Therefore, even if the speed of the movement of the performance position P[t] fluctuates irregularly, the playback position R[t] that follows the performance by the user with high accuracy can be identified.

[0047] The observation likelihood L[t,n] and transition probability τ[n1,n2] are described in detail below.

[0048] (1) Calculation of observation likelihood L[t,n] (Sd1) As described above, the observation likelihood L[t,n] is the likelihood that the unit period U[n] of the acoustic signal X should be reproduced at each time point t on the time axis. The determination unit 331 calculates the observation likelihood L[t,n] for each of the multiple time points t on the time axis by calculating the following formula (3).

number

[0049] Formula (3)means that the observation likelihood L[t,n] follows a normal distribution (Normal) with the number n of the unit period U[n] as a random variable. The average of the probability distribution of the observation likelihood L[t,n] is set to a numerical value F(P[t]) obtained by converting the performance position P[t] estimated by the acquisition unit 32 into the number n of the unit period U[n]. In other words, the average of the probability distribution of the observation likelihood L[t,n] is set according to the performance position P[t]. With the above configuration, the possibility that the playback position R[t] deviates excessively from the performance position P[t] within the analysis period q is reduced.

[0050] Moreover, the variance σ(Wb[n],O) of the probability distribution of the observation likelihood L[t,n] is expressed by a function with the above-mentioned fluctuation index Wb[n] and the sounding point group O as variables. The sounding point group O is a set of time points t corresponding to the performance positions P[t] corresponding to the sounding points of the audio signal X. In other words, each time point t constituting the sounding point group O satisfies the following formula (4a) and formula (4b).

number

[0051] The variance σ(Wb[n],O) of the probability distribution regarding the observation likelihood L[t,n] is expressed, for example, by the following equation (5).

number

[0052] As can be seen from formula (5), when time t corresponds to the sounding point (t∈O), the second term on the right side of formula (5) is eliminated, and therefore the variance σ(Wb[n],O) is set to a sufficiently small value ε. On the other hand, when time t does not correspond to the sounding point, the first term on the right side of formula (5) is eliminated, and therefore the variance σ(Wb[n],O) is set to a value 1 / Wb[n] according to the fluctuation index Wb[n]. The value ε of the variance σ(Wb[n],O) when time t corresponds to the sounding point is lower than the value 1 / Wb[n] of the variance σ(Wb[n],O) when time t does not correspond to the sounding point. The variance ε of the probability distribution when time t corresponds to the sounding point is an example of the "first variance", and the variance 1 / Wb[n] of the probability distribution when time t does not correspond to the sounding point is an example of the "second variance".

[0053] Therefore, at time t (t∈O) corresponding to the sound generation point, the observation likelihood L[t,n] is locally high near the average F(P[t]) of the random variable n. In other words, at time t corresponding to the sound generation point, the possibility that the playback position R[t] is close to or coincides with the performance position P[t] is sufficiently high compared to the possibility that the playback position R[t] deviates from the performance position P[t]. This has the advantage that the playback of the audio signal X can easily follow the user's performance of the target music piece.

[0054] However, when a period in which the acoustic characteristics of the audio signal X significantly fluctuate is expanded or contracted on the time axis, the reproduced sound may give an unnatural auditory impression. On the other hand, when a period in which the acoustic characteristics of the audio signal X are stably maintained is expanded or contracted on the time axis, the unnaturalness of the reproduced sound is unlikely to become apparent.

[0055] Considering the above tendency, as understood from the above-mentioned formula (5), the determination unit 331 of the first embodiment sets the variance σ(Wb[n],O) of the probability distribution of the observation likelihood L[t,n] when the time t does not correspond to the sounding point to a numerical value according to the fluctuation index Wb[n]. Specifically, the smaller the fluctuation index Wb[n], the larger the numerical value of the variance σ(Wb[n],O) is set. That is, compared to the case where the time t corresponds to the sounding point, the possibility of identifying a playback position R[t] that deviates from the performance position P[t] increases. As described above, the more stably the acoustic characteristics of the audio signal X are maintained, the smaller the numerical value of the fluctuation index Wb[n] is set. Therefore, the longer the period during which the acoustic characteristics of the audio signal X are maintained stably (i.e., the period during which the fluctuation index Wb[n] is small), the more likely the playback position R[t] will deviate from the performance position P[t]. According to the above configuration, a period in which the acoustic characteristics of the audio signal X are stably maintained is easily expanded or contracted on the time axis, whereas a period in which the acoustic characteristics fluctuate unstably is not easily expanded or contracted. Therefore, a reproduced sound with a natural auditory impression can be reproduced.

[0056] (2) Calculation of transition probability τ[n1,n2] (Sb) As described above, the transition probability τ[n1, n2] means the likelihood that the playback position R[t] will transition from the unit period U[n1] to the rear unit period U[n2] of the audio signal X. The determination unit 331 calculates the transition probability τ[n1, n2] for all combinations of selecting two unit periods U[n] (U[n1], U[n2]) from the N unit periods U[1] to U[N] of the audio signal X.

[0057] 7 and 8 illustrate a specific procedure of a process (hereinafter referred to as a "probability setting process") Sb in which the determination unit 331 calculates the transition probability τ[n1, n2]. When the probability setting process Sb is started, the determination unit 331 selects a combination of two unit periods U[n] (U[n1], U[n2]) from the N unit periods U[1] to U[N] of the acoustic signal X (Sb1).

[0058] The determination unit 331 determines whether or not the unit period U[n1] before the transition corresponds to the last unit period U[n] of the inter-utterance period V (Sb2). The inter-utterance period V is a period obtained by dividing the audio signal X on the time axis with each utterance point as a boundary. In Fig. 9, two inter-utterance periods V (V1, V2) that are successive on the time axis are illustrated, and it is assumed that the unit period U[n1] is located at the end of the inter-utterance period V1 (Sb2: YES).

[0059] If the unit period U[n1] before the transition is located at the end of the inter-tone period V1 (Sb2: YES), the determination unit 331 determines whether a predetermined condition is satisfied (Sb3). Specifically, the determination unit 331 determines whether a first condition (n1=n2) that the unit period U[n1] and the unit period U[n2] match, or a second condition that the unit period U[n2] after the transition is the unit period U[n1+1] immediately following the unit period U[n1] before the transition, is satisfied. The first condition means that the playback position R[t] stays in the last unit period U[n] of the inter-tone period V1. The second condition means that the playback position R[t] transitions from the last unit period U[n] of the inter-tone period V1 to the unit period U[n+1] in the inter-tone period V2 immediately following it.

[0060] When the first condition or the second condition is satisfied (Sb3: YES), the determination unit 331 sets the transition probability τ[n1, n2] according to the following rule (Sb4). Specifically, when the first condition is satisfied, the determination unit 331 sets the transition probability τ[n1, n2] (n1 = n2) to a predetermined value αH. On the other hand, when the second condition is satisfied, the determination unit 331 sets the transition probability τ[n1, n2] (n2 = n1 + 1) to a predetermined value αL. The predetermined value αH and the predetermined value αL are predetermined positive numbers. The predetermined value αH is set to a value sufficiently larger than the predetermined value αL (αH >> αL). For example, the predetermined value αH is set to a positive number equal to or smaller than "1" and sufficiently close to "1", and the predetermined value αL is set to a value obtained by subtracting the predetermined value αH from "1" (αL = 1 - αH).

[0061] As can be understood from the above description, the transition probability τ[n1,n2] (=αH) of the playback position R[t] staying at the last unit period U[n1] of the inter-utterance period V1 is sufficiently higher than the transition probability τ[n1,n2] (=αL) of the playback position R[t] transitioning from the last unit period U[n1] of the inter-utterance period V1 to the first unit period U[n2] of the inter-utterance period V1 immediately following the last unit period U[n1] of the inter-utterance period V1. According to the above configuration, the transition of the playback position R[t] across the sound generation points of the audio signal X is suppressed, so that the possibility of the audio component corresponding to one sound generation point being repeatedly reproduced multiple times is reduced. For example, the possibility that the singing voice, which is the reproduced sound of the audio signal X, is perceived by the listener as stuttering is reduced. In other words, a reproduced sound with a natural auditory impression can be reproduced. Note that, when the playback position R[t] stays continuously at one unit period U[n], the volume of the reproduced sound of the audio signal X may be reduced over time.

[0062] On the other hand, if the unit period U[n1] does not correspond to the last unit period U[n] of the inter-sound period V (Sb2: NO), or if the predetermined condition is not satisfied (Sb3: NO), the determination unit 331 determines whether or not the unit period U[n2] after the transition is within a predetermined range on the time axis with respect to the unit period U[n1] before the transition (Sb5), as illustrated in FIG. 8. Specifically, the determination unit 331 determines whether or not the unit period U[n2] is located within a range of a predetermined length Δn starting from the unit period U[n1]. If the number n2 of the unit period U[n2] after the transition is equal to or greater than the number n1 and equal to or less than (n1+Δn) (n1≦n2≦n1+Δn), the result of the determination is positive. If the number n2 of the unit period U[n2] exceeds the predetermined value (n1+Δn), this means that the playback position R[t] moves excessively far backward from the unit period U[n1].

[0063] If the unit period U[n2] is within the predetermined range (Sb5: YES), the determination unit 331 determines whether the audio signal X is silent in both the unit period U[n1] before the transition and the unit period U[n2] after the transition (Sb6). That is, it is determined whether both the sound presence indicator Wa[n1] and the sound presence indicator Wa[n2] are the numerical value "0" indicating the silence. If both the unit period U[n1] and the unit period U[n2] are silent (Sb6: YES), the determination unit 331 sets the transition probability τ[n1,n2] by the following formula (6) (Sb7).

number

[0064] In the formula (6), the symbol β means a predetermined positive number, and the symbol τ0 means a predetermined threshold value. As can be understood from the formula (6), when the absolute value |n1-n2| of the difference between the number n1 and the number n2 is below the threshold value τ0, the transition probability τ[n1,n2] is set to a predetermined value β. On the other hand, when the absolute value |n1-n2| is equal to or greater than the threshold value τ0, the transition probability τ[n1,n2] is set to "0". As can be understood from the above explanation, within a range in which the transition amount |n1-n2| on the time axis is below the threshold value τ0, the transition of the playback position R[t] is permitted with the transition probability τ[n1,n2] set to the predetermined value β. On the other hand, the transition of the playback position R[t] in which the transition amount |n1-n2| on the time axis exceeds the threshold value τ0 is prohibited (τ[n1,n2]=0).

[0065] On the other hand, if the acoustic signal X is in the presence of sound in one or both of the unit periods U[n1] and U[n2] (Sb6: NO), the determination unit 331 sets the transition probability τ[n1, n2] according to the following equation (7) (Sb8).

number

[0066] Equation (7) means that the transition probability τ[n1,n2] follows a normal distribution (Normal) with the difference (n1-n2) between the number n1 and the number n2 as a random variable. The difference (n1-n2) corresponds to the amount of movement of the playback position R[t] between the time (t-1) and the time t, that is, the movement speed of the playback position R[t].

[0067] The average of the probability distribution of the transition probability τ[n1, n2] is set to the standard speed P0 mentioned above. The standard speed P0 corresponds to the standard playback speed of the audio signal X, and is set to a predetermined positive number. Specifically, the standard speed P0 means the amount of change in the number n between time point (t-1) and time point t when the playback position R[t] of the audio signal X moves on the time axis at the standard speed. For example, the standard speed P0 is set to the ratio of the hop length Hn to the hop length Ht (P0=Hn / Ht).

[0068] The variance of the probability distribution of the transition probability τ[n1,n2] is set to a value P0 / Wb[n1] according to the fluctuation index Wb[n]. Specifically, the smaller the fluctuation index Wb[n1], the larger the variance P0 / Wb[n1] of the probability distribution is set to. That is, the smaller the fluctuation index Wb[n1], the more likely it is that the moving speed of the playback position R[t] will deviate from the standard speed P0. As described above, the more stably the acoustic characteristics of the audio signal X are maintained, the smaller the fluctuation index Wb[n] is set to. Therefore, for example, in a period in which the acoustic characteristics of the audio signal X are stably maintained (i.e., a period in which the fluctuation index Wb[n] is small), the variance P0 / Wb[n1] in the probability distribution of the transition probability τ[n1,n2] is set to a large value, and as a result, the moving speed of the playback position R[t] is allowed to deviate from the standard speed P0. On the other hand, in a period in which the acoustic characteristics of the audio signal X fluctuate unstably (i.e., a period in which the fluctuation index Wb[n] is large), the variance P0 / Wb[n1] in the probability distribution of the transition probability τ[n1.n2] is set to a small value, and as a result, the moving speed of the playback position R[t] is maintained at a speed close to the standard speed P0. In other words, a period in which the acoustic characteristics of the audio signal X are stably maintained is easily stretched on the time axis, and a period in which the acoustic characteristics fluctuate unstably is not easily stretched. Therefore, a reproduced sound with a natural auditory impression can be reproduced.

[0069] In addition, the transition probability τ[n1,n2] (=β) when the audio signal X is silent in both the unit period U[n1] and the unit period U[n2] (Wa[n1]=Wa[n2]=0) exceeds the transition probability τ[n1,n2] when the audio signal X is silent in one or both of the unit periods U[n1] and U[n2]. Under the above conditions, transitions of the playback position R[t] in the silent period of the audio signal X are more likely to occur than transitions of the playback position R[t] between silent periods and silent periods, or transitions of the playback position R[t] in a silent period. Therefore, compared to a form in which transitions of the playback position R[t] occur frequently in a silent period, a reproduced sound with a natural auditory impression can be reproduced.

[0070] If the unit period U[n2] is not within a predetermined range relative to the unit period U[n1] (Sb5: NO), the determination unit 331 sets the transition probability τ[n1,n2] to a predetermined value γ (Sb9). The predetermined value γ is set to a positive number that is sufficiently smaller than the predetermined value β in the formula (6). In other words, the transition of the playback position R[t] from the unit period U[n1] to a unit period U[n2] outside the predetermined range is permitted, although with a lower probability (predetermined value γ) compared to the transition of the playback position R[t] within the range.

[0071] After calculating the transition probability τ[n1, n2] for the current combination (U[n1], U[n2]) by the above processing (Sb4, Sb7, Sb8, Sb9), the identification unit 331 determines whether or not the transition probabilities τ[n1, n2] have been set for all combinations of two units selected from N unit periods U[1] to U[N] of the acoustic signal X (Sb10), as illustrated in FIG. 7. If there is an unset transition probability τ[n1, n2] (Sb10: NO), the identification unit 331 moves the processing to step Sb1. That is, the identification unit 331 newly selects two unit periods U[n] (U[n1], U[n2]) for which the transition probability τ[n1, n2] has not been set (Sb1), and sets the transition probability τ[n1, n2] for the combination (Sb2 to Sb9). On the other hand, if all of the transition probabilities τ[n1, n2] have been set (Sb10: YES), the specification unit 331 ends the probability setting process Sb.

[0072] B: Second embodiment In a form in which the sound of the sound signal X reproduced by the sound emitting device 23 and the sound emitted by the keyboard instrument 10 are dissociated in volume, there is a possibility that a sense of musical unity between the two cannot be generated. In consideration of the above circumstances, in the second embodiment, the volume of the reproduced sound of the sound signal X (hereinafter referred to as "reproduced volume") is linked to the strength of the operation of the keyboard instrument 10 by the user (hereinafter referred to as "operation strength"). Specifically, the reproduction unit 332 controls the reproduced volume of the sound signal X according to the strength of the operation by the user. The configuration and operation of each element other than the reproduction unit 332 are the same as in the first embodiment. Therefore, the second embodiment also achieves the same effects as the first embodiment.

[0073] 10 is a flowchart illustrating a specific procedure of the process (hereinafter referred to as the "playback process") Se executed by the playback unit 332 in the second embodiment. When the playback process Se is started, the playback unit 332 calculates the operation strength Λ[k] (Se1) using the following formulas (8a) and (8b). The operation strength Λ[k] is a numerical value (velocity) designated by the performance data D.

number

[0074] FIG. 11 is an explanatory diagram of the operation strength Λ[k]. The symbol k in the formula (8) is a number for identifying each operation (specifically, a key press) on the keyboard instrument 10. The symbol t[k] means the time when the operation k occurs. As illustrated in FIG. 11, it is assumed that an operation (k-1) with an operation strength λ[k-1] occurs at a time t[k-1], and an operation k with an operation strength λ[k] occurs at a time t[k] after the time t[k-1]. The operation k is, for example, a key press immediately after the operation (k-1). The time t[k-1] is an example of a "first time point", and the operation (k-1) is an example of a "first operation". The time t[k] is an example of a "second time point", and the operation k is an example of a "second operation".

[0075] As can be seen from Equation (8a), the playback unit 332 selects the larger (max) of the operation strength z[k] and the operation strength λ[k] as the operation strength Λ[k] at time t[k]. As can be seen from Equation (8b), the operation strength z[k] is the strength obtained by reducing the operation strength λ[k-1] of the operation (k-1) over time from time t[k-1] to time t[k]. The symbol λ in Equation (8b) is a predetermined positive number indicating the degree to which the operation strength λ[k-1] attenuates over time. The operation strength z[k] is an example of a "first strength", and the operation strength λ[k] is an example of a "second strength".

[0076] After calculating the operation strength Λ[k] by the above calculation, the playback unit 332 calculates an adjustment value G according to the operation strength Λ[k] (Se2). The adjustment value G is a coefficient (gain) by which the portion Y of the audio signal X to be played back is multiplied. Specifically, the playback unit 332 calculates the adjustment value G by the following formula (9).

number

[0077] In the second embodiment, the playback volume of the audio signal X is controlled according to the larger of the operation strength z[k] obtained by reducing the operation strength λ[k-1] of the operation (k-1) over time until the time t[k] and the operation strength λ[k] of the operation k at the time t[k] (i.e., the operation strength Λ[k]). Therefore, even if the operation strength λ[k] is sufficiently smaller than the operation strength λ[k-1], for example, if the operation strength Λ[k] obtained by reducing the operation strength λ[k-1] over time until the time t[k] is sufficiently large, the playback volume of the audio signal X is sufficiently maintained. Therefore, compared to a configuration in which the playback volume is controlled according to the operation strength λ[k] for each operation, the playback volume can be appropriately controlled for the user's performance.

[0078] C: Modification Specific modified embodiments added to each of the above-mentioned embodiments are exemplified below. Two or more embodiments selected from the following examples may be appropriately combined as long as they are not mutually contradictory.

[0079] (1) In each of the above-described embodiments, a keyboard instrument 10 is exemplified, but the type of instrument on which the user plays the target piece of music is not limited to the keyboard instrument 10. For example, any type of instrument, such as a string instrument, a wind instrument, or a percussion instrument, may be used by the user to play the target piece of music. For example, the acquisition unit 32 estimates the performance position P[t] by analyzing performance data D supplied from any instrument. In addition, the device that generates the performance data D may be a device in a form other than a musical instrument. For example, an information device such as a smartphone or a tablet terminal, or an operating device such as a keyboard, or any other type of device that accepts performance instructions from a user, may be used in place of the above-described keyboard instrument 10.

[0080] In the above-mentioned embodiments, the instruction data representing the performance instruction by the user is exemplified as the performance data D, but the type of the performance data D used for the analysis of the performance (estimation of the performance position P[t]) is not limited to the instruction data. For example, sound data representing the waveform of the sound generated by the performance by the user may be used as the performance data D for the analysis of the performance.

[0081] (2) In each of the above-described embodiments, a part of the processing period Q is set as the analysis period q to identify the playback position R[t], but the identification unit 331 may identify the playback position R[t] by setting the entire processing period Q as the analysis period q. In other words, time t2 and time t3 may coincide on the time axis, and the distinction between the processing period Q and the analysis period q is omitted.

[0082] (3) In each of the above-described embodiments, the variance σ(Wb[n],O) in the probability distribution of the observation likelihood L[t,n] is changed according to the fluctuation index Wb[n], but the variance of the probability distribution of the observation likelihood L[t,n] may be set to a predetermined value that is independent of the fluctuation index Wb[n]. Similarly, in each of the above-described embodiments, the variance P0 / Wb[n1] in the probability distribution of the transition probability τ[n1,n2] is changed according to the fluctuation index Wb[n], but the variance of the probability distribution of the transition probability τ[n1,n2] may be set to a predetermined value that is independent of the fluctuation index Wb[n].

[0083] (4) The moving speed of the playback position R[t] may be limited within a predetermined range. For example, if the amount of movement of the playback position R[t] between time (t-1) and time t exceeds a predetermined upper limit, the determination unit 331 sets the playback position R[t] to a value corresponding to the upper limit. On the other hand, if the amount of movement of the playback position R[t] between time (t-1) and time t falls below a predetermined lower limit, the determination unit 331 sets the playback position R[t] to a value corresponding to the lower limit. With the above configuration, it is possible to suppress an excessive deviation between the performance position P[t] and the playback position R[t].

[0084] (5) When the difference between the performance position P[t] and the playback position R[t] exceeds a predetermined threshold, the determination unit 331 may initialize the playback position R[t] to the performance position P[t] (R[t]=P[t]). With the above configuration, excessive deviation between the playback position R[t] and the performance position P[t] is suppressed. In addition, the playback position R[t] may be changed at a standard speed P0 within a predetermined period from the point in time when the playback position R[t] is initialized to the performance position P[t]. In other words, the playback position R[t] does not need to reflect the performance position P[t] within that period.

[0085] (6) In each of the above-described embodiments, the analysis unit 31 generates the index W[n] by analyzing the acoustic signal X stored in the storage device 22. However, in an embodiment in which the index W[n] related to the acoustic signal X is stored in advance in the storage device 22, the analysis unit 31 may be omitted. For example, in an embodiment in which the index W[n] related to the acoustic signal X is provided to the signal processing system 20 from an external device, the analysis unit 31 is omitted.

[0086] (7) As exemplified in each of the above-mentioned embodiments, various conditions (hereinafter referred to as "search conditions") are applied to the path search Sd2 in each of the above-mentioned embodiments. The search conditions are conditions that are set according to the characteristics of the audio signal X. The search conditions include the constraint conditions regarding the playback position R[t] as well as the numerical values ​​of the variables applied to the path search Sd2. As exemplified in the above-mentioned embodiments, the constraint conditions are, for example, a condition that the playback position R[t1] at time t1 of the analysis period q is fixed to the performance position P[t1] at the time t1, and the playback position R[t2] at time t2 of the analysis period q is fixed to the performance position P[t2] at the time t2. In addition, examples of search conditions regarding variables applied to the path search Sd2 include indices such as the observation likelihood L[t,n], the transition probability τ[n1,n2], and the fluctuation index Wb[t]. In other words, any variable applied to the path search Sd2 is included in the concept of the search conditions.

[0087] (8) In each of the above-mentioned embodiments, the acquisition unit 32 identifies the performance position P[t] of the target song by the user, but the information used to identify the playback position R[t] is not limited to the performance position P[t]. For example, a position that changes in the target song in response to an operation on an operation device such as a mouse or a touch panel may be substituted for the performance position P[t]. For example, a position that the user indicates and changes in the target song is replaced with the performance position P[t]. As can be understood from the above examples, the position used to identify the playback position R[t] is comprehensively expressed as a position that changes on the time axis in the target song in response to the user's actions (hereinafter referred to as the "indicated position"). The performance position P[t] in each of the above-mentioned embodiments and the position indicated by the user by operating the operation device are specific examples of the indicated position. Note that, as the operation device used by the user to indicate the indicated position, for example, a DJ controller in which a disk-shaped turntable rotates in response to the user's operation may be used. The acquisition unit 32 identifies the indicated position in response to the angle of rotation of the turntable.

[0088] (9) In the above-described embodiments, the audio signal X representing the performance sound of the target piece of music is expanded or contracted in response to the performance of the keyboard instrument 10 by the user, but the time-series signal to be expanded or contracted is not limited to the audio signal X. For example, a video signal representing a video related to the target piece of music may be expanded or contracted on the time axis in response to the performance of the user. The video signal represents, for example, a video or other video to be displayed in parallel with the performance of the target piece of music.

[0089] In the embodiment for processing a video signal, the estimation of the performance position P[t] by the acquisition unit 32 and the identification of the playback position R[t] by the identification unit 331 are the same as in the above-mentioned embodiments. The playback unit 332 displays a portion of the video signal corresponding to the playback position R[t] on a display device. The fluctuation index Wb[n] calculated by the analysis unit 31 through analysis of the video signal is, for example, a variable representing the degree of fluctuation of the video characteristics in the video signal. The video characteristics are, for example, the brightness of an image. The analysis unit 31 may also calculate an index (motion vector) representing a change in images that precede and follow each other on the time axis as the fluctuation index Wb[n].

[0090] As can be understood from the above description, the signal to be processed by the signal processing system 20 is comprehensively expressed as a time-series signal (e.g., audio signal X or video signal) representing audio or video related to the target music piece. The playback unit 332 is an element that causes a playback device to play a portion of the time-series signal corresponding to the playback position R[t]. The playback device includes a sound emitting device 23 that plays the audio represented by the audio signal X, or a display device that displays the video represented by the video signal.

[0091] (10) The signal processing system 20 may be implemented by a server device that communicates with an information device such as a smartphone or a tablet terminal. For example, performance data D generated by a keyboard instrument 10 connected to the information device is transmitted from the information device to the signal processing system 20. In the signal processing system 20, similar to the above-described various forms, estimation of the performance position P[t] by the acquisition unit 32 and specification of the reproduction position R[t] by the specification unit 331 are executed. The reproduction unit 332 transmits a portion Y corresponding to the reproduction position R[t] in the acoustic signal X to the information device. The information device includes a sound playback device 23 that plays back the portion Y received from the signal processing system 20. Even in the above configuration, the same effects as those of the above-described various forms are realized. The operation of the reproduction unit 332 transmitting the portion Y of the acoustic signal X to the information device is expressed as an operation of causing the information device to reproduce the portion.

[0092] (11) As described above, the functions of the signal processing system 20 according to the above-described various forms are realized by the cooperation of one or more processors constituting the control device 21 and a program stored in the storage device 22. The program according to the present disclosure can be provided in a form stored in a computer-readable recording medium and installed in a computer. The recording medium is, for example, a non-transitory recording medium, and an optical recording medium (optical disk) such as a CD-ROM is a preferred example, but any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium is also included. Note that the non-transitory recording medium includes any recording medium excluding a transitory, propagating signal, and a volatile recording medium is not excluded. Also, in a configuration where a distribution device distributes a program via a communication network, the recording medium storing the program in the distribution device corresponds to the above-described non-transitory recording medium.

[0093] D: Supplementary Note From the forms exemplified above, for example, the following configurations can be understood.

[0094] A signal processing system according to one aspect (aspect 1) of the present disclosure is a signal processing system that causes a playback device to play a time-series signal in response to the playback of a piece of music, and includes an acquisition unit that acquires a position designated by a user in the playback of the piece of music, and a control unit that performs time stretching of the time-series signal in response to the designated position. According to the above aspect, the time-series signal is time-stretched in response to the position designated by a user in the playback of the piece of music. Therefore, it is possible to cause the playback of the time-series signal to follow the user's instruction.

[0095] A "designated position" is a position designated by a user within a piece of music. Specifically, a position that changes within a piece of music in response to an action by the user is exemplified as a "designated position". A typical example of a "designated position" is, for example, a position on the time axis within a piece of music where the user plays (playing position). However, the action of the user that is reflected in the designated position is not limited to "playing". For example, a form in which the "designated position" changes in response to an operation on an operating device such as a mouse or a touch panel (another example of an "action") is also envisioned. Furthermore, a "designated position" includes not only a position currently designated by the user, but also a position that the user is predicted to designate in the future.

[0096] A "time-series signal" is a time-domain signal to be reproduced. Specifically, a "time-series signal" is a time-domain signal representing, for example, sound or video. Specifically, typical examples of a "time-series signal" are an audio signal representing the sound of a musical piece being played, or a video signal representing a video to be displayed in parallel with the performance of a musical piece. Therefore, a "playback device" is, for example, a sound emitting device that emits the sound represented by the audio signal, or a display device that displays the video represented by the video signal.

[0097] The performance sound represented by the "audio signal" includes not only musical sounds produced by musical instruments through performance, but also voices produced by singers (singing voices). The performance sound represented by the audio signal and the performance sound produced by a user's performance correspond to a common musical piece, but the specific relationship between the two is arbitrary. For example, it does not matter whether the performance part of the performance sound represented by the audio signal is different from the performance part played by the user. In other words, assuming that the user plays one or more of the multiple performance parts of a musical piece, the audio signal represents the performance sound of the one or more performance parts, or the performance sound of one or more performance parts other than the one or more performance parts.

[0098] In a specific example (Aspect 2) of Aspect 1, the time series signal is a signal representing sound or video, the acquisition unit acquires a plurality of designated positions over time, and the control unit executes the time warping by a route search that applies two or more distinct designated positions among the plurality of designated positions and search conditions according to characteristics of the time series signal. The "search conditions" are conditions that are set according to characteristics of the time series signal and are applied to the route search. The "search conditions" include constraint conditions related to the playback position (e.g., Aspect 7), as well as numerical values ​​of variables applied to the route search (e.g., Aspects 8, 10, and 11).

[0099] In a specific example (Aspect 3) of Aspect 1 or Aspect 2, the playback of the music piece is a performance of the music piece by the user. According to the above aspect, the playback of the time-series signal can be made to follow the performance of the music piece by the user.

[0100] "Performance" refers to the action of a user progressing music, and is a broad concept that includes not only the action of operating an instrument or other device to produce sound on that instrument (performance in the narrow sense), but also the action of a user singing a piece of music. The instruction position (performance position) is identified by analyzing the performance by the user. "Analysis of performance" is realized, for example, by analyzing performance data that represents the performance by the user. The performance data is instruction data (e.g., MIDI data) that represents the performance instructions by the user, or audio data (e.g., a sample series) that represents the waveform of the sound produced by the performance by the user.

[0101] In a specific example (Aspect 4) of Aspect 1, the control unit includes an identifying unit that identifies a playback position in the time-series signal corresponding to the designated position, and a playback unit that executes the time warping by having a playback device play a portion of the time-series signal corresponding to the playback position. According to the above aspect, time warping of the time-series signal following a change in the designated position is realized by having a playback device play a portion of the time-series signal corresponding to the playback position. The "playback position" is a position on the time axis in the time-series signal.

[0102] In a specific example (Aspect 5) of Aspect 4, the acquisition unit sequentially identifies the designated position for each of a plurality of time points on a time axis, and the identification unit performs a route search in each of a plurality of processing periods on the time axis, applying two or more designated positions identified for two or more time points within the processing period among the plurality of time points and a search condition according to a characteristic of the time-series signal, thereby identifying a time series of two or more playback positions corresponding to different time points within at least a portion of the processing period, and the playback unit causes the playback device to play back a portion of the time-series signal corresponding to each of the two or more playback positions. According to the above aspect, since the route search for identifying a time series of two or more playback positions is performed for each processing period on the time axis, it is possible to identify a playback position that accurately follows an instruction from a user, even if, for example, the speed of movement of the designated position fluctuates irregularly.

[0103] In a specific example (aspect 6) of aspect 5, the processing period is a period between a first time point among the multiple time points and a second time point subsequent to the first time point, and the at least a portion of the processing period is an analysis period from the first time point to a third time point between the first time point and the second time point. According to the above aspect, the time series of two or more playback positions within the analysis period from the first time point to the third time point is estimated according to the time series of the pointed position within the processing period from the first time point to the second time point. Therefore, it is possible to reduce the influence (noise) of the estimation error of the pointed position in a period near the end point of the processing period (for example, the period from the third time point to the second time point). In other words, it is possible to appropriately identify the playback position compared to a configuration in which the time series of the pointed position within the processing period is used to identify the time series of the playback position throughout the processing period.

[0104] In a specific example (Aspect 7) of Aspect 6, the search conditions include a condition for fixing the playback position at the first time point to the designated position at the first time point, and fixing the playback position at the second time point to the designated position at the second time point. According to the above aspect, the playback position at the first time point is fixed to the designated position at the first time point, and the playback position at the second time point is fixed to the designated position at the second time point. This reduces the possibility that the playback position will deviate excessively from the designated position within the analysis period.

[0105] In a specific example (Aspect 8) of Aspect 5, the search condition includes an observation likelihood at each of the plurality of time points, and the observation likelihood is a probability that each of a plurality of unit periods obtained by dividing the time series signal on a time axis corresponds to the playback position at the corresponding time point, and a probability distribution of the observation likelihood is specified by an average according to the designated position. In the above aspect, the average of the probability distribution of the observation likelihood applied to the route search is set according to the designated position. Therefore, the possibility that the playback position will deviate excessively from the designated position within the analysis period is reduced.

[0106] In a specific example (aspect 9) of aspect 8, the time-series signal is an audio signal representing the performance sound of the music piece, and a probability distribution of the observation likelihood at a time point among the multiple time points where the designated position corresponds to a sound generation point of the audio signal is determined by a first variance, and a probability distribution of the observation likelihood at a time point among the multiple time points where the designated position does not correspond to a sound generation point of the audio signal is determined by a second variance that is greater than the first variance. According to the above aspect, the variance (first variance) of the probability distribution used to identify the playback position for a time point corresponding to the sound generation point of the audio signal is lower than the variance (second variance) of the probability distribution used to identify the playback position for a time point not corresponding to the sound generation point. Therefore, at the time point corresponding to the sound generation point, the observation likelihood becomes a locally high value in the vicinity of a value corresponding to the designated position. In other words, at the time point corresponding to the sound generation point, the possibility that the playback position will approximate or match the designated position is higher than the possibility that the playback position will deviate from the designated position. This has the advantage that it is easier to make the playback of the audio signal follow the performance by the user.

[0107] In a specific example (aspect 10) of aspect 8 or aspect 9, the search condition includes a fluctuation index indicating a degree of fluctuation in characteristics of the time-series signal, and the variance of the probability distribution of the observation likelihood is set according to the fluctuation index. According to the above aspect, the variance of the probability distribution of the observation likelihood is set according to the fluctuation index of the time-series signal. For example, at a time when the characteristics of the time-series signal fluctuate unstably, the variance is set to a small value, and as a result, the playback position is close to the indicated position. On the other hand, at a time when the fluctuation in characteristics of the time-series signal is small, the variance is set to a large value, and as a result, it is permitted to identify a playback position that deviates from the indicated position. In other words, a playback sound that gives an auditory impression of naturalness can be reproduced.

[0108] A "variation index" is any index according to the degree of variation of a characteristic in a time series signal. The degree of variation of a characteristic is, for example, the frequency at which the characteristic varies or the amount of variation of the characteristic. Therefore, the variation index can also be said as an index of the stability or instability of a characteristic in a time series signal. A variation index for an audio signal represents the degree of variation of an audio characteristic such as a fundamental frequency or a frequency characteristic (e.g., an amplitude spectrum or MFCC). A variation index for a video signal represents the degree of variation of a video characteristic such as brightness.

[0109] In a mode in which the fluctuation index is set to a larger value as the degree of fluctuation in the characteristic increases (i.e., as the characteristic fluctuates more unstably on the time axis), the fluctuation index is expressed as an index representing how easily the characteristic fluctuates. On the other hand, in a mode in which the fluctuation index is set to a larger value as the degree of fluctuation in the characteristic decreases (i.e., as the characteristic is maintained more stably on the time axis), the fluctuation index is expressed as an index representing how difficult the characteristic is to fluctuate.

[0110] In a specific example (Aspect 11) of any one of Aspects 4 to 10, the search condition is set for each combination of two unit periods among a plurality of unit periods obtained by dividing the time-series signal on a time axis, and includes a transition probability indicating a likelihood that the playback position will transition between the two unit periods. According to the above aspect, the time series of the playback position can be appropriately identified by a path search that applies the transition probability for each combination of two unit periods in the time-series signal 2.

[0111] "Two unit periods" includes two different unit periods on the time axis as well as a common unit period on the time axis. When the two unit periods are different, the transition probability means the probability that the playback position moves on the time axis. On the other hand, when the two unit periods are common, the transition probability means the probability that the playback position stays at one unit period on the time axis.

[0112] In a specific example (Aspect 12) of Aspect 11, the time-series signal is an audio signal representing the performance sound of the music piece, and the transition probability (first transition probability) when the audio signal is silent in both of the two unit periods exceeds the transition probability (second transition probability) when the audio signal is silent in one or both of the two unit periods. According to the above aspect, transitions of the playback position within silent periods of the audio signal are more likely to occur than transitions of the playback position between silent periods or transitions of the playback position within a silent period. Therefore, compared to a form in which transitions of the playback position occur frequently within a silent period, it is possible to reproduce a playback sound that gives an auditory impression of being more natural.

[0113] In a specific example (aspect 13) of aspect 12, the probability distribution of the transition probability when the audio signal is in the presence of sound in one or both of the two unit periods is specified by an average set to a predetermined value and a variance according to a fluctuation index representing the degree of fluctuation of the audio characteristic in the audio signal. In the above aspect, the variance in the probability distribution of the transition probability is set according to the fluctuation index of the audio signal. For example, in a period in which the audio characteristic of the audio signal is stably maintained, the variance in the probability distribution of the transition probability is set to a large value, and as a result, the moving speed of the playback position is allowed to deviate from the predetermined value. On the other hand, in a period in which the audio characteristic of the audio signal fluctuates unstably, the variance in the probability distribution of the transition probability is set to a small value, and as a result, the moving speed of the playback position approaches the predetermined value. In other words, a period in which the audio characteristic of the audio signal is stably maintained is easily expanded or contracted on the time axis, and a period in which the audio characteristic fluctuates unstably is not easily expanded or contracted. Therefore, a reproduced sound with a natural auditory impression can be reproduced.

[0114] In a specific example (Aspect 14) of any of Aspects 11 to 13, the transition probability that the playback position stays at the last time point of a first inter-utterance period among a plurality of inter-utterance periods obtained by dividing the audio signal on the time axis by a plurality of utterance points exceeds the transition probability that the playback position transitions from the last time point to a time point within a second inter-utterance period immediately following the first inter-utterance period. In the above aspects, transitions of the playback position across utterance points are suppressed, so that the possibility of an audio component corresponding to one utterance point being repeatedly played back is reduced. In other words, a playback sound that gives an auditory impression of being natural can be generated.

[0115] In a specific example (Aspect 15) of any one of Aspects 4 to 14, the indicated position is a performance position estimated by the acquisition unit analyzing the performance of the music piece by the user. According to the above aspects, the performance position of the music piece by the user is identified as the indicated position. Therefore, it is possible to make the playback of the time-series signal by the playback device follow the performance of the music piece by the user.

[0116] In a specific example (aspect 16) of aspect 15, when a first operation occurs at a first time point in the performance and a second operation occurs at a second time point after the first time point, the reproduction unit selects the larger of a first intensity obtained by reducing the intensity of the first operation over time from the first time point to the second time point and a second intensity of the second operation (i.e., the maximum value) as the operation intensity at the second time point, and controls the volume of the reproduced sound of the time-series signal according to the operation intensity. In the above aspect, the volume of the reproduced sound of the audio signal is controlled according to the maximum value (control value) of multiple intensities including the first intensity obtained by reducing the intensity of the first operation over time to the second time point and the second intensity of the second operation at the second time point. Therefore, even if the second intensity is sufficiently smaller than the first intensity, for example, if the first intensity obtained by reducing the first intensity over time to the second time point is sufficiently large, the volume of the reproduced sound is sufficiently maintained. Therefore, compared to a configuration in which the volume of the reproduced sound is controlled according to the intensity of each operation, the volume of the reproduced sound can be appropriately controlled for the user's performance.

[0117] A signal processing method according to one aspect (aspect 17) of the present disclosure is a method for having a playback device play a time series signal in response to playback of a piece of music, which method obtains a position designated by a user during playback of the piece of music, and performs time stretching of the time series signal in accordance with the designated position.

[0118] In a specific example (Aspect 18) of Aspect 17, the time series signal is a signal representing sound or video, and in acquiring the designated position, a plurality of designated positions are acquired over time, and in time warping, the time warping is performed by a route search that applies search conditions according to two or more different designated positions among the plurality of designated positions and characteristics of the time series signal. The playback of a piece of music is, for example, a performance of the piece of music by a user.

[0119] A program according to one embodiment (embodiment 20) of the present disclosure is a program for causing a playback device to play a time series signal in response to playback of a piece of music, and causes a computer to function as an acquisition unit that acquires a position indicated by a user during playback of the piece of music, and a control unit that performs time stretching of the time series signal in accordance with the indicated position. [Explanation of symbols]

[0120] 100... performance system, 10... keyboard instrument, 20... signal processing system, 21... control device, 22... storage device, 23... sound emission device, 31... analysis unit, 32... acquisition unit, 33... control unit, 331... identification unit, 332... playback unit.

Claims

1. A signal processing system for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: an acquisition unit that acquires a position designated by a user within the piece of music; a control unit that identifies a playback position in the time-series signal corresponding to the designated position, and causes a playback device to play back a portion of the time-series signal corresponding to the playback position, thereby performing time expansion / contraction of the time-series signal; Equipped with the acquiring unit acquires a plurality of designated positions over time; The control unit executes a route search using two or more different designated positions among the plurality of designated positions and a search condition according to a characteristic of the time-series signal, thereby identifying the time series of the playback positions. Signal processing system.

2. A signal processing system for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: an acquisition unit that sequentially acquires a position designated by a user for each of a plurality of time points on a time axis within the music piece; a determination unit that determines a playback position in the time-series signal according to the designated position; a reproduction unit that performs time expansion / contraction of the time-series signal by causing a reproduction device to reproduce a portion of the time-series signal corresponding to the reproduction position; Equipped with the identification unit performs a path search by applying two or more designated positions identified for two or more time points within each of a plurality of processing periods on a time axis among the plurality of time points and a search condition according to a characteristic of the time-series signal, thereby identifying a time series of two or more playback positions corresponding to different time points within at least a portion of the processing period; The reproduction unit causes the reproduction device to reproduce portions of the time-series signal corresponding to each of the two or more reproduction positions. Signal processing system.

3. the processing period is a period between a first time point among the plurality of time points and a second time point that is located after the first time point, The at least a portion of the processing period is an analysis period from the first time point to a third time point between the first time point and the second time point.

3. The signal processing system of claim 2.

4. The search condition includes a condition that fixes the playback position at the first time point to the designated position at the first time point and fixes the playback position at the second time point to the designated position at the second time point.

4. The signal processing system of claim 3.

5. the search conditions include an observation likelihood at each of the plurality of time points; the observation likelihood is a likelihood that each of a plurality of unit periods obtained by dividing the time series signal on a time axis corresponds to the playback position at that time point; The probability distribution of the observation likelihood is determined by an average according to the pointing position.

3. The signal processing system of claim 2.

6. the time-series signal is an audio signal representing a performance sound of the musical piece, a probability distribution of the observation likelihood at a time point where the instruction position corresponds to an outgoing point of the acoustic signal among the plurality of time points is defined by a first variance; A probability distribution of the observation likelihood at a time point where the pointing position does not correspond to an emission point of the acoustic signal among the plurality of time points is defined by a second variance that is greater than the first variance.

6. The signal processing system of claim 5.

7. the search condition includes a fluctuation index representing a degree of fluctuation in a characteristic of the time-series signal; The variance of the probability distribution of the observation likelihood is set according to the fluctuation index.

7. The signal processing system according to claim 5 or 6.

8. The search condition is set for each combination of two unit periods among a plurality of unit periods obtained by dividing the time-series signal on a time axis, and includes a transition probability that indicates a likelihood that the playback position will transition between the two unit periods.

3. The signal processing system of claim 2.

9. the time-series signal is an audio signal representing a performance sound of the musical piece, A transition probability in a case where the audio signal is silent in both of the two unit periods is greater than a transition probability in a case where the audio signal is active in one or both of the two unit periods.

9. The signal processing system of claim 8.

10. The probability distribution of the transition probability when the acoustic signal is in the state of being in the state of being sound in one or both of the two unit periods is defined by an average set to a predetermined value and a variance according to a fluctuation index representing a degree of fluctuation of the acoustic characteristics in the acoustic signal.

10. The signal processing system of claim 9.

11. A transition probability that the playback position stays at a last time point of a first inter-tone period among a plurality of inter-tone periods obtained by dividing the time series signal on a time axis by a plurality of sounding points is greater than a transition probability that the playback position transitions from the last time point to a time point within a second inter-tone period immediately following the first inter-tone period.

11. The signal processing system according to claim 8.

12. The designated position is a performance position estimated by the acquisition unit analyzing the performance of the music piece by the user.

3. The signal processing system of claim 2.

13. The reproducing unit includes: when a first operation occurs at a first time point in the performance and a second operation occurs at a second time point after the first time point has elapsed, selecting a first intensity obtained by reducing an intensity of the first operation over time from the first time point to the second time point, or a second intensity of the second operation, whichever is larger, as an operation intensity at the second time point; The volume of the reproduced sound of the time-series signal is controlled according to the strength of the operation.

13. The signal processing system of claim 12.

14. A method for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: acquiring a position designated by a user within the piece of music; A playback position corresponding to the designated position is identified in the time-series signal, and a portion of the time-series signal corresponding to the playback position is played back by a playback device, thereby performing time expansion / contraction of the time-series signal.

1. A computer-implemented signal processing method, comprising: In acquiring the designated position, a plurality of designated positions are acquired over time; In the time warping, a route search is performed by applying search conditions according to two or more different designated positions among the plurality of designated positions and characteristics of the time-series signal, thereby identifying the time series of the playback position. Signal processing methods.

15. A method for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: Sequentially acquiring positions designated by a user for each of a plurality of time points on a time axis within the piece of music; identifying a playback position in the time-series signal corresponding to the designated position; The portion of the time-series signal corresponding to the playback position is played back by a playback device, thereby performing time expansion / contraction of the time-series signal.

1. A computer-implemented signal processing method, comprising: In identifying the playback positions, a route search is performed in each of a plurality of processing periods on a time axis, using two or more designated positions identified for two or more time points within the processing periods among the plurality of time points and search conditions according to characteristics of the time series signal, thereby identifying a time series of two or more playback positions corresponding to different time points within at least a portion of the processing periods; In the time warping, the reproduction device reproduces portions of the time-series signal corresponding to each of the two or more reproduction positions. Signal processing methods.

16. A program for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: an acquisition unit that acquires a position designated by a user within the piece of music; and a control unit that identifies a playback position in the time-series signal corresponding to the designated position, and causes a playback device to play back a portion of the time-series signal corresponding to the playback position, thereby performing time expansion / contraction of the time-series signal; A program for causing a computer to function as the acquiring unit acquires a plurality of designated positions over time; The control unit executes a route search using two or more different designated positions among the plurality of designated positions and a search condition according to a characteristic of the time-series signal, thereby identifying the time series of the playback positions. program.

17. A program for causing a playback device to play back a time-series signal in response to playback of a piece of music, comprising: an acquisition unit that sequentially acquires a designated position designated by a user for each of a plurality of time points on a time axis within the music piece; a determination unit that determines a playback position in the time-series signal according to the designated position; and a reproduction unit that performs time expansion / contraction of the time-series signal by causing a reproduction device to reproduce a portion of the time-series signal corresponding to the reproduction position; A program for causing a computer to function as the identification unit performs a path search by applying two or more designated positions identified for two or more time points within each of a plurality of processing periods on a time axis among the plurality of time points and a search condition according to a characteristic of the time-series signal, thereby identifying a time series of two or more playback positions corresponding to different time points within at least a portion of the processing period; The reproduction unit causes the reproduction device to reproduce portions of the time-series signal corresponding to each of the two or more reproduction positions. program.

Citation Information

Patent Citations

  • Musical performance clock generating device, data reproducing device, musical performance clock generating method, data reproducing method, and program

    JP2009014923A

  • Score alignment device and score alignment program

    JP2015079183A

  • Reproduction control method and reproduction control device

    JP2019056871A

  • Real-Time Music to Music-Video Synchronization Method and System

    US20110230987A1

  • Musical performance analysis method, automatic music performance method, and automatic musical performance system

    WO2018016582A1